Reactive maintenance is the most expensive way to run a plant — yet it remains the default culture in many high-volume operations. Drawing on fifteen years leading engineering teams across brewing, pharmaceutical manufacturing and large-scale logistics, Lazarus Maha sets out a practical route from breakdown-driven firefighting to reliability-led asset performance, and the results that follow when the transition takes hold.
By Lazarus G. Maha CMRP
Regional Engineering & Reliability Leader | Multi-Site Maintenance & Asset Performance Project Manager
Go into any high-volume, stressed plant and you’ll encounter the same scenario: highly competent and experienced technicians constantly playing catch-up, maintenance backlog outpacing the work completed, expensive and rushed parts orders, and leaders of production who have subconsciously lost faith in the engineering department. This can be seen as a foresight as people in reactive cultures always work harder than anyone else, where the plant is not short of effort.
However, the price of that deficit is not always apparent in one line of the profit-and-loss account, which is why it stays that way. The evidence base behind the issues of unplanned maintenance work is compelling: research shows that unplanned tasks can cost organisations three to five times more per intervention than planned tasks (Weidner, 2023), with industry data revealing that on average, unplanned downtime costs manufacturers $125,000 per hour across all sectors (Liu et al., 2012). The deficit can be seen in premium freight (expedited shipping to make up lost time), overtime (to make up for missed schedules), reduced asset life (by running equipment to failure) and the gradual erosion of engineers who are “burning out” battling the same fires. In my previous years working in food and beverage manufacturing, pharmaceutical production and fulfilment logistics, the operations that change this cycle are the ones that actively and step-by-step move to performance-based maintenance (Introna and Santolamazza, 2024).
What ‘reliability-led’ actually means
Reliability-led maintenance is often boiled down to acronyms (TPM, RCM, FMEA, RCFA), but the concept is very easy to explain: Maintenance effort should be allocated based on the consequences of failure, not on which asset complained the loudest last week. Total Productive Maintenance (TPM) makes the operating teams part and parcel of owning the basic asset care (Nakajima, 1988); Reliability-Centred Maintenance (RCM) asks what each asset actually needs to continue to do its job (Nowlan and Heap, 1978); Failure Mode and Effects Analysis (FMEA) ranks where the real risk lies (Huang et al., 2020); and Root Cause Failure Analysis (RCFA) makes sure that when an asset fails it fails for the last time (Latino and Latino, 2006).
The sequencing matters. But when an organisation moves quickly to advanced condition monitoring, and their asset register is incomplete, their criticality ranking is outdated, and the discipline of planned-work is not strong, then they simply create more data about a plant that they still cannot control (Introna and Santolamazza, 2024). The basis is always the same: be familiar with your assets, rank them by criticality (Lopes et al., 2020), stabilise the preventive maintenance schedules on the critical assets and implement structured root cause analysis to prevent repeat failures from eating up the freed-up capacity.
Case one: rebuilding reliability in a large-scale brewing operation
In a production brewery with more than 1.3 million hectolitres per year, I was responsible for the reliability strategy for the steam generation and ammonia refrigeration systems, compressed air and carbon dioxide recovery systems, and the utilities backbone that powers every bottle that a brewery packages. The starting point was that it was acceptable in paper but had been done through heroics and not planning, and that the maintenance spend was climbing year on year and was controllable.
This involved vibration analysis of rotating equipment (Vishwakarma et al., 2017), disciplined RCFA on all major loss events (Latino and Latino, 2006), and RCM reviews on the critical utility systems (Nowlan and Heap, 1978). The elimination of defects was considered a production objective rather than an engineering pastime; each failure mode that is eliminated is recorded, valued, and reported in conjunction with the production figures. During the programme, a 20% reduction in mean time between failures on production and utilities assets, a 93% increase in equipment availability and a 60 per cent fall in controllable spend on the targeted asset groups were documented with the saving of approximately £49,000 in a single operating period as a result of lifecycle cost optimisation not deferred care (Hartman and Tan, 2014).
Equally significant was the fact that none of the specific ATEX-rated grain handling and storage equipment involved underwent statutory or safety-critical maintenance during the process, which was carried out at one hundred per cent adherence to service-level. As soon as it is apparent that saving money is being done at the cost of compliance, the reliability programmes lose their right to operate and keeping them on track is a leadership issue, not a technical one.
Case two: an eighty per cent reduction in reactive callouts
The second example was in a live fulfilment operation with over two thousand associates, during which I was responsible for the reliability and maintenance engineering teams and integrated service providers for the site. Fulfilment logistics is a strict reliability environment: materials-handling equipment, HVAC and building management systems run close to continuously and every unexpected disruption has a direct impact on the customer promise.
The strategy was to take the reactive workloads head on. The most dangerous equipment was identified via asset lifecycle assessments (Hartman and Tan, 2014), the worst offenders were targeted for retrofit and early equipment management (EEM) interventions, preventive maintenance schedules were reworked around real duty and not manufacturer schedules and contracts included measurable key performance indicators (KPIs) for contractors that were monitored with the same level of rigour as internal teams.
Eighty per cent of the reactive engineering callouts were reduced in a year. The documentation of savings from reliability-led asset performance management and lifecycle planning and maintenance cost optimisation was valued at £144,610, and one hundred per cent statutory compliance demonstrated by internal audit scores of over ninety per cent for the last twelve months – not just an accident of the numbers.
Making it stick: systems, standards and people
All the reliability transformations I have led have ended up on three no-nonsense disciplines. The first one is data hygiene – a computerised maintenance management system is only as good as its master data (Labib, 1998); reinforcing asset hierarchies, bills of materials and failure coding is what turns work orders into actionable information (Hamodi and Aljumaili, 2017). The second is standardisation, such as 5S in the workshops (Randhawa and Ahuja, 2017), the standard operating procedures for repeated tasks, and governance processes for capital replacement decisions (Hartman and Tan, 2014) to ensure that good practice lasts beyond the life of the original perpetrator.
The third is and is most often underestimated, people. Just as much as any analytical tool, succession planning, building skills matrices, and purposeful coaching of first-line engineering leaders are important to building reliability over the long term—not over a few months (Dereje et al., 2025). To ensure that the value of defect elimination is recognized over reactive improvisation, technicians who have been rewarded for years for dramatic breakdown recoveries must be convinced and put into practice that they need to know that the defect that was eliminated is more valuable than the breakdown that was made impossible by defect elimination (Dereje et al., 2025).
The problem is not just a matter of individual behaviour – it is usually structural. Many organisations have created systems which are essentially engineered and ingrained to reward unreliable maintenance. This includes overtime policies that encourage emergency calls out and contractor billing policies which reward more ‘reactive’ work than ‘planned’ work, and KPIs that reward fast breakdowns but fail to recognise defect elimination. These do not last by malicious intent, they are built up over the years where the only viable answer was to respond reactively. However, as soon as a reliability programme starts to work they are its biggest enemies. Just as with any technical change — a change of the contractor performance framework that incentivises the availability of services instead of the volume of interventions, and a change in the internal recognition system so that the failure that was averted — the failure that did not occur — is what constitutes engineering excellence and the measure of that success. Callout and overtime schemes need to be changed as well, to ensure that the people who can most stop the breakdown culture don’t make it financially attractive to them.
Well, nothing of this is exotic, and that’s the idea. What is observed is that the plant which reaches a step change in reliability is not always the most technologically advanced plant, but is the plant that has the discipline to follow well-known methods thoroughly, in the proper sequence and for a sufficiently long period of time to change the culture (Nakajima, 1988; Nowlan and Heap, 1978). It’s hard to be more effective at preventing fires than at putting them out. Engineering leadership is about making that discipline more valuable, and then proving it, in availability, in cost and the unobtrusive way in which an operation functions without relying on adrenaline.
About the Author
Lazarus G. Maha CMRP is a regional engineering and reliability leader with 15 years’ experience across FMCG, pharmaceutical and logistics operations, currently working as an EU/UK Engineering Project Manager overseeing a multi-site capital engineering portfolio. He is a Certified Maintenance & Reliability Professional (CMRP), holds IPMA Level D in project management, an MBA and a BSc in Mechanical Engineering, and has led reliability, maintenance and utilities teams in the United Kingdom and internationally.
References
Anderson, R.T. and Neri, L. (eds.) (2012) Reliability-centered maintenance: management and engineering methods. Springer Science & Business Media.
Dereje, M., Tilahun, S., Mekuria, G. and Negesse, Y. (2025) ‘Exploring the role of human factors and organizational culture in maintenance management practice within large processing industries’, Public Organization Review, pp. 1–32.
Hamodi, H. and Aljumaili, M. (2017) ‘Data quality of maintenance data: a case study in MAXIMO CMMS’, in Maintenance Performance Measurement and Management 2016 (MPMM 2016), 28 November, Luleå, Sweden. Luleå: Luleå tekniska universitet, pp. 105–110.
Hartman, J.C. and Tan, C.H. (2014) ‘Equipment replacement analysis: a literature review and directions for future research’, The Engineering Economist, 59(2), pp. 136–153.
Huang, J., You, J.X., Liu, H.C. and Song, M.S. (2020) ‘Failure mode and effect analysis improvement: a systematic literature review and future research agenda’, Reliability Engineering & System Safety, 199, p. 106885.
Introna, V. and Santolamazza, A. (2024) ‘Strategic maintenance planning in the digital era: a hybrid approach merging Reliability-Centered Maintenance with digitalization opportunities’, Operations Management Research, pp. 1–24.
Kasim, N.I., Musa, M.A., Razali, A.R., Mohamad Noor, N. and Wan Saidin, W.A.N. (2015) ‘Improvement of overall equipment effectiveness (OEE) through implementation of total productive maintenance (TPM) in manufacturing industries’, Applied Mechanics and Materials, 761, pp. 180–185.
Labib, A.W. (1998) ‘World-class maintenance using a computerised maintenance management system’, Journal of Quality in Maintenance Engineering, 4(1), pp. 66–75.
Latino, R.J., Latino, M.A., Latino, K. and Latino, K.C. (2002) Root cause analysis: improving performance for bottom-line results. Boca Raton: CRC Press.
Liu, J., Chang, Q., Xiao, G. and Biller, S. (2012) ‘The costs of downtime incidents in serial multistage manufacturing systems’.
Lopes, I., Figueiredo, M. and Sá, V. (2020) ‘Criticality evaluation to support maintenance management of manufacturing systems’, International Journal of Industrial Engineering and Management, 11(1), p. 3.
Rajput, H. (2012) A total productive maintenance (TPM) approach to improve overall equipment efficiency.
Randhawa, J.S. and Ahuja, I.S. (2017) ‘5S–a quality improvement tool for sustainable performance: literature review and directions’, International Journal of Quality & Reliability Management, 34(3), pp. 334–361.
Vishwakarma, M., Purohit, R., Harshlata, V. and Rajput, P. (2017) ‘Vibration analysis and condition monitoring for rotating machines: a review’, Materials Today: Proceedings, 4(2), pp. 2659–2664.
Weidner, T.J. (2023) ‘Planned maintenance vs unplanned maintenance and facility costs’, IOP Conference Series: Earth and Environmental Science, 1176(1), p. 012037



