Predictive Maintenance Works. Most Programmes Still Fail.
The 30 percent and 45 percent savings figures in every predictive maintenance deck trace to one uncited line in a 2010 US government guide. Here is what the primary sources actually say.
Predictive maintenance is not a hard technical problem any more. Vibration analysis on rotating equipment was mature in the 1990s. The statistical methods sit in libraries any engineer can install in an afternoon. A wireless accelerometer costing a few hundred dollars does what a wired channel costing thousands did fifteen years ago. And most predictive maintenance programmes still stall, get quietly defunded or end up as a dashboard nobody opens.
The reasons are almost never algorithmic. Machines looked after properly produce very few examples of failure, which starves the supervised learning everyone assumes is happening. Alerts that turn out to be nothing burn trust faster than correct alerts earn it. An alert that never becomes a work order with a part attached is a notification, not maintenance.
What changed in 2026 is the target. The industry stopped selling prediction and started selling prescription. Not "this pump will fail in eleven days" but "replace the bearing on pump P-114 during Thursday night changeover, part reserved, technician assigned." A better product. Also a much harder one to install, for reasons that have nothing to do with the model.
The maintenance ladder, rung by rung
Reactive maintenance is running the machine until it stops. The US Department of Energy's O&M Best Practices Guide cites a winter 2000 study putting the average American facility at more than 55 percent reactive, 31 percent preventive and 12 percent predictive. Its profile for top-performing plants is inverted: under 10 percent reactive, 25 to 35 percent preventive and 45 to 55 percent predictive.
Preventive maintenance services on a schedule. Calendar time, run hours, cycles. It is the default in most plants and it rests on an assumption disproved in 1978. That December, F. Stanley Nowlan and Howard F. Heap published their analysis of United Airlines fleet reliability data. Sorting components by how failure probability moved with age, they found six patterns. Only 6 percent of items showed pronounced wear-out. Another 5 percent grew steadily more likely to fail with no defined wear-out age. Their own exhibit is blunt: 11 percent might benefit from a limit on operating age, 89 percent cannot. Scheduled overhaul is the wrong tool for roughly nine failure modes in ten, and that has been documented for nearly fifty years.
Condition-based maintenance measures something and acts when it crosses a limit, under the umbrella of ISO 17359. Unglamorous, effective and the right place for many plants to stop.
Predictive maintenance adds the time axis. Instead of "this reading is high" it attempts "this reading hits the limit in roughly N days." That estimate is remaining useful life, and it is where the data problem bites.
Prescriptive maintenance adds the decision: which action, on which date, with which part, weighed against the production schedule. Here is the detail that should embarrass the industry. ISO 13374-1 specified this in June 2003. Its six functional blocks run data acquisition, data manipulation, state detection, health assessment, prognostic assessment and advisory generation. Advisory generation is prescription. It is the sixth block, and it is the one nearly everyone skipped for twenty years while shipping the first five.

Prediction is a model problem. Prescription is a distributed transaction problem.
What each rung requires you to sense
Four modalities carry almost all industrial condition monitoring. Vibration is the workhorse for rotating equipment; bearing defect frequencies live in the kilohertz range, so a useful channel needs bandwidth out to roughly 10 kHz, which rules out cheap accelerometers sampled at 100 Hz. Thermal catches loose connections, misalignment and anything dissipating energy it should not. Acoustic emission, covered by ISO 22096, picks up structure-borne stress waves at ultrasonic frequencies and gives the earliest available warning of lubrication failure and crack initiation, though it is very hard to read. The fourth is underused: motor current signature analysis infers broken rotor bars, air-gap eccentricity and mechanical looseness from the current waveform at the motor control centre, without putting a sensor on the machine at all. No cabling, no intrinsic safety paperwork, no confined space entry.
Cost per monitored point decides coverage. Wiring an accelerometer back to a central system is commonly quoted at $1,000 to $5,000. Industrial wireless accelerometers now land around $200 to $800 installed. I live in Shenzhen, and I have spent many afternoons in Huaqiangbei with a shopping list. A three-axis MEMS accelerometer module goes over the counter there for the price of lunch, and the wideband industrial parts that genuinely reach 6 kHz sit in trays two floors up. The sensor stopped being the constraint years ago.
What plants collect is thinner than the marketing suggests. Siemens' downtime survey found 87 percent of large manufacturers gathering some machine-health data but only half collecting at least one of current, vibration or temperature, with close to three-quarters still on factory historians. China has attacked this from the standards side. GB/T 43555-2023, the algorithm evaluation method for predictive maintenance, took effect on 1 July 2024. GB/T 45507-2025, the performance evaluation method, took effect on 1 October 2025. First standardise how you judge the model, then whether the programme delivered anything. The second question kills projects.
Where the 30 percent and 45 percent came from
Two figures appear in nearly every predictive maintenance pitch: a 30 percent cut in maintenance costs and a 45 percent cut in downtime. I went looking for the study. There isn't one.
The trail ends at the US Department of Energy's Federal Energy Management Program, in the Operations and Maintenance Best Practices Guide, Release 3.0, August 2010. Chapter 5, section 5.4. The wording is "independent surveys indicate the following industrial average savings resultant from initiation of a functional predictive maintenance program", followed by five bullets: return on investment 10 times, maintenance costs down 25 to 30 percent, breakdowns eliminated 70 to 75 percent, downtime down 35 to 45 percent, production up 20 to 25 percent.
The chapter does not say which independent surveys. Its reference list holds exactly two entries: a NASA facilities guide from 2000 and a 2001 trade article about pumps published on Pump-Zone.com. Neither is cited against the savings bullets. No sample size, no year, no sector, no method and no name.
Two things follow. The numbers everyone repeats are the top of a range restated as a point estimate: 25 to 30 becomes 30, 35 to 45 becomes 45. And the same document, speaking in its own voice, makes a far more modest claim. A properly functioning predictive maintenance programme saves 8 to 12 percent over preventive maintenance alone. That is the honest headline. Nobody prints it.
The other canonical sources are softer than their reputations. McKinsey's much quoted 30 to 50 percent downtime reduction comes from an August 2017 article with no disclosed method. Deloitte's more conservative set, 20 to 50 percent less planning time and 5 to 10 percent lower maintenance costs, is footnoted to internal Deloitte analysis derived from work with clients. Siemens' True Cost of Downtime 2024, source of the $1.4 trillion Fortune Global 500 figure, rests on 181 online interviews covering April 2019 to March 2023, published by a company that sells predictive maintenance software.
None of this means the technology fails. It means nobody has published a defensible number for how well it works, and a 2010 restatement of an uncited survey has been laundered through fifteen years of slide decks into something people treat as measured.
Healthy machines do not generate training data
You want a model that predicts gearbox failure, so you need examples of gearbox failure. A well-maintained gearbox might fail once every seven years. Assembling 200 labelled examples of one failure mode needs roughly 1,400 machine-years, recorded in a form that survives contact with a model.
They were not. Maintenance history lives in CMMS work orders as free text typed at the end of a shift. "Replaced bearing, machine running OK" is not a label. No fault mode, no severity, no onset time and usually no accurate timestamp for when the symptom started.
The academic literature shows how bad this is. The most widely used public benchmark in the field is the AI4I 2020 Predictive Maintenance Dataset, donated to the UCI repository by S. Matzka in 2020: 10,000 rows, 339 labelled as machine failure, 3.39 percent. It is synthetic. The default dataset for a field about physical machines is generated, because real run-to-failure data at scale is not obtainable.
So what gets deployed is not failure classification. It is anomaly detection. Autoencoder reconstruction error, isolation forests, one-class methods, Mahalanobis distance over spectral features. The model learns what normal looks like for one asset and flags departures. Genuinely useful and strictly weaker than what was sold, because it cannot name the fault mode and cannot give a remaining useful life with any honesty. "This gearbox does not look like it did for the past ninety days" is a real signal. It is not a prediction.
Year one is therefore a labelling exercise, not a modelling one. Label Studio for the event set, Edge Impulse or ONNX Runtime for pushing feature extraction onto a gateway: that work matters more than the choice of architecture.
The alert that never becomes a work order
Two failure modes kill deployed systems and they are not symmetric. A missed failure costs a lot, once, visibly. A false positive costs a little, quietly, every single time, and it is the false positives that end programmes. Process control learned this decades ago. EEMUA 191, fourth edition November 2024, targets roughly one alarm per operator per ten minutes in steady state, because past that humans stop reading them. A nuisance alarm in a control room is an acknowledged beep. A nuisance alert in maintenance is a technician up a ladder with a torque wrench and a line waiting on him.
Trust is the actual product. Once a crew decides the system cries wolf, alerts stop converting and the programme is dead regardless of model accuracy. I have watched a well-tuned system get switched off after four wrong calls in a fortnight, on a line where it had also caught a coupling failure that would have cost a full shift.
For an alert to matter it must become a work order in the CMMS, with a part reserved in the ERP, a window agreed with production planning and a named technician. That chain crosses four systems never designed to speak to each other: the historian, the CMMS, the ERP or MRP and the MES. Prediction is a model problem. Prescription is a distributed transaction problem with compensating actions, which is why durable workflow engines like Temporal now show up in stacks that used to be pure OT and why telemetry lands in analytic stores like Apache Doris.
The best illustration sits inside the McKinsey article everyone quotes for the opposite reason. An offshore operator built a model that flagged compressor failures weeks ahead. It could not prevent them. It cut downtime from 14 days to 6 by pre-staging people and parts. The model did not create the value. The logistics did.
The strongest case against everything above
Stated as strongly as I can: you have described a marketing problem and called it a technology problem. Badly sourced savings figures do not mean savings are absent, only that nobody audited them. The failure-data objection is a 2018 objection, since transfer learning, fleet pooling, physics-informed models and simulation-generated data have all moved on. And the organisational objection is self-liquidating. If the system files the work order, reserves the part and books the window itself, the trust problem disappears, because no human is left in the loop to lose faith.
Most of that is correct. Fleet pooling genuinely works, and it works exactly where you would expect: wind turbines, elevators, rooftop HVAC units, anywhere one operator owns thousands of near-identical assets. Simulation helps for kinematics and structural modes, and environments like NVIDIA Isaac Sim and MuJoCo now generate failure trajectories that would be ruinous to produce on real hardware.
The last step is where it breaks. Automating the work order does not remove the trust problem, it raises the stakes and moves it upstream. A false positive wasting a technician's hour is annoying. One that reserves a $40,000 spare, books a four-hour production window and fires a purchase order is expensive before any human has looked at it. Gartner's June 2025 forecast that more than 40 percent of agentic AI projects will be cancelled by the end of 2027, for escalating costs, unclear business value or inadequate risk controls, was not about maintenance. It names the failure mode prescription creates, precisely.
Fleet pooling also only helps if you own a fleet. Most manufacturers own one of each machine, bought across three decades from vendors who no longer exist.
Detection improved. Repair got slower.
The most interesting number in the Siemens survey is one nobody quotes. Between 2019 and 2023 the average plant cut unplanned downtime incidents from 42 a month to 25 and hours lost from 39 to 27. Over the same period mean time to repair rose from 49 minutes to 81.
Siemens attributes this partly to skilled labour lost after 2020 and partly to something sharper: the easy failures got predicted away, leaving a population that is hard to detect and hard to fix. Prediction did not merely reduce the failure count, it selected for difficulty on both axes at once. Prescription is now sold into that residual population as a scheduling optimisation. Scheduling is not the constraint when the person who knew how to rebuild the machine took redundancy in 2022.
There is a further assumption in the ladder, visible only from outside the countries that wrote it. It assumes the machine is the least reliable thing in the building. In Nigeria that is not true. The Manufacturers Association of Nigeria put manufacturers' spending on alternative energy at 1.34 trillion naira in 2025, up 71.4 percent from 781.68 billion naira in 2023, after tariffs had already jumped from around 68 naira per kilowatt-hour to between 209 and 225. On a plant like that the dominant failure mode is the grid and the critical asset is a diesel generator everyone already knows is the bottleneck. A remaining-useful-life estimate for an extruder bearing is a precise answer to a question nobody in the building asked.
Prescriptive maintenance is the right direction. It is also sold as a way around organisational discipline that it actually demands more of. You cannot automate a work order into a CMMS whose asset register is wrong.
Tools referenced
Edge Impulse, reviewed here: Edge Impulse review.
ONNX Runtime, reviewed here: ONNX Runtime review.
Label Studio, reviewed here: Label Studio review.
Temporal, reviewed here: Temporal review.
Apache Doris, reviewed here: Apache Doris review.
NVIDIA Isaac Sim, reviewed here: NVIDIA Isaac Sim review.
Sources
US DOE Federal Energy Management Program, O&M Best Practices Guide Release 3.0, Chapter 5: Types of Maintenance Programs (August 2010): https://www1.eere.energy.gov/femp/pdfs/om_5.pdf
Nowlan, F.S. and Heap, H.F., Reliability-Centered Maintenance, United Airlines / US Department of Defense, December 1978: https://reliabilitywebfiles.s3.amazonaws.com/Reliability+Centered+Maintenance+by+Nowlan+and+Heap.pdf
Siemens / Senseye, The True Cost of Downtime 2024 (methodology: 181 online interviews, April 2019 to March 2023): https://assets.new.siemens.com/siemens/assets/api/uuid:1b43afb5-2d07-47f7-9eb7-893fe7d0bc59/TCOD-2024_original.pdf
UCI Machine Learning Repository, AI4I 2020 Predictive Maintenance Dataset (S. Matzka, 2020): https://archive.ics.uci.edu/dataset/601/ai4i+2020+predictive+maintenance+dataset
Deloitte Insights, Using predictive technologies for asset maintenance (9 May 2017): https://www.deloitte.com/us/en/insights/industry/manufacturing-industrial-products/industry-4-0/using-predictive-technologies-for-asset-maintenance.html
McKinsey & Company, Manufacturing: Analytics unleashes productivity and profitability (14 August 2017): https://www.mckinsey.com/capabilities/operations/our-insights/manufacturing-analytics-unleashes-productivity-and-profitability
Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (25 June 2025): https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Vanguard, Manufacturers' alternative energy expenditure surges 71.4% to N1.34trn (Manufacturers Association of Nigeria data, June 2026): https://www.vanguardngr.com/2026/06/economic-reforms-manufacturers-alternative-energy-expenditure-surges-71-4-to-n1-34trn-in-two-years/
Frequently Asked Questions
How does predictive maintenance work?
Predictive maintenance works by continuously measuring a physical signal from a machine, most often vibration, temperature, acoustic emission or motor current, extracting features from that signal and comparing them against a learned model of normal behaviour for that specific asset. When the deviation crosses a threshold the system estimates how long the machine can keep running before failure, a figure called remaining useful life, and raises an alert. In practice most deployed systems use anomaly detection rather than failure classification, because well-maintained machines produce too few failure examples to train a supervised model. That means they can tell you something has changed, but usually not which fault mode it is or how long you have.
What is the difference between predictive and preventive maintenance?
Preventive maintenance services equipment on a fixed schedule based on calendar time, run hours or cycles, regardless of the machine's actual condition. Predictive maintenance measures condition and acts only when the data says intervention is needed. The distinction matters because the 1978 Nowlan and Heap reliability study of United Airlines fleet data found that only 11 percent of component types showed a wear-out pattern where an age limit helps, while 89 percent did not. For most failure modes, servicing on a schedule performs unnecessary work without lowering the failure rate, and can introduce faults through the intervention itself.
Is the 30 percent maintenance cost reduction from predictive maintenance real?
The commonly quoted figures of a 30 percent maintenance cost reduction and a 45 percent downtime reduction trace to the US Department of Energy's Operations and Maintenance Best Practices Guide, Release 3.0, published August 2010. That guide gives ranges of 25 to 30 percent and 35 to 45 percent and attributes them only to unnamed independent surveys, with no sample size, date, sector or method. The chapter's reference list contains just two entries, neither cited against those figures. Speaking in its own voice the same guide makes a more modest claim: 8 to 12 percent savings over a preventive maintenance programme alone.
Read next
China's University Major Cuts Are AI Policy, and Nigeria Should Read the Fine Print
The latest analysis essay.
Working on something in this space?
If this analysis is close to a problem you're thinking about, say so. I read every message personally.
Start a conversation