If a chiller fails without warning, whose fault is it — the equipment, or the maintenance strategy?
1. Introduction
Every facilities engineer has lived through this moment: a chiller, an AHU, or a pump fails mid-operation, seemingly out of nowhere. Occupants notice before the BAS does. A technician gets dispatched. Parts get expedited. The building runs uncomfortably, sometimes for days, while the problem gets diagnosed.
Was the failure unavoidable — or was it actually building up for weeks, unnoticed, because nothing was watching for it?
That question is what this article answers.
2. How Maintenance Has Traditionally Worked
HVAC maintenance strategies have generally followed one of two models, and most buildings run some mix of both.
Time-based (scheduled) maintenance replaces or services components at fixed intervals — every 3 months, every year — regardless of actual condition. A belt gets replaced on schedule whether it has 10% or 90% of its life left. This is safe, predictable, and easy to budget for. It's also wasteful: a lot of perfectly good parts get replaced early, and a lot of labor gets spent servicing equipment that didn't need it yet.
Reactive (run-to-failure) maintenance is the opposite: equipment runs until it breaks, then gets repaired or replaced. This avoids unnecessary early replacement, but the cost shows up elsewhere — unplanned downtime, emergency labor rates, expedited parts shipping, and in HVAC specifically, occupant discomfort and sometimes cascading damage to adjacent equipment.
Both approaches have real, well-understood costs. Neither one is "wrong" — they're the best available options when you don't have visibility into how a piece of equipment is actually degrading in real time.
3. Where Traditional Maintenance Falls Short
The core problem with both models is the same: they aren't responding to actual equipment condition, because until recently, nobody had a practical way to continuously measure it.
A bearing beginning to wear doesn't announce itself. Long before it fails audibly, it may already be running with abnormal vibration, drawing slightly more current, or running a few degrees hotter than it should. None of that is visible on a maintenance calendar, and none of it triggers an alarm on a conventional BAS, because conventional systems are built to flag values that are already out of acceptable range — not values that are gradually drifting toward one.
By the time a fault is severe enough to trip an alarm or become audible, the window for a low-cost, planned intervention has usually already closed. What's left is an emergency repair, not a scheduled one.
4. What "Predictive" Actually Means
Predictive maintenance shifts the basis for action from time or failure to condition — using continuously monitored data to catch a developing problem while it's still developing, not after it's already failed.
This isn't a new idea in industrial equipment monitoring generally — vibration analysis and thermal imaging have existed in rotating-equipment maintenance for decades. What's changed for HVAC specifically is that AI now makes it practical to apply this continuously, automatically, and at scale across every piece of major equipment in a building, rather than relying on periodic manual inspections by a specialist.
The mechanism is straightforward: establish what "normal" operation actually looks like for a specific piece of equipment, then continuously watch for the moment its behavior starts to drift away from that baseline.
Figure 1 — From Schedules and Failures to Condition-Based Action
Figure 1 — AI shifts maintenance from fixed intervals or reactive repair to condition-based action, acting on actual equipment degradation rather than arbitrary schedules.
5. How AI Detects Developing Faults
The same historical operating data introduced in Part 2 — past temperatures, loads, and equipment behavior — plays a different role here. Instead of predicting building load, it's used to establish a normal operating signature for each specific piece of equipment: its typical vibration levels, current draw, discharge temperature, and efficiency, under its typical operating conditions.
Once that baseline exists, the system continuously compares live sensor readings against it, watching for gradual deviation rather than a single threshold breach. A compressor's current draw creeping upward over three weeks, a pump's vibration signature shifting slightly at a specific frequency, an efficiency curve that's quietly dropped a few percentage points since last quarter — these are exactly the kinds of slow, easy-to-miss trends that a fixed-threshold alarm was never designed to catch, but that a continuously learning baseline can flag early.
Figure 2 — From Sensor Data to Work Order
Figure 2 — AI continuously compares live sensor data against a learned baseline, detects anomalies early, and generates work orders before failure occurs — but a technician still diagnoses and repairs.
6. Why This Doesn't Replace Technicians
A predictive maintenance system doesn't diagnose or repair anything. What it does is narrow down where and when a technician's attention is actually needed — flagging a specific piece of equipment, a specific developing issue, and a rough timeline, instead of waiting for a scheduled inspection or an outright failure to surface the problem.
The technician still does what only a technician can do: confirm the fault, decide on the repair, and carry it out. AI changes the timing and targeting of that work — it doesn't remove the expertise required to act on it.
7. A Practical Example
Consider a chiller compressor operating normally for months. Around week 8, its vibration signature begins a slow upward drift above baseline — small enough week to week that it wouldn't stand out on a walkthrough, and nowhere near severe enough to trip a conventional alarm.
The drift continues gradually for several weeks. By week 13, it has been trending upward consistently enough that the system flags it as a developing issue — well before failure, but only once the trend is clearly distinguishable from normal variation, not from the very first small deviation.
A technician inspects the compressor during a planned maintenance window, confirms early bearing wear, and replaces it — on a schedule the team chose, using a part that was ordered in advance, without an unplanned outage.
Without that signal, the same trend continues to worsen, and the bearing is projected to fail outright by around week 20 — at a moment nobody chose, probably during occupied hours, and possibly taking the compressor's other components down with it.
Vibration Signature: Bearing Wear Progression
Chiller compressor — weeks 0 to 20
Figure 3 — AI flags the developing vibration trend at Week 13, well before the failure threshold is reached at Week 20. The detection window allows planned intervention instead of emergency repair.
📊 View data table for Figure 3
| Week | Vibration Signature | Failure Threshold | AI Early Warning Threshold |
|---|---|---|---|
| 0 | 1.00 | 3.40 | 1.55 |
| 2.5 | 1.00 | 3.40 | 1.55 |
| 5 | 1.05 | 3.40 | 1.55 |
| 7.5 | 1.10 | 3.40 | 1.55 |
| 10 | 1.20 | 3.40 | 1.55 |
| 12.5 | 1.35 | 3.40 | 1.55 |
| 13 | 1.55 | 3.40 | 1.55 — AI flags |
| 15 | 1.60 | 3.40 | 1.55 |
| 17.5 | 2.80 | 3.40 | 1.55 |
| 20 | 3.40 | 3.40 | 1.55 — Failure |
Values shown for illustrative purposes — actual results vary by equipment and operating conditions.
8. Where the Value Comes From
- Avoiding unplanned downtime, since the highest-cost failures are the ones nobody saw coming
- Avoiding cascading damage, since a failing component often takes adjacent equipment down with it the longer it runs undetected
- Better parts and labor planning, since a known, scheduled repair costs less than an expedited, emergency one
- Extending equipment life, by catching and correcting abnormal operating conditions before they accelerate wear on the rest of the unit
As with Part 2's energy savings, none of this is magic — it's the direct result of knowing something sooner than a conventional system would have told you.
9. Real-World Evidence
Predictive maintenance platforms across industrial and commercial HVAC have reported similar outcomes: fewer emergency callouts, and equipment faults caught weeks ahead of failure rather than discovered at the point of breakdown. As one illustration, several commercial and industrial deployments of vibration- and current-signature-based monitoring on rotating HVAC equipment (compressors, pumps, and fan motors) have reported measurable reductions in unplanned downtime after early-stage faults were flagged and addressed proactively.
As with the energy optimization case in Part 2, results like these depend heavily on sensor coverage, data history, and how well a specific deployment was configured — they illustrate the mechanism, not a guaranteed outcome for every building.
10. Limitations
- It cannot predict sudden, precursor-free failures. Some failures — a power surge, a control-wiring fault, physical damage — happen without any gradual signature to detect in advance
- It depends entirely on sensor coverage. Equipment without adequate instrumentation simply can't be monitored this way, no matter how good the underlying algorithm is
- It depends on enough operating history. A newly commissioned system, or one that's just had major work done, hasn't yet built the baseline needed to detect meaningful drift
- False positives carry a real cost. An overly sensitive system that flags too many non-issues erodes technician trust in the alerts — and once that trust is gone, the alerts get ignored, defeating the purpose
In short: predictive maintenance narrows the gap between problem and detection. It doesn't eliminate the need for good instrumentation, good data history, or good judgment.
Key Takeaways
- Time-based and reactive maintenance both have real costs — one wastes labor and parts, the other creates expensive emergencies
- Both models fail because they don't respond to actual equipment condition
- AI establishes a normal operating baseline for each piece of equipment and watches for gradual drift
- Predictive maintenance doesn't replace technicians — it flags what and when to inspect
- Early detection enables planned repairs, avoiding unplanned downtime and cascading damage
- Limitations include sudden failures, sensor coverage gaps, insufficient history, and the cost of false positives
11. Looking Ahead
Energy optimization and predictive maintenance both come from the same underlying shift: using data the building already generates to act earlier and more precisely. In Part 4, we'll look at a different application of that same shift — how AI supports indoor air quality and intelligent ventilation control.
In your experience, what's harder to get right in practice: getting enough reliable sensor coverage, or trusting the alerts once you have them?