If a chiller fails without warning, whose fault is it — the equipment, or the maintenance strategy?

1. Introduction

Every facilities engineer has lived through this moment: a chiller, an AHU, or a pump fails mid-operation, seemingly out of nowhere. Occupants notice before the BAS does. A technician gets dispatched. Parts get expedited. The building runs uncomfortably, sometimes for days, while the problem gets diagnosed.

Was the failure unavoidable — or was it actually building up for weeks, unnoticed, because nothing was watching for it?

That question is what this article answers.

2. How Maintenance Has Traditionally Worked

HVAC maintenance strategies have generally followed one of two models, and most buildings run some mix of both.

Time-based (scheduled) maintenance replaces or services components at fixed intervals — every 3 months, every year — regardless of actual condition. A belt gets replaced on schedule whether it has 10% or 90% of its life left. This is safe, predictable, and easy to budget for. It's also wasteful: a lot of perfectly good parts get replaced early, and a lot of labor gets spent servicing equipment that didn't need it yet.

Reactive (run-to-failure) maintenance is the opposite: equipment runs until it breaks, then gets repaired or replaced. This avoids unnecessary early replacement, but the cost shows up elsewhere — unplanned downtime, emergency labor rates, expedited parts shipping, and in HVAC specifically, occupant discomfort and sometimes cascading damage to adjacent equipment.

Both approaches have real, well-understood costs. Neither one is "wrong" — they're the best available options when you don't have visibility into how a piece of equipment is actually degrading in real time.

3. Where Traditional Maintenance Falls Short

The core problem with both models is the same: they aren't responding to actual equipment condition, because until recently, nobody had a practical way to continuously measure it.

A bearing beginning to wear doesn't announce itself. Long before it fails audibly, it may already be running with abnormal vibration, drawing slightly more current, or running a few degrees hotter than it should. None of that is visible on a maintenance calendar, and none of it triggers an alarm on a conventional BAS, because conventional systems are built to flag values that are already out of acceptable range — not values that are gradually drifting toward one.

By the time a fault is severe enough to trip an alarm or become audible, the window for a low-cost, planned intervention has usually already closed. What's left is an emergency repair, not a scheduled one.

4. What "Predictive" Actually Means

Predictive maintenance shifts the basis for action from time or failure to condition — using continuously monitored data to catch a developing problem while it's still developing, not after it's already failed.

This isn't a new idea in industrial equipment monitoring generally — vibration analysis and thermal imaging have existed in rotating-equipment maintenance for decades. What's changed for HVAC specifically is that AI now makes it practical to apply this continuously, automatically, and at scale across every piece of major equipment in a building, rather than relying on periodic manual inspections by a specialist.

The mechanism is straightforward: establish what "normal" operation actually looks like for a specific piece of equipment, then continuously watch for the moment its behavior starts to drift away from that baseline.

Figure 1 — From Schedules and Failures to Condition-Based Action

⏱ Time-Based / Reactive Fixed intervals, or repair after failure AI 🔮 Predictive Maintenance Acts on actual equipment condition

Figure 1 — AI shifts maintenance from fixed intervals or reactive repair to condition-based action, acting on actual equipment degradation rather than arbitrary schedules.

The goal isn't to catch the very first data point that deviates — a single reading can be noise. It's to recognize a sustained trend early enough that the gap between when a problem starts and when someone finds out about it shrinks from months to weeks.

5. How AI Detects Developing Faults

The same historical operating data introduced in Part 2 — past temperatures, loads, and equipment behavior — plays a different role here. Instead of predicting building load, it's used to establish a normal operating signature for each specific piece of equipment: its typical vibration levels, current draw, discharge temperature, and efficiency, under its typical operating conditions.

Once that baseline exists, the system continuously compares live sensor readings against it, watching for gradual deviation rather than a single threshold breach. A compressor's current draw creeping upward over three weeks, a pump's vibration signature shifting slightly at a specific frequency, an efficiency curve that's quietly dropped a few percentage points since last quarter — these are exactly the kinds of slow, easy-to-miss trends that a fixed-threshold alarm was never designed to catch, but that a continuously learning baseline can flag early.

Figure 2 — From Sensor Data to Work Order

📊 Sensor Data 📈 Compare to Learned Baseline ⚠ Anomaly Detected 📋 Work Order Generated before failure occurs A technician still diagnoses and repairs — AI flags what and when

Figure 2 — AI continuously compares live sensor data against a learned baseline, detects anomalies early, and generates work orders before failure occurs — but a technician still diagnoses and repairs.

6. Why This Doesn't Replace Technicians

A predictive maintenance system doesn't diagnose or repair anything. What it does is narrow down where and when a technician's attention is actually needed — flagging a specific piece of equipment, a specific developing issue, and a rough timeline, instead of waiting for a scheduled inspection or an outright failure to surface the problem.

The technician still does what only a technician can do: confirm the fault, decide on the repair, and carry it out. AI changes the timing and targeting of that work — it doesn't remove the expertise required to act on it.

7. A Practical Example

Consider a chiller compressor operating normally for months. Around week 8, its vibration signature begins a slow upward drift above baseline — small enough week to week that it wouldn't stand out on a walkthrough, and nowhere near severe enough to trip a conventional alarm.

The drift continues gradually for several weeks. By week 13, it has been trending upward consistently enough that the system flags it as a developing issue — well before failure, but only once the trend is clearly distinguishable from normal variation, not from the very first small deviation.

A technician inspects the compressor during a planned maintenance window, confirms early bearing wear, and replaces it — on a schedule the team chose, using a part that was ordered in advance, without an unplanned outage.

Without that signal, the same trend continues to worsen, and the bearing is projected to fail outright by around week 20 — at a moment nobody chose, probably during occupied hours, and possibly taking the compressor's other components down with it.

Vibration Signature: Bearing Wear Progression

Chiller compressor — weeks 0 to 20

Actual Vibration Signature Failure Threshold AI Early Warning Threshold Detection Window (Week 13)

Figure 3 — AI flags the developing vibration trend at Week 13, well before the failure threshold is reached at Week 20. The detection window allows planned intervention instead of emergency repair.

📊 View data table for Figure 3
Week Vibration Signature Failure Threshold AI Early Warning Threshold
01.003.401.55
2.51.003.401.55
51.053.401.55
7.51.103.401.55
101.203.401.55
12.51.353.401.55
13 1.55 3.40 1.55 — AI flags
151.603.401.55
17.52.803.401.55
20 3.40 3.40 1.55 — Failure

Values shown for illustrative purposes — actual results vary by equipment and operating conditions.

AI doesn't replace the technician. It changes the timing and targeting of their work — flagging what and when, not diagnosing or repairing.

8. Where the Value Comes From

As with Part 2's energy savings, none of this is magic — it's the direct result of knowing something sooner than a conventional system would have told you.

9. Real-World Evidence

Predictive maintenance platforms across industrial and commercial HVAC have reported similar outcomes: fewer emergency callouts, and equipment faults caught weeks ahead of failure rather than discovered at the point of breakdown. As one illustration, several commercial and industrial deployments of vibration- and current-signature-based monitoring on rotating HVAC equipment (compressors, pumps, and fan motors) have reported measurable reductions in unplanned downtime after early-stage faults were flagged and addressed proactively.

As with the energy optimization case in Part 2, results like these depend heavily on sensor coverage, data history, and how well a specific deployment was configured — they illustrate the mechanism, not a guaranteed outcome for every building.

10. Limitations

In short: predictive maintenance narrows the gap between problem and detection. It doesn't eliminate the need for good instrumentation, good data history, or good judgment.

Key Takeaways

  • Time-based and reactive maintenance both have real costs — one wastes labor and parts, the other creates expensive emergencies
  • Both models fail because they don't respond to actual equipment condition
  • AI establishes a normal operating baseline for each piece of equipment and watches for gradual drift
  • Predictive maintenance doesn't replace technicians — it flags what and when to inspect
  • Early detection enables planned repairs, avoiding unplanned downtime and cascading damage
  • Limitations include sudden failures, sensor coverage gaps, insufficient history, and the cost of false positives

11. Looking Ahead

Energy optimization and predictive maintenance both come from the same underlying shift: using data the building already generates to act earlier and more precisely. In Part 4, we'll look at a different application of that same shift — how AI supports indoor air quality and intelligent ventilation control.

In your experience, what's harder to get right in practice: getting enough reliable sensor coverage, or trusting the alerts once you have them?

Part 2: AI-Based Energy Optimization