Unplanned equipment downtime in manufacturing costs an average of $260,000 per hour across industries — and most of it is preventable. Predictive maintenance (PdM) uses sensor data and machine learning to predict failures before they happen, shifting maintenance from reactive (break-fix) to proactive (predict-and-prevent). This guide details the engineering approach Zyllo Tech uses to build these systems.
Phase 1 — Which sensors does predictive maintenance need?
The quality of your predictive models depends entirely on the quality and coverage of your sensor data. We start with a sensor audit and gap analysis:
- Identify critical equipment assets and failure modes — not everything needs sensors; focus on highest-impact machines first.
- Vibration sensors for rotating equipment (motors, pumps, compressors) — the most predictive signal for bearing and shaft failures.
- Temperature sensors for thermal anomaly detection — motors running hot indicate lubrication or overload issues.
- Current/power monitoring for electrical machinery — power consumption spikes often precede failure.
- Acoustic sensors for high-frequency anomaly detection in pneumatic systems and gearboxes.
- OPC-UA protocol for modern industrial equipment; Modbus RTU/TCP for legacy PLCs and SCADA systems.
Phase 2 — Why can't all sensor data go to the cloud?
High-frequency sensor data (vibration at 10kHz+ sampling rates) cannot all be sent to the cloud — the bandwidth and storage costs are prohibitive. Edge computing filters and compresses the data:
- Edge gateway (industrial PC or AWS IoT Greengrass) runs feature extraction locally — RMS vibration, kurtosis, crest factor — reducing data volume by 95%+ before cloud transmission.
- Keep the raw waveform on the edge for a short retention window anyway: when a failure does occur, the seconds either side of it are the labelled training data you cannot reconstruct later.
- Time-series database on edge for 7-day local buffering in case of connectivity loss.
- Cloud time-series database: InfluxDB, TimescaleDB, or AWS Timestream for historical storage and querying.
- Data lake in S3/GCS for raw sensor dumps from initial deployment — useful for training models on historical failure events.
import numpy as np
from scipy.stats import kurtosis
def extract_features(window: np.ndarray, fs: int = 10_000) -> dict:
"""window: one second of single-axis vibration sampled at fs Hz."""
rms = float(np.sqrt(np.mean(window ** 2)))
peak = float(np.max(np.abs(window)))
spectrum = np.abs(np.fft.rfft(window))
return {
"rms": rms, # overall energy — climbs as a bearing wears
"peak": peak,
"crest": peak / rms, # peak-to-RMS — early impacts appear here first
"kurtosis": float(kurtosis(window)), # 'spikiness' — the classic bearing-defect signal
"dominant_hz": float(np.argmax(spectrum) * fs / len(window)),
}
# 10,000 float32 samples per second per axis = 40 KB/s raw.
# Five float32 features per second = 20 bytes/s.
# The model needs the TREND of these features across weeks, not the
# waveform — which is why this runs at the edge and not in the cloud.Phase 3 — ML Model Development
How do you detect anomalies without labelled failure data?
Most manufacturing facilities don't have labelled failure data for every machine and failure type. We start with unsupervised anomaly detection:
- Isolation Forest or Autoencoder-based anomaly detection trained on normal operating data — no labelled failures required.
- Statistical control charts (CUSUM, EWMA) as interpretable baselines that plant engineers trust.
- Seasonal decomposition to separate normal cyclical patterns (production shifts, temperature cycles) from genuine anomalies.
How do you predict a machine's remaining useful life?
Where historical failure data exists, we build supervised RUL models:
- LSTM or Transformer networks for time-series RUL prediction — trained on run-to-failure datasets.
- Survival analysis models (Cox Proportional Hazards) for probabilistic failure risk estimation.
- Model calibration: RUL predictions need confidence intervals, not just point estimates — a maintenance decision based on 'failing in 7 ± 2 days' is very different from '7 ± 30 days'.
Phase 4 — How do predictions become maintenance work orders?
- Work order creation in CMMS (SAP PM, IBM Maximo, Infor EAM) triggered automatically when failure probability exceeds a configured threshold.
- Maintenance scheduling optimised against production calendar — don't trigger maintenance during peak production windows.
- Spare parts inventory integration: validate that required replacement parts are in stock before creating a work order.
- OEE (Overall Equipment Effectiveness) dashboard combining PdM insights with production data.
Phase 5 — How do you prevent alert fatigue for technicians?
- Mobile app for technicians: work order details, equipment history, sensor readings, maintenance procedures, and POD capture.
- Alert fatigue management: threshold tuning and alert grouping to prevent technicians ignoring notifications.
- Explainability: show technicians which sensor readings triggered an alert — not just 'anomaly detected'.
- Feedback loop: technicians confirm or reject alerts after inspection — this feedback retrains the model.
- Unplanned Downtime Reduction: −42%
- Maintenance Cost Reduction: −28%
- False Alert Rate: < 8%
- MTBF Improvement: +35%
