Skip to content

Reference architecture

Industrial IoT predictive maintenance MLOps platform

An MLOps platform that flags likely machinery failures early enough to schedule repairs, instead of finding them as unplanned downtime.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Flag bearing and motor anomalies with enough lead time to plan repairs (target: 48 hours)
  • 02Process 50,000 sensor readings per second with automated feature computation
  • 03Automated model evaluation and retraining when drift is detected
  • 04Edge inference that keeps working in plants with intermittent or no connectivity

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Industrial IoT predictive maintenance MLOps platform

Illustrative reference architecture

  1. 01

    Telemetry ingestion

    Vibration, temperature and acoustic sensor streams

    MQTT + Apache Kafka

  2. 02

    Feature store

    Historical and real-time sensor features

    Feast + Redis / Snowflake

  3. 03

    ML training & experimentation

    Time-series anomaly detection and remaining useful life (RUL) models

    PyTorch + MLflow

  4. 04

    Edge inference engine

    Local model evaluation on factory hardware

    ONNX Runtime + NVIDIA Jetson

Subsystem 01

Telemetry ingestion

Vibration, temperature and acoustic sensor streams

Typical stack

MQTT + Apache Kafka

Subsystem 02

Feature store

Historical and real-time sensor features

Typical stack

Feast + Redis / Snowflake

Subsystem 03

ML training & experimentation

Time-series anomaly detection and remaining useful life (RUL) models

Typical stack

PyTorch + MLflow

Subsystem 04

Edge inference engine

Local model evaluation on factory hardware

Typical stack

ONNX Runtime + NVIDIA Jetson

Data lifecycle

End-to-end data flow

  1. Vibration sensors stream 1 kHz readings to edge gateways on the factory floor.

  2. Each gateway computes FFT frequency features and runs ONNX inference every 5 seconds.

  3. Features and anomalies sync to central Kafka topics in the cloud.

  4. Model performance is tracked in MLflow, and detected concept drift triggers a retraining pipeline.

  5. Maintenance managers receive work order recommendations on a mobile dashboard.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Sensor drift causing false alarms

Mitigation

Baseline normalization that adapts to seasonal changes in ambient temperature.

Failure mode 02

Factory network disconnections

Mitigation

Gateways buffer telemetry on local NVMe disks and keep alerting on their own while offline.

Failure mode 03

Concept drift after an equipment overhaul

Mitigation

Retraining triggered by Kolmogorov–Smirnov tests on the feature distributions.

Questions

What teams ask about this design

Typically autoencoders for unsupervised anomaly detection, with LSTM or transformer models and XGBoost for remaining useful life (RUL) estimation. The choice depends on how much labeled failure history you have.

Yes. The edge inference stack runs on standalone containerized servers with no internet connection; updated models are brought in through the plant's approved transfer process.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.