← All writing
articleFeb 17, 202518 min read

Industrial IoT Platforms: Where the Cloud Stops and the Plant Starts

The IT and OT boundary, the edge, OPC-UA versus Modbus, historians, digital twins, and why a control loop must never depend on a network it does not own.

IoTEdgeOPC-UAArchitecture
Industrial IoT Platforms: Where the Cloud Stops and the Plant Starts cover illustration

A consumer IoT platform assumes a device can be offline for an hour and nothing much happens. An industrial platform cannot make that assumption, because the device is usually part of a process where the physical consequence of stopping is measured in tonnes, and sometimes in injuries.

That single difference propagates through every design decision. It is why industrial IoT has an edge, why the protocols are strange, why the historian exists as a separate concept from the database, and why a whole category of “optimisation” is off limits.

This article works through a platform spanning 100 factories and about 100,000 sensors.

Two systems, not one

The first thing to get right is that a plant contains two systems with incompatible requirements, and most of the difficulty comes from pretending they are one.

IT OT
Restarting acceptable may stop production
Latency seconds is fine milliseconds, bounded
Downtime cost lost revenue damaged product, safety risk, contractual penalty
Deployment frequent rare, scheduled, during a shutdown
Data accuracy usually good enough must be traceable and defensible
Lifetime 3 to 5 years 15 to 25 years
Protocols HTTP, SQL, gRPC Modbus, OPC-UA, PROFINET, EtherNet/IP

The last row is the one that surprises people joining industrial work from a web background. A PLC that went into service in 2004 will still be running in 2029, and its protocol will not have changed. Anything you build has to coexist with that reality for longer than your product’s roadmap.

The hierarchy is real

Plants have a layered structure that predates every vendor’s product. It maps closely onto the ISA-95 and IEC 62264 models, and it is worth internalising because most integration pain comes from flattening it.

Level 4   ERP          business planning, orders, purchasing
Level 3   MES          production execution, work orders, genealogy
Level 2   SCADA        supervisory control, HMI, historians
Level 1   Control      PLCs, safety controllers, drives
Level 0   Process      sensors, transmitters, actuators

Level 0 produces data. Level 1 acts on it in a control loop measured in milliseconds. Level 2 aggregates and displays. Level 3 coordinates production. Level 4 plans.

The critical property is that each layer must keep working if the layer above it disappears. A plant where losing the ERP connection stops the line is badly designed, and in a mature plant it is designed that way precisely because the failure has happened before.

Your platform is an observer of this hierarchy. It should not be a new layer in it.

Why the edge exists

“Put it at the edge” is a phrase that hides four separate requirements, each of which is independently sufficient to justify edge compute.

Control-loop latency. A PID loop on a temperature jacket closes every ten to fifty milliseconds. Network jitter alone makes this impossible over a WAN, and a cloud round trip of twenty to two hundred milliseconds is one to three orders of magnitude outside the budget. A loop that depends on a cloud call is a loop that will eventually fail open, fail closed, or oscillate.

Bandwidth. Assume 100 factories with 1,000 sensors each. A realistic mix:

2,000 vibration sensors at ~1,600 Hz  = 3.2 million readings/sec
98,000 process sensors at ~1 Hz        = 98,000 readings/sec
                                        -------------------
                                        ~3.3 million readings/sec
3.3 million/sec x 86,400 sec x 60 bytes = 17 TB per day

Seventeen terabytes a day, six petabytes a year, from a plant network that is typically a 1 Gbps industrial uplink shared with everything else on site. The uplink is not the constraint. The fact that nobody wants to pay for it is the constraint.

Availability. When the WAN goes down, and it will, the plant does not stop. An edge gateway that buffers locally and forwards when the link returns is the difference between a lost morning of telemetry and a lost week of production.

17 TB/day / 100 plants = 170 GB per plant per day

A plant server can hold thirty days of that. The buffering requirement is genuinely affordable at the edge and genuinely unaffordable in the cloud.

Protocol translation. A cloud region cannot hold a TCP connection to a PLC on a private network that nobody will expose to the internet. The bridge from Level 0 and 1 protocols into anything a cloud service can consume is physically required, and that is what the edge is.

The other reason, which is not technical, is that plants frequently will not send their process data outside the network at all. Oil, defence, pharmaceuticals, and some food production have real constraints. A platform that assumes egress will be told no, and should be designed so that the answer degrades the product rather than removing it.

Protocols, honestly

Modbus is register-oriented. A slave device exposes a table of memory addresses holding numbers, and the master reads and writes them by address. There is no authentication, no encryption, no session concept, and no notion of what a value means. It is ubiquitous because it is simple, cheap, and predates everything else, not because it is good.

read holding register 40001 from unit 1

That is the whole protocol. Which means the gateway has to be told what register 40001 is on each device model, and that mapping is per-model configuration, not code. On a real fleet it is a large, ugly, essential table that someone maintains for years.

OPC-UA is a different category. It is not a better Modbus, it is a different model:

  • an address space of nodes, each with a NodeId, a browse name, a data type, and attributes;
  • references linking nodes, so a temperature is a child of a tag which is a child of a machine;
  • typed data, so a value is a float rather than a 16-bit register holding half of one;
  • subscriptions with a publishing interval, where the server pushes only what changed rather than the client polling;
  • alarms and conditions as a first-class model, not a log of strings;
  • security designed in, with signed and encrypted messages and separate application and user identity.
server
  ns=2;s=Machine.Pump01.Tag.Temperature
       NodeId        BrowseName path
  value: 71.4 degC   (data type is declared, not inferred)

Security is where OPC-UA is genuinely different from Modbus, and it is also where deployments get it wrong. The protocol supports None, Sign, and SignAndEncrypt, and a large fraction of deployed OPC-UA servers run in None because that is what worked during commissioning. Securing an OPC-UA deployment means configuring message security mode, security policy, and both application and user identity, and treating each as a separate change with its own test.

The trust model is worth knowing too: OPC-UA uses trust lists of certificate thumbprints rather than a public certificate authority, which is the right call for an isolated plant network and the wrong call if you assume a certificate can be revoked centrally.

The migration is a modelling project. Converting a Modbus plant to OPC-UA means building a meaningful information model, not rewriting an address. Budget for the modelling, because that is where the time goes.

The data model

Three things are stored, and they are different things.

Telemetry, high volume, time-ranged, written and rarely updated. Temperature, pressure, vibration, current. This is the time-series store, and the whole of the earlier discussion about downsampling and cardinality applies.

Events, low volume, high value, never expire. State changes, alarms, operator actions, configuration modifications, firmware versions. This is a transactional store, and it is the part you will be asked to produce evidence from.

Assets, slow changing, relational, deeply nested. A plant has lines, lines have machines, machines have tags, tags have sensors. This is a graph or a relational hierarchy, and it is what every query joins against.

telemetry  ->  time series     -> sensor_id + recorded_at
events     ->  transactional   -> asset_id + recorded_at, immutable, retained
assets     ->  relational      -> plant -> line -> machine -> tag -> sensor

The common mistake is collapsing telemetry into events. A temperature change is not an event, and a million of them a day will make an event store either unusable or unmanageable. They have different lifetimes, different query patterns, and different retention requirements, and merging them means the merged version is bad at both.

The other common mistake is treating assets as a flat list of sensors. A sensor that belongs to a pump on line 3 in plant 7 is not describable by a tag string, and every dashboard that needs to show “line 3” depends on that hierarchy existing properly.

The historian

The historian predates most vendor products and is worth understanding on its own terms, because it answers a question the other stores do not.

A historian is an immutable, append-only, time-stamped record of what a plant actually did, written as it happened and never altered. Not a database you can query for current state. A record.

Three properties matter.

It is the legal and contractual record. When a pharmaceutical batch is rejected, the historian is what proves what the temperatures were. This is why OT data is “traceable and defensible” in the requirements table, and why you do not get to correct it later.

It is write-optimised and reads are historical. Everything goes in the order it happened. Nobody was querying it while it was being written.

It must not block the plant. A historian that stalls and causes a control system to fault is a historian that gets replaced, and getting replaced with a worse product is a normal and sad outcome. Its write path has a bounded, explicitly lossy mode where it drops data rather than propagating backpressure into the control network.

That last property is the one that distinguishes a real historian from a database someone connected to a PLC. Data loss is acceptable, production is not.

The digital twin

A digital twin is a virtual representation of a physical asset, and the useful versions of the concept are narrower than the marketing suggests.

The distinctions worth enforcing:

It is a copy, and the copy is labelled as a copy. If a dashboard cannot tell you whether it is showing the physical asset or the model of it, it is dangerous.

It is versioned against the asset’s configuration. A twin built when the pump had a different impeller is a twin of a machine that no longer exists. Pair every twin with the configuration it was built from.

It has a timestamp, and it is not now. A twin computed from data that is two hours old is a two-hour-old twin. A platform that presents it as live is lying.

It never becomes authoritative for the physical world. A twin that predicts a bearing will fail at 40 hours is a prediction. If someone acts on it and the bearing fails at 30, the prediction was useful. If someone skips an inspection because the twin said it was fine, the prediction has been used as a safety case it was never qualified for.

The genuinely valuable twins are not the marketing ones. They are:

  • a configuration and topology model the platform can query, so “which sensors on line 3 changed in the last firmware rollout” is answerable;
  • a behavioural model used for what-if analysis, kept clearly separate from live data;
  • a maintenance history model, which is just the event history structured as an asset timeline.

The last of those is often the highest value and the least glamorous, and it is entirely built from the event store that already exists.

Predictive maintenance

The pattern is straightforward and the honest version has an edge component that is not optional.

At the edge: feature extraction. A vibration sensor producing 1,600 samples per second cannot have meaningful features computed in a cloud region, and it does not need to.

raw 1,600 Hz
    -> remove DC, filter
    -> RMS, peak, crest factor, kurtosis
    -> FFT, band energies
    -> 1 Hz feature vector

That is a reduction of about 1,600 to 1, and the output is a 30-number vector per second per sensor. The cloud receives 1,600 times less data and the model gets features rather than raw signal.

In the cloud: the model. Trained on labelled failure data, which is expensive and rare, which is the real constraint on this whole category of product.

In the gap: the data you discarded. If you only kept features, you cannot go back and re-extract with a better method. Keeping a short rolling raw window on the edge, say one minute, means a new feature can be computed retrospectively for the last minute. Keeping the full raw stream is the alternative and it costs 1,600 times the bandwidth for a benefit most programmes never use.

The honest limitation is that these models are usually trained on one asset class and one failure mode, and a pump bearing model will not transfer to a gearbox. Every deployment is a project, not a configuration, and a platform promising otherwise is selling something it cannot deliver.

Safety is not a feature

This is the part where an IT-shaped instinct causes real harm, so it is worth being blunt.

Industrial safety is governed by functional safety standards, principally IEC 61508 and IEC 61511. They define safety integrity levels, and a safety instrumented function is required to perform correctly under a specified probability of dangerous failure. Emergency stops, interlocks, and overpressure trips are implemented in dedicated, independently certified hardware for this reason.

never:
    cloud call -> PLC input  -> decides to stop the line

That design is unsafe for reasons that are not subtle. The network is not in your control. The latency is not bounded. The availability is not guaranteed. The system is not certified. And the failure mode is that the safety function silently does nothing during the exact conditions it exists for.

The rule is simple: safety lives in the control layer, and the cloud observes it. Your platform reads the state of a safety system. It does not decide anything about it.

This is not a technical preference. Connecting a cloud service into a safety path invalidates the certification of the safety function and, in most jurisdictions, is unlawful.

Where cloud involvement is genuinely wanted, the legitimate patterns are advisory and supervisory:

  • an operator sees a recommendation and approves it;
  • a control system pulls a plan from the cloud and executes it locally with its own safety interlocks still in force;
  • a historian records what happened and a cloud system analyses it afterwards.

All three keep the safety function local. The distinction is not whether the cloud is involved, it is whether a cloud failure can cause an unsafe physical state.

The network boundary

The IT and OT boundary is a real security boundary, and the design choices are visible choices about risk.

  [ corporate IT ]  ---  firewall / DMZ  ---  [ control network ]  ---  [ plant floor ]
                                                    |
                                              field devices

Reasonable controls, roughly in order of how much they matter:

Unidirectional or filtered data flow out. A one-way gateway from the control network to the DMZ removes an entire class of incident. Where regulations require it, this is mandatory.

A protocol allowlist at the gateway. Only OPC-UA, only specific node ids. This turns a flat network into one where the industrial gear is not reachable by anything that can run a port scan.

No inbound paths to Level 1. Remote access to a PLC should be brokered, time-boxed, and recorded, not opened. Vendor remote support is a real operational need and a real risk, and the answer is a session broker rather than a firewall exception.

Separate management and data networks. Configuration and telemetry on the same network means a configuration mistake is also a data plane incident.

The control network is not flat. Modern plants use cells and zones with switches that enforce boundaries, because a flat Level 1 network is one compromised laptop away from a plant.

None of this is exotic, and all of it is regularly skipped by platforms that arrived from IT and assumed the customer already had it sorted. Assume nothing, and audit it.

Failure stories worth testing

The plant runs normally. The edge buffers, the data arrives late, and nothing is lost. This is the single most important test in the whole platform.

Inject 200 milliseconds of latency into the control network

Anything that depended on a round trip should fail loudly. If nothing breaks, you have verified that the control loop is local, and that is worth knowing.

Lose the SCADA server

The line keeps running on the PLC. Level 2 is a display and a historian, and it is not in the control path.

Fail a historian mid-write

Confirm it enters its lossy mode rather than stalling. Then confirm you can tell how much data was dropped, because an unexplained gap in a batch record is a compliance problem.

Untrust an OPC-UA server

Reconnecting a compromised server to the gateway is the nightmare scenario, and a gateway that re-establishes sessions blindly will do it. Verify the deny list works while a device is connected.

Run a firmware rollout across 100 plants with one bad image

The staging gate has to work at fleet scale, not per plant. Confirm the rollout can be halted globally in minutes.

Lose the time sync source

Everything downstream becomes uncorrelated. Industrial sites running PTP or NTP with a local holdover should keep a usable clock for longer than the outage.

Let an operator change a tag description

Confirm the change is an event with an identity and a timestamp, not an update. In a regulated environment, “who changed this and when” is the question you will be asked.

A production-ready architecture

  [ plant floor ]
    sensors, actuators, PLCs, drives
            |
    field network  (Modbus, OPC-UA, PROFINET)
            |
  +---------v----------+
  | Edge gateway       |
  | protocol bridge    |
  | feature extraction |
  | local buffering    |
  | store-and-forward  |
  +---------+----------+
            |  MQTT / OPC-UA PubSub, one-way where possible
      [ DMZ / site server ]
            |  mTLS, allowlisted
  +---------v----------+
  | Cloud ingestion    |
  +---------+----------+
            |
   +--------+-----------+------------+-----------------+
   |                    |                |                 |
+--v---------+  +-------v-------+  +-----v------+  +-------v--------+
| Time series|  | Event store   |  | Asset/CMDB |  | Object storage |
| telemetry  |  | immutable     |  | hierarchy  |  | raw archive    |
+------------+  +---------------+  +------------+  +----------------+
   |                    |
   +--------+-----------+
            |
     models, dashboards, MES integration

Sits outside the loop, always:

  Safety systems      certified, independent, local
  Control loops       PLC, local, bounded latency

A sensible delivery checklist:

  1. Audit the actual plant topology and protocols before designing anything, including the register maps.
  2. Prove the control loop is local by injecting latency and confirming nothing breaks.
  3. Put feature extraction at the edge, and decide explicitly what raw data you keep.
  4. Size edge buffering for a multi-day WAN outage, and test that outage.
  5. Separate telemetry, events, and assets into three stores with three retention policies.
  6. Make the historian write path bounded and explicitly lossy, and record what it dropped.
  7. Treat every twin as a versioned, timestamped copy, and never as authority.
  8. Keep safety systems out of every cloud path, and audit that the boundary holds.
  9. Secure OPC-UA properly: message security mode, security policy, and both identity types.
  10. Assume a non-flat control network, a brokered path for vendor access, and no inbound routes to Level 1.
  11. Keep MES and ERP integration explicit and late, because ISA-95 boundaries are where the modelling effort belongs.

Common mistakes

Mistake What actually happens Better decision
A cloud call in the control loop A network spike stops a production line, or worse Keep the loop local; the cloud observes
Safety functions in the platform Certification is invalid and the design is unsafe Safety stays in certified local hardware
Sending raw waveforms to the cloud 1,600x the bandwidth for analysis nobody does remotely Extract features at the edge, keep a short raw window
Flattening the ISA-95 hierarchy “Line 3” becomes unanswerable and integration becomes guesswork Model the hierarchy properly from the start
Telemetry and events in one store Either unusable volume or lost audit value Three stores, three retention policies
A historian that can backpressure the plant Data loss is preferred, and the historian gets replaced Bounded lossy write path that records what it dropped
Calling a digital twin live A two-hour-old model is read as present fact Timestamp and label every twin as a copy
Modbus register maps hard-coded Every device model is a code change forever Per-model configuration, maintained deliberately
OPC-UA left in None security mode The information model is there and the security is not SignAndEncrypt, and both identity types configured
Assuming the plant has a DMZ The control network is one scan away Audit it; segment the control network into cells
Opening inbound paths for vendor support Remote access becomes permanent access Brokered, time-boxed, recorded sessions
Depending on the corporate NTP server A WAN outage silently corrupts every correlation Local holdover on the control network
Correcting a historian record Batch traceability and audit are gone Append-only, with corrections as new events
Uniform retention across data types Cheap telemetry kept at forensic cost, or evidence deleted Retention per data class, justified per class

The complete story in one minute

A plant has layers, and each one has to keep working when the layer above it fails. A sensor produces data, a PLC closes a control loop in milliseconds, a SCADA system displays and records, an MES coordinates production.

Your platform observes that hierarchy. An edge gateway at each plant bridges Modbus, OPC-UA, and proprietary field protocols into MQTT, extracts features from high-frequency signals so a 1,600 Hz sensor sends a 1 Hz vector instead of raw samples, and buffers locally for a WAN outage that is going to happen.

The gateway forwards north over an allowlisted, authenticated, preferably one-way path. Cloud ingestion stores three different things: telemetry in a time series store, immutable events in a transactional store, and the asset hierarchy relationally.

Analytics run on the features, models run in the cloud, and MES integration happens at the ISA-95 boundary where the modelling belongs. A historian keeps the append-only record of what the plant actually did, and its write path is allowed to drop data rather than ever stall a controller.

Safety systems and control loops remain local, certified, and independent. The platform reads their state and never writes to them.

That is the whole path:

sensor -> PLC (local loop) -> edge gateway -> MQTT -> telemetry | events | assets
                |
                +-> safety systems, never connected to the cloud

The hard parts were never the dashboard. They were respecting the fact that the plant is a safety-critical system with a twenty-year asset lifetime, that the control loop has a latency budget the internet cannot meet, and that the most valuable thing the platform can do is prove it is not in the way.

What this team still owns

The plant control and safety layers remain authoritative. The platform may observe, aggregate, simulate, and propose; a cloud timeout must not become a control input. Any write path is separately authenticated, allow-listed by operation, bounded by equipment state, recorded with operator intent, and rejected safely at the edge. OPC UA security modes and transport encryption do not replace network zoning, asset inventory, certificate lifecycle, or a physical hazard analysis.

Technical references

Keep reading
Browse everything