IoT Device Management: Registry, Configuration, and Firmware at Fleet Scale
Desired state versus reported state, bulk configuration, staged firmware rollouts, and why a two million device fleet is a control plane problem before it is a console problem.

A device management platform answers questions that ingestion cannot. Which devices exist, who owns them, what firmware are they on, what configuration should they have, and is the version they are running one we would still support.
It sounds like administration. It is not. Once a fleet is large enough, every one of those questions becomes an availability problem, because a mistake in the control plane reaches every device at once.
This article works through a fleet of two million devices across roughly five hundred customers.
What the control plane actually owns
Four things, and it is worth being precise about them because they have different failure characteristics.
registry -> identity, ownership, hardware revision, provisioning state
configuration -> the settings a device should be running
firmware -> the software version a device should be running
operations -> commands, jobs, and their outcomes
The registry is the control-plane source of identity and ownership. A rollout must tolerate it being temporarily unavailable: devices keep their last known-good state, artifact delivery uses already-issued authorization, and the controller pauses new assignments instead of taking telemetry down.
Configuration and firmware are large, frequently read, rarely written, and highly cacheable. This is what makes the desired-state model practical.
Operations is a job system wearing a device-shaped hat. Every operation is asynchronous, idempotent, and has an outcome that someone will eventually ask about.
Keeping these on one transactional database is completely reasonable. Keeping them on the same database as the telemetry write path is a mistake with a predictable bad day.
The boundary that matters
+-------------------------+ +--------------------------+
| Control plane | | Data plane |
| | | |
| registry | events | ingestion |
| configuration |<-------->| broker, bridge, Kafka |
| firmware manifests | | time series |
| operations, audit | config | alerting |
+------------+------------+ +------------+-------------+
| |
+--------- separate stores ---------+
separate failure domains
The control plane emits events. The data plane does the high-volume work. A rollout publishes a rollout event; the data plane’s device-facing service observes it and pushes the desired configuration to devices.
The reverse direction is the fleet’s own reported state, arriving as ordinary device messages and consumed by the control plane to update the registry.
control plane --(rollout started, desired version 4.2.0)--> device-facing service
device --(reporting version 4.2.0, healthy)---------> registry
Neither side calls the other synchronously. This is what lets you take the control plane down for maintenance without interrupting telemetry, and it is the difference between a rollout and an outage.
Desired state versus reported state
The central design decision. Get this right and configuration becomes almost trivially correct.
A naïve model is imperative. The operator clicks “set report interval to 30 seconds” and the platform sends a command to every device. Every device either receives it or does not, and the platform has to track which.
The declarative model is much better behaved:
desired state -> { reportIntervalSec: 30, sampleHz: 25, logLevel: "warn" }
reported state -> { reportIntervalSec: 30, sampleHz: 25, logLevel: "info", version: "4.2.0" }
The platform stores a versioned desired document and hashes a canonical serialization. The device reports both the version/hash it accepted and the version/hash it is actually running. Hashes are useful only if both sides agree on field ordering, defaults, normalization, and excluded local state.
device boots
-> fetch desired configuration for its identity
-> compute desired hash
-> compute hash of local configuration
-> equal? do nothing at all
-> different? apply the delta, reboot if required, report the new hash
Every property you want falls out of this.
The protocol is idempotent only when each apply step is designed that way and partial application is recoverable. It can self-heal if the last known-good state, desired document, credentials, and reconnect path survive the failure. Offline devices are not free: desired state must be retained, jobs need expiry semantics, credentials may rotate while the device is away, and reconnects must be rate-limited to avoid a fleet-wide thundering herd. Desired/reported history makes convergence auditable; one current hash does not.
The uncomfortable consequence is that a rollout is not complete when the operator clicks. It is complete when the last device reports a matching hash. Those are different moments and the interface has to show the second one.
The other uncomfortable consequence is that a device stuck on old configuration is invisible unless you are comparing desired against reported continuously. A configuration change that fails on a subset of devices is a silent partial rollout otherwise, and it will be found by a customer rather than by you.
Provisioning and claiming
A device arrives from a factory with an identity and no owner. Somebody has to attach it to a customer.
The flow that works at this scale:
- The device is manufactured with a hardware root of trust and a factory identity burned into secure storage.
- It ships. On first boot it presents that factory identity to the bootstrap endpoint.
- After proof of possession of a per-device bootstrap identity, the service returns a short-lived credential restricted to the claim workflow.
- The installer, or the customer through a portal, enters a claim code.
- The claim endpoint issues a certificate scoped to the actual tenant and device, and the registry records ownership.
That boundary depends on manufacturing identity and anti-cloning controls. A shared bootstrap secret turns one extracted device into a fleet credential. A per-device key in a hardware root of trust, proof of possession, one-time claim state, rate limits, and revocation are what make the restricted certificate meaningful.
The failure to plan for is double claiming. Two installers scan the same device within the same second. The claim must be a compare-and-swap on the registry’s ownership field, so the second attempt fails cleanly and visibly rather than both succeeding.
The other failure is unclaimed device drift. Devices that ship and are never installed sit in the registry forever. They need a lifecycle state, an expiry, and a sweeper that reclaims their credentials.
Configuration at fleet scale
Configuration is a tree. A tenant sets a default, a site overrides it, a device overrides that.
tenant default reportIntervalSec: 60
site "warehouse-3" reportIntervalSec: 30
device "sensor-4412" sampleHz: 25
The resolution has to be done in one place, and the result has to be what gets hashed. If the device resolves the tree locally and the platform resolves it centrally, the two can disagree, and a rollout computed from the platform’s view will look complete while devices are running something else.
The effective configuration is a resolved document, and the hash is over the resolved document. Not the delta, not the intent, the resolved result.
Two decisions worth making explicitly:
Blast radius. A tenant-wide configuration change touches every device for that tenant, so it needs the same staged rollout and soak discipline as a firmware update. Most platforms treat configuration as safe and firmware as dangerous, and that is backwards for large fleets. A bad config change can take a site offline exactly as effectively as a bad image.
Config size. A resolved document per device is small, but the control plane holds the effective configuration for two million devices, and it is recomputed whenever anything in the tree changes. Materialise it per device and update incrementally, or a single tenant-level change becomes a two-million-row rewrite.
Firmware updates are a bandwidth problem
This is the part that surprises people, so it is worth doing the arithmetic carefully.
A fleet of two million devices and a firmware image of 40 MB:
2,000,000 x 40 MB = 80 TB of image transferred for one full rollout
Now the constraint. Suppose a rollout can sustain 5,000 devices per hour, which is a modest CDN and network budget.
5,000 devices/hour x 40 MB = 200 GB per hour
2,000,000 devices = 400 hours for a full sweep
Sixteen days to update the whole fleet, if you go as fast as you can.
That number is the reason staged rollouts are not optional politeness. They are the mechanism that lets you detect a bad image on 1% of the fleet instead of on all of it.
A workable schedule:
stage 0 internal and canary devices soak 24h
stage 1 1% of the fleet soak 24h
stage 2 5% of the fleet soak 48h
stage 3 25% of the fleet soak 72h
stage 4 100% of the fleet
At 5,000 devices per hour, stage 1 is 20,000 devices and about four hours of transfer, and stage 4 is the sixteen-day sweep. The soak periods are where you find the problems, and they are the part that gets cut when a release is late.
The automated gate between stages should not be a human clicking approve. It should be a health comparison of the staged population against a control population:
boot success rate new firmware vs previous
crash rate within 24h new firmware vs previous
time to first message new firmware vs previous
reported error rate new firmware vs previous
Comparing against the previous version rather than an absolute threshold is what makes this work, because a fleet that is already unhealthy will fail an absolute threshold on a perfectly good image.
Rollback is a design requirement, not a contingency
A fleet-wide bad image is one of the few genuinely catastrophic IoT incidents, and its severity depends entirely on whether rollback is fast.
For unattended devices, an A/B layout is a strong rollback design: write the inactive slot, verify it, boot it provisionally, and revert if health confirmation never arrives. It is not the only design—some constrained devices use a recovery image or bootloader-assisted swap—but every design needs an independently bootable recovery path and an anti-rollback policy.
partition A current running image
partition B staging area for the new image
The sequence is: write the new image to the inactive partition, set it as the pending boot target, reboot, and confirm a successful boot within a short window. If the confirmation never arrives, the boot loader reverts to the other partition on the next restart.
With a valid previous slot, rollback can be a boot-selection change instead of redistributing the full image. The slot must still be intact, compatible with current data/config schema, and allowed by security rollback counters.
Without dual partitions, a rollback is another full image transfer, which on the fleet above is weeks. With them, a rollback is a message to two million devices and completes in hours.
One more rule: never make a new image the only bootable option for a device that cannot physically be reached. A device installed inside a sealed enclosure or on a rooftop with no access is a device that must be able to return to a working state without human contact, and that is what the revert-on-failed-boot counter guarantees.
Version skew is permanent
There will never be a moment when the whole fleet is on one version. There will be a distribution, and it will have a long tail.
4.2.0 62%
4.1.7 28%
4.1.3 7%
4.0.9 3%
Every design decision that touches version has to accept this:
- Queries. “Show me all devices on the latest firmware” is a real query and the answer is a list nobody wants to read.
- Dashboards. A metric that changed meaning between versions produces a chart with a discontinuity and no explanation.
- APIs. A device API that has a new field in 4.2 and not in 4.1 needs the field to be genuinely optional, and the platform needs to know not to send it.
- Support. When a customer reports a problem, the first question is which version, and the answer needs to be one lookup rather than an investigation.
The most useful thing a device management platform provides is a compatibility matrix: which versions this fleet has, what changed between them, and which of those changes matter. Without it, every cross-version question becomes archaeology.
The second most useful thing is refusing to let a version go below a floor. A device on a version that has been withdrawn from support is a device that will eventually be unfixable, and you want to know that in advance, as a scheduled report, not as an incident.
Group operations and bulk actions
A real platform needs “update these 40,000 devices”, “apply this config to everything in site 7”, “reclaim these 900 unclaimed devices”. These are jobs, and they need job semantics.
a job has: target set, operation, parameters, state, per-item outcome
The parts that are easy to get wrong:
Freezing the target set. A job targeting “all devices at site 7” evaluated at execution time picks up devices installed since the operator chose it. Materialise the target list when the job is created.
Per-item outcomes, not a single result. A bulk operation across 40,000 devices will partially succeed. The job’s status is a summary, and the detail is per device.
Cancelling. Some devices will already be updated. Cancel means “stop starting new work”, and the interface should say how many were already done.
Rate limits. A bulk job that immediately pushes to 40,000 devices is a self-inflicted denial of service against your own broker. Bulk operations go through the same staged mechanism as a rollout, with the same soak logic.
Audit. Who ran it, against what, when, with what outcome. For anything that changes configuration or firmware, this is not optional.
Health, and why quarantine must keep the device talking
Devices fail in ways that are not hardware failures. Wrong credentials, a clock far off, a certificate about to expire, impossible reported values, a device that has been reassigned to another tenant but is still authenticated with the old identity.
The response is quarantine: mark the device, restrict what it can do, and keep it reporting so the problem is diagnosable.
The mistake is treating quarantine as silence.
wrong: quarantine -> stop accepting its messages
right: quarantine -> accept and store its messages, flag them, restrict commands
A silent device is a black box. You know it is misbehaving and you cannot see how, which means you cannot tell whether the fix is a re-provision, a firmware revert, or a factory reset at the customer’s site. It also means the fleet’s reported state becomes unreliable exactly when you most need it to be accurate.
Quarantine should be scoped and gradual:
allow telemetry -> yes
allow state reporting -> yes
allow commands -> no
allow OTA -> no
And it should expire. A device quarantined for a credential problem should be released when the credential is fixed, and a device quarantined for ninety days should be escalated rather than left in a state nobody is maintaining.
The signals worth watching across the fleet:
| Signal | What it usually means |
|---|---|
| Time since last check-in | Network, power, or a device that has been physically removed |
| Version distribution | Rollout progress, and whether the long tail is growing |
| Configuration hash mismatch | A failed or partial rollout |
| Boot counter or failed boots | A bad image, or a failing boot loader |
| Certificate expiry horizon | Predictable outages, if you look far enough ahead |
| Time offset from the server | A device whose clock has drifted badly |
That last one is worth calling out. Clock skew quietly breaks ordering, retention, and every time-window query, and it is invisible until someone tries to correlate two devices’ events and gets an answer that makes no sense.
Failure stories worth testing
Ship an image that fails to boot
Confirm the boot loader reverts on its own, that the platform learns about it, and that the stage gate stops the rollout. Then repeat it at stage 1 and watch the gate block.
Kill the control plane mid-rollout
Telemetry should continue uninterrupted. The rollout should pause and resume. This is the test that proves the control plane and data plane are actually separate.
Make an unclaimed device attempt a claim twice concurrently
Confirm one claim wins and the other fails cleanly, with no possibility of two owners.
Push a configuration change to a site with poor connectivity
Half the devices will not receive it. The fleet view has to show that as a partial state rather than as success.
Retire a version that devices are still running
Confirm the platform reports the affected population instead of silently breaking their next config fetch.
Let a device report with a stale certificate after it was reassigned
Confirm the old identity is rejected and the event is quarantined rather than applied to the new owner.
Fill the CDN or egress budget mid-rollout
The rollout should slow down and remain correct, not fail. Bandwidth exhaustion on a sixteen-day sweep is a budget question that will happen.
A production-ready architecture
+----------------------+ events +----------------------+
| Control plane | -----------> | Device-facing service|
| | | (data plane) |
| registry | +----------+-----------+
| desired config | |
| firmware manifests | v
| rollout scheduler | +----------+-----------+
| operations, audit | | broker / ingestion |
+----------+-----------+ +----------------------+
^ |
| reported state (events) |
+------------------------------------+
Artifact store: firmware images, per-tenant CDN
Config store: resolved configuration per device, hash-indexed
Job store: per-item outcomes for every bulk operation
A sensible delivery checklist:
- Separate the control plane from the data plane in storage, in deployment, and in failure domain.
- Store desired state and compare hashes, rather than sending imperative configuration commands.
- Hash the resolved configuration, computed in one place, and compare against what the device reports.
- Materialise rollout and bulk-operation target sets when the job is created.
- Budget OTA bandwidth before committing to a schedule, and size the sweep from that number.
- Give devices dual boot partitions and a revert-on-failed-boot rule.
- Gate rollout stages on a health comparison against the previous version, not an absolute threshold.
- Treat configuration changes with the same staging discipline as firmware.
- Expect permanent version skew, publish a compatibility matrix, and set a support floor.
- Quarantine by restricting commands, never by silencing the device.
- Watch clock offset, certificate expiry horizon, and configuration hash mismatch as fleet metrics.
- Keep every configuration and firmware change in the audit log, per device, with the operator identity.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| Control plane sharing a database with ingestion | A rollout write load competes with telemetry writes and both degrade | Separate stores and separate failure domains |
| Imperative configuration commands | Offline devices need a separate retry path and never converge | Desired state compared by hash |
| Hashing the intent rather than the resolved config | Platform and device can disagree about what a device should be running | Hash the resolved document, computed centrally |
| No dual boot partitions | A bad image needs another 80 TB to roll back | A/B partitions with revert-on-failed-boot |
| Absolute health thresholds as rollout gates | An already-unhealthy fleet fails a good image | Compare the staged population against the previous version |
| Sizing a rollout by device count | Bandwidth, not devices, sets the schedule | Compute the sweep time from bytes and a transfer budget |
| Treating config changes as safe | A bad config takes a site offline as effectively as a bad image | Staged rollout and soak for configuration too |
| Assuming a single fleet version | Queries, dashboards, and APIs break on the long tail | Publish a compatibility matrix and a support floor |
| Silencing a quarantined device | The failure becomes undiagnosable and reported state goes unreliable | Restrict commands, keep accepting telemetry |
| Evaluating a job’s targets at run time | New devices join a job they were never meant to be part of | Materialise the target set at creation |
| Bulk jobs with no rate limit | 40,000 devices are pushed at once and the broker self-DoSes | Bulk operations use the staged mechanism |
| No clock monitoring | Ordering, retention, and time-window queries quietly break | Track time offset as a fleet metric |
| Assuming device ids in topics are identity | A device can act as another device | Derive identity from the authenticated connection |
| No audit on config or firmware changes | Every incident becomes an investigation into who changed what | Per-device audit with operator identity |
The complete story in one minute
Every device has an identity in a registry, and that registry is the source of truth for ownership, hardware revision, and current reported state. It lives in a control plane with its own store and its own failure domain, separate from the ingestion path.
Configuration is declarative. The control plane resolves tenant, site, and device layers into a versioned canonical document. The device fetches it, applies it through an atomic or resumable state machine, and reports accepted and running versions. A hash accelerates comparison; it does not replace versioning, apply status, or last-known-good rollback.
Firmware moves in stages with soak periods between them, and each stage is gated by comparing the boot success rate, crash rate, and time to first message of the staged population against the previous version. Devices have two boot partitions, so a failed boot reverts by itself and a rollback is a message rather than another 80 terabytes.
Bulk operations are jobs with a frozen target set, per-device outcomes, rate limits, and a full audit trail.
Devices that misbehave are quarantined by having their commands restricted, not by being silenced, because a silent device cannot be diagnosed and the fleet’s reported state is exactly what you need when something is wrong.
That is the whole path:
operator intent -> desired state -> hash -> device converges -> reported state -> registry
|
+-> firmware manifest -> staged rollout -> gated by health -> rollback path
The hard parts were never the console. They were keeping control-plane failure out of the data path, securing bootstrap and artifact metadata, making convergence a versioned fact, and designing recovery before shipping the image that needs it.


