Connected Vehicle Platforms: A Client That Lives Longer Than Your Backend
The fifteen year fleet problem, what a CAN bus will actually let you log, command safety on an intermittently connected car, and why a connected vehicle has four data subjects and not one.

A refrigerator has a three year product life. A car has fifteen, and a manufacturer that signs a ten year warranty has committed to supporting the vehicle after its own backend was last rebuilt twice.
That mismatch is the defining constraint of connected vehicle work, and it is not a scalability problem. You can buy throughput for a connected fleet. You cannot buy backward compatibility with a car that left the factory in 2026 and will still be on the road in 2041.
The fleet, for scale
Ten million connected vehicles, roughly 60 percent moving at any moment.
6M moving, reporting every 30s = 200,000 msg/sec
4M parked, reporting hourly = 1,111 msg/sec
-----------
~200,000 msg/sec average
Morning commute concentration pushes that to around a million messages per second for a couple of hours, because vehicles follow the same roads on the same schedule.
200,000/sec x 86,400 = 17.3 billion messages/day
x 365 = 6.3 trillion messages/year
The throughput is the easy part. Ten million devices is a fleet, and the ingestion article’s architecture absorbs it.
The firmware side is where the arithmetic is uncomfortable:
10,000,000 vehicles x 3 GB image = 30 PB per complete OTA cycle
at ~1 MB/s over cellular = ~50 minutes of download per vehicle
Thirty petabytes, and a vehicle that has to sit still, on power, with a good signal, for the duration. An OTA campaign is not a deployment, it is a logistics operation with a fifteen year tail.
What you can actually log
The most persistent disappointment in connected vehicle work is discovering that the interesting data is not available.
The vehicle has internal networks. Older cars use CAN, a broadcast bus from 1986 running at 250 kbit/s to 1 Mbit/s, where every frame is eight bytes and the message identifier is also its priority, with lower IDs winning arbitration. Trucks use J1939, a 250 kbit/s extension of CAN with 29-bit identifiers and a transport protocol for messages longer than eight bytes. Newer platforms add CAN FD, Automotive Ethernet, and SOME/IP.
The diagnostic standards sit on top. OBD-II is the mandated external port, and it was mandated to be minimally useful rather than maximally revealing. It exposes a small fixed set of standardised parameters:
engine speed
vehicle speed
coolant temperature
engine load
fuel level
intake air temperature
throttle position
That is the list. And the parameters manufacturers most want to keep to themselves are powertrain calibration, individual cylinder behaviour, battery state of health, and anything about how the driver behaves.
UDS, ISO 14229, is how a diagnostic tester reads anything the manufacturer chooses to unlock, usually to a supplier or under a court order, and not on the public port.
So a platform that wants to offer a maintenance feature faces this: the data needed for it exists in the vehicle, is already on a bus, and the manufacturer is not publishing it. The realistic options are a manufacturer partnership, a data-sharing agreement, or building the feature from what OBD-II does expose, which is enough for some things and not for others.
Designing as though every field is available is the most common early failure, and it is discovered after the pilot when the interesting charts are empty.
The vehicle is offline most of the time
A car is a metal box in a tunnel. It loses signal in car parks, on rural roads, in underground structures, and behind buildings. A long journey crosses several dead zones.
The communication design has to accept this rather than work around it.
Store and forward on the vehicle. The telematics unit buffers locally and uploads when it has a connection. A buffer sized for a few hours of driving is cheap and removes an enormous class of problem.
Session resumption, not reconnection. MQTT’s persistent sessions and connection expiry interval are exactly designed for this. A vehicle that reconnects after three minutes should resume its subscriptions and receive its queued commands, not re-register and start over.
Reconnect backoff with jitter. Ten million vehicles in a city cell, all of them losing signal when the base station reboots, is a reconnect storm. Full jitter on backoff and a randomised boot delay are as necessary here as they were in the ingestion platform.
Idempotent uploads. A buffered batch uploaded after a half-acknowledged connection will arrive twice. Every upload has to be safe under repetition.
The pattern to avoid is treating the car as an HTTP client. Stateless request and response feels clean and it is wrong for a device that is offline for 12 percent of its life and cannot be asked to change its mind about it.
Commands are the hard direction
Telemetry flows from the car to the cloud and is comparatively forgiving. Commands flow the other way and can cause physical harm.
Three command classes, in increasing order of risk:
Informational. “Where is my car?”, “What is the state of charge?”, “Is it locked?” Read-only and safe to retry.
Convenience. Remote lock, remote unlock, remote start, climate control. A failed command is an inconvenience; a wrong command is a security or safety event.
Safety-relevant. Anything that starts the engine, releases a brake, or moves the vehicle. These need conditions checked on the vehicle at the moment of execution, not at the moment the command was sent.
The remote start case is the clearest illustration, and it is a documented real-world hazard: a remote start executed in a closed garage, hours after the request, with the engine running and nobody there. The mitigation is entirely on the vehicle side:
command received
-> check expiry timestamp reject if too old
-> check state of charge reject if insufficient
-> check local safety conditions reject if in an enclosed space
or if a fault is present
-> execute
-> report result
The server cannot check whether the car has since been driven into a garage. Only the vehicle can, and the vehicle is the only place where the check means anything.
Every command therefore needs:
- a correlation id, so the result can be matched to the request;
- an issued-at and an expiry, enforced on the vehicle;
- a precondition set, checked at execution time;
- a durable record, written before the command is sent, so “did this get delivered” is answerable months later.
And the interface should distinguish four states, not two:
acknowledged and executed
acknowledged and refused (with the reason: expired, precondition failed)
not acknowledged (unknown — the vehicle never confirmed)
expired before delivery (and therefore never executed)
Collapsing “not acknowledged” into “failed” is how a support conversation turns into an argument, because a failure has a cause and an unknown does not.
Firmware over a fifteen year fleet
The device management article covered staged rollouts and A/B boot partitions. A vehicle makes both mandatory rather than advisable.
The vehicle must be able to survive a bad image alone. No technician, no tow, no workshop. If a campaign bricks 200,000 cars, that is a recall, and the cost curve is completely different. Dual boot partitions with revert-on-failed-boot are the mechanism, and a vehicle that fails to confirm a successful boot within a window reverts by itself.
Campaigns are opportunistic. A vehicle will not sit still, on power, with signal, for fifty minutes during the day. The realistic design is that the vehicle downloads in the background whenever conditions allow, stages, and then installs at the next opportunity: parked, engine off or in a specific state, sufficient battery.
Partial progress is the normal case. A vehicle with 300 MB of a 3 GB image has to resume. Range requests over a moving connection, with a signature check over the whole image rather than the chunk, so a corrupt partial cannot be installed.
The fleet will never be homogeneous. A fifteen year fleet has vehicles from every model year, on hardware that was current then. Capability discovery is mandatory: a vehicle that cannot support a feature must be excluded from the campaign by capability, not by an assumption baked into a query.
staged 1% -> soak -> 5% -> soak -> 25% -> soak -> 100%
gates: boot success, crash rate, connectivity loss, battery delta
The gate is a comparison against the previous version, for the same reasons as everywhere else: an absolute threshold fails a fleet that was already unhealthy.
A twin of a car is a snapshot of a moving object
A digital twin is a natural fit for a vehicle and nearly useless as a live thing, because the thing being modelled is moving and the data is a few seconds to a few minutes old.
What makes vehicle twins genuinely useful is not the 3D model. It is the configuration and state history:
- what was fitted when (tyres, battery, sensors, firmware)
- what has been serviced and when
- what the vehicle has reported about itself over time
- which faults appeared, in which order, and what happened next
That last one is worth building deliberately. Fault codes do not arrive alone. A sequence of events before a fault is almost always more diagnostic than the fault code itself, which is why capturing CAN frames with their timestamps, not just the decoded final value, has real forensic value for warranty and accident reconstruction.
And the discipline from the industrial article applies exactly: the twin is a copy, timestamped, and never authoritative about the physical car. A model predicting a battery failure at 1,000 km is a prediction. A tow truck dispatched on it is a business decision with a fallback.
Predictive maintenance, honestly
The commercial appeal is obvious and the data reality is constrained.
What works well: state of charge and charge cycles, tyre pressure, mileage, service intervals, driving pattern, fault code sequences. These are available from OBD-II or from the vehicle’s own telematics unit, and they support useful models.
What does not: anything requiring powertrain internals, unless there is a manufacturer partnership. A knock detector, an individual cylinder’s behaviour, or a detailed combustion signature is behind the manufacturer’s data, not behind your API key.
The sample problem. A model predicting rare failures needs rare examples. A fleet of ten million vehicles might have a few hundred instances of a specific bearing failure mode per year, scattered across makes and models, and labelled inconsistently. Collecting them is a data engineering project with a physical logistics component, because the useful examples are the ones from vehicles that were subsequently repaired, and getting a diagnostic snapshot from a vehicle that has already left the workshop requires having asked for it in advance.
Models in this space are usually a rules engine over known failure signatures with a learned component on top. Calling that machine learning in a pitch is how the field ended up with a reputation.
Four data subjects, one car
This is the part that gets a connected vehicle platform into regulatory difficulty, and it is genuinely unusual.
A car’s data has at least four interested parties, and they are not the same person:
- The owner, who paid for the car and has a property interest in its data.
- The driver, who may or may not be the owner. A company fleet, a rental, a shared household car, a colleague’s commute.
- The account holder, who pays for the connected services. Often the manufacturer or the fleet operator, sometimes the owner, and under some subscription models a third party entirely.
- The other people in the car. Passengers, family, a taxi’s customers, and anyone whose location is incidentally recorded because the vehicle is moving.
A single consent checkbox cannot cover this, and a platform that offers one is making a legal decision by accident.
The design consequences:
- Purpose separation. Data collected to run the car, data collected to provide a service to the owner, and data collected for fleet analytics are different purposes with different consent requirements. Store the purpose with the data.
- Location history is the sensitive part. Where the car has been is a behavioural profile of every journey, and it identifies passengers as much as the driver.
- Erasure has genuine conflicts. A subject access request to delete data collides with accident reconstruction records, warranty evidence, and regulatory retention. This is a real collision, not an excuse, and the platform needs a documented policy rather than a hopeful one.
- Handover on sale. The car is sold. The owner’s account, the driver’s history, the subscription, and the accumulated location data all have to be dealt with. Most platforms handle this badly, and it is a predictable, high-volume support problem.
- Retention per data class. Operational telemetry, location history, and diagnostic snapshots should not share a retention period, and a car’s fifteen year life does not mean its location history should exist for fifteen years.
The cross-border dimension makes it worse rather than better. A car crosses borders continuously, and a single vehicle can generate data in dozens of jurisdictions in a week, with different consent and residency rules applying to each.
What this looks like end to end
+-------------------+
| Vehicle |
| telematics unit |
| GNSS, LTE/5G |
| local buffer |
| CAN gateway |
+---------+---------+
| MQTT over TLS, persistent session
| store-and-forward, jittered backoff
v
+-------------------+ +-------------------+
| Command service | | Ingestion |
| issue, expiry, | | validate, key by |
| preconditions, | | vin, fan out |
| durable record | +---------+---------+
+---------+---------+ |
| v
| +---------+---------+
| | Kafka |
| +---------+---------+
| |
| +------------------+------------------+
| | | |
| +-----v------+ +-------v-------+ +-------v--------+
| | Telemetry | | Position | | Diagnostic |
| | time series| | time series | | event store |
| +-----+------+ +---------------+ +----------------+
| |
| +-----v------------------------------+
| | OTA campaign service |
| | capability filter, staged rollout, |
| | signed manifests, A/B verify |
| +------------------------------------+
Control plane, kept deliberately boring:
vehicle registry -> VIN, model year, hardware capability, feature flags
command ledger -> every command, its state, its result, its operator
consent store -> per-purpose, per-data-subject, revocable
Failure stories worth testing
Send an unlock command to a car in a tunnel
It should queue within its expiry window, and if the car does not surface in time, it should expire unexecuted with a visible reason. A stale unlock is a security incident, not a convenience.
Send a remote start and then move the car into a closed garage
The vehicle should refuse on its own precondition check. The server cannot prevent this; only the vehicle can.
Lose cellular coverage for three hours mid-upload
The buffer should hold, the upload should resume, and nothing should arrive twice. Test the half-acknowledged case, not just the clean one.
Restart a base station serving a dense city
Ten thousand vehicles lose signal together. Confirm the reconnect is jittered and the device-facing tier does not fall over.
Ship a campaign with a fault in 1% of vehicles
Confirm the health gate stops the campaign, the vehicles revert by themselves, and the revert is visible in the fleet view.
Serve a fifteen year old vehicle the current API
This is the test that matters. A vehicle built in 2011 speaking the protocol it shipped with must still be able to register, report, and receive updates. If the current API only works with the current generation, the platform has a fifteen year problem it has not solved.
Sell a car with an active subscription
Confirm the account, the location history, the consent records, and the telemetry access are all handled as one workflow, with an audit trail.
Have a passenger exercise a subject access request
Confirm the platform can respond about location data that identifies a person who never interacted with it.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| Designing for the data you want | The powertrain fields are behind the manufacturer’s data, not the API | Scope the feature to what is actually exposed |
| Treating a car as an HTTP client | It is offline 12 percent of the time and cannot change that | Persistent session, store and forward, jittered backoff |
| Command expiry enforced on the server | The command executes hours later in a garage it was never meant for | Expiry and preconditions checked on the vehicle |
| Two command states, success and failure | “Never acknowledged” is treated as failed and has no cause | Four states, including unknown and expired |
| Single-sector OTA without A/B partitions | A bad image is a recall rather than a revert | Dual boot, revert on failed boot confirmation |
| Assuming the fleet will be homogeneous | Fifteen years of model years with different capabilities | Capability discovery and campaign filtering |
| One consent checkbox | The owner, driver, subscriber, and passengers are not the same person | Purpose-separated, per-subject consent |
| Uniform retention for all vehicle data | Location history kept for fifteen years by inertia | Retention per data class, decided explicitly |
| Ignoring the resale handover | High-volume support problem with a privacy dimension | One workflow for account, data, and access |
| Calling a rules engine machine learning | The field earned a reputation it did not deserve | Describe what the model actually does |
| A twin treated as live state | A minutes-old model drives a physical decision | Timestamp it, label it a copy, never act without a fallback |
| Backing up the protocol with the fleet | A fifteen year fleet meets a three year API | Version the wire protocol and never remove a version |
| No fault-sequence capture | Only the final fault code is available, which is the least useful part | Capture sequenced CAN events with timestamps |
| Relying on a single carrier | Coverage holes in exactly the places that matter | Multi-network fallback and local buffering |
The complete story in one minute
A vehicle reports telemetry over MQTT with a persistent session, buffering locally and uploading when it has a signal. A reconnect after three minutes resumes rather than restarting. A ten million vehicle fleet produces around 200,000 messages a second on average and a million during the morning commute, and the fleet is keyed by VIN so one vehicle’s history stays ordered.
The data available is what the manufacturer exposes, which is a modest fixed set of diagnostic parameters plus whatever a data-sharing agreement provides. Designing as though the powertrain internals are available is the most common early mistake, and it is discovered after the pilot.
Commands travel the other way with a correlation id, an expiry, and a precondition set, all checked on the vehicle at the moment of execution rather than on the server when it was sent. That is what prevents a three-hour-old remote start from running an engine in a closed garage, and it is the only place the check means anything.
Firmware updates are staged with health gates and dual boot partitions, because a vehicle that bricks itself is a recall rather than a rollback. The fleet will never be on one version, so campaigns filter by hardware capability rather than by assumption.
And a connected car has four data subjects rather than one, which is why consent is per-purpose and per-person, retention is per data class, and the resale of a vehicle is treated as a first-class workflow rather than a support ticket.
That is the whole path:
vehicle -> MQTT -> ingest -> telemetry | position | diagnostics
|
+-> commands with expiry and vehicle-side preconditions
+-> staged OTA with A/B rollback
+-> purpose-separated consent and per-class retention
The hard parts were never the message rate. They were a client that must still work in fifteen years, commands that can start an engine, and the fact that the people in the car are not always the person who consented.
What this team still owns
The cloud owns command admission, identity, audit, expiry, and campaign targeting; the vehicle owns the final safety precondition and must remain safe when the cloud is wrong or absent. MQTT session state can resume delivery, but it does not make a remote command current, authorised, or safe. Every command therefore carries a unique ID, an expiry, the intended vehicle capability/version, and a terminal outcome that distinguishes rejected, expired, executed, and unknown.


