← All writing
articleApr 30, 202518 min read

Fleet Management Systems: The Platform Is Telematics, the Product Is Exceptions

Hours-of-service compliance, driver behaviour scoring as an employment decision, fuel leakage, and why a fleet system's real value sits in the five percent of cases that are not normal.

IoTTelematicsDataArchitecture
Fleet Management Systems: The Platform Is Telematics, the Product Is Exceptions cover illustration

Everything about fleet management that is technically interesting is telematics, and the telematics is the easy half. Position, fuel level, engine hours, fault codes, and driver behaviour signals all arrive through a pipeline that the GPS tracking article already covered.

The half that is hard is that a fleet is an operating business, and operating businesses are made of exceptions, paperwork, and people who need to be told something before it becomes a problem.

This article is about the second half.

The shape of a customer

A fleet management platform is multi-tenant SaaS, and its tenants are small.

2,000 fleet customers
average 150 vehicles each
                              = 300,000 vehicles
average 1.3 drivers per vehicle
                              = ~400,000 drivers

Compare that to the industrial platform in this set: three hundred thousand vehicles across two thousand customers is a mid-sized platform, and the ratio of customers to vehicles is what shapes the design.

A customer with 150 vehicles has a fleet manager who uses it daily, an owner who looks at it weekly, and drivers who interact with it once a month. The user base is small, the data volume is modest, and the number of distinct workflows is large. This is a business application with a data feed attached, not the other way round.

Compliance is a specification

Hours of service is the clearest example of a feature that is really a specification, and the distinction matters more than it sounds.

United States, FMCSA property-carrying drivers, roughly:

- 11 hours maximum driving after 10 consecutive hours off duty
- a 14-hour on-duty window, of which 11 may be driving
- a 30-minute break required after 8 cumulative hours of driving
- 60 hours in 7 consecutive days, or 70 hours in 8
- a 34-hour restart that resets the weekly count

European Union, tachograph, roughly:

- 4.5 hours maximum continuous driving, then a 45-minute break
- 9 hours maximum daily driving, extendable to 10 hours, at most twice a week
- 11 consecutive hours of daily rest, splittable into 9 plus 2
- 45 consecutive hours of weekly rest, extendable to 60

The exact numbers vary by jurisdiction, by vehicle class, and by date, and the short version is that a rules change is a product change with a compliance deadline attached.

That is the design requirement:

ruleset:  US-FMCSA-PROPERTY
version:  2024-11-01
jurisdiction: US
applies_to: [class: property, gvwr_over: 12000]

Not:

if (driving_hours > 11 && rest_hours < 10) { ... }

The second version is code, and code does not have an effective date, an audit trail, or a way to answer “which rule was in force at 03:00 on Tuesday”. When a driver is cited, someone will ask exactly that question, and the answer needs to be a lookup rather than a git archaeology exercise.

Store the ruleset with an effective date, evaluate the version in force at the time of the driving, and record which ruleset version produced every conclusion. A violation that is recorded against the wrong ruleset is a product defect with legal consequences.

The ELD and the audit trail

The electronic logging device is a regulatory device with a specific job: produce a defensible record of driving time.

on duty        driving
|--------|------|     |---------|------|     |------|--------|
        14-hour on-duty window
              30-minute break

The requirements that shape the software:

  • The record is append-only. A duty status change is a fact about time and cannot be edited after the fact. Corrections are new records that reference the original.
  • Edits carry a reason. A missed break or a wrong status needs a documented reason and an identity, and the reason is part of the record.
  • The sequence must be complete. A gap in the duty status log is a violation in itself, which means the UI has to make partial periods visible rather than letting a user file a clean-looking day that is actually missing three hours.
  • Certification is periodic. Drivers certify their logs and the certification is a distinct, timestamped act.
  • Export must be reproducible. An enforcement request asks for a specific driver’s log over a specific window in a specific format, and it has to be produced from the underlying records rather than from a report someone generated earlier.

A fleet platform that treats the ELD record as a display concern rather than a compliance record will eventually fail an audit, and the failure will be expensive.

Driver scoring is an employment decision

This is where a lot of fleet platforms get it wrong, and the mistake is treating behaviour scoring as a data product rather than as a system that makes decisions about people’s employment and pay.

The signals are straightforward to collect:

harsh braking        deceleration beyond a threshold
harsh acceleration   forward acceleration beyond a threshold
speeding             above the posted limit, with tolerance
cornering           lateral g beyond a threshold
idling               engine on, vehicle stationary
distraction          driver-facing camera or sensor inference

The engineering problem is that none of these are objective. A threshold of minus 0.4 g for harsh braking will flag a driver on a wet roundabout and miss a genuinely dangerous event on a dry straight. GPS-derived speed is wrong at urban canyons and in tunnels. The thresholds are a product decision with a real-world consequence, and they are almost always set by someone who does not drive the vehicles.

Then there is the consequence problem. A score that affects a driver’s pay, a bonus, or their employment is:

  • legally constrained in many jurisdictions, and increasingly so;
  • contestable, which means the driver has to be able to see the event, the data behind it, and how to dispute it;
  • anonymised in aggregate for the fleet-level view, because a leaderboard of named drivers is how a data tool becomes a culture problem.

A defensible design treats the score as a proposal, not a verdict:

event     -> the specific event, with location, time, and the raw signal
context   -> road, weather, vehicle, and driver history
rule      -> which threshold fired, and its version
decision  -> who decided, when, and what was decided
appeal    -> the driver's response, and the outcome

And the product requirement that goes with it: if the platform can affect someone’s pay, it can afford to explain itself. That means a driver-facing view, not just an admin dashboard, showing the same events and the same reasoning the manager sees. The asymmetry where a manager sees a score and a driver sees nothing is the thing that turns a data product into a labour dispute.

Fuel is where the money is

Fuel is typically the largest controllable cost in a fleet, and it is already fully instrumented by the vehicle.

For 300,000 trucks averaging 40 km a day:

12 million km/day
at ~3.5 km per litre                  = 3.4 million litres/day
at $1.30 per litre                    = ~$4.5 million/day
                                        ~$1.6 billion/year

Two leaks are visible without any new hardware.

Idling. A truck idles at roughly three to four litres an hour. If the average vehicle idles an hour a day:

300,000 x 1 hour x 3.5 litres        = 1.05 million litres/day
                                        ~$500 million/year

Cutting idle by a third is worth something north of a hundred and fifty million dollars a year across the platform, and it requires nothing more than a fuel-level-over-time signal and a driver who will be told why their week looks like that.

Unauthorised purchase. Fuel is bought before it is consumed, which creates an obvious opportunity that is visible as a fuel purchase with no corresponding consumption, or consumption with no corresponding purchase. Card-level data makes this straightforward.

Both of these are good features because they are objective, financially legible, and not contentious — which is a useful contrast with driver scoring, where the same data is contested.

The failure mode worth naming is a fuel number nobody trusts. Tank size, sensor calibration, refuelling patterns, and driver behaviour all affect the calculation, and if the platform shows a driver an efficiency figure that is wrong by fifteen percent, the whole module is dismissed as a gimmick. Publish the confidence interval, or do not publish the number.

Maintenance: three triggers

Maintenance scheduling comes in three flavours and the useful system combines them.

Calendar-based. Service every 10,000 km or every six months. Simple, predictable, and always slightly wrong.

Distance and engine hours based. Driven from odometer and engine hours, which is more accurate for mixed duty cycles where distance and wear diverge.

Condition-based. Driven by fault codes, sensor thresholds, and degradation signals. The genuinely valuable one and the one with the least data, because most fleets do not have the history to train on and most parts fail without a fault code first.

next service = max( distance trigger, calendar trigger, condition trigger )

A condition signal should be able to bring a vehicle forward, and the interesting cases are all overrides: a driver reports a noise, a diagnostic tech reads codes on a drop-off, a sensor fires outside its learned range. That last one is worth having as a first-class input, because the most expensive maintenance failures are usually the ones nobody was looking for.

The operational point is that the maintenance module’s output is a work order and a parts reservation, not an alert. A system that emails a fleet manager saying a vehicle is due for service has not finished the job. A system that holds the booking, checks parts availability, and tells the driver is the system that changes behaviour.

Dispatch is an optimisation with a person in it

Route planning for a fleet is a genuinely hard combinatorial problem: capacity, time windows, service durations, driver hours, vehicle types, tolls, and the fact that a real depot has specific vehicles with specific constraints.

The temptation is to present it as an optimisation to be solved automatically. The reality is that dispatchers use the output as a draft and override it constantly, because the model does not know that the customer at that address always takes an hour, or that the delivery window is flexible on Fridays, or that the driver prefers that route.

The design that works treats the solver as a proposal:

1. dispatcher selects the work to plan
2. solver returns an assignment with the objective value
3. dispatcher adjusts, pins, or excludes
4. the difference is stored and fed back

That last step is the valuable one. The overrides are labelled data about constraints the model does not know, and over time they are the difference between a system that gets better and one that gets ignored.

The other half is the live case: a vehicle breaks down, a road closes, a customer cancels, and the plan has to change with people in the field acting on a phone. That workflow is not the same product as the nightly optimisation and it should not be a degraded version of it.

Documents, and the expiry problem

A real fleet system manages paperwork: registrations, insurance, licences, driver qualifications, vehicle inspections, permits, and the certificates each one depends on.

The value is almost entirely in expiry alerts with a real workflow behind them. A document expiring is not information, it is a deadline, and the difference between a product that says “insurance expires in 12 days” and one that says “insurance expires in 12 days, here is the renewal, here is who owns it, here is the last three renewals, and here is the button” is the difference between a dashboard and a product.

The two things that make document systems fail:

  • Owner ambiguity. Every document needs exactly one accountable owner with a notification, and a shared inbox is not an owner.
  • Expiry without renewal. An alert that fires and is ignored teaches people to ignore alerts. Track acknowledgement and make the ignored case escalate.

Exceptions are the product

A fleet is a normal case and an exception case, and the normal case is easy.

Normal: a vehicle runs a route, arrives within its window, logs its hours, gets serviced on schedule. A dashboard shows this and it is fine.

The product is the rest:

- a driver hits their HOS limit 200 km from anywhere
- a vehicle faults in a tunnel with no connectivity
- a customer refuses a delivery and the driver needs a signature decision
- a fuel card is declined at 23:00
- a licence lapses and the driver is already on the road
- a temperature-controlled load breaches its range in transit
- a vehicle is stolen and recovery is needed in the next hour
- a driver disputes a harsh-braking event
- a customer wants proof of delivery for a load delivered three weeks ago

Each of these is a workflow: a trigger, a decision, an action, an audit record, and a resolution. That is the product. Everything else is the data feed that makes the workflow possible.

The design consequence is that exception handling deserves the same care as the happy path, and in most fleet products it does not get it, because exception workflows are irregular and hard to demo.

Multi-tenancy, and the data boundary that matters

Two thousand customers in one platform means the tenancy boundary is a first-class security concern rather than a schema detail.

The things that go wrong:

  • A missing tenant predicate in one report. One query without the filter and one customer sees another’s fleet. This is the single most common multi-tenancy breach and the cheapest to prevent with a repository layer that cannot express an unscoped query.
  • A shared lookup table. Device type catalogues, integration credentials, and notification templates that are accidentally global.
  • Per-customer configuration drift. If a customer’s integration behaves differently and nobody can say why, the answer is usually a configuration value nobody documented.

The useful discipline is that every table is either tenant-scoped or explicitly global, there is no third category, and the global tables are a short, reviewed list.

Integrations

A fleet platform is rarely the system of record for anything. It integrates with:

accounting        fuel and maintenance cost posting
payroll           driver pay, especially where behaviour affects it
telematics/ELD    the device feeds
CRM and orders    the work to be done
telephony         driver communication
maintenance       workshop systems, parts

Each integration is a place where a partial failure looks like success, which is where the operational pain lives. The specific failure to design for is the idempotent replay: an accounting export posts fuel costs, and re-running it after a partial failure must not double-post. That means idempotency keys agreed with the receiving system, not a flag in the source.

The other one is schema drift in the other direction. An accounting system that renames a field will break the export at 3am on a Sunday, and the detection has to be an alert rather than a stack trace nobody reads.

Failure stories worth testing

A driver’s HOS is calculated with yesterday’s ruleset

Confirm the version in force at the time of driving is used, and that the audit shows which one.

A duty log has a three-hour gap

The record must show the gap. Confirm the UI makes a partial period visible rather than letting someone file a clean day.

A driver disputes a harsh-braking event

Confirm the driver sees the event, the raw signal, the rule version, and has a route to appeal, and that the appeal is recorded.

A fuel purchase posts twice after a retry

Confirm the export is idempotent and that a duplicate cannot reach the ledger.

A vehicle faults with no connectivity for four hours

The fault must be on record when it reconnects, with timestamps that show when it actually happened rather than when the vehicle got signal.

A customer’s insurance expires and nobody acts

Confirm it escalates rather than sitting as an ignored notification.

The overnight optimisation produces a plan the dispatcher rejects entirely

Confirm overrides are captured and that the dispatcher can pin work the solver will not move.

One tenant’s data appears in another’s report

Test the tenancy boundary with a deliberate missing predicate, and make the failure impossible rather than merely unlikely.

A licence lapses while the driver is mid-route

Confirm the platform can notify the dispatcher, the driver, and the compliance owner, and record who was told and when.

A production-ready architecture

  [ telematics devices, ELDs ]
            |  MQTT / vendor APIs
            v
  +---------------------------+
  | Telematics ingestion      |   (the GPS tracking pipeline)
  +-------------+-------------+
                |  position, fuel, faults, behaviour signals
                v
  +---------------------------+   +---------------------------+
  | Time series (state)       |   | Event store (facts)       |
  +-------------+-------------+   +-------------+-------------+
                |                             |
                +--------------+--------------+
                               v
                  +------------+------------+
                  |  Operational core        |
                  |  fleet, vehicle, driver, |
                  |  job, document, HOS      |
                  +-----+------------+------+
                        |            |
          +-------------+            +--------------+
          |                            |              |
  +-------v--------+        +----------v-----+  +-----v------+
  | Compliance     |        | Dispatch       |  | Documents  |
  | versioned      |        | solver +       |  | expiry     |
  | rulesets       |        | overrides      |  | workflow   |
  +----------------+        +----------------+  +------------+
          |
  +-------v--------+
  | Scoring        |
  | events, rules, |
  | appeals        |
  +----------------+

  Every table tenant-scoped, or explicitly and briefly global.
  Every HOS conclusion stores the ruleset version that produced it.

A sensible delivery checklist:

  1. Model hours-of-service rules as versioned, dated rulesets, never as code.
  2. Make the ELD record append-only, with reasoned corrections and reproducible exports.
  3. Show drivers the same behaviour events and rules the manager sees.
  4. Build an appeal path into the scoring module, not as a support process.
  5. Publish the confidence of a fuel or efficiency figure, or do not publish it.
  6. Treat maintenance output as a booked work order, not an alert.
  7. Capture dispatcher overrides as labelled constraint data.
  8. Build the exception workflows to the same standard as the normal path, because they are the product.
  9. Make every table tenant-scoped or explicitly global, and keep the global list short.
  10. Make every outbound export idempotent with a key agreed by the receiving system.
  11. Give every document exactly one accountable owner and an escalation path.
  12. Give the fleet a printable, locally-served runbook that works without the platform.

Common mistakes

Mistake What actually happens Better decision
HOS rules hard-coded A rules change is a code change with no effective date and no audit Versioned, dated rulesets per jurisdiction
The current ruleset applied to past driving A citation is computed against the wrong law Evaluate the version in force at the time of driving
An editable duty log The record stops being defensible Append-only, with reasoned corrections
A behaviour score with no driver view The driver sees nothing and the tool becomes a labour dispute Show drivers the events, the rules, and an appeal path
One fixed harsh-braking threshold Flags wet roundabouts, misses real events Context, calibration, and an explainable rule version
A named driver leaderboard Culture damage and no operational improvement Aggregate for fleets, specific for coaching
Fuel efficiency shown to one decimal place A wrong number discredits the whole module Publish the confidence or publish nothing
Maintenance alerts by email A manager now owns a notification nobody acts on Booked work orders with parts reserved
Automatic dispatch with no override Dispatchers ignore the plan entirely Solver as proposal, with overrides captured as data
Happy-path workflows only The product is useless during the actual problem Exception workflows built to the same standard
A shared global configuration table One customer’s change silently affects another Tenant-scoped, or explicitly and briefly global
A nightly export with no idempotency key A retry double-posts to the ledger Idempotent, with a key agreed by the receiver
Document alerts with a shared inbox Nothing is owned and alerts are ignored One accountable owner, with escalation
Alerting on idling without context Legitimate waiting time is treated as waste Separate mandated waiting from unnecessary idling
No offline runbook The platform is useless exactly when it is needed A printable, locally-served version of the procedure
Demos of the dashboard as the product Buyers see charts and miss the workflows Demo an exception, live, end to end

The complete story in one minute

A fleet management platform ingests position, fuel, fault, and behaviour signals through the pipeline from the GPS tracking article, and then does the harder half: compliance, dispatch, maintenance, documents, and people.

Hours of service is evaluated against versioned, dated rulesets per jurisdiction, using the version in force at the time of the driving, and every conclusion records which ruleset produced it. The electronic log is append-only, with reasoned corrections, visible gaps, and reproducible exports, because it is a regulatory record rather than a report.

Driver behaviour is scored against versioned, explained rules, and because that score can affect someone’s pay it is treated as an employment decision. Drivers see the same events and rules the manager sees, and there is an appeal path with a recorded outcome.

Fuel, the largest controllable cost, is already fully instrumented. Idling and unauthorised purchases are objective, financially legible, and uncontroversial, and they are worth more than the entire scoring module. Efficiency figures are published with their confidence or not at all.

Maintenance produces booked work orders with parts reserved. Dispatch treats the solver as a proposal and captures the dispatcher’s overrides as labelled constraint data, because the overrides are what makes the model better over time.

And the product is the exception workflows, built to the same standard as the happy path, because the dispatcher handling a road closure at six in the morning is the person actually using the software.

That is the whole path:

vehicles -> telematics -> operational core -> (compliance | dispatch | maintenance | exceptions)
                                        |
                                        +-> every HOS conclusion records its ruleset version
                                        +-> every driver-facing decision is explainable and appealable

The hard parts were never the position stream. They were treating the compliance rules as a specification with a version and a date, and recognising that a number which changes someone’s pay is not a data product.

What this team still owns

Compliance and payroll outputs need reproducible evidence, not only a current aggregate. Retain the source event, correction history, actor, device identity, rule-set version, jurisdiction, and effective time that produced each duty-status or pay decision. Telematics can propose classifications; it cannot silently rewrite a driver’s legal record. Corrections are append-only and reviewable, and exports must reproduce the view that was valid for the requested period.

Technical references

Keep reading
Browse everything