Logistics Package Tracking: The Event Stream Behind a Progress Bar
Package tracking as a scan-event pipeline — status models, exception handling, the ETA problem, and why carrier webhooks are the hard part.

A tracking page looks trivial. Behind it is an event pipeline that ingests scan data from dozens of carriers in inconsistent formats, maintains a state machine per parcel, predicts a delivery window, and survives the fact that the carriers themselves are the least reliable part of the system.
This is an article about that pipeline, and about the specific failure modes that appear once a parcel crosses three or four different organisations’ hands.
The scale
a national carrier network
40M parcels per year
~110,000 per day
peak: 400,000 (December)
scans per parcel: 8-15
total scan events: ~1.5M per day
peak: 5M per day
event volume is not the hard part
the hard parts:
- 14 carrier formats, none documented
- scan events arrive out of order
(a hub scan arriving after a
delivery attempt)
- the same scan arrives 3 times
- ~2% of parcels go silent for
5+ days with no event at all
- "out for delivery" does not mean
it is on the van
The volume is unremarkable. Everything that makes this hard is about correctness under partial, late, and contradictory information.
The state machine
CREATED
-> INFORMATION_RECEIVED
-> PICKED_UP (origin scan)
-> IN_TRANSIT (line-haul hub scans)
-> AT_DESTINATION_HUB
-> OUT_FOR_DELIVERY (loaded on van)
-> DELIVERED (proof: signature/photo)
-> DELIVERY_ATTEMPTED
-> RETURNED_TO_SENDER
-> EXCEPTION: LOST / DAMAGED / HELD / RETURNED
-> CANCELLED
illegal transitions that arrive in real life
- DELIVERED -> IN_TRANSIT
(a second carrier's late scan)
- IN_TRANSIT -> INFORMATION_RECEIVED
(out-of-order replay)
- DELIVERED -> OUT_FOR_DELIVERY
(mis-scan at a hub)
- anything -> anything, from a manual
carrier override in their own UI
The most common integration bug is not a missing field. It is an out-of-order or duplicate event producing an impossible status that then gets shown to a customer. So the pipeline has to be defensive in a specific way:
on each event
1. dedupe by (parcel, event_id)
2. if event carries a scan timestamp
older than current state's timestamp
-> log it, do not apply it
(or route to a "late event" review)
3. validate the transition against the
state machine
4. if illegal
-> record it in an exceptions table
-> decide: ignore, or correct the
state with a repair rule
5. if legal -> apply, emit a new state
event, update projections
The rule about event timestamps rather than arrival order is the important one. Scan events carry their own timestamps from the carrier’s handheld or hub system, and those timestamps can be hours apart from arrival. Applying by arrival order means a stale event overwrites a fresher state; applying by scan timestamp means the most recent scan wins, which is what the customer expects to see.
Status normalisation
carrier A: "DPU" "OTD" "ARR" "IT" "DEL"
carrier B: "PICKUP_COMPLETE" "OUT_FOR_DELIVERY"
"ARRIVED_AT_FACILITY" "IN_TRANSIT"
"DELIVERED"
carrier C: numeric codes, a different
code for "attempted, customer
absent" vs "attempted, address
incomplete"
carrier D: free text, occasionally
translated
the mapping table is the product
200+ source codes
-> 12 canonical statuses
-> 40+ customer-facing strings
(different wording per market)
Every carrier integration is a mapping problem and the mapping is versioned, because carriers change their codes without notice. The robust pattern is a mapping layer with a default bucket plus an explicit unknown-value path:
unknown code
-> canonical status UNKNOWN_EVENT
-> logged with the raw value
-> surfaced in a monitoring dashboard
-> never silently mapped to a
customer-facing status
because the alternative is a parcel
that shows "in transit" forever
because of a code you never mapped
That monitoring dashboard is not optional. A carrier that adds a code on a Tuesday and ships a million parcels through it means your customers are seeing stale or wrong statuses for a day, and the only way to know is to watch the unknown-rate.
The ETA problem
a parcel with a good scan history:
created 09:00
picked up 14:00
in transit 16:00
at destination hub 04:00 (+1d)
out for delivery 08:00 (+2d)
ETA for delivery: ?
naive: hub scan + historical average
-> 14:00 (wrong, 40% of the time)
good: a model conditioned on
- origin/destination pair
- service level
- day of week
- time of day of the last scan
- destination's local cut-off time
- current weather (a real factor
for road networks)
-> distribution over delivery windows
The distribution is the product. Two design commitments follow:
Predict a window, not a timestamp. “Between 1pm and 5pm” is honest and is what the customer actually uses. A precise time that is wrong 40% of the time trains people to distrust the number entirely, and once they distrust it they stop using the number at all.
Recalculate on every state change. The ETA is a function of the current state, so it changes as the parcel moves. A parcel at the destination hub at 04:00 has a different and much tighter ETA than the same parcel that was in transit yesterday. Recompute it on the event, not on a nightly batch, because the customer is watching the page while the parcel moves.
The specific failure that annoys customers most is the ETA that gets worse as the parcel gets closer. That happens when the model is a simple average of historical transit time and does not condition on the current position. Conditioning on position — a parcel already at the destination hub is categorically different from one in line-haul — is what makes the ETA feel intelligent.
Exceptions are statuses, not errors
exception types
STUCK - no scan in N hours,
location unchanged
LOST - no scan in M days,
trajectory impossible
DELAYED - missed a transit scan
deadline
DAMAGED - carrier-reported
RETURNED - sender/address issue
HELD - customs, weather, address
verification
MISSORT - destination hub could
not process it
ADDRESS_ISSUE - incomplete or
contradictory address
each one needs
- a customer-facing message that is
actually useful ("we've contacted
the recipient", "we need more
information from the sender")
- an owner (who is working on it)
- a resolution path (retry, reroute,
return, contact)
- an SLA
The design principle is that an exception is a state the parcel is in, not a failure of the tracking system. Modelling it as a nullable error means the customer page shows the last known status — “in transit” for a parcel that has been sitting in a sorting office for a week. The exception state is the single most valuable thing on the tracking page when something is genuinely wrong, and it is the thing most implementations get wrong by treating it as an error code.
The stuck/lost detection is worth doing carefully because it is the case that actually matters:
stuck
no scan in 6 hours AND
last known location has not changed
AND
not in a known quiet period
(a hub that legitimately does
not scan overnight)
lost
no scan in 3 days AND
the last location is inconsistent
with the destination
(e.g. still at the origin country
when the destination is domestic)
AND
the carrier has not acknowledged it
The “quiet period” exemption is the detail that makes this usable. A naive stuck detector fires every night for every parcel in a hub that does not scan overnight, and once it fires habitually everyone ignores it. The detector has to know which facilities have quiet periods and which do not.
Carrier webhooks: assume the worst
what a carrier integration actually does
- sometimes delivers via webhook
- sometimes via polling
- sometimes via SFTP file drop
- sometimes via an email that
someone parses (yes)
- retries with duplicates
- delivers out of order
- delivers late (hours)
- delivers gaps (missing scans)
- changes the code table without notice
- sometimes goes down for a day
- has no rate limit, then rate-limits
you hard
design for all of it
- idempotent consumers (event_id dedupe)
- a per-carrier raw event store, kept
for debugging and replay
- polling as a backstop for webhook
carriers (detect gaps, backfill)
- a dead-letter queue with alerting
- per-carrier lag monitoring
The raw event store is the component that people skip and then wish they had. When a customer reports that a parcel showed the wrong status for six hours, the only way to answer is to look at what the carrier actually sent — and if you normalised and discarded on ingest, that information is gone. Keeping the raw payload per carrier is cheap relative to the debugging cost of not having it.
The polling backstop matters more than it looks: webhooks are the fast path, but the gap detection (“we have not received a scan for this parcel in 4 hours, the carrier says it should have moved”) is what catches a webhook that silently stopped working. Without polling, a broken integration looks exactly like a carrier having a bad day.
Retention, and why it is a product decision
tracking data ages differently
- standard parcels: useful for ~90 days
- high-value / signed: longer, often
1-2 years for claims and disputes
- international: customs records have
statutory retention requirements
- failed/lost: needs to be retrievable
as long as a claim is possible
- analytics aggregates: keep forever
(anonymised)
the trap
deleting the raw event log too early
-> a claim 6 months later has no
evidence trail
-> refunds get issued on the basis
of a customer's word
the right shape
- hot: current state + recent events
(fast, queryable)
- warm: full event history per parcel
(slower, still retrievable)
- cold: anonymised aggregates
(analytics)
with a per-class retention policy
A tracking system that deletes its history to save storage has quietly turned every dispute into an argument, and the argument is always resolved in favour of the customer, because that is the commercially correct decision when you have no evidence. Retention tiers cost very little and are the difference between a defensible claims process and an expensive one.
Failure stories worth testing
Deliver the “out for delivery” scan twice
The status must not change and the event must not double-apply. This is the idempotency test.
Send a hub scan with a timestamp earlier than the current state’s scan
The state must not regress. This is the out-of-order test and it is the most common real-world failure.
Send a “delivered” event followed by an “in transit” event
The system must not show a delivered parcel in transit. Whatever repair rule you use, it must be explicit.
Remove a carrier’s code mapping and replay a day of events
Confirm the unknown-rate alert fires and the raw payload is available. This is the “carrier changed their codes” test.
Deliver a parcel to a hub at 04:00 and check the ETA trajectory
The ETA must tighten as the parcel approaches. If it does not, the model is not conditioning on position.
Set the ETA to a single precise time and measure accuracy
Measure the fraction within 1 hour, within 2 hours, within 4 hours. Then compare against a window model with the same underlying predictions — the window is more useful at every horizon.
Go silent on a parcel for 5 days
Confirm the stuck/lost detector fires, that it respects quiet-period facilities, and that the customer page shows an exception with a next action.
Break the webhook endpoint and confirm polling backfills the gap
Without polling, a broken integration is invisible until someone complains.
Flood the webhook with duplicates at 10x rate
Confirm idempotency holds and the lag monitor catches the rate.
Change a carrier’s status wording mid-flight
The canonical mapping must hold; only the customer-facing string should change. Version the string tables separately from the code tables.
Route a parcel through an international leg with customs holds
Confirm the exception state explains the delay in terms the customer can act on, and that the sender — not the customer — is the party being asked for information.
Delete a parcel’s history and then try to resolve a claim
This is the retention test, and it demonstrates exactly why the warm tier exists.
A production-ready architecture
carrier sources
- webhooks (many, unreliable)
- polling backstops
- file drops, legacy integrations
|
v
+----------------------------------------------------------+
| INGEST |
| - raw event store per carrier (immutable) |
| - dedupe by (carrier, event_id) |
| - normalise: carrier code -> canonical status |
| - unknown codes -> UNKNOWN_EVENT + alerting |
+----------------------------+-----------------------------+
|
v
+----------------------------------------------------------+
| STATE MACHINE per parcel |
| - apply by scan timestamp, not arrival order |
| - validate transitions; illegal -> repair rule + log |
| - exceptions as first-class states |
+----------------------------+-----------------------------+
|
state change
|
+----------------------+----------------------+
v v v
+-------------+ +----------------+ +------------------+
| ETA model | | projections | | exception engine |
| conditioned | | (current view, | | stuck / lost / |
| on position | | index, search)| | delayed / held |
+-------------++ +----------------+ +------------------+
| | |
+----------------------+----------------------+
v
+----------------------------------------------------------+
| TRACKING API + PAGE |
| - status history, ETA WINDOW, exceptions with next step |
| - language/locale per market |
+----------------------------+-----------------------------+
|
v
+----------------------------------------------------------+
| RETENTION TIERS |
| hot: current state + recent events |
| warm: full per-parcel event history (claims evidence) |
| cold: anonymised aggregates |
+----------------------------------------------------------+
watch: unknown-code rate, event lag by carrier, webhook uptime,
stuck-detection precision, ETA accuracy by horizon
Delivery checklist:
- Make the tracking flow a state machine with explicit legal transitions, and treat every other transition as something to log and repair deliberately.
- Apply events by scan timestamp, not arrival order. This is the single most important correctness rule.
- Dedupe on a stable event id, and keep the raw carrier payload in a separate store so debugging is possible.
- Version the carrier code mapping, monitor the unknown-code rate, and never silently map unknown codes to a customer-facing status.
- Predict a window, not a timestamp, and recalculate on every state change. Conditioning on the parcel’s current position is what makes the ETA feel right.
- Make exceptions first-class states with a customer-facing message, an owner, a resolution path, and an SLA. The stuck/lost detector needs to know which facilities have quiet periods.
- Build polling backstops for webhook carriers so a broken integration is detectable before a customer reports it.
- Monitor per-carrier event lag and webhook uptime as first-class metrics.
- Use retention tiers, with a longer window for high-value, international, and exception parcels, so claims have an evidence trail.
- Keep the aggregates anonymised and forever, and keep them separate from the per-parcel history.
- Separate the canonical status codes from the customer-facing strings so wording can change per market without touching logic.
- Give every exception a “what happens next” line. A status that tells a customer nothing is worse than no status at all.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| Applying events by arrival order | Stale scans overwrite fresh state | Apply by scan timestamp |
| No event dedupe | Duplicate scans cause phantom transitions | Dedupe on (carrier, event_id) |
| Illegal transitions applied | Customers see delivered-then-in-transit | Validate, log, repair deliberately |
| Normalise and discard raw events | No way to debug a bad status | Keep a per-carrier raw event store |
| Unmapped carrier codes | Parcels stuck “in transit” forever | Unknown bucket + alert on the rate |
| A single ETA timestamp | False precision, distrusted numbers | A window, recalculated per event |
| ETA not conditioned on position | ETA gets worse as the parcel nears | Model position explicitly |
| Exceptions as error/nulls | Page shows the last known status forever | Exceptions as first-class states |
| Stuck detector with no quiet periods | Fires nightly, gets ignored | Per-facility quiet periods |
| Webhook-only integration | A broken webhook is invisible | Polling backstop and lag monitor |
| Single flat retention | Claims have no evidence | Retention tiers by parcel class |
| Canonical status tied to wording | Market wording changes break logic | Separate code and string tables |
| No “what happens next” on exceptions | A status that informs nobody | Owner, action, and SLA per exception |
| Aggregates mixed with per-parcel data | Deleting for privacy loses claims data | Anonymised cold tier, kept separate |
The complete story in one minute
A tracking page is an event pipeline in disguise, and the volume is not the hard part — 1.5 million scan events a day is unremarkable. The hard part is that a parcel crosses three or four organisations and the scan events arrive late, duplicated, and out of order, from fourteen carriers with 200+ undocumented status codes between them. So the core is a state machine per parcel, applied by scan timestamp rather than arrival order, deduplicated on a stable event id, with every illegal transition logged and repaired by an explicit rule rather than applied. And the single most valuable debugging asset is the per-carrier raw event store: if you normalise and discard on ingest, you will never be able to answer “what did the carrier actually send” when a customer says the status was wrong for six hours.
The ETA is a prediction whose distribution is the product. A window of 1pm to 5pm that hits 75% of the time beats a precise 4:07pm that is wrong 40% of the time, because a false-precision number trains people to stop trusting the whole feature. Recalculate on every state change and condition on the parcel’s current position — the worst customer experience is an ETA that gets worse as the parcel gets closer, which happens the moment the model is a flat historical average.
Exceptions are statuses, not errors. A package stuck for a week should show “we’ve contacted the recipient”, not the last known status of “in transit”. Stuck and lost detection has to know which facilities have quiet periods, or it fires every night for every parcel and gets ignored — which is worse than not having it. Retention is a product decision with a legal dimension: per-class tiers where high-value, international, and exception parcels keep their history long enough to settle a claim, and aggregates kept anonymised and separate, because a system that deletes its history settles every dispute in favour of the customer.
state machine, applied by scan timestamp, deduped, illegal transitions logged
per-carrier raw events kept; unknown-code rate monitored
ETA as a window, recalculated, conditioned on position
exceptions as first-class states with an owner and a next action
polling backstop on every webhook; retention tiers by parcel class
The hard part was never displaying the status. It was deciding, when the data disagrees with itself, which of the fourteen sources to believe — and admitting that a package sitting in a sorting office for a week is information the customer is owed.
What this team still owns
Keep observations immutable and derive customer status from them. A scan has source, facility, device, event time, receipt time, carrier code, and confidence; a corrected status adds evidence rather than deleting the inconvenient scan. The projection rule is versioned so support can explain why “out for delivery” became “exception,” and the system must represent unknown, stale, and conflicting—not collapse all three into “in transit.”


