← All writing
articleJul 04, 202519 min read

Logistics Package Tracking: The Event Stream Behind a Progress Bar

Package tracking as a scan-event pipeline — status models, exception handling, the ETA problem, and why carrier webhooks are the hard part.

LogisticsEvent StreamingArchitectureData Modeling
Logistics Package Tracking: The Event Stream Behind a Progress Bar cover illustration

A tracking page looks trivial. Behind it is an event pipeline that ingests scan data from dozens of carriers in inconsistent formats, maintains a state machine per parcel, predicts a delivery window, and survives the fact that the carriers themselves are the least reliable part of the system.

This is an article about that pipeline, and about the specific failure modes that appear once a parcel crosses three or four different organisations’ hands.

The scale

  a national carrier network
    40M parcels per year
    ~110,000 per day
    peak: 400,000 (December)

    scans per parcel: 8-15
    total scan events: ~1.5M per day
    peak: 5M per day

    event volume is not the hard part
    the hard parts:
      - 14 carrier formats, none documented
      - scan events arrive out of order
        (a hub scan arriving after a
         delivery attempt)
      - the same scan arrives 3 times
      - ~2% of parcels go silent for
        5+ days with no event at all
      - "out for delivery" does not mean
        it is on the van

The volume is unremarkable. Everything that makes this hard is about correctness under partial, late, and contradictory information.

The state machine

  CREATED
    -> INFORMATION_RECEIVED
    -> PICKED_UP            (origin scan)
    -> IN_TRANSIT            (line-haul hub scans)
    -> AT_DESTINATION_HUB
    -> OUT_FOR_DELIVERY      (loaded on van)
    -> DELIVERED             (proof: signature/photo)
    -> DELIVERY_ATTEMPTED
    -> RETURNED_TO_SENDER
    -> EXCEPTION: LOST / DAMAGED / HELD / RETURNED
    -> CANCELLED

  illegal transitions that arrive in real life
    - DELIVERED -> IN_TRANSIT
      (a second carrier's late scan)
    - IN_TRANSIT -> INFORMATION_RECEIVED
      (out-of-order replay)
    - DELIVERED -> OUT_FOR_DELIVERY
      (mis-scan at a hub)
    - anything -> anything, from a manual
      carrier override in their own UI

The most common integration bug is not a missing field. It is an out-of-order or duplicate event producing an impossible status that then gets shown to a customer. So the pipeline has to be defensive in a specific way:

  on each event
    1. dedupe by (parcel, event_id)
    2. if event carries a scan timestamp
       older than current state's timestamp
       -> log it, do not apply it
       (or route to a "late event" review)
    3. validate the transition against the
       state machine
    4. if illegal
       -> record it in an exceptions table
       -> decide: ignore, or correct the
          state with a repair rule
    5. if legal -> apply, emit a new state
       event, update projections

The rule about event timestamps rather than arrival order is the important one. Scan events carry their own timestamps from the carrier’s handheld or hub system, and those timestamps can be hours apart from arrival. Applying by arrival order means a stale event overwrites a fresher state; applying by scan timestamp means the most recent scan wins, which is what the customer expects to see.

Status normalisation

  carrier A: "DPU"  "OTD"  "ARR"  "IT"  "DEL"
  carrier B: "PICKUP_COMPLETE"  "OUT_FOR_DELIVERY"
             "ARRIVED_AT_FACILITY"  "IN_TRANSIT"
             "DELIVERED"
  carrier C: numeric codes, a different
             code for "attempted, customer
             absent" vs "attempted, address
             incomplete"
  carrier D: free text, occasionally
             translated

  the mapping table is the product
    200+ source codes
    -> 12 canonical statuses
    -> 40+ customer-facing strings
       (different wording per market)

Every carrier integration is a mapping problem and the mapping is versioned, because carriers change their codes without notice. The robust pattern is a mapping layer with a default bucket plus an explicit unknown-value path:

  unknown code
    -> canonical status UNKNOWN_EVENT
    -> logged with the raw value
    -> surfaced in a monitoring dashboard
    -> never silently mapped to a
       customer-facing status

  because the alternative is a parcel
  that shows "in transit" forever
  because of a code you never mapped

That monitoring dashboard is not optional. A carrier that adds a code on a Tuesday and ships a million parcels through it means your customers are seeing stale or wrong statuses for a day, and the only way to know is to watch the unknown-rate.

The ETA problem

  a parcel with a good scan history:
    created 09:00
    picked up 14:00
    in transit 16:00
    at destination hub 04:00 (+1d)
    out for delivery 08:00 (+2d)

  ETA for delivery: ?

  naive: hub scan + historical average
    -> 14:00 (wrong, 40% of the time)

  good: a model conditioned on
    - origin/destination pair
    - service level
    - day of week
    - time of day of the last scan
    - destination's local cut-off time
    - current weather (a real factor
      for road networks)
    -> distribution over delivery windows

The distribution is the product. Two design commitments follow:

Predict a window, not a timestamp. “Between 1pm and 5pm” is honest and is what the customer actually uses. A precise time that is wrong 40% of the time trains people to distrust the number entirely, and once they distrust it they stop using the number at all.

Recalculate on every state change. The ETA is a function of the current state, so it changes as the parcel moves. A parcel at the destination hub at 04:00 has a different and much tighter ETA than the same parcel that was in transit yesterday. Recompute it on the event, not on a nightly batch, because the customer is watching the page while the parcel moves.

The specific failure that annoys customers most is the ETA that gets worse as the parcel gets closer. That happens when the model is a simple average of historical transit time and does not condition on the current position. Conditioning on position — a parcel already at the destination hub is categorically different from one in line-haul — is what makes the ETA feel intelligent.

Exceptions are statuses, not errors

  exception types
    STUCK          - no scan in N hours,
                     location unchanged
    LOST           - no scan in M days,
                     trajectory impossible
    DELAYED        - missed a transit scan
                     deadline
    DAMAGED        - carrier-reported
    RETURNED       - sender/address issue
    HELD           - customs, weather, address
                     verification
    MISSORT        - destination hub could
                     not process it
    ADDRESS_ISSUE  - incomplete or
                     contradictory address

  each one needs
    - a customer-facing message that is
      actually useful ("we've contacted
      the recipient", "we need more
      information from the sender")
    - an owner (who is working on it)
    - a resolution path (retry, reroute,
      return, contact)
    - an SLA

The design principle is that an exception is a state the parcel is in, not a failure of the tracking system. Modelling it as a nullable error means the customer page shows the last known status — “in transit” for a parcel that has been sitting in a sorting office for a week. The exception state is the single most valuable thing on the tracking page when something is genuinely wrong, and it is the thing most implementations get wrong by treating it as an error code.

The stuck/lost detection is worth doing carefully because it is the case that actually matters:

  stuck
    no scan in 6 hours AND
    last known location has not changed
    AND
    not in a known quiet period
      (a hub that legitimately does
       not scan overnight)

  lost
    no scan in 3 days AND
    the last location is inconsistent
      with the destination
      (e.g. still at the origin country
       when the destination is domestic)
    AND
    the carrier has not acknowledged it

The “quiet period” exemption is the detail that makes this usable. A naive stuck detector fires every night for every parcel in a hub that does not scan overnight, and once it fires habitually everyone ignores it. The detector has to know which facilities have quiet periods and which do not.

Carrier webhooks: assume the worst

  what a carrier integration actually does
    - sometimes delivers via webhook
    - sometimes via polling
    - sometimes via SFTP file drop
    - sometimes via an email that
      someone parses (yes)
    - retries with duplicates
    - delivers out of order
    - delivers late (hours)
    - delivers gaps (missing scans)
    - changes the code table without notice
    - sometimes goes down for a day
    - has no rate limit, then rate-limits
      you hard

  design for all of it
    - idempotent consumers (event_id dedupe)
    - a per-carrier raw event store, kept
      for debugging and replay
    - polling as a backstop for webhook
      carriers (detect gaps, backfill)
    - a dead-letter queue with alerting
    - per-carrier lag monitoring

The raw event store is the component that people skip and then wish they had. When a customer reports that a parcel showed the wrong status for six hours, the only way to answer is to look at what the carrier actually sent — and if you normalised and discarded on ingest, that information is gone. Keeping the raw payload per carrier is cheap relative to the debugging cost of not having it.

The polling backstop matters more than it looks: webhooks are the fast path, but the gap detection (“we have not received a scan for this parcel in 4 hours, the carrier says it should have moved”) is what catches a webhook that silently stopped working. Without polling, a broken integration looks exactly like a carrier having a bad day.

Retention, and why it is a product decision

  tracking data ages differently
    - standard parcels: useful for ~90 days
    - high-value / signed: longer, often
      1-2 years for claims and disputes
    - international: customs records have
      statutory retention requirements
    - failed/lost: needs to be retrievable
      as long as a claim is possible
    - analytics aggregates: keep forever
      (anonymised)

  the trap
    deleting the raw event log too early
    -> a claim 6 months later has no
       evidence trail
    -> refunds get issued on the basis
       of a customer's word

  the right shape
    - hot: current state + recent events
      (fast, queryable)
    - warm: full event history per parcel
      (slower, still retrievable)
    - cold: anonymised aggregates
      (analytics)
    with a per-class retention policy

A tracking system that deletes its history to save storage has quietly turned every dispute into an argument, and the argument is always resolved in favour of the customer, because that is the commercially correct decision when you have no evidence. Retention tiers cost very little and are the difference between a defensible claims process and an expensive one.

Failure stories worth testing

Deliver the “out for delivery” scan twice

The status must not change and the event must not double-apply. This is the idempotency test.

Send a hub scan with a timestamp earlier than the current state’s scan

The state must not regress. This is the out-of-order test and it is the most common real-world failure.

Send a “delivered” event followed by an “in transit” event

The system must not show a delivered parcel in transit. Whatever repair rule you use, it must be explicit.

Remove a carrier’s code mapping and replay a day of events

Confirm the unknown-rate alert fires and the raw payload is available. This is the “carrier changed their codes” test.

Deliver a parcel to a hub at 04:00 and check the ETA trajectory

The ETA must tighten as the parcel approaches. If it does not, the model is not conditioning on position.

Set the ETA to a single precise time and measure accuracy

Measure the fraction within 1 hour, within 2 hours, within 4 hours. Then compare against a window model with the same underlying predictions — the window is more useful at every horizon.

Go silent on a parcel for 5 days

Confirm the stuck/lost detector fires, that it respects quiet-period facilities, and that the customer page shows an exception with a next action.

Break the webhook endpoint and confirm polling backfills the gap

Without polling, a broken integration is invisible until someone complains.

Flood the webhook with duplicates at 10x rate

Confirm idempotency holds and the lag monitor catches the rate.

Change a carrier’s status wording mid-flight

The canonical mapping must hold; only the customer-facing string should change. Version the string tables separately from the code tables.

Route a parcel through an international leg with customs holds

Confirm the exception state explains the delay in terms the customer can act on, and that the sender — not the customer — is the party being asked for information.

Delete a parcel’s history and then try to resolve a claim

This is the retention test, and it demonstrates exactly why the warm tier exists.

A production-ready architecture

   carrier sources
     - webhooks (many, unreliable)
     - polling backstops
     - file drops, legacy integrations
        |
        v
  +----------------------------------------------------------+
  |  INGEST                                                  |
  |  - raw event store per carrier (immutable)              |
  |  - dedupe by (carrier, event_id)                        |
  |  - normalise: carrier code -> canonical status          |
  |  - unknown codes -> UNKNOWN_EVENT + alerting            |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  STATE MACHINE per parcel                                |
  |  - apply by scan timestamp, not arrival order            |
  |  - validate transitions; illegal -> repair rule + log    |
  |  - exceptions as first-class states                      |
  +----------------------------+-----------------------------+
                               |
             state change
                               |
        +----------------------+----------------------+
        v                      v                      v
  +-------------+      +----------------+     +------------------+
  | ETA model   |      | projections    |     | exception engine |
  | conditioned |      | (current view, |     | stuck / lost /   |
  | on position |      |  index, search)|     | delayed / held   |
  +-------------++      +----------------+     +------------------+
        |                      |                      |
        +----------------------+----------------------+
                               v
  +----------------------------------------------------------+
  |  TRACKING API + PAGE                                     |
  |  - status history, ETA WINDOW, exceptions with next step  |
  |  - language/locale per market                            |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  RETENTION TIERS                                          |
  |  hot: current state + recent events                      |
  |  warm: full per-parcel event history (claims evidence)   |
  |  cold: anonymised aggregates                             |
  +----------------------------------------------------------+

  watch: unknown-code rate, event lag by carrier, webhook uptime,
        stuck-detection precision, ETA accuracy by horizon

Delivery checklist:

  1. Make the tracking flow a state machine with explicit legal transitions, and treat every other transition as something to log and repair deliberately.
  2. Apply events by scan timestamp, not arrival order. This is the single most important correctness rule.
  3. Dedupe on a stable event id, and keep the raw carrier payload in a separate store so debugging is possible.
  4. Version the carrier code mapping, monitor the unknown-code rate, and never silently map unknown codes to a customer-facing status.
  5. Predict a window, not a timestamp, and recalculate on every state change. Conditioning on the parcel’s current position is what makes the ETA feel right.
  6. Make exceptions first-class states with a customer-facing message, an owner, a resolution path, and an SLA. The stuck/lost detector needs to know which facilities have quiet periods.
  7. Build polling backstops for webhook carriers so a broken integration is detectable before a customer reports it.
  8. Monitor per-carrier event lag and webhook uptime as first-class metrics.
  9. Use retention tiers, with a longer window for high-value, international, and exception parcels, so claims have an evidence trail.
  10. Keep the aggregates anonymised and forever, and keep them separate from the per-parcel history.
  11. Separate the canonical status codes from the customer-facing strings so wording can change per market without touching logic.
  12. Give every exception a “what happens next” line. A status that tells a customer nothing is worse than no status at all.

Common mistakes

Mistake What actually happens Better decision
Applying events by arrival order Stale scans overwrite fresh state Apply by scan timestamp
No event dedupe Duplicate scans cause phantom transitions Dedupe on (carrier, event_id)
Illegal transitions applied Customers see delivered-then-in-transit Validate, log, repair deliberately
Normalise and discard raw events No way to debug a bad status Keep a per-carrier raw event store
Unmapped carrier codes Parcels stuck “in transit” forever Unknown bucket + alert on the rate
A single ETA timestamp False precision, distrusted numbers A window, recalculated per event
ETA not conditioned on position ETA gets worse as the parcel nears Model position explicitly
Exceptions as error/nulls Page shows the last known status forever Exceptions as first-class states
Stuck detector with no quiet periods Fires nightly, gets ignored Per-facility quiet periods
Webhook-only integration A broken webhook is invisible Polling backstop and lag monitor
Single flat retention Claims have no evidence Retention tiers by parcel class
Canonical status tied to wording Market wording changes break logic Separate code and string tables
No “what happens next” on exceptions A status that informs nobody Owner, action, and SLA per exception
Aggregates mixed with per-parcel data Deleting for privacy loses claims data Anonymised cold tier, kept separate

The complete story in one minute

A tracking page is an event pipeline in disguise, and the volume is not the hard part — 1.5 million scan events a day is unremarkable. The hard part is that a parcel crosses three or four organisations and the scan events arrive late, duplicated, and out of order, from fourteen carriers with 200+ undocumented status codes between them. So the core is a state machine per parcel, applied by scan timestamp rather than arrival order, deduplicated on a stable event id, with every illegal transition logged and repaired by an explicit rule rather than applied. And the single most valuable debugging asset is the per-carrier raw event store: if you normalise and discard on ingest, you will never be able to answer “what did the carrier actually send” when a customer says the status was wrong for six hours.

The ETA is a prediction whose distribution is the product. A window of 1pm to 5pm that hits 75% of the time beats a precise 4:07pm that is wrong 40% of the time, because a false-precision number trains people to stop trusting the whole feature. Recalculate on every state change and condition on the parcel’s current position — the worst customer experience is an ETA that gets worse as the parcel gets closer, which happens the moment the model is a flat historical average.

Exceptions are statuses, not errors. A package stuck for a week should show “we’ve contacted the recipient”, not the last known status of “in transit”. Stuck and lost detection has to know which facilities have quiet periods, or it fires every night for every parcel and gets ignored — which is worse than not having it. Retention is a product decision with a legal dimension: per-class tiers where high-value, international, and exception parcels keep their history long enough to settle a claim, and aggregates kept anonymised and separate, because a system that deletes its history settles every dispute in favour of the customer.

state machine, applied by scan timestamp, deduped, illegal transitions logged
per-carrier raw events kept; unknown-code rate monitored
ETA as a window, recalculated, conditioned on position
exceptions as first-class states with an owner and a next action
polling backstop on every webhook; retention tiers by parcel class

The hard part was never displaying the status. It was deciding, when the data disagrees with itself, which of the fourteen sources to believe — and admitting that a package sitting in a sorting office for a week is information the customer is owed.

What this team still owns

Keep observations immutable and derive customer status from them. A scan has source, facility, device, event time, receipt time, carrier code, and confidence; a corrected status adds evidence rather than deleting the inconvenient scan. The projection rule is versioned so support can explain why “out for delivery” became “exception,” and the system must represent unknown, stale, and conflicting—not collapse all three into “in transit.”

Technical references

Keep reading
Browse everything