Real-Time GPS Tracking: From Device Ping to Live Map
How a fleet of GPS devices turns location pings into a live map, without losing positions, overwriting history, or melting the database.

Imagine a logistics company with five thousand vans. Every van has a tracker bolted under the dashboard. Every van reports its position every five seconds, whether or not anyone is watching.
An operations manager opens a map. They want to see where the vans are, which ones have stopped moving, and which ones entered a restricted zone. The map should feel live, so a position that is four seconds old is already stale on screen.
That is the whole product. It is also a write-heavy distributed system, and almost every design decision follows from one number: how many pings arrive per second.
The two questions behind a tracking screen
Before choosing a database or a message broker, separate the two things a tracking product actually needs.
The first question is where is it now. That is a single row per asset, overwritten several times a minute. Only the latest value matters. If you miss an update, the next one supersedes it.
The second question is how did it get here. That is a trail. The previous value is not superseded, because the whole point is the sequence. A missed update here is data loss.
These two questions want opposite things. The first wants a small, hot, cheaply overwritten record. The second wants an append-only log that never discards anything.
current position -> one record per asset, overwritten
historical trail -> one record per observation, appended forever (or until retention says otherwise)
Most painful tracking systems exist because both questions were pointed at the same table.
Start with the numbers
Assume five thousand vehicles, reporting every five seconds.
5,000 vehicles / 5 seconds = 1,000 pings per second average
One thousand writes a second does not sound dangerous until you do it daily.
1,000/sec x 86,400 sec = 86.4 million pings per day
86.4 million x 365 days = 31.5 billion pings per year
At roughly 150 bytes per position payload, that is about 4.7 TB of raw data a year, before indexes, before replication, before any of it is replicated three times for durability.
Now consider the read side. If two hundred people have the map open and the map refetches every second, that is 200 requests a second against whatever is serving positions. The write load and the read load are almost entirely unrelated, which means one of them is always going to be the wrong shape for a single system.
A peak factor matters too. Vehicles report on a timer, so a fleet tends to synchronise. A restart at 09:00 can produce a burst where every device reports within the same second. Plan for two to three times the average, and design so that a burst is absorbed by a queue rather than by your database.
Why the relational table becomes the bottleneck
The obvious first design is one table:
CREATE TABLE positions (
asset_id uuid NOT NULL,
recorded_at timestamptz NOT NULL,
latitude double precision NOT NULL,
longitude double precision NOT NULL,
speed_kph double precision NOT NULL,
heading double precision NOT NULL,
PRIMARY KEY (asset_id, recorded_at)
);
This table is not wrong. It is a perfectly good audit log. The problems appear when you try to also use it to answer the current-position question.
A query for current position becomes a window function over a table that is growing by 86 million rows a day:
SELECT DISTINCT ON (asset_id) *
FROM positions
ORDER BY asset_id, recorded_at DESC;
That query cannot use the primary key efficiently for a fleet-wide view, because it needs the newest row per asset across the entire table. You will add an index, the index will be large, and every insert will pay to maintain it. You will end up with a hot index on the highest-write table in the system.
The second problem is retention. To drop old rows from a table that is receiving 86 million inserts a day, you run large deletes. On most databases, that means vacuum churn, bloat, and a maintenance window you did not plan for.
The third problem is that writes and history have different lifetimes. When an asset is sold, you probably want to keep its historical trail for compliance. When a vehicle is reassigned, you probably do not want its last known position to keep updating under a new identity. Same table, different rules.
Split the two questions into two stores and both problems get smaller.
The shape of the system
+--------------------------+
| Current position (hot) |
| Redis hash per asset |
+------------^-------------+
|
Vehicle --cellular--> Ingest API --> Kafka topic
|
+------------v-------------+
| Position history |
| time-series database |
+------------^-------------+
|
+------------v-------------+
| WebSocket gateway |
+------------^-------------+
|
map clients
There is also a slower path that is easy to forget:
Kafka --> Geofence evaluator --> Alert service --> notification
And a batch path for anything that does not need to be live:
Kafka --> object storage (Parquet) --> analytics and reporting
Keep those three consumers on one topic rather than chaining services together. A WebSocket gateway and an alert evaluator have completely different latency requirements and failure modes, and they should be able to fall over without affecting each other.
The device side
A tracker is a small embedded device with a GPS receiver, a modem, and a battery. Its constraints shape the protocol.
The report interval is a business decision with a direct cost. Dropping from five seconds to fifteen seconds cuts ingest volume by two thirds. Do that before optimising the ingest path.
The device should send a compact payload. Binary encodings such as protobuf, or a fixed-layout binary format, cut the 150 bytes above to a fraction of that. On a metered cellular connection with a battery budget, this matters more than server-side throughput.
Devices are offline regularly. A van in a tunnel, a yard with poor coverage, a truck in a basement car park. The tracker needs a small buffer and must acknowledge what it stores, or it will silently lose the period it was offline.
The device also needs an explicit clock story. The recorded_at value is produced by the device clock, which drifts. The server’s receive time is a separate value. Store both. Once you have only one of them you cannot tell the difference between a vehicle that was stationary and a device with a badly wrong clock.
Ingestion: the part that has to survive
The ingest API is the front door, and it is the component most likely to be measured by how fast it accepts a write.
Its job is narrow:
- Authenticate the device.
- Validate the payload shape.
- Stamp the server receive time.
- Produce to Kafka with the asset id as the message key.
- Return quickly.
It should not write to Redis, should not write to the time-series database, and should not evaluate geofences. Every one of those is a failure the device does not need to know about.
The message key is the interesting decision. Using asset_id routes all pings for one vehicle to the same partition, which gives per-vehicle ordering. That matters for geofence entry and exit logic, where an out-of-order pair of positions can produce a nonsense crossing.
The cost is that one partition holds one vehicle’s traffic, and a skewed fleet means a few very hot partitions. If a single asset is genuinely hot, you need per-vehicle sub-keying, and you have to accept that entry and exit events for that vehicle may now be processed out of order.
Return a 202 with a record identifier rather than a 200 that implies the position is queryable yet. The device should not wait for the time-series write. It should not wait for Redis either. A tracker with a two-second TCP timeout will retry the whole payload, and you have now doubled your ingest volume for no benefit.
device -> POST /ingest -> Kafka (ack) -> [async] current position
-> [async] history
-> [async] geofence evaluation
Current position: the hot store
Current position is a perfect fit for an in-memory key-value store. It is small, it is read constantly, and it is overwritten constantly.
A Redis hash per asset:
HSET pos:{assetId} lat 51.5074 lon -0.1278 speed 42 heading 180 recordedAt 1730000000000 receivedAt 1730000012000
Add a TTL, and the TTL is the interesting part. Set it to two or three report intervals. Set it to an hour and a decommissioned vehicle stays on the map forever. Set it to exactly one interval and a single dropped ping makes the asset disappear, which looks identical to a real problem and is much more annoying.
TTL = 3 x report interval
A map that shows 5,000 assets is reading 5,000 keys. That is nothing for Redis. The design decision that matters is not throughput, it is what you do when the hot store is empty or stale for an asset, which is the next problem worth writing down.
History: the time-series store
Every ping goes to a time-series database, partitioned or indexed by asset and time. The reason is simple: the access pattern is always a time range over a set of series, and the write pattern is always append.
Store more than the coordinates. speed_kph, heading, acceleration, ignition, and odometer are all cheap to store at ingest and expensive to reconstruct later. Forwarded geofence crossings and driver behaviour events also belong here, computed once at ingest rather than recomputed from raw coordinates every time someone opens a report.
This is the store you downsample later, and downsampling only works if you stored the raw value with a known timestamp. The moment you store a pre-aggregated value, the aggregation is permanent.
Ordering, lateness, and the out-of-order ping
Here is the part that quietly breaks tracking systems.
Two pings for the same vehicle can arrive in the wrong order. The vehicle loses signal, buffers locally, and reconnects. The buffered pings are then delivered after newer pings that already went out over a different network path.
If your current-position update is a blind SET, a late ping overwrites a fresh one and the vehicle jumps backwards on the map.
The fix is a conditional write on recorded_at:
read existing recordedAt
if incoming recordedAt > existing recordedAt
overwrite
else
drop for the current-position store
The late ping is still valid for history, where it lands in the right time slot. It is only rejected for the latest-value view. Those are two different decisions and they should be made in two different places.
Make both writes idempotent. Give every position a deterministic identifier, and treat a repeated delivery as a no-op rather than an error. Devices retry. Network stacks retry. Kafka redelivers after a consumer crash. All three are normal.
For the history store, decide what happens when two records claim the same asset and timestamp. Last-write-wins, first-write-wins, or an explicit conflict record. Pick one, write it down, and make sure the geofence evaluator knows which one it will see.
The live map
Now the read side.
A map does not want the whole fleet in one response. It wants a viewport. A bounding box plus a zoom level tells you which assets matter, and it is the difference between returning 5,000 positions and returning 80.
GET /map/positions?bbox=minLon,minLat,maxLon,maxLat&zoom=12
Use a geospatial index for the bounding box filter. If the current-position store is Redis, a geo set per city or region is a reasonable structure. If it is PostgreSQL, PostGIS with a GiST index is the straightforward answer.
The WebSocket gateway exists for the same reason: polling a moving map feels broken, and a poll every second across 200 open maps is a lot of repeated identical work.
The gateway’s job is subscription management, not data storage. A client subscribes to a viewport and receives positions for assets inside it, plus an immediate snapshot on connect.
Then handle the failure case, because it always happens. A phone in a lift loses its socket. On reconnect, the client sends the time it was last updated, and the gateway sends everything it missed before resuming live updates. Without that, a reconnecting client shows a map frozen at the moment it dropped.
A useful test is to kill the connection on purpose and watch what the operator sees. If the map silently freezes, the catch-up path is missing.
Geofences
A geofence is a polygon or circle with an associated action. A depot, a restricted zone, a customer site.
Evaluation is straightforward and expensive if done naively. Checking five thousand assets against two hundred polygons on every ping is a million polygon tests a second, and it is completely avoidable.
The standard optimisation is to only evaluate assets that moved far enough to matter:
received position
-> has the asset moved more than the geofence error margin since last evaluation?
no -> skip
yes -> which fences does the current point fall in?
which fences did the previous point fall in?
compare the two sets -> ENTER, EXIT, or nothing
The margin should be larger than the expected GPS error, which is single-digit metres for a decent consumer receiver and worse when the vehicle is under cover or in a canyon of tall buildings. Testing fences that are smaller than your positioning noise produces alerts that operators learn to ignore, and ignored alerts are worse than no alerts.
Keep a small in-memory index of fences and their bounding boxes. A fence that is not in the same city as the asset cannot possibly contain it, and rejecting on the bounding box first turns most checks into a dictionary lookup.
Alerts
An alert is a geofence result plus context and a delivery decision.
The three questions worth answering before building:
- Deduplicate by what? One alert per vehicle per fence per day, or one per entry event? Entry events are usually correct for geofences and completely wrong for a vehicle that is idling just inside a boundary.
- Suppress how? A vehicle that loses signal inside a restricted zone will not generate an exit event, and an operator who sees “entered” with no matching “exited” will stop trusting the system.
- Deliver how? In-app, email, SMS, webhook. The delivery mechanism should not be in the geofence evaluator.
A silent alert is a design failure. If the system knows it cannot deliver, it must still record the event so an operator can see that something was attempted.
Retention and the cost of storing everything
Keeping every five-second ping for five thousand vehicles for a year is 31.5 billion points. It is also mostly redundant, because a vehicle travelling at 60 km/h moves about 83 metres in five seconds, and asking for the same location twice is not useful.
Downsampling by age is the standard answer:
0 to 7 days -> keep raw 5-second points
7 to 90 days -> 1-minute points
90 days onward -> 5-minute points
For the fleet above that turns roughly 31.5 billion points a year into about 1.6 billion, a reduction of around 19 times.
The numbers are worth checking rather than trusting, because they depend entirely on the report interval and the retention window:
raw 5s, 7 days -> 5,000 x 17,280 x 7 = 604.8 million
1 min, 83 days -> 5,000 x 1,440 x 83 = 597.6 million
5 min, 275 days -> 5,000 x 288 x 275 = 396.0 million
total ~ 1.60 billion points
Two rules keep this honest. First, raw data must be kept long enough that the downsampled series is still useful for the questions people actually ask, and a 30-day-raw policy quietly destroys any investigation into something that happened six weeks ago. Second, downsampling should keep a mean and a min and a max, not only a mean, because an operator chasing a speed spike needs the spike, not the average around it.
When a device goes quiet
Stale data is a real state, and it needs its own representation on the map.
A device that stops reporting is not a vehicle at its last known position. It is a vehicle of unknown position, last seen at a time. Those should look different, and the API should say so:
GET /assets/{id}
position -> { lat, lon, speed, recordedAt, ageSeconds }
status -> "moving" | "stationary" | "stale"
staleAfter -> 3 x report interval
Deriving “stationary” from a low speed value is also wrong at the boundary. Decide stationary from sustained low speed over a window, not from a single sample, and make the window long enough to survive one missing ping.
Alert on the aggregate rate of stale devices, not on individual ones. A single stale device is often a parked van in a basement. Thirty percent of the fleet going stale in five minutes is an ingest or carrier problem, and it deserves a page rather than a map colour change.
Failure stories worth testing
A vehicle reconnects after an hour offline
The device flushes its buffer. Confirm the late pings land in history in the correct time order, are rejected by the current-position store, and do not move the marker backwards.
Two pings share a timestamp
Decide what history does before this happens in production rather than after an operator reports a missing point.
The time-series database is unavailable
The current-position store should keep working. The map should keep moving. History should buffer in Kafka and catch up. If the map stops when the history store has a bad afternoon, the two are too tightly coupled.
Redis is unavailable
Decide what the API returns. Returning stale positions with an age is honest. Returning a 500 is a worse experience than a slightly old marker.
A geofence polygon is edited while vehicles are inside it
Decide whether editing a fence replays entries and exits for everyone currently inside. It probably should, and that means the fence edit path needs the same evaluation code as the normal path.
A vehicle is decommissioned
Its key should expire, its trail should stay, and its historical data should not keep accruing pings from a device that still has power.
A single partition becomes hot
One very active asset, such as a vehicle idling in one spot sending constantly, can dominate a partition. Watch per-partition skew rather than average throughput.
A production-ready architecture
+----------+ MQTT/HTTP +------------+ produce +--------+
| Tracker | ------------> | Ingest API | -----------> | Kafka |
+----------+ +------------+ +---+----+
| device auth | |
| validation | |
| receive timestamp | |
+---------+---+----------+
| | |
+--------v--+ +---v--------+ +-v-----------+
| Current | | Position | | Geofence |
| position | | history | | evaluator |
| (Redis) | | (TSDB) | | -> alerts |
+--------+--+ +------------+ +-------------+
|
+--------v--------+
| WebSocket |
| gateway |
+---------+--------+
|
map clients
Slow path: Kafka -> object storage -> analytics and compliance reporting
A sensible delivery checklist:
- Separate current position from history before choosing any component.
- Measure ingest volume as devices times report interval, not as user traffic.
- Give each device a stable identity and a rotating credential.
- Use the asset id as the message key and accept the partition skew that follows.
- Acknowledge the device after Kafka, not after the downstream stores.
- Make every downstream write idempotent and keyed on a deterministic position id.
- Guard the current-position store against late and out-of-order pings.
- Give stale data its own state on the map instead of a stale-looking position.
- Define the WebSocket catch-up path before the first field report of a frozen map.
- Downsample by age, and keep min and max as well as the mean.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| One table for current position and history | The newest-row-per-asset query scans a table that grows by 86 million rows a day | Separate a hot current store from an append-only history store |
A blind SET for current position |
A late ping moves the vehicle backwards on the map | Compare recordedAt and only accept a newer value |
| Only storing the device timestamp | Clock drift becomes indistinguishable from a stationary vehicle | Store server receive time as well |
| Acknowledging only after the database write | Device timeouts cause the whole payload to be retried | Acknowledge after Kafka and absorb the rest asynchronously |
| Checking every fence on every ping | A million pointless polygon tests a second | Gate on distance moved, then on fence bounding box |
| Alerts on any geofence contact | Idling on a boundary generates alert noise operators learn to ignore | Require an entry event and a margin larger than GPS error |
| Missing exit event after signal loss | Operators see entries with no exits and stop trusting the system | Detect stale devices and reconcile the fence state |
| WebSocket with no catch-up | A reconnecting client sees a map frozen at the moment it dropped | Send a snapshot plus everything missed since the client’s last update |
| Retaining five-second data for a year | 31.5 billion points, mostly redundant | Downsample by age and keep min and max |
| Assuming the fleet reports evenly | A fleet restart produces a burst of synchronised pings | Absorb bursts in the queue and plan for two to three times average |
| Deleting old rows in large batches | Index churn and maintenance pressure on the highest-write table | Let retention be the store’s job, not a scheduled DELETE |
| Treating a stale device as a parked one | An ingest outage looks like a fleet full of stationary vehicles | Expose ageSeconds and a separate stale state |
The complete story in one minute
A tracker wakes, gets a GPS fix, and posts a compact position with its own timestamp. The ingest API authenticates the device, validates the shape, stamps the server receive time, and produces the record to Kafka using the asset id as the key. It returns as soon as Kafka accepts the write.
Three consumers read that topic independently. One writes history to a time-series database, which keeps the full trail and downsamples it as it ages. One maintains current position in Redis, accepting the write only when the incoming timestamp is newer than what is already there, with a TTL of roughly three report intervals. One evaluates geofences, but only for assets that have moved far enough to change the answer, and compares the fences the previous position was inside with the fences this one is inside to produce entries and exits.
A WebSocket gateway reads the hot store and pushes positions to map clients filtered by viewport, and a client that reconnects asks for everything it missed before live updates resume.
Where a device has gone quiet, the API says so with an age rather than showing a confident old position.
That is the whole path:
device fix -> ingest -> Kafka -> (history | current position | geofences) -> map
The hard parts were never the map. They were deciding that current position and history are different problems, that a late ping is a normal event, and that a device which stops reporting is a state the system has to be able to represent honestly.
What this team still owns
“Current position” is a projection selected from observations, not the row with the greatest arrival time. Carry event time, receive time, accuracy, fix source, sequence/session, and motion context; define how far out of order an update may replace the live marker. History can accept later corrections without making the live marker jump backward. Every consumer sees age and accuracy, because a precise-looking stale dot is a product lie.


