Health Checks and Heartbeats: Three Different Questions That Look Identical
Liveness, readiness, and heartbeat are not the same test, and conflating them is how a slow dependency restarts your entire fleet.

Every service has a health check, most systems have a heartbeat, and almost nobody can explain the difference on request. They are three different mechanisms answering three different questions, they have opposite failure semantics, and treating them as one thing produces a specific and extremely common outage: a slow database restarts every replica of every service that depends on it.
This is a short article because the design is small. The reason it matters is that the small design, done wrong, takes out production.
The three questions
+----------------+--------------------------------------------------+
| mechanism | question |
+----------------+--------------------------------------------------+
| liveness | is this process wedged and in need of a restart? |
| readiness | should traffic be sent to this instance right now? |
| heartbeat | is this instance still there at all? |
+----------------+--------------------------------------------------+
liveness failure -> orchestrator RESTARTS the container
readiness failure -> load balancer STOPS sending requests
heartbeat missing -> node is declared DOWN and evicted
The failure semantics are opposites, which is why mixing them is dangerous:
- A liveness failure causes a restart. A restarting process stops working.
- A readiness failure removes it from rotation. A removed instance keeps running and recovers on its own.
- A missing heartbeat declares it dead and triggers replacement elsewhere.
A process that fails liveness is expensive to fix. A process that fails readiness is cheap to fix. Design liveness to be as dumb as possible.
Liveness: the dumbest check that can detect a wedge
Liveness exists for one failure: the process is alive but no longer able to do anything. Deadlocked threads, a wedged event loop, a corrupted state machine, an exhausted file descriptor table that cannot accept a connection.
liveness probe
GET /live
-> 200 OK, immediately
the process cannot serve this request
-> the probe times out
-> the orchestrator restarts the container
Two properties matter and both are about restraint.
It must not check dependencies. A liveness probe that queries the database is a mechanism for turning a database slowdown into a full restart of your service. Consider what happens:
database p99 goes from 5ms to 4 seconds
readiness fails everywhere (correct)
liveness fails everywhere (if you wired it to the db)
every replica restarts at once
-> connection storm at cold start
-> all replicas hammering the database
-> database gets slower still
-> positive feedback, total outage
That loop is not hypothetical and it is one of the most common self-inflicted outages in containerised systems. The database is not something a liveness probe should be able to restart.
It must not be expensive. A liveness probe runs every few seconds across every replica, so 1,000 replicas at a 5-second interval is 200 probes per second for the lifetime of the cluster. If the probe touches the database, that is 200 database queries per second of pure overhead. If the probe allocates and triggers a garbage collection, you have built a periodic memory pressure event into your service.
The recommended shape is a handler that returns 200 without doing work. The one thing it must actually detect is wedgedness, and there are two ways to do that. One is a select(2) or equivalent on a file descriptor that the event loop should be servicing — a kernel-level check that costs nothing. The other is a watchdog: the application updates a timestamp on its main loop, and the handler compares it to now, failing if the loop has not progressed in a threshold. The watchdog detects actual wedging, which is the thing you want to catch, and it costs one comparison.
Readiness: honest about what this instance can do
Readiness is the check that should be thorough, because it is the one that actually protects the system. A failed readiness check removes an instance from rotation, which reduces load on the struggling part of the system.
GET /ready
- am I past startup initialisation? -> else not ready
- am I still accepting connections? -> else not ready
- is the critical dependency reachable? -> else not ready
- am I within my concurrency limit? -> else not ready
- am I not draining? -> else not ready
all pass -> 200
any fail -> 503
Several things in that list earn their place.
Startup. An instance that has started its process but not finished warming caches, connecting pools, or loading configuration is not ready, and sending it traffic produces errors that look like application bugs. This is the check that prevents a rollout from being worse than the previous version.
Draining. When an instance is shutting down, readiness must fail before the process stops accepting connections, and it should keep failing while in-flight work completes. This is the check that makes graceful shutdown work.
Concurrency limits. Reporting “not ready” when the instance is at its concurrency ceiling is load shedding, and it is far better than accepting a request you cannot serve. This turns a queue in your process into a 503 at your load balancer, which is honest and which clients can handle.
Which dependencies to include is a judgement call, and the rule is: include a dependency only if the instance genuinely cannot serve useful traffic without it, and only if the dependency failing is likely to be persistent. A cache that is down degrades performance; a primary database that is down means you cannot serve at all. And a transient blip should not flap readiness, so add a small number of consecutive failures before flipping to not-ready, and a smaller number of successes before flipping back. Without hysteresis, readiness oscillates and the flapping is worse than either state.
Heartbeat: presence, not health
A heartbeat is a periodic signal that says the instance is still there. It is a different mechanism from a health check and it fails in the opposite direction.
heartbeat
every 10s, lease renewed
(etcd / ZooKeeper / Consul / service registry)
three missed heartbeats (30s) -> instance declared DOWN
-> removed from discovery
-> its pods rescheduled elsewhere
The key difference: a heartbeat failure means something is wrong that the instance cannot fix by itself. A process that is alive but unable to renew a lease is either partitioned from the registry or completely wedged, and in both cases the correct response is to stop trusting it and replace it. A health check that fails and recovers on its own is a normal event. A heartbeat that is missed and comes back is a failed node, and the reason it matters is that the replacement already started while the original might still be running.
That leads to the requirement every heartbeat-based system needs and most forget: the dead instance must be fenced out, not just removed from the registry.
node A loses network to the registry
node A is still running, still serving its
local clients, still writing to the database
orchestrator sees the missing lease
-> starts node B to replace it
-> A and B are now both running
-> A still holds its lease ID (it does not
know the lease expired on the server)
-> A believes it is the legitimate instance
consequence: two nodes, same identity,
both doing work, one of them fenced
A service instance that renews a lease must stop work when renewal fails, not retry forever. A process that notices it has lost its lease and continues is a split-brain bug, and the discipline of “fence yourself” is what makes a registry safe to depend on. This is the same fencing-token idea from distributed locking, applied to service identity: the thing that detects the split is a monotonic value, and the thing that enforces it is the resource, not the registry.
Cascading health, and why you should resist it
There is a temptation to make readiness smarter: check everything the instance depends on, transitively, and mark the instance unhealthy if the whole graph is degraded. This feels thorough and it is a mistake.
naive cascading health
service A depends on B
B depends on database C
C is slow
A checks B: degraded -> A not ready
everything upstream of A also checks A
-> they all go not ready too
result: the whole system reports unhealthy
because one database is slow, and the
instances that could still serve cached
data are taken out of rotation too
The problem is that unready means “do not send me traffic”, and removing every instance in a dependency chain from rotation does not protect the struggling component. It removes the load-handling capacity, and the components that would have kept serving from cache or from partial data are also removed.
The better model is that readiness reflects this instance’s ability to serve, and an instance that can serve degraded responses should report ready and return degraded responses. Reserve “not ready” for the case where the instance should not receive new work at all, which is startup, shutdown, and genuine unavailability. A dependency being slow is a reason to shed load deliberately and return errors you can reason about, not a reason to disappear.
This is worth stating because the instinct to make health checks comprehensive is strong, and the outcome is a design where one slow component removes the whole fleet from service simultaneously. The system gets smaller exactly when it needs to be bigger.
The cost of a check
Every check is load on the thing it checks, and this is routinely underestimated.
1,000 replicas, liveness every 5s
= 200 probes/sec, forever
if the probe opens a DB connection
= 200 extra connections/sec of churn
against a system with a connection limit
if the probe allocates
= a small GC every 5s per replica
Two rules follow. A health endpoint should be able to answer without touching a dependency, where possible — a flag set by the background health check thread is better than a live query, because the flag read is free and the check runs at a controlled rate. And the check should be able to detect a hung dependency without waiting for its own timeout, which means a separate background loop with a short timeout that updates state, rather than a synchronous request that inherits the dependency’s latency.
A synchronously-checked health endpoint with a 10-second timeout is a design where the health check itself takes 10 seconds when the system is unhealthy, which means the health checker fires concurrent probes and can become the load that takes the system down.
Failure stories worth testing
Make the database 10x slower
Confirm readiness goes false, liveness stays true, and replicas are not restarted. This is the single most important test in this article.
Deadlock the main thread
Liveness should fail and the container should restart. If it does not fail, your liveness check is not actually detecting wedging.
Start an instance with an unreachable dependency
It must report not ready, not crash-loop, and not accept traffic. Then bring the dependency up and confirm it becomes ready without a restart.
Kill the network between a node and the registry
The node must stop working once its lease renewal fails. If it keeps serving, you have a split-brain waiting to duplicate work.
Freeze a node for 60 seconds, past three heartbeats
The orchestrator replaces it. Then confirm the frozen node cannot resume and write to shared storage, because otherwise the replacement was not a replacement.
Make a cache unavailable
Confirm the instance stays ready and serves degraded responses, if it can. If it goes unready, you have cascading health and a fleet-wide removal from rotation.
Introduce a startup delay
Confirm the instance is not in rotation until it is actually ready, and that this eliminates the error spike at the start of a rollout.
Send a request during graceful shutdown
It must succeed or fail cleanly, never hang. This is the preStop-plus-readiness-failing sequence and it is worth testing with a slow endpoint in particular.
Oscillate a dependency for two seconds, repeatedly
Confirm readiness hysteresis prevents flapping. Flapping readiness is worse for load balancers than steady unavailability because every transition churns connections.
Run 5,000 health probes in a tight loop
Confirm the probes do not become the load. If they do, the endpoint is doing work per request and it should read a flag.
A production-ready architecture
orchestrator / load balancer / registry
| | |
liveness readiness heartbeat
| | |
v v v
+--------------------------------------------+
| /live /ready /healthz /metrics |
| |
| /live |
| no dependency checks |
| watchdog on main loop progress |
| or kernel-level socket check |
| -> 200 or timeout (restart) |
| |
| /ready |
| startup complete? |
| draining? (fails FIRST on shutdown) |
| at concurrency ceiling? |
| critical dependency OK? |
| hysteresis on transitions |
| -> 200 or 503 (remove from rotation) |
| |
| /healthz |
| deep, for humans and dashboards |
| no timeout budget smaller than |
| the timeout that made it slow |
+--------------------------------------------+
background check thread (1 per instance)
every 2s: probe dependencies with a
1s timeout, set flags, never block /ready
heartbeat: lease renewal every 10s
failure -> stop serving AND exit
(never renew-and-continue)
A sensible delivery checklist:
- Liveness checks nothing but wedging. No database, no cache, no downstream call, no allocation worth measuring.
- Readiness is thorough, because a readiness failure is cheap and a liveness failure is not.
- Dependencies in readiness, never in liveness. Write it in the health check policy so nobody has to remember it.
- A background thread probes dependencies with a short timeout and sets flags; the HTTP handler reads a flag. Never a synchronous deep check on a hot path.
- Fail readiness before shutdown, not after. The endpoint must go unready while in-flight work drains.
- Add hysteresis to readiness transitions in both directions, and pick the counts deliberately.
- Make heartbeat lease loss fatal to the process. A node that cannot renew must exit, not retry.
- Ensure the shared resource enforces identity, so a node that lost its lease cannot write after being replaced.
- Keep health endpoints out of the load-balancer’s request accounting and out of the same pool as real traffic.
- Report degraded separately from unready. A fleet that can serve from cache should not be taken out of rotation.
- Set the check intervals from a cost calculation — probe rate times instance count — not from a default in a framework.
- Keep a deep health endpoint for humans that is allowed to be slow, and never let a service’s own restart logic call it.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| Liveness checks the database | A database slowdown restarts every replica at once | Liveness checks nothing but wedging |
| Liveness and readiness are the same handler | One slow dependency becomes a restart storm | Separate endpoints, separate semantics |
| Deep synchronous readiness with a long timeout | The probe itself becomes the load during an outage | Background thread, short timeout, flag reads |
| Readiness fails on every transient blip | Flapping churns load balancer connections | Hysteresis on transitions |
| Heartbeat failure logged, then ignored | Two instances, one identity, both working | Lease loss is fatal to the process |
| No fencing on the shared resource | A replaced instance keeps writing | Resource enforces monotonic identity |
| Cascading health across the dependency graph | Whole fleet leaves rotation for one slow component | Report degraded, stay ready |
| Health endpoint in the production request pool | Probes consume user capacity | Separate handling, excluded from accounting |
| Health check interval from framework defaults | 200 unnecessary probes per second at scale | Derive the interval from cost |
| Liveness handler that allocates | Periodic GC pressure in every replica | Read a flag, return immediately |
| Unready only at process exit | In-flight requests get connection resets | Fail readiness first, drain, then exit |
| Probe timeout larger than the health checker’s own | Health checking times out during the incident | Probe timeout well inside the budget |
| No distinction between startup and readiness failure | Rollouts produce an error spike | Separate startup gate from steady state |
| Monitoring the health endpoint’s own availability | A useful signal nobody can act on | Alert on the thing operators can change |
| Assuming /healthz proves the service works | Deep check green while users see errors | Alert on user-visible error rate too |
The complete story in one minute
Three mechanisms, three questions. Liveness asks whether the process is wedged and needs a restart. Readiness asks whether this instance should receive traffic. A heartbeat asks whether the instance is still there at all.
The one rule that prevents the most damage: dependencies go in readiness and never in liveness. A liveness check that queries the database means a database slowdown fails liveness everywhere, every replica restarts at once, the connection storm from cold starts makes the database slower still, and the whole system goes down for a dependency that was only slow. Readiness failing is cheap — the instance leaves rotation and keeps working. Liveness failing is not, and it should be reserved for the one thing it can safely detect: a process that is alive but no longer progressing.
Keep both checks cheap. A background thread probes dependencies every couple of seconds with a short timeout and sets a flag; the HTTP handler reads the flag. A synchronously-checked health endpoint with a ten-second timeout becomes the load that takes the system down, because every probe now takes ten seconds when the system is unhealthy.
A heartbeat is not a health check. It fails in the opposite direction, and missing one means something the instance cannot fix by itself. That is why lease loss must be fatal to the process: a node that cannot renew its lease must exit rather than retry, and the shared resource must enforce identity so a node that lost its lease cannot write after being replaced.
And resist cascading health. Unready means “do not send traffic”, so making readiness depend transitively on a slow database removes the whole fleet from rotation at once — the instances that could still serve cached data included. Report degraded, stay ready, shed load deliberately.
/live -> restart me, and only if I am wedged
/ready -> send me traffic, and only if I can serve it
lease -> renew or die, never renew and continue
The hard part was never writing a handler that returns 200. It was making sure the check that is allowed to kill a process is the one that has no way to be affected by anything outside that process.


