← All writing
articleMar 18, 202419 min read

Container Orchestration: The Scheduler Is Easy, the Controller Is the Product

Two-phase scheduling, level-triggered reconciliation, the pod lifecycle, and why a Kubernetes-style control loop is the reusable idea underneath every orchestrator.

KubernetesSchedulingControl LoopsArchitecture
Container Orchestration: The Scheduler Is Easy, the Controller Is the Product cover illustration

A container orchestrator does two things that people constantly conflate: it decides where something runs, and it keeps reality matching a declaration of intent. The first is a scheduling problem and it is interesting for about an hour. The second is a control loop, it is the reason these systems self-heal, and it is the part that will occupy the rest of your career.

Getting this distinction right is what turns a bag of scaling tricks into an architecture you can reason about.

The scale, and the two numbers that matter

A cluster with five thousand nodes and one hundred thousand pods.

5,000 nodes x 110 pods/node (30 per core)
  = 550,000 pod slots at theoretical max
  realistically ~100,000 running

a rolling deployment of a 200-pod service
  = 200 pods, each with a start, a health check,
    a readiness gate, and a shutdown sequence

state churn per rolling deploy of 200 pods
  200 pod creates
  200 status updates (pending -> running -> ready)
  200 lease renewals, every 10s
  ...
  = well over 1,000 API writes per deployment,
    seconds apart, and there are always several
    deployments in flight

Two numbers drive the design.

The registry is read-heavy and write-heavy at the same time. Controllers list pods, watch pods, and write status constantly. A scheduler that lists all pods on every placement decision is O(pods) per decision and is the classic way to build an orchestrator that collapses under load. This is why informers and a scheduler cache exist at all.

A rolling deployment is a distributed transaction you did not write. The system must start new pods before killing old ones, wait for health, and roll back if the new version is bad. Every one of those steps is a separate API write that can fail independently, and the recovery from a partial failure is where the design is tested.

Everything is a declared desired state

The core abstraction is a pair of states:

spec (desired)  what the user asked for
status (observed) what is actually true

A controller’s entire job is to make status converge to spec, and the loop is the same shape in every case:

    +----------------------------------------+
    |  1. observe: read status                |
    |  2. diff:   status vs spec              |
    |  3. act:    do the smallest thing       |
    |              that moves status          |
    |  4. wait:   for the next event, or      |
    |              after a resync period      |
    +----------------^-----------------------+
                     |  (loop forever)

The word that makes this work is level-triggered. A level-triggered controller acts on what is true now, not on what changed since last time.

level-triggered (Kubernetes)
  "there should be 3 replicas running right now"
  -> count actual replicas
  -> if 1, create 2

edge-triggered (a naive event handler)
  "I received a 'pod died' event, so create a pod"
  -> if the event is lost or the create fails, nothing
     ever corrects it

This is not a stylistic preference. It is the property that makes the system survive missed events, dropped connections, controller restarts, and clock skew. An edge-triggered design needs every event to be delivered exactly once, in order, forever. A level-triggered design only needs the current state to be readable, and it repairs itself from any starting point.

The corollaries fall out:

  • Controllers must be idempotent. Running the loop twice must be safe. This rules out append-only side effects in the reconcile path, which is why the pattern is “declare the desired state of a child object” rather than “call the cloud API to create a machine.”
  • Status is a subresource. Controllers write status, users write spec, and the two do not overwrite each other. Without that split, a controller that rewrites the whole object clobbers user intent.
  • Resync is mandatory. Even with watches, a periodic full resync bounds how long a controller can be wrong if an event was somehow missed. Watches are an optimisation that reduces latency, not the source of truth.

The scheduler: propose, then bind

Placement is a constrained optimisation. You have a pod with requirements and a set of nodes with different properties, and you need to pick one that is feasible, and if several are feasible, pick a good one.

The two-phase design is what makes this safe.

        unscheduled pod
              |
              v
  +---------------------------+
  |  PHASE 1: filter          |  remove nodes that cannot run it
  |  - enough CPU?           |  resource fit
  |  - enough memory?         |  node selectors
  |  - taints tolerated?      |  taints and tolerations
  |  - affinity satisfied?    |  inter/intra-pod affinity
  |  - pod affinity/anti?     |  topology spread
  |  - PV available here?     |  volume topology
  +------------+--------------+
               |  viable nodes
               v
  +---------------------------+
  |  PHASE 1b: score          |  rank what survived
  |  - least requested        |  bin packing
  |  - least allocated         |
  |  - image locality         |  already have the image
  |  - topology spread        |  zone balance
  |  - node affinity pref     |  soft preferences
  +------------+--------------+
               |  ranked list
               v
  +---------------------------+
  |  PHASE 2: reserve + bind  |  optimistic concurrency
  |  write "pod scheduled    |  in the API server
  |   to node X" atomically   |
  |  if it fails, re-run      |
  +---------------------------+

The critical detail is that phase 1 and phase 1b have no side effects. Filtering and scoring read a cached view of the cluster and produce a ranked list. Nothing is reserved. That is what makes it safe to run them again.

Phase 2 writes the decision, and that write is where concurrency bites. Two schedulers must not both decide that node 47 is the right place for two pods that will not both fit. The write is conditional:

pod.spec.nodeName is still empty
  AND the node still has room
  -> bind: set pod.spec.nodeName = "node-47"
  -> the API server checks both conditions atomically
  -> if another pod claimed the space first, this
     write fails and the scheduler re-runs from phase 1

This is optimistic concurrency, and the retry rate is the metric to watch. A high bind-conflict rate does not mean the scheduler is broken; it means the cache is stale relative to the cluster, which usually means the scheduler is too slow or the API server is saturated. The fix is a faster scheduler or a fresher cache, never a lock.

The scheduler cache is the real scaling problem. A scheduler that queries the API server for every pod on every node on every decision is unusable. Real schedulers keep an informer-backed cache of all pods and nodes and do filtering and scoring entirely in memory. That means the scheduler’s view is slightly stale, which is exactly why phase 2 has to be a conditional write. The two decisions are coupled: the cache is an optimisation, and the conditional bind is what makes the optimisation correct.

Scoring has one more property worth knowing. leastRequested and leastAllocated are bin packing heuristics, and they exist to reduce the number of nodes with any capacity left. That is about cost, not about performance. Spreading pods thinly across many nodes for “balanced” graphs is a different objective and it costs money.

The API server is the bottleneck by design

Every component — kubectl, the scheduler, every controller, admission webhooks, aggregation layers — talks to one component: the API server. And the API server has exactly one real-time store underneath it: etcd.

        kubectl        controllers       webhooks
             \             |               /
              \            |              /
               +-----------+---------------+
                           |
                    +------v------+
                    |  API server |   authn/authz
                    |             |   admission chain
                    |             |   validation
                    |             |   version conversion
                    |             |   optimistic concurrency
                    +------+------+
                           |
                         etcd

The important design decision is optimistic concurrency on an API object. A client can submit the resourceVersion it read; an update based on a stale version is rejected with a conflict. Kubernetes does not expose a general transaction across several API objects, so multi-object workflows have to record progress and reconcile partial completion. This avoids holding application-visible locks through controller work, but it is not by itself the reason the API server scales.

The cost is that there are no transactions across objects. There is no way to atomically say “create these three pods.” Every multi-object operation has to be expressed as a sequence of individual conditional writes, and each one can succeed or fail on its own. This is the single most important constraint for anyone writing a controller: your reconcile loop is a series of individually-atomic steps against a system that will let you observe it half-finished, and it must be written to be correct from any intermediate state.

This is also why List is discouraged for large collections. A List is served from etcd’s watch cache, but the response is built in memory, and a list of a hundred thousand pods is a real problem for the API server’s memory. Paginated and chunked lists exist for this reason, and a controller that lists everything in one call is a controller that will take down the cluster.

Pods are disposable, and the API makes that clear

The pod abstraction is deliberately weak. A pod is a group of containers that share a network namespace and can share volumes. It has no identity beyond its name in a namespace, it is not addressable in a durable way, and it is expected to die.

pod lifecycle
  Pending       scheduled, images pulling, volumes attaching
  Running       all containers created, at least one running
  Ready         passed its readiness probe
  Terminating   received SIGTERM, grace period counting down
  (gone)        kubelet confirmed the containers are gone

The design is right about this, and the way it is right is instructive. The reason pods are disposable is that any stateful thing in a pod will be destroyed, and the system gives you two tools for that state: a PersistentVolumeClaim, which is deliberately outside the pod’s lifecycle, and no other one.

The common failure is putting a database’s data directory in an emptyDir volume because it is convenient during development. That works until the node dies, and the failure is total. The lesson is not “use a real database” — it is that a design should make the mistake hard. A platform that requires persistent state to be declared as a claim has made the ephemeral path the default and the dangerous path explicit.

Readiness and liveness are different probes and confusing them is an outage. Liveness says “restart this container.” Readiness says “stop sending traffic.” A slow database under load makes a service fail its readiness check; if that same check is wired to liveness, the orchestrator restarts every replica at once, which makes the database slower, which fails readiness, restarting more replicas. A positive feedback loop with a restart loop in it. The rule is that liveness should be as dumb as possible — “is the process wedged” — and anything that depends on another service belongs in readiness, not liveness.

Kubelet: a node agent with a loop

Kubelet is the piece that actually runs on each node, and it is the same control loop pattern applied locally.

  +------------------------------------------------+
  |  kubelet on node-47                            |
  |                                                |
  |  pod workers, one per pod                      |
  |    - sync pod from API (informer)              |
  |    - ensure containers running                 |
  |    - mount volumes                              |
  |    - manage probe results -> pod status        |
  |    - on termination: SIGTERM, grace, SIGKILL   |
  |                                                |
  |  static pod / mirror pods                      |
  |  image management, log rotation, GC            |
  +------------------------------------------------+

Two things follow from this being a controller rather than a daemon.

A kubelet restart does not lose state. It re-lists, re-syncs, and converges. Any cleanup that must happen on eviction is derived from observed state rather than from a sequence of in-memory actions.

The eviction path is where node problems become cluster problems. Kubelet monitors node resources and starts evicting pods when memory or disk crosses a threshold. A node under pressure evicts, the pods land on other nodes, those nodes come under pressure, and you have a cluster-wide cascade. Mitigations are priority classes that let important pods resist eviction, PodDisruptionBudgets to keep a minimum available during voluntary disruptions, and headroom so that losing one node’s worth of capacity is not already an emergency.

Deployments are a state machine wearing YAML

A Deployment does not do rolling updates. It maintains a number of replicas, and a separate controller turns a change in that number into new pods. The rollout itself is a control loop that is easy to get wrong and instructive to read.

  spec.replicas: 200 -> 400
        |
        v
  maxSurge=25%  -> allow up to 450 pods
  maxUnavailable=25% -> never drop below 150 available
        |
        v
  new ReplicaSet scaled up, old ReplicaSet scaled down
  in steps, waiting for readiness between steps
        |
        v
  if new pods fail readiness, the loop stalls
  progress stops; the old ReplicaSet is not scaled down
  -> the rollout halts rather than making things worse

The halting behaviour is the point. A rolling update is not a transaction; it is a controller that stops advancing when the world stops cooperating. maxUnavailable is a promise about availability during the rollout and maxSurge is a promise about cluster capacity, and setting them is the act of deciding which of those you are willing to trade.

The counterpart is the pod termination sequence, which is where most shutdown bugs live.

SIGTERM
  -> endpoint removed from service
  -> preStop hook runs
  -> SIGKILL after terminationGracePeriodSeconds

There is a race here that catches real systems. Pod termination, endpoint updates, load-balancer convergence, and application shutdown are coordinated by different components. Kubernetes marks terminating endpoints as not ready for ordinary traffic, but every proxy and external load balancer still has its own propagation and connection-draining behaviour. The application should stop accepting new work, drain in-flight requests inside terminationGracePeriodSeconds, and expose readiness honestly. A short preStop delay can cover a measured propagation gap, but a blind sleep is not a delivery guarantee and it consumes the same termination grace period as shutdown.

Failure stories worth testing

Delete the API server’s etcd

Nothing creates or updates. Existing pods keep running, because a running container does not need the control plane. This is the most reassuring failure in the platform and it is worth demonstrating once so people trust it.

Kill a node hard

The node controller notices the heartbeat stops, marks the node NotReady, and the pods are rescheduled elsewhere. Confirm the rescheduled pods can actually attach their volumes, because a pod that reschedules to a zone without its PV is a stuck pod and a much worse outcome than a delay.

Fill a node’s memory to 95%

Kubelet begins evicting. Confirm that eviction respects priority classes and that a critical service is not the first thing to die.

Roll out a version that fails readiness immediately

The rollout should stall with old pods still serving. Then confirm the rollback path. This test is the one that most teams have never actually run.

Restart every kubelet on a node at once

The node re-converges from scratch. The interesting question is how long it takes and whether it duplicates work, since a controller that is not idempotent will double-create here.

Saturate the API server with watches

Everything slows down together, because the API server is the shared dependency. Confirm the client-side rate limiters keep one misbehaving controller from starving everything else.

Create a Deployment with an impossible spec

An unschedulable pod with no events going stale. Confirm that unschedulable pods surface clearly in describe output, because a silent pending pod is the most expensive kind of “my deploy is stuck.”

Take a node out with a PodDisruptionBudget in place

Voluntary eviction is refused; involuntary eviction is not. Confirm the distinction is understood by whoever is on call.

Roll a deployment during a node failure

The new pods land on a different set of nodes than the old ones, and the surge limit interacts with the reduced capacity. This is where maxSurge and capacity planning meet.

A production-ready architecture

   clients (kubectl, CI, controllers)
              |
              v
  +---------------------------+
  |  API server (n replicas)  |  authn -> authz -> admission
  |  etcd behind it           |  -> validation -> conversion
  |  optimistic concurrency   |  resourceVersion on every write
  |  watch cache              |  list served from memory
  +------------+--------------+
               |
    +----------+-----------+------------+
    |          |           |            |
    v          v           v            v
 scheduler  controllers  admission   aggregation
 (cached,   (one per     webhooks     (metrics,
  bind via   resource)                  custom)
  conditional
  write)
               |
               v
  +---------------------------+
  |  etcd cluster (3 or 5)    |  linearizable, watched
  |  single source of truth   |  by every informer
  +------------+--------------+
               |
      watches, patch, status
               |
               v
  +---------------------------+
  |  kubelet on each node     |  pod workers, one loop each
  |  + CRI -> containerd      |  run / stop / probe
  |  + cni -> pod network     |  volume mount, GC
  |  + eviction manager       |
  +---------------------------+

  every controller:
    informer cache -> reconcile -> status write
    level-triggered, idempotent, resync

A sensible delivery checklist:

  1. Every controller you write is level-triggered, idempotent, and safe to interrupt at any point. If it needs to know what changed, that is a design smell.
  2. Cache everything with an informer; do not list the API server inside a reconcile loop.
  3. Use a conditional write for anything where two writers could race, and treat the failure as normal flow rather than an exception.
  4. Read from a cache for decisions, but make the write that commits the decision conditional and on the API server. That pair is the pattern.
  5. Schedule as an optimisation over a stale cache, never as a source of truth.
  6. Set maxSurge and maxUnavailable deliberately for every workload, and remember they are an availability-versus-capacity decision.
  7. Test termination through every proxy and load balancer. Add a bounded preStop delay only for a measured propagation gap, leaving enough grace time for the application to drain.
  8. Wire liveness to the dumbest possible check and put every external dependency in readiness.
  9. Declare persistent state as a claim. Do not make an emptyDir a supported way to keep data.
  10. Set priority classes and PodDisruptionBudgets for anything whose loss takes the cluster with it.
  11. Watch the metric that matters: bind conflicts, reconcile rate, queue depth per controller, and API request latency. A controller with a growing queue is the first sign of trouble.
  12. Assume the API server is a shared dependency and rate-limit clients accordingly, so one bad controller cannot starve the rest.

Common mistakes

Mistake What actually happens Better decision
Edge-triggered reconcile A lost event is never corrected Level-triggered, act on current state
Non-idempotent reconcile Restarting the controller duplicates work Same input, same result, every time
Skipping resync A missed event leaves the system wrong forever Periodic full resync, watch as latency only
Scheduling from live API queries O(pods) per decision, unusable at scale Informer cache, conditional bind to stay correct
Unconditional bind write Two schedulers place pods that will not both fit Conditional write on nodeName empty plus capacity
Treating bind conflicts as errors Correct-by-construction behaviour looks like a bug Retry; watch the conflict rate as a cache-staleness signal
Serialising multi-object operations No atomicity across objects, and a partial failure is possible Assume any step can fail and reconcile from any state
Large unpaginated lists API server memory blows up, cluster stalls Paginate, chunk, cache
Liveness probe that calls a dependency Restart loop, positive feedback, cluster-wide outage Liveness dumb, readiness thorough
Untested termination path Rollouts reset or strand in-flight requests Drain, verify endpoint propagation, then add only the delay you measured
Durable state in a pod It is deleted eventually, without warning PersistentVolumeClaim, declared explicitly
No priority classes Eviction picks victims by luck Classify and protect what matters
Assuming schedulable capacity is free capacity A node failure becomes an immediate emergency Headroom for a full node’s worth of load
Many API server replicas with one etcd The replicas are redundant but the store is one thing Size etcd properly; it is the real bottleneck
Controller with an unbounded list in its loop One controller starves the whole cluster Cache, rate-limit, and instrument queue depth
Watching instead of writing status Observed state never converges Reconcile must write status back
Treating NotReady as “not running” A node with a network partition cascades evictions Understand the eviction thresholds and priorities
Optimising for balanced graphs Bin-packing works better and costs less Decide the objective, then pick the score

The complete story in one minute

A container orchestrator has two halves and they deserve different amounts of attention.

Scheduling is a constrained optimisation. Filter nodes that cannot run the pod, score the ones that can, and pick the best. The design that makes it work is splitting that into propose and bind: filtering and scoring read a cache and have no side effects, and only the final bind writes to the API server, conditionally. The conditional write is what lets you schedule against a slightly stale cache and still be correct, and the bind-conflict rate is a staleness metric, not an error count. Everything happens in memory, and bin-packing heuristics exist to reduce cost, not to look pretty.

The controller half is the product. Desired state versus observed state, act on the difference, loop forever, level-triggered rather than edge-triggered. Level-triggered is the whole trick: it is why a missed event, a controller restart, or a dropped connection repairs itself instead of leaving the system permanently wrong. It is also why every controller must be idempotent, and why status is a subresource so a controller cannot clobber user intent.

The API server and etcd underneath it impose the constraint everything else bends around. Optimistic concurrency on individual objects, no transactions across objects. Every multi-step operation is a sequence of individually atomic steps that can be observed half-finished, and any controller must be correct from any intermediate state. That is also why informers exist: the watch cache makes reads fast enough that nobody needs to list inside a reconcile loop.

Pods are disposable by design, and the API makes the durable path explicit rather than implicit. Liveness and readiness are separate because wiring an external dependency into liveness creates a restart loop that a slow database turns into a cluster-wide outage. A rolling deployment is a control loop that halts rather than proceeding when the new version is unhealthy, which is the right behaviour and which nobody believes until they have watched it happen.

spec vs status -> reconcile -> status
informer cache -> filter + score in memory -> conditional bind
kubelet -> the same loop, locally

The hard part was never placing a pod on a node. It was writing controllers that are correct from any state, including the states where a previous attempt died halfway.

Technical references

Keep reading
Browse everything