← All writing
articleJun 14, 202418 min read

Blue-Green Deployment: The Deployment Where Your Rollback Plan Is a DNS TTL

Two identical environments, database migrations that break the contract, and the failure modes that turn a supposedly-safe switchover into an outage.

DeploymentRelease EngineeringDatabasesArchitecture
Blue-Green Deployment: The Deployment Where Your Rollback Plan Is a DNS TTL cover illustration

Blue-green is described as the safe deployment strategy: run two identical environments, deploy to the idle one, test it, then flip a switch. Rollback is flipping the switch back. The pitch is that downtime is zero and risk is near zero.

The pitch is right about the application and mostly wrong about the system, because the interesting parts are the database, the switch itself, and the fact that “identical” stops being true the moment you have connections, caches, and third-party state.

The shape of it

        +-----------------------+
        |  load balancer / DNS  |
        |  "where does prod go"  |
        +-----------+-----------+
                    |
          +---------+---------+
          |                   |
      live colour          other colour
          |                   |
    +-----+------+      +-----+------+
    | blue      |      | green      |
    | version 4 |      | version 5  |
    +------------+      +------------+

  deploy:  update green, verify, switch
  rollback: switch back
  in theory: instant, zero downtime

Everything hard about blue-green is in the three words “verify” and “switch”. Deploying to green is a normal deployment and inherits all the normal problems. The switch is a change to a system that is currently serving every request you have. And “blue” is not actually identical to “green” the first time you try it, because green is cold.

The database is the whole problem

Say version 5 renames a column.

-- version 4
CREATE TABLE accounts (
  id    uuid PRIMARY KEY,
  name  text,
  email text
);

-- version 5
CREATE TABLE accounts (
  id         uuid PRIMARY KEY,
  full_name  text,        -- was "name"
  email      text
);

The switch looks instant. It is not, because there is state you did not switch.

  t=0    apply migration to the database
         -> column "name" is gone
         -> blue is still serving version 4
         -> every version 4 query now fails
         -> switch traffic to green
         -> green works
         -> "successfully" deployed with an
            outage in the middle

  t=2min realise something is wrong
         -> switch back to blue
         -> blue is version 4 and the column
            it needs no longer exists
         -> rollback does not work either

Both versions are now broken and the database is in a state neither understands. This is not a hypothetical migration mistake; it is what a rename does by default, and renames are routine.

Expand and contract: the rule that makes this safe

The fix is to stop doing breaking migrations in the same release as the code that stops using them, and to split every schema change into phases that are each independently compatible.

  release N     expand
    add the new column, nullable
      ALTER TABLE accounts ADD COLUMN full_name text;
    old code ignores it, new code can read it

  release N+1   dual write
    code writes both name and full_name
    reads still prefer "name"
    backfill existing rows

  release N+2   switch reads
    code reads full_name, falls back to name
    writes both
    -> old code still works: it reads "name",
       which is still being written

  release N+3   drop
    only now: ALTER TABLE accounts DROP COLUMN name;
    -> safe only when no running version
       reads "name" at all

The compatibility target this gives you is the one blue-green actually needs:

Every version that can still receive traffic works against every schema phase that can coexist with it.

Do not generalize that into “any version, any schema, any order.” Each phase has a compatibility window, and the deployment controller must know when an old binary is no longer eligible. Forward-fix schema evolution plus reversible traffic routing is usually safer than assuming a destructive DDL rollback will restore data.

Two rules that go with it:

Never make a destructive change in the same release as the code that stops using the column. If a release both stops reading a column and drops it, a rollback of that release is impossible. Separate them by at least one release, and know which running version is in each environment.

Backfills are not migrations. A backfill of a million rows is a long-running operation that competes with production traffic, takes locks, and cannot be rolled back when it is half done. It runs as a job, in batches, with its own observability, and the code must work correctly with the backfill partially complete — which means dual write from the start, so there is no window where a row is missing from the new column.

The switch, and why it is the risky part

Once the application is safe, the switch is a write to the load balancer or a DNS record, and it has properties worth being precise about.

  L7 load balancer (target group swap)
    - one API call, takes effect in ~1-5s
    - existing connections may persist to blue
    - reversible in the same time

  DNS change
    - TTL 30s-5min
    - clients cache, so the change is gradual
    - not instant, and not uniformly so
    - the rollback is another TTL

This matters more than it sounds. A DNS-based blue-green with a five-minute TTL does not have instant rollback; it has a five-minute window during which a bad release is serving some fraction of traffic, and a rollback that takes another five minutes to fully apply. During a bad deploy that is ten minutes of errors. If you want seconds, you need something in the request path that can be switched, which means a load balancer or a service mesh routing rule rather than DNS.

Other things that break the clean switch:

Connection draining. Established TCP connections and, worse, long-lived WebSockets and streaming responses do not move when you switch. Blue keeps serving existing connections for as long as they last. If version 4 has a long-lived connection protocol, “switching back” does not bring those connections to version 5, and the rollback is partial in a way that is very hard to reason about.

Sticky sessions. If the load balancer is configured for session affinity and you switch target groups, existing sessions break. Either the switch handles it or your sessions are stateless, and you want the second.

Cache warming. Green is cold. Its local caches, its JIT, its connection pools, and any connection to a third-party API with a cold-start penalty are all starting empty. This is a real problem, and it is why blue-green is a strategy for warm green: pre-warm green before switching, or accept an error and latency spike that is indistinguishable from a bad release.

Asymmetric traffic during the switch. For the seconds the switch takes, some requests go to blue and some to green. If blue and green are not behaviourally identical — different config, different feature flags, different dependency versions — you have both versions of your system running simultaneously in production, which is the thing the strategy was supposed to prevent. This is the argument for the database work above: if both versions are correct against the current schema, asymmetric traffic is fine.

The pre-switch verification, done properly

“Verify green” needs more than a health endpoint. Readiness is still valuable—it proves the instance should receive traffic under the probe’s contract—but it cannot prove changed business behavior, peak capacity, or side-effect isolation.

  functional
    - run the real smoke suite against green,
      against production-shaped dependencies
    - a synthetic transaction that exercises
      the code path you actually changed

  load
    - drive green to your peak traffic profile
      while it is still not receiving user traffic
    - a cold environment that has never seen
      peak load is the most common cause of a
      failed switchover

  data
    - confirm the schema version green expects
    - confirm the backfill completed or the
      fallback path is correct
    - verify green is not writing to anything
      that only blue can read

  side effects
    - third-party integrations: does green
      have its own credentials, its own queue,
      its own webhook registrations?
    - a green environment writing to the same
      third-party account as blue is a
      double-processing bug that looks like
      a successful deployment

That last one is the most dangerous and the least commonly checked. Any external system that identifies you by credential, by hostname, or by queue is now receiving traffic from two environments, and if the deploy changes behaviour that produces side effects, both are producing them. Give green its own sandbox credentials for verification and switch them at the same time you switch traffic, or do not switch traffic at all.

What blue-green actually buys, honestly

The genuine wins:

  • Rollback is a config change, not a build-and-deploy. If the failure is a code bug, recovery is seconds instead of the fifteen to forty minutes a forward fix takes. That is worth a lot.
  • Green is a real pre-production environment. It is the same size, the same config, the same code path as production, which is a materially better test than a staging box with a synthetic dataset.
  • The desired route can change atomically. Existing connections, propagation through proxies, retries, and load-balancer convergence can still produce an overlap window where both versions serve work.

The costs, honestly:

  • Two viable environments during the cutover. Keeping both permanently at full peak capacity is the simplest and most expensive implementation; temporary provisioning or pre-switch scaling trades money for slower rollback readiness.
  • Database state is shared, which is the whole problem. You have the rollback benefit of blue-green and none of the isolation, because the thing that breaks is the thing you cannot switch.
  • Stateful services are awkward. Anything with persistent local state, WebSockets, or a warm cache makes the switch partial.
  • Config drift between colours is a real failure mode. Blue and green must be configuration-identical apart from the version. A difference discovered after the switch is an outage caused by your deployment system.

When it is worth it: a service where a bad release is expensive and the traffic is not so stateful that the switch becomes partial. Payments, signup flows, anything with a rollback that takes minutes. When it is not: a stateless internal service behind a mesh, where a canary or a rolling update gives you most of the safety at a fraction of the infrastructure cost.

Failure stories worth testing

Rename a column and deploy

Confirm blue breaks. Then do it properly with expand and contract and confirm blue keeps working through every phase. This test is the entire reason the pattern exists.

Roll back after a destructive migration

Confirm that rolling back is impossible and that you knew it in advance. Then confirm the expand and contract sequence has no point where a rollback breaks.

Make green’s health check fail after the switch

Confirm the switch back is a single action and takes effect in seconds. Measure the real time, not the intended time.

Break the code so it 500s on 5% of requests

The realistic rollback test. Confirm the detection is fast enough that the error budget survived, and measure how long each request served the bad version.

Leave blue running with production traffic and a sticky session

Confirm what happens on switch. If sessions break, the strategy is not compatible with your session model.

Switch with a cold green

Confirm the error and latency spike, and that it is distinguishable from a genuinely bad release. This is the case where people roll forward into a worse version.

Give green the same third-party credentials as blue

Confirm the double side effects. Both environments writing to the same payment provider or queue is a data corruption bug that the deployment system caused.

Run a long backfill during the switch

Confirm what happens to the lock and to production latency, and confirm the new code is correct with the backfill half done.

Switch while a rolling update is in flight on blue

Confirm the two deployment systems do not fight. Two controllers writing the same target group is an outage that looks like a config error.

Scale green to zero after a failed verification

Confirm the rollback is not complicated by green having been resized. This is trivial to get wrong and awkward to debug.

A production-ready architecture

   pipeline
      |
      v
   +------------------+
   |  build once      |  the artefact is identical
   |  promote the     |  in both colours
   |  same image      |
   +--------+---------+
            |
      +-----+------+
      |            |
      v            v
   +-----+    +-----+
   |blue |    |green|    same size, same config,
   | v4  |    | v5  |    different image tag
   +--+--+    +--+--+
      |          |
      |  pre-warm green: caches, pools,
      |  JIT, sandbox credentials
      |  run the smoke suite + peak load test
      |          |
      |          v
      |    +----------+
      |    |  verify  |  functional, load, data,
      |    |          |  side effects
      |    +-----+----+
      |          |
      |     pass | fail -> fix, or abandon green
      v          v
   +---------------------------+
   |  switch                  |  target group swap, seconds
   |  - drain blue            |  NOT DNS with a long TTL
   |  - point traffic at green|
   +-------------+-------------+
                 |
                 v
   +---------------------------+
   |  database                |  shared by both colours
   |  expand / contract only  |  any version works against
   |  forward-only migrations |  any reached schema state
   +---------------------------+

   rollback: switch back. seconds, not a rebuild.

A sensible delivery checklist:

  1. Every schema change is expand and contract, split across at least two releases, and backward compatible at every point.
  2. No release both stops using a column and drops it. Know which running version reads which column, and keep that knowledge current.
  3. Migrations are forward-only. Rollback moves the code, not the schema.
  4. Backfills are jobs, in batches, with dual write running first so that partial completion is safe.
  5. The switch is a load balancer target change, not a DNS change, unless you have accepted a multi-minute rollback window.
  6. Drain existing connections on the old colour and confirm long-lived connections do not prevent a clean rollback.
  7. Pre-warm green before switching: caches, connection pools, and anything with a cold-start cost.
  8. Verification includes peak load against green, not a health endpoint.
  9. Green gets its own third-party credentials for verification, and they are switched with the traffic or not at all.
  10. Blue and green are configuration-identical apart from the image. Drift between them is a deployment outage waiting to happen.
  11. Lock down concurrent deployment controllers so two switches cannot interleave.
  12. Document the exact rollback action and its measured time. A rollback plan nobody has executed is a plan that does not work.

Common mistakes

Mistake What actually happens Better decision
Rename a column and deploy together Blue breaks, and the rollback breaks too Expand and contract across releases
Treat the migration as reversible Rollback of a destructive migration is impossible Forward-only schema, backward-compatible code
Use DNS with a long TTL for the switch Rollback takes minutes, not seconds Load balancer target change
Verify only with a health endpoint A correctly-implemented health check proves nothing Smoke suite plus peak load against green
Switch to a cold green Error and latency spike, indistinguishable from a bad release Pre-warm before switching
Health check on green is the gate A green-only failure mode is invisible Real transactions on production-shaped data
Blue and green share third-party credentials Double writes, double charges, double emails Separate credentials, switched with traffic
Config drift between colours Asymmetric traffic, undefined behaviour Enforce identical config except the image
Ignore long-lived connections Rollback is partial and hard to reason about Drain, or use stateless protocols
Two deployment controllers on one target group The switch fights itself One controller, serialised deploys
Assume blue-green includes the database The only stateful thing is the thing you cannot switch Expand and contract, deliberately
Run the backfill during the switch Lock contention, latency spike, partial data Backfill first, dual write during code deploy
No drain step In-flight requests hit a closing socket Explicit drain with a window
Sticky sessions plus target swap Every session breaks on switch Stateless sessions, or affinity that follows the switch
Roll back without measuring Discovered that recovery takes four minutes Time the rollback in a drill
One colour only in staging Staging tests a different system than production Staging exercises the same two-colour mechanism
Consider it cost-free Two environments is a real, permanent line item Justify it against rollback cost
No feature-flag independence Colours must be identical, which limits the strategy Flags, or accept the restriction
“It is just a config change” Untested until the incident Run the switchover drill quarterly

The complete story in one minute

Blue-green is two identical environments, a deploy to the idle one, verification, and a switch. Rollback is a switch back, which is genuinely valuable when a bad release is expensive: recovery in seconds rather than a forward fix.

The database is where the strategy meets reality. Rename a column in the same release that stops using it and blue breaks the moment the migration lands, and then rolling back does not work either because the column is gone. The fix is expand and contract: add the new column, dual write, backfill, switch reads, and only then drop the old one — spread across at least two releases. The property that buys blue-green is that any running version of the code works against any schema state reached that way, so the code can move back and the schema only moves forward.

The switch itself is the risky part. DNS with a five-minute TTL does not give you instant rollback; it gives you a five-minute window of bad traffic and another five minutes to undo it. Use a load balancer target change. Then handle the things that make the switch partial: long-lived connections that do not move, sticky sessions that break, and a cold green whose cache and connection pools have never seen peak load.

Verification has to be more than a health check, because a correctly-implemented health endpoint is green almost by definition. Run the real smoke suite, drive green to peak load, confirm the schema and backfill state, and check the side effects. That last one is the dangerous one: if green uses the same third-party credentials as blue, both environments are writing to the same payment provider, and the deployment has caused data corruption while reporting success.

The costs are honest: double the infrastructure, permanently, and no isolation for the database — the one thing you most wanted to switch.

blue <-> green, same artefact, same config
switch on a load balancer, not DNS
schema: expand, dual write, backfill, switch reads, drop
rollback = switch back, in seconds

The hard part was never flipping a switch. It was defining the compatibility window across every shared dependency, then proving routing, connection draining, and rollback while both colors were alive.

Technical references

Keep reading
Browse everything