Chaos Engineering Platform: Breaking Things on Purpose Is a Measurement Problem
Steady-state hypotheses, blast radius controls, abort conditions, and the discipline that separates experiments from outages with extra steps.

Chaos engineering gets described as “breaking production on purpose”, which is an accurate summary of the mechanics and a completely misleading summary of the intent. The intent is to find out whether your assumptions about failure behaviour are true, and to do it in a way that produces an answer rather than an incident.
The difference between the two is entirely in the preparation. A platform that can inject arbitrary faults anywhere at any time is a disaster generator. A platform that can only run approved experiments against a declared hypothesis within a bounded scope is an instrument.
The scale, and why chaos is usually done by a handful of people
This is worth stating because it recalibrates expectations. Chaos engineering is not a high-volume activity.
2,000 services
realistically 2-10 experiments per week
50 experiments per quarter
each experiment:
- a written hypothesis (30 min to hours of design)
- a scope definition and an abort condition
- a review before running
- 5-60 min of runtime
- a writeup
the ratio is roughly 1 hour of preparation
per minute of experiment
That ratio explains why the platform matters more than the technique. If experimenting costs an hour of preparation, an organisation does ten of them a quarter and its reliability work is driven by anecdotes. If the platform handles hypothesis templating, scope enforcement, abort conditions, and result capture, the cost per experiment drops and the volume rises by an order of magnitude.
It also means the platform’s job is mostly to make the safe path easy and the unsafe path impossible, not to make fault injection powerful.
The hypothesis is the experiment
hypothesis
"while serving 2x normal traffic, if the
recommendation service returns 500s,
then the checkout page will still complete
because it falls back to the cached list,
and p99 will stay under 800ms"
NOT a hypothesis
"let's see what happens if the
recommendation service goes down"
the difference
- measurable (a specific metric, a threshold)
- falsifiable (it can be wrong)
- about user-visible behaviour, not
about whether something breaks
The three properties in that example are what make it an experiment:
Measurable. There is a number, and a threshold, decided before the experiment. “Still works” is not measurable. “Completes at p99 under 800ms” is.
Falsifiable. The hypothesis has to be able to come out wrong, and it has to be wrong sometimes. A hypothesis that is always confirmed is not a hypothesis, it is a demonstration, and it teaches nothing. The platform should be tracking a hit rate, and a team with a 100% success rate is either running trivial experiments or not measuring properly.
About outcomes, not mechanisms. “The fallback path still works” is a claim about user-visible behaviour. “The circuit breaker opens” is a claim about an internal component, which is useful to confirm but is not what you are protecting. The strongest experiments are written in terms of what the user gets.
There is a fourth property that is easy to skip and that experienced practitioners insist on: the hypothesis should say what you expect and what would surprise you. If nothing in the experiment could surprise you, you have written a test for a behaviour you already believe, and the belief is untested.
Making it safe: the platform’s actual job
Everything below is a control that makes an experiment safe. If any one of them is optional, the platform is not finished.
+---------------------------------------------------------------+
| chaos platform |
| |
| +------------------+ +----------------------------------+ |
| | experiment spec | | scope controls (ENFORCED) | |
| | - hypothesis | | - service allowlist | |
| | - steady state | | - blast radius: max N pods | |
| | - fault | | - traffic percentage | |
| | - abort condition| | - environment allowlist | |
| | - blast radius | | - max duration (hard cap) | |
| +--------+---------+ +----------------+-----------------+ |
| | | |
| v v |
| +------------------+ +----------------------------------+ |
| | approval | | safety net (AUTOMATIC) | |
| | - requires signoff| | - abort on SLO burn | |
| | - review history | | - abort on alert firing | |
| +--------+---------+ | - auto-revert on timeout | |
| | | - kill switch, one click | |
| | +----------------------------------+ |
| v |
| +------------------------------------------------------+ |
| | experiment engine | |
| | injects the fault within scope, | |
| | records events, aborts on condition | |
| +------------------------------------------------------+ |
| |
| results -> steady state confirmed | falsified |
| -> writeup required, both outcomes |
+---------------------------------------------------------------+
The design points that matter:
Scope is enforced by the platform, not by the operator’s discipline. The experiment declares a service allowlist, a maximum number of pods, a maximum traffic percentage, and a hard duration cap. The engine refuses anything outside those bounds. An operator who can inject a fault into 100% of production traffic does not need a discipline; they need a permission model, and the two are not equivalent.
Abort conditions are automatic and independent of the fault. If the experiment triggers any customer-facing alert, or burns SLO budget faster than a threshold, the engine reverts the fault without waiting for a human. The person running it may not be watching, may be asleep, or may be the one who wrote the faulty code. The abort has to be a machine decision.
A hard duration cap exists and is enforced by the clock, not by the fault’s own lifecycle. A fault that hangs around because a cleanup step failed is exactly the failure you cannot afford. The platform turns it off at the deadline regardless of what the experiment believes about its own state.
A kill switch is one action and works immediately. Global halt, not per-experiment. This is the control that determines whether the team can run experiments at all, because it is the difference between a platform people use and a platform people are afraid of.
Non-production is not a substitute. The whole premise is that production has properties staging does not: real traffic shape, real data volume, real configuration, real dependencies. A chaos platform that can only run in staging has missed the point. It is a fault injection tool, not a chaos platform.
What to break, in order
The ordering matters enormously for both safety and value.
First: timeouts and retry behaviour. A dependency that hangs instead of failing fast is one of the most common causes of cascading failure, and almost nobody has tested it because it is harder to inject than a 500. Injecting latency rather than errors is often more valuable than injecting errors.
Second: retry budgets. Set a dependency’s retry count absurdly high and watch what happens to the dependency. This routinely reveals that the retry amplification you calculated is worse than you calculated, and that the circuit breaker is not where you thought it was.
Third: backpressure and queue limits. Fill a queue faster than it drains, or set a timeout so aggressive that legitimate traffic is rejected. This tests the overload path, which is the one that runs during incidents and is the one least covered by tests.
Fourth: single-instance failure. Kill a pod, a node, or one instance of a service. Boring, essential, and it finds broken readiness logic more often than anything else.
Fifth: dependency degradation. A slow or partially failing dependency, at a rate you control. This is where fallback paths get validated, and fallback paths are almost always wrong until tested.
Sixth: partial network failure. Latency, packet loss, and connection resets between two components. Requires more care to scope, and is where the interesting multi-hop behaviour lives.
Last: zone or region loss. The highest-value experiment and the highest-consequence one. It requires a documented, tested failover story before you start, and the prerequisite is a real dependency on the thing that fails over. Many teams discover during this experiment that their regional failover requires a manual step nobody knew about.
The order is a risk gradient, and skipping it is how people get a reputation for chaos engineering and a ban from the SRE team.
Scheduling, exclusions, and honest measurement
the confound
chaos experiment at 14:00 Tuesday
...during the weekly traffic peak
...during Black Friday
...during a partner's incident
...while a large deploy is rolling out
the result is uninterpretable
Three controls make a result mean something:
Exclusion windows. No experiments during peak trading, during a major campaign, during a freeze, or when any major incident is open. The system should refuse to start an experiment that overlaps an exclusion window rather than trusting the operator to check.
A change freeze for the duration. Nothing else deploys to the affected service while an experiment runs. This is the same problem as the canary attribution issue, and it has the same solution.
Blast radius that does not scale the blast. The percentage of traffic affected is not the only knob. Halving the availability of a service that is already the dependency of everything else is a larger blast radius than losing 5% of a leaf service, and the engine should reason about dependency position rather than only about percentage.
A useful addition: run a control. If the experiment is going to affect 5% of traffic, keep 5% of traffic unaffected in a comparable way, so you can distinguish the experiment’s effect from what was already happening. This is the same discipline as a canary, applied to fault injection.
And the most underrated control: the platform should record what it injected, when, and to where, in a timeline that overlays every dashboard. So the first thing anyone does when an alert fires is check whether an experiment was running. A platform where the answer to “what changed” sometimes being “an experiment” trains people to distrust alerts, which is a worse outcome than any single experiment.
Failure stories worth testing
Run an experiment while a large deploy is rolling
Confirm the platform blocks it, and if it does not, confirm the resulting incident is diagnosable. This is the attribution test and it is worth doing deliberately once.
Double the experiment scope in the config
Confirm the engine refuses it. If a scope violation is accepted, your platform is a script with good documentation rather than a safety mechanism.
Kill the operator’s connection mid-experiment
Confirm the abort condition and the duration cap both revert the fault. An experiment that only stops when its operator can click stop is not safe to run unattended.
Make a dependency hang instead of returning errors
Confirm the timeout path works. Latency injection finds more problems than error injection does, and almost nobody does it.
Set a dependency’s retry count to 100
Measure the amplification against the dependency. This is where the real number differs from the calculated one.
Trigger a customer-facing alert during an experiment
Confirm the auto-abort fires without a human. Then confirm someone noticed. An abort that is silent teaches the team that alerts are not reliable.
Set the duration cap to 2 hours and make the fault’s cleanup fail
Confirm the clock stops it at the cap. This is the control that keeps a failed experiment from becoming an outage that lasts until someone finds it.
Run an experiment in staging only
Confirm the team recognises that staging told them nothing about production behaviour. This is the point of the whole practice, and it is worth making explicit at least once.
Hit the global kill switch during an experiment
Confirm everything reverts within seconds. If it does not, the switch is decorative and the platform is not ready for real use.
Take away a service that three other services depend on
Observe the fan-out. Almost every team discovers one downstream service has no timeout and blocks indefinitely. This is the highest-value-per-minute experiment in the list.
Run twenty experiments and record the hit rate
A 100% confirmed rate means the hypotheses were too weak. The platform should make it easy to see this, because it is the clearest signal that the practice has become theatre.
A production-ready architecture
+----------------------------------------------------------+
| experiment spec (YAML / UI) |
| hypothesis, steady state, fault, abort, scope |
+------------------------+---------------------------------+
|
v
+----------------------------------------------------------+
| pre-flight gate |
| - hypothesis is measurable and falsifiable? |
| - steady state is currently green? |
| - exclusion window clear? (peak, freeze, incident) |
| - scope within platform caps? |
| - approver signed off, is not the author? |
+------------------------+---------------------------------+
|
v
+----------------------------------------------------------+
| engine |
| +----------+ +------------+ +-----------+ |
| | latency | | errors | | resource | |
| | loss | | kill | | cpu/mem | |
| | partition | | throttle | | disk | |
| +----------+ +------------+ +-----------+ |
| |
| enforces: service allowlist, pod count cap, |
| traffic %, hard duration cap, auto-revert |
+------------------------+---------------------------------+
|
+------------------+------------------+
| |
v v
+--------------+ +------------------+
| safety net | | timeline |
| - SLO burn | | inject events |
| -> abort | | overlaid on |
| - any page | | every dashboard |
| -> abort | +------------------+
| - duration |
| -> abort |
| - kill switch| global, one action, immediate
+--------------+
+----------------------------------------------------------+
| results |
| confirmed | falsified | inconclusive |
| hit rate tracked: a 100% rate is a warning sign |
| falsified -> engineering work, not blame |
| writeup required either way |
+----------------------------------------------------------+
A sensible delivery checklist:
- Write the hypothesis with a metric and a threshold before the experiment is approved. If the hypothesis is not falsifiable, do not run it.
- Enforce scope in the engine, not in a review checklist. Allowlist, pod cap, traffic percentage, and a hard duration cap that the platform controls.
- Make abort conditions automatic and independent of the fault: any page, any SLO burn, and the duration cap each revert the experiment on their own.
- Build a global kill switch and test it. It is the control that determines whether the team will allow experiments at all.
- Record every injection on a timeline that overlays dashboards, so “was an experiment running” is the first question anyone asks when an alert fires.
- Block experiments during exclusion windows at the platform level.
- Freeze other changes to the affected service for the duration.
- Start with timeouts, retries, backpressure, and single-instance failures. Earn the right to run zone-loss experiments.
- Include a control group in the fault, so the experiment’s effect is distinguishable from ambient conditions.
- Track the hypothesis hit rate. Consistently confirming hypotheses is a signal that the practice has stopped testing anything.
- Treat a falsified hypothesis as the valuable outcome. If nobody ever finds out the assumptions were wrong, the experiments were not experiments.
- Review the experiment backlog quarterly and retire the ones that keep confirming themselves, and add experiments for the assumptions nobody has written down.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| No measurable hypothesis | The result is an anecdote | Metric and threshold, decided upfront |
| Hypothesis about mechanisms, not outcomes | “The circuit breaker opened” proves nothing | “Users still completed checkout” |
| Blast radius left to operator discipline | One experiment becomes an outage | Enforce caps in the engine |
| No automatic abort | The experiment runs until someone notices | Abort on any page or SLO burn |
| Duration cap controlled by the fault | A failed cleanup leaves the fault running | Platform clock stops it regardless |
| No kill switch | People are afraid to run experiments, correctly | Global halt, one action, tested |
| Chaos in staging only | Production behaviour is never tested | Run against production, bounded |
| Skipping to zone loss first | A high-consequence experiment with no foundation | Work up the risk gradient |
| Running during peak or a freeze | The result is uninterpretable | Exclusion windows, enforced |
| No change freeze during an experiment | Concurrent changes confound the result | Freeze the affected service |
| No control group | Ambient effects are attributed to the fault | Keep a comparable unaffected slice |
| Injections not on the dashboard timeline | “What changed” has no answer | Overlay inject events on every view |
| 100% hypothesis confirmation rate | The practice has become theatre | Track the rate, write falsifiable claims |
| No post-experiment monitoring window | The damage is measured before it appears | Watch for hours after reverting |
| Falsified hypothesis treated as blame | Nobody writes an honest hypothesis | Incidents are information, not culpability |
| One giant experiment testing five things | The result attributes nothing | One hypothesis per experiment |
| No experiment review after the fact | Nothing is learned beyond the result | Require a writeup, feed findings to backlog |
| Retrying a falsified experiment unchanged | Same result, more confidence, no progress | A falsified hypothesis is a work item |
| Chaos tooling built but never used | Platform exists, practice does not | Reduce friction until volume increases |
The complete story in one minute
Chaos engineering is a measurement practice, not a destruction practice, and the difference is entirely in the preparation. The unit of work is a falsifiable hypothesis with a specific metric and a threshold, written before anything is broken. “Let’s see what happens if the recommendation service goes down” is a stress test; “while serving double normal traffic, if recommendations return 500s, checkout still completes from the cached list and p99 stays under 800ms” is an experiment, and the second one is falsifiable on a Tuesday.
The ratio of preparation to runtime is roughly an hour per minute, which is why the platform matters more than the technique. Its job is to make the safe path easy and the unsafe path impossible: scope enforced in the engine rather than by operator discipline, abort conditions that fire automatically on any page or SLO burn, a duration cap the platform controls rather than the fault’s own lifecycle, and a global kill switch. If a scope violation is accepted, you have a script with documentation rather than a safety mechanism.
Then order the work as a risk gradient. Timeouts and latency injection first, because a dependency that hangs instead of failing is a common cause of cascading failure and the hardest fault to inject. Then retry budgets, where the real amplification is usually worse than the calculated one. Then backpressure, single-instance failure, dependency degradation, partial network failure, and only at the end zone loss, which requires a tested failover story before it starts.
And control the confounders: exclusion windows enforced at the platform, a change freeze on the affected service, a control group so the experiment’s effect is separable from ambient conditions, and every injection recorded on a timeline that overlays every dashboard. That last one is what keeps alerts trustworthy, because the first question during any incident becomes “was an experiment running” with an answer.
The outcome is not more reliability. It is less uncertainty, and the honest measurement is your hypothesis hit rate — a team confirming everything has stopped testing anything.
hypothesis: metric + threshold, falsifiable, user-visible
safety: enforced scope, auto-abort, duration cap, kill switch
controls: exclusion windows, change freeze, control group, timeline
The hard part was never injecting the fault. It was making sure the answer you got was about the thing you asked about.


