← All writing
articleApr 09, 202419 min read

Service Mesh: Every Request Now Goes Through a Proxy, So What Did You Buy

xDS control plane, sidecar versus ambient, mTLS with SPIFFE identity, and the honest reasons a mesh is worth it — and the ones that are not.

Service MeshmTLSxDSArchitecture
Service Mesh: Every Request Now Goes Through a Proxy, So What Did You Buy cover illustration

A service mesh is a decision to put a proxy on the network path of every service-to-service call, and to put a control plane in charge of configuring those proxies. That first part is a real cost: an extra process, extra memory, an extra hop with its own latency and its own failure modes, added to every single request in the system.

It is a decision many large organisations have made deliberately. This article is about what you get for that cost, and about the several things people expect to get that you do not.

The scale the numbers actually imply

Ten thousand pods, and every one of them making outbound calls.

10,000 pods x sidecar proxy
  50 MB RSS per proxy (idle)
  = 500 GB of memory just for proxies
  = ~125 nodes' worth of capacity, gone

  per-request added latency
  +0.1 - 0.5ms per hop on the same node
  with 4 hops, that is the tail latency of your p50

  request rate
  20,000 requests/sec across the mesh
  every one is now two TLS sessions,
  two policy evaluations, two spans

Those numbers explain almost every design decision in the mesh ecosystem. The memory figure is why ambient data paths exist. The latency figure is why sidecars are often considered unacceptable for RPC-heavy internal traffic. The request rate is why the control plane has to be push-based with a pull fallback rather than a chatty request-response protocol.

What the proxy actually does

The sidecar sits next to the application in the same pod, and it intercepts traffic in both directions.

  +-----------------------------------------------+
  |  pod                                         |
  |                                               |
  |   +----------------+     +----------------+   |
  |   | application    | --> |  sidecar proxy  |   |
  |   | :8080          |     |  inbound:       |   |
  |   |                | <-- |   authz, metrics |   |
  |   +----------------+     |                 |   |
  |                          | outbound:        |   |
  |   +----------------+     |   discovery,     |   |
  |   | sidecar proxy   | --> |   LB, retry,    |   |
  |   +----------------+     |   mTLS, tracing  |   |
  |                          +----------------+   |
  +-----------------------------------------------+

Inbound sidecar work:

  1. The connection arrives, the proxy identifies the caller from its mTLS certificate.
  2. It evaluates an authorization policy against that identity, the path, the method, and the headers.
  3. It applies rate limits if configured.
  4. It records metrics and a trace span.
  5. It forwards to the application, or rejects.

Outbound sidecar work:

  1. The application resolves a logical name like payments.default.svc.cluster.local.
  2. The proxy translates it to an endpoint using the discovery information the control plane gave it.
  3. It applies load balancing, outlier detection, circuit breaking, retries, and timeouts from configuration.
  4. It establishes mTLS to the destination’s sidecar.
  5. It forwards.

The application is supposed to be oblivious to all of it. That is the appeal and it is real: a Go service can gain mTLS and retries without a line of code, which is a genuine advantage when you have four hundred services written in four languages by four different teams.

Ambient: the same idea, amortised per node

Sidecar’s memory cost scales with pod count, and that is its structural weakness. The ambient data path moves the proxy out of the pod and up to the node.

  node
  +--------------------------------------------------+
  |                                                  |
  |   +--------+  +--------+  +--------+             |
  |   | pod A  |  | pod B  |  | pod C  |             |
  |   +--------+  +--------+  +--------+             |
  |        \           |           /                |
  |         +----------v-----------+                 |
  |         |  ztunnel / node proxy|                 |
  |         |  L4 aware, mTLS,     |  one per node   |
  |         |  no L7 parsing       |                |
  |         +----------+-----------+                 |
  |                    |                             |
  |             optional L7 proxy                   |
  |             only where HTTP policy is needed      |
  +--------------------------------------------------+

The gains are substantial: one proxy per node rather than per pod, no sidecar injection into the application, and a smaller blast radius when the proxy has a problem. Node-level proxying also makes the security story simpler, because you have far fewer components to patch and audit.

The loss is equally substantial: L7 policy has to be paid for separately. A pure L4 proxy can do identity-aware mTLS, connection-level metrics, and network policy, but it cannot read an HTTP path or a header. If you want per-route authorization, per-route timeouts, or header-based routing, you need an L7 proxy in the path, and the architecture becomes “ambient for the common case, an L7 proxy for the exceptions.” That is a real design and it is more complex than “sidecar everywhere,” not less.

The honest summary: sidecar is a stronger per-pod feature set with a memory cost you pay per pod. Ambient is a cheaper, simpler data path with a weaker feature set that you extend where needed.

Identity is the actual product

The thing that makes a mesh more than a load balancer is that a request carries a cryptographically verifiable identity that the receiver can check, without trusting the network, without trusting a header, and without a shared secret in configuration.

  +----------------+                    +----------------+
  |  payments pod  |                    |  checkout pod  |
  |                |                    |                |
  |  SVID:         |   mTLS handshake    |  SVID:         |
  |  spiffe://     |-------------------->|  spiffe://     |
  |  cluster.local/|  presents cert     |  cluster.local/|
  |  ns/default/   |  for the *caller*  |  ns/default/   |
  |  sa/payments   |                    |  sa/checkout   |
  +----------------+                    +----------------+

The order is what matters. The client presents a certificate for its own identity, and the server verifies it against a trust domain. The result is a channel where both ends know who they are talking to over a network they do not control.

SPIFFE defines the identity format and SVID is the credential, with X.509 certificates for mTLS and JWTs for other uses. The critical property is short lifetime: certificates live for minutes to hours, and the workload obtains them automatically from a certificate authority in the control plane. Nothing long-lived is stored in a file in the pod.

This is where the security claim is often overstated. A mesh provides identity only if obtaining that identity is harder than writing an arbitrary string into a header. If a workload can simply set a header and the receiving sidecar trusts headers, then mTLS is decorative. In practice the failure mode looks like this: someone adds a debugging header, a sidecar is configured to trust headers from inside the mesh, and now every pod in the cluster can impersonate every other pod. Peer authentication in a mesh should default to mutual TLS and identity should come from the certificate, always, with no header path.

The authorization model is then a policy attached to a workload, evaluated at the proxy:

allow
  principal: spiffe://cluster.local/ns/default/sa/checkout
  action: ALLOW
  paths: ["/v1/payments/*"]
  methods: ["GET", "POST"]

And the policy lives in the control plane, versioned and reviewed, rather than in every service’s code. For an organisation with hundreds of services and a real problem with “who is allowed to call the payments API,” that is a genuine win.

xDS: push config, and the failure mode nobody warns about

The control plane does not sit in the request path. It configures proxies out of band, which is what keeps the data path fast. The protocol is xDS: a set of discovery services, each streaming updates to clients over gRPC.

  +---------------------------+
  |  control plane            |
  |  - service registry (EDS) |
  |  - route config (RDS)     |
  |  - cluster config (CDS)   |
  |  - listeners (LDS)        |
  |  - secrets (SDS)          |
  +------------+--------------+
               |
     gRPC streams, one per type
     push on change, keepalive every N seconds
               |
               v
  +---------------------------+
  |  sidecar (data plane)     |
  |  holds last-known config  |
  |  serves traffic with it   |
               |
  when the control plane is
  unreachable: the proxy keeps
  serving. Forever, if needed.

That last line is the property to internalise. A mesh degrades to stale configuration, not to failure. The proxy does not stop working when the control plane is unreachable; it continues using whatever it last received. That is deliberate and it is what makes the data path independent and available.

It also creates the worst class of mesh incident: a policy change that never reaches the proxies. A revocation of a service account’s access is pushed, some proxies receive it, some do not, and the difference between a compliant and a non-compliant pod is invisible until you compare the config versions. Any serious mesh deployment needs a way to answer “what configuration is proxy X actually running”, which means a way to compare the intended config version against what each proxy has acknowledged. Without that, policy changes are best-effort, and “we revoked it” is not the same as “it is revoked everywhere.”

The other xDS problem is the connection storm. When the control plane comes back after being down, every one of ten thousand proxies reconnects at once. A correct implementation uses jittered backoff and connection deduplication. An incorrect one turns a control plane recovery into a thundering herd that knocks the control plane over again, which knocks the proxies out, and you have built an oscillator.

Retries and timeouts: one owner, or a 30x outage

This is the mesh feature most likely to hurt you, and it has nothing to do with security.

If the mesh retries and the application also retries, every failure is multiplied.

  app retry: 3 attempts
  mesh retry: 3 attempts, per hop
  4-hop path

  worst case
  3 x 3 x 3 x 3 = 81 requests
  for one user click

And retries against a service that is already overloaded make the overload worse, which makes it retry more. The death spiral is not subtle; it is the most common self-inflicted outage in meshed systems.

The rules that keep it sane:

  1. One policy owns the retry budget. The application usually knows whether an operation is semantically safe and can provide an idempotency key; the proxy has better transport signals. Either layer may execute retries, and some systems intentionally use both, but their combined attempts must fit one end-to-end budget rather than multiplying independently.
  2. Never retry a non-idempotent request without an idempotency key. A retried POST that succeeded and whose response was lost is a duplicate charge, and the mesh does not know which requests were safe.
  3. Retry only on specific conditions. Connection refused and 503 are retryable. 400 and most 4xx are not. Retrying a validation error just wastes capacity.
  4. Budget retries globally, not per request. A retry budget as a fraction of total requests, or a ratio of retries to originals, is the standard defence. Cap it and the amplification is bounded by construction.
  5. Derive deadlines from the caller’s end-to-end budget. Each downstream attempt, retry, queue, and network hop consumes part of it, so inner timeouts must leave time for the caller to react. “Longer than the callee’s p99” is not a safe rule when the user budget is smaller, and p99 itself moves during an incident.
  6. Only retry on the same host or a different one, deliberately. Retrying the same broken host gives you three times the load on the instance that is failing. Outlier detection, which ejects hosts that fail health checks, is the counterpart to this.

What you do not get

A mesh does not give you these, and the confusion about this is the source of a lot of disappointment.

It does not make your services more reliable. A proxy is another process that can fail, another hop, another thing that consumes CPU on an already-loaded node. A mesh that retries a service into overload is a mesh that reduced your availability. Reliability still comes from the services, their dependencies, and their capacity.

It does not replace application-level observability. The mesh gives you golden signals per hop: request count, duration, and status. It does not tell you what your code was doing. Traces from the mesh are the connective tissue, not the trace of your business logic.

It does not remove the need to design your APIs for failure. Timeouts, idempotency, bulkheads, backpressure, and circuit breakers are still application decisions. The mesh can enforce a retry budget; it cannot decide that a given endpoint is safe to retry.

It does not make deployment safer. Progressive delivery is a separate concern, and a mesh rollout of proxies is itself a risky operation that needs its own canary.

Failure stories worth testing

Take the control plane down

Nothing breaks immediately, and everything keeps working on stale config. This is the most important test in the whole product: it tells you what your true availability floor is and whether you understood it.

Push a policy change and immediately probe from a specific pod

Find out how long convergence actually takes and how much it varies. If you cannot answer this question about your mesh, you do not have a security posture, you have an aspiration.

Kill a sidecar in a pod

The application loses outbound connectivity and possibly inbound. This is a real failure mode of the sidecar model and the test is whether your applications have any behaviour that is better than a hard failure.

Return 500s from a dependency for 30 seconds

Watch the retry volume. This is where amplification shows up, and the number should be bounded by a retry budget you can point at.

Make a 4-hop chain with retries configured at every layer

Measure the actual request count. If the worst case is 81 requests for one user action, the mesh’s defaults are doing something your organisation has not agreed to.

Push a config update faster than the proxies can apply it

The control plane must coalesce. An implementation that queues every update and applies them in order will fall arbitrarily far behind, and a burst of config changes becomes a slow-motion rollout nobody can reason about.

Reconnect 10,000 proxies simultaneously

Reconnect the control plane after an outage. If the recovery knocks it over, the jitter and deduplication are wrong.

Rotate the mesh’s root certificate

Every workload’s SVID changes. If this is a scary operation in your deployment, the certificate lifetime is too long and your rotation story needs work.

Send a request with a forged identity header

It must be rejected. If it is not, the mesh is authenticating connections and then trusting a string, which is worse than having no identity at all because it looks secure.

Make L4-only ambient proxying handle an L7 policy

It cannot, and the failure should be a clear configuration error at rollout time, not a policy that silently never applies.

A production-ready architecture

  +-----------------------------------------------------+
  |  data plane (on every node)                          |
  |                                                     |
  |   +--------+   +--------+   +--------+              |
  |   | pod A  |   | pod B  |   | pod C  |              |
  |   +---+----+   +---+----+   +---+----+              |
  |       \           |            /                   |
  |        +----------v-----------+                     |
  |        |  node proxy          |  mTLS origination   |
  |        |  (L4, identity,      |  workload identity   |
  |        |   network policy)    |  from node agent     |
  |        +----------+-----------+                     |
  |                   |                                 |
  |            +------v------+  where L7 policy exists  |
  |            | L7 gateway  |                          |
  |            +-------------+                          |
  +---------------------|--------------------------------+
                        |
  service A ----mTLS---- service B
       every hop presents a workload
       identity, authorized by policy
                        |
                        v
  +-----------------------------------------------------+
  |  control plane (separate, not in the data path)      |
  |                                                     |
  |   cert authority  -> SVIDs, short-lived             |
  |   config store    -> policies, routes, endpoints    |
  |   xDS server      -> LDS, RDS, CDS, EDS, SDS        |
  |   push on change, keepalive, coalesce, jitter       |
  +-----------------------------------------------------+

  observability: RED metrics per hop, trace spans per hop,
  access logs, config-version acknowledgement per proxy

A sensible delivery checklist:

  1. Name the specific feature you are buying — identity, policy, telemetry — before you install anything. If the answer is vague, the cost is not justified.
  2. Choose the data path deliberately. Sidecar for per-pod L7 features and isolation; ambient for a cheaper, node-amortised path with L7 only where needed.
  3. Identity always from the mTLS certificate, short-lived SVIDs, no header-based identity path. Ever.
  4. Peer authentication defaults to strict mutual TLS. Nothing trusts a header it did not cryptographically verify.
  5. Keep the control plane out of the data path and prove the data path survives its absence with a real test.
  6. Instrument config convergence: the version the control plane intended, and the version each proxy acknowledged. Without it you cannot claim a policy is enforced.
  7. Exactly one layer owns retries. Document which. Disable the other.
  8. Put an end-to-end retry budget in place before an incident, count attempts from every layer, and require idempotency keys for operations that may be retried safely.
  9. Make the timeout hierarchy monotonic, shorter inside than outside, and derived from real latency data rather than round numbers.
  10. Rate limit per principal, not per IP. An IP-based limit punishes NAT and does nothing against a compromised workload.
  11. Roll the mesh itself out canary-first, with a way to eject a proxy to the pod’s original networking. A mesh rollout is a networking change to every request in the system.
  12. Keep a documented bypass. A mechanism to disable the mesh for a single namespace during an incident is worth more than any feature on the list.

Common mistakes

Mistake What actually happens Better decision
Mesh installed with no stated goal Cost without benefit, and no way to evaluate it Name identity, policy, or telemetry first
Long-lived static service certificates Revocation is manual and slow Short-lived SVIDs, automatic renewal
Trusting identity headers Any pod impersonates any other pod Identity from the mTLS certificate only
Strict mTLS turned off “to debug” Becomes permanent and invisible Debug with scoped exceptions, alert on the setting
Retries at both mesh and application 3x3x3x3 amplification on failure One owner, explicitly chosen
Retrying non-idempotent requests Duplicate charges, duplicate orders Idempotency key, or no retry
Retrying the same failing host Triple load on the instance that is broken Outlier detection, then retry elsewhere
Timeout inside longer than outside Slow responses become errors at the wrong layer Monotonic timeout hierarchy
No retry budget Retry storms amplify an overload into an outage Global cap on retries relative to originals
No config-convergence metric Policy changes are best-effort, invisibly Track intended versus acknowledged version
Control plane with no backpressure A config burst becomes a slow-motion rollout Coalesce and rate-limit updates
Reconnects without jitter Control plane recovery oscillates Jittered exponential backoff
IP-based rate limits NAT is punished, compromised workload is not Rate limit per workload identity
Sidecar per pod at 50 MB 500 GB of memory and per-pod CPU overhead Evaluate the ambient data path
Rolling the mesh to every pod at once A networking change to every request, unobserved Canary, with a bypass
L7 expectations on an L4 ambient path Policy silently never applies Fail configuration at rollout time
Mesh treated as reliability The extra hop reduces availability Reliability still comes from the services
Disabling the mesh with no plan A bad rollout has no way out Documented per-namespace bypass

The complete story in one minute

A service mesh is a proxy on the request path of every call, plus a control plane that configures it out of band. The memory cost and latency are real — a sidecar at 50 MB per pod is 500 GB across ten thousand pods — which is exactly why the ambient data path moves the proxy to the node and pays for L7 only where HTTP policy actually needs it.

The product you are actually buying is identity and policy. mTLS with SPIFFE SVIDs gives a request a cryptographically verifiable workload identity, evaluated at the receiving proxy against a policy that lives in the control plane. The security claim only holds if identity comes from the certificate and never from a header, because a header is a string any pod can write.

The control plane is not in the data path, and that is deliberate: a proxy that loses its xDS connection keeps serving the last configuration it received, forever. That makes the data path available and it makes policy propagation unverifiable unless you track the version the control plane intended against the version each proxy acknowledged. Without that metric, revocation is best-effort.

Retries are the thing most likely to hurt you. Three attempts at multiple layers and hops can multiply one user action into a retry storm against the service that is already failing. Define one end-to-end attempt budget, decide which layer has the semantic authority to spend it, require idempotency for retryable side effects, and propagate a deadline that decreases toward each callee.

And the mesh gives you none of reliability, no observability into your business logic, no safety from deployment, and no substitute for designing your services for failure. It is a networking layer with a security model attached, and it earns its cost only when that security model is the thing you needed.

workload identity -> mTLS -> policy at the proxy -> metrics, traces
control plane -> xDS push -> proxy config (never in the request path)
one layer owns retries, and the budget is global

The hard part was never proxying a request. It was owning the failure behaviour you just multiplied by adding a hop.

Technical references

Keep reading
Browse everything