← All writing
articleFeb 25, 202419 min read

Secrets Management: The Design Where Decryption Should Be the Rare Event

Envelope encryption, HSM-backed roots, dynamic database credentials, and why the hard part of secrets management is revocation, not storage.

SecurityEncryptionArchitecture
Secrets Management: The Design Where Decryption Should Be the Rare Event cover illustration

Every team stores secrets badly at first. A password goes in an environment variable, then in a config file, then in the repository, then in the repository’s history forever. The fix is usually presented as an encryption story: put secrets in a vault, encrypt them at rest, use TLS in transit, done.

That gets you the part that is easy to measure and the part that rarely causes an incident. The incidents come from elsewhere — a secret that cannot be revoked, a read that was not logged, a rotation that silently broke one instance out of ten thousand, or a vault that became a single point of failure nobody planned for.

The design goal is worth stating up front, because it inverts the naive approach:

Encryption at rest is a solved problem. What is expensive is decryption. So design so that decrypt happens as little as possible, for as short a time as possible, and is recorded every time.

The scale, and where the cost actually is

A platform with ten thousand service instances, several thousand distinct secrets, and a steady stream of rotation.

10,000 instances x secrets read at startup    = tens of thousands of reads/day
                                               = a few requests/sec sustained

one secret materializes into ~10,000 plaintext
copies on disk across the fleet, in memory,
in crash dumps, in core dumps, in swap

rotation events
a security requirement, not a preference:
  database passwords every 90 days
  API keys per team, on a schedule
  certificates on renewal
  emergency revocation, on demand

Two things fall out of those numbers.

The read path is cheap and the write path is where the ceremony belongs. A few requests per second of secret reads is nothing. Rotation and revocation are the operations that must not be slow, must not require downtime, and must be audited with precision.

Plaintext is the real exposure surface. One secret decrypted to one process that then writes it to a log, a crash dump, or a support bundle multiplies the exposure by ten thousand. Every design decision below ultimately reduces either the lifetime of a plaintext secret or the number of copies of one.

Envelope encryption: the core of the design

The key insight is that you should not encrypt a secret with the key that protects the master key. You encrypt with a data key, and you protect the data key with a root key.

write path
  plaintext secret
        |
        v
  generate random data key (DEK)     256-bit, fresh per secret version
        |
        +--> encrypt secret with DEK (AES-256-GCM)  -> ciphertext + nonce
        |
        +--> wrap DEK with root key (KEK)            -> wrapped DEK
        |
        v
  store ciphertext, wrapped DEK, nonce, metadata
  store the plaintext DEK in no database, ever

read path
  ciphertext + wrapped DEK
        |
        v
  ask the HSM to unwrap the DEK
        |
        v
  decrypt ciphertext with DEK
        |
        v
  plaintext, in memory, briefly

What this buys:

Key rotation does not require re-encrypting data. The root key wraps data keys. Rotating the root key means re-wrapping the data keys, which is cheap, not re-encrypting every secret. Re-encrypting a large corpus is the thing that turns a routine rotation into a maintenance window.

Blast radius per key is small. A data key protects one secret version. Compromise of one data key exposes one secret, not all of them.

The root key can be non-exportable. If the KEK lives inside an HSM and the API offers only unwrap operations with no extract, then there is no value an attacker can steal from a backup of the vault database. Backups become ciphertext, which is exactly what you want.

Encryption and decryption are cheap. AES-GCM on a modern core runs at multiple gigabytes per second. A 4 KB secret is not a rounding error in any request’s budget. The expensive operation in the read path is the HSM round trip, which is why data keys are cached with a short TTL rather than fetched on every read.

The nonce handling is where hand-rolled implementations get themselves into trouble. AES-GCM requires a unique nonce per key, and a repeated nonce with the same key is catastrophic: it leaks the XOR of plaintexts and destroys the authentication key. Generate a 96-bit random nonce per encryption and store it with the ciphertext, or use a counter you are certain is durable. Random 96-bit is fine at this volume; the failure mode to fear is a random source that silently returns a constant.

The AEAD tag also matters. AES-GCM is authenticated, and the tag must be verified on decrypt. A system that decrypts and ignores authentication errors has not verified anything.

What the HSM is actually for

A hardware security module is a tamper-resistant device that holds a key and performs cryptographic operations without ever returning the key material.

  application
       |  "unwrap this blob"
       v
  +----------------------+
  |  HSM                 |
  |  - KEK never leaves  |
  |  - unwrap is an      |
  |    operation, not    |
  |    an export         |
  |  - tamper-evident    |
  |  - rate limited      |
  |  - audited internally|
  +----------+-----------+
             |
      vendor / network API

The value is not performance, which is poor, and it is not that the software is better. The value is that the key is not in the process memory of the vault, not in a heap dump, not in a core dump, not in a swap file, and not in a backup. Software key storage has to be correct about not leaking memory; an HSM removes the question.

The costs are real and need to be in the design: a network round trip of 0.5-5ms to the HSM, an availability dependency on a device and a vendor API, and a capacity limit. A design that HSM-unwraps on every secret read is HSM-bound by definition. The mitigation is caching unwrapped data keys in the vault process for a bounded period — 30 seconds to a few minutes — and accepting that a compromise of a vault process within that window exposes the data keys it was holding. That is a deliberate, bounded, and much smaller exposure than the alternative.

If an HSM is not available, a software key store protected by a KMS with envelope encryption is a genuinely good answer. Do not let a desire for the strongest option delay having a working one.

The design that actually prevents incidents

Static secrets — a database password sitting in a vault that gets handed out — have an inherent problem. It is out there. It does not expire on its own. If it leaks at month nine, it is valid until someone notices, and noticing is the hard part.

Dynamic credentials invert this. For databases, the sequence is:

  1. A service authenticates to the vault using its own identity, not a shared password.
  2. The vault verifies that identity against an auth method, usually a signed JWT from the workload’s own platform.
  3. The vault generates a new database role with a random password and a short TTL.
  4. The credential is created in the database and leased for one hour.
  5. The service connects, does its work, and the credential is revoked by the lease at the end.

The consequence is that the database password is never stored anywhere, never in a repo, never in an environment variable, and never needs rotating. It is generated fresh and dies on its own. The only remaining question is revocation, and revocation is now a fast, single call rather than a change propagated to ten thousand instances.

  service                 vault                  database
     |                      |                       |
     |  JWT identity         |                       |
     |--------------------->|                       |
     |                      |  verify identity      |
     |                      |  generate role+pass   |
     |                      |  CREATE ROLE ...      |
     |                      |---------------------->|
     |                      |  GRANT, set TTL       |
     |                      |<----------------------|
     |  credential, 1h TTL  |                       |
     |<---------------------|                       |
     |                                              |
     |  connect with the generated credential       |
     |--------------------------------------------->|
     |                                              |
     |                    1 hour passes              |
     |                      |  DROP ROLE            |
     |                      |---------------------->|

This requires the database to support short-lived authenticated users, which Postgres and MySQL both do. It also requires the service to tolerate the credential expiring mid-operation, which is the real work: pooled connections have to be recycled, and long-running transactions have to be shorter than the lease. That constraint is often worth having anyway.

Certificates should be short-lived too. Thirty-day certificates with automated renewal is a common target; one-hour certificates with automation is better and is what internal PKI increasingly issues. The design rule is the same as for credentials: make the validity window shorter than your detection window and you never have to reason about a leaked credential you do not know about.

Rotation is a property of the data, not a job. The right mental model is that secrets have a creation time and an expiry, and anything that needs rotation is just something with a short expiry. A rotation job is a workaround for a schema where the secret has no lifetime.

Auditing: reads matter more than writes

Knowing who changed a secret is the easy half. The half that catches incidents is knowing who read it.

audit record
  timestamp
  actor identity (workload, not user, for machine access)
  secret path
  operation: read | write | delete | lease | unwrap
  client: pod id, node, IP, user agent
  token or connection id
  decision: allowed | denied
  reason for denial: no policy, expired lease, wrong path

A read audit is what answers the question that matters during an incident: which workloads had access to a credential in the window before it leaked. Without read records, you can list the people who could have changed it and you still cannot scope the exposure.

Two practical details. Audit records must be tamper-evident — hash-chain each entry to the previous one and ship them off-host, because an audit log that a compromised vault can edit is not evidence. And the audit contract has to be explicit. HashiCorp Vault, for example, blocks a request when no enabled audit device can persist its record; multiple devices provide redundancy, but a request must not be reported as successful while its evidence is silently dropped. A home-grown asynchronous pipeline can reduce latency, but unless it durably admits the record before responding, it is best-effort telemetry rather than a fail-closed audit boundary.

Denials are the most valuable records. A denied read means something tried to reach a secret it had no business reaching, which is either a misconfiguration, a bug, or an intrusion, and it is a strong signal precisely because it is rare.

Availability: the vault is a dependency of everything

If the vault is unreachable and services cannot start, you have converted a secrets problem into an outage. The design has to assume the dependency exists and degrade deliberately.

startup, good case
  fetch secret -> decrypt -> connect -> serve

startup, vault unreachable
  ???

There are three options, and they differ mostly in what you give up.

Cache secrets locally with a long TTL. Each service keeps a copy on disk, encrypted with a key that is not in the vault, and starts from that copy when the vault is down. Simple, and it reintroduces a long-lived local copy of a secret — which is the thing dynamic credentials were meant to eliminate. It is a reasonable fallback for development and staging and a poor one for production.

Make secrets optional at startup when they are not needed yet. A service that reads a secret lazily, on first use, can start and become healthy without the vault. It fails on the specific operation that needs the secret instead, which is a much smaller blast radius. This requires not having a static secret hardcoded as a boot dependency, which is a good discipline regardless.

Cache unwrapped data keys, not secrets. The vault serves from a local cache of data keys. If the vault process is up but the HSM is not, most reads still work. This handles the most common real failure, which is a dependency of the dependency being down, and it does not require storing plaintext secrets on client machines.

Choose the fallback on purpose per secret class. The database credential that lives for an hour and the third-party API key that never expires do not deserve the same caching policy, and a system with one global policy will make the wrong trade for one of them.

Failure stories worth testing

Lose the HSM

The vault can still serve reads for data keys it has cached, and cannot unwrap anything new. Measure how long the cache covers you and make sure the expiry is a decision rather than a discovery.

Restore a vault backup onto new hardware

The new instance has ciphertext and no root key. This is the case envelope encryption was designed for, and the test is that it fails with a clear error rather than returning garbage. A backup is not a restore without the KEK.

Revoke a secret while ten thousand instances hold it

The next renewal fails. The point of the test is what each instance does: crash, retry forever, or drop the capability. Decide, and then make the decision visible.

Compromise one service’s identity

It should get exactly that service’s secrets and nothing else. Test with a policy that denies a neighbouring namespace, because a policy engine that allows broadly on a failed lookup will quietly hand out everything.

Make the audit pipeline unavailable

For a fail-closed audit boundary, reads must fail when every audit sink is unavailable; confirm redundant sinks prevent a single-sink outage from taking the service down. If the product deliberately chooses buffered best-effort telemetry instead, verify durable admission or a bounded and visible rejection mode—never silent drops—and do not describe it as a complete audit log.

Issue a certificate with a clock that is slightly wrong

Clock skew in a PKI produces certificates that are valid from the future or already expired, and both look like an application bug. Verify the whole chain, not just the leaf.

Generate a very large number of secrets

Confirm data key generation, wrapping, and the read path stay within budget. The failure mode is a design that assumed secrets are small, so a 4 MB certificate bundle makes the vault a throughput bottleneck.

Run an HSM-unavailable drill with a service that must start

Confirm the actual startup path. A service that only reads secrets lazily starts fine, and one that reads at boot does not. Knowing which of your services are which is the point of the drill.

Attempt a read with a revoked token

It must be denied and audited, and the denial must not be cacheable by any layer in between.

A production-ready architecture

   10,000 workloads
        |  workload identity (JWT / SPIFFE / cloud IMDS)
        v
  +---------------------------+
  |  secrets service          |  policy engine per request
  |  auth + policy            |  short-lived tokens only
  |  cache: unwrapped DEKs    |  bounded TTL, memory only
  +------------+--------------+
               |
     +---------+---------+
     |                   |
     v                   v
  +--------+        +------------+
  | HSM    |        | database   |  roles, leases,
  | KEK    |        | (dynamic  |  revocation
  | wrap / |        |  creds)   |
  | unwrap |        +------------+
  +--------+
               |
               v
  +---------------------------+
  |  storage                  |  ciphertext only
  |  versioned secrets        |  wrapped DEK
  |  no plaintext at rest     |  nonce, tag, metadata
  +---------------------------+

  every read/write/denial
        |
        v
  +---------------------------+
  |  audit log                |  hash-chained
  |  shipped off-host         |  denials surfaced
  +---------------------------+

A sensible delivery checklist:

  1. Envelope encryption everywhere, with a fresh data key per secret version and the KEK in an HSM or KMS with no export.
  2. Verify the AEAD tag on decrypt and fail closed on mismatch. Never log the reason a decrypt failed in a way that reveals plaintext.
  3. Treat data key reuse across secrets as a critical bug, and make the code path that generates them impossible to get wrong.
  4. Prefer dynamic credentials over stored passwords wherever the backing system supports them.
  5. Give every credential a lifetime, and make anything long-lived a deliberate exception with a review date.
  6. Audit every read, write, delete, and denial, with the workload identity rather than a shared username.
  7. Hash-chain the audit log and ship it somewhere the vault cannot rewrite.
  8. Cache unwrapped data keys for a bounded period in the vault process, and accept that exposure deliberately.
  9. Keep secrets off the synchronous read path where a failure would take down a service; read them lazily.
  10. Separate policies by blast radius, and make a policy lookup failure deny rather than allow.
  11. Never put a secret in a log, an error message, a crash dump, a metric label, or a URL.
  12. Run the HSM-unavailable and revoke-under-load drills, because those are the two failures that actually happen.

Common mistakes

Mistake What actually happens Better decision
Same key encrypts every secret One key leak exposes everything Fresh data key per secret version
Rotating the KEK by re-encrypting all data A routine rotation becomes a maintenance window Rotate the KEK, re-wrap the data keys
Reusing an AES-GCM nonce Plaintext XOR leaks and the auth key is destroyed Random 96-bit nonce per encryption, stored with ciphertext
Decrypting and ignoring the AEAD tag Integrity is not verified at all Verify the tag, fail closed
HSM unwrap on every read The vault is HSM-bound by design Cache data keys in memory, bounded TTL
Software key store with a plaintext KEK file The KEK is in every backup and core dump KMS or HSM, or a key file that is itself protected
Storing a static database password It is valid forever once leaked Dynamic roles with a short lease
Rotation job instead of a credential lifetime Rotations get skipped silently Expiry is the primary mechanism
Auditing writes but not reads Exposure cannot be scoped during an incident Audit reads with workload identity
Audit log writable by the vault A compromised vault destroys its own evidence Hash-chain, ship off-host
Policy lookup failure falls back to allow One bug hands out the whole namespace Fail closed, and alert on the denial
Unbounded audit buffer on the read path A logging outage becomes an application outage Bounded buffer, async, visible drops
Secrets read at boot, synchronously A vault blip stops every service starting Read lazily where possible
Long local cache of plaintext secrets The long-lived copy the design was meant to remove Cache data keys, not secrets
One global fallback policy A 1-hour dynamic credential and a permanent API key get the same treatment Per-class caching and retry policy
Secrets in metrics labels or URLs They end up in a monitoring system nobody threat-modelled Structured, redacted, never in telemetry
HSM as the only source of truth Vendor API outage is a total outage Cached data keys, tested degradation path

The complete story in one minute

Start with the insight that encryption at rest is only the first layer. Services still need plaintext or a signing/decryption operation at runtime, so design for short-lived access, bounded caches, revocation behaviour, and audited use rather than pretending decryption is rare everywhere.

Envelope encryption is the design. A fresh random data key encrypts each secret version with AES-GCM; the data key is then wrapped by a root key that lives inside an HSM and is never exportable. That gives small blast radius per key, rotation that costs a re-wrap rather than a re-encrypt, backups that are pure ciphertext, and a data key that is not in any process’s memory except briefly.

What the HSM actually buys is not performance. It is that the root key is not in the vault’s heap, not in a core dump, not in swap, and not in a backup. The cost is a round trip, so the vault caches unwrapped data keys in memory for a bounded window and accepts that exposure as a deliberate, small, and finite risk.

Then the parts that reduce blast radius. Dynamic credentials — for example, a generated database role with an hour-long lease — replace application-password rotation with issuance, renewal, revocation, and cleanup. They do not remove rotation entirely: the secrets system’s privileged database credential, signing keys, role definitions, and revocation path still need ownership and rehearsal. Short-lived certificates make the same trade for TLS. Audit reads as well as writes when exposure analysis requires knowing which workloads obtained a credential, and make the audit failure policy honest.

Availability is the constraint that shapes everything. A vault that is a boot-time dependency of ten thousand services has turned a security control into a single point of failure, so serve from a bounded data key cache, read secrets lazily, and be explicit per secret class about what happens when the HSM, the vault, or the network is down.

write: plaintext -> DEK -> ciphertext, DEK -> HSM -> wrapped DEK, store ciphertext
read:  ciphertext + wrapped DEK -> HSM -> DEK -> decrypt -> plaintext, briefly
every read -> audit record

The hard part was never encrypting a password. It was making sure the password was only ever in the clear at the two places that needed it, for the shortest time, with a record saying so.

Technical references

Keep reading
Browse everything