fsync: The Flush Between a Commit and Stable Storage
How PostgreSQL separates writes, WAL flushes, synchronous commit, hardware truthfulness, and group commit—and where each durability promise can fail.

A commit returns success. The client gets 200 OK. Then the machine loses power, and the transaction is gone. Under PostgreSQL’s durable defaults and a storage stack that truthfully honors flushes, that should not happen. When it does, the investigation begins with asynchronous/non-durable settings, the exact wal_sync_method, and hardware or virtual storage that acknowledged persistence it did not provide.
The gap between a successful write and a durable one is the entire subject of this article, and it is smaller than it looks and larger than most teams budget for.
The other end of the same cost is the network, and the two get confused in almost every postmortem. A commit record, a group commit, and a batched write are all the same idea applied to a different bottleneck, which is why batching round trips is worth reading alongside this. If you replicate, the durability question is also a placement question, and the standby has a lag of its own: replica lag and read-your-writes.
What a successful write actually means
When your database calls write(), it is asking the kernel to put bytes in the page cache. That is a memory copy. The call returns, and the durability question has not been asked yet.
application kernel disk
| | |
| write(data) | |
|----------------->| copy into page cache |
| | | <-- returns here
|<-----------------| |
| "success" | |
| | ... some time later |
| |------------------------>|
| | | write-back, maybe
Two separate things can be lost here, and they are worth separating because they have different causes.
The operating system may not have flushed its buffer. Background writeback timing is workload and kernel dependent and is not a durability contract.
The storage stack can add controller, drive, hypervisor, and network caches. A reliable path must either make those caches non-volatile or honor the issued flush/barrier all the way down. Product labels are not evidence; use provider durability guarantees and destructive power-loss testing where the operating model permits it.
So an ordinary successful write is not yet a durability promise. PostgreSQL uses the configured WAL synchronization method—sometimes an fsync/fdatasync call, sometimes synchronous-open semantics—to request stable storage. The request is only as truthful as every layer below it.
The two knobs people confuse
Postgres exposes two settings and they are routinely discussed as one. They are not, and the difference between them is the difference between losing data and losing a database.
fsync = on (default)
the operating system is instructed to put WAL on stable storage
before the commit is acknowledged
synchronous_commit = on (default)
the fsync happens at COMMIT, not later at a checkpoint
fsync = off
PostgreSQL stops issuing required synchronization requests
-> an operating system crash can leave the WAL inconsistent
-> the database may not start, or may be corrupt
-> unrecoverable, not merely missing rows
synchronous_commit = off
commits return before the WAL is flushed
-> a crash loses the transactions not yet flushed
-> recovery still works, nothing is corrupt
-> PostgreSQL bounds the normal risk window; the business loss count varies with throughput
The distinction is the whole answer to “can I turn it off”. Turning off synchronous_commit makes the database faster and risks losing recent committed transactions. Turning off fsync makes the database faster and risks a database you cannot start.
fsync = off belongs only where the entire cluster can be discarded and rebuilt from another durable authority. Labeling a node “replica” is insufficient if it might be promoted or contains the only copy of acknowledged work.
How much you actually lose
fsync is not the setting that governs how much you lose. The WAL writer delay is.
The WAL writer periodically flushes unwritten WAL. The documented default wal_writer_delay is 200 milliseconds, but PostgreSQL states that the maximum asynchronous-commit risk window is three times that delay because the writer can favor whole WAL pages during busy periods. That is a time window, not “three transactions”; the number of lost commits depends on throughput inside it.
synchronous_commit = on
COMMIT -> WAL sync -> return 200 durable if the storage contract is truthful
synchronous_commit = off, wal_writer_delay = 200ms
COMMIT -> return 200 <--------+
COMMIT -> return 200 |
COMMIT -> return 200 | all three un-fsynced
... |
(WAL writer wakes) ------------+
WAL sync; commits in the risk window may be lost after a crash
So the exposure is roughly the writer delay plus whatever is sitting in the buffers, and it is a number you can set and reason about. The Postgres documentation is careful to say the loss is limited to transactions not yet flushed, which is the important part: you lose recent writes, not the database, and recovery is clean.
The cost of that setting is real and it is worth measuring rather than assuming. An fsync is a round trip to the storage, so removing it from the commit path is the difference between a commit that takes a storage latency and one that takes a memory write. On a local NVMe drive that is a few hundred microseconds. On a network-attached cloud volume it can be several milliseconds, and there the setting is the difference between a service that handles its traffic and one that falls over at 200 requests per second per connection.
Group commit is the real answer
The instinct is to make the flush cheaper. The better instinct is to make fewer flushes do more work, because a flush is not expensive per byte, it is expensive per call.
Group commit is the mechanism every serious database uses, and it is why a thousand concurrent commits do not cost a thousand flushes.
no group commit
COMMIT -> fsync (1.2 ms) -> return
COMMIT -> fsync (1.2 ms) -> return
COMMIT -> fsync (1.2 ms) -> return
1000 commits: 1200 ms, 1000 flushes
group commit
COMMIT 1 begins the WAL flush
COMMIT 2..40 arrive while that flush is in progress
the flush completes, all 40 are released together
1000 commits: ~30 flushes, same wall clock, 1.2 ms each
The arithmetic is the important part. Group commit does not make a flush faster. It makes the flush rate the limit rather than the transaction rate, so throughput scales with how many commits you can fit into one flush window.
This has a direct consequence for how you load test. A benchmark that commits serially measures flush latency. A benchmark with realistic concurrency measures flush throughput, which is often an order of magnitude higher per core. Testing durability on a single-threaded client tells you almost nothing about how the database behaves under load, and it makes durability look far more expensive than it is.
It is also why write-ahead logging matters. Commit durability advances a sequential log position rather than forcing every changed data page to stable storage before commit. The workload is still sensitive to sync latency and WAL volume; it is not “key order,” and checkpoints later own data-page writeback.
Managed storage changes the evidence
Local and managed volumes implement persistence differently. A sync may traverse a filesystem, hypervisor, network, replicated storage service, controller, and device cache. Location and replication guarantees are provider- and volume-specific, so latency cannot be inferred from the word “cloud.”
compare on the exact production storage class
p50 / p95 / p99 WAL sync latency
sustained rather than burst throughput
behavior during provider throttling and failover
documented acknowledgement and replication semantics
Two consequences follow. First, storage sync latency becomes part of synchronous commit latency. Second, provisioned limits, burst models, noisy neighbors, or failover can change the tail without any schema change. A short benchmark may measure burst entitlement rather than sustained capacity.
The honest summary is that a durable commit waits for the storage contract behind the WAL volume. On network-backed storage that includes a remote service; on local durable media it may not. Measure the deployed path instead of substituting either architecture by assumption.
The benchmark that measures nothing
The most common durability experiment is to run a commit loop against a database on a filesystem mounted tmpfs, or to copy the data directory onto a memory-backed volume, or to benchmark a container with no real disk in the write path.
On tmpfs, a sync cannot establish survival across loss of the memory-backed filesystem. The result can be a useful upper bound on non-storage overhead, but it is not a durability benchmark.
Worth being explicit about what a real measurement needs:
on a real block device, flushed, no write-back cache
time N commits at concurrency C, report p50 and p99
vary C from 1 to 64
the curve is the group commit behaviour
a flat curve means flush throughput is the limit
If the number barely moves as concurrency rises, the flush is the bottleneck and the setting matters enormously. If throughput scales nearly linearly with concurrency, the flush cost is being amortised and the durability decision is cheap. Either way you learn something. The tmpfs number teaches you nothing.
What this means to the application
There is a part of this that is a correctness problem rather than a performance one, and it does not belong in the database configuration discussion.
If a commit can be acknowledged and then lost, then some user-visible statement was a lie. A checkout that returns a confirmation and then vanishes after a crash is not a performance incident, it is a broken promise, and the customer impact lands on support rather than on the on-call rotation.
// the application cannot tell the difference between these
// on a connection with synchronous_commit = off
await using var tx = await conn.BeginTransactionAsync();
// safe: this returns only after the WAL is on stable storage
await tx.CommitAsync();
// this returns before durability, and a crash can undo it
// -> the order row is gone
// -> the payment was captured
// -> the inventory decrement is gone
// -> reconciliation finds a paid order with no stock movement
The uncomfortable middle ground is a payment flow. A card charge that succeeded and an order row that was lost is a real operational problem with a real cost, and it is worth being explicit that synchronous_commit = off is what makes it possible. It is a good setting for a read-heavy catalogue and a dangerous one for a write that represents a financial event.
The general rule is that the durability decision belongs to the business, not to the performance budget. It is a decision about what a customer is told.
A production-ready architecture
COMMIT
|
v
+---------------------------+
| synchronous_commit = on |
| durability before ack |
+---------------------------+
| signal WAL writer
v
+---------------------------+ group commit:
| flush WAL to stable | one flush releases
| storage, release all | every waiter in the
| waiting committers | window
+---------------------------+
|
v
+---------------------------------------------------------------+
| storage requirements: |
| sequential WAL, so provisioned IOPS, not throughput |
| power loss protection on the volume |
| p99 flush latency measured, not p50 |
+---------------------------------------------------------------+
+---------------------------------------------------------------+
| per-workload decision: |
| catalogue / analytics -> synchronous_commit = off is fine |
| orders / payments -> on, and it is not a tuning knob |
+---------------------------------------------------------------+
A delivery checklist:
- Decide the durability requirement per workload, not once for the instance. A catalogue and a payment ledger have different correct answers.
- If you turn off
synchronous_commit, setwal_writer_delaydeliberately and document the exposure window you have accepted. - Turn off
fsynconly for a cluster you are prepared to discard and recreate; do not promote or treat that node as durable. - Measure flush latency at your p99 on the actual volume, with the actual provider, with the credit bucket in whatever state production will be in.
- Benchmark durability at realistic concurrency. A serial commit loop measures flush latency and tells you nothing about throughput.
- Verify the provider/device durability contract and flush behavior; power-loss protection is one implementation, not the only possible storage contract.
- Provision for measured WAL bytes, writes, sync latency, and sustained limits rather than reducing the requirement to one IOPS number.
- Alert on flush latency, not only on commit latency. The difference between them is the durability budget you are spending.
- If you accept bounded loss, make the reconciliation path real. Something must find the transactions that were acknowledged and never persisted.
- Re-measure after any storage migration. Flush latency is a property of the volume, and a new volume type is a new durability profile.
Failure stories worth testing
Kill the machine after a commit returns and check whether the row exists
Use a provider where you can hard-stop a VM, or an fsync wrapper that discards unflushed writes. The result is the single most useful number in this article: the exact set of acknowledged-and-lost transactions. Measure it under concurrency, because the loss window is a function of how many commits are in flight.
Benchmark commits at concurrency 1 and then at 64
Two numbers from the same loop. The ratio between them is the group commit effect, and it is usually a large multiple. This also tells you whether your storage is the bottleneck or your application is.
Set synchronous_commit = off, commit at concurrency, then cut power
Confirm the database starts and recovers cleanly rather than failing to start. That difference, clean recovery versus an unrecoverable cluster, is the entire reason the two settings are not interchangeable.
Benchmark on a data directory on tmpfs and on a real volume
Run the identical loop against both. If the tmpfs result is an order of magnitude better, every durability number you have collected on it is fiction. Do this once so nobody has to argue about it again.
Fill a cloud volume’s IOPS credit bucket during a test, then re-run
Throughput drops to the baseline rate, and it looks like a database regression because the application and the schema did not change. This is a real and recurring cause of “it was fast last week”.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| Treating a successful write as durable | The bytes are in the page cache, and the drive may not have moved | Ask what the storage guarantees on power loss |
Setting synchronous_commit and fsync together |
One risks lost commits, the other risks an unrecoverable database | Treat them as two separate decisions |
fsync = off on a primary |
An OS crash can leave WAL inconsistent and the database unopenable | Replica or rebuildable node only |
| Treating async-commit loss as one writer interval | PostgreSQL documents a maximum risk window of three times wal_writer_delay |
Document the time window and business events per second |
| Benchmarking commits serially | You measure flush latency and conclude durability is unaffordable | Benchmark at realistic concurrency |
Measuring fsync on tmpfs |
It is nearly a no-op, so the number is meaningless | Measure on a real flushed block device |
| Assuming a local disk’s flush cost applies to a cloud volume | Network storage is a round trip to another service, with a long tail | Measure p99 on the actual provider and volume type |
| Trusting a write-back cache that lies about flush completion | PostgreSQL cannot make a false hardware acknowledgement durable | Use a storage path with a verified persistence contract |
| Loading a credit-bucketed volume for a benchmark | The test spends burst credits and reports throughput you cannot sustain | Test at the sustained baseline rate |
| Acknowledging a payment before durability | A crash produces charged cards with no order row | Keep synchronous commit on for financial writes |
| Assuming more IOPS means faster commits | WAL is sequential, so commits are latency-bound | Provision for committed IOPS, not throughput |
| No plan for the transactions that were acknowledged | Lost commits are discovered by customers or by nobody | Build the reconciliation path before you need it |
The complete story in one minute
An ordinary buffered write() can complete in the operating-system cache. PostgreSQL must then use its configured sync method, and the storage stack must honor that request. There are two separate settings because there are two separate risks. Turning off synchronous_commit allows a crash to lose acknowledged transactions whose WAL was not flushed; PostgreSQL documents a normal maximum window of three times wal_writer_delay, while recovery remains consistent. Turning off fsync disables the coordination that keeps WAL and data writes recoverable and can corrupt the cluster after an OS or hardware crash.
The reason durability is affordable at all is group commit, not a cheaper flush. A flush is expensive per call and cheap per byte, so a database that batches the commits arriving during a flush window pays for one flush and releases every waiter together. Throughput therefore scales with the flush rate rather than the transaction rate, which is why concurrent commits cost far less than the flush latency multiplied out, and why benchmarking durability with a single-threaded client produces a number that is both pessimistic and irrelevant.
That number belongs to the deployed storage path, PostgreSQL sync method, WAL volume, and workload. Measure p50 through p99 sync time using pg_stat_wal, test sustained rather than only burst capacity, and use pg_test_fsync only as one storage diagnostic rather than an application-throughput forecast. Then decide per workload whether an acknowledged transaction may disappear, because a checkout confirmation followed by a missing order is not a tuning detail. It is a promise the system did not keep.


