RabbitMQ Between Microservices
Designing RabbitMQ service boundaries with owned queues, publisher confirms, transactional outbox and inbox records, bounded retries, backpressure, and recovery tests.

Moving a call from HTTP to RabbitMQ removes a service from the caller’s latency path. It does not remove coupling. The producer still depends on a contract, the consumer still depends on delivery, and somebody still owns what happens when the database commits but the acknowledgement does not.
That is the useful framing for RabbitMQ between services: not “make it asynchronous,” but “move work across a durable boundary while making every ambiguous outcome recoverable.”
Consider checkout. The Orders service owns order acceptance. Billing, inventory, and notifications react afterward. Checkout should not wait for three downstream APIs, but each downstream action needs independent failure and recovery.
Draw ownership before topology
I start with the business messages:
OrderAcceptedis a fact: order123was accepted under contract version2.ReserveInventoryis a command: the inventory capability is asked to perform one action.PaymentCapturedis another fact produced by billing.
Events are named in the past tense and can have several interested services. Commands have one logical owner. RabbitMQ cannot enforce those semantics; exchange and queue topology should reflect them.
Orders outbox
|
v
orders.events (topic exchange)
| order.accepted.v2
+--------------------+----------------------+------------------+
v v v
inventory.order-events billing.order-events notifications.order-events
| | |
Inventory service Billing service Notification service
Each consumer capability owns its queue, binding, concurrency, retry policy, dead-letter destination, and alert. A shared queue would load-balance one copy between consumers; that is correct for several instances of the same logical consumer, not for three services that all need the event.
The producer owns the event meaning and compatibility policy. Platform automation may create broker objects, but it should not invent bindings without the consuming team’s contract.
Choose the exchange for routing, not durability
Exchanges route publications to queues:
- a direct exchange matches an exact routing key;
- a topic exchange matches patterns such as
order.*; - a fanout exchange copies to every binding regardless of key;
- a headers exchange routes from header values.
Durability comes from the queue type, durable topology, message persistence, and broker configuration—not from choosing a topic exchange.
An unroutable message is another failure boundary. If a publisher sends a routing key with no matching queue and does not request returned unroutable messages, the publication can disappear even though the broker connection was healthy. Use the mandatory publishing flag (or the client equivalent), handle returned messages, and treat an unexpected return as a production error.
Publisher confirms answer one narrow question
A publisher confirm tells the producer that RabbitMQ accepted responsibility for a publication. It is not the consumer’s acknowledgement, and it does not mean the business action completed.
For a persistent message routed to a durable classic queue, acknowledgement depends on the broker persisting it according to its rules. For a quorum queue, a confirm is issued after the message has been replicated to a quorum. The application still has to correlate confirms, negative acknowledgements, channel closure, and timeouts with the outstanding batch.
There is an unavoidable ambiguity:
publisher sends -> broker accepts -> network drops -> publisher sees no confirm
The publisher cannot know whether retrying will duplicate the message. Therefore every message needs a stable ID, and consumers must be idempotent. Confirms reduce loss; they do not eliminate duplicates.
Avoid publishing one message and synchronously waiting for one confirm on the hot path. Batch or pipeline confirms with a bounded in-flight set so throughput does not become one network round trip per message. Bound it because an unlimited outstanding set turns a broker slowdown into application-memory growth.
The producer’s database and broker are not one transaction
The dangerous sequence is straightforward:
1. commit order to database
2. publish OrderAccepted
If the process dies between those lines, the order exists and no event follows. Reversing the order publishes an event for a database transaction that may roll back.
Use a transactional outbox in the Orders database:
begin;
insert into orders (order_id, status, accepted_at_utc)
values (:order_id, 'accepted', :accepted_at_utc);
insert into outbox_messages (
message_id,
aggregate_id,
message_type,
contract_version,
payload,
occurred_at_utc
)
values (
:message_id,
:order_id,
'order.accepted',
2,
:payload,
:accepted_at_utc
);
commit;
A relay claims unpublished rows, publishes them with confirms, then records publication. It may crash after the confirm and before marking the row published, so it will publish again. That duplicate is expected.
The outbox also needs operations: claim leases that expire after a crashed relay, bounded batches, oldest-unpublished age, attempt count, and retention after confirmed publication. A table that grows forever eventually turns reliability into database pressure.
Acknowledge only after durable consumer work
RabbitMQ consumer acknowledgements tell the broker whether it may remove a delivery. They are independent of publisher confirms.
The safe consumer sequence is:
delivery
-> begin database transaction
-> insert message ID into inbox (unique constraint)
-> apply business state change
-> commit database transaction
-> acknowledge delivery
If the process dies before commit, the database rolls back and RabbitMQ redelivers. If it dies after commit but before acknowledgement, RabbitMQ redelivers and the unique inbox record identifies the duplicate. The consumer acknowledges without applying the state change twice.
The inbox record and business update must share the same local database transaction. Recording the ID in Redis first, or in a separate database after the write, reopens a failure window. Keep the record long enough to cover broker redelivery and operational replay policy; deleting it too soon converts an old redelivery into new work.
Do not hold a database transaction open while calling an unrelated remote service. If consuming an event must cause another asynchronous side effect, update local state and write another outbox message in the same transaction.
Retry is a classification, not a loop
Failures fall into different classes:
| Failure | Example | Action |
|---|---|---|
| Transient infrastructure | database failover, brief dependency timeout | delayed, bounded retry |
| Rate or capacity pressure | downstream quota, local saturation | backoff and reduce concurrency |
| Invalid contract | unreadable JSON, missing required field | dead-letter immediately |
| Permanent business rejection | referenced entity cannot legally transition | record outcome; usually no retry |
| Consumer defect | repeatable null reference | stop churn, alert, quarantine |
Immediate requeue can create a hot loop: the same poison message is delivered, fails, and is placed back at the queue head repeatedly. Use delayed retry queues, dead-letter routing with expiry, or a retry scheduler. Preserve the original message ID and increment trusted retry metadata at infrastructure boundaries; do not let arbitrary producer headers bypass limits.
After the last attempt, move the delivery to a dead-letter queue or parking-lot queue with enough context to diagnose it. Do not put secrets or raw credentials in failure metadata.
A dead-letter queue is not recovery by itself. Define who is paged, how the message is inspected, whether the underlying state is still valid, how replay is authorized, and how inbox idempotency behaves on replay.
Backpressure begins with prefetch
Consumer prefetch limits the number of unacknowledged deliveries RabbitMQ sends. It should be large enough to keep workers busy and small enough to bound memory and redelivery after a crash.
Start from measured processing time and concurrency, not a copied constant. If one instance runs 16 handlers and each message can occupy several megabytes, a prefetch of 1,000 creates very different risk than it does for 1 KB jobs.
Keep handler concurrency bounded. An asynchronous consumer that starts an unlimited task per delivery merely moves the queue into application memory. When a dependency slows, outstanding work, connection-pool demand, and acknowledgement latency grow together.
Observe consumer capacity/utilization, unacknowledged count, acknowledgement rate, processing duration, and the oldest ready message. Queue depth alone misses a small number of very old stuck messages and a large in-flight set.
Ordering is local and expensive
RabbitMQ queues are ordered, but observable processing order changes with multiple consumers, redelivery, priorities, and retry paths. If one aggregate requires ordered transitions, make that requirement explicit.
Options include routing an aggregate to one shard queue, adding an aggregate version and rejecting/gapping out-of-order messages, or designing commutative/idempotent state changes. A single global consumer preserves more order by sacrificing throughput and availability; it is rarely the right default.
The event should carry the producer’s aggregate version when ordering matters. The consumer can then distinguish a duplicate, the next transition, and a gap that requires delayed retry or reconciliation.
Choose the queue type from the workload
Quorum queues are the replicated queue type for data-safety-focused workloads. They trade additional resources and replication work for stronger failure tolerance. They are not the answer for every temporary queue, extremely long backlog, lowest-latency path, or high-fanout use case.
Classic queues may remain appropriate where their documented durability model and operating constraints match. RabbitMQ streams fit append-oriented, replayable, high-throughput logs better than treating a queue as indefinite history.
Make the decision from:
- acceptable message loss and node-failure behavior;
- backlog length and message size;
- replay requirement;
- publish and consume throughput;
- number and independence of consumers;
- storage and network capacity;
- RabbitMQ version and supported queue features.
Do not use a message broker as archival storage accidentally. Retention and replay are product requirements, not consequences of a consumer being offline for months.
Contract evolution
Messages cross deployment boundaries and can wait in queues while producers and consumers run different releases.
Give the envelope a stable message ID, message type, schema/contract version, occurrence time, correlation/causation IDs, and trace context. Keep business fields in a documented payload. Consumers should ignore additive fields they do not understand and reject incompatible meaning explicitly.
Changing an enum’s meaning, reusing a field for a new unit, or making a previously optional value mandatory is not additive. When semantics change, publish a new version and run compatibility tests against retained examples from supported producers.
Production evidence
The dashboard should make the whole flow visible:
- outbox rows and oldest-unpublished age;
- publish confirm latency, negative confirms, returned messages, and channel failures;
- ready and unacknowledged messages by queue;
- oldest-message age and redelivery rate;
- consumer processing and acknowledgement latency;
- retry traffic and dead-letter rate by reason;
- inbox duplicate count;
- end-to-end age from business commit to consumer completion;
- node, disk, memory, file descriptor, and network alarms.
Then rehearse the commit boundaries:
- Kill the API after the order/outbox transaction commits.
- Kill the relay after the broker confirms but before the outbox row is updated.
- Publish with a missing binding and verify the mandatory return is handled.
- Kill a consumer after its database commit but before acknowledgement.
- Lose a broker node while publishing to the selected queue type.
- Slow the consumer database and confirm prefetch plus concurrency remain bounded.
- Send a poison message and prove retries stop.
- Replay a dead-letter twice and verify the business outcome stays singular.
- Deploy an older and newer contract consumer against the same queue examples.
- Fill disk or trigger the broker’s resource alarms in a controlled environment and verify publishers back off.
The boundary
RabbitMQ provides routing, buffering, acknowledgements, confirms, and several durability models. It cannot make two independent databases and a broker commit as one unit. The application closes those gaps with stable identities, outbox and inbox transactions, idempotent state transitions, and reconciliation.
Moving the HTTP call off the request path was one task. Making every publish, delivery, retry, and replay safe under an ambiguous crash was the service boundary.
