Replacing Self-Managed Kafka with Azure Event Hubs
An end-to-end migration article covering compatibility, partition mapping, dual running, cutover, rollback, and the operational model after Kafka.

The system was a Media/Workstation application platform. Workstation clients and backend services produced a continuous stream of operational events: workstation status, media-ingest progress, render and processing state, license activity, and audit records. Other applications consumed those events to update operational views, search indexes, reporting stores, notification services, and the central data platform.
Exact production counts are omitted here. What matters is the shape of the system: many producers, many consumer groups, many topics and partitions, sustained traffic with peaks, and downstream processes that could not all be stopped at the same time.
Kafka was the transport between those systems.
The migration did not begin because Kafka had stopped carrying the traffic. It began because the organization was operating every part of the Kafka platform: brokers, storage, replication, upgrades, security, monitoring, capacity, and disaster recovery. The same teams also had to develop the applications that depended on it.
The assignment was therefore not “replace Kafka with an Azure product.” It was to change the operating model without breaking the event contracts and processing guarantees already in use.
How the existing system worked
The original data path was conventional Kafka, but its behaviour matters to the migration.
Workstation clients Media and backend services
- health/status - asset lifecycle
- usage/telemetry - render/processing state
- user activity - licensing and audit
| |
+----------- Kafka clients --------+
|
v
+-----------------------------+
| Self-managed Kafka platform |
| |
| topics -> partitions |
| leaders -> replicas |
| retained event log |
+-----------------------------+
| | |
group:ops group:search group:data
| | |
v v v
Operations Search Data platform
|
+----> notification/integration services
A producer selected a topic and, where ordering mattered, supplied a business key such as workstation ID, asset ID, or processing-job ID. The Kafka client used that key to select a partition. Records with the same key therefore reached the same partition, and Kafka preserved their order within that partition.
Each downstream application used its own consumer group. The operations service, search indexer, reporting pipeline, and data-platform ingestion process could all read the same event stream independently because each group maintained its own position. If a consumer stopped, Kafka retained the records and the group could resume from its committed offset. If a downstream store had to be rebuilt, the consumer could replay retained events.
This gave the platform four properties that could not be lost during migration:
- independent consumption by multiple applications;
- order within a business-key partition;
- recovery from committed positions after restarts;
- replay while records remained inside the retention window.
Kafka did more than move bytes. It defined the failure and recovery behaviour of the applications around it. Any replacement had to be assessed against those behaviours, not against a product-feature checklist.
What had become difficult
The problem was the amount of infrastructure ownership behind that data path.
Traffic growth affected broker CPU, network, disk throughput, disk capacity, partition placement, and replication. Retention was not only a topic setting; it became a storage forecast. Adding broker capacity also meant redistributing partitions without overwhelming the network or disks. Maintenance required rolling changes and observation of replica health throughout the change.
The platform team was responsible for:
- broker and controller availability;
- operating-system, JVM, and Kafka maintenance;
- disk expansion, replacement, and retention capacity;
- partition balance, leader distribution, and replica health;
- client certificates or SASL credentials and Kafka ACLs;
- network access to every broker endpoint;
- cross-zone or cross-region recovery architecture;
- broker, producer, consumer, and consumer-lag monitoring.
Kafka could continue to support the workload. The concern was whether maintaining all of that infrastructure remained the best use of the engineering team.
That produced a precise requirement: retain partitioned streaming and Kafka-client compatibility where possible, but move broker operations to a managed service.
The service decision: Event Hubs, Event Grid, or Confluent?
These services solve different problems. Treating them as interchangeable would have produced the wrong architecture.
Why Event Grid was not the Kafka replacement
Event Grid is suited to discrete notifications: a resource was created, an asset finished processing, or a workflow changed state. Subscribers react to the notification, normally through push delivery or an event subscription.
The Kafka platform, however, carried ordered event series, supported independent consumer positions, and allowed replay from a retained log. Event Grid does not provide the same partitioned-log and consumer-offset model. It was therefore not a replacement for the main stream.
Event Grid still had a role in the new design. Once a stream-processing service had consumed and validated a business state change, such as AssetProcessingCompleted, it could publish a small CloudEvent to Event Grid for Azure Functions, webhooks, or other reactive handlers. The raw processing events and their replayable history remained in Event Hubs.
The separation was deliberate:
- Event Hubs: high-volume ordered streams, consumer groups, retention, and replay;
- Event Grid: discrete notifications and event routing after a meaningful state change had been established;
- Service Bus, where required: commands or work items needing queue semantics, dead-lettering, duplicate detection, or transactional messaging features.
Using each service for its own delivery model kept the Event Hubs migration focused.
Why Confluent Cloud was considered
Confluent Cloud was a credible option because it provides managed Apache Kafka. It removes broker ownership while preserving native Kafka behaviour and the wider Kafka ecosystem. It is the stronger candidate when applications depend on Kafka Streams, ksqlDB, a substantial Kafka Connect estate, Kafka transactions, specific broker or topic behaviour, or other features for which protocol compatibility is insufficient.
Its Azure integration also means that choosing Confluent is not the same as rejecting Azure. Confluent resources can be provisioned through Azure integrations, and the service supports Azure-oriented identity, billing, and connector workflows.
The deciding question was not vendor preference. It was dependency depth.
The migration scope was dominated by applications using the standard producer API, consumer API, partitions, consumer groups, and committed offsets. Those applications needed Kafka protocol compatibility, not ownership of Kafka brokers or the complete Kafka platform. Event Hubs could satisfy that narrower requirement while placing the streaming resource, identity, network controls, diagnostics, and capacity management in the Azure operating model already used by the surrounding applications.
Workloads that failed that compatibility test should not be forced onto Event Hubs. They should be redesigned deliberately or remain on native Kafka, including Confluent Cloud if managed Kafka is the appropriate destination.
Why Event Hubs was selected for the main streaming path
Event Hubs provided the combination the main workload required:
- a Kafka-compatible endpoint for existing producer and consumer clients;
- partitioned, retained event streams with independent consumer groups;
- managed service capacity instead of broker and disk administration;
- Microsoft Entra ID and Azure RBAC for application identities;
- private networking through Azure network controls;
- Azure Monitor metrics and diagnostic logs;
- Event Hubs Capture for writing new stream data to Blob Storage or Azure Data Lake Storage;
- AMQP and Kafka access to the same event hubs, allowing existing Kafka clients and Azure-native consumers to coexist.
This did not make Event Hubs equivalent to Kafka. It made it a viable target for the part of the estate that used Kafka as a partitioned event transport.
Old architecture and target architecture
The old architecture concentrated transport, retention, consumer position, and operational responsibility in the Kafka platform.
OLD
Producers
|
| Kafka protocol + existing credentials
v
+--------------------------------------------------+
| Self-managed Kafka |
| brokers | replication | disks | topics | offsets |
+--------------------------------------------------+
| | |
v v v
Operational Business Data/analytics
consumers consumers consumers
Platform team owns:
broker lifecycle, storage, replication, upgrades,
ACLs, capacity, monitoring, and recovery platform
The target architecture separated streaming, archival, and discrete notifications.
NEW
Workstation and service producers
|
| Kafka protocol, SASL_SSL, Entra ID where supported
v
Private endpoint / controlled network path
|
v
+--------------------------------------------------+
| Azure Event Hubs |
| namespaces -> event hubs -> partitions |
| retained streams -> Kafka consumer groups |
+--------------------------------------------------+
| | |
| | +--> Capture --> ADLS
| |
v v
Existing Kafka State/notification
consumers publisher
|
v
Event Grid
|
Functions / webhooks
Azure operates:
service infrastructure, storage layer, replication,
patching, and broker-equivalent service maintenance
Platform team still owns:
event contracts, namespaces, capacity, partitions,
identities, network policy, clients, monitoring,
recovery procedures, and cost controls
Namespaces were designed as operational boundaries, not as folders. Workloads with different security rules, traffic profiles, criticality, or failure impact could be placed in separate namespaces. That prevented a high-volume telemetry stream from sharing every capacity and access decision with a lower-volume but business-critical control stream.
Kafka topics mapped initially to event hubs on a one-to-one basis. That made application configuration and validation easier. It was a starting rule, not an absolute one; unused topics were excluded, and related low-volume event types could be consolidated only if their contracts, access, retention, and partitioning requirements were compatible.
Partition counts were recalculated rather than copied. The design considered peak ingress, fan-out egress, required consumer parallelism, key distribution, and ordering. A topic with many historical partitions was not automatically provisioned with the same count. Conversely, a low average rate did not justify fewer partitions if peak traffic or required parallelism demanded more. Flows replicated with MirrorMaker were the exception during transition: their source partition identity had to be preserved. If a different partition model was required, it was introduced as a separate target stream and populated through an intentional repartitioning process rather than hidden inside mirroring.
How the work was planned
The work started with five engineering artefacts. None required changing production.
- Dependency register: every producer, topic, consumer group, owner, downstream effect, client library, and deployment location.
- Compatibility record: serializers, compression, message-size percentiles, partitioner, authentication, commit strategy, transaction use, Streams or Connect use, AdminClient calls, and retention requirements.
- Target design: namespace boundaries, event hubs, partitions, capacity tier, identity model, private connectivity, archive path, and monitoring.
- Migration-wave plan: flows grouped by dependency and risk, with a defined producer switch, consumer start position, data boundary, rollback point, and owner.
- Evidence plan: baseline metrics, test cases, acceptance thresholds, reconciliation queries, and cutover dashboards.
The dependency register was the most important. A topic could not be migrated safely because its name and partition count were known. The team also needed to know who wrote to it, who read it, whether the consumers produced another event, and which business state changed when those events were processed.
That graph determined migration order. A leaf telemetry flow with no business side effects could move early. A stream that fed several services and then produced further events moved only after every downstream dependency was understood.
Kafka compatibility was tested, not assumed
Event Hubs exposes a Kafka endpoint at <namespace>.servicebus.windows.net:9093. Clients use TLS through SASL_SSL. Authentication can use a shared access signature with PLAIN, or Microsoft Entra ID through OAUTHBEARER.
The connection change is small:
# Existing Kafka
bootstrap.servers=<kafka-brokers>
security.protocol=SASL_SSL
sasl.mechanism=<existing-mechanism>
# Event Hubs
bootstrap.servers=<namespace>.servicebus.windows.net:9093
security.protocol=SASL_SSL
sasl.mechanism=OAUTHBEARER # or PLAIN when SAS is deliberately used
The assessment around it is not small.
For each client family, the test covered client version, token handling, idle connections, metadata refresh, request timeout, maximum request size, compression, producer idempotence, partition selection, group coordination, offset commits, and rebalance behaviour.
Kafka features were placed into three groups:
| Classification | Meaning | Migration action |
|---|---|---|
| Supported and verified | Behaviour passed functional and load tests | Configuration-led migration |
| Supported with a change | Identity, compression, payload, timeout, or consumer logic required modification | Change application, then retest |
| Broker/platform dependent | Application relied on a Kafka capability that was unsuitable or unavailable on the selected Event Hubs tier | Redesign or retain on native Kafka |
This prevented the phrase “Kafka compatible” from becoming an architecture decision by itself.
Designing capacity and retention
Broker count was not converted into Event Hubs capacity. The workload was measured from the clients and the streams.
For each event hub:
peak ingress = events per second x encoded event size
peak egress = ingress x independently reading consumer groups
retained data = average ingress x retention duration
The model also included producer bursts, batch overhead, compression, retry traffic, consumer catch-up after an outage, active connections, and uneven partition keys.
Fan-out was especially important. A stream written once and read completely by four consumer groups creates roughly four times the stream volume on the read side. Sizing only for producer ingress understates the service demand.
Retention was divided into two needs:
- operational replay: recent data required for consumers to recover or rebuild;
- historical record: data retained for audit, analytics, or long-range reprocessing.
Event Hubs retention served the first need. ADLS served the second. New events could be written to ADLS with Event Hubs Capture. Existing Kafka history required a separate migration decision; Capture does not reach backward into the Kafka cluster.
Moving old Kafka data
Not every retained record belonged in the new Event Hubs namespace. The data was classified before it was copied.
Data still needed for active replay
For topics whose retained records were required by live consumers, Kafka MirrorMaker 2 could replicate from Kafka into the green Event Hubs environment through the Kafka endpoint. Replication used an explicit allow-list of topics; it did not mirror the entire cluster by default.
The process was:
- create the target event hub with the source partition count and approved retention, preserving partition identity for replication;
- start MirrorMaker for the selected topic;
- record the source start boundary;
- allow it to copy the retained window and continue following new records;
- compare source and target counts by topic and time window;
- sample event keys, timestamps, headers, and payload hashes;
- measure replication lag until it remained within the cutover threshold.
The target offsets were treated as a new coordinate system. A Kafka offset identifies a position within one partition of one log; it is not a portable business identifier. MirrorMaker checkpoints can assist with position translation, but the translated start point still has to be tested for the specific client and flow. Target consumers were initialized at an agreed Event Hubs position, and reconciliation used event IDs, source timestamps, and business sequence numbers rather than assuming that numeric offsets would match.
Data needed only for history
Long-term audit or analytical history was exported by a controlled Kafka consumer pipeline to ADLS with its original event ID, source topic, partition, offset, event timestamp, key, headers, and schema version. That preserved provenance without consuming the operational retention window in Event Hubs.
The archive was validated independently and given a documented replay procedure. If historical events ever had to be reintroduced, a controlled replay producer would publish them with their original event identity so consumers could detect duplicates.
Data no longer required
Expired, test, orphaned, or unowned topics were not migrated merely because they existed. Their owners had to confirm retention or approve deletion through the normal governance process. Migration was not used as an excuse to copy an unknown estate into a new platform.
The red/green deployment model
For this migration, red referred to the existing Kafka path and green to the Event Hubs path. It followed the same principle commonly called blue/green deployment: build the replacement beside the live system, validate it, shift controlled traffic, and keep the old path available until rollback is closed.
RED: live GREEN: being proved
Producers -> Kafka Kafka -> MirrorMaker -> Event Hubs
| |
v v
active consumers shadow consumers
business writes on no business writes
Green consumers initially ran in shadow mode. They deserialized records, applied validation logic, and produced comparison results, but they did not update production databases, send notifications, or trigger user-visible actions. This allowed the team to compare behaviour without creating duplicate side effects.
The cutover unit was one complete event flow. Moving producers independently of the consumers and downstream applications that depended on them would have created an unowned data boundary.
For order-sensitive flows, a short controlled gate was used:
1. Stop or pause writes for the selected flow.
2. Wait until Kafka is stable and MirrorMaker lag reaches the agreed boundary.
3. Record the final red watermark for every source partition.
4. Stop red consumers at that boundary.
5. Change all producers for the flow to Event Hubs.
6. Start green consumers from the approved target position.
7. Resume writes and validate the first production events end to end.
That brief gate avoided concurrent writes for the same ordering key through both Kafka and Event Hubs. For flows without ordering or duplicate-sensitive side effects, producers could move gradually, but only when stable event IDs and consumer idempotency made overlap safe.
Deployment pipelines
The migration was not executed through portal changes during a maintenance window. It used separate but coordinated pipelines.
Infrastructure pipeline
The infrastructure pipeline deployed and reviewed:
- namespaces and capacity settings;
- event hubs, partition counts, and retention;
- private endpoints and private DNS integration;
- Entra ID role assignments;
- Capture destinations where required;
- diagnostic settings, alert rules, and dashboards.
The same definitions were promoted through development, test, pre-production, and production. Production resource creation was complete before any production client changed endpoints.
Application pipeline
Producer and consumer builds did not contain environment-specific broker addresses or secrets. Red and green connection profiles were external configuration. The same tested application artifact could therefore be pointed at Kafka or Event Hubs without rebuilding it.
The pipeline verified required configuration before deployment: bootstrap address, authentication mechanism, token handler, client timeouts, maximum request size, compression, group ID, and offset-reset policy. A missing value failed deployment instead of becoming a runtime discovery.
Replication and validation pipeline
MirrorMaker configuration was versioned by topic allow-list and deployed separately. Automated checks compared replication progress and data samples. Validation results were stored with the migration wave so the change approval referenced evidence rather than a verbal confirmation.
Promotion gates
A wave could move forward only when:
- infrastructure and access tests passed;
- compatibility tests passed for every client in the flow;
- historical replication or archive work reached its boundary;
- performance and recovery tests met the agreed thresholds;
- dashboards and alerts were live;
- the cutover and rollback runbooks had named operators;
- a production-like rehearsal had completed.
Testing the complete system
A successful produce and consume test was only the first layer.
Contract tests
The producer emitted representative events for every schema version still in use. Consumers verified payload, key, headers, timestamp, and serialization behaviour. Events near the maximum observed size were included. Unsupported compression or an incompatible serializer had to fail here, not during the production wave.
Partition and ordering tests
Known sequences were published for a set of business keys from every producer implementation. The test verified that each client used a compatible partitioning strategy and that the consumer observed monotonically increasing sequence values per key.
This mattered when the estate contained different Kafka client implementations. Two clients can accept the same key but use different default partitioners.
Consumer recovery tests
Consumers were restarted during load, scaled out and in, and deliberately paused long enough to accumulate backlog. Tests verified group rebalances, committed-position recovery, duplicate handling, and catch-up while new traffic continued.
auto.offset.reset was set explicitly. A new Event Hubs consumer group has no previous position, so choosing earliest or latest is a migration decision, not a harmless client default.
Performance and soak tests
The load profile used production-shaped event sizes and key distributions. It covered normal load, sustained peak, short bursts, full consumer fan-out, and backlog recovery. The test measured producer acknowledgement latency, request errors, retries, throttle time, consumer lag, processing latency, rebalance frequency, and partition skew.
A soak test then ran long enough to expose token renewal, idle connections, memory growth, connection recycling, and periodic traffic patterns that a short benchmark would miss.
Failure tests
The team interrupted consumers, restarted producer instances, denied network access in a controlled environment, tested expired credentials, and exercised downstream timeouts. The purpose was to verify application recovery and alerts, not to demonstrate that the service never fails.
The absence of a broker-managed dead-letter queue also had to be handled explicitly. Consumers that could encounter poison events required their own failure store, quarantine topic, or Service Bus hand-off, depending on the processing contract.
Migration rehearsal
The entire red/green sequence was rehearsed with production-like topology: establish the replication boundary, stop red processing, enable green, validate, then reverse the change. The rollback rehearsal was mandatory because a rollback that has never moved real events is only a document.
Monitoring during migration and after cutover
Monitoring was built before traffic moved. It had four views.
1. Event Hubs service view
Azure Monitor tracked incoming and outgoing events and bytes, successful requests, user and server errors, quota or throttling indicators, connections, and namespace capacity. Diagnostic settings routed the required logs to the approved destination.
2. Kafka client view
Producer metrics covered request latency, record error rate, retries, batching, local queue size, and throttle time. Consumer metrics covered records received, processing rate, lag, commit failures, poll interval, and rebalances.
Client telemetry remained essential. A managed service removes broker monitoring; it does not reveal every problem inside a producer or consumer.
3. Migration view
During coexistence, a dedicated dashboard showed MirrorMaker state, source and target event counts, replication lag, the latest replicated source timestamp, shadow-consumer results, duplicate IDs, sequence gaps, and reconciliation failures.
4. Business view
The final proof came from the applications: workstations appeared online, media jobs advanced through valid states, search indexes updated, license activity reconciled, notifications fired once, and data-platform ingestion met its freshness target.
Infrastructure metrics answered whether events moved. Business metrics answered whether the system still worked.
Step-by-step production migration
The complete sequence was:
- Define the migration boundary. Identify which Kafka workloads were in scope and which native-Kafka dependencies were excluded.
- Baseline the red system. Record throughput, latency, lag, errors, traffic peaks, retention, message sizes, and application SLOs.
- Map the dependency graph. Connect every producer, topic, consumer group, downstream write, and responsible owner.
- Complete the compatibility assessment. Test clients, authentication, compression, payloads, partitions, commits, and Kafka-specific features.
- Approve the service split. Use Event Hubs for streams, Event Grid for derived discrete notifications, and retain native Kafka or another messaging service where its semantics were required.
- Design the green platform. Define namespaces, event hubs, partitions, capacity, retention, identities, private networking, archive, and recovery.
- Deploy through pipelines. Build the complete production target, monitoring, and access controls without moving traffic.
- Migrate required history. Replicate active replay data with MirrorMaker and export long-term history to ADLS.
- Run shadow consumers. Compare data and processing decisions without production side effects.
- Complete load, soak, recovery, and security tests. Do not approve the wave from smoke tests alone.
- Rehearse cutover and rollback. Execute both directions against the production-like environment.
- Cut over one flow. Apply the ordering gate where required, change producers, start green consumers, and validate downstream outcomes.
- Observe before the next wave. Keep the blast radius limited; do not move unrelated flows simply because the first one connected successfully.
- Stabilize the green system. Operate through representative peak and business cycles while red remains recoverable.
- Close rollback and decommission red. Confirm no clients or required history remain, revoke access, archive evidence, and then retire Kafka infrastructure.
Rollback was part of the data design
Before green consumers created production side effects, rollback was straightforward: stop the green test, return clients to red, and investigate.
After producers had written directly to Event Hubs and green consumers had updated downstream systems, rollback required data decisions. The runbook specified:
- whether producers returned immediately to Kafka or were paused;
- how Event Hubs records accepted before rollback were drained or copied;
- where red consumers resumed;
- how duplicate event IDs were detected;
- how downstream state was reconciled;
- which system was authoritative at each point in the sequence.
This is why a configuration switch alone is not a rollback strategy. Once an event changes business state, the recovery boundary extends beyond the streaming platform.
What changed after replacement
The new system did not eliminate platform engineering. It changed its subject.
The team no longer managed broker operating systems, disks, replica placement, controller maintenance, or Kafka upgrades. It still managed event contracts, client standards, partitions, namespace capacity, identities, private networking, retention, archives, dashboards, recovery tests, and cost.
The result should be judged with measured evidence rather than a claim that a managed service is automatically simpler. Useful measures include:
- hours spent on infrastructure maintenance before and after;
- streaming incidents by cause;
- producer and consumer SLO compliance;
- time required to onboard a new event flow;
- capacity headroom and throttling frequency;
- recovery time from consumer and regional failure scenarios;
- total service, networking, monitoring, and storage cost.
If those results do not improve, the operating-model argument needs to be examined again.
Closing view
The Kafka platform in this system was functional. The migration was justified by the responsibility required to keep it functional as the application estate grew.
Event Hubs was selected because the main workloads depended on Kafka’s producer and consumer protocol, partitioning, consumer groups, and retained streams, but did not all require a native Kafka broker or its wider ecosystem. Confluent Cloud remained the correct comparison for workloads that did. Event Grid entered the architecture only where discrete business notifications were needed; it did not replace the stream.
The migration succeeded or failed at the boundaries: the partition key, the consumer start position, the last red event, the first green event, the historical-data decision, and the downstream side effect.
Those boundaries had to be designed, tested, monitored, and rehearsed before production traffic moved.
Changing bootstrap.servers was one task. Establishing and validating everything around that change was the migration.
Technical references
- What is Azure Event Hubs for Apache Kafka?
- Migrate to Azure Event Hubs for Apache Kafka
- Apache Kafka client configurations for Azure Event Hubs
- Replicate Kafka data to Event Hubs with MirrorMaker 2
- Choose between Event Grid, Event Hubs, and Service Bus
- Monitor Azure Event Hubs
- Azure Event Hubs quotas and limits
- Apache Kafka and Apache Flink on Confluent Cloud: Azure Native Integrations


