← All writing
articleApr 23, 202512 min read

Long Polling in .NET

Designing a cursor-based long-poll endpoint in ASP.NET Core with bounded waits, replay, replica coordination, overload control, and failure testing.

.NETC#Architecture
Long Polling in .NET cover illustration

The requirement sounds small: when an export finishes, update the browser within a few seconds. Ordinary polling every ten seconds feels late. Polling every second produces mostly empty requests. WebSockets would work, but the product needs one server-to-client notification path, not a bidirectional session protocol.

Long polling is useful in that gap. The client sends an ordinary HTTP request; the server returns immediately when data exists or ends the wait after a bounded interval; the client processes the response and starts the next request.

The transport is simple. The delivery contract is not. A production design has to answer what happens when the update is committed just before timeout, when the response is lost, when the next request lands on another replica, and when ten thousand clients reconnect together.

Define the contract before holding requests open

Use a monotonic cursor scoped to the stream the caller is authorized to read:

GET /projects/7f3a/updates?after=1842&limit=100

The response contains ordered records and the next cursor:

{
  "items": [
    {
      "position": 1843,
      "type": "export.completed",
      "resourceId": "01K6...",
      "occurredAt": "2026-09-28T08:41:12.381Z"
    }
  ],
  "next": 1843
}

The client advances only after it has processed the response. If the server produced position 1843 but the connection died before the client received it, the next request still asks for everything after 1842.

That makes delivery at least once from the client’s point of view. Client-side handling must tolerate the same position more than once. Exactly-once UI effects come from applying a position idempotently, not from assuming one HTTP response per event.

Do not use a timestamp alone as the cursor. Multiple changes can share a timestamp, clock precision varies, and an inclusive/exclusive comparison is easy to get wrong. A database sequence, ordered event ID, or compound (timestamp, unique_id) cursor gives a total order. Treat the cursor as opaque if its representation may change.

Query first, then wait, then query again

A signal is not the source of truth. It only tells a waiter that querying again may now return data.

request(after=1842)
    |
    +-- query durable update store
    |       |
    |       +-- rows found -> return them
    |
    +-- subscribe/wait for signal
    |
    +-- query durable store again
    |       |
    |       +-- rows found -> return them
    |
    +-- signal or timeout -> query and return

The second query closes a race: an update can commit after the first query but before the waiter subscribes. If the implementation only queries and then waits, that notification can be missed until some later update wakes the request.

The durable store owns replay. An in-memory channel or distributed pub/sub message owns wake-up latency. Mixing those roles is how a transient notification loss becomes a missing business update.

A bounded ASP.NET Core endpoint

The empty wait is a normal outcome, not an exception that should become a 500.

app.MapGet("/projects/{projectId:guid}/updates", async Task<IResult> (
    Guid projectId,
    long after,
    UpdateFeed feed,
    CurrentActor actor,
    CancellationToken requestAborted) =>
{
    if (!await actor.CanReadProjectAsync(projectId, requestAborted))
        return Results.Forbid();

    var existing = await feed.ReadAsync(
        projectId, after, limit: 100, requestAborted);

    if (existing.Count > 0)
        return Results.Ok(UpdateBatch.From(existing));

    using var waitBudget = new CancellationTokenSource(
        TimeSpan.FromSeconds(25));
    using var linked = CancellationTokenSource.CreateLinkedTokenSource(
        requestAborted,
        waitBudget.Token);

    try
    {
        await feed.WaitForChangeAsync(projectId, linked.Token);
    }
    catch (OperationCanceledException) when (
        waitBudget.IsCancellationRequested &&
        !requestAborted.IsCancellationRequested)
    {
        return Results.NoContent();
    }

    var updates = await feed.ReadAsync(
        projectId, after, limit: 100, requestAborted);

    return updates.Count == 0
        ? Results.NoContent()
        : Results.Ok(UpdateBatch.From(updates));
});

There are two cancellation reasons. The application wait budget produces a normal empty response. requestAborted means the caller disconnected or upstream canceled the request; no response needs to be manufactured, and downstream work should stop.

An alternative is ASP.NET Core request-timeout middleware with a named policy. It can set RequestAborted and produce a timeout response when application code does not handle the cancellation. Whichever mechanism owns the budget, use one owner per endpoint and test its response. Stacking an ad hoc token, timeout middleware, proxy timeout, and client timeout at the same duration makes failures indistinguishable.

Order the timeouts

I want the application to end a quiet wait while the HTTP path is still healthy:

application wait       25s
proxy idle timeout     35s or more
client request timeout 40s or more

Those are illustrative, not universal values. The important property is ordering plus margin for network and response processing.

If the proxy closes at 25 seconds while the application also waits 25 seconds, some clients see 204 No Content and others see a transport error. Both reconnect, but dashboards and retry behavior become noisy. Infrastructure limits also vary by environment, so verify them with the deployed route rather than only Kestrel locally.

On an empty response, reconnect immediately with a small randomized delay. On transport or 5xx failure, use capped exponential backoff with jitter. Honor Retry-After for overload responses. Without jitter, a deployment or network recovery can synchronize every client into a request wave.

Multiple replicas change the wake-up path

An in-process waiter registry works only while the update producer and waiting request share one process. With several replicas, the update may be written by instance B while the request waits on instance A.

The cross-replica design needs two pieces:

  • a durable ordered update log that every instance can query;
  • a broadcast or pub/sub wake-up signal that reaches the relevant instances.

The signal may be lost or duplicated. Correctness still comes from re-reading the durable log by cursor. The signal should contain enough routing information—such as tenant or project ID—to avoid waking every request for every change, but it should not contain secrets that bypass the authorized query.

Sticky sessions do not solve this. They may keep a browser on one replica, but producers, background workers, failover, and scaling still move work across instances.

Bound the waiting population

Long polling reduces empty responses but increases concurrent open requests. Each waiter consumes a socket, request state, cancellation registrations, and some memory. HTTP/2 multiplexing reduces client connection pressure; it does not make server-side waiters free.

Capacity starts with Little’s Law. If new long-poll requests arrive at 400 per second and remain open for an average of 20 seconds, the system holds roughly 8,000 concurrent requests before accounting for bursts or retries.

Protect the service with:

  • a global concurrent-waiter limit;
  • a per-tenant and per-user limit;
  • a bounded registration structure;
  • a maximum batch size and response size;
  • cancellation cleanup that removes every waiter;
  • load shedding with 429 or 503 and Retry-After before saturation;
  • a cursor-retention policy and explicit response when a cursor is too old.

A stale cursor cannot replay forever if the update log has finite retention. Return a stable error that tells the client to fetch a current snapshot and resume from a new cursor. Silently starting at the oldest retained event can create a partial state that looks complete.

Authorization continues for the life of the stream

The route must scope both wait registration and replay query by the authorized resource. Never register on a globally supplied channel name and assume the initial URL check protects subsequent reads.

If access can be revoked during a 25-second wait, decide whether the next query re-evaluates authorization. For most business systems that extra check is worth the bounded delay. At minimum, every new request must authenticate and authorize again; a long-poll loop is not a permanent session grant.

Avoid putting sensitive event payloads in the distributed signal. Use the signal to wake the endpoint, then load authorized data from the system of record.

Observe the delivery path

Request latency is supposed to be high here, so a generic p95 HTTP-duration alert is misleading. Separate useful wait from failure.

Track:

  • current waiters globally and by tenant;
  • waits completed by data, empty timeout, client cancellation, and error;
  • time from update commit to client-visible response;
  • replay batch size and oldest unread position;
  • cursor-too-old responses;
  • waiter registration and cleanup counts;
  • reconnect rate, overload responses, and proxy disconnects;
  • durable-log query latency and signal delivery lag.

The business signal is not “requests stayed open for 24 seconds.” It is “completed exports became visible within the target delay without losing positions.”

Failure rehearsal

Before choosing long polling as the modest option, rehearse the non-modest cases:

  1. Commit an update between the initial query and waiter registration.
  2. Produce an update, then drop the response before the client receives it.
  3. Kill the replica holding thousands of waiters.
  4. Drop the pub/sub wake-up signal and confirm the next empty-timeout cycle still replays the update.
  5. Revoke access while a request waits.
  6. Return more than one batch while a client is offline.
  7. Restart all clients together and verify jitter plus load shedding.
  8. Present a cursor older than retention and verify snapshot recovery.
  9. Cancel requests repeatedly and check that waiter counts return to baseline.
  10. Set the proxy timeout below the application timeout in staging and confirm monitoring distinguishes the failure.

When another transport is the better boundary

Use server-sent events when delivery is continuously server-to-client, intermediaries support streaming, and keeping one response open is operationally acceptable. Use WebSockets when the interaction is genuinely bidirectional or needs lower per-message overhead. Use ordinary polling when changes are rare and several seconds of delay is harmless.

Long polling is not a primitive WebSocket. It is a cursor-based read model whose request is allowed to wait for a wake-up. The cursor, durable replay, authorization, and capacity limits are the design. Holding an HTTP request open is only the mechanism.

Reducing empty responses was one task. Proving that a client can disconnect, change replicas, and recover every authorized update was the feature.

Technical references

Keep reading
Browse everything