← All writing
articleFeb 11, 202513 min read

Parallel Programming in .NET

Choosing and operating bounded parallelism in .NET across CPU work, asynchronous I/O, channels, dependency limits, cancellation, and production measurement.

.NETC#Performance
Parallel Programming in .NET cover illustration

The first parallel version of a batch processor often looks excellent on a developer machine. Ten thousand items that took eight minutes now take ninety seconds. In production, the same change exhausts the database connection pool, increases API throttling, allocates thousands of pending tasks, and makes unrelated requests wait.

The loop became faster at submitting work. The system became slower at completing it.

Parallel programming is therefore a capacity decision before it is a syntax choice. I start by naming what each item consumes and which resource is allowed to saturate.

Separate CPU work from waiting work

Asynchrony and parallelism solve different constraints.

  • Asynchronous I/O lets a thread serve other work while the operation waits for a socket, database, file, or broker.
  • Parallel CPU work executes computations concurrently, usually across multiple cores.
  • Concurrent I/O allows several independent waits to be in flight, which can increase throughput until a dependency limit is reached.

await does not create another CPU core. Task.Run does not make network I/O faster. Task.WhenAll does not impose a safe concurrency limit; it only completes when the supplied tasks complete.

The workload can also be mixed. An image pipeline may download, decode, transform, and upload. Treating the whole item as one concurrency unit makes the CPU and network stages share a limit even though they have different capacities. A staged pipeline can bound each resource independently.

Pick the primitive from the work

Work shape Typical primitive Primary limit
Finite independent CPU collection Parallel.For / Parallel.ForEach cores and memory bandwidth
Finite asynchronous collection Parallel.ForEachAsync downstream concurrency
Small known set of independent async calls Task.WhenAll known fan-out size
Continuous producer/consumer flow bounded Channel<T> queue capacity and worker count
CPU work from UI/client code Task.Run at the boundary thread-pool/CPU capacity
CPU work inside ASP.NET Core request usually direct or offloaded to a worker service request latency and host CPU

The table is a starting point. A call being asynchronous does not prove it is safe to fan out. An EF Core DbContext is not thread-safe, many clients have rate limits, and a single request opening hundreds of connections can starve every other request.

Bound a finite asynchronous batch

Parallel.ForEachAsync expresses bounded item processing directly:

var options = new ParallelOptions
{
    MaxDegreeOfParallelism = 12,
    CancellationToken = cancellationToken
};

await Parallel.ForEachAsync(
    documentIds,
    options,
    async (documentId, token) =>
    {
        await processor.ProcessAsync(documentId, token);
    });

Twelve is not a magic number and Environment.ProcessorCount is not automatically right. For I/O work, derive an initial bound from the constrained dependency:

  • available database connections after reserving headroom for interactive traffic;
  • external API quota and observed latency;
  • per-operation memory;
  • storage throughput;
  • broker in-flight limits;
  • tenant fairness requirements.

Then load-test the complete path. If the dependency permits 120 requests per second and p95 latency is 250 ms, Little’s Law suggests roughly 30 in-flight operations just to sustain that rate under those conditions. Retries, bursts, and tail latency need headroom; they do not justify unbounded fan-out.

When result order matters, index inputs and write results to their assigned slots. ConcurrentBag<T> protects mutation but does not preserve source order.

var indexed = documents.Select((item, index) => (item, index));
var results = new Result[documents.Count];

await Parallel.ForEachAsync(indexed, options, async (entry, token) =>
{
    results[entry.index] = await TransformAsync(entry.item, token);
});

Task.WhenAll is appropriate when the fan-out is already bounded

For three independent service calls, Task.WhenAll is clear:

Task<Customer> customerTask = customers.GetAsync(customerId, token);
Task<Balance> balanceTask = balances.GetAsync(customerId, token);
Task<Offers> offersTask = offers.GetAsync(customerId, token);

await Task.WhenAll(customerTask, balanceTask, offersTask);

For 50,000 items, this pattern creates all operations immediately:

await Task.WhenAll(items.Select(item => ProcessAsync(item, token)));

Even if a connection pool limits actual database calls, the application has already allocated tasks, captured state, and queued demand. The pool becomes an accidental concurrency controller with a large invisible backlog.

Use a bounded iterator or producer/consumer pipeline when the input can be large or continuous.

Bound the backlog with a channel

A bounded Channel<T> controls both active workers and queued items:

var channel = Channel.CreateBounded<Job>(new BoundedChannelOptions(500)
{
    FullMode = BoundedChannelFullMode.Wait,
    SingleWriter = false,
    SingleReader = false
});

Task[] workers = Enumerable.Range(0, 12)
    .Select(_ => ConsumeAsync(channel.Reader, cancellationToken))
    .ToArray();

await foreach (var job in source.WithCancellation(cancellationToken))
    await channel.Writer.WriteAsync(job, cancellationToken);

channel.Writer.Complete();
await Task.WhenAll(workers);

With FullMode.Wait, a fast producer slows when the 500-item buffer is full. That is real backpressure. Choosing DropOldest, DropNewest, or DropWrite changes business semantics and needs metrics plus a documented loss policy.

For a hosted background service, finish the lifecycle: complete the writer, drain or abandon according to shutdown policy, propagate the host cancellation token, and make redelivered work idempotent. A bounded in-memory channel is not durable. If work must survive process loss, the durable queue or database remains the source of truth.

CPU parallelism follows a different budget

For sufficiently large, independent CPU operations, Parallel.ForEach can partition work across thread-pool threads:

var options = new ParallelOptions
{
    MaxDegreeOfParallelism = Math.Max(1, Environment.ProcessorCount - 1),
    CancellationToken = cancellationToken
};

Parallel.ForEach(images, options, image =>
{
    image.ApplyTransform();
});

Leaving one core is only an experiment, not a universal formula. Container CPU limits, simultaneous requests, garbage collection, native libraries with their own threading, and memory bandwidth all change the useful degree.

Do not parallelize tiny iterations. Partitioning, scheduling, synchronization, and cache misses can cost more than the work. Do not nest parallel loops unless the combined maximum is intentionally bounded. A library method may already parallelize internally, so application-level fan-out can oversubscribe the machine without obvious nested code.

Inside ASP.NET Core, Task.Run for CPU-bound work still uses the process thread pool and CPU. It may move work off the current execution segment, but it does not protect request latency under load. Long or bursty CPU jobs often belong behind a bounded durable queue in separately scaled workers.

Shared state changes the result and the scaling

The easiest parallel operation is pure: each input produces an independent output. Shared mutable state requires synchronization, which can erase the speedup or introduce nondeterminism.

Prefer local accumulation followed by one merge. For counters, Interlocked may be enough. For maps and sets, concurrent collections protect their own operations but not an arbitrary multi-step invariant.

This is not atomic:

if (!dictionary.ContainsKey(key))
    dictionary[key] = BuildValue();

Another worker can enter between the calls. Use an atomic collection operation such as GetOrAdd, while remembering that a value factory may run more than once under contention. If construction has side effects, separate value creation from effect commitment.

Also verify whether dependencies are thread-safe. One DbContext cannot be used by parallel operations. Creating one context per item can then exhaust the connection pool. The dependency lifetime and concurrency budget have to agree.

Failure semantics must be chosen

When several items run concurrently, more than one can fail before cancellation propagates.

Decide whether the batch is:

  • fail-fast, canceling pending work after the first failure;
  • best-effort, collecting every item result;
  • transactional, which usually requires a different design because external side effects cannot be rolled back as one unit;
  • resumable, storing checkpoints and retrying failed item identities.

Do not catch and discard exceptions inside each worker merely to keep the loop alive. Return a typed outcome or write a durable failure record. Preserve the item ID, stage, attempt, and safe error classification.

Cancellation is cooperative. Stop accepting new work, pass the token to every cancellable dependency, and let in-flight operations reach a safe boundary. If an item may be retried after cancellation or process death, its side effects need an idempotency key or transactional claim.

Remember that awaiting Task.WhenAll surfaces a failure, while multiple tasks may have faulted. Inspect and record all relevant failures without logging the same exception at every layer.

Retries multiply concurrency

A worker count of 20 with three immediate retries can create much more than 20 attempts against a struggling dependency over time. Retries also extend occupancy, so the same fixed worker set processes fewer new items while the backlog grows.

Classify transient failures, back off with jitter, honor server retry guidance, and enforce one total attempt/time budget. A circuit breaker can stop repeated work when the dependency is clearly unavailable, but it does not replace queue bounds or admission control.

Keep retry policy close to the dependency owner. Stacking an HTTP-client retry, repository retry, item retry, and batch retry creates multiplicative attempts that no dashboard explains easily.

Measure the constrained resource

A microbenchmark of the loop body answers only whether local computation got faster. Production evidence includes:

  • completed items per second and end-to-end batch duration;
  • active workers, queued items, and queue age;
  • CPU utilization, allocation rate, GC pause, and working set;
  • thread-pool thread count and queue length;
  • database pool usage and acquisition time;
  • downstream latency, throttles, timeouts, and retry attempts;
  • item failure/cancellation count;
  • fairness and latency of unrelated requests.

Increase concurrency in steps under a representative workload. Throughput normally rises, then flattens, while latency and errors start increasing. Operate below that knee with headroom for bursts and partial dependency degradation.

Failure rehearsal

Test the system while one resource slows:

  1. Add 500 ms to the downstream call and confirm active work stays bounded.
  2. Reduce the database pool and verify the batch does not starve interactive traffic.
  3. Return 429 with Retry-After and observe retry rate and backlog age.
  4. Cancel halfway through input production and confirm workers terminate and claims recover.
  5. Fault several items simultaneously and retain every required failure identity.
  6. Kill the process after a side effect and before checkpointing; replay must remain safe.
  7. Run inside the real container CPU and memory limits.
  8. Compare ordered output with the sequential baseline.
  9. Load multiple tenants and verify one large batch cannot own every worker.
  10. Roll deployment shutdown while the channel is full.

The decision

Parallelism spends capacity to reduce elapsed time. It succeeds only when items are independent, coordination cost is smaller than useful work, the constrained dependency has headroom, and admission remains bounded when that dependency slows.

Starting more tasks was one code change. Proving that throughput rises without moving the backlog into memory, the thread pool, or somebody else’s service was the performance work.

Technical references

Keep reading
Browse everything