Table of Contents

WAL tuning for durable backends

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at wal-tuning.md, and llms.txt lists every page.

This document explains how the WAL partition grain's concurrency knobs interact with a durable backend's throughput envelope. The default values are tuned for a single Azure Tables Standard storage account on a 4-vCPU silo; lifting them above the documented envelope without a matching fan-out change collapses per-flush latency under provider-side throttling.

If you're sizing for a different backend or a different silo SKU, the Benchmarks doc and the benchmark/azure-throughput harness show the measurement protocol used to derive these numbers.

The knobs

The WAL partition grain's pipeline depth against IWalStorageProvider.AppendEncodedBatchAsync is bounded by two independent caps - WalMaxPendingBatches and WalPartitions - and shaped by a third knob, WalAppendCoalescingInFlightThreshold, that decides how full each flush gets:

Knob Default What it bounds
LatticeOptions.WalMaxPendingBatches 16 Maximum number of in-flight + just-started batches per shard.
LatticeOptions.WalAppendCoalescingInFlightThreshold 4 In-flight depth at or above which an arriving batch's final entry stops kicking its own flush, so small fanned-out slices accumulate into the next flush window instead of each paying a round trip. 0 disables.
LatticeOptions.WalPartitions 8 Number of WAL partition grains per tree the producer fans out across (every shard of the tree shares them).

The combined ceiling on simultaneous provider calls for one tree is therefore WalPartitions * WalMaxPendingBatches - at the defaults, 8 * 16 = 128 concurrent provider appends. Each tree's WAL partitions carry their own caps, so concurrently written trees add up.

Why 16 is the default

The previous default of 8 was picked when the foreground commit path was bounded by per-batch latency rather than throughput. Re-measuring on Standard_D4as_v5 in the same region as the Tables account (the canonical operator profile for the benchmark/azure-throughput harness, June 2026) showed the cap=8 regime spent most of its wall time waiting on the admission gate rather than on the provider RTT:

Instrument (p99, last full reporter window) cap=8 cap=16 Direction
wal.writer.partition.pending_appends 7 15 doubled (cap took effect)
wal.writer.append.admission_wait ms 2,000-2,555 1,196-1,389 -40 to -53%
wal.shard.dispatch.duration ms 1,233 702-1,075 -13 to -43%
leaf.commit.duration (step=wal) ms 2,281 1,296-1,479 -35 to -43%
wal.append.provider.duration ms 65-120 85-96 unchanged (Tables RTT floor)
provider.commit.duration (phase2) ms 50-57 51-56 unchanged

The mechanical story is that doubling the cap lifted the admission gate (admission_wait halved), the leaf observed it (leaf.commit step=wal halved), and per-flush provider duration was unchanged - so the saved wall time was pure throughput. The campaign recorded a +57% increase in steady-state silo throughput at the 4k:5 rung with no reliability regression (failed=0 across three runs, no [stall-watchdog] firings, no [wal-admission-timeout] lines).

CPU efficiency improved alongside throughput (the silo did ~10% more useful work per percent of CPU under cap=16), which rules out the "spinning more, not doing more" confound and points at the upstream admission gate as the binding constraint rather than the provider itself.

When lifting the cap stops helping

The cap=16 default sits at the inflection point beyond which the producer side stops benefiting and the storage account becomes the bottleneck. Two independent ceilings appear when the cap is lifted further in combination with the silo's other concurrency knobs:

  1. Producer-side offered rate. Once the silo is no longer admission-gated at the 4k:5 rung, the runner's offered-rate ceiling (~20,000 messages/sec from a single co-located producer process) becomes the binding constraint. The silo's measured steady-state throughput is a lower bound on its true capacity, not a measurement of it. Pushing past the producer floor requires either a larger silo SKU or a partitioned producer.

  2. Storage-account throughput. Azure Tables Standard has a sustained per-account throughput threshold around 2,500 transactions/sec. At WalPartitions = 8, a WalMaxPendingBatches of 16 produces a steady-state pressure of 8 * 16 = 128 concurrent flushes - at ~50 ms/flush this lands at ~2,560 ops/sec, right at the threshold. Doubling both knobs together (so WalPartitions = 8 and WalMaxPendingBatches = 16 becomes some combined configuration that drives a higher fan-out, for example raising the silo's own dispatch concurrency to 16 alongside cap=16 at 6,000 keys/s offered load, or raising WalPartitions to 16 against a single account) pushes the account above its budget and surfaces in one of two manifestations:

    Manifestation A: explicit throttling. Under sustained pressure the storage account responds with 429 TooManyRequests carrying a Retry-After header, the Azure.Data.Tables SDK back-off engages, per-flush provider duration spikes from ~50 ms to ~8 s, the writer's WalAppendDispatchTimeout (default 30 s), which bounds both the admission wait and each shard dispatch, exhausts, and [wal-admission-timeout] lines fire on the slowest partition.

    **Manifestation B: TableTransactionFailedException bursts (409

    • 504).** Under burst pressure the storage account does not always issue a clean 429; instead, individual transactions accumulate retry attempts within the SDK's default per-attempt deadline, succeed server-side on the first attempt, and the SDK's retry then races into a 409 Conflict ("The specified entity already exists") on the second attempt - or the per-attempt deadline (AzureTableWalStorageOptions.RetryNetworkTimeout, 10 s by default; null restores the SDK's ~100 s) exhausts and surfaces as "Operation could not be completed within the specified time" (a 504-style provider timeout). Both fault paths surface to the silo's foreground commit path as Azure.Data.Tables.TableTransactionFailedException carrying one of those two messages, not as RequestFailedException(StatusCode=429). This was the dominant manifestation observed on D8as_v5 at 6k:5 with WalPartitions=16 (256 concurrent transactions against one account); see benchmark/azure-throughput/throughput.md section 31 for the cohort.

    Both manifestations produce the same downstream symptoms: per-flush provider duration climbing into seconds, a drain wedge whose phenotype differs cleanly from the wedges covered by the existing WalFlushTimeout and WalAppendDispatchTimeout bounds, and non-zero failed=N entries surfacing to the foreground commit path. The recovery path is the same for both: partition the storage, not raise the per-grain timeouts.

The recovery is not to raise the per-grain timeouts further - the underlying constraint is the storage account, not the grain. The recovery is to partition the storage: spread the tree's WAL partitions across multiple accounts - register one provider per account under its own key and pin partitions to them (see Multi-account fan-out) - or move to a Premium account with a higher per-account throughput target. WalPartitions adds partitions to spread only for trees registered after the change; an existing tree keeps the count pinned at its first registration until a resize, a shadow-cutover restore or a schema remediation moves it onto a new physical copy, which pins the silo-wide WalPartitions value in force when that copy is registered.

Throttle the producer before the regime fires. The per-tree saturation back-pressure signal (IWalSaturationSignal, IWalSaturationObserver; see WAL Saturation Signal) exposes the writer-side admission gate as a typed, Healthy / Throttled / Saturated per-tree state, so callers driving offered load into the silo can slow down or pause before the failure tail above surfaces. The state is keyed by the id the tree's WAL is written under, so on an aliased tree a read by the logical id returns Healthy (see Resolution and scope). The signal is silo-scoped, and a healthy partition's SetAsync / SetManyAsync hot path pays only two concurrent-dictionary lookups per append to consult it, one each at the writer's admission gate and throttle pace - it is the leading-edge surface that pairs with the structural fix above (multi-account fan-out) and the shutdown drain below (bounded SIGTERM). The storage-account ceiling is still the binding constraint; the signal stops callers from offering past it.

Bounded shutdown when the writer is wedged

The above section describes the steady-state wedge: the silo is running and the storage account is saturating. A second wedge surface exists at silo shutdown: when the host receives SIGTERM while the writer's per-(tree, partition) admission semaphore has parked callers (callers awaiting a WalAppendDispatchTimeout-bounded admission slot whose downstream dispatches are themselves parked in the SDK retry loop), the silo cannot finish its bounded deactivation drain without releasing those parked callers.

The library handles this automatically: WalCommitLogWriter exposes a per-silo drain entry that the host's StopAsync lifecycle stage invokes via a registered IHostedService. The drain signals every parked AcquireAsync caller on the owning silo's writer; each parked caller surfaces a typed LatticeShuttingDownException (a sealed InvalidOperationException subclass whose message names WalDrainBudget for grep-attribution and whose InnerException preserves the legacy TimeoutException(WalDrainBudget) shape so existing diagnostic tooling continues to work), and a counter sample lands on orleans.lattice.wal.writer.append.drain.releases tagged with (tree, partition) so dashboards can graph "how many parked callers were released on this silo's shutdown". Post-drain AppendAsync / AppendManyAsync calls on the draining silo fail fast with the same LatticeShuttingDownException rather than blocking on a drained admission gate. See API Reference - Shutdown back-pressure for the caller contract.

The drain is per-silo, local-only. Each silo process in a multi-silo cluster has its own WalCommitLogWriter singleton with its own drain state; a drain on silo A does not touch silo B's admission semaphore and does not interrupt any in-flight IWalShardGrain activation that silo B is dispatching to. Rolling restarts settle cleanly because each silo drains its own writer independently when its turn arrives.

The recovery for the steady-state wedge is still to partition the storage (above); the drain seam closes the shutdown half of the same problem so a saturated silo terminates inside its bounded deactivation drain instead of needing SIGKILL. The two halves are complementary - the structural fix (multi-account fan-out) reduces the rate at which the steady-state wedge fires; the drain seam ensures that when it does fire, shutdown still settles inside the host's grace window.

Sizing rules of thumb

For a single Azure Tables Standard storage account:

Silo SKU Recommended WalMaxPendingBatches Notes
2 vCPU (Standard_D2as_v5 and smaller) 8 The silo is CPU-bound at the 4k:5 rung; the admission gate is not the binding constraint. Default 16 wastes admission depth on a CPU that cannot pull faster.
4 vCPU (Standard_D4as_v5) 16 (default) The sweet spot the default is tuned for. Silo CPU sits 55-75% of box at peak; admission depth is the binding constraint and 16 unblocks it without saturating the storage account.
8+ vCPU (Standard_D8as_v5 and larger) 16 (still default) The single-account ceiling, not the silo, is the binding constraint at this SKU. Measured envelope on D8as_v5 + single Azure Tables Standard account: ~22-24 ke/s at 6k:5 with WalMaxPendingBatches=16, WalPartitions=8 - only ~10-15% above the D4as_v5 baseline at the same defaults (~21 ke/s at 4k:5). Lifting WalMaxPendingBatches to 32 against the same account doubles the concurrent flushes to 256, a pressure the section 31 A/B measured as strictly worse when it reached the same 256 by raising WalPartitions to 16: per-partition throughput collapsed 64% and about 36.9k entries failed in a 45 s run. The recovery is spreading the tree's WAL partitions across accounts with named providers and pinned placement (see WAL Storage Providers), not a higher cap.

For a Premium Azure Tables account, or a fan-out of a tree's partitions across multiple Standard accounts through named providers and pinned placement, the available throughput ceiling is higher in proportion and the cap can be lifted accordingly. The mechanical rule does not change: keep the combined WalPartitions * WalMaxPendingBatches * average flush rate below the aggregate storage budget.

What to measure

Four instruments tell you which regime you are in:

  • wal.writer.append.admission_wait - time spent waiting at the per-shard admission gate. If p99 is on the order of seconds and wal.append.provider.duration is on the order of tens of milliseconds, you are admission-bound and lifting the cap helps. If admission_wait is sub-millisecond and provider duration is seconds, you are storage-bound and lifting the cap will not help.

  • wal.append.provider.duration - per-flush wall time against IWalStorageProvider.AppendEncodedBatchAsync. If this climbs from the ~50-100 ms Tables RTT floor into the seconds, the storage account is throttling. Lifting WalMaxPendingBatches will not help; the recovery is WalPartitions fan-out across accounts.

  • wal.writer.partition.pending_appends - the number of dispatches already in flight on the partition when each append was admitted, so it tops out at WalMaxPendingBatches - 1 by construction (the 7 and 15 in the table above). If p99 is consistently pinned there, one below the cap, the cap is the binding constraint and there may be headroom to lift it (subject to the storage envelope). If p99 sits well below the cap, the cap is not binding and lifting it is a no-op.

  • wal.saturation.state - the per-tree saturation regime, emitted as a 0 / 1 / 2 step function for Healthy / Throttled / Saturated. A tree at 1 is the leading edge - admission depth at or above the throttled ratio, or a sustained drain-lag or pin-latency condition; a tree at 2 means an acute cause is firing (dispatch timeouts, provider failures, or sustained flush latency, plus an at-cap partition only when WalSaturationAcuteOnly is false). Pair with wal.saturation.transitions to plot how often the regime flips. The same signal is exposed to callers via IWalSaturationSignal / IWalSaturationObserver so producers can throttle without scraping the meter.

See Metrics for the full set of WAL-side instruments and their tags.

See also

  • WAL - the foreground commit pipeline and how the per-shard grain enforces the bounds.
  • WAL Storage Providers - the IWalStorageProvider seam and the Azure Tables provider's two-phase batch protocol.
  • WAL Saturation Signal - the per-tree back-pressure surface that lets callers throttle their offered load before the saturation regime's failure tail surfaces.
  • Configuration - the full options reference, including the validator rules that reject non-positive values.
  • Benchmarks - the measurement harness and how to reproduce the numbers above on your own SKU.