WAL tuning for durable backends
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at wal-tuning.md, and llms.txt lists every page.This document explains how the WAL partition grain's concurrency knobs interact with a durable backend's throughput envelope. The default values are tuned for a single Azure Tables Standard storage account on a 4-vCPU silo; lifting them above the documented envelope without a matching fan-out change collapses per-flush latency under provider-side throttling.
If you're sizing for a different backend or a different silo SKU, the
Benchmarks doc and the
benchmark/azure-throughput
harness show the measurement protocol used to derive these numbers.
The knobs
The WAL partition grain's pipeline depth against
IWalStorageProvider.AppendEncodedBatchAsync is bounded by two independent
caps - WalMaxPendingBatches and WalPartitions - and shaped by a third
knob, WalAppendCoalescingInFlightThreshold, that decides how full each
flush gets:
| Knob | Default | What it bounds |
|---|---|---|
LatticeOptions.WalMaxPendingBatches |
16 |
Maximum number of in-flight + just-started batches per shard. |
LatticeOptions.WalAppendCoalescingInFlightThreshold |
4 |
In-flight depth at or above which an arriving batch's final entry stops kicking its own flush, so small fanned-out slices accumulate into the next flush window instead of each paying a round trip. 0 disables. |
LatticeOptions.WalPartitions |
8 |
Number of WAL partition grains per tree the producer fans out across (every shard of the tree shares them). |
The combined ceiling on simultaneous provider calls for one tree is
therefore WalPartitions * WalMaxPendingBatches - at the defaults,
8 * 16 = 128 concurrent provider appends. Each tree's WAL
partitions carry their own caps, so concurrently written trees add up.
Why 16 is the default
The previous default of 8 was picked when the foreground commit path
was bounded by per-batch latency rather than throughput. Re-measuring
on Standard_D4as_v5 in the same region as the Tables account (the
canonical operator profile for the
benchmark/azure-throughput
harness, June 2026) showed the cap=8 regime spent most of its wall
time waiting on the admission gate rather than on the provider RTT:
| Instrument (p99, last full reporter window) | cap=8 | cap=16 | Direction |
|---|---|---|---|
wal.writer.partition.pending_appends |
7 | 15 | doubled (cap took effect) |
wal.writer.append.admission_wait ms |
2,000-2,555 | 1,196-1,389 | -40 to -53% |
wal.shard.dispatch.duration ms |
1,233 | 702-1,075 | -13 to -43% |
leaf.commit.duration (step=wal) ms |
2,281 | 1,296-1,479 | -35 to -43% |
wal.append.provider.duration ms |
65-120 | 85-96 | unchanged (Tables RTT floor) |
provider.commit.duration (phase2) ms |
50-57 | 51-56 | unchanged |
The mechanical story is that doubling the cap lifted the admission
gate (admission_wait halved), the leaf observed it (leaf.commit
step=wal halved), and per-flush provider duration was unchanged - so
the saved wall time was pure throughput. The campaign recorded a +57%
increase in steady-state silo throughput at the 4k:5 rung with no
reliability regression (failed=0 across three runs, no
[stall-watchdog] firings, no [wal-admission-timeout] lines).
CPU efficiency improved alongside throughput (the silo did ~10% more useful work per percent of CPU under cap=16), which rules out the "spinning more, not doing more" confound and points at the upstream admission gate as the binding constraint rather than the provider itself.
When lifting the cap stops helping
The cap=16 default sits at the inflection point beyond which the producer side stops benefiting and the storage account becomes the bottleneck. Two independent ceilings appear when the cap is lifted further in combination with the silo's other concurrency knobs:
Producer-side offered rate. Once the silo is no longer admission-gated at the 4k:5 rung, the runner's offered-rate ceiling (~20,000 messages/sec from a single co-located producer process) becomes the binding constraint. The silo's measured steady-state throughput is a lower bound on its true capacity, not a measurement of it. Pushing past the producer floor requires either a larger silo SKU or a partitioned producer.
Storage-account throughput. Azure Tables Standard has a sustained per-account throughput threshold around 2,500 transactions/sec. At
WalPartitions = 8, aWalMaxPendingBatchesof16produces a steady-state pressure of8 * 16 = 128concurrent flushes - at ~50 ms/flush this lands at ~2,560 ops/sec, right at the threshold. Doubling both knobs together (soWalPartitions = 8andWalMaxPendingBatches = 16becomes some combined configuration that drives a higher fan-out, for example raising the silo's own dispatch concurrency to 16 alongside cap=16 at 6,000 keys/s offered load, or raisingWalPartitionsto 16 against a single account) pushes the account above its budget and surfaces in one of two manifestations:Manifestation A: explicit throttling. Under sustained pressure the storage account responds with
429 TooManyRequestscarrying aRetry-Afterheader, the Azure.Data.Tables SDK back-off engages, per-flush provider duration spikes from ~50 ms to ~8 s, the writer'sWalAppendDispatchTimeout(default 30 s), which bounds both the admission wait and each shard dispatch, exhausts, and[wal-admission-timeout]lines fire on the slowest partition.**Manifestation B:
TableTransactionFailedExceptionbursts (409- 504).** Under burst pressure the storage account does not always
issue a clean 429; instead, individual transactions accumulate
retry attempts within the SDK's default per-attempt deadline,
succeed server-side on the first attempt, and the SDK's retry
then races into a
409 Conflict("The specified entity already exists") on the second attempt - or the per-attempt deadline (AzureTableWalStorageOptions.RetryNetworkTimeout, 10 s by default;nullrestores the SDK's ~100 s) exhausts and surfaces as"Operation could not be completed within the specified time"(a 504-style provider timeout). Both fault paths surface to the silo's foreground commit path asAzure.Data.Tables.TableTransactionFailedExceptioncarrying one of those two messages, not asRequestFailedException(StatusCode=429). This was the dominant manifestation observed on D8as_v5 at 6k:5 withWalPartitions=16(256 concurrent transactions against one account); seebenchmark/azure-throughput/throughput.mdsection 31 for the cohort.
Both manifestations produce the same downstream symptoms: per-flush provider duration climbing into seconds, a drain wedge whose phenotype differs cleanly from the wedges covered by the existing
WalFlushTimeoutandWalAppendDispatchTimeoutbounds, and non-zerofailed=Nentries surfacing to the foreground commit path. The recovery path is the same for both: partition the storage, not raise the per-grain timeouts.- 504).** Under burst pressure the storage account does not always
issue a clean 429; instead, individual transactions accumulate
retry attempts within the SDK's default per-attempt deadline,
succeed server-side on the first attempt, and the SDK's retry
then races into a
The recovery is not to raise the per-grain timeouts further - the
underlying constraint is the storage account, not the grain. The
recovery is to partition the storage: spread the tree's WAL
partitions across multiple accounts - register one provider per account
under its own key and pin partitions to them (see
Multi-account fan-out) -
or move to a Premium account with a higher per-account throughput target.
WalPartitions adds partitions to spread only for trees registered after
the change; an existing tree keeps the count pinned at its first
registration until a resize, a shadow-cutover restore or a schema
remediation moves it onto a new physical copy, which pins the silo-wide
WalPartitions value in force when that copy is registered.
Throttle the producer before the regime fires. The per-tree
saturation back-pressure signal (IWalSaturationSignal,
IWalSaturationObserver; see WAL Saturation Signal)
exposes the writer-side admission gate as a typed, Healthy /
Throttled / Saturated per-tree state, so callers driving offered
load into the silo can slow down or pause before the failure tail
above surfaces. The state is keyed by the id the tree's WAL is written
under, so on an aliased tree a read by
the logical id returns Healthy
(see Resolution and scope).
The signal is silo-scoped, and a healthy partition's
SetAsync / SetManyAsync hot path pays only two concurrent-dictionary
lookups per append to consult it, one each at the writer's admission
gate and throttle pace - it is the leading-edge surface
that pairs with the structural fix above (multi-account fan-out) and
the shutdown drain below (bounded SIGTERM). The storage-account
ceiling is still the binding constraint; the signal stops callers
from offering past it.
Bounded shutdown when the writer is wedged
The above section describes the steady-state wedge: the silo is
running and the storage account is saturating. A second wedge surface
exists at silo shutdown: when the host receives SIGTERM while
the writer's per-(tree, partition) admission semaphore has parked
callers (callers awaiting a WalAppendDispatchTimeout-bounded
admission slot whose downstream dispatches are themselves parked in
the SDK retry loop), the silo cannot finish its bounded deactivation
drain without releasing those parked callers.
The library handles this automatically: WalCommitLogWriter exposes
a per-silo drain entry that the host's StopAsync lifecycle stage
invokes via a registered IHostedService. The drain signals every
parked AcquireAsync caller on the owning silo's writer; each
parked caller surfaces a typed LatticeShuttingDownException (a
sealed InvalidOperationException subclass whose message names
WalDrainBudget for grep-attribution and whose InnerException
preserves the legacy TimeoutException(WalDrainBudget) shape so
existing diagnostic tooling continues to work), and a counter
sample lands on orleans.lattice.wal.writer.append.drain.releases
tagged with (tree, partition) so dashboards can graph "how many
parked callers were released on this silo's shutdown". Post-drain
AppendAsync / AppendManyAsync calls on the draining silo fail
fast with the same LatticeShuttingDownException rather than
blocking on a drained admission gate. See
API Reference - Shutdown back-pressure
for the caller contract.
The drain is per-silo, local-only. Each silo process in a
multi-silo cluster has its own WalCommitLogWriter singleton with
its own drain state; a drain on silo A does not touch silo B's
admission semaphore and does not interrupt any in-flight
IWalShardGrain activation that silo B is dispatching to. Rolling
restarts settle cleanly because each silo drains its own writer
independently when its turn arrives.
The recovery for the steady-state wedge is still to partition the
storage (above); the drain seam closes the shutdown half of the
same problem so a saturated silo terminates inside its bounded
deactivation drain instead of needing SIGKILL. The two halves
are complementary - the structural fix (multi-account fan-out)
reduces the rate at which the steady-state wedge fires; the drain
seam ensures that when it does fire, shutdown still settles
inside the host's grace window.
Sizing rules of thumb
For a single Azure Tables Standard storage account:
| Silo SKU | Recommended WalMaxPendingBatches |
Notes |
|---|---|---|
| 2 vCPU (Standard_D2as_v5 and smaller) | 8 |
The silo is CPU-bound at the 4k:5 rung; the admission gate is not the binding constraint. Default 16 wastes admission depth on a CPU that cannot pull faster. |
| 4 vCPU (Standard_D4as_v5) | 16 (default) |
The sweet spot the default is tuned for. Silo CPU sits 55-75% of box at peak; admission depth is the binding constraint and 16 unblocks it without saturating the storage account. |
| 8+ vCPU (Standard_D8as_v5 and larger) | 16 (still default) |
The single-account ceiling, not the silo, is the binding constraint at this SKU. Measured envelope on D8as_v5 + single Azure Tables Standard account: ~22-24 ke/s at 6k:5 with WalMaxPendingBatches=16, WalPartitions=8 - only ~10-15% above the D4as_v5 baseline at the same defaults (~21 ke/s at 4k:5). Lifting WalMaxPendingBatches to 32 against the same account doubles the concurrent flushes to 256, a pressure the section 31 A/B measured as strictly worse when it reached the same 256 by raising WalPartitions to 16: per-partition throughput collapsed 64% and about 36.9k entries failed in a 45 s run. The recovery is spreading the tree's WAL partitions across accounts with named providers and pinned placement (see WAL Storage Providers), not a higher cap. |
For a Premium Azure Tables account, or a fan-out of a tree's partitions
across multiple Standard accounts through named providers and pinned
placement, the available throughput ceiling is higher in proportion and
the cap can be lifted accordingly. The mechanical rule does not change: keep the
combined WalPartitions * WalMaxPendingBatches * average flush rate
below the aggregate storage budget.
What to measure
Four instruments tell you which regime you are in:
wal.writer.append.admission_wait- time spent waiting at the per-shard admission gate. If p99 is on the order of seconds andwal.append.provider.durationis on the order of tens of milliseconds, you are admission-bound and lifting the cap helps. If admission_wait is sub-millisecond and provider duration is seconds, you are storage-bound and lifting the cap will not help.wal.append.provider.duration- per-flush wall time againstIWalStorageProvider.AppendEncodedBatchAsync. If this climbs from the ~50-100 ms Tables RTT floor into the seconds, the storage account is throttling. LiftingWalMaxPendingBatcheswill not help; the recovery isWalPartitionsfan-out across accounts.wal.writer.partition.pending_appends- the number of dispatches already in flight on the partition when each append was admitted, so it tops out atWalMaxPendingBatches - 1by construction (the 7 and 15 in the table above). If p99 is consistently pinned there, one below the cap, the cap is the binding constraint and there may be headroom to lift it (subject to the storage envelope). If p99 sits well below the cap, the cap is not binding and lifting it is a no-op.wal.saturation.state- the per-tree saturation regime, emitted as a0/1/2step function forHealthy/Throttled/Saturated. A tree at1is the leading edge - admission depth at or above the throttled ratio, or a sustained drain-lag or pin-latency condition; a tree at2means an acute cause is firing (dispatch timeouts, provider failures, or sustained flush latency, plus an at-cap partition only whenWalSaturationAcuteOnlyisfalse). Pair withwal.saturation.transitionsto plot how often the regime flips. The same signal is exposed to callers viaIWalSaturationSignal/IWalSaturationObserverso producers can throttle without scraping the meter.
See Metrics for the full set of WAL-side instruments and their tags.
See also
- WAL - the foreground commit pipeline and how the per-shard grain enforces the bounds.
- WAL Storage Providers - the
IWalStorageProviderseam and the Azure Tables provider's two-phase batch protocol. - WAL Saturation Signal - the per-tree back-pressure surface that lets callers throttle their offered load before the saturation regime's failure tail surfaces.
- Configuration - the full options reference, including the validator rules that reject non-positive values.
- Benchmarks - the measurement harness and how to reproduce the numbers above on your own SKU.