---
title: "WAL tuning for durable backends"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/wal-tuning.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/wal-tuning.md"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# WAL tuning for durable backends

Part of the [Orleans.Lattice documentation](architecture.md).

This document explains how the WAL partition grain's concurrency knobs
interact with a durable backend's throughput envelope. The default
values are tuned for a single Azure Tables Standard storage account on
a 4-vCPU silo; lifting them above the documented envelope without a
matching fan-out change collapses per-flush latency under provider-side
throttling.

If you're sizing for a different backend or a different silo SKU, the
[Benchmarks](benchmarks.md) doc and the
[`benchmark/azure-throughput`](https://github.com/NSTA1/Orleans.Lattice/tree/release/9.9/benchmark/azure-throughput/)
harness show the measurement protocol used to derive these numbers.

## The knobs

The WAL partition grain's pipeline depth against
`IWalStorageProvider.AppendEncodedBatchAsync` is bounded by two independent
caps - `WalMaxPendingBatches` and `WalPartitions` - and shaped by a third
knob, `WalAppendCoalescingInFlightThreshold`, that decides how full each
flush gets:

| Knob | Default | What it bounds |
|---|---|---|
| `LatticeOptions.WalMaxPendingBatches` | `16` | Maximum number of in-flight + just-started batches **per shard**. |
| `LatticeOptions.WalAppendCoalescingInFlightThreshold` | `4` | In-flight depth at or above which an arriving batch's final entry stops kicking its own flush, so small fanned-out slices accumulate into the next flush window instead of each paying a round trip. `0` disables. |
| `LatticeOptions.WalPartitions` | `8` | Number of WAL partition grains per tree the producer fans out across (every shard of the tree shares them). |

The combined ceiling on simultaneous provider calls for one tree is
therefore `WalPartitions * WalMaxPendingBatches` - at the defaults,
`8 * 16 = 128` concurrent provider appends. Each tree's WAL
partitions carry their own caps, so concurrently written trees add up.

## Why 16 is the default

The previous default of 8 was picked when the foreground commit path
was bounded by per-batch latency rather than throughput. Re-measuring
on Standard_D4as_v5 in the same region as the Tables account (the
canonical operator profile for the
[`benchmark/azure-throughput`](https://github.com/NSTA1/Orleans.Lattice/tree/release/9.9/benchmark/azure-throughput/)
harness, June 2026) showed the cap=8 regime spent most of its wall
time waiting on the admission gate rather than on the provider RTT:

| Instrument (p99, last full reporter window) | cap=8 | cap=16 | Direction |
|---|---|---|---|
| `wal.writer.partition.pending_appends` | 7 | 15 | doubled (cap took effect) |
| `wal.writer.append.admission_wait` ms | 2,000-2,555 | 1,196-1,389 | -40 to -53% |
| `wal.shard.dispatch.duration` ms | 1,233 | 702-1,075 | -13 to -43% |
| `leaf.commit.duration` (step=wal) ms | 2,281 | 1,296-1,479 | -35 to -43% |
| `wal.append.provider.duration` ms | 65-120 | 85-96 | unchanged (Tables RTT floor) |
| `provider.commit.duration` (phase2) ms | 50-57 | 51-56 | unchanged |

The mechanical story is that doubling the cap lifted the admission
gate (admission_wait halved), the leaf observed it (leaf.commit
step=wal halved), and per-flush provider duration was unchanged - so
the saved wall time was pure throughput. The campaign recorded a +57%
increase in steady-state silo throughput at the 4k:5 rung with no
reliability regression (`failed=0` across three runs, no
`[stall-watchdog]` firings, no `[wal-admission-timeout]` lines).

CPU efficiency improved alongside throughput (the silo did ~10% more
useful work per percent of CPU under cap=16), which rules out the
"spinning more, not doing more" confound and points at the upstream
admission gate as the binding constraint rather than the provider
itself.

## When lifting the cap stops helping

The cap=16 default sits at the **inflection point** beyond which the
producer side stops benefiting and the storage account becomes the
bottleneck. Two independent ceilings appear when the cap is lifted
further in combination with the silo's other concurrency knobs:

1. **Producer-side offered rate.** Once the silo is no longer
   admission-gated at the 4k:5 rung, the runner's offered-rate ceiling
   (~20,000 messages/sec from a single co-located producer process)
   becomes the binding constraint. The silo's measured steady-state
   throughput is a **lower bound** on its true capacity, not a
   measurement of it. Pushing past the producer floor requires either
   a larger silo SKU or a partitioned producer.

2. **Storage-account throughput.** Azure Tables Standard has a
   sustained per-account throughput threshold around 2,500
   transactions/sec. At `WalPartitions = 8`, a `WalMaxPendingBatches`
   of `16` produces a steady-state pressure of `8 * 16 = 128`
   concurrent flushes - at ~50 ms/flush this lands at ~2,560 ops/sec,
   right at the threshold. Doubling **both** knobs together (so
   `WalPartitions = 8` and `WalMaxPendingBatches = 16` becomes some
   combined configuration that drives a higher fan-out, for example
   raising the silo's own dispatch concurrency to 16 alongside cap=16
   at 6,000 keys/s offered load, or raising `WalPartitions` to 16
   against a single account) pushes the account above its budget
   and surfaces in one of two manifestations:

   **Manifestation A: explicit throttling.** Under sustained pressure
   the storage account responds with `429 TooManyRequests` carrying a
   `Retry-After` header, the Azure.Data.Tables SDK back-off engages,
   per-flush provider duration spikes from ~50 ms to ~8 s, the
   writer's `WalAppendDispatchTimeout` (default 30 s), which bounds both
   the admission wait and each shard dispatch, exhausts,
   and `[wal-admission-timeout]` lines fire on the slowest
   partition.

   **Manifestation B: `TableTransactionFailedException` bursts (409
   + 504).** Under burst pressure the storage account does not always
   issue a clean 429; instead, individual transactions accumulate
   retry attempts within the SDK's default per-attempt deadline,
   succeed server-side on the first attempt, and the SDK's retry
   then races into a `409 Conflict` ("The specified entity already
   exists") on the second attempt - or the per-attempt deadline
   (`AzureTableWalStorageOptions.RetryNetworkTimeout`, 10 s by default;
   `null` restores the SDK's ~100 s) exhausts and surfaces as
   `"Operation could not be completed within the specified time"`
   (a 504-style provider timeout). Both fault paths surface to the
   silo's foreground commit path as
   `Azure.Data.Tables.TableTransactionFailedException` carrying one
   of those two messages, **not** as
   `RequestFailedException(StatusCode=429)`. This was the dominant
   manifestation observed on D8as_v5 at 6k:5 with
   `WalPartitions=16` (256 concurrent transactions against one
   account); see `benchmark/azure-throughput/throughput.md` section
   31 for the cohort.

   Both manifestations produce the same downstream symptoms: per-flush
   provider duration climbing into seconds, a drain wedge whose
   phenotype differs cleanly from the wedges covered by the existing
   `WalFlushTimeout` and `WalAppendDispatchTimeout` bounds, and
   non-zero `failed=N` entries surfacing to the foreground commit
   path. The recovery path is the same for both: partition the
   storage, not raise the per-grain timeouts.

The recovery is **not** to raise the per-grain timeouts further - the
underlying constraint is the storage account, not the grain. The
recovery is to **partition the storage**: spread the tree's WAL
partitions across multiple accounts - register one provider per account
under its own key and pin partitions to them (see
[Multi-account fan-out](wal-storage-providers.md#multi-account-fan-out-named-providers-and-pinned-placement)) -
or move to a Premium account with a higher per-account throughput target.
`WalPartitions` adds partitions to spread only for trees registered after
the change; an existing tree keeps the count pinned at its first
registration until a resize, a shadow-cutover restore or a schema
remediation moves it onto a new physical copy, which pins the silo-wide
`WalPartitions` value in force when that copy is registered.

**Throttle the producer before the regime fires.** The per-tree
saturation back-pressure signal (`IWalSaturationSignal`,
`IWalSaturationObserver`; see [WAL Saturation Signal](wal-saturation-signal.md))
exposes the writer-side admission gate as a typed, `Healthy` /
`Throttled` / `Saturated` per-tree state, so callers driving offered
load into the silo can slow down or pause *before* the failure tail
above surfaces. The state is keyed by the id the tree's WAL is written
under, so on an [aliased](tree-registry.md#tree-aliasing) tree a read by
the logical id returns `Healthy`
(see [Resolution and scope](wal-saturation-signal.md#resolution-and-scope)).
The signal is silo-scoped, and a healthy partition's
`SetAsync` / `SetManyAsync` hot path pays only two concurrent-dictionary
lookups per append to consult it, one each at the writer's admission
gate and throttle pace - it is the leading-edge surface
that pairs with the structural fix above (multi-account fan-out) and
the shutdown drain below (bounded SIGTERM). The storage-account
ceiling is still the binding constraint; the signal stops callers
from offering past it.

## Bounded shutdown when the writer is wedged

The above section describes the **steady-state** wedge: the silo is
running and the storage account is saturating. A second wedge surface
exists at **silo shutdown**: when the host receives `SIGTERM` while
the writer's per-(tree, partition) admission semaphore has parked
callers (callers awaiting a `WalAppendDispatchTimeout`-bounded
admission slot whose downstream dispatches are themselves parked in
the SDK retry loop), the silo cannot finish its bounded deactivation
drain without releasing those parked callers.

The library handles this automatically: `WalCommitLogWriter` exposes
a per-silo drain entry that the host's `StopAsync` lifecycle stage
invokes via a registered `IHostedService`. The drain signals every
parked `AcquireAsync` caller on the owning silo's writer; each
parked caller surfaces a typed `LatticeShuttingDownException` (a
sealed `InvalidOperationException` subclass whose message names
`WalDrainBudget` for grep-attribution and whose `InnerException`
preserves the legacy `TimeoutException(WalDrainBudget)` shape so
existing diagnostic tooling continues to work), and a counter
sample lands on `orleans.lattice.wal.writer.append.drain.releases`
tagged with `(tree, partition)` so dashboards can graph "how many
parked callers were released on this silo's shutdown". Post-drain
`AppendAsync` / `AppendManyAsync` calls on the draining silo fail
fast with the same `LatticeShuttingDownException` rather than
blocking on a drained admission gate. See
[API Reference - Shutdown back-pressure](api/shutdown-back-pressure-latticeshuttingdownexception.md)
for the caller contract.

The drain is **per-silo, local-only**. Each silo process in a
multi-silo cluster has its own `WalCommitLogWriter` singleton with
its own drain state; a drain on silo A does not touch silo B's
admission semaphore and does not interrupt any in-flight
`IWalShardGrain` activation that silo B is dispatching to. Rolling
restarts settle cleanly because each silo drains its own writer
independently when its turn arrives.

The recovery for the **steady-state** wedge is still to partition the
storage (above); the drain seam closes the **shutdown** half of the
same problem so a saturated silo terminates inside its bounded
deactivation drain instead of needing `SIGKILL`. The two halves
are complementary - the structural fix (multi-account fan-out)
reduces the rate at which the steady-state wedge fires; the drain
seam ensures that when it does fire, shutdown still settles
inside the host's grace window.

## Sizing rules of thumb

For a single Azure Tables Standard storage account:

| Silo SKU | Recommended `WalMaxPendingBatches` | Notes |
|---|---|---|
| 2 vCPU (Standard_D2as_v5 and smaller) | `8` | The silo is CPU-bound at the 4k:5 rung; the admission gate is not the binding constraint. Default 16 wastes admission depth on a CPU that cannot pull faster. |
| 4 vCPU (Standard_D4as_v5) | `16` (default) | The sweet spot the default is tuned for. Silo CPU sits 55-75% of box at peak; admission depth is the binding constraint and 16 unblocks it without saturating the storage account. |
| 8+ vCPU (Standard_D8as_v5 and larger) | `16` (still default) | The single-account ceiling, not the silo, is the binding constraint at this SKU. Measured envelope on D8as_v5 + single Azure Tables Standard account: **~22-24 ke/s** at 6k:5 with `WalMaxPendingBatches=16`, `WalPartitions=8` - only ~10-15% above the D4as_v5 baseline at the same defaults (~21 ke/s at 4k:5). Lifting `WalMaxPendingBatches` to 32 against the same account doubles the concurrent flushes to 256, a pressure the section 31 A/B measured as strictly worse when it reached the same 256 by raising `WalPartitions` to 16: per-partition throughput collapsed 64% and about 36.9k entries failed in a 45 s run. The recovery is spreading the tree's WAL partitions across accounts with named providers and pinned placement (see [WAL Storage Providers](wal-storage-providers.md#multi-account-fan-out-named-providers-and-pinned-placement)), not a higher cap. |

For a Premium Azure Tables account, or a fan-out of a tree's partitions
across multiple Standard accounts through named providers and pinned
placement, the available throughput ceiling is higher in proportion and
the cap can be lifted accordingly. The mechanical rule does not change: keep the
combined `WalPartitions * WalMaxPendingBatches * average flush rate`
below the aggregate storage budget.

## What to measure

Four instruments tell you which regime you are in:

- **`wal.writer.append.admission_wait`** - time spent waiting at the
  per-shard admission gate. If p99 is on the order of seconds and
  `wal.append.provider.duration` is on the order of tens of
  milliseconds, you are admission-bound and lifting the cap helps.
  If admission_wait is sub-millisecond and provider duration is
  seconds, you are storage-bound and lifting the cap will not help.

- **`wal.append.provider.duration`** - per-flush wall time against
  `IWalStorageProvider.AppendEncodedBatchAsync`. If this climbs from the
  ~50-100 ms Tables RTT floor into the seconds, the storage account
  is throttling. Lifting `WalMaxPendingBatches` will not help; the
  recovery is `WalPartitions` fan-out across accounts.

- **`wal.writer.partition.pending_appends`** - the number of dispatches
  already in flight on the partition when each append was admitted, so it
  tops out at `WalMaxPendingBatches - 1` by construction (the 7 and 15 in
  the table above). If p99 is consistently pinned there, one below the
  cap, the cap is the binding constraint
  and there may be headroom to lift it (subject to the storage
  envelope). If p99 sits well below the cap, the cap is not binding
  and lifting it is a no-op.

- **`wal.saturation.state`** - the per-tree saturation regime,
  emitted as a `0` / `1` / `2` step function for `Healthy` /
  `Throttled` / `Saturated`. A tree at `1` is the leading edge -
  admission depth at or above the throttled ratio, or a sustained
  drain-lag or pin-latency condition; a tree at `2` means an acute
  cause is firing (dispatch timeouts, provider failures, or sustained
  flush latency, plus an at-cap partition only when
  `WalSaturationAcuteOnly` is `false`). Pair with
  `wal.saturation.transitions` to plot how often the regime flips. The same signal is exposed to callers via
  `IWalSaturationSignal` / `IWalSaturationObserver` so producers
  can throttle without scraping the meter.

See [Metrics](metrics.md) for the full set of WAL-side instruments
and their tags.

## See also

- [WAL](wal.md) - the foreground commit pipeline and how the
  per-shard grain enforces the bounds.
- [WAL Storage Providers](wal-storage-providers.md) - the
  `IWalStorageProvider` seam and the Azure Tables provider's
  two-phase batch protocol.
- [WAL Saturation Signal](wal-saturation-signal.md) - the per-tree
  back-pressure surface that lets callers throttle their offered
  load before the saturation regime's failure tail surfaces.
- [Configuration](configuration.md) - the full options reference,
  including the validator rules that reject non-positive values.
- [Benchmarks](benchmarks.md) - the measurement harness and how to
  reproduce the numbers above on your own SKU.
