---
title: "WAL saturation back-pressure - Lattice Public API Reference"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/api/wal-saturation-back-pressure.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/api.md?plain=1#L1087-L1238"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# WAL saturation back-pressure

Part of [Lattice Public API Reference](../api.md).

Lattice publishes a per-tree, three-state saturation signal so callers driving offered load into `ILattice` can throttle their own input *before* the saturation regime's failure tail surfaces to them as a `TimeoutException` from `SetAsync` / `SetManyAsync`. The surface has three shapes - polling, await, and push - all backed by a single silo-scoped sampler that ticks at `LatticeOptions.WalSaturationSampleInterval` (default 200 ms).

| Type | Shape | Description |
|------|-------|-------------|
| `WalSaturationState` | `enum` (`Healthy`, `Throttled`, `Saturated`) | The classification a tree is in. `Healthy` = admit without waiting; `Throttled` = a back-pressure input is up (admission depth near cap, or a sustained materialiser drain-lag or pin-latency run), callers should slow down; `Saturated` = an acute input is up (dispatch timeouts, provider failures, or sustained flush latency - plus, only when `WalSaturationAcuteOnly = false`, an admission semaphore at its cap with parked callers), callers should pause new appends. |
| `IWalSaturationSignal` | DI singleton (polling + await) | `GetCurrentState(treeId)` returns the cached per-tree state in one concurrent-dictionary lookup; `GetAggregateState()` returns the worst case across every observed tree; `WaitForHealthyAsync(treeId, ct)` returns a task that completes when the tree returns to `Healthy` (synchronous fast-path when already `Healthy`). Each lookup uses exactly the id it is passed, while the sampler records a verdict under the id the WAL is written under, which is the physical copy's id once the tree is [aliased](../tree-registry.md#tree-aliasing), so on an aliased tree a lookup by the logical id reads `Healthy` (see [Resolution and scope](../wal-saturation-signal.md#resolution-and-scope)). |
| `IWalSaturationObserver` | DI hook (push) | `OnStateChangedAsync(change, ct)` is invoked once per transition with a `WalSaturationStateChange` payload. Registered via `services.AddSingleton<IWalSaturationObserver, MyObserver>()`. Exceptions are caught and logged; the dispatcher continues to the next observer. |
| `WalSaturationStateChange` | `readonly record struct` | `TreeId`, `PreviousState`, `NewState`, `AttributedPartition?`, `AttributedShard?`, `ObservedAt`, `Cause`. `TreeId` is the id the WAL is written under, which is the physical copy's id on an aliased tree. The two attribution slots are best-effort and neither is filtered by `Cause`: `AttributedPartition` is the partition with the highest admission depth in the window (`null` when every partition was idle), and `AttributedShard` is a shard that recorded a dispatch-timeout, provider-failure, flush-latency or pin-latency trip in the window (`null` when none did). `Cause` is the `WalSaturationCause` the transition was attributed to. |
| `WalSaturationCause` | `enum` (`None`, `DispatchTimeouts`, `ProviderFailures`, `FlushLatency`, `AdmissionDepth`, `MaterialiserDrainLag`, `MaterialiserPinLatency`) | Which sampler input drove a transition, so an observer can route an alert at the subsystem actually under pressure. Best-effort and single-valued: when several inputs cross in the same window the first evaluated is reported. `None` on every transition back to `Healthy`, and on a `Throttled` reached only through the recovery-window hold. |

The signal is driven by six sampler inputs:

- **Admission-semaphore depth.** Per-partition `in_flight / WalMaxPendingBatches` ratio. `>= WalSaturationThrottledRatio` (default 0.75) raises a tree to `Throttled`. A semaphore at its cap with parked callers raises it to `Saturated` only when `WalSaturationAcuteOnly` is `false`; under the default (`true`) an at-cap partition reads `Throttled`, because the semaphore already applies that back-pressure itself.
- **Dispatch-timeout rate.** Trips of `WalAppendDispatchTimeout` observed within a single sample window. `>= WalSaturationDispatchTimeoutThreshold` (default 1) raises a tree to `Saturated` regardless of admission depth.
- **Provider-failure rate.** Exceptions surfaced from a downstream WAL shard append (single or batched) dispatch within a single sample window, other than a cancellation (any `OperationCanceledException`), a shutdown refusal (`LatticeShuttingDownException`), or the writer's own dispatch-deadline expiry, which counts on the dispatch-timeout input instead. `>= WalSaturationProviderFailureRateThreshold` (default 1; set to `0` to disable) raises a tree to `Saturated` regardless of admission depth and dispatch-timeout trips. Captures the regime where the provider's commit calls return quickly but terminally fail (e.g. the Azure Tables single-account 409-Conflict burst), so the silo surfaces the saturation regime instead of silently leaking entries.
- **Flush latency.** Opt-in through `WalSaturationFlushLatencyThreshold` (default `null`, disabled): provider flushes at or above the threshold in `WalSaturationFlushLatencySampleWindows` consecutive windows raise a tree to `Saturated`.
- **Materialiser drain lag.** A tree whose slowest eligible WAL consumer cursor (leaf materialisers and tree-wide consumers alike) stays more than `WalSaturationMaterialiserLagThreshold` (default 30 s) behind the WAL head for `WalSaturationMaterialiserLagSampleWindows` consecutive windows is held at `Throttled` - never `Saturated`.
- **Materialiser pin latency.** Opt-in through `WalSaturationMaterialiserPinLatencyThreshold` (default `null`, disabled): durable leaf-materialiser pin writes that are slow or fault in `WalSaturationMaterialiserPinLatencySampleWindows` consecutive windows hold a tree at `Throttled` - never `Saturated`.

The tree's state is the worst case across these inputs across every partition / shard, and `WalSaturationRecoveryWindow` holds a tree at `Throttled` for a short window after its last `Saturated` observation.

## Polling (canonical TCP-read-loop pattern)

```csharp verify
var signal = client.ServiceProvider.GetRequiredService<IWalSaturationSignal>();
while (!cancellationToken.IsCancellationRequested)
{
    if (signal.GetCurrentState("my-tree") == WalSaturationState.Saturated)
    {
        await Task.Delay(TimeSpan.FromMilliseconds(50), cancellationToken);
        continue;
    }
    // Continue reading from the producer transport (TCP, gRPC, queue, ...)
    // and dispatching SetAsync / SetManyAsync calls.
}
```

The polling getter costs one concurrent-dictionary lookup returning an `enum` - it is safe to call inside any per-message check on the producer hot path. The `SetAsync` / `SetManyAsync` hot path on `ILattice` is unchanged: the sampler runs on its own timer and never adds per-call work.

## Canonical per-call back-pressure helper (`ApplyBackPressureAsync`)

For TCP listeners, saga coordinators, and other consumers that drive offered load through a per-call decision loop, the library packages the three-state response pattern in a single call:

```csharp verify
var signal = client.ServiceProvider.GetRequiredService<IWalSaturationSignal>();
while (!cancellationToken.IsCancellationRequested)
{
    // Per-call back-pressure: no-op on Healthy, brief delay on
    // Throttled (1 ms default - tunable via the overload), full
    // park-until-Healthy on Saturated.
    await signal.ApplyBackPressureAsync("my-tree", cancellationToken);

    // Continue reading from the producer transport and dispatching
    // SetAsync / SetManyAsync calls.
}
```

The helper centralises the canonical response pattern so consumers do not roll their own (a recurring source of "the signal fires but back-pressure is too soft" when the consumer's Throttled response is `Task.Yield()` instead of an honest delay). The default Throttled delay (1 ms) slows a 10 k events/sec offered stream to ~1 k events/sec - enough to give the writer's admission gate time to drain before the regime escalates to Saturated, without producing a perceptible per-call latency penalty when the regime is transient. Pass an explicit `TimeSpan` to the four-argument overload to tune the strength of the Throttled response. Both overloads are extension methods on `WalSaturationSignalExtensions`, whose `DefaultThrottledDelay` (1 ms) is the delay the three-argument overload applies. The helper reads the verdict through `GetCurrentState`, so when it is passed the logical id of an aliased tree it reads `Healthy` and applies no back-pressure (see [Resolution and scope](../wal-saturation-signal.md#resolution-and-scope)).

## Await (single recovery wait)

```csharp verify
var signal = client.ServiceProvider.GetRequiredService<IWalSaturationSignal>();
// Block until the tree returns to Healthy, or until the caller's CT fires.
await signal.WaitForHealthyAsync("my-tree", cancellationToken);
```

`WaitForHealthyAsync` returns `Task.CompletedTask` synchronously when the tree is already `Healthy`, so the fast-path is allocation-free. When the tree is not `Healthy`, the awaiter completes on a sample tick that observes the tree at `Healthy` - at worst one `WalSaturationSampleInterval` beyond the underlying recovery, unless more than `WalSaturationRecoveryReleaseBatch` callers are parked on the tree, in which case the backlog is released oldest-first across consecutive ticks.

## Push (subscription model)

Implement the observer:

```csharp verify
public sealed class MyBackPressurePolicy : IWalSaturationObserver
{
    public ValueTask OnStateChangedAsync(WalSaturationStateChange change, CancellationToken cancellationToken)
    {
        // React to the transition. Keep this fast - long-running work
        // belongs on a channel drained by an IHostedService.
        return ValueTask.CompletedTask;
    }
}
```

Register it on the silo's DI container:

```csharp verify
public sealed class MyBackPressurePolicy : IWalSaturationObserver
{
    public ValueTask OnStateChangedAsync(WalSaturationStateChange change, CancellationToken cancellationToken)
    {
        // React to the transition. Keep this fast - long-running work
        // belongs on a channel drained by an IHostedService.
        return ValueTask.CompletedTask;
    }
}

public static void Register(ISiloBuilder siloBuilder)
{
    siloBuilder.Services.AddSingleton<IWalSaturationObserver, MyBackPressurePolicy>();
}
```

Multiple observers may coexist; they are invoked in registration order. Exceptions thrown by one observer are logged as a warning and suppressed - the others continue to run. Observers do not see the per-tick state; they see only the transitions, so a tree that holds `Throttled` for an hour produces one `Healthy -> Throttled` callback at the start and one `Throttled -> Healthy` callback at the end.

## Metrics

| Instrument | Type | Tags | Description |
|------------|------|------|-------------|
| `orleans.lattice.wal.saturation.state` | observable gauge (long) | `tree` | Current per-tree state as a step function, unit `{state}`. Values: `0` = Healthy, `1` = Throttled, `2` = Saturated. The regime is deliberately not also carried as a label: the value already encodes it, and a state label fragmented the per-tree series on every transition. A tree appears only once the sampler has observed a signal for it. |
| `orleans.lattice.wal.saturation.transitions` | counter (long) | `tree`, `state`, `previous_state`, `cause`, optional `partition`, optional `shard` | Incremented once per per-tree transition, unit `{transition}`. The state tag values are lowercased enum names (`healthy`, `throttled`, `saturated`); `cause` is the attributed `WalSaturationCause` in lowercase snake case (`none`, `dispatch_timeouts`, `provider_failures`, `flush_latency`, `admission_depth`, `materialiser_drain_lag`, `materialiser_pin_latency`). |

Both instruments also carry the derived `tenant` dimension that every Lattice instrument carries.

A flat-zero series on `transitions` is the healthy steady state. A rising rate of `state=throttled` transitions on a tree is the leading edge of the saturation regime; `state=saturated` is the regime itself. Pair with the `state` observable gauge for "what is the current regime" and with the `transitions` counter for "how often is the regime changing" - flapping between `Throttled` and `Saturated` is a different operational signal from a sustained `Saturated`.

## Options

The sampler reads every option below from the global (unnamed) options except the five budgets - `WalAdmissionSaturationWaitBudget`, `WalAdmissionSaturationCallBudget`, `WalThrottledAdmissionPace`, `SetManyFanOutBudget` and `SetManyEnvelopeBudget` - which resolve per tree, so a per-tree override of a classification input has no effect. Two options are read on both sides: `WalSaturationAcuteOnly` also decides, per tree, when a caller parked at the writer admission gate resumes, and a per-tree `WalSaturationFlushLatencyThreshold` counts slow flushes but feeds the classifier only while the global value is also set. [Configuration](../configuration.md) marks each option's scope.

| Option | Default | Description |
|--------|---------|-------------|
| `WalSaturationSampleInterval` | `200 ms` | Sampler cadence. The worst-case observer / await transition latency is one interval beyond the underlying signal crossing the threshold. Set to `Timeout.InfiniteTimeSpan` to disable the sampler entirely (signal pins to `Healthy`). |
| `WalSaturationThrottledRatio` | `0.75` | Per-partition admission-depth ratio at or above which the signal raises a tree to `Throttled`. Range `[0.0, 1.0]`. |
| `WalSaturationDispatchTimeoutThreshold` | `1` | Minimum dispatch-timeout trips per sample window that raise a tree to `Saturated` regardless of admission depth. |
| `WalSaturationProviderFailureRateThreshold` | `1` | Minimum provider-side commit failures per sample window that raise a tree to `Saturated` regardless of admission depth and dispatch-timeout trips. Captures the regime where provider commit calls return quickly but terminally fail (e.g. Azure Tables 409-Conflict bursts) so the silo surfaces the saturation regime instead of silently leaking entries. Set to `0` to disable the trigger entirely. |
| `WalSaturationFlushLatencyThreshold` | `null` (disabled) | Per-provider-flush wall-clock latency at or above which the WAL writer increments a per-(tree, shard) flush-latency trip counter that feeds the saturation classifier. Closes the small-batch blind spot the other three inputs cannot see (slow-but-successful flushes against a saturating storage account never fill the admission semaphore, trip the dispatch deadline, or tally a provider-failure). Must be positive when set; leaving it `null` is a zero-cost no-op. Pair with `WalSaturationFlushLatencySampleWindows`. |
| `WalSaturationFlushLatencySampleWindows` | `3` | Number of consecutive sample windows that must each observe a non-zero flush-latency-threshold trip-counter delta before the classifier upgrades the tree to `Saturated`. Noise floor for the flush-latency input. Minimum 1; has no effect when `WalSaturationFlushLatencyThreshold` is left at its default `null`. |
| `WalSaturationMaterialiserLagThreshold` | `30 s` | Materialiser drain-lag input: a tree whose slowest eligible WAL consumer cursor stays more than this far behind the WAL head for `WalSaturationMaterialiserLagSampleWindows` consecutive windows is held at `Throttled` (never `Saturated`), so a write burst that outruns the materialiser slows producers instead of faulting them. Set to `null` to disable the input; must be positive when set. |
| `WalSaturationMaterialiserLagSampleWindows` | `3` | Consecutive sample windows the drain lag must stay over `WalSaturationMaterialiserLagThreshold` before the tree is held at `Throttled`. Minimum 1; has no effect when the threshold is `null`. |
| `WalDrainLagConsumerFreshness` | `5 min` | Freshness window for the drain-lag input. A WAL consumer whose last cursor report is older than this, or a leaf materialiser whose cursor has not advanced within it and whose cursor itself predates it, is left out of the lag measurement, so an idle leaf cannot hold a live tree at `Throttled`. Such consumers still pin the WAL GC trim floor. `TimeSpan.Zero` disables the exclusion; negative values are rejected. |
| `WalDrainLagHolderLogInterval` | `10 min` | Minimum interval between repeat drain-lag holder warnings while a tree stays over `WalSaturationMaterialiserLagThreshold`. The first over-threshold window logs a warning naming up to three eligible consumers with the lowest cursors - in the log only, never as a metric tag - and it repeats at most once per interval while the tree stays over; recovery resets it, so the next crossing logs at once. `null` keeps edge-only logging; must be positive when set. Read from the global (unnamed) options only. |
| `WalSaturationMaterialiserPinLatencyThreshold` | `null` (disabled) | Materialiser pin-latency input: durable leaf-materialiser pin writes that take at least this long, or fault, in `WalSaturationMaterialiserPinLatencySampleWindows` consecutive windows hold the tree at `Throttled` (never `Saturated`). The only input measured against durable storage rather than in-memory progress. Must be positive when set. |
| `WalSaturationMaterialiserPinLatencySampleWindows` | `3` | Consecutive sample windows that must each record a slow or faulted pin write before the pin-latency input holds the tree at `Throttled`. Minimum 1; has no effect when the threshold is `null`. |
| `WalSaturationRecoveryWindow` | `1 s` | Window after the most-recently observed `Saturated` transition during which the classifier holds a tree at or above `Throttled` even if the current sampler tick's depth observation classifies it as `Healthy`. Defends against bursty per-partition WAL drain where the per-tick `max(depth_ratio)` oscillates `~1.0 <-> ~0.0` and the classifier would otherwise flap `Healthy <-> Saturated` at the sampler cadence with `Throttled` never observed as a stable state. Set to `TimeSpan.Zero` to disable the upgrade (per-tick depth observation drives the regime directly); set to `Timeout.InfiniteTimeSpan` to hold `Throttled` forever after the first `Saturated` observation. |
| `WalSaturationAcuteOnly` | `true` | When `true`, only acute causes (dispatch timeouts, provider failures, sustained flush latency) classify a partition `Saturated`; an admission semaphore merely at its cap reads `Throttled`, and a caller parked at the writer admission gate resumes once its partition leaves `Saturated` rather than waiting for `Healthy`. Set `false` to restore the historical at-cap `Saturated` classification. |
| `WalSaturationRecoveryReleaseBatch` | `16` | Maximum number of parked WAL-admission waiters a recovered partition admits per sampler tick. Bounds the burst a recovery hands back to the admission pipeline so the released population cannot immediately re-saturate the partition it was waiting on, which would leave the gate flapping with no net progress. Waiters are released oldest-first, and the release is level-triggered so a paced release never strands its residue: under the default `WalSaturationAcuteOnly = true` a gate waiter only needs its partition to leave `Saturated`, so every non-`Saturated` tick (`Throttled` included) drains a further batch; with `WalSaturationAcuteOnly = false` every `Healthy` tick does. Set to `0` to release every parked waiter at once. |
| `WalAdmissionSaturationWaitBudget` | `5 s` | Wall-clock budget the WAL writer's admission gate spends waiting, when the target partition reads `Saturated`, for that partition to leave `Saturated` (to return to `Healthy` when `WalSaturationAcuteOnly = false`) before refusing the dispatch with [`LatticeSaturatedException`](saturation-back-pressure-latticesaturatedexception.md). Sized shorter than `WalAppendDispatchTimeout` (so the saturation refusal wins over the dispatch timeout) and longer than one `WalSaturationSampleInterval` (so a transient classifier flap does not surface as a refusal). Set to `TimeSpan.Zero` to disable the gate entirely (the historical pre-admission-gate behaviour). Set to `Timeout.InfiniteTimeSpan` to wait forever on recovery. |
| `WalAdmissionSaturationCallBudget` | `Timeout.InfiniteTimeSpan` | Wall-clock budget **one top-level call** may spend waiting at the WAL admission saturation gate, summed across every append and every retry layer. `WalAdmissionSaturationWaitBudget` bounds one *wait*, not one *call*: the write path holds three nested retry layers and each re-dispatch opened a fresh budget, so a call could accumulate a multiple of it while every individual wait stayed correctly bounded ([#3348](https://github.com/NSTA1/Orleans.Lattice/issues/3348) remedy 3). The call is identified by a start instant its first public entry point stamps into `RequestContext`; nested entry points inherit it rather than re-stamping. The gate takes the smaller of the remaining allowance and the per-append budget, and refuses immediately once the allowance is spent. Writes with no ambient call (convergence-only, background) keep the per-append bound unchanged. `TimeSpan.Zero` means "never wait within a call"; it **defaults to `Timeout.InfiniteTimeSpan`**, so the bound is opt-in and upgrading changes no behaviour. 15 s (3x the per-append default) is the recommended finite value; making that the default is deferred to the next major ([#3390](https://github.com/NSTA1/Orleans.Lattice/issues/3390)). |
| `WalThrottledAdmissionPace` | `25 ms` | Per-append pacing delay the WAL writer applies before admission while the target partition reads `Throttled`. It gives the `Throttled`-only inputs (such as drain lag) teeth on the local write path, where the `Saturated`-only admission gate never engages. A pure back-off: it never throws, and it is skipped on `Saturated` (the admission gate governs that case). `TimeSpan.Zero` disables it; negative values are rejected. |
| `SetManyFanOutBudget` | `Timeout.InfiniteTimeSpan` | Wall-clock budget `SetManyAsync` spends awaiting its shard fan-out before refusing the call with [`LatticeSaturatedException`](saturation-back-pressure-latticesaturatedexception.md) (`SaturationSource` = `SetManyFanOut`). Bounds the *slowest branch*, which is what a batch actually pays: without it the call tracks the branch p99 and degrades as the shard count rises ([#3348](https://github.com/NSTA1/Orleans.Lattice/issues/3348)). Refusal sheds the caller and rolls nothing back - already-committed branches stay committed. Unlike `WalAdmissionSaturationWaitBudget`, `TimeSpan.Zero` is **rejected** rather than treated as a disable sentinel (it would refuse every batch); it **defaults to `Timeout.InfiniteTimeSpan`** (unbounded), so the bound is opt-in and upgrading an existing deployment changes no behaviour. 30 s is the recommended finite value; making that the default is deferred to the next major ([#3386](https://github.com/NSTA1/Orleans.Lattice/issues/3386)). |
| `SetManyEnvelopeBudget` | `Timeout.InfiniteTimeSpan` | Wall-clock budget for the **whole** `SetManyAsync` call - its `gate`, `route`, `bucket` and `fanout` stages together - before refusing it with [`LatticeSaturatedException`](saturation-back-pressure-latticesaturatedexception.md) (`SaturationSource` = `SetManyEnvelope`). Bounds the *sum*, which `SetManyFanOutBudget` structurally cannot: [#2685](https://github.com/NSTA1/Orleans.Lattice/issues/2685) measured a `gate` of 4,108.96 ms plus a `fanout` of 26,709.17 ms totalling 30,818 ms against a 30,000 ms response timeout, with **neither stage breaching alone**, so a fan-out budget sized for the fan-out never fired and the caller saw an anonymous Orleans timeout. The two compose - the fan-out waits for the narrower of them - and unlike the fan-out budget this one also bounds single-shard batches, which have an envelope but no slowest branch. The refusal carries a per-stage breakdown so the stage that *moved* is named rather than the stage that is merely largest. `TimeSpan.Zero` is **rejected**; defaults to `Timeout.InfiniteTimeSpan`, so the bound is opt-in. Size it below the governing `ResponseTimeout` (25 s against the 30 s Orleans default) so the refusal reaches a caller that is still listening. |

See [WAL Saturation Signal](../wal-saturation-signal.md) for the full design including the per-tree resolution contract, the multi-tree aggregate view, and the bench-side adoption pattern.

## Consumers

In addition to application callers, the `Orleans.Lattice.Replication` package consumes this signal automatically: a receiver's `WalSaturationReceiverFlowControlPolicy` (registered by `AddLatticeReplication` by default) reads `IWalSaturationSignal.GetCurrentState` under the push's logical tree name after each applied push and translates the regime into the backoff hints carried on the `ReplicationAck`, so a saturated receiver asks the sender to ship smaller batches and pause before its local admission gate faults the apply. On an aliased receiver tree that lookup reads `Healthy`, so the receiver sends no back-off hint (see [Resolution and scope](../wal-saturation-signal.md#resolution-and-scope)). See [Receiver flow control](../../lattice.replication/receiver-flow-control.md#built-in-wal-saturation-policy).

Previous: [Caller-credential propagation (LatticeCredentialContext)](caller-credential-propagation-latticecredentialcontext.md). Next: [WAL consumer cursors - IWalCursorRegistry](wal-consumer-cursors-iwalcursorregistry.md). Contents: [Lattice Public API Reference](../api.md).
