---
title: "Saturation back-pressure - LatticeSaturatedException - Lattice Public API Reference"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/api/saturation-back-pressure-latticesaturatedexception.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/api.md?plain=1#L1353-L1401"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Saturation back-pressure - `LatticeSaturatedException`

Part of [Lattice Public API Reference](../api.md).

Public typed exception thrown when an operation is refused because the tree's storage layer is back-pressured - most commonly by the WAL writer admission gate and the atomic-write saga coordinator, when the per-tree `IWalSaturationSignal` reported `WalSaturationState.Saturated` for longer than the caller's configured wait budget. Distinct from [`LatticeShuttingDownException`](shutdown-back-pressure-latticeshuttingdownexception.md): saturation is a *recoverable* steady-state regime (offered load is exceeding the storage layer's sustained drain rate), not a one-way silo shutdown. Seven seams raise it - the WAL writer admission gate, the atomic-write saga coordinator, the snapshot-cursor open path, the per-silo replay-permit admission gate, the `SetManyAsync` shard fan-out, the `SetManyAsync` whole-call envelope, and the transaction-registry row-size admission bound - and the `SaturationSource` property names which one, because their triggers and their safe recovery differ. The typed slot lets callers that care about the saturation regime explicitly distinguish it from genuine `InvalidOperationException` failures. It derives from `InvalidOperationException` for backwards compatibility, but that inheritance is a hazard rather than a convenience, and this exception is the worked example: a broad `catch (InvalidOperationException)` elsewhere in the routing layer once absorbed this back-pressure signal, discarded the whole routing cache and re-fanned-out every shard, amplifying load on a tree that was already saturated in proportion to the shard count. The type implements [`ILatticeDomainFault`](../api.md#domain-faults---ilatticedomainfault), so a broad handler declines it with `catch (InvalidOperationException ex) when (ex is not ILatticeDomainFault)`; catching it by name and backing off remains the correct handling.

Surfaces from seven distinct saturation failure shapes that share the same operational meaning ("this tree's storage layer is back-pressured; the operation was refused"):

- **Writer-side admission refusal.** The WAL commit-log writer consults the saturation verdict for the target partition before each per-partition admission acquire. On `Saturated`, the writer waits up to `LatticeOptions.WalAdmissionSaturationWaitBudget` (default 5 seconds) for that partition to leave `Saturated` (to return to `Healthy` when `WalSaturationAcuteOnly = false`) and, if the regime persists, throws this exception so callers observe the back-pressure in budget time instead of parking on the admission semaphore for up to `WalAppendDispatchTimeout` (default 30 seconds). Reported as `LatticeSaturationSource.WalAdmission`.

  The gate is additionally bounded **per top-level call** by `LatticeOptions.WalAdmissionSaturationCallBudget`. The per-append budget bounds one *wait*; it does not bound a *call*, because the write path holds three nested retry layers (the stale-routing retry, the shard-activation retry, and the per-leaf batch-dispatch retry) and each re-dispatch opened a fresh one, so a single call could accumulate a multiple of the configured budget while every individual wait stayed correctly bounded. The gate now takes the smaller of the call's remaining allowance and the per-append budget, and refuses without waiting once the allowance is spent. It defaults to `Timeout.InfiniteTimeSpan`, so the per-call bound is opt-in on the 9.x line ([#3390](https://github.com/NSTA1/Orleans.Lattice/issues/3390)); left at the default, only the per-append budget applies and the multiplication measured in [#3348](https://github.com/NSTA1/Orleans.Lattice/issues/3348) is unmitigated.
- **Saga-coordinator quiesce refusal.** The atomic-write saga's quiesce-on-saturated preflight runs before each batched dispatch and parks the saga on `WaitForHealthyAsync` up to the lesser of a fixed 30-second saga quiesce cap and the tree's `WalAppendDispatchTimeout`. On budget expiry with the tree still `Saturated`, the saga's fast-path refuses with this exception rather than re-dispatching the same RowKeys into a still-throttled storage account (the canonical 409-Conflict amplification regime); the saga's persisted state stays at `Execute` with the current `NextIndex` so the caller's next retry on the same `operationId` resumes from where the refusal stopped. The preflight looks the signal up under the tree id the saga addressed, so on an aliased tree it reads `Healthy` and does not engage (see [Resolution and scope](../wal-saturation-signal.md#resolution-and-scope)). Reported as `LatticeSaturationSource.AtomicWriteSaga`.
- **Snapshot-cursor open shed.** Opening a snapshot cursor is expensive, so a tree whose signal already reads `Saturated` sheds the open immediately - before routing resolution and the capture fan-out - rather than amplifying its own collapse. Unlike the two gates above there is no wait budget: the shed is immediate. Only `Saturated` sheds, so a `Throttled` tree stays browsable. Gated by the default-on `ShedSnapshotOpensWhenSaturated` option. Like the saga preflight, the shed looks the signal up under the tree id the open addressed, so it does not engage on an aliased tree (see [Resolution and scope](../wal-saturation-signal.md#resolution-and-scope)). Reported as `LatticeSaturationSource.SnapshotCursorOpen`.
- **Replay-permit admission refusal.** The per-silo WAL replay-permit gate refuses an activation when the permit queue is at or above its admitted bound *and* the queue is failing to drain: either the smoothed queue wait is at or above `LatticeOptions.WalReplayPermitMaxQueueWait`, or no queued activation has acquired a permit for that long. This seam does **not** consult `IWalSaturationSignal`: it is driven by replay-permit queue depth, so it can fire on a tree whose signal reads `Healthy`. Reported as `LatticeSaturationSource.ReplayPermitAdmission`. Background starvation drives - the WAL GC sweep's blocked-leaf drives and a leaf's own coverage-lag timer drives - are refused under the same source when the gate's share for them has no free permit: they never queue, they share at most half the replay permits still in circulation (the configured ceiling less any the memory-adaptive back-pressure is withholding; at least one), and while that share holds more than one slot a timer drive never takes its last free slot, which is kept for WAL GC sweep drives (a single-slot share is shared, and the timer yields it while a refused sweep drive is outstanding). Those refusals stay inside the silo (the sweep records an `admission_refused` outcome and the refused leaf backs off) and never reach an `ILattice` caller.
- **Batch fan-out budget expiry.** `SetManyAsync` splits a batch across the shards its keys route to and awaits every branch. A batch therefore pays its *slowest* branch, so as the shard count rises the call tracks the branch p99 rather than the branch median - the collapse measured in [#3348](https://github.com/NSTA1/Orleans.Lattice/issues/3348), where widening from four silos to eight *improved* the branch median 7.7x to 386 ms while the branch p99 degraded 11x to 94 s. The fan-out can now be bounded by `LatticeOptions.SetManyFanOutBudget`, which defaults to `Timeout.InfiniteTimeSpan` so the bound is opt-in on the 9.x line: once a finite budget expires with branches still outstanding, the call is refused with this exception instead of blocking indefinitely, which is what makes the fan-out shed load rather than queue it. Like the other seams this reports back-pressure to the caller only - it is **not** a rollback. `SetManyAsync` is not atomic across shards, so branches that already committed stay committed and outstanding branches run to completion; only the moment the caller is told changes. Reported as `LatticeSaturationSource.SetManyFanOut`.
- **Batch envelope budget expiry.** The fan-out budget above bounds one stage. `SetManyAsync` runs four in sequence - `gate` (registering the tree's compaction reminder and arming its autonomic loops), `route`, `bucket`, `fanout` - and a call can breach the caller's deadline by *summing* them while no single stage breaches. [#2685](https://github.com/NSTA1/Orleans.Lattice/issues/2685) measured exactly that: a `gate` of 4,108.96 ms (12,085x its 0.34 ms healthy baseline, an uncached arming path running on every write) plus a `fanout` of 26,709.17 ms, totalling 30,818 ms against a 30,000 ms Orleans response timeout. A fan-out budget never fires at 26.7 s, so the caller received an anonymous `TimeoutException` naming no stage. `LatticeOptions.SetManyEnvelopeBudget` bounds the running total instead, defaults to `Timeout.InfiniteTimeSpan` so the bound is opt-in, composes with the fan-out budget (the fan-out waits for the narrower of the two), and also bounds single-shard batches, which the fan-out budget deliberately skips. The refusal carries a per-stage breakdown, so the next occurrence names the stage that moved rather than the stage that is largest. Like the other seams it is **not** a rollback: already-committed branches stay committed and outstanding ones run to completion. Reported as `LatticeSaturationSource.SetManyEnvelope`.
- **Transaction-registry capacity refusal.** Each shard of a tree's transaction registry (`TxRegistryShardCount`, default 1) persists its whole state as one grain-state row, and each completed saga leaves a tombstone in its shard's row for `TxDecisionRetention`. Before a new atomic-write saga does any work, its shard estimates its row size and refuses the saga when that estimate is at or above `LatticeOptions.TxRegistryAdmissionBudgetBytes` (default 768 KiB) even after purging expired tombstones. This keeps the row under the storage provider's per-row limit. In-flight sagas, status reads and recovery are never refused. A refused saga discards the transaction id it drew, so a retry mints a fresh id and, when `TxRegistryShardCount` is above 1, may land on a shard with room; a single shard's capacity returns only as its tombstones age out, so on an unsharded tree back off for seconds, not milliseconds. A single-tree write is never retried by the library. A cross-tree write has already recorded its preparing phase and armed its one-minute keepalive reminder before it dispatches the sub-sagas, so the caller still receives the refusal, but the keepalive re-dispatches the prepare phase about once a minute until the sub-sagas are admitted, and a retry with the same `operationId` re-attaches to that saga. Reported as `LatticeSaturationSource.TxRegistryCapacity` ([#3475](https://github.com/NSTA1/Orleans.Lattice/issues/3475)).

Carries the originating tree id in the `TreeId` property so caller-side diagnostics can attribute the back-pressure to the specific tree without parsing the exception message. When the writer-side admission gate refused, the property names the tree whose admission semaphore observed the saturation; when the saga refused, the property names the saga's tree (or, when an underlying writer-side `LatticeSaturatedException` was wrapped through `SetManyAsync`'s leaf fan-out, the extracted tree id of the writer-side refusal).

Also carries `SaturationSource`, a `LatticeSaturationSource` value naming which of the seven seams refused, or `Unspecified` when the refusal was not attributed - an exception built through a constructor that takes no source, or deserialised from a host that predates source attribution - which a handler should treat as not automatically retryable. Because the seams differ in what is safe to do next, a handler that retries must branch on this property rather than on the exception type alone.

Caller contract: treat as back-pressure and **retry after backing off**. Typical recovery is 1-10 seconds (until the underlying storage account or per-partition WAL admission gate drains). Long-lived consumers should also reduce offered load on the affected tree until the per-tree signal returns to `Healthy`. Unlike `LatticeShuttingDownException`, retries against the same silo activation can succeed once the regime clears.

Part of that contract is now honoured inside the library. The shard-dispatch envelope retries a `ReplayPermitAdmission` refusal on a jittered 2 s / 4 s ladder, so a refused leaf activation is shed and re-attempted instead of failing the caller ([#3294](https://github.com/NSTA1/Orleans.Lattice/issues/3294)). It retries **only** that member: re-attempting a `WalAdmission` refusal at that seam would re-issue the batch across every shard of an already-saturated tree, the amplification [#3348](https://github.com/NSTA1/Orleans.Lattice/issues/3348) removed one layer above. That is why the discriminator is load-bearing rather than diagnostic decoration. Note the retry does not suppress `orleans.lattice.leaf.activation.failures{reason=refused_replay_admission}`: the activation still fails and is still counted, and the dispatch envelope recovers above it.

```csharp verify
var entries = new List<KeyValuePair<string, byte[]>>
{
    new("k1", new byte[] { 0x01 }),
    new("k2", new byte[] { 0x02 }),
};

try
{
    await lattice.SetManyAsync(entries);
}
catch (LatticeSaturatedException ex)
{
    // The per-tree saturation signal stayed Saturated past the
    // admission budget. The entries did not commit. Back off
    // (typical 1-10s), then retry. The TreeId property attributes
    // the back-pressure to the specific tree.
    Console.WriteLine($"tree {ex.TreeId} is saturated; backing off");
    await Task.Delay(TimeSpan.FromSeconds(2));
    // retry...
}
```

The writer-side admission refusal is also recorded on the `orleans.lattice.wal.writer.append.admission_saturation_refusals` counter (tagged `tree`, `partition`) so operators can dashboard the back-pressure rate separately from the dispatch-timeout counter (`orleans.lattice.wal.writer.append.admission_timeouts`) and the drain-release counter (`orleans.lattice.wal.writer.append.drain.releases`). Every refusal from any of the seven seams is also counted on `orleans.lattice.saturation.refusals`, tagged `tree`, `source` (the snake_case `SaturationSource`, for example `wal_admission` or `set_many_fan_out`) and `tenant`, together with a background starvation drive's refused replay permit, which the library absorbs rather than throws - so a refusal rate can be attributed to the seam that raised it. A `replay_permit_admission` refusal also carries an `arm` tag naming the predicate that refused (`wait_exceeded`, `no_progress`, or `gc_share`), and the `tree` value is the id held at the refusing seam, so on an aliased tree a `wal_admission` or `replay_permit_admission` refusal is counted under the physical copy's id. Each refusing seam records its own refusal, so the count is not one per caller-visible exception: a refusal raised by a write inside an atomic-write saga (a `wal_admission` refusal, for example) is counted under its own source and again under `atomic_write_saga`, and a saga refusal because its own quiesce budget elapsed is counted twice under `atomic_write_saga`. See [Metrics](../metrics.md).

Previous: [Shutdown back-pressure - LatticeShuttingDownException](shutdown-back-pressure-latticeshuttingdownexception.md). Next: [Admission back-pressure - LatticeQuotaExceededException](admission-back-pressure-latticequotaexceededexception.md). Contents: [Lattice Public API Reference](../api.md).
