Table of Contents

Shutdown back-pressure - LatticeShuttingDownException

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at shutdown-back-pressure-latticeshuttingdownexception.md, and llms.txt lists every page.

Part of Lattice Public API Reference.

Public typed exception thrown by any ILattice operator (and by the internal saga coordinator on its caller-facing throw path) when the operation cannot complete because the owning silo's write-ahead-log writer is draining as part of host shutdown. The typed slot lets callers that care about the shutdown regime explicitly distinguish it from genuine InvalidOperationException failures (which are not back-pressure). It derives from InvalidOperationException for backwards compatibility, but that inheritance is a hazard rather than a convenience: a broad catch (InvalidOperationException) absorbs the refusal and typically retries, which cannot succeed because the writer drain is a one-way transition. The type implements ILatticeDomainFault, so a broad handler declines it with catch (InvalidOperationException ex) when (ex is not ILatticeDomainFault); catching it by name remains correct.

Surfaces from four distinct shutdown failure shapes that share the same operational meaning ("this silo is going away; the operation was refused"):

  • Lifetime-aware pre-dispatch fast-fail. The public write entry points on ILattice - SetAsync (both overloads), SetIfVersionAsync, ApplyCrdtDeltaAsync (both overloads), ApplyCrdtDeltaManyAsync, GetOrSetAsync, SetManyAsync, SetManyWherePredicateAsync, SetManyAtomicAsync, SetManyAtomicWhereAsync, DeleteAsync, DeleteRangeAsync, and DeleteRangeWherePredicateAsync (the bulk-load entry points do not take this pre-check) - check IHostApplicationLifetime.ApplicationStopping first and throws this exception before it touches the activation directory or dispatches to the WAL writer once the host has begun shutting down. The same pre-check guards the internal shard-root write path and the operator-driven tombstone-compaction pass, so a single catch/is check covers every write path rather than only the atomic-write saga. The check is null-tolerant: non-hosted test activations (which do not register IHostApplicationLifetime) are unaffected.
  • Writer-side drain refusal. The WAL commit-log writer's drain path flips the per-instance drain flag; any new append (single or batched) dispatch after the flip throws this exception inline.
  • Admission-semaphore drain release. A caller already parked on the per-partition admission semaphore when the drain fired sees the release-by-drain surface as this exception (rather than the legacy TimeoutException(WalDrainBudget) shape).
  • Saga-coordinator short-circuit. When the atomic-write saga coordinator detects any of the above (or the Orleans OrleansMessageRejectionException "Unable to create local activation" / "invalid activation" shape that fires when a leaf grain has been deactivated as part of the same shutdown), the saga skips its retry loop and its whole compensation path - no abort is recorded or broadcast, and no further saga state is written - and wraps the cause in this exception so consumers can detect the regime via a single is check.

Caller contract: treat as back-pressure, not as a real failure. The entries the operation carried were never durably committed, but the silo refused to accept them because the host is going away rather than because the storage layer rejected them. A caller observing this exception should abandon the operation rather than retry it - every subsequent attempt against the same silo activation in this lifetime will fail with the same exception, because the writer drain is a one-way transition. Long-lived clients should either fail over to a peer silo (if the cluster is multi-node) or surface the back-pressure to upstream callers (drop the request, queue it to a side outbox, or rate-limit). Re-issuing the same operation after the host restarts is the normal recovery path; the previously failed entries are not durable, so the re-issue is a fresh attempt against a fresh silo activation.

var entries = new List<KeyValuePair<string, byte[]>>
{
    new("k1", new byte[] { 0x01 }),
    new("k2", new byte[] { 0x02 }),
};

try
{
    await lattice.SetManyAsync(entries);
}
catch (LatticeShuttingDownException)
{
    // The host is shutting down. The entries did not commit.
    // Queue them to a side buffer / fail over to a peer silo /
    // surface back-pressure to the caller; do NOT retry against
    // this lattice activation.
}

On the SetManyAtomicAsync path, a shutdown refusal records no terminal outcome. The saga does not complete: it persists nothing further, records no abort, and leaves its persisted state at the execute phase with its keepalive reminder still registered, and it throws this exception to the caller. A later reminder-driven resume on a live activation re-runs the saga from its persisted position and drives it to a terminal outcome, and only then does the orleans.lattice.atomic_write.completed counter record it - as committed, failed, or compensated. The counter's shutdown_refused outcome value is not produced by any code path, so a dashboard cannot use it to count shutdown refusals; count this exception at the caller instead.

Quieter shutdown logs

During the host deactivation window the Orleans runtime emits a Warning per in-flight grain call from two transport tear-down categories (Orleans.Messaging - "the silo is blocking application messages" - and Orleans.Runtime.Placement.PlacementService - "Unable to create local activation"). That is expected shutdown back-pressure, not a fault, but at steady-state verbosity it floods the silo log on every clean stop. AddLattice installs an in-library logger filter that demotes only those two categories' Warning records, and only while IHostApplicationLifetime.ApplicationStopping is signalled; on a healthy host the categories keep their Warning floor, and Error/Critical always survive even during shutdown so a genuine transport fault is never hidden. The filter is a Microsoft.Extensions.Logging LoggerFilterRule (not an IGrainCallFilter) because the records originate inside the Orleans runtime's own logging rather than in the grain-call pipeline. No host configuration is required; the demotion is automatic with AddLattice.