---
title: "Instrument catalog: Shard-root SetManyAsync decomposition (sourced from ShardRootGrain) to WAL append pipeline (sourced from WalShardGrain and WalCommitLogWriter) - Metrics"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/metrics/instrument-catalog-6.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/metrics.md?plain=1#L530-L607"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Instrument catalog: Shard-root `SetManyAsync` decomposition (sourced from `ShardRootGrain`) to WAL append pipeline (sourced from `WalShardGrain` and `WalCommitLogWriter`)

Part of [Instrument catalog](instrument-catalog.md), in [Metrics](../metrics.md).

## Shard-root `SetManyAsync` decomposition (sourced from `ShardRootGrain`)

Inside the `LatticeGrain.SetManyAsync` `stage=fanout` span, every shard runs
its own `ShardRootGrain.SetManyAsync` slice. The instruments below split
that per-shard slice into the local-apply work, the online-resize
shadow-forward path, and the per-leaf RPC fan-out. Together with
`leaf.commit.duration` per step, this gives an end-to-end attribution from
the lattice grain boundary down to the leaf commit pipeline.

The conditional batched write (`ILattice.SetManyWherePredicateAsync`) runs
the same per-shard shape and records into the same three instruments, but
it has no `stage=fanout` span: its envelope is
`set_many_where_predicate.duration`, and the `operation` tag separates its
samples on the two instruments that carry that tag.

One caveat bounds that arithmetic on the failure path. The fan-out settles
on the first faulted branch and does not cancel its siblings, so on a
failed call the surviving shard slices run on past the close of the
`stage=fanout` span and record into the instruments below afterwards.
Their samples can therefore exceed the parent span they are nominally
nested inside. Treat the containment relationship as holding on the
success path only, and scope any attribution sum to calls that succeeded.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.shard_root.set_many.local_apply.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside a shard root's local apply of one batched-write slice, recorded whether or not the local apply throws. The span starts after the shard root's own preparation and reject checks and covers the per-entry leaf routing, the parallel per-leaf RPC fan-out, the split promotion that follows it and - while an adaptive split is in progress or slots have moved away - the per-key forward of affected writes to their new owner shard. Recorded by both batched-write paths and tagged `tree` and `operation` - `set_many` for a slice of an `ILattice.SetManyAsync` batch and `set_many_where_predicate` for a slice of an `ILattice.SetManyWherePredicateAsync` batch; the instrument carries no `shard` tag. Compare it against `set_many.duration` only when filtered to `operation="set_many"`, and against `set_many_where_predicate.duration` only when filtered to `operation="set_many_where_predicate"` - the unfiltered series mixes both paths. Includes per-leaf RPC scheduling, leaf turn-queue wait, leaf commit, WAL append, and WAL phase-2 commit. Excludes the lattice-grain's bucket build (covered by `set_many.stage`) and excludes the online-resize shadow-forward path (covered separately below). |
| `orleans.lattice.shard_root.set_many.shadow_forward.duration` | `Histogram<double>` | `ms` | Wall-clock ms a shard root's batched-write slice spent awaiting its trailing shadow-forward task after the local apply succeeded (nothing is recorded when the local apply throws). Recorded by both batched-write paths - the conditional path forwards its whole slice too, and the destination re-evaluates the predicate against its own copy - with no `operation` tag, so the series mixes both. The shadow forward mirrors the slice, in parallel with the local apply, to the same shard of the destination tree while an online snapshot - including the one an online resize runs - is draining, so this is the residual tail of that forward rather than its full duration; it is not the adaptive-split migration forward. Tagged `tree` only - the instrument carries no `shard` tag. Near zero whenever no shadow forward is active, because the forward then completes synchronously. |
| `orleans.lattice.shard_root.set_many.leaf_rpc.duration` | `Histogram<double>` | `ms` | Per-leaf wall-clock ms inside one batched-write RPC dispatched by the shard-root local-apply fan-out. Recorded for every attempt that returns and for every attempt whose fault is retried - a retried attempt's sample also includes the back-off it waits when the leaf refused the write mid-retirement - while the attempt whose fault ends the dispatch records nothing. Tagged `tree` and `operation` - `set_many` for the unconditional per-leaf batch and `set_many_where_predicate` for the conditional one; the instrument carries no `shard` tag. The gap between this and `shard_root.set_many.local_apply.duration`, read within one `operation`, is the max-of-N tail across the parallel per-leaf calls plus the local apply's own routing and split-promotion work. |
| `orleans.lattice.shard_root.forward.timeouts` | `Counter<long>` | `{timeout}` | Count of outbound shard-to-shard write forwards (the online-resize shadow forward and the adaptive-split migration forward) abandoned after exceeding `ShardForwardTimeout`. Tagged `tree`. A non-zero value indicates a forward parked against a sibling shard whose ownership was changing during a reshard swap; the forward was faulted as a `TimeoutException` so the foreground write pipeline could make progress and the operation be retried against refreshed routing. Expected to be zero in steady state. |
| `orleans.lattice.shard_root.activation_ready.timeouts` | `Counter<long>` | `{timeout}` | Count of `ShardRootGrain` activation-readiness seeds abandoned after exceeding `ActivationReadyTimeout`. Tagged `tree`. A non-zero value indicates a first-activation seed (registry registration or root-leaf init) parked - typically because a startup reshard or membership change left the target activation not-yet-visible - and was faulted as a `TimeoutException` so the held activation gate could release and the foreground write pipeline make progress, with the seed retried against refreshed routing. Expected to be zero in steady state. |
| `orleans.lattice.shard_root.scan_page.stalls` | `Counter<long>` | `{stall}` | Count of shard-root range-scan page fills abandoned after exceeding the effective `MaxScanPageStallDuration` ceiling (by default derived to sit below the silo's Orleans `ResponseTimeout`). Tagged `tree`, `shard`, and `phase` (`prologue`, `descent`, `leaf-walk`, `baseline-fold`). `MaxScanPageDuration` is a cooperative budget sampled between leaf reads, so it can only stop a walk somewhere it can resume from and cannot bound a prologue that parks or a single leaf read that never returns; this counter fires when the hard end-to-end ceiling releases the deliberately non-reentrant shard so its queue can drain, faulting the call as a `ScanPageStalledException` the caller retries from its last continuation token. The `phase` tag makes the stall self-diagnosing: `prologue` blames shard preparation (registry or options resolution), `descent` blames the traversal to the start leaf, `leaf-walk` blames a single in-flight leaf read, and `baseline-fold` blames the second, fanned-out pass of a snapshot baseline capture, which folds the frozen leaves' WAL tails back onto their frozen caches. Since issue 2585 the ceiling faults the call only when it caught the walk with **nothing to bank**; a fire that had completed work banks it as a short page the caller resumes from and never reaches this counter, so read it as the discarding subset of ceiling fires and `shard_root.scan_page.ceiling_outcomes` for the whole. What counts as bankable work widened in issue 2807: a page fill banks its rows, and an aggregate walk - a count, an emptiness probe, a diagnostics, storage-usage, projection-rebuild or materialiser-lag sweep - banks a finished partial page carrying its running total and a resume key, so a count-heavy or diagnostics-heavy tree that used to record every fire here now records only the fires that caught it inside its first leaf read. Expected to be zero in steady state. |
| `orleans.lattice.shard_root.flush.retries_suspended` | `Counter<long>` | `{suspension}` | Count of `ShardRootGrain` coalescing background-flush loops that reached the consecutive-failure ceiling and suspended themselves for the remainder of the activation. Tagged `tree`, `shard`, and `kind` (`leaf-access`, `dirty-leaves`). Both loops previously re-armed on every failure and so retried for the life of the activation, which turned a permanently-failing write - most commonly an activation whose ETag no longer matches its stored row - into an unbounded write load that logged only at `Debug`. A non-zero value names the shard whose background flushes have given up, and is the operator-visible signal that accompanies the suspension warning. What follows depends on the failure. On a transient storage fault the loop suspends and the activation stays up: pending state is retained in memory, deactivation still attempts a final best-effort write, and both subsystems self-heal. On a repeated version (optimistic-concurrency) conflict - another writer advanced the stored row, so no retry from this activation can be accepted - the shard root also requests its own deactivation so the next activation re-reads the row; the pending state it held is dropped rather than written, which both subsystems already recover from (dirty marks are rediscovered by the leaf-chain walk, and the leaf-access model rebuilds from live traffic), so the cost is a colder shard, never a wrong answer. Expected to be zero in steady state. |
| `orleans.lattice.shard_root.scan_page.ceiling_outcomes` | `Counter<long>` | `{fire}` | Count of `ShardRootGrain` page-fill stall-ceiling fires, tagged `tree`, `shard`, and `outcome` by what the fire did with the work the walk had already done. `outcome=banked` means the walk had completed work to show, so it was returned as a short page the caller resumes from. For a page fill that is rows, returned with `HasMore` set and no `ResumeFromKey` - the shape the cooperative `MaxScanPageDuration` budget already emits, which every cursor advances past by taking the last row's key as its next continuation token. For an aggregate walk, which has no rows to return, it is a finished partial page carrying the running total and the exclusive high bound of the last leaf completed as its `ResumeFromInclusive` (issue 2807). `outcome=discarded` means the fire caught the walk with nothing collected, so the call faulted with `ScanPageStalledException`; banking an empty page there would be strictly worse than the fault, because a page with no rows and no resume key is how every caller recognises the end of a scan, so it would silently truncate the result set instead. Before issue 2585 every fire discarded, which made the ceiling a livelock rather than a bound: the retry it invites re-walked the same leaves, hit the same ceiling and discarded the same work, so a page that could not fill in one attempt could not fill in any number of them. **Reading a zero:** both arms are emitted from the one site a fire passes through, so `banked + discarded` is the fire count and neither arm needs a denominator from elsewhere - a zero `banked` beside a non-zero `discarded` has **three distinct causes and does not separate them on its own**, and it specifically does *not* indicate a parking prologue or descent - an earlier revision of this row said so, and the `phase` tag refutes it, because a fire that parks before the leaf chain never reaches `ScanPagePhase.LeafWalk` and is tagged `prologue` or `descent` rather than `leaf-walk`. **Cause 1, and the one seen in the field:** the fire landed inside the page's *first* leaf read. The accumulator is published before the phase flips to `leaf-walk`, and a leaf contributes its rows only once its read returns in full, so a read still in flight has banked nothing; a `leaf-walk` stall with `banked=0` on a paging operation therefore means no leaf read completed at all. Banking cannot help this shape, and where the range fits in a single leaf it never can, since there is no earlier completed leaf to have contributed rows however small that leaf is - the mitigation is the read coalescing reported by `shard_root.scan_page.leaf_read_outcomes`, not banking, so read its `joined` arm against `shard_root.scan_page.stalls` to tell a retry that attaches to the in-flight read from one that enqueues another behind it. **Cause 2:** the operation cannot bank at all. Two stall-guarded operations still record `discarded` however many leaves had completed, and both are deliberate rather than unwired: `CaptureSnapshotBaselineAsync` has no meaningful partial, since a baseline covering part of a chain is not a baseline, and `DeleteRangeBoundedAsync` publishes its replication notification after the walk, so a banked resume key would carry the caller past a prefix whose tombstones were applied locally and never published, orphaning that closure permanently - the fault is retried from the range start instead, which re-publishes it. On those two series `banked=0` is not a statement about the walk at all. Before issue 2807 this cause covered ten of the sixteen guarded operations, because banking required the core method to publish a row accumulator through `BeginScanPageRows` and to return `KeysPage` or `EntriesPage`, the only two shapes `TryBankPartialScanPage` could construct, which the counting, any, diagnostics, storage-usage, projection-rebuild and materialiser-lag pages all failed; those eight now publish a finished partial page at each leaf boundary through `PublishScanPagePartial`, so a dashboard or alert written against the older reading - that a count-heavy or diagnostics-heavy tree necessarily shows a zero `banked` arm - is measuring the old behaviour and will now misreport. **Cause 3:** every row read so far was filtered out, since moved-away slots are skipped before they reach the accumulator, so a page that has walked several completed leaves holding only moved-away entries still banks nothing - expect this only while a shard is consolidating. It follows that the banked-to-discarded ratio is mix-dependent and is not a health signal on its own: a shard whose scan traffic is dominated by range deletes or snapshot baselines reports a zero `banked` arm however healthy its leaf reads are, while one serving key, entry, count and diagnostics pages reports a ratio that does track first-leaf health, so comparing it across trees compares their operation mixes as much as their leaf latency. Both arms at zero means simply that no ceiling fired. The misreading this closes off is the series being **absent**: no points here while `shard_root.scan_page.stalls` climbs means the banking path is not wired up, which is a broken measurement rather than a clean shard. `discarded` equals `shard_root.scan_page.stalls` by construction - the same fire raises both - so a divergence is itself a wiring fault. Expected to be zero in steady state. |
| `orleans.lattice.shard_root.scan_page.leaf_read_outcomes` | `Counter<long>` | `{read}` | Count of bounded leaf reads issued by a **stall-guarded** `ShardRootGrain` page fill, tagged `tree`, `shard`, and `outcome` by whether the read reached the leaf, attached to one already in flight, or was served again from a settled read whose leaf revision cookie had not moved. This is the convergence counterpart to `shard_root.scan_page.ceiling_outcomes`, and the two are not substitutes: that series reports whether a ceiling fire kept the rows it had, this one reports whether the **next** attempt had to pay for them again. Banking without this is only half a fix, because a walk that banks nothing on every attempt - the `after 0 leaf/leaves` case, where the ceiling fires before the first leaf read returns - has no rows to bank, so the banking path is never reached and the retry restarts at the same continuation token forever. `outcome=issued` means no argument-identical read was held, so a fresh call went to the leaf; this is the ordinary steady state of a healthy scan as well as the signature of the defect, so it is diagnostic only when read *against* `shard_root.scan_page.stalls` and never on its own. `outcome=joined` means an argument-identical read was still in flight and this attempt attached to it rather than enqueuing a second one. That matters because `Task.WaitAsync` ends the *wait* and never the *call*, and `IBPlusLeafGrain` is deliberately non-reentrant: an abandoned read keeps its place in the leaf's queue, so before this counter existed each retry enqueued another argument-identical read **behind** the one it had just abandoned. The retry loop therefore diverged rather than merely failing to converge - after N attempts the leaf held N queued reads for the same rows - and the identical-looking failures were each strictly worse than the last, alike only because the ceiling truncated them all at the same number. Joining is not a staleness compromise: because the leaf is non-reentrant the joined read holds that leaf's turn for its whole duration and every write is ordered strictly before or after it, so a joiner observes exactly what the original caller would have. **The third arm, `outcome=served`, is real and is not recency-based reuse.** An earlier revision of this fix did retain settled results for one ceiling and emit a `served` arm, reasoning that the window was short; that is wrong in kind rather than in degree, because once the leaf's turn ends later writes are ordered after the read, so replaying its rows returns a scan page that misses committed writes - incorrect at any window length, so both the clause and the arm were removed. Issue #2786 reinstated `served` on a different and sufficient basis: the leaf's activation-fenced revision cookie, sampled before the read is issued and compared at attach time, so a settled read is served again only when the leaf published no mutation across a window that strictly contains it, no row it surfaced has since expired, and it resolved no uncommitted transactional write, whose registry decision can change with nothing written to the leaf (issue #2823). Read the two arms differently - `joined` is serialisable by construction, whereas `served` rests entirely on that cookie comparison and is therefore the arm that falls to zero if the cookie stops being published on some mutation path. **Reading a zero:** all three outcome arms are primed at zero from shard-root **activation**, through the same recorder the live path uses, so a zero here is a measured absence rather than an absent measurement. The priming point is load-bearing and was moved deliberately (issue #2809): it was once taken on the guarded read path, below the `!IsStallGuarded` early return, so an absent series was ambiguous between *the build does not carry the instrument* and *it does, but no stall-guarded scan-page leaf read ever ran on that `(tree, shard)`*. Primed from activation the series exists for every activated `(tree, shard)` with no scan traffic of any kind, so an absent series has exactly one cause again and is usable as a build-provenance signal. Note what that costs the older reading: a flat zero across all three outcome arms no longer says the walks on this tree are unguarded, because it is now also what an idle shard looks like - take the guarded-or-not question from `shard_root.scan_page.stalls` instead. Priming is per activation, so partial presence across `(tree, shard)` pairs means only that some shard roots have not activated. The livelock signature is `issued` climbing in step with `shard_root.scan_page.stalls` while `joined` stays at zero: every retry re-reading the same leaf from scratch. Expected to be non-zero on `issued` and zero on `joined` and `served` in steady state, since a healthy page fill never stalls and so never has a read to rejoin or to re-serve. |
| `orleans.lattice.shard_root.scan_page.zero_progress_stalls` | `Counter<long>` | `{stall}` | Count of `ShardRootGrain` page-fill ceiling fires that completed **no leaf at all**, tagged `tree`, `shard`, and `outcome` by whether the leaf the fire named in flight is merely slow this once (`outcome=slow`) or has now been classified unreadable after `ShardRootGrain.StrandedLeafStallThreshold` consecutive zero-progress fires naming that same leaf (`outcome=stranded`). This is the series that separates a shard that is retrying from a shard that is wedged, and neither of the two counters beside it can do so. `shard_root.scan_page.stalls` counts abandoned page fills but says nothing about whether successive fills are the *same* fill over again, so 307 stalls on one leaf and 307 stalls spread over 307 leaves that each then recovered read identically on it; `shard_root.scan_page.ceiling_outcomes{outcome=discarded}` equals that series by construction and so inherits the same blindness. Consecutiveness on one leaf identity is the whole signal, which is why a single fire is deliberately recorded as `slow` and never as `stranded`: a leaf replaying a long WAL window from cold and a leaf that will never answer produce byte-identical first attempts, and recovering the first out from under itself would be a regression. **Reading a zero:** both arms are primed at zero from shard-root **activation** through the same recorder the live path uses, so an absent series means the build does not carry the instrument rather than that no fire ever happened, and a zero `stranded` beside a non-zero `slow` is a measured statement that stalls are occurring and recovering. The reading that matters operationally is `stranded` climbing at all: it says a specific leaf has been unreadable across a run of attempts and that the shard root has stopped waiting on the read it was parked on, so every attempt after the classification is a genuinely fresh read rather than a re-attach to a read that will never return. Sustained `stranded` on one `(tree, shard)` after that point localises the fault to inside the leaf activation itself, which the shard root cannot repair and which no retry will clear, and is the cue to escalate rather than to keep retrying. Read it against `shard_root.scan_page.leaf_read_outcomes`: its `joined` arm is healthy coalescing when stalls recover and is the mechanism of the livelock when they do not, and the two arms here are what tell those apart. Expected to be zero in steady state. |
| `orleans.lattice.shard_root.scan_page.chain_regressions` | `Counter<long>` | `{regression}` | Count of range-scan rows suppressed from leaves that regressed the scan's chain watermark, tagged `tree`, `shard`, and `outcome` by whether the event is one more suppression (`outcome=suppression`) or the first time this activation has seen that particular leaf regress (`outcome=distinct-leaf`). A leaf whose keys sit at or behind the watermark is not reachable by descent for the range it claims, so its rows duplicate a live leaf and are dropped (issue 3271). This counter exists because the suppression used to be reported **only** in the log, and the log could not carry it: one field burst emitted 110,322 warnings in about eight minutes, roughly 2,354 lines a second and 99.3% of all warning output, which filled half of a 100 MB container log ring and cut that host's log retention to about 108 seconds. Nothing on the box could then be diagnosed after the fact. Read the two arms apart, because they answer different questions and their ratio is the diagnosis. `distinct-leaf` is the **census**: how many leaves in this shard are damaged, which is the number an operator acts on and the number that does not change however often the shard is scanned. `suppression` is the **rate**: how much read work is currently being spent stepping over them, which rises with scan traffic on an unchanged set of damaged leaves. A large `suppression` over a small `distinct-leaf` therefore means a few damaged leaves re-encountered on every page fill rather than spreading damage - exactly the field shape, where 2,886 distinct leaves were met about 43 times each because the in-page dedupe only holds within a single page. `distinct-leaf` climbing is the reading that matters: it is new damage, and it is the cue to escalate. Both arms count per shard-root activation, so `distinct-leaf` restarts from zero when the shard root reactivates and is a census of *this* activation rather than a lifetime total. The log is bounded to one detailed line per distinct leaf up to `ShardRootGrain.ChainRegressionWarnDetailCap` plus one summary line, so beyond that cap this counter is the **only** report of the condition (issue 3341). It is deliberately not tagged by leaf identity: leaf ids are unbounded cardinality, and a label set that grows with the damage is a second outage waiting behind the first. **Reading a zero:** both arms are primed at zero from shard-root activation through the same recorder the live path uses, so an absent series means the build does not carry the instrument rather than that the chain is intact, and a measured zero is a positive statement that this shard suppressed nothing. Expected to be zero in steady state; any non-zero reading means the shard's leaf chain needs repair (issue 3269). |
| `orleans.lattice.scan.stall_resumptions` | `Counter<long>` | `{resumption}` | Count of client-side resilient scans (`ScanKeysAsync` / `ScanEntriesAsync`, on both `ILattice` and `ILatticeView`) that met a `ScanPageStalledException`, tagged `tree`, `phase` (carried through from the stall) and `outcome`. `outcome=resumed` means the scan resumed from its last continuation token and still completed in full - read `resumed` against the terminal outcomes *within this same series*, because a scan that succeeded only after N resumptions would otherwise be indistinguishable from one that never stalled, and a worsening contention trend would be hidden by its own recovery. Do NOT read it against `shard_root.scan_page.stalls`: **the two series count different populations and no ratio between them is meaningful.** This series counts only stalls met by the resilient scan extensions named above, whereas `shard_root.scan_page.stalls` counts every abandoned page fill regardless of caller - and the scan surface has many direct consumers (the backup sink, the grain index, membership, replication, and internal registry walks) that receive a `ScanPageStalledException` and increment nothing here. A deployment can therefore show tens of shard-root stalls against a handful of resumptions with nothing wrong: the difference is caller coverage, not lost work, since a direct caller sees the fault thrown rather than silently truncated. Dividing one by the other yields a "recovery rate" that is not one. `outcome=budget-exhausted` means the scan had already resumed its permitted number of times (`LatticeExtensions.DefaultScanStallResumeAttempts`, or a lower caller `maxAttempts`) and rethrew without banking a record between those consecutive stalls. `outcome=ceiling-exhausted` is the opposite diagnosis with the same symptom: the scan *was* banking records between stalls and instead exceeded the lifetime resumption ceiling (`LatticeExtensions.DefaultScanStallResumeCeiling`), so it reports that the ceiling is too small for this workload rather than that the source is dead. Those two are the terminal outcomes and the strongest single contention signal on this surface; read them apart, because `budget-exhausted` calls for relieving contention on the source while `ceiling-exhausted` calls for raising the ceiling or for the caller to bank partial progress so a terminated walk resumes rather than restarts. `ceiling-exhausted` has not been observed in the field, which is not evidence that the lifetime ceiling is correctly sized: a walk can only reach it by surviving the consecutive-futility bound repeatedly, and walks are removed at that bound far earlier, so the ceiling remains untested rather than validated. A resumption never truncates: the scan either yields its full range or rethrows the last stall verbatim, so every `budget-exhausted` is a scan whose caller saw the failure. An `outcome=no-progress` label was emitted by earlier builds when a progress gate refused a stall that still had budget; that gate refused every stall that occurred in practice and has been removed, so the label no longer exists. Expected to be zero in steady state. |
| `orleans.lattice.scan.stall_futility_outcomes` | `Counter<long>` | `{outcome}` | Sibling series to `scan.stall_resumptions{outcome=budget-exhausted}`, answering the one question that counter is structurally unable to: whether the source a futility termination gave up on was merely busy or genuinely not yielding. Tagged `tree`, `phase` (carried through from the stall that ended the walk) and `outcome`. Under load those two conditions are the same observation at the consecutive-futility bound, so `budget-exhausted` on its own records only that a walk was killed, never whether the shard it abandoned served that same region moments later. Each `budget-exhausted` termination therefore opens a short-lived watch keyed by the abandoned source and position; the watch never re-drives the abandoned work, so this series observes the termination policy without altering it, and it resolves to exactly one of four outcomes. `outcome=recovered` means a later walk against that same source yielded a record at or beyond the abandoned position inside the window - the bound cut off a recoverable source, and a sustained non-zero rate is the signal that `LatticeExtensions.DefaultScanStallResumeAttempts` is too tight for this workload. `outcome=still-stalled` means a later walk against that source was itself killed for futility before reaching the abandoned position - the source really was not yielding and the bound is doing its job. `outcome=unobserved` means nothing revisited the source before the window closed, so the run says nothing either way. `outcome=dropped` means the watch was evicted under the registry capacity bound before it could resolve. Read the last two before concluding anything from the first: a zero `recovered` rate is evidence the bound is correct only when `still-stalled` is non-zero, because a run whose watches all end `unobserved` or `dropped` reads exactly the same zero for an entirely different reason. A low reading is not a verdict unless the run was strained: this series can only speak about a bound that came under pressure, so before reading a small `recovered` count as evidence the bound is well sized, confirm that `scan.stall_resumptions{outcome=budget-exhausted}` is non-zero (otherwise no watch was ever opened), that `recovered + still-stalled` is non-zero (otherwise nothing revisited the abandoned sources), (a third check on `shard_root.scan_page.stalls` was previously listed here and has been removed: it is strictly implied by the first check, since a `budget-exhausted` termination cannot occur without stalls, and it invited a comparison across two different caller populations - see that series' entry above). Where any of those fails, the finding is that the bound was not exercised, not that it is correct - a bound that was never strained is untested, in exactly the way a `ceiling-exhausted` of zero does not validate the lifetime ceiling. The sum across the four outcomes is bounded above by `scan.stall_resumptions{outcome=budget-exhausted}` and is not expected to equal it, since a watch still open when the process ends is never resolved. Expected to be zero in steady state, because a steady-state scan never exhausts its consecutive resume budget at all. |
| `orleans.lattice.internal.digest_publish.timeouts` | `Counter<long>` | `{timeout}` | Count of internal-node upward digest publishes (the `ChildDigestSnapshot` propagation from a `BPlusInternalGrain` to its parent) abandoned after exceeding `DigestPublishTimeout`. Tagged `tree`. A non-zero value indicates a publish parked against a parent internal node that was mid-mutation; the publish was faulted as a `TimeoutException` so the holding turn released the non-reentrant split gate instead of pinning it. The digest is staleness-tolerant, so the next mutation's publish re-drives convergence. Expected to be zero in steady state. |
| `orleans.lattice.wal.append_dispatch.timeouts` | `Counter<long>` | `{timeout}` | Count of writer-side WAL partition dispatches - the single-entry and the batched append the writer sends to a WAL partition grain - abandoned after exceeding `WalAppendDispatchTimeout` (default 30 seconds; `Timeout.InfiniteTimeSpan` disables the deadline and with it this counter). Tagged `tree` and `shard`. The dispatch is the writer-side cross-grain RPC into the per-shard WAL grain; it was historically unbounded on the writer side, so a wedged shard activation would hold every caller's dispatch parked until the Orleans response deadline (`MessagingOptions.ResponseTimeout`, 30 seconds by default) expired. A non-zero value attributes the wedge to a specific `(tree, shard)` pair and counts the trip toward the saturation signal's dispatch-timeout input, and the parked dispatch is faulted as a `TimeoutException` so the request pipeline releases its slot. At the defaults the two deadlines are equal, so the counter adds attribution rather than an earlier fault; it faults a parked dispatch sooner than Orleans would only where `WalAppendDispatchTimeout` is set below the silo's `ResponseTimeout`. Sustained non-zero counts on a specific `(tree, shard)` identify the wedged shard for follow-up investigation. Expected to be zero in steady state. |
| `orleans.lattice.wal.flush.preflight.timeouts` | `Counter<long>` | `{timeout}` | Count of WAL partition `FlushAsync` preflight regions (the synchronous setup and initial scheduler yield that precede the bounded provider call) abandoned after exceeding `WalFlushPreflightTimeout`. Tagged `tree` and `shard`. The preflight region is normally microseconds; a non-zero count indicates the activation's grain scheduler did not resume the flush's post-yield continuation within the deadline, leaving the in-flight slot pinned with no provider-call deadline armed (`WalFlushTimeout` only covers the provider call itself, which has not yet been issued). The faulted preflight surfaces as a `TimeoutException` routed through the normal failure handler, the slot drains, and this counter attributes the trip. Sustained non-zero counts indicate the activation's scheduler is being held by a startup reshard / membership change, a non-cooperative work item, or a mid-flush activation tear-down. Expected to be zero in steady state. |
| `orleans.lattice.wal.shard.deactivate.in_flight` | `Histogram<long>` | `{slot}` | Per-WAL-shard in-flight slot count observed at `OnDeactivateAsync` time. Tagged `tree` and `shard`. Recorded exactly once per `OnDeactivateAsync` call. A zero observation is the healthy steady-state shape (the grain drained cleanly); a non-zero observation means the activation was torn down with in-flight flushes still pending, the slot population that defines the post-#568 residual phase-1/activation wedge fingerprint. Combined with `orleans.lattice.wal.flush.preflight.timeouts`, a deactivation with non-zero in-flight count immediately followed by a preflight timeout on a successor activation is the smoking gun for the "mid-call deactivation orphan" hypothesis. |
| `orleans.lattice.wal.shard.drain.budget.expirations` | `Counter<long>` | `{expiration}` | Count of per-shard `WalShardGrain` deactivation drains that exceeded `WalDrainBudget` and had to force-fault one or more in-flight slots so the activation could finish tearing down. Tagged `tree` and `shard`. Reliability intent: under a saturating-storage-account wedge, the provider call's await can park behind an SDK retry loop in pre-attempt back-off where the per-flush `WalFlushTimeout` deadline does not fire promptly (the SDK observes cancellation only between attempts, not during back-off), so a chain with N in-flight slots could otherwise hold the deactivation indefinitely. With the drain budget the deactivation force-faults any slot that has not unlinked within the deadline; this counter names the wedged shard so operators can attribute the trip without source-walking the silo log. Zero on a healthy drain; any non-zero rate identifies a shard whose provider call could not be cancelled inside the drain budget. |
| `orleans.lattice.wal.shard.drain.budget.force_faulted_slots` | `Histogram<long>` | `{slot}` | Per-WAL-shard in-flight slot count force-faulted by a deactivation drain after `WalDrainBudget` expired. Tagged `tree` and `shard`. Recorded exactly once per drain that hit the budget; the value is the number of slots that had not unlinked when the budget fired and were force-faulted to release the activation. Only fires on the `wal.shard.drain.budget.expirations` path, so the histogram's count and the counter's count are the same number. |
| `orleans.lattice.wal.shard.start_flush.calls` | `Counter<long>` | `{call}` | Count of `WalShardGrain.StartFlush` invocations per `(tree, shard)`. Incremented once at the top of every `StartFlush` call, including the follow-on flushes a completing flush kicks off. Tagged `tree` and `shard`. Diagnostic intent: under the residual phase-1/activation WAL wedge, if `start_flush.calls` keeps incrementing throughout the wedge then new flushes ARE being kicked off, so the wedge is a slot-leak in the in-flight chain's `finally` (slots never removed even after the flush's task settles); if `start_flush.calls` stops incrementing during the wedge then the cap-cutover loop in `AppendBatchAsync` is itself blocked and no new flush ever kicks off. |
| `orleans.lattice.wal.shard.pending_segments` | `Histogram<long>` | `{segment}` | Per-WAL-shard `_pendingSegments.Count` observed at every `StartFlush` entry, sampled *before* the pending list is captured into the new in-flight slot. Tagged `tree` and `shard`. Diagnostic intent: under the wedge, a growing distribution indicates callers are still arriving and enqueueing into `_pendingSegments` even though the chain cannot drain (back-pressure absorbing everything but never releasing). A stuck-at-zero distribution combined with a `start_flush.calls` trickle indicates the cap-cutover loop blocked itself; combined with a healthy `start_flush.calls` rate it indicates the wedge is downstream of the flush kick-off. |
| `orleans.lattice.shard_root.reshard.initiated` | `Counter<long>` | `{reshard}` | Count of reshard requests (`ILattice.ReshardAsync`) that passed validation and started the tree's reshard coordinator, or that took the empty-tree fast path instead - re-pinning the shard count and rebuilding the default shard map in the registry without starting one. Tagged `tree`. Diagnostic intent: the residual WAL wedge is correlated with the `reshard ... REJECTED (Forwarding failed)` log storm; pairing this counter with `reshard.completed` and `reshard.rejected` lets a dashboard correlate reshard activity with wedge onset directly without grepping a rotated silo log. Note: Orleans-side message-routing rejections ("Forwarding failed") are emitted by Orleans's own router and are not captured here - they remain log-only until a separate diagnostic source is added. |
| `orleans.lattice.shard_root.reshard.rejected` | `Counter<long>` | `{rejection}` | Count of reshard requests (`ILattice.ReshardAsync`) the tree's reshard coordinator refused before starting a reshard. Tagged `tree` and `reason`, where `reason` enumerates the rejection class (`argument_out_of_range_min`, `argument_out_of_range_max`, `already_in_progress`, `resize_in_flight`, `state_write_failed`); every `reason` is primed at zero whenever a request reaches the coordinator (issue #2918), so a flat zero is a measured zero. A call refused before it reaches the coordinator - on a reserved system tree, on a materialised-view tree, or by authorization - is not counted, and neither is an in-range request the tree already satisfies (the shard count it already has, or the target of the reshard already in progress), which is a no-op. Excludes Orleans-side message-routing rejections; see `reshard.initiated`. |
| `orleans.lattice.shard_root.reshard.completed` | `Counter<long>` | `{reshard}` | Count of reshards whose coordinator reached its terminal phase successfully - for a shrink, only once no fold it started is still running - recorded after the reshard's terminal state write succeeds (and the empty-tree fast-path reshards, counted in lockstep with `reshard.initiated`). Tagged `tree`. The difference between this and `reshard.initiated` over a window is the number of reshards still in flight or that failed mid-coordinator. |
| `orleans.lattice.shard_root.reshard.in_flight` | `Histogram<long>` | `{reshard}` | Per-tree reshard in-flight state observation, emitted at every `ReshardAsync` entry as either `0` (idle) or `1` (a reshard is already in progress for this tree). Tagged `tree`. A non-zero observation immediately preceding wedge onset is the same signal a periodically-polled gauge would provide. |
| `orleans.lattice.wal.writer.append.dispatched` | `Counter<long>` | `{dispatch}` | Count of `WalCommitLogWriter` per-partition append dispatches that started (incremented at the `Enqueued` lifecycle stamp, before the shard RPC is invoked). Tagged `tree` and `partition`. Diagnostic intent: the writer-layer kick-off signal for the saturation-rung WAL wedge's dominant mode (5 of 7 wedged cohorts on 2026-06-03), where every shard's in-flight chain is empty yet hundreds of callers are parked in `WalCommitLogWriter.AppendForPartitionAsync`. A sustained dispatched rate combined with stale `wal.writer.partition.pending_appends` p99 readings localises the stall to the awaited shard-grain RPC; a collapse to zero ranges the wedge upstream of the writer itself. |
| `orleans.lattice.wal.writer.partition.pending_appends` | `Histogram<long>` | `{dispatch}` | Per-writer-partition pending-append-dispatch depth observed at every `WalCommitLogWriter` append entry, sampled *before* the new pending stamp is linked into the partition's tracker. Tagged `tree` and `partition`. Diagnostic intent: a growing distribution under the wedge confirms the writer is the choke (callers enqueuing into a tracker that cannot drain); a stuck-at-zero distribution combined with sustained `wal.writer.append.dispatched` rules out a writer-layer dispatch lifecycle stall and points the next bisect downstream of the `SentToShard` stage. Mirrors `wal.shard.pending_segments` one layer up. |
| `orleans.lattice.wal.writer.append.admission_timeouts` | `Counter<long>` | `{timeout}` | Count of `WalCommitLogWriter` append dispatches whose per-partition admission wait exceeded `WalAppendDispatchTimeout`. Tagged `tree` and `partition`. Reliability intent: the per-partition admission semaphore caps `PartitionTracker._inFlight` depth at `WalMaxPendingBatches`, mirroring the shard-side ceiling. When the shard cannot drain, callers awaiting an admission slot are released with a typed `TimeoutException` at the deadline rather than silently parking forever in an unbounded writer queue. A non-zero counter under steady-state operation is the signal that the offered rate exceeds the shard's drain rate - the saturation regime previously hidden as a silent wedge. Pair with `wal.writer.append.admission_wait` to distinguish back-pressured-cleanly (wait p99 elevated, zero timeouts) from back-pressure-exceeded-deadline (non-zero timeouts). |
| `orleans.lattice.wal.writer.append.admission_wait` | `Histogram<double>` | `ms` | Wall-clock ms a `WalCommitLogWriter` dispatch waited for a per-partition admission slot before linking a new `PendingAppend` stamp. Tagged `tree` and `partition`. Reliability intent: under healthy operation this histogram sits at the floor (a sub-microsecond uncontended semaphore acquire). A spreading distribution indicates the per-partition tracker is approaching its `WalMaxPendingBatches` ceiling, surfacing back-pressure as an honest tail-latency signal long before any caller hits the `wal.writer.append.admission_timeouts` deadline. Recorded for every dispatch that successfully acquired a slot (timed-out dispatches feed the counter only). |
| `orleans.lattice.wal.writer.append.drain.releases` | `Counter<long>` | `{release}` | Count of writer-side parked admission callers released by a silo-drain signal on host shutdown. Tagged `tree` and `partition`. One sample per parked caller faulted out of `PartitionTracker.AcquireAsync` when the owning `WalCommitLogWriter` drains on host shutdown; zero on a healthy shutdown that has no parked callers. Distinct from `wal.writer.append.admission_timeouts` (per-call deadline expiries during steady-state operation) and from `wal.shard.drain.budget.expirations` (shard-grain deactivation drains that had to force-fault). This counter names writer-side parked callers released by the silo's drain on shutdown - the surface that closes the writer-admission-semaphore-wedged-at-SIGTERM phenotype documented in `benchmark/azure-throughput/throughput.md` section 32.6. A non-zero rate on shutdown is normal when the silo was under storage saturation at drain entry; a non-zero rate during steady-state operation indicates the drain hook fired spuriously and is a regression signal. Per-silo: each silo process emits its own samples for the trackers its `WalCommitLogWriter` owns. |
| `orleans.lattice.wal.saturation.state` | `ObservableGauge<long>` | `{state}` | Current per-tree WAL saturation regime as an ordinal step function: `0` = Healthy, `1` = Throttled, `2` = Saturated. Tagged `tree` and the derived `tenant` label only - the ordinal value already encodes the regime, so the state is deliberately not also carried as a label (doing so fragmented the per-tree series on every transition and left the prior state's series lingering at its last elevated value under scrape staleness, making a recovered tree read as Saturated). Published from the per-tree state the silo-scoped saturation sampler maintains; a tree contributes a measurement only after the sampler has observed at least one signal for it, so an unwritten tree does not appear in the series. Pair with `wal.saturation.transitions` (which carries the per-state breakdown) to plot regime flips alongside the current regime. See [WAL Saturation Signal](../wal-saturation-signal.md). |
| `orleans.lattice.wal.saturation.transitions` | `Counter<long>` | `{transition}` | Count of per-tree saturation-state transitions observed by the silo-scoped sampler. Tagged `tree`, `tenant`, `state` (the new state, lowercased: `healthy`, `throttled` or `saturated`), `previous_state` (the same spelling), and `cause`, which names the single sampler input the transition was attributed to: `dispatch_timeouts`, `provider_failures`, `flush_latency` or `admission_depth` on a move to `saturated`; `admission_depth`, `materialiser_drain_lag` or `materialiser_pin_latency` on a move to `throttled`; and `none` on every move back to `healthy` and on a `throttled` held only by the recovery window. `flush_latency` and `materialiser_pin_latency` need their opt-in thresholds (`LatticeOptions.WalSaturationFlushLatencyThreshold`, `LatticeOptions.WalSaturationMaterialiserPinLatencyThreshold`), and `admission_depth` reaches `saturated` only with `LatticeOptions.WalSaturationAcuteOnly` switched off. Attribution follows the classifier's own evaluation order, so it never disagrees with `state`, and when several inputs cross in one window the first evaluated wins. Two location tags are added independently of `cause`: `partition` whenever a WAL partition carried admission depth or parked callers on that tick (the one with the highest depth ratio, otherwise the first with parked callers), and `shard` whenever one of the four shard-attributed inputs - dispatch timeouts, provider failures, flush-latency trips and durable pin-write latency trips - advanced in the window (the first such shard observed, taking the inputs in that order). Materialiser drain lag is tree-wide and adds neither. A transition can therefore carry both, or a location unrelated to its `cause`: read `cause` for what drove it and `partition` / `shard` only as where pressure was seen. A flat-zero series is the healthy steady state. A rising rate of `state=throttled` is the leading edge of a saturation episode; `state=saturated` is the regime itself. Flapping between Throttled and Saturated is a different operational signal from a sustained Saturated. |

## WAL append pipeline (sourced from `WalShardGrain` and `WalCommitLogWriter`)

The WAL partition grains expose per-flush instruments that distinguish
grain-side scheduling cost from storage-provider cost; every histogram
carries the Phase A attribution tags `wal_partitions` and
`wal_max_pending_batches` so cross-configuration comparisons are direct.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.wal.shard.dispatch.duration` | `Histogram<double>` | `ms` | Caller-side wall-clock duration of the cross-grain `IWalShardGrain.AppendAsync` / `AppendBatchAsync` RPC, observed by `WalCommitLogWriter`. Tagged `tree`, `shard` (WAL partition index), `wal_partitions`, `wal_max_pending_batches`. Subtracting `wal.append.turn_wait` from this isolates the Orleans scheduling tax on the single WAL activation per partition. |
| `orleans.lattice.wal.shard.dispatch.entries` | `Histogram<int>` | `{entry}` | Per-dispatch entry count handed to `IWalShardGrain.AppendAsync` / `AppendBatchAsync` by `WalCommitLogWriter`. Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. Single-key sends record `1`; batched sends record the per-partition slice size. |
| `orleans.lattice.wal.append.batch_entries` | `Histogram<int>` | `{entry}` | Per-flush packing inside the WAL grain: how many entries the grain's cutover loop accumulated into the batch that the storage provider ultimately sees. Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. Pair with `wal.shard.dispatch.entries` to detect a missing cross-`AppendBatchAsync` coalescing window. |
| `orleans.lattice.wal.append.batch_bytes` | `Histogram<long>` | `By` | Per-flush size of the packed batch handed to the storage provider, in bytes. Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. |
| `orleans.lattice.wal.append.in_flight` | `Histogram<int>` | `{flush}` | `_inFlight.Count` (in-flight flushes against the provider) sampled at the moment a new flush is admitted. Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. p99 sitting at `wal_max_pending_batches - 1` indicates the cap is the binding constraint. |
| `orleans.lattice.wal.append.provider.duration` | `Histogram<double>` | `ms` | Wall-clock duration of one storage-provider flush, measured inside the WAL grain (`FlushAsync` body). Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. This is the pure provider-RTT signal; the gap between `wal.shard.dispatch.duration` and this is everything Orleans + the grain do on top. |
| `orleans.lattice.wal.append.turn_wait` | `Histogram<double>` | `ms` | Despite its name, not the wait for the WAL activation's turn: the wall-clock duration from the moment an append's body starts inside the WAL grain - after the call has been given the turn - to the moment its entry's acknowledgement completes, so it covers any wait for an in-flight flush slot, the time the entry sits in the pending batch before cutover, and the provider flush itself (`wal.append.provider.duration`). Time queued for the turn before the body starts is not included; it shows up only in the gap between `wal.shard.dispatch.duration` and this. Recorded only by the exclusive-turn per-entry `IWalShardGrain.AppendAsync` overload, so with `WalBatchedSingleEntryAppends` on (the default) point and bulk appends, which take the interleaving `AppendBatchAsync`, are not observed (issue #812). Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. |
| `orleans.lattice.wal.append.queue_depth` | `Histogram<int>` | `{entry}` | Despite its name, not the turn queue: the number of entries in the WAL grain's pending (not yet flushed) batch, sampled when an append adds its own entry and counting it, so `1` means the entry arrived to an empty pending batch. Recorded only by the exclusive-turn per-entry `AppendAsync` overload, like `wal.append.turn_wait`. Tagged `tree`, `shard`, `wal_partitions`, `wal_max_pending_batches`. |
| `orleans.lattice.wal.writer.append.admission_saturation_refusals` | `Counter<long>` | `{refusal}` | Writer-side admission dispatches refused with `LatticeSaturatedException` because the per-tree saturation signal stayed `Saturated` beyond `WalAdmissionSaturationWaitBudget`. Tagged `tree` and `partition` (WAL partition index). |
| `orleans.lattice.saturation.refusals` | `Counter<long>` | `{refusal}` | Every `LatticeSaturatedException` raised, recorded at the seam that raises it (issue #3761). Tagged `tree`, `tenant` and `source`, whose value is the snake_case form of the `LatticeSaturationSource` the exception carries: `unspecified`, `wal_admission`, `atomic_write_saga`, `snapshot_cursor_open`, `replay_permit_admission`, `set_many_fan_out`, `set_many_envelope` or `tx_registry_capacity`. `unspecified` is mapped, but no seam records it. `set_many_envelope` is a `SetManyAsync` batch whose fan-out was still outstanding when the opt-in `LatticeOptions.SetManyEnvelopeBudget` ran out across all of the call's stages together, and `set_many_fan_out` one whose fan-out alone outlasted `SetManyFanOutBudget`; the two are kept apart because an envelope refusal can follow a perfectly healthy fan-out. `replay_permit_admission` also counts background starvation drives refused a replay permit, which are **not** raised: a WAL GC sweep's drive returns the `drove_admission_refused` verdict on `orleans.lattice.wal.gc.blocked_leaf_reactivations`, and a leaf's own coverage-lag timer drive is counted as `recheck_drive_refused` on `orleans.lattice.leaf.snapshot.driver.declines`, so a `replay_permit_admission` rate with no caller-visible exceptions is routine background refusal. A `replay_permit_admission` refusal also carries an `arm` tag naming the predicate arm that fired (issue #3921): `wait_exceeded` means queued waits are completing but slowly, so too much is queued; `no_progress` means no permit was released to the queue for `WalReplayPermitMaxQueueWait`, so the replays holding the permits are slow, typically bound on the grain store, and raising `WalMaterialiserMaxConcurrentReplays` makes it worse; `gc_share` is a background starvation drive refused a GC slot. **It counts refusals raised, not exceptions a caller sees.** A saga refusal that wraps an inner refusal counts under both sources, and one raised because the saga's own quiesce budget elapsed counts twice under `atomic_write_saga`; when the library retries a `replay_permit_admission` refusal internally, every refused attempt counts. When the saga wraps a refusal, an atomic-write caller sees only the saga's, whose `SaturationSource` is `AtomicWriteSaga` and whose `InnerException` is the wrapped one (a `wal_admission` or `set_many_fan_out` refusal, for example). In a cross-tree atomic write, a refusal raised while a participant stages its writes becomes a failed prepare vote, so that caller sees the write abort with an `InvalidOperationException` instead. `tree` is the raw id the refusing seam holds, not the resolved logical id most per-tree series use: `wal_admission` and `replay_permit_admission` carry the id of the WAL or leaf that refused, which is the physical copy's id on a resized, restored or remediated tree; `snapshot_cursor_open`, `set_many_fan_out`, `set_many_envelope`, `tx_registry_capacity` and the saga's own quiesce refusal carry the id the tree was addressed by; and a saga refusal that wraps another refusal carries that refusal's id (see [The `tree` dimension across aliasing](tag-conventions.md#the-tree-dimension-across-aliasing)). Not zero-primed, because `tree` is a runtime value. |

Previous: [Instrument catalog: WAL compaction (sourced from FileWalShard) to Foreground read envelopes (sourced from LatticeGrain)](instrument-catalog-5.md). Next: [Instrument catalog: Storage-provider commit pipeline to Grain-call observation (opt-in)](instrument-catalog-7.md). Contents: [Metrics](../metrics.md).
