---
title: "Instrument catalog: Leaf-level (sourced from BPlusLeafGrain) - Metrics"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/metrics/instrument-catalog-2.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/metrics.md?plain=1#L209-L295"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Instrument catalog: Leaf-level (sourced from `BPlusLeafGrain`)

Part of [Instrument catalog](instrument-catalog.md), in [Metrics](../metrics.md).

## Leaf-level (sourced from `BPlusLeafGrain`)

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.leaf.write.duration` | `Histogram<double>` | `ms` | Duration of a leaf grain's persisted state-row write (`IPersistentState.WriteStateAsync`) - i.e. storage-provider write latency - recorded with no `kind` tag. Those writes come from the leaf's topology and bookkeeping paths: sibling, parent, key-range, tree-id and shard-index updates, splits, moved-away slot marks, successor reclaim, snapshot load-hint banking, and projection checkpoint flushes and rebuilds. The same instrument also times three WAL appends, each tagged `kind`: `compact` (the append a tombstone compaction issues), `merge` (the append for a batched last-writer-wins merge into the leaf - for example replication apply, a tree merge, a snapshot copy, or a shard split or consolidation migration - and for the entries a leaf split transfers into its new sibling) and `backstop` (the cross-migration backstop append on a transaction terminal), so filter on `kind` rather than reading the unfiltered distribution as one population - the bundled Overview and CommitPath panels aggregate over it, and a query for the state-row writes alone selects `kind=""`. Tagged `tree`. |
| `orleans.lattice.leaf.commit.duration` | `Histogram<double>` | `ms` | Per-step latency on the leaf commit path. Tagged `tree` and `step=wal` (commit-log-writer append), `step=apply` (the in-memory merge, plus any relocation of out-of-span rows or leaf split the write triggers; the merge itself persists no grain state), `step=observer` (`IMutationObserver` fan-out), or `step=digest`, which hands the write's projection-digest change to the parent internal node: with the default [`DigestCoalescingWindowMs`](../configuration/options-reference-1.md#digestcoalescingwindowms) (5 ms) it schedules, or joins, a publish the leaf sends when the window elapses, outside the commit, so the step includes the awaited cross-grain publish itself only when the window is `0`. When the write leaves the digest unchanged, the leaf has no parent, or digest maintenance is off, the step instead covers the leaf's best-effort byte-footprint publish to its shard root. Emitted from every foreground commit-path code path (single-key `SetAsync` / `DeleteAsync`, the per-leaf batched `SetManyAsync` and conditional-batch commits, per-leaf `DeleteRangeAsync`, and saga prepare writes, which ride the same set paths) so operators can attribute latency between durability, projection, observer overhead and the digest hand-off independently. A per-leaf `DeleteRangeAsync` records no `observer` step: its observer publish is made once per shard, by the shard root after the leaf-chain walk, rather than per leaf. The structural digest publishes of a projection-checkpoint flush, a saga terminal, a tombstone-reap compaction or a cross-shard merge run outside these paths and are not recorded here, but a leaf split that a write triggers runs inside that write's `apply` step, its own inline digest publish included. |
| `orleans.lattice.leaf.commit.in_flight` | `Histogram<int>` | `{commit}` | Concurrent foreground commits in flight on a single leaf at the moment a new commit enters the commit path. Tagged `tree` only (no shard tag - leaves are routed per-key so the same leaf grain can be addressed by multiple shards under online reshard). The leaf's mutator surface is `[AlwaysInterleave]`, so a commit that arrives while another is still awaiting its WAL append enters the commit path rather than queueing behind the turn: a flat-zero series means commits never overlapped on one leaf activation, and a sustained tail is that overlap - saga-driven prepare arrivals or producer-side fan-in landing on the same leaf. |
| `orleans.lattice.leaf.scan.duration` | `Histogram<double>` | `ms` | Duration of leaf-level range scans. Tagged `tree` and `operation=keys` (a keys scan) or `operation=entries` (an entries scan). |
| `orleans.lattice.leaf.compaction.duration` | `Histogram<double>` | `ms` | Duration of one leaf's tombstone-compaction pass, clocked from the start of the grain call - the activation replay barrier included - to the end of its reap loop, so it measures the same quantity the pass's work budget bounds. A visit that returns early because the leaf is unchanged since its last complete compaction (counted `outcome=noop`), and a call that throws before its reap loop ends, record no sample. Every pass that truncated records a sample at or above [`BackgroundDrainMaxDuration`](../configuration/options-reference-1.md#backgrounddrainmaxduration), but the converse does not hold: the budget is checked only at intervals during the scan once an entry has been condemned, and after each removal while condemned entries remain, so a pass that condemned nothing (a cold activation's replay barrier alone can spend the budget) or whose budget ran out after its last check finishes, and records a sample above the budget as well. `outcome=partial` on `orleans.lattice.compaction.leaves.visited` is the exact count of truncated passes; a non-positive `BackgroundDrainMaxDuration` disables the bound, and then no pass truncates. Tagged `tree`, plus `trigger` (as on `orleans.lattice.compaction.pass.duration`) when at least one compaction policy knob (`MinTombstoneRatioForCompaction` or `MaxLeafEntriesBeforeForcedCompaction`) is non-default; with every knob at its default the tag is omitted. |
| `orleans.lattice.leaf.tombstones.reaped` | `Counter<long>` | `{tombstone}` | Tombstones (from explicit `DeleteAsync` / `DeleteRangeAsync`) permanently removed by compaction. Tagged `tree`, plus `trigger` under the same condition as `orleans.lattice.leaf.compaction.duration`. |
| `orleans.lattice.leaf.tombstones.created` | `Counter<long>` | `{tombstone}` | Tombstones newly written by `DeleteAsync` (1) or `DeleteRangeAsync` (N). Tagged `tree`. |
| `orleans.lattice.leaf.tombstones.expired` | `Counter<long>` | `{tombstone}` | Live entries reaped by compaction because their per-entry TTL (set via `SetAsync(key, value, TimeSpan)`) elapsed past the configured grace period. Separate from `reaped` so operators can distinguish TTL churn from explicit-delete throughput. Tagged `tree`, plus `trigger` under the same condition as `orleans.lattice.leaf.compaction.duration`. |
| `orleans.lattice.compaction.pass.duration` | `Histogram<double>` | `ms` | End-to-end wall-clock duration of one tombstone-compaction pass over the tree's shards, recorded once when the pass completes and measured from when the current coordinator activation started or resumed it, so a pass resumed after its coordinator reactivated reports only the resumed part. Tagged `tree`, `trigger=reminder` (baseline reminder-driven sweep), `trigger=ratio` (policy trigger fired by `MinTombstoneRatioForCompaction`), `trigger=size` (policy trigger fired by `MaxLeafEntriesBeforeForcedCompaction`), or `trigger=operator` (out-of-cycle pass requested via `ILattice.CompactShardAsync`). |
| `orleans.lattice.compaction.leaves.visited` | `Counter<long>` | `{leaf}` | Per-leaf compaction outcome. Tagged `tree`, `outcome=reaped` (at least one tombstone or TTL-expired entry physically removed and the leaf finished) or `outcome=noop` (nothing was removed: the leaf was unchanged since its last complete compaction, or every tombstone and expired entry it held was still inside the grace window) or `outcome=partial` (the leaf reaped what its bounded turn allowed and stopped on its work budget with condemned entries still outstanding, so its dirty mark is retained and a later pass re-nominates it - a sustained rate here means compaction is reclaiming ground more slowly than the tree is condemning it, which is a capacity signal rather than a fault) or `outcome=skipped` (the coordinator's call to the leaf threw - any exception, a request timeout included; on the dirty-set fast path the fault is absorbed once the leaf's dirty mark is retained above the pass watermark and the walk continues, so this arm, recorded once per blocking leaf per pass, is the only signal that a specific leaf is wedged; on the legacy chain walk, or when that mark cannot be retained, the fault fails the batch and the shard's retry-then-skip handling (`orleans.lattice.compaction.shard.retries`, `orleans.lattice.compaction.shard.skipped`) takes over, and a retry that faults again records the leaf again), `path=walk` (legacy chain walk) or `path=dirty-set` (dirty-leaves fast path) when a path is in scope, and - when at least one compaction policy knob (`MinTombstoneRatioForCompaction` or `MaxLeafEntriesBeforeForcedCompaction`) is non-default, on reminder and operator passes as well as policy-triggered ones - `trigger=reminder\|ratio\|size\|operator`. |
| `orleans.lattice.compaction.shard.retries` | `Counter<long>` | `{retry}` | Compaction-coordinator transient retries during a tombstone-compaction pass. Tagged `tree`. A non-zero rate that does not lead to skips is healthy resilience; a sustained rate is worth investigating. |
| `orleans.lattice.compaction.shard.skipped` | `Counter<long>` | `{shard}` | Shards the coordinator gave up on after exhausting retries. **Any non-zero rate is alert-worthy** - the affected shard's tombstones will not be reaped until the next pass. Tagged `tree`. |
| `orleans.lattice.compaction.shard.dirty_leaves` | `Histogram<int>` | `{leaf}` | Per-shard dirty-leaf snapshot size at the moment the compaction coordinator enters a shard. Tagged `tree`. A flat-zero series indicates the dirty-leaves fast path saw nothing to do and the coordinator fell back to the legacy chain walk for that shard. |
| `orleans.lattice.leaf.tombstone.ratio` | `Histogram<double>` | `1` | A leaf's tombstone-to-total-entry ratio, sampled at the entry of every tombstone-compaction pass over that leaf (including passes that reap nothing). Tagged `tree` and `tenant` only: the leaf's identity is deliberately not a tag, so every leaf of a tree records into one series and the family holds **at most one series per tree**, independent of how many leaves the store holds (issue #2518 - a per-leaf tag previously gave it one series per leaf grain). The distribution across a tree's leaves is what the quantiles report; the p95 is the headroom signal for the `MinTombstoneRatioForCompaction` threshold. |
| `orleans.lattice.leaf.splits` | `Counter<long>` | `{split}` | Leaf-node splits triggered by a leaf exceeding its key-count bound (`MaxLeafKeys`) or its byte bound (`MaxLeafBytes`). Tagged `tree`. |
| `orleans.lattice.leaf.split.completion.in_flight` | `ObservableGauge<long>` | `{completion}` | Leaf-split completions currently suspended inside `CompleteSplitAsync`, per `tree` (issue #2967). The division outcome counters (`orleans.lattice.leaf.split_attempts`: `divided`, `faulted`, and the declining arms) are a partition of *terminated* divisions, so a division still executing the completion body at scrape time is a member of none of them - and because the `catch` that records `faulted` is unconditional, a division carrying neither `divided` nor `faulted` did not throw but is genuinely in flight. That un-terminated state had no instrument before this gauge, so a division wedged forever inside the completion read as healthy concurrency and nothing separated `busy` from `stuck`. Covers both the forward path (`SplitAsync`) and the recovery path (`CompleteRecoverySplitUnderGateAsync`), because both reach the completion through the same seam; it is orthogonal to the recovery *outcome* metering of issue #2860. Read it with `orleans.lattice.leaf.split.completion.oldest_age`: a non-zero count is only actionable once the oldest age is climbing. There is no per-tree pre-mint - a tree with nothing in flight emits no series, so an absent series is the healthy steady state rather than a build that lacks the instrument. |
| `orleans.lattice.leaf.split.completion.oldest_age` | `ObservableGauge<double>` | `s` | Age in seconds of the oldest leaf-split completion currently suspended inside `CompleteSplitAsync`, per `tree`, or no series for a tree with none in flight (issue #2967). This is the signal that separates a division that is merely momentarily in flight (age near zero, falling as scrapes advance) from one that is stuck (age climbing without bound), which the bare in-flight count cannot do. It is deliberately a **measurement, not a threshold**: the code encodes no age at which a completion is declared abandoned, because the honest answer to "how long is too long" cannot be chosen without observing real completion durations first, so the gauge supplies the observation and leaves the judgement to an operator or an alert reading real values. A registry-backed observable gauge rather than a completion-duration histogram, and the distinction is load-bearing: a histogram only records on completion, so a division that never completes - the exact wedge this exists to surface - would never appear in one. Its declared unit is `s`, so a host exporting through `.AddPrometheusExporter()` publishes it with a `_seconds` suffix, while the repository-context container's own exposition publishes the bare name; the bundled panel queries the `_seconds` family, matching `orleans.lattice.backup.inventory.oldest_age` (see [How an instrument name becomes a PromQL series name](../../lattice.dashboards/metrics-to-panel-map.md#how-an-instrument-name-becomes-a-promql-series-name)). |
| `orleans.lattice.leaf.bisect_refusals` | `Counter<long>` | `{refusal}` | Leaf splits that could not place their pivot from the snapshot frame and fell back to the ordered key view. Tagged `tree`, `reason` (`no_snapshot_attached`, `too_few_rows`, `frame_key_unreadable`, `no_key_sorts_below_pivot`) and `detach_seam` - the cache surface that had already released the frame: `none`, `clear`, `keys_accessor`, `enumerate_rows_accessor`, `frame_decode_fallback`, `underlying_rows_accessor` or `range_hydration_completed`. What the fallback costs is decided by whether a frame is still attached at the refusal, and `reason` says which (issue #2856). `reason=no_snapshot_attached` is reported exactly when no frame is attached, so every row is already resident and the fallback materialises nothing: with `detach_seam=none` no frame was ever attached (a leaf replayed from the WAL); with any other seam a frame was attached and the named surface had already materialised or discarded it - earlier, typically on the read path - so the cost was paid there, not by the split, and the seam is where to look. `too_few_rows` is cheap (the attached frame holds under two rows). `frame_key_unreadable` and `no_key_sorts_below_pivot` are the expensive refusals: a frame is still attached, so the fallback itself materialises the whole remainder of the leaf and detaches the frame, resident and unsheddable for the life of the activation, on precisely the oversized leaf that can least afford it. A sustained non-zero rate carrying a seam other than `none` is operator-actionable and names the read-path seam that forfeited the frame; any non-zero rate of the two expensive reasons (`frame_key_unreadable` and `no_key_sorts_below_pivot`) is a division paying for the whole leaf. Two `detach_seam` values are never produced by a deployed process, so their absence is expected and carries no information: `underlying_rows_accessor`, whose only producer is reached solely through a deliberately retained test-only accessor, and `range_hydration_completed`, which no surface has assigned since issue #2843 and which survives only so the enum ordinals do not shift. Scoped to those two by name (issue #2865); this is not a claim that every other value is reachable on every deployment. |
| `orleans.lattice.leaf.byte.overflow` | `Counter<long>` | `{leaf}` | Leaves found to exceed `MaxLeafBytes`, the byte bound that complements the key-count bound `MaxLeafKeys`. Tagged `tree` and `outcome`: `split` (the leaf was divided to bring it back under the bound) or `irreducible` (the leaf holds a single entry larger than the bound, so dividing it cannot terminate and it is left intact). A sustained `irreducible` rate is alert-worthy: a leaf whose snapshot cannot be captured pins the whole tree WAL trim floor at zero, because the floor is a minimum over leaves. See [Write-Ahead Log](../wal.md). Pre-minted at zero for **both** outcomes whenever a leaf reaches the byte-bound check on the snapshot-capture path, carrying the exact tag set a real emission carries, so an absent series means the capture seam was never reached rather than that no leaf overflowed (issue #2756). That ambiguity is not hypothetical: the absence of this series on a live deployment was read as evidence the capture-seam pre-split had not landed, when the same absence was equally consistent with the pre-split having landed and found nothing to divide. A counter that is only ever incremented cannot separate "not happening" from "not deployed" from "nothing to do", which makes it unfit as a stop criterion however good its evidence once it fires. Because the series is minted, a flat zero here is a **measured** zero, and the bundled panel omits the `or vector(0)` fallback its neighbours use. |
| `orleans.lattice.leaf.split_attempts` | `Counter<long>` | `{attempt}` | Leaf divisions **sought** on an over-capacity leaf. Tagged `tree` and `outcome`: `divided` (the division ran to completion), `gate_contended` (another turn held the split gate, so this turn returned without evaluating the leaf - no division was attempted and none was refused), `already_under_capacity` (the in-gate re-check found the leaf back under threshold, because a concurrent turn had already divided it), `faulted` (the division began and threw before it could complete, additionally tagged `failure_class`: `unaffordable` when it could not be paid for in memory, `timeout` when it exceeded a deadline, `other` for anything else - the two named classes have opposite remedies, a smaller leaf or bigger hydration budget against examining the storage path under it), `no_admissible_pivot` (the leaf held no row strictly inside its own declared `[LowKeyInclusive, HighKeyExclusive)` range, so there was no key it was entitled to divide at and the division was declined rather than taken at an out-of-span pivot - issue #3117), and `recovered` (a division whose intent was already durable - stranded by an earlier fault, a deactivation, or a crash - was completed later, either by the recovery path before the next write was admitted or by an over-capacity check that resumed the half-finished division instead of starting a fresh one; issue #2860). Exists to make `orleans.lattice.leaf.bisect_refusals` interpretable: a zero refusal count carries two opposite meanings - no division ever forfeited its fast path, or no division was ever attempted - and only the attempt count separates them. That distinction is the whole question on an oversized leaf, because a leaf sitting far above `MaxLeafBytes` with zero refusals **and** zero attempts is not a leaf whose divisions are succeeding, it is a leaf nothing is trying to divide, which is a different defect with a different fix. Pre-minted at zero for **all six** outcomes - and, on `faulted`, for each `failure_class` arm separately, since the class tag is part of the series identity - whenever a leaf reaches the byte-bound check on the snapshot-capture path, through the same recording helper a real emission uses so the tag shape matches by construction, so a flat zero here is a **measured** zero and an absent series means the capture seam was never reached (issue #2756). The bundled panel therefore omits the `or vector(0)` fallback its neighbours use. Sustained `gate_contended` with no `divided` is one pathological shape: divisions are being sought and none completes. Any non-zero `faulted` is the other, but it must be read against `orleans.lattice.leaf.splits` rather than on its own. That counter increments only *after* the split intent is persisted, so `splits - divided` is the number of divisions that stranded a leaf holding a half-finished division for the recovery path to re-enter on the next write, whereas a `faulted` with no matching `splits` increment threw before the intent was persisted and stranded nothing. The subtraction is the reading, not the fault count. A recovery that completes lands on `recovered` - never on `divided`, and never on a second `orleans.lattice.leaf.splits` increment, since that counter already counted the division when its intent was persisted and a grain reactivates for idle collection with the process alive - so `splits - divided - recovered` is the divisions still stranded over the same window. A recovery can complete a division initiated before a restart, so on a freshly restarted process that difference can read negative. Before the `recovered` arm existed a recovered completion recorded nothing here (issue #2860), so in the restart regime, where recoveries are the only split activity, the instrument read as though no division had been sought. The distinction is not academic: the hydration admission gate declines *in advance*, so `failure_class=unaffordable` is both the arm least likely to have wedged anything and the arm an operator is most likely to meet first on an oversized leaf. Before the `faulted` arm existed such a division recorded nothing at all here (issue #2845), so the instrument read identically to a tree on which no division had ever been sought - the exact collapse it exists to prevent. |
| `orleans.lattice.leaf.digest.publishes` | `Counter<long>` | `{publish}` | Leaf-side projection-digest publish decisions, partitioned by path. Tagged `tree` and `path`, where `path` is one of: `coalesced_scheduled` (a fresh one-shot timer was registered - first dirtying mutation inside a new coalescing window), `coalesced_skipped` (a dirtying mutation arrived while a coalesced publish was already pending, so the cross-grain hop was deferred onto the existing window - this is the `publishes saved` surface that justifies the coalescing default), `coalesced_fired` (the coalescing timer tick issued the cross-grain `OnChildDigestPublishedAsync` RPC - one per window per leaf unless an inline publish or graceful flush cancelled the timer first), `inline` (the leaf issued the cross-grain publish synchronously - either because `DigestCoalescingWindowMs` is zero, the timer registration failed in a test harness, or a structural caller routed through `PublishDigestUpwardInlineAsync`), or `deactivation_flush` (the leaf's graceful `OnDeactivateAsync` drained a pending coalesced publish before the activation tore down). The headline coalescing invariant - `N writes inside one window produce one cross-grain hop` - translates to `coalesced_scheduled + coalesced_fired` per window regardless of write count, with `coalesced_skipped` absorbing the remaining `N - 1` dirtying mutations. A regression to the c2-xxix shape (resolver dropping `DigestCoalescingWindowMs` on the floor) shows up as zero `coalesced_*` increments and `inline` growing in lockstep with writes - the propagation-guard regression gate catches the property-level drop, this counter catches the behaviour-level drop. |
| `orleans.lattice.leaf.activation_replays` | `Counter<long>` | `{replay}` | Activation-time leaf-materialiser replays started, counted once per activation that begins a replay. Tagged `tree` and `activation_temperature`: `cold` when the activation had neither a snapshot rehydrate nor a populated entry cache and therefore replays the whole readable WAL window, `warm` when it resumed above one of those anchors and replays only the tail. Both arms are the same counter, so any single scrape yields a coherent cold:warm ratio - two independently scraped counters would not. An activation with no tree id bound takes no replay permit and is counted on neither arm, which excludes both arms identically and so cannot bias the ratio. The same two cumulative totals are also sampled into a rate-limited `Information` log line (once per tree per minute) for hosts that expose no metrics scrape (issue #2148). |
| `orleans.lattice.leaf.activation_replays_over_budget` | `Counter<long>` | `{replay}` | Activation-time leaf replays whose own **post-range-filter applied-entry count** exceeded `MaxLeafReplayEntries` while the WAL still covered the whole needed window. Since issue #2149 the count is the leaf's exact applied-entry total, taken during the replay; it was previously the classifier's partition-wide WAL gap, which spans every leaf pinned to that partition and so overstated a single leaf's work by up to the partition's leaf fan-out (measured at ~1,350 on one deployment). The counter is therefore now in the same units as the budget it is named after. These converge because the replay flushes its checkpoint incrementally, so every attempt makes durable progress even when torn down early, and a sustained rate is a capacity signal rather than a fault. That reading holds only while the persisted checkpoint is advancing: a sustained rate at a frozen checkpoint is a replay livelock (issue #1819) and IS a fault. Tagged `tree` and `partition` (the WAL partition ordinal, bounded by `WalPartitions`). The counter is deliberately not tagged by leaf, because the leaf count is unbounded - so it measures the rate only and cannot by itself separate many leaves each replaying once from one leaf replaying forever. Make that call from the accompanying warning log, which names the leaf and its persisted checkpoint; comparing checkpoints across warnings is only valid when they name the same leaf and partition (issue #2023). That log is a deliberately bounded sample and this counter remains the exact census: repeat warnings back off per leaf and are capped per tree, because one warning per occurrence made this line 46% of one deployment's container log (issue #2100). Withheld repeats are reported as a summary line while that tree keeps replaying; that summary is carried by a later occurrence, so a tree's final withheld tally goes unreported and this counter is the exact census for that case. A leaf partition reporting over budget for the first time is never withheld, so a newly appearing condition still surfaces promptly. A leaf that re-enters replay from an **unchanged** checkpoint is reported separately as a stalled-replay fault, so that shape no longer has to be recovered by hand from the cost warnings (issue #2149). |
| `orleans.lattice.wal.replay.permit_adaptations` | `Counter<long>` | `{permit}` | Adaptations to the per-silo WAL replay concurrency gate under memory pressure (issue #2781). Tagged `outcome`: `withheld` (a permit the replay held was not returned to the gate) or `restored` (a replay completed cleanly with the heap relieved, so one previously withheld permit was returned). **The `withheld` arm is additionally tagged `trigger`** (issue #2883), naming which of the two mechanisms withheld: `occupancy` for the proactive one and `fault` for the reactive one. Before that dimension existed both wrote the same untagged series, so `withheld = N` was a **sum no scrape could attribute** - and that ambiguity produced a wrong published conclusion when acceptance run 12's `withheld = 6` was read as evidence that the reactive trigger fires, a reading equally consistent with it firing zero times. **Summing over `trigger` recovers the historical untagged total**, so comparisons against runs that predate the tag remain valid; that is why the fix was a dimension rather than a second counter. **The `restored` arm deliberately carries no `trigger`, and a per-trigger level is therefore not derivable.** Withheld permits are fungible - the accounting is one process-wide count, not a per-trigger ledger - so a restore cannot know which mechanism withheld the permit it hands back, and `withheld{trigger="fault"} - restored` is **meaningless**. Read the trigger split as *which mechanism is doing the work*, never as *how much each is currently holding*. **Two triggers withhold** (issue #2862): *proactively*, a replay that finishes while managed heap occupancy is at or above 75% of the GC hard limit withholds its permit whether or not it faulted; *reactively*, a replay that fails for memory pressure withholds its permit whatever the occupancy. A permit is restored only when a replay completes cleanly **and** occupancy has receded below 60%, so the two figures form a hysteresis band the gate settles inside rather than a single edge it chatters across. The occupancy denominator is the **GC hard limit** (`GC.GetGCMemoryInfo().TotalAvailableMemoryBytes`, floored by the cgroup memory limit), not the container grant: under a container memory limit .NET applies its default `GCHeapHardLimitPercent` to the grant, so a 12 GiB grant yields a 9 GiB managed ceiling and a threshold written against the grant would sit 33% above the limit that actually throws. Read the **difference** `withheld - restored` as the number of permits currently out of circulation, making the effective ceiling `configured - (withheld - restored)`; read the rates to see how hard the gate is oscillating. The gate admits `WalMaterialiserMaxConcurrentReplays`, defaulting to the lesser of `Environment.ProcessorCount` and the enforced container CPU grant (issue #2816), cold WAL replays at once, and a tree whose entries are large makes every such replay hold large buffers - so on a constrained host the admitted set can exhaust the heap, at which point the replays fail slowly **while holding permits**, other leaves time out queued for a permit and never replay at all, no snapshot is banked, the durable materialiser pin never advances, and WAL garbage collection reclaims nothing, which makes the next replay window larger still. This counter measures the mechanism that breaks that loop. **Deliberately not tagged by tree**: the gate is process-wide, so a per-tree tag would imply a per-tree ceiling that does not exist and would invite a reader to sum arms sharing one resource. Backpressure never withholds the last permit, so the effective ceiling has a floor of one and the mechanism cannot latch. It is also incapable of raising the gate above the configured ceiling for any input, because it works by declining to return a permit already taken rather than by resizing - a dynamic floor beneath the operator's setting is not an override of it (issues #2278/#2279). Both arms are **primed at zero** when the gate is sized - and since issue #2883 that means **all three** primed series, each `trigger` value on the withheld arm as well as the restored arm, because priming the counter as a whole but not each trigger would make "this trigger never fired" and "this build predates the trigger split" identical on a scrape. A flat zero means "measured, never engaged" while an absent series means the build does not carry the mechanism. |
| `orleans.lattice.wal.replay.permits_withheld` | `ObservableGauge<int>` | `{permit}` | Replay permits currently withheld from the per-silo WAL replay concurrency gate by memory backpressure, so the effective ceiling is the configured ceiling minus this figure (issue #2784). Tagged with the platform tenant label and nothing else, matching its sibling counter: the gate is a process-wide static, so every tree on the silo draws from the same pool and a per-`tree` series would repeat one figure under many names and invite a reader to sum them. It exists because the two arms of `orleans.lattice.wal.replay.permit_adaptations` cannot show what a restart destroyed. The level is derivable from them as `withheld - restored`, but both are process-lifetime counters: when a process dies with adaptation in force they reset to zero **together**, so the derived difference silently becomes zero too and nothing in the series marks the discontinuity - "the gate had adapted and the restart threw that away" and "the gate never adapted" render identically. What this gauge adds is therefore the **boundary**, not a new number: the last sample before the gap is the level that was discarded, and the step to zero across the gap is the discard. That reading cannot be produced by the successor process, which has no access to its predecessor's static, so it has to be published continuously by the process holding the state rather than emitted once by the one that inherits none of it. It is **not** a substitute for the counters: a gauge is sampled, so a withhold and its matching restore landing between two scrapes are invisible here and visible there, and the counters remain the record of how often the mechanism fired. This series answers only "how much headroom is being declined right now". Published unconditionally once the gate is sized, so a flat zero is a measured zero - the mechanism is present and has never engaged - while an absent series means the build does not carry it. |
| `orleans.lattice.wal.replay.permit_ceiling` | `ObservableGauge<int>` | `{permit}` | Ceiling the per-silo WAL replay concurrency gate was sized to, or `0` before the first activation sizes it (issue #3047). This is the **denominator** the rest of the permit family was missing. One permit withheld from a ceiling of sixteen is noise; one withheld from a ceiling of two is half this silo's replay throughput; on a scrape those two readings are identical, and until this series existed nothing in the estate could tell them apart. The figure is not recoverable after the fact: a `SemaphoreSlim` does not expose its own maximum, and the gate is sized once and never re-created or topped up, so the only way to publish it is to capture it at sizing time and report it continuously. Tagged with the platform tenant label and nothing else, matching the rest of the family, because the gate is a process-wide static and a per-`tree` series would repeat one figure under many names. A `0` here is an unambiguous **not yet sized** sentinel rather than a degenerate reading: the sizing resolver returns the configured value when it is positive and the derived default otherwise, and both are at least one, so a sized gate cannot publish zero. |
| `orleans.lattice.wal.replay.permits_available` | `ObservableGauge<int>` | `{permit}` | Permits currently available on the per-silo WAL replay concurrency gate - the headroom a new activation arriving now would find (issue #3047). **Never read alone.** A zero here is ambiguous between a gate that has not been sized and a gate that is fully saturated, and only `orleans.lattice.wal.replay.permit_ceiling` separates the two; saturation is the reading this instrument exists for, which is why the pair landed together and is documented as a pair rather than as two independent series. It also does **not** measure the queue: the underlying count saturates at zero and reports the same figure whether one activation or a thousand are blocked behind it, so a saturated gate looks the same here however deep the backlog is. `orleans.lattice.wal.replay.permits_queued` is the instrument for that. Tagged with the platform tenant label only, as the gate is process-wide. |
| `orleans.lattice.wal.replay.permits_queued` | `ObservableGauge<int>` | `{activation}` | Activations that have entered the wait for a permit on the per-silo WAL replay concurrency gate and not yet left it (issue #3047). This counts **arrivals in flight, not activations blocked** (issue #3290): the count is incremented before the wait, so an activation that acquires without ever blocking is included for its acquire window, and the blocked set is a subset of it, exceeded by at most `orleans.lattice.wal.replay.permit_ceiling`. This is the only series in the estate that can observe a gate **admitting nothing**, and the reason it is worth a third instrument is structural rather than a matter of resolution. Every other permit series records at a *terminal* outcome - `orleans.lattice.wal.replay.permit_queue_wait` records on acquisition and on cancellation, and the activation-replay counters record once a permit is held - so all of them sit downstream of the wait. An activation that never acquires and is never cancelled therefore records **nothing anywhere**: not its wait, not its replay start. A permanently saturated gate consequently renders byte-identically to a gate no tree has ever asked for a permit from, which is precisely the ambiguity that left 47% of a deployed estate's stranded-leaf population uninterpretable. Read the three together: the ceiling says how wide the door is, availability says whether it is open, and this says how many have arrived at it and not yet got through. Tagged with the platform tenant label only, as the gate is process-wide. |
| `orleans.lattice.wal.replay.permit_queue_wait` | `Histogram<double>` | `ms` | Time an activation spent queued on the per-silo WAL replay concurrency gate, recorded once per wait as it ends (issue #2873). Tagged `tree`, `tenant`, and `outcome`: `acquired` (the wait ended in a permit) or `canceled` (the activation was cancelled while still queued). It exists to **discriminate** a reading that `orleans.lattice.leaf.activation.failures{reason="canceled_awaiting_permit"}` cannot make on its own: that counter is honest about *where* an activation was cancelled and silent about *why*. The Orleans request deadline spans the whole grain call, so an activation that burned its budget upstream arrives here already doomed and is cancelled within seconds, and the count is identical whether the gate was saturated or idle. Only the duration separates them, and it separates them by orders of magnitude: **genuine permit starvation** shows a `canceled` arm whose wait approaches the request budget, whereas **upstream budget exhaustion** shows a `canceled` arm at or near zero, because the activation was already out of time when it arrived. Read the two arms together - a large `acquired` wait alongside a large `canceled` wait is a saturated gate serving a queue, while a near-zero `canceled` wait says the gate is not the constraint no matter how high that counter climbs. **A lifetime `_sum / _count` mean is NOT a present-state discriminator, and reading it as one has already produced a misdiagnosis on this epic.** The sum and count are cumulative over the process lifetime, so a pool that saturated once and has long since drained still reports the mean of that episode indefinitely - a twelve-second mean read as a current wait against a pool that was empty at the time of reading. Use this histogram to characterise the **shape** of waits that have completed; to answer whether the gate is saturated **now**, read `orleans.lattice.wal.replay.permits_queued`, `orleans.lattice.wal.replay.permit_waits_in_flight`, and `orleans.lattice.wal.replay.permit_wait.oldest_age`, which are levels and carry no history. Never conclude health or absence from a single sample of any of them. **Read it as `_sum / _count`, not as a quantile.** The repository-context container's own exposition renders a `Histogram<T>` as a Prometheus summary carrying `_sum` and `_count` and **no** `_bucket`, so `histogram_quantile` over this series returns nothing; the usual `or vector(0)` repair would substitute a literal zero that reads exactly like a genuine near-zero wait, which is one of the two conclusions this instrument exists to distinguish, so it must not be used here. The mean is sufficient for that discrimination because the two regimes differ by orders of magnitude, not by percentile shape. **Since issue #3044 this histogram declares explicit `InstrumentAdvice` bucket boundaries, and that does NOT retire the paragraph above.** `InstrumentAdvice` is a .NET-level hint that only an exporter which reads it can honour, and the exposition this container runs honours none: a full scrape of it contains **zero** `_bucket` lines, types **all 98** of its histogram families as `summary`, and carries no `quantile=` labels and none of the OpenTelemetry exporter fingerprints (`target_info`, `otel_scope_name`). The advice is therefore **necessary and not sufficient** - it makes buckets available to an exporter that honours them, such as the OpenTelemetry `AddPrometheusExporter` used by `samples/McpTelemetry`, and changes nothing under this adapter. **Keep reading the mean until a scrape actually shows `_bucket` series for this family.** Switching to `histogram_quantile` on the strength of boundaries existing in source returns nothing and invites exactly the `or vector(0)` repair the paragraph above forbids, so the boundary is part of the reading rather than a footnote to it. **Tagged by `tree`, unlike its sibling `orleans.lattice.wal.replay.permit_adaptations`**, and the difference is not an inconsistency: that counter measures a process-wide *ceiling*, which has no per-tree value, whereas a wait is one caller's own experience and is attributable to the tree that waited. Scoping a reading to a single tree is also what the runs 10/11 scoring defect could not do. **Deliberately not primed at zero**, unlike its sibling: priming a histogram means fabricating a `0 ms` sample, which is not a neutral "measured, never engaged" marker but a datum reading as "the gate was instant" - it would bias the instrument toward one of the two conclusions it exists to distinguish. An absent series is therefore **UNINTERPRETABLE rather than innocent** and must be diagnosed, not read as a pass: corroborate against `orleans.lattice.wal.replay.permit_adaptations` (which *is* primed) and the activation counters to tell "no build" from "no waits". |
| `orleans.lattice.wal.replay.permit_hold` | `Histogram<double>` | `ms` | Time a replay **held** a permit on the per-silo WAL replay concurrency gate, from acquisition to the end of the replay, recorded once per hold for an activation replay and for a WAL GC starvation drive alike (issue #3921). Tagged `tree` and `tenant`. This is the measurement that separates the gate's two regimes, which have opposite remedies: when replays are CPU-bound, raising `WalMaterialiserMaxConcurrentReplays` finishes more replays per second and hold time stays flat; when they are bound on a grain store that serialises its writes, the service rate on `orleans.lattice.wal.replay.permits_served` stays flat and hold time grows, and the admission gate's `no_progress` arm fires harder. Declares explicit bucket boundaries from 1 ms to 30 minutes. Not zero-primed, for the reason given for `orleans.lattice.wal.replay.permit_queue_wait`: a fabricated 0 ms hold would bias the distribution toward "the store is fast". |
| `orleans.lattice.wal.replay.permits_served` | `Counter<long>` | `{permit}` | Permit holds that ended on the per-silo WAL replay concurrency gate, whether the permit was then returned or withheld by memory backpressure (issue #3921). Its rate is the gate's **service rate** in replays per second, and by Little's law the mean in-flight replay count is that rate multiplied by the mean of `orleans.lattice.wal.replay.permit_hold`. Tagged with the platform tenant label only, since the gate is process-wide. **Zero-primed when the gate is sized**, so a sized gate whose service count has stopped moving while activations are queued reads as a measured flat line, which is the regime the `no_progress` refusal arm reports. |
| `orleans.lattice.wal.replay.permit_waits_in_flight` | `ObservableGauge<long>` | `{activation}` | Activations **currently queued** on the per-silo WAL replay concurrency gate (issue #3044). Tagged `tree` and the `tenant` derived from it. It exists because `orleans.lattice.wal.replay.permit_queue_wait` is **structurally blind to a wait that never ends**: both of that histogram's arms are terminal - `acquired` is recorded after the semaphore is entered, `canceled` from the catch around the wait - so an activation parked on a saturated gate and never admitted records on **neither**, and the instrument falls silent in exactly the condition it was built to measure. **No amount of zero-priming reaches that gap**, which is the part worth keeping: a primed `canceled` arm reading zero asserts only that no cancellation *completed*, and the missing observable here is a **level**, not a terminal event. That generalises - priming is free for an instrument whose empty state is a level, and wrong for one whose empty state is a shape, which is why this gauge is primed and its sibling histogram deliberately is not. **Primed at zero per tree once that tree has queued at least once in this process**, so a flat zero here is a *measured* zero rather than an absent build. **That priming has a boundary which is part of the reading and is not severable from it**: it covers trees this process has *observed* queueing, not every tree in the estate, so a zero means "has queued before, is not queued now" and never "cannot queue" - and a tree that has never queued has no series at all, an absence which is correct and carries no information about that tree. **Tagged by `tree`, like `permit_queue_wait` and unlike `permit_adaptations`**, for the same reason: a ceiling is process-wide and has no per-tree value, whereas a wait is one caller's own experience and is attributable to the tree that waited. **Useless on its own** - a count of three cannot separate three healthy 200 ms waits from one activation parked for twelve minutes - so read it with `orleans.lattice.wal.replay.permit_wait.oldest_age`, which supplies the half it cannot. |
| `orleans.lattice.wal.replay.permit_wait.oldest_age` | `ObservableGauge<double>` | `s` | Age of the **oldest** activation currently queued on the per-silo WAL replay concurrency gate (issue #3044). Tagged `tree` and the `tenant` derived from it. This is the half of the pair that separates **busy** from **stuck**: a steady `permit_waits_in_flight` whose oldest age oscillates near zero is a gate serving its queue promptly, while the same count with this age climbing without bound is an activation that will never be admitted. That second shape closes the loop described under `permit_adaptations` - a leaf that cannot obtain a permit cannot replay, cannot bank a snapshot, cannot advance the durable materialiser pin, so WAL garbage collection finds the pin unusable and reclaims nothing, which makes the next replay window larger still. **No threshold is encoded here deliberately**; "stuck" is read off the slope against real data, because an honest threshold differs per host and per WAL size, and a fixed one would be a number that looked authoritative while meaning nothing. **Deliberately NOT primed at zero, unlike its sibling count, and the asymmetry is the design rather than an oversight.** The age of the oldest waiter when there is no waiter is **undefined**, not zero: a fabricated zero would publish the healthiest possible value for the emptiest possible state, making an idle gate indistinguishable from one whose waits are all instant - which is precisely one of the two conclusions this pair exists to let a reader choose between. An **absent series here is therefore correct** and means nothing is currently queued; that is safe to read only because the *primed* count beside it does separate a measured zero from a build that does not carry the mechanism, so the pair is interpretable where neither member is alone. |
| `orleans.lattice.wal.replay.slice_narrowings` | `Counter<long>` | `{narrowing}` | Activation-time replay slice-width narrowings forced by memory pressure on a commit-log read (issue #2867). Tagged `tree`, `partition`, `tenant`. **This is the second of the two factors that set peak replay memory.** Peak memory is the *product* of how many replays run at once and how much each one buffers. The first factor is `WalMaterialiserMaxConcurrentReplays` - an option, surfaced on the container tuning overlay as `REPOCONTEXT_MAX_CONCURRENT_REPLAYS`, and already measured from both sides by `orleans.lattice.wal.replay.permit_adaptations` and `orleans.lattice.wal.replay.permit_queue_wait`. The second is the per-replay slice width: `WalReplaySliceBudget` (issue #2898) sets the width a replay starts from, and the only thing that moves it below that at run time is the reactive narrowing this counter records - a read refused for memory pressure is retried at a quarter of the current width. Without this counter that run-time half of the factor is invisible, because the narrowing is otherwise reported by a warning log alone and a log is not a series. **What it discriminates:** a managed `OutOfMemoryException` during a mass reactivation is consistent with two stories whose remedies point in opposite directions - either the narrowing engaged and was not enough, so the width is genuinely too coarse for the host, or it never engaged at all because the allocation that failed was not the slice read the clause guards, in which case lowering the width would change nothing. Read it against `orleans.lattice.leaf.activation.failures` on the same tree: failures climbing while this stays at zero is the second story. **Tagged by tree deliberately**, unlike the process-wide `permit_adaptations`: a width is a per-replay property and a replay belongs to the tree whose leaf is activating, so this genuinely has a per-tree value - and that tag is what lets one tree be shown buffering harder than its siblings under identical cycling. **Zero-primed per (tree, partition)** when a partition replay begins, because the healthy steady state is never to narrow, so an unprimed counter would leave the common case indistinguishable from a build with no narrowing at all. |
| `orleans.lattice.wal.replay.starvation_drive_abandonments` | `Counter<long>` | (none) | WAL GC starvation drives abandoned after exhausting their budget, recorded on the grain side as the drive gives up (issue #3065). Tagged `tree` and the universal derived `tenant` dimension, which the leaf grain derives from the same tree id the scheduler does, so on a tree that has not been resized, restored or remediated the zero the scheduler primes and the one the grain records share a single series identity; on such an aliased tree the scheduler primes under the physical copy's id while the grain records under the logical id (see [The `tree` dimension across aliasing](tag-conventions.md#the-tree-dimension-across-aliasing)). **This is the sole discriminator between a bounded-but-slow drive and a drive that would have parked forever**, and that is the whole reason it exists. Bounding the drive necessarily destroys the diagnostic that found the defect: before issue #3065 a parked drive held one of two per-silo replay permits indefinitely, and the scheduler's own arms rendered that as `attempted` climbing, `undelivered` climbing just behind it (the touch times out at the Orleans response-timeout default long before the drive's own budget elapses), and `drove_already_driving` collecting the successors that bounced off the in-flight latch. The measured signature on two independently frozen containers was `attempted` 81 = `undelivered` 78 + `completed` 3, with all three completions returning `AlreadyDriving`. A drive that is merely *slow* now produces that same shape, byte for byte, because the touch still times out at the same deadline and successors still bounce off the same latch - so **none of the scheduler's arms can separate the two any more** and only this counter can. An operator reading a frozen-looking reactivation panel must read this series before concluding anything: advancing means drives are hitting the budget, and the paired warning log says why: it reports how the budget divided between admission to the replay gate and the replay itself, and how far the replay moved the leaf's checkpoint, so a slow host-supplied commit-log read or an oversized replay gap is visible as such (issue #3479), whereas flat-at-zero means the drives are completing and the shape is ordinary sweep contention. **Primed at zero beside `attempted` on the scheduler, not at the grain's drive entry point.** Priming at the drive would make an absent series mean "no drive has ever run" as well as "this build is not deployed", which is exactly the ambiguity the priming exists to remove; co-primed with `attempted` the reading is unambiguous, because `attempted` greater than zero with this series present is a real measured zero, while `attempted` greater than zero with this series absent proves the build carrying the fix never landed. That makes the prime a free deployment proof as well as a zero-disambiguator. **One arm covers both abandonment shapes** - a drive that timed out queued for a permit, and one that timed out deep inside the replay - because they are the same verdict on the sweep; they call for different operator responses, so they are separated on the paired warning log, which splits the drive's time into admission and replay (issue #3479), and the first also records on `orleans.lattice.wal.replay.permit_queue_wait`'s `canceled` arm while the second does not. The drive's elapsed time is carried on the paired warning log rather than as a histogram, because an abandoned drive's duration is approximately the budget by construction and would carry almost no information as a series. |
| `orleans.lattice.leaf.activation_stalled_replays` | `Counter<long>` | `{replay}` | Activation-time leaf replays that re-entered from a persisted checkpoint that had **not advanced** since the same leaf partition's previous replay on this silo, meaning the previous activation banked no durable forward progress at all (issue #2285). This is the **fault** arm and is a different condition from `orleans.lattice.leaf.activation_replays_over_budget`, which is a capacity signal: an over-budget replay converges, just slowly, whereas a non-advancing one has not converged at all. Before this counter the fault was reported only by a warning log throttled to one line per (tree, leaf, partition) per minute, so the rate an operator could observe was the throttle's and not the condition's; the counter is incremented outside that throttle and is the exact census, with the log a bounded sample of it. Read it as a rate over time rather than as a level: the condition is transient whenever an activation is torn down mid-replay, so a burst of cancellations or timeouts produces a cluster and then stops, and it is persistent only when the same leaf partition keeps reporting across many minutes. The two are indistinguishable in a single sample and were conflated in issue #2285, where a 70-second burst of 47 occurrences was read as a permanent convergence defect. Alert on the condition continuing, not on its appearance. Tagged `tree` and `partition`; deliberately not tagged by leaf, because leaf count is unbounded - the accompanying warning names the leaf. |
| `orleans.lattice.leaf.activation_cursor_publish_failures` | `Counter<long>` | `{failure}` | Activation-time eager cursor-publish failures, emitted when the post-replay cursor publish fails during a leaf activation. Tagged `tree`. |
| `orleans.lattice.leaf.deactivation.checkpoint_delta` | `Histogram<long>` | `{offset}` | Projection-checkpoint offsets a leaf activation banked **durably** while gracefully deactivating: the sum across WAL partitions of the persisted checkpoint on leaving `OnDeactivateAsync` minus the same sum on entering it (issue #2280). Tagged `tree`, `deactivation_reason` and `activation_temperature`; deliberately **not** tagged by leaf, because the live-leaf population is unbounded - per-leaf detail goes to the accompanying `Debug` line, which also names the entries applied through the projection seam. It measures the **persisted** mark on purpose: the in-memory accessor returns `max(persisted, pending)`, so an offset advanced but not yet flushed is already present on entry and differencing it would report zero for every deactivation at every rate of occurrence. **This is a LOWER BOUND on deactivations and must not be quoted as a census.** Crash, process kill and silo failure bypass the hook by design, and an activation that THREW never reaches it at all - Orleans does not run `OnDeactivateAsync` when `OnActivateAsync` throws, measured with a positive control on Orleans 10.2.2 - so that population is carried by `orleans.lattice.leaf.activation.failures` instead and the two must be read together. A **zero on the `cold` arm is expected and is not by itself a fault**: a cold replay restarts from the `-1` sentinel and the checkpoint is strictly monotonic, so nothing can be banked until the scanned-through offset passes the mark the activation started above. Cold progress below that mark is not discarded, it is unrepresentable, which makes an unchanged cold checkpoint arithmetically forced rather than symptomatic. The offsets are **scanned-through, not applied-through** (issue #2270): replay advances the checkpoint over entries it skips as another leaf's work, deliberately, because the WAL retention floor is the minimum of these offsets. A non-zero delta is durable forward progress through the log - the quantity issue #2280 is about - and is **not** a count of mutations this leaf applied; the accompanying log line carries that separately as `EntriesApplied`, counted as each entry is applied to the leaf's projection. |
| `orleans.lattice.leaf.deactivation.barrier.failures` | `Counter<long>` | `{failure}` | Graceful-deactivation durability barriers that **faulted**, tagged `tree` and `reason`: `digest_publish`, `checkpoint_flush`, `snapshot_capture`, `frontier_pin` (issue #3366). Since issue #3393 the `snapshot_capture` and `frontier_pin` arms also count a barrier that **skipped** because the deactivation deadline had already torn the activation down: a barrier that could not do its work is counted under its own reason whether it threw or declined, and only the paired warning log tells the two apart. Each barrier is contained independently, so a fault in one no longer cancels the barriers after it - previously the four ran as one unguarded sequence and the first fault silently skipped the rest, which is why a digest publish failing could cost a checkpoint flush. Since issue #3393 they run as `checkpoint_flush`, `snapshot_capture`, `frontier_pin`, then `digest_publish` last, because the digest publish is the slow, staleness-tolerant step and running it first spent the deadline the durable barriers needed; when the final checkpoint flush deferred its own inline digest publish, that publish takes the last slot instead and a failure of it is counted as `inline_digest_publish` on `orleans.lattice.leaf.checkpoint.flush.tail.failures`, not here. **The arms are not equally serious and must not be alerted on as one series.** `digest_publish` is benign in isolation: the digest is staleness-tolerant and the next mutation republishes it. `checkpoint_flush` above zero is a **durability fault** - the activation's checkpoint progress was not banked, so the next activation replays from an older offset and any progress the WAL no longer carries is gone. `snapshot_capture` and `frontier_pin` sit between the two: neither loses data on its own, but both leave the WAL retention floor held below where it should be. Read it against `orleans.lattice.leaf.deactivation.checkpoint_delta`, which shares the same population and the same blind spot. **A zero is NOT evidence of health, and on some trees it is structurally unreachable.** This counter is only written from inside the deactivation hook, so a tree whose leaves never deactivate - a hot tree read continuously, with no residency sheds - emits **no series at all** while being exactly the population most at risk, because nothing else captures its snapshots either. Establish that a tree runs the deactivation path (a live `orleans.lattice.leaf.deactivation.checkpoint_delta` series) **before** reading this counter's silence as anything. Crash, process kill and silo failure bypass the hook by design, so like `checkpoint_delta` this is a **lower bound**. |
| `orleans.lattice.leaf.deactivation.barrier.duration` | `Histogram<double>` | `ms` | Wall-clock duration of each graceful-deactivation barrier, tagged `tree` and `reason` with the same four barrier values as `orleans.lattice.leaf.deactivation.barrier.failures` (issue #3628). The failure counter says which barrier did not complete; this says where a drain spends its time, so a rise in per-activation drain cost can be attributed to one barrier rather than guessed at. One sample per barrier that ran, **whatever its outcome**: a barrier that faulted or skipped after consuming the deadline spent that time all the same. `digest_publish` also covers the teardown persist's deferred inline digest publish, which is the same upward publish. Read it as an interval mean per `reason`, delta `_sum` over delta `_count`. The per-barrier means summed approximate one activation's serial teardown cost; Orleans deactivates concurrently, so a drain's wall-clock is less than that sum multiplied by the resident count. Crash, process kill and silo failure bypass the hook, so like its siblings this covers graceful deactivations only. |
| `orleans.lattice.leaf.deactivation.barrier.elided` | `Counter<long>` | `{barrier}` | Graceful-deactivation barriers **skipped** because the pin store had already acknowledged, in the same deactivation, a pin dominating everything the barrier would publish, tagged `tree` and `reason` (today only `frontier_pin`) (issue #3643). The teardown persist tail publishes the frontier pin and the `frontier_pin` barrier then resolves the same batch again; when every partition of that batch is dominated on both axes (frontier HLC and checkpoint offset) by what the pin store acknowledged for the tail, the redundant pin-store round trip is skipped. Acknowledged means every shard write for the batch landed: a faulted or shed shard write does not count, and a deactivation whose tail did not run always publishes. The `frontier_pin` series of `orleans.lattice.leaf.deactivation.barrier.duration` is still recorded for an elided barrier, so the ratio of this rate to that histogram's count rate is the fraction of drains whose pin publish was redundant. Not zero-primed, matching its sibling barrier instruments: absent means no barrier has elided yet, not that the instrument is missing. |
| `orleans.lattice.leaf.checkpoint.flush.tail.failures` | `Counter<long>` | `{failure}` | Post-flush notification steps that **threw after the checkpoint flush itself succeeded**, tagged `tree` and `reason`: `cursor_report`, `inline_digest_publish`, `snapshot_recheck` (issue #3393). Read it beside `orleans.lattice.leaf.deactivation.barrier.failures`, not instead of it, because the two measure different halves of one named thing: that counter covers the flush **body**, the durable write, and stands down once the write lands, while the tail runs its post-write notifications after it, each contained separately. That containment is deliberate and correct - under issue #2220 a failed notification tearing down the activation produced a replay loop - and this counter does not change it; what it changes is that the fault is countable. **The consequence for reading the barrier counter is the whole point: a `checkpoint_flush` count of zero there does NOT establish that the checkpoint path is healthy.** On the drain captured under issue #3393 the barrier read zero for `checkpoint_flush` while 1,699 dirty digests died inside this tail, so barrier counts are a **floor** on the damage rather than a total. **The arms are not interchangeable and must not be alerted on as one series.** `inline_digest_publish` is the arm seen in production and means a leaf's parent was never told its digest changed, so the upward digest chain is stale until the next mutation re-drives it - on a quiescing tree, possibly never. `snapshot_recheck` means the post-flush re-evaluation of snapshot need did not run (on the ordinary tail the arm also covers the coverage-lag timer arming that precedes it), so a block pin can retain the tree's shared WAL, making it a candidate contributor to unbounded WAL growth (issue #3094). `cursor_report` carries a specific warning: each per-partition reporter call is **also** wrapped in its own `try`/`catch` inside the cursor report, so an ordinary report failure is logged there and never reaches this counter; only what surrounds those calls can - the grain-state and option reads, consumer-id construction, and the durable materialiser pin publish that follows the report. Since issue #3599 the same arm also counts a failed pin republish after a snapshot-recheck capture the store kept and, on a graceful deactivation's final flush, a failed awaited pin publish before the recheck or after a kept capture. A zero on `cursor_report` is therefore **not** evidence that cursor reporting is healthy, because the per-partition reports are double-contained and their ordinary failures are invisible here by construction. **Not zero-primed**, matching its barrier sibling: a series appears on the first fault, so an absent line does **not** establish that the build lacks the instrument. Confirm deployment from the image digest, never from the presence of this series. |
| `orleans.lattice.leaf.activation.failures` | `Counter<long>` | `{activation}` | Leaf activations that ended by **throwing** out of `OnActivateAsync` rather than coming online (issue #2280). Tagged `tree`, `activation_temperature` and `reason`: `canceled` when a replay already in progress was cancelled, `canceled_awaiting_permit` when the cancellation arrived while the activation was still queued for the replay concurrency permit and no replay had begun, `canceled_resolving_options` when it arrived while the activation was still resolving its tree options and had not yet requested a permit, `canceled_rehydrating_snapshot` when it arrived during the snapshot rehydrate, `refused_replay_admission` when the activation was turned away before it ever joined the permit queue because the admitted-waiter bound of issue #3284 was already reached, and `faulted` for any other failure. **`refused_replay_admission` and `canceled_awaiting_permit` are a pair and must be read as one.** The second means an activation was admitted, waited, and lost; the first means it was never admitted, because the queue was already deeper than it could outlast. Before admission control existed the second population did not, so the whole backlog surfaced under `canceled_awaiting_permit` and a queue nobody was bounding was indistinguishable from a gate that was merely busy. A rise in `refused_replay_admission` is **the bound working**, not a fault: it is the cheap failure, since a refusal is immediate and the caller retries, whereas an admitted-but-doomed waiter holds an activation for its whole request deadline and then enqueues its replacement. The cancellation arms are kept apart deliberately, and issue #2770 added the last two after the original pair proved too coarse to act on: `canceled_awaiting_permit` was selected by `replayPermit is null`, which is true for a cancellation in the tree-options resolve and in the permit queue alike, so one arm was unreachable and a cold-start stall on the shared registry singleton was reported for six deployments as replay-gate saturation (issue #2768). Read the arms as distinct causes with distinct remedies: a rise in `canceled_resolving_options` means activations are serialised behind a shared dependency and the replay gate may be entirely idle; a rise in `canceled_awaiting_permit` means the gate itself is saturated (issues #2279, #2256); a rise in `canceled` means a single replay is too long to finish inside the deadline (issue #2411). Do not infer gate saturation from a cancellation that never reached the gate. This counter exists because `orleans.lattice.leaf.deactivation.checkpoint_delta` is structurally blind to this population: a failed activation never runs the deactivation hook, so without it a cancelled cold replay would read as zero at every rate of occurrence including the highest, an absence indistinguishable from health. The same blindness hid how far a failed replay got, so since issue #2411 every counted failure is also paired with an unthrottled `Information` line naming the WAL entries the activation had applied through the projection seam, with its temperature, failure arm and admission phase; read those lines as a distribution filtered to cold activations that entered replay, because one cancelled while queued for a permit or resolving tree options reports zero by construction. **Still a lower bound**, though a tighter one than it was: the permit acquisition and, since issue #2770, the snapshot rehydrate both sit inside the guarded region, so the queue wait and the rehydrate are now covered where a cancellation during the rehydrate previously escaped uncounted entirely - it did not even appear in the split as a zero. A failure raised before that region is entered, during the coherence reset, is still not counted, and an outright process kill reaches no observation site anywhere. |
| `orleans.lattice.leaf.activation.cold_replay_loop` | `Counter<long>` | `{activation}` | Cold leaf-activation cancellations at or past the escalation threshold (3) of a **consecutive** cold-cancellation streak on this silo, meaning the same leaf's cold activation has been cancelled at least three times in a row with no successful activation in between (issue #2280). The streak counts a cold cancellation wherever it landed - mid-replay, queued for a replay permit, or still resolving tree options or rehydrating its snapshot - and the paired warning splits the run by those phases; every cancellation at or past the threshold counts, so a streak that runs to five contributes three increments. Tagged `tree` only - the leaf's identity is on the paired warning, never a tag. It exists because every other signal for this condition is an **aggregate over leaves**: `orleans.lattice.leaf.activation.failures` and the cold-replay log line both sum across the leaf population, and a sum cannot separate one leaf cancelled five times, which is a self-reinforcing loop and a defect, from five leaves cancelled once each, which is a cost. Because the streak is tracked per leaf, this counter can. The streak is **consecutive and resets on the leaf's first successful activation**, warm or cold, which is load-bearing and not an implementation detail: a cumulative count against a fixed threshold is a function of process age rather than leaf health, so at the measured rate it would eventually fire on a healthy silo no matter how high the threshold was set. Calibrated from issue #2278's corrected field measurement - 79 runtime cancellations over roughly 40 minutes, all `bplusleaf`, across 67 distinct leaves, of which 55 were cancelled once, 12 twice, and none more than twice - which reads as leaves occasionally repeating rather than leaves trapped. A threshold of 2 would fire on 12 of those 67 leaves in a normal window and would be pure noise, and a warning that always fires gets muted, taking the real signal with it; 3 is one above the observed maximum and is also robust to the log never having recorded whether the 12 twice-cancelled leaves succeeded in between, since neither reading of that ambiguity reaches 3. Do **not** recalibrate from the figures originally published on issue #2278 (27 cancelled more than once, one four times): they were impossible on their own arithmetic, because 79 occurrences over 67 distinct leaves leaves a surplus of only 12. A paired structured `Warning` line names the pathology and the leaf and is throttled to one per leaf per minute; this counter is incremented outside that throttle and is the exact census, with the log a bounded sample of it. A cancellation that lands mid-replay banks the prefix it had already re-read as snapshot coverage, so successive attempts are expected to advance; the warning therefore names the three shapes that still reach the threshold - a replay cut before its first slice boundary, a banked claim refused because it would regress a partition's existing coverage, or too little progress per attempt to converge - and notes that a cancellation queued for a permit or still resolving options or rehydrating never entered replay and banks nothing. Observation only - nothing here bounds the readable replay window, checkpoints incrementally, or relaxes the snapshot-capture gate, which together are the bounding design tracked as issue #2411. |
| `orleans.lattice.leaf.replay_barrier_outcomes` | `Counter<long>` | `{replay}` | Terminal states of a leaf's **deferred** WAL replay (issue #2871), tagged `tree` and `outcome`: `completed` when the replay applied fully and the barrier is satisfied for the rest of the activation, `faulted` when it threw, and `canceled` when it was cut short by deactivation or by an operation that discards the state it was rebuilding (`ClearGrainStateAsync`, `RebuildProjectionFromWalAsync`). This series exists because #2871 moved the replay OFF the activation critical path, and that trade replaces a loud failure with a quiet one: before, a failed replay destroyed the activation and was counted on `orleans.lattice.leaf.activation.failures`; now the activation comes online and stays healthy-looking while its data operations fail, so **a fault here does not appear on that counter and the two must be read together rather than one substituted for the other**. All three outcome arms share one instrument and each is **zero-primed per tree when a barrier is armed**, so a flat zero on `faulted` is a measured zero and not an unpublished series - the distinction the priming exists to preserve, because `no leaf on this tree ever armed a replay` and `replays were armed and none failed` are different facts that an absent series cannot tell apart. `canceled` is expected in bounded quantity and is not on its own a defect; a sustained non-zero `faulted` is, and its symptom is data-operation failure on a live leaf. A replay that faults or is cancelled **disarms itself**, so the next data operation or WAL GC touch re-arms a fresh one: the count is of replay attempts, not of leaves, and one leaf may contribute several. |
| `orleans.lattice.leaf.deferred_terminals_dropped_at_cap` | `Counter<long>` | `{terminal}` | **Durable ledger records refused, not terminals lost** (issue #3190). Pass 1 retains each `TxCommit`, `TxAbort` or `DeleteRange` in `deferredTerminals` before attempting the capped `UnresolvedReplayWork` ledger write. A refusal falls back to the in-memory checkpoint clamp; pass 2 still applies the terminal and drains its ledger record if one was admitted. Interruption before that drain can require re-reading unbanked work. Also counts refusals when the ledger is disabled. The head-of-window liveness admission is unchanged. Tagged `tree`, `partition` and `tenant`, not leaf; a cumulative count or a sustained aggregate rate cannot establish that one leaf stays full, fails to drain, or holds the WAL floor. The #3190 vector-index observations are consistent with expected saturation in a range-delete-heavy replay: `LatticeVectorIndexStore.DeletePrefixAsync` normally issues `DeleteRangeAsync`, and non-last partitions defer those deletes even without sagas. Regression coverage at cap 1 and the shipping cap 1024 shows saturation followed by a fully drained ledger when pass 2 completes; the aggregate field counts do not establish a separate drain defect. Raising the cap may reduce re-reads after interruption but increases persisted row size (notably constrained on Azure Table); any finite cap merely moves the threshold for a long enough WAL window. No default increase is justified by these counts alone. The uncapped prepare-recording arm (`leaf.unresolved_prepare_ledger_beyond_cap`) is not a proxy for terminal-record refusals. Zero-primed on partition replay entry. The shipped name, unit, tags, emission sites and replay behaviour are unchanged. |
| `orleans.lattice.leaf.span_fail_open_commits` | `Counter<long>` | `{key}` | Keys a leaf committed **locally** although its declared `[LowKeyInclusive, HighKeyExclusive)` span excludes them, because declared-span admission found no neighbouring leaf to forward them to (issue [#2125](https://github.com/NSTA1/Orleans.Lattice/issues/2125)). Tagged `tree`, `reason` (`no_sibling`: the chain pointer on the key's side is null; `self_reference`: it names the leaf itself) and `origin` (`client_write`, `merge`, or `cross_shard_migration`), so a deliberate migration graft reads apart from an accidental fall-back. It matters because WAL replay admits by declared span: the leaf holding the row will not reinstate it on its own rebuild, so its survival across a cold restart depends on the checkpoint position of the leaf that does declare it. A leaf that declares no span at all (a single-leaf tree or a bulk-loaded leaf) owns every key and never advances it. Each fail-open also logs a warning naming the leaf and tree, rate-limited to one line per silo every 10 seconds. Not pre-minted, so an absent series is the healthy reading and any value is worth investigating. Observability only: the fail-open behaviour is unchanged. See [Span Admission](../tree-structure.md#span-admission). |
| `orleans.lattice.leaf.snapshot.captures` | `Counter<long>` | `{capture}` | Leaf-snapshot capture **attempts**, counted once each at the point the leaf's snapshot-capture core commits to a capture (issue #2696). Tagged `tree` and `outcome`: `succeeded` once the blob has landed and durable coverage has advanced, `abandoned` when the caller's token was cancelled, `failed` for anything else. It exists because the capture path was **wholly uninstrumented**: the advisory capture wrapper catches every exception and only logs, and the sole capture-derived series, `orleans.lattice.storage.snapshot_bytes`, is fed only after a *successful* save. So a deployment in which every capture was failing and one in which no capture had ever been attempted exported byte-identical metrics, and those are opposite operational states - a broken storage provider against a correctly idle deployment. Read `attempted` as the **sum across the `outcome` tag**, which keeps that total a within-family comparison that cannot drift from its parts; a bare failure counter would not have fixed the defect, because zero failures is itself ambiguous between "all succeeded" and "none ran". The `abandoned` split is load-bearing rather than cosmetic: a fleet-wide shutdown abandons thousands of captures **by design** (issue #1965), and folded into `failed` that routine event would present as a mass provider outage. Declines are deliberately **not** counted - no tree id, nothing checkpointed and no live data, a capture already in flight, or no coverage it could claim (see `orleans.lattice.leaf.snapshot.capture.declines`) - so the total is attempts, not invocations. |
| `orleans.lattice.leaf.snapshot.capture.declines` | `Counter<long>` | `{decline}` | Leaf-snapshot capture invocations that **declined before the attempt boundary**, tagged `tree` and `reason`: `no_tree_id`, `not_eligible`, `already_in_flight`, `no_coverage_claim` (issues #2696, #2725). It is a **separate instrument** rather than a fourth value of the `outcome` tag on `orleans.lattice.leaf.snapshot.captures`, and the separation is load-bearing: that counter and `orleans.lattice.leaf.snapshot.capture.duration` are written at one boundary so they share a population exactly, and a decline is never timed, so folding declines into the attempt counter would make the two families silently differ. It exists because without it the capture instruments reproduce, one gate higher, the ambiguity they were added to remove - a deployment in which every capture is *declined* and one in which the capture path is never reached at all would both leave the attempt counter at zero. A declined capture is a **third state**, distinct from a failed capture and from an idle deployment, and it is the most likely way a self-heal silently never runs. The four reasons call for different responses: `already_in_flight` dominating is a **contention** signal, captures being requested faster than the shared snapshot storage provider retires them; `not_eligible` dominating is the **starved-leaf** signal, a leaf with nothing checkpointed and no live data that will never cover itself; `no_tree_id` above zero is a **bug**, capture invoked on a leaf never attached to a tree; `no_coverage_claim` is the **unclaimable-rows** signal (issue #2725), a leaf holding live rows that has never checkpointed, so any blob it wrote would claim no coverage and be refused by the load gate - benign, since WAL replay covers such a leaf completely, but it names the population whose retained WAL cannot shrink until it checkpoints. A `no_tree_id` decline carries **no** `tree` tag, because the leaf has no tree identity at that point, but it still carries the derived `tenant` label with the reserved `_platform_` value - read it as a global count and do not expect it under a per-tree filter. |
| `orleans.lattice.leaf.snapshot.driver.declines` | `Counter<long>` | `{decline}` | Leaf-snapshot capture **drivers** that declined to drive a capture, tagged `tree` and `reason` (issue #3185). Deliberately a **separate instrument** from `orleans.lattice.leaf.snapshot.capture.declines`, and the separation is load-bearing in exactly the way that counter's own separation from `orleans.lattice.leaf.snapshot.captures` is: those declines are raised *inside* `CaptureSnapshotCoreAsync`, so declines and attempts together partition its invocations exactly and stay co-populated with `orleans.lattice.leaf.snapshot.capture.duration`, whereas the declines counted here happen in a **caller that never reached that method**. Folding them in would count a strict superset and silently corrupt the decline-versus-attempt ratio the other counter exists to report. It exists because the capture *drivers* were the last unlit segment of the snapshot-coverage path: the zero-coverage repair reports six outcomes and `CaptureSnapshotCoreAsync` reports four declines, but the graceful-deactivation hook and the periodic recheck returned in **complete silence** on every gate. Ten arms: `deactivate_unproven_coverage_stale` (the #1535 no-loss gate declined while some partition's checkpoint had outrun its durable coverage, so this leaf leaves behind a pin frozen below its own checkpoint), `deactivate_unproven_coverage_current` (the same gate declining harmlessly, coverage already current everywhere), `deactivate_unproven_unclassified` (the classification itself threw - expected to be zero outside a test harness, and kept separate so the two informative arms never absorb an unknown), `recheck_cadence_not_reached`, `recheck_capture_in_flight` and `recheck_coverage_current` for the three periodic-recheck gates, `recheck_no_durable_checkpoint` (issue #3300) and `recheck_checkpoint_stalled` (issue #3389) when the coverage-lag timer instead routes a leaf holding rows to the starvation drive, because no partition has ever checkpointed or because its checkpoint has stopped advancing, and `recheck_drive_refused` and `recheck_drive_deferred` (issue #3575), which **qualify** those two routing arms rather than adding to them. Each qualifier is recorded on the same tick as one of the routing arms - refused when the routed drive was refused admission to the per-silo WAL replay gate before replaying anything, deferred when it was skipped because the leaf is backing off after a refusal - so the reasons on this instrument do not partition ticks. The timer's drives never take the last free slot of the gate's GC share, which is kept for the WAL GC sweep, and while that share is a single slot they yield it to a refused sweep drive; a refused leaf backs off by a count of drive opportunities that doubles with consecutive refusals up to four to seven, jittered per leaf, and a drive that is admitted ends the backoff. Before issue #3575 the refusal escaped the timer callback as an unhandled exception, logged twice by the runtime and counted nowhere. **Reading it: `deactivate_unproven_coverage_stale` is the minting rate of the frozen floor-holder population and is the arm that matters.** The decline itself is *correct* - a leaf that neither advanced a checkpoint over cache-resident applies nor cold-rebuilt its cache from the WAL start cannot honestly stamp coverage - so what is being measured is the **rate**, not a fault. Read it against the WAL GC reactivation drive's completion rate (`orleans.lattice.wal.gc.blocked_leaf_reactivations`): sustained above it, the frozen population grows without bound and **no drive-side tuning can converge**, because the mismatch is asymptotic rather than a matter of constants; below it, the drive is merely slow and tuning applies. Until this instrument existed only the accumulated *consequence* of that minting was observable and never its rate, which is why a flow presented as a static population. Like `orleans.lattice.leaf.snapshot.capture.declines` it is **not zero-primed**, so a series appears on first decline and absence does **not** establish that a build lacks the instrument - the usual "presence of a series is the deployment fact" reading does not apply to this family. A decline on a leaf with no `tree` identity carries no `tree` tag, for the same reason `no_tree_id` does not. |
| `orleans.lattice.leaf.snapshot.capture.duration` | `Histogram<double>` | `ms` | Wall-clock duration of a leaf-snapshot capture attempt, recorded at the same boundary and with the same `tree`/`outcome` tags as `orleans.lattice.leaf.snapshot.captures`, so the two share a population exactly: every counted attempt is timed and every timed attempt is counted (issue #2696). Timing starts at the single-flight boundary rather than at method entry, which is a correctness requirement and not a detail: the decline gates return in microseconds and are taken far more often than a capture runs, so timing them would let near-zero no-ops dominate the sample count and report "captures are fast" precisely when none are happening. Capture is `await`ed inline inside `OnActivateAsync`, and Orleans delivers no request to a grain until activation completes, so this duration is **paid by every caller of that leaf** - it is activation latency, not background work. All per-leaf snapshot grains write through one shared storage provider with no cross-leaf admission control (issue #2696), so a mass capture event is the shape a latency cliff takes. The Prometheus exporter renders this as an explicit-bucket histogram, exactly as it does every other `ms`-unit histogram on this meter, so `histogram_quantile` over `_bucket` is available and is what the bundled CommitPath panel charts. A delta of `_sum` divided by the delta of `_count` gives an interval mean as a companion reading; do not read either cumulative total as a level. |
| `orleans.lattice.leaf.snapshot.capture.concurrency_peak` | `ObservableGauge<int>` | `{capture}` | Greatest number of leaf-snapshot captures observed executing **concurrently on this silo** since process start, as a monotone non-decreasing high-water mark (issue #2696). It measures the quantity every other capture instrument is blind to by construction: the single-flight guard those are written around is an instance field on one activation, so it stops a leaf capturing twice at once and says nothing about how many *different* leaves are capturing against the one shared snapshot storage provider. A thousand leaves each passing their own guard is the fan-out issue #2696 describes, and it registers on no per-leaf series. The shape is chosen because both obvious alternatives fail **silently** on this deployment. An *instantaneous* gauge of the current depth is sampled only when the endpoint is scraped, so the transient spike that makes the metric worth having is exactly what falls between two scrapes - the objection PR #2723 raised when it deferred this metric, and it is correct. A `Histogram<T>` recorded at each entry is event-sampled and so cannot miss the spike, but the repository-context container's own exposition renders every histogram as `_sum`/`_count` with no buckets and no quantiles, leaving only the mean, and a mean dilutes a spike to nothing: one capture at depth 64 among ten thousand at depth 0 reads as `0.0064`. A monotone peak is immune to both, because a value that never falls is reported by the scrape that sees it **and by every scrape after it**, so scrape timing stops being a risk to manage and becomes irrelevant. The cost is that it is a high-water mark for the life of the process and does not decay; that is the intended reading, and a restart resets it. Registered at host build from `AddLattice` rather than on first capture, so a silo that has never captured still publishes the series and a scraped `0` means **measured none** rather than **no detector**. Carries the constant platform `tenant` sentinel and no `tree` tag: the peak is a maximum across every leaf on the silo, so it spans trees and tenants, and the contended resource - one storage provider per silo - is itself silo-wide. Splitting it per tree would report several smaller numbers, none of which is the depth the provider actually saw. Pair it with `orleans.lattice.leaf.snapshot.capture.concurrent_entries` for the per-tree and frequency halves of the question: the peak says how bad it got, the counter says how often it happens, and a peak alone cannot tell one transient spike at startup from sustained contention. |
| `orleans.lattice.leaf.snapshot.capture.concurrent_entries` | `Counter<long>` | `{capture}` | Leaf-snapshot captures that crossed the attempt boundary while **at least one other capture was already in flight** on the same silo, tagged `tree` (issue #2696). It is the frequency half of the cross-leaf concurrency reading whose magnitude half is `orleans.lattice.leaf.snapshot.capture.concurrency_peak`, and the pairing is load-bearing rather than decorative: a high-water mark alone cannot distinguish a single transient depth of 3 during startup from a silo sitting 3 deep continuously for an hour, and those are opposite operational states. Note it is **not** a duplicate of the `already_in_flight` reason on `orleans.lattice.leaf.snapshot.capture.declines`: that reason counts a capture the per-leaf guard **turned away** on a leaf already capturing, whereas this counts a capture that was **admitted** while a *different* leaf was capturing. The two populations are disjoint - one never reaches the attempt boundary, the other is past it - and only this one describes real concurrent load on the shared provider. Emitted on **every** attempt, adding `0` for an uncontended one, so each tree gets a zero-primed series from its first capture; the zero-add is deliberate and must not be optimised away, because without it a tree with no contention is indistinguishable from a tree whose captures never ran. |
| `orleans.lattice.leaf.snapshot.load_failures` | `Counter<long>` | `{load}` | Snapshot loads that failed, distinct from genuine snapshot absence (issues #2364, #2404). Tagged `tree`, `reason`, `tenant`. The three reasons are `resource_exhausted` (an observable `OutOfMemoryException`), `contiguity_exhausted` (that OOM while hydration was admitted as the sole occupant), and `unclassified` (no OOM observable in the exception chain). Both observed-memory arms decline activation instead of escalating to a more expensive WAL replay. For `contiguity_exhausted`, divide the leaf or lower `MaxLeafBytes`; increasing the memory grant does not repair a contiguous-allocation failure. `unclassified` retains the existing cold-replay fallback but **does not rule out memory pressure**: a snapshot-grain activation failure can hide the original OOM. Correlate storage/activation logs and memory headroom rather than inferring a storage-provider fault. Read alongside activation failures and `orleans.lattice.leaf.snapshot.hydration_admissions`; a load failure alone does not establish whether activation succeeded. `unclassified` replaces the historical `faulted` value; update reason-filtered alerts to include both during rollout. The public `SnapshotLoadFailureFaulted` field is retained for source compatibility and now exports `unclassified`. This is a vocabulary correction, not a claim of live production OOM attribution. |
| `orleans.lattice.leaf.snapshot.hydration_admissions` | `Counter<long>` | `{hydration}` | Activation-time leaf-snapshot hydrations that passed the per-silo admission gate (issue #2765). Tagged `tree` and `outcome`: `immediate` when the claim fitted the budget straight away, `queued` when it had to wait behind it, and `sole_occupancy` when the contiguity rule serialised it (issue #2844). These three strings are the literal tag **values** the gate emits, from `LatticeMetrics.SnapshotHydrationAdmittedImmediately`, `SnapshotHydrationQueued` and `SnapshotHydrationSoleOccupancy`; document them exactly as emitted and never derive them from the constant's field name, because a query or alert matches on the value, so a plausible near-miss spelled from the field name instead of the value matches no series at all while looking correct (issue #2854). The gate exists because nothing bounded how many leaves materialised their persisted state at once on a cold start, and the cost of one such materialisation is **several multiples of the stored frame**, by an amount that depends on which read path the payload takes (issue #2858). A blob written by a current build carries the Orleans binary `LGB1` marker and costs about **2x** the frame: the provider's column read plus the decoded rows. A legacy JSON blob written before issue #2516 still decodes as JSON, because reads route on the stored payload rather than on the type, and costs roughly **10-13x** the frame: the base64 document is itself 4/3 of the frame, and reading it back allocates a UTF-16 `System.String` of the whole document, Newtonsoft's growing character buffer, and the decoded rows on top. The gate prices every claim at **5x** the stored frame whichever path it will take, because it sizes the claim before the row is read. That factor covers the binary path with headroom and deliberately under-prices the legacy one, whose excess the budget absorbs: an all-legacy cold start, fully admitted, still peaks at about a third of the heap hard limit. Do not lower it to the binary figure while legacy blobs remain, and do not raise it to the legacy figure, which would halve admission for every current blob to protect a population that shrinks as leaves re-snapshot. The aggregate in the original incident crossed the .NET heap **hard limit**, which the runtime sizes from the container's cgroup, so the process threw a **managed** `OutOfMemoryException`, exited 0 and was never reported as OOM-killed - a restart loop that read for a long time as an unexplained clean exit. The budget is derived from that same heap hard limit and is deliberately **not configurable**, because a bound an operator has to set cannot rescue a process that is already restart-looping. Admission is first-in-first-out, and a claim larger than the entire budget is admitted as soon as it is the only claimant - without that rule an oversized leaf could never activate, so it could never be divided back under bound, and the gate would turn a crash loop into a permanent stall. Read a climbing `queued` rate during a cold start as the gate **working**: hydrations are being serialised instead of exhausting the heap together, and it should decay as leaf division shrinks the corpus. A `queued` rate that never decays across many restarts means division is not progressing; correlate with the leaf byte-overflow and capture instruments, not with container memory headroom. All three outcome arms are pre-minted at zero the first time a tree is seen in the process, so a flat zero is a measured zero and an absent series means the build does not carry the gate. The `sole_occupancy` arm reports a different predicate and is read differently. The budget above is denominated in **total** bytes, and the allocation that actually failed in issue #2844 is a **single contiguous** one, which must be met by an unbroken run on the large object heap and which no amount of aggregate headroom guarantees. How large that run is depends on which read path the payload takes, and the gate cannot know that before it sizes the claim: a blob written by the Orleans binary serializer returns as a `byte[]` of roughly the stored frame, while a legacy JSON blob written before issue #2516 returns as a contiguous UTF-16 `System.String` costing **8/3** of the frame (base64 inflates it by 4/3, then each character costs 2 bytes). Reads route on the stored payload rather than on the type, so both paths stay live and the gate assumes the worse of the two. A claim whose stored frame exceeds the default leaf size bound is therefore admitted only as sole occupant and excludes every other claim while it is held; that ratio applies equally to the claim and to the ceiling, so it cancels out of the comparison and governs only the contiguous figure the exception reports. That ceiling is deliberately **absolute rather than derived from the memory grant**, because every other bound here - this budget, the resident working-set budget, the replay gate - rises with the grant and admits *more* concurrent work, while contiguous feasibility does not improve with it at all. Read a sustained `sole_occupancy` rate as a signal to divide leaves or lower `MaxLeafBytes`, and specifically **not** to provision more memory, which would make it more frequent. All three outcome arms are pre-minted at zero. |
| `orleans.lattice.leaf.snapshot.segment_reads` | `Counter<long>` | `{read}` | Individual snapshot **segment** frames read during a segmented hydration (issue #2914). Tagged `tree` and `outcome`: `loaded` when the frame read back and validated, `missing` when it was absent or failed validation, `failed` when the read threw. It exists because a snapshot whose encoded frame exceeds `LeafSnapshotSegmentBytes` is no longer stored as one row: the row payload is split into row-aligned segments, each a standalone frame in its own grain-state row, and hydration folds them one at a time. The split is what makes the bound real. The contiguous allocation that failed in issue #2844 is materialised by the storage provider's **column read**, before any lattice code runs, so bounding it requires bounding the column, and the only way to bound a column is to stop putting the whole payload in one. `missing` and `failed` are deliberately separate arms because a torn durable snapshot and a transient storage fault call for different responses: `missing` means the manifest references a segment that is not there, which is a durability question, while `failed` is an I/O question. Either one **fails the hydration closed** - the cache is emptied and the rehydrate declines, so the leaf replays its whole readable WAL rather than activating with a snapshot that silently holds fewer rows than its coverage claims, which is the one shape that would let the coverage-gated WAL GC trim the last durable copy of a prefix. All three outcome arms are pre-minted at zero the first time a segmented hydration is observed for a tree in the process, so a flat zero is a **measured** zero (an absent series means no snapshot for that tree has yet exceeded the segment window, or the build lacks the instrument); without that priming an uninstrumented branch and a never-taken branch export byte-identical metadata. |
| `orleans.lattice.leaf.snapshot.segmented_hydrations` | `Counter<long>` | `{hydration}` | Activation-time leaf-snapshot hydrations that folded a **segmented** snapshot to completion (issue #2914). Tagged `tree`. Read it as the population for which the contiguity bound was actually exercised: an ordinary leaf whose snapshot fits inside `LeafSnapshotSegmentBytes` is stored and hydrated inline and is never counted here, so this counter names exactly the leaves that would previously have demanded a single contiguous array of the whole payload. It is counted only on a **complete** fold, so it is not a proxy for attempts; a hydration that declined partway is visible on the `missing` or `failed` arms of `orleans.lattice.leaf.snapshot.segment_reads` and not here, which keeps "how many segmented snapshots hydrated successfully" and "how many segment reads went wrong" from sharing a population they do not share. Read it alongside `orleans.lattice.leaf.snapshot.segment_peak_bytes`, which reports whether the bound held, and against `orleans.lattice.leaf.snapshot.hydration_admissions`, whose `sole_occupancy` arm should **fall** as segmentation takes effect, since a segmented snapshot no longer presents an oversized contiguous claim to the gate. Pre-minted at zero per tree. |
| `orleans.lattice.leaf.snapshot.segment_peak_bytes` | `ObservableGauge<long>` | `By` | **High-water** mark of the largest single contiguous segment frame this process has materialised while hydrating a segmented snapshot, per `tree` (issue #2914). This is the instrument that says whether the bound held: it must stay at or below the configured `LeafSnapshotSegmentBytes` window however large the underlying snapshot grows, so a value above the window is the direct signal that segmentation is not bounding what it claims to bound. It is a **monotonic high-water gauge rather than a histogram**, and that choice is load-bearing on the repository-context container's own exposition, which renders every histogram as a summary carrying only `_sum` and `_count`, with no buckets and no quantiles, so the only reading available from one there would be a mean - and a mean is exactly the wrong statistic for a peak, diluting the single large allocation that matters into an average of many small ones until it disappears. Because the value never decreases within a process lifetime, read it as "the worst this host has seen since it started", not as a current level; it resets on restart, which is the intended behaviour for a high-water mark. Pre-minted at zero per tree, but only when the tree's first segmented hydration is observed, so the series is absent until a snapshot for that tree has exceeded the segment window, and a flat zero means the segmented path was entered but no segment frame has yet read back. |
| `orleans.lattice.leaf.residency.sheds` | `Counter<long>` | `{activation}` | Leaf activations deactivated by the per-silo **resident leaf working set** so the hydrated population stays under its derived byte budget (issue #2767). Tagged `tree` and `kind`: `banked` when the shed leaf had a snapshot to reload from, `unbanked` when it did not. It exists because `orleans.lattice.leaf.snapshot.hydration_admissions` bounds only how many hydrations run **concurrently**, and the heap cost of a hydrated leaf is **resident, not transient**: `LeafSnapshotHydrationSource` retains the whole encoded frame for the activation's lifetime, so completed hydrations accumulate at roughly one frame each and a forced blocking compacting gen2 collection reclaims none of it while the activations live. Concurrency-bounding therefore controls the peak of each hydration and nothing at all about the steady state, which is why the heap hard limit was still reached with that gate provably binding. Shedding is a **graceful** deactivation, so the leaf captures a snapshot on the way out and an `unbanked` shed converts into a `banked` candidate for next time - the reason the class preference drains rather than recirculates. Read a steady `banked` rate as the bound working at its cheap end. A sustained `unbanked` rate is the one to act on: it means the working set has run out of cheap candidates and is paying whole-window replays to stay under budget, so the answer is snapshot coverage, not this bound. Both arms are pre-minted at zero the first time a tree registers, so a flat zero is a measured zero and an absent series means the build does not carry the bound. |
| `orleans.lattice.leaf.residency.budget_bytes` | `ObservableGauge<long>` | `By` | The byte budget the per-silo **resident leaf working set** is actually enforcing (issue #2788). Carries the platform tenant sentinel (`tenant="_platform_"`) and no other tag: the working set is per-silo and spans every tree, so it has no tenant to derive and a `tree` tag would be a category error, and the silo dimension arrives from the exporter's resource attributes. The sentinel is emitted rather than omitted so that an unattributable measurement is distinguishable from a site that simply failed to attribute one. **Read it against the container's memory grant.** This gauge exists because `orleans.lattice.leaf.residency.sheds` reads zero in two situations with opposite remedies - the bound is populated and correctly under budget, or the bound cannot engage at all - and no amount of priming separates them, because priming makes a counter's *silence* readable while a counter observes an **action** and this question is about **state**. A budget of the same order as the whole container grant is the second case: the process reaches its heap limit and dies before the threshold can ever be crossed, so the bound is present, plausible-looking, and doing nothing. That was a real defect, not a hypothetical: the budget was originally derived from `GCMemoryInfo.TotalAvailableMemoryBytes` alone on the belief that it reports zero when no heap hard limit is configured, whereas it reports **host physical memory**, so a large host running a small container derived a budget larger than the grant. It is now the smaller of the heap hard limit and the cgroup memory limit, divided by `HeapBudgetDivisor` (4) and floored at 64 MiB. A flat line at exactly 1073741824 is the both-ceilings-unknown fallback and should be investigated as a **detection** failure rather than read as a tuned value. |
| `orleans.lattice.leaf.residency.resident_bytes` | `ObservableGauge<long>` | `By` | Bytes currently accounted to live, un-shed leaf registrations in the per-silo resident leaf working set (issue #2788). Carries the platform tenant sentinel, for the same reason as the budget gauge. Read as a ratio against `orleans.lattice.leaf.residency.budget_bytes`: that ratio is the headroom, and it is the only reading that separates an estate comfortably under bound from one about to begin shedding. Resident pinned near the budget with a steady shed rate is the bound working. Resident at a small fraction of the budget **while the process is exhausting memory** does not mean leaf residency is not the cost; it means the budget is too large to be a bound, so compare the budget line against the container grant. Resident flat at zero while leaves are demonstrably activating is a wiring defect rather than quiescence, and `orleans.lattice.leaf.residency.registrations` is what tells the two apart. |
| `orleans.lattice.leaf.residency.registrations` | `ObservableGauge<long>` | `{registration}` | Leaf registrations currently held by the per-silo resident leaf working set (issue #2788). Carries the platform tenant sentinel, for the same reason as the budget gauge. **This is the instrument that settles whether a zero shed count is quiescence or a dead ledger.** An empty ledger and a populated ledger correctly under budget are indistinguishable on the shed counter, because a counter reports events and both of those are the *absence* of an event; this gauge reports the ledger's occupancy directly. A non-zero count with zero sheds is the bound populated and under budget, which is healthy and needs no action. A count flat at zero while the silo is activating leaves means the registration path is not running and the bound is inert. All three residency gauges are registered eagerly from `AddLattice` rather than on first registration, so a present-and-zero reading is a **measured** zero and an absent series means the build does not carry them. |
| `orleans.lattice.leaf.snapshot.coverage_repairs` | `Counter<long>` | `{evaluation}` | Zero-coverage repair **evaluations** driven by the leaf's zero-coverage snapshot repair (issues #2692, #2940). Tagged `tree` and `outcome` with seven arms in two classes, and the distinction is load-bearing rather than taxonomic. **Six are terminal outcomes, and those six partition every invocation of the repair** - the denominator is named deliberately, because it includes the two invocations that decline before any repair work happens, so "invocations that reach the repair" would be the wrong denominator and would leave the declines outside the partition: `repaired` (a capture left every checkpointed partition covered), `unsatisfied` (a capture ran and a checkpointed partition is **still** uncovered - including a capture that threw or timed out, because the advisory capture wrapper swallows every exception, so a leaf whose snapshot store is failing is counted here rather than going silent), `exhausted` (the attempt budget was spent with a partition still uncovered), `capture_in_flight` (declined, a snapshot capture was already running on this leaf), `no_checkpointed_uncovered_partition` (declined, no checkpointed partition lacks coverage - the healthy majority, recorded on every activation and every checkpoint persist of a leaf that is not in the repairable population, so its magnitude tracks write volume rather than severity; it is not comparable with the blocked arm of `orleans.lattice.wal.gc.blocking_pin_state`, whose population is disjoint from this one by construction, but it does intersect that instrument's floor-holder arm by design, because a pin classified `checkpointed_uncovered` there is driven into this very repair - so that classification standing against a climbing `no_checkpointed_uncovered_partition` for the same leaf is a direct contradiction rather than two unrelated readings, issue #3168) and `backing_off` (suppressed, an earlier exhaustion armed a re-arm backoff that has not yet elapsed). **The seventh, `rearmed`, is a lifecycle transition and NOT a terminal outcome**: it records that a spent attempt budget was reset after its backoff, and the invocation that records it then proceeds to the repair and goes on to record `repaired` or `unsatisfied` as well. Those two are named rather than counted deliberately: by the time an invocation records `rearmed` it has passed both entry guards and both the abandon and backoff branches, so four of the six terminal arms are unreachable from there, and "one of the six" would be true only in the weak sense of not excluding them. It therefore **co-occurs** with a terminal arm rather than excluding one, exactly as `rearmed` does on the sibling `blocked_leaf_reactivations_total`, whose documentation separates its lifecycle events from the terminal outcomes that partition for the same reason. Do not read the seven as a partition; read the six as one, and `rearmed` as an event alongside them. A leaf that holds a checkpointed partition with no durable snapshot coverage publishes the `(Zero, -1)` block pin, and because that leaf's Zero-HLC block pin disables the cursor trim in every partition it registers one in, **one such leaf disables cursor-based WAL trimming for its entire tree** - which is how a single-silo deployment reached tens of gigabytes of retained WAL. **Attribute that to the right axis.** The disabling happens on the **frontier** axis, in the WAL GC durable-materialiser floor's `pin <= HybridLogicalClock.Zero` branch, which blocks the cursor branch for the pin's own partition; a leaf seeds a Zero pin in *every* partition at birth, and that is what carries a single leaf to the whole tree. The **offset** floor does the opposite of folding: the offset-floor computation *skips* a `-1` reporter and lets only real checkpoints (`offset >= 0`) constrain the floor, specifically so that a `-1` cannot collapse it and wedge the trim tree-wide - a skipped reporter is deliberately left out of the covered set and protected by its cursor or its block pin instead (issue #3172). The distinction is load-bearing rather than pedantic: attributing the fold to the offset floor makes an offset-based release read as impossible when it is in fact the designed path. **The two rejection arms are the discriminator, and they are why this instrument was widened.** Until issue #2940 the guard exit and the attempted-but-unsatisfied exit both recorded nothing, so "the repair never ran" and "the repair ran and failed" were byte-identical silence while calling for opposite responses: `no_checkpointed_uncovered_partition` says the repair is structurally unable to address the fault and routes the remedy to the pin/guard seam, whereas `unsatisfied` says the repair is reaching the fault and failing and routes it to the capture seam. **`exhausted` and `rearmed` are a pair, and the gap between them is the reading that matters.** The attempt budget was documented as a per-activation ceiling, which is self-limiting only while activations turn over - false for a leaf held permanently active by read traffic, where it was in effect a per-process ceiling that, once spent, was never reset at all. It is now abandoned on exhaustion and re-armed in place after a backoff that DOUBLES per abandonment to a ceiling, so an `exhausted` count with no matching `rearmed` is a leaf still abandoned, whereas the two advancing in step is a leaf retrying on an ever-slower schedule - the intended shape for a leaf whose captures keep failing, since the underlying fault is a capture timeout rather than a transient one and a flat backoff would have it paying a timing-out capture at constant cost for ever. **All seven outcome arms are zero-primed per tree** the first time a tree is seen in this process, so an **absent** series means the repair path never ran for that tree, or the running binary predates the instrument - it does not mean healthy - and a **zero is a measured zero**. Since issue #3195 this instrument also has a **cadence floor**: the coverage-lag timer drives the repair once per `LeafSnapshotMaxCoverageLagSeconds` (default 300s) per ACTIVE leaf, above the cadence gate and the debounce, so an arm is recorded regardless of write traffic - which is what makes a zero here interpretable at all, unlike a purely traffic-driven counter, provided the observation window is at least the cadence (the tick phase is jittered uniformly across the period) and the option has not been set to `0` or below, which disables the timer and removes the floor. That guarantee was previously conditional on the tree having already emitted `repaired`, which is precisely the trees not under diagnosis. The priming walk iterates the single declared arm set `LatticeMetrics.CoverageRepairArms`, pinned by a reflection test to every `CoverageRepair*` arm declared on the class, so a new arm declared the way every existing arm is declared is forced into the primed set. Be exact about the residual gap rather than importing the sibling's stronger claim: a recording site that inlines a tag pair instead of declaring a static would evade both the reflection test and the priming walk, so "an arm added later CANNOT ship unarmed" is true of the sibling and NOT true here. This is a weaker construction than that sibling `blocked_leaf_reactivations_total`, which walks an enum through a switch that throws on an unmapped member, because these arms are tag-value statics rather than enum members - the guarantee is enforced by a test rather than by the compiler. One boundary now, and the history of the other is worth keeping because it is the reason this instrument is trusted. A re-arming invocation records `rearmed` **in addition to** its terminal arm, so summing all seven over-counts on exactly the leaves this repair is working hardest on. **Sum the six terminal arms if you want an invocation tally**, and read `rearmed` separately as an event count. That sum is now EXACT rather than a lower bound: every path below the two entry guards records exactly one terminal arm, which the fixtures pin as an equality (41 evaluations, 41 increments) rather than as an inequality, so a new silent return added below the guards fails the build instead of quietly reappearing as an under-count. Until `backing_off` existed the sum WAS a lower bound, and the documented reason was wrong as well as incomplete: the shortfall was attributed to the once-per-activation exhaustion dedup latch, but the suppressed invocations never reached that latch at all - they took the backoff-suppression branch, which recorded nothing. That branch is the one that needed an arm, and on an eight-hour backoff against a 300-second recheck cadence it was hiding roughly ninety-six invocations behind a single `exhausted`. The dedup latch is in fact unreachable in the present flow (its one call site is guarded by a null re-arm deadline that the same branch immediately sets, and only the re-arm clears it, in the same step); it is retained as a defence because that reachability argument is a property of the call-site guards rather than of the latch. Read `repaired` as the heal rate while a stranded deployment drains, expected to fall to zero once every leaf is covered; read a sustained `exhausted` or `unsatisfied` rate as a snapshot-store or capture-seam fault; read a tree whose only non-zero arm is `no_checkpointed_uncovered_partition`, while its WAL is still pinned, as evidence the pin has a cause this repair does not reach. |
| `orleans.lattice.leaf.unresolved_prepare_ledger_beyond_cap` | `Counter<long>` | `{prepare}` | Resident unresolved saga prepares recorded into the leaf's durable replay-work ledger once that ledger holds at least `MaxDurableUnresolvedReplayWork` entries (issue #2183). The threshold is inclusive (issue #2756), so the prepare that fills the ledger to the cap is counted as well as every prepare recorded past it: the cap is the resting value at which the ledger starts refusing deferred terminals, so that is the value the signal must see. A resident prepare must never be dropped - dropping one pins the incremental flush ceiling at `(prepare - 1)` permanently and the leaf banks zero forward progress, which is the frozen-checkpoint livelock this counter's fix removes (a latent defect proven by a two-arm unit control, not the deployed-leaf freeze tracked as issue #2220) - so past the cap the prepare is recorded unconditionally and the row is allowed to grow for as long as a saga leaves a prepare unresolved (registry status InFlight: the residual population after issue #2190's self-terminalisation, whose orphan source is tracked as issue #2304). **Persist risk:** Azure Table Storage rejects oversized writes at its ~960 KB grain-state limit (see [Storage provider per-row limits](../tree-storage.md#storage-provider-per-row-limits)), bounding persisted row growth while leaving the previous row intact. **Read risk:** the larger SQLite limit on the default `local` profile permits growth that can exhaust memory or the read budget during activation, before grain-level repair can run. A successful persist is not proof of a safe activation read; write amplification is not the only cost. Alert on every profile, including `local`. It is observability only - nothing here caps or drops a prepare, because a behavioural cap would reintroduce the exact drop-and-freeze defect issue #2183 removes. A paired warning fires once per activation when the crossing first occurs. Tagged `tree` and `partition` (the WAL partition ordinal, bounded by `WalPartitions`). |
| `orleans.lattice.materialiser.drain_lag` | `Histogram<double>` | `ms` | Materialiser drain lag, in milliseconds, recorded by the WAL saturation sampler on every tick (default 200 ms) while the drain-lag input is enabled, as the WAL head wall-clock timestamp minus the slowest eligible in-memory WAL consumer cursor - leaf materialisers and tree-wide consumers alike - clamped at zero. Recorded for every checked tree, not only over-threshold ones. A sustained breach of `WalSaturationMaterialiserLagThreshold` holds the tree at `Throttled` and never escalates to `Saturated`. Measures the in-memory cursor, so it does not observe the durable materialiser pin store that sets the WAL retention floor (issue #2015); `orleans.lattice.materialiser.pin.durable_write_latency` is the input that does. The frontier is a minimum over the consumers `WalDrainLagConsumerFreshness` admits: a consumer whose last report is older than that window is excluded (issue #2446), and so is a leaf materialiser whose position has not advanced within it and predates it, because a leaf's cursor covers only its own key range and a leaf whose range received no writes is caught up however far its cursor trails the tree-wide head (issue #3131). A leaf that is draining, or a tree-wide tailer that has stalled, still sets the reading on its own, and the aggregate carries no dimension naming which consumer contributed it, so a large value is a prompt to identify the contributor - the sampler logs up to three eligible holders on the crossing and periodically during a standing breach (see [holder observations](#drain-lag-holder-observations)) - rather than evidence that projection is falling behind ingest (issue #2433). Tagged `tree`. |
| `orleans.lattice.materialiser.lagging_consumers` | `Histogram<int>` | `{consumer}` | Count of WAL cursor consumers whose own reported cursor trails the WAL head wall clock by more than `WalSaturationMaterialiserLagThreshold`, recorded by the WAL saturation sampler for a tree whose aggregate drain lag is already over that same threshold on the same tick. Counts exactly the consumers the aggregate was computed over: consumers that have never reported a cursor (HLC zero), cold consumers, and position-stale leaf materialisers are excluded by the same predicate as the `min(cursor)` meet (issues #2446, #3131). It exists because `orleans.lattice.materialiser.drain_lag` above is a minimum and therefore reads identically for one legitimately dormant consumer holding the frontier down and for many consumers genuinely falling behind, two conditions whose correct responses are opposite (issue #2444). It makes that aggregate triageable, not diagnosable: it separates those two shapes but still does not name the contributing consumer, which is named out of band by the [holder observations](#drain-lag-holder-observations). Sampled only for trees already found over threshold, so a healthy estate adds no per-tick cost and emits nothing. Tagged `tree`. |
| `orleans.lattice.materialiser.pin.durable_writes` | `Counter<long>` | `{write}` | Durable writes to the leaf-materialiser pin store: one increment per persist that issued at least one durable write, so a persist that throws records nothing and, once a shard's pins span several slots, one persist can rewrite more than one slot. Tagged `tree` and `outcome`: `birth` (the awaited seed that makes a newly born leaf's block pin durable before its data becomes reachable) or `coalesced` (every other persist - chiefly the coalesced flush that makes merged pin reports durable, plus a pin removal, the tree-deletion clear, a re-layout and the deactivation flush). It counts **writes, not advancement**, and the two are not proportional: pins are bucketed, so one advancing consumer rewrites its whole bucket, and a merge that moves only the HLC frontier rewrites the pin at the same checkpoint offset. Offset advancement is the quantity that lets the WAL GC offset floor move, so read `orleans.lattice.materialiser.pin.advances` for that (issue #2694), not this. |
| `orleans.lattice.materialiser.pin.advances` | `Counter<long>` | `{report}` | Merged leaf-materialiser pin reports partitioned by **which of the two axes advanced**, emitted once per merged report. Tagged `tree` and `outcome`: `both` (the HLC frontier and the durable checkpoint offset moved together), `offset_only` (the offset moved and the frontier did not), `frontier_only` (the frontier moved and the offset did not), or `none` (a report at or behind the existing pin on both axes). The four arms are the complete truth table of the two axes, so exactly one is recorded per report and each marginal is recoverable by summing: offset advanced is `both` + `offset_only`, frontier advanced is `both` + `frontier_only`. They were three **first-match** arms until issue #3163: `offset` absorbed every merge that also advanced the frontier, so it meant "the offset advanced, the frontier is unknown", and an absent `frontier_only` did not show a flat frontier - only one that never advanced alone. Read `both` + `offset_only` when asking "why is retained WAL not being reclaimed?", and read a healthy `offset_only` rate with `both` flat as a consumer advancing in offset space alone, which cannot release a WAL entry because the offset floor only lowers a trim point the HLC clauses already authorised. It exists because `durable_writes` above cannot separate any of these (issue #2694). Every arm is zero-primed once per pin-shard activation, so all four series exist for a tree with a live pin grain and a flat arm is a measured zero rather than an absence. |
| `orleans.lattice.materialiser.pin.durable_write_latency` | `Histogram<double>` | `ms` | Caller-observed duration of a durable leaf-materialiser pin write, recorded by `LeafCursorReporter` around each pin grain call that reached `WalSaturationMaterialiserPinLatencyThreshold` or faulted. The only materialiser instrument that observes the durable retention floor rather than in-memory progress, so it is what makes a stalled pin store visible (issue #2015). Measured at the call site, so it includes time queued ahead of the shard's non-reentrant activation. Emitted only when the threshold option is set; it defaults to null. Tagged `tree`. |
| `orleans.lattice.materialiser.pin.reports_shed` | `Counter<long>` | `{report}` | Steady-state per-checkpoint pin reports a leaf's cursor reporter declined to enqueue because the target shard's last durable write showed the store is not keeping up. Only that path is shed; the birth block-pin seed and the deactivation flush are last-chance writes and are always issued. Shedding is safe for **durability** - a shed report leaves the durable pin staler, which only ever retains more WAL - but it is **not** safe for **boundedness**, and the two were conflated until issue #3310. The shed path is the only one that carries an advancing checkpoint offset, so for as long as it is suppressed the WAL GC offset floor cannot restamp and retained WAL grows without bound. A sustained non-zero rate means the pin store is the bottleneck; pair it with `durable_write_latency` and `shed_stall_seconds`, and consider raising `WalMaterialiserPinBuckets`. Tagged `tree` and `pin_shard`. **`pin_shard` is not `shard`**: it is the durable-pin routing shard (`WalMaterialiserPinShards`), a different hash of a different key from the physical WAL partition tag, and both default to 8. A join between them looks well-formed and means nothing. **Zero-primed**: the WAL GC scheduler emits a single `Add(0)` for each tree on its first pass, so a tree that has never shed reports an explicit `0` rather than no series at all. That primed series carries `tree` and `tenant` but no `pin_shard`, so it sits beside the per-pin-shard series a real shed records rather than priming them - aggregate over `pin_shard` to read the two together - and its `tree` is the registry id the scheduler walks, the physical copy's id included on a resized, restored or remediated tree, while a real shed records under the logical id (see [The `tree` dimension across aliasing](tag-conventions.md#the-tree-dimension-across-aliasing)). Before that priming the healthy case and a dead reporter were both exported as nothing (issue #2694). |
| `orleans.lattice.materialiser.pin.shed_forced` | `Counter<long>` | `{report}` | Steady-state pin reports **forced through** a live shed window because that pin shard had been shedding continuously for longer than `WalMaterialiserPinShedCeiling`. Structurally zero unless that option is set; it defaults to unset (disarmed), so an unconfigured host reports no forcing at all. Every increment is one deliberate override of the issue #2014 back-pressure, so a sustained rate is an alarm rather than routine: it means the pin store has been unable to keep up for longer than the ceiling and WAL retention was being held hostage to it. Forcing cannot overstate durability - `ResolveDurablePinForPartition` clamps the reported offset to `min(checkpoint, covered)` inside the leaf, so a forced report publishes a true value sooner and never a higher one, which is what keeps this fix on the safe side of issue #3300. Tagged `tree` and `pin_shard`. **`pin_shard` is not `shard`**: it is the durable-pin routing shard (`WalMaterialiserPinShards`), a different hash of a different key from the physical WAL partition tag, and both default to 8. A join between them looks well-formed and means nothing. |
| `orleans.lattice.materialiser.pin.shed_stall_seconds` | `ObservableGauge<long>` | `s` | Age, in seconds, of each pin shard's current **unbroken** run of shed reports. It resets to zero the moment any report gets through, and the series is omitted entirely for a shard that is not currently shedding, so an idle host publishes nothing rather than a wall of zeros. This is the instrument that separates a healthy burst of shedding from a latch: `reports_shed` climbs at the same rate under both, whereas a rising `shed_stall_seconds` means the window has never once closed, and so the durable pin offset - and with it the WAL GC offset floor - has not restamped for that long (issue #3310). It is emitted whether or not `WalMaterialiserPinShedCeiling` is armed, which is what keeps the default-disarmed configuration observable instead of silent. Tagged `tree` and `pin_shard`. **`pin_shard` is not `shard`**: it is the durable-pin routing shard (`WalMaterialiserPinShards`), a different hash of a different key from the physical WAL partition tag, and both default to 8. A join between them looks well-formed and means nothing. |

### Drain-lag holder observations

The WAL saturation sampler attributes an over-threshold `orleans.lattice.materialiser.drain_lag` reading through structured Warning logs, never consumer-id metric tags. It emits an observation immediately on the crossing tick and repeats while the tree stays strictly over `WalSaturationMaterialiserLagThreshold`, no more often than [`WalDrainLagHolderLogInterval`](../configuration/options-reference-4.md#waldrainlagholderloginterval) (10 minutes by default; `null` means edge-only). Recovery clears the interval. This is per-tree, per-silo diagnostic state, not a durable assignment of blame; a later snapshot may name different consumers.

Each observation contains up to **three lines**, naming the three eligible consumers with the lowest HLC cursors in ascending order (ties retain snapshot order). All lines share a generated `ObservationId` and the sampler tick's `ObservedAtUtc`. Each carries `TreeId`, `DrainLagSeconds` (the aggregate), `LaggingConsumers` (the full eligible over-threshold count, not capped at three), `HolderRank` (1 is the minimum), `HolderCount`, `ConsumerId`, `Cursor`, `ReportAgeSeconds` and `PositionAgeSeconds`. Group by `ObservationId` rather than counting lines as separate breaches. `ConsumerId` is operator-facing identity: restrict access to these logs as you would other tree diagnostics. No new admin endpoint or metric instrument is involved.

A recent report with an old position age means a consumer is re-reporting without advancing; a recent position age means it has just advanced, even if its cursor still trails the head. `PositionAgeSeconds = null` means no advance was observed since registration, or the registry does not track it. Cold consumers, never-reported cursors, and position-stale leaf materialisers are excluded by the same lag-plane predicate as the histogram; tree-wide tailers remain eligible on report freshness alone. Multiple consumer ids can belong to one leaf (one per WAL partition), so three holders are not necessarily three distinct leaves. The snapshot is read after the aggregate: concurrent progress may change the holders or leave none eligible, in which case no holder is invented or logged. Compare successive observations and the lagging-consumer count before inferring stalled work versus active catch-up.

Selection reuses the snapshot already read in the over-threshold branch. Only a due warning selects the top three, using three local nullable value types and no sorting, LINQ or per-tick holder collection. Monotonic rate-limit state is bounded by over-threshold tree cardinality and removed on recovery, disappearance from the sampler or input disablement. Logging arguments allocate only when an observation is emitted; the leaf-materialiser drain loop is unchanged.

Previous: [Instrument catalog: Process identity to Shard-level and tree registry](instrument-catalog-1.md). Next: [Instrument catalog: Snapshot cursors (sourced from SnapshotLeafGrain / snapshot-cursor open path)](instrument-catalog-3.md). Contents: [Metrics](../metrics.md).
