---
title: "Instrument catalog: WAL garbage collector (sourced from LatticeWalGc) - Metrics"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/metrics/instrument-catalog-4.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/metrics.md?plain=1#L304-L353"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Instrument catalog: WAL garbage collector (sourced from `LatticeWalGc`)

Part of [Instrument catalog](instrument-catalog.md), in [Metrics](../metrics.md).

## WAL garbage collector (sourced from `LatticeWalGc`)

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.wal.entries_trimmed` | `Counter<long>` | `{entry}` | WAL entries removed by the per-tree garbage collector. Tagged `tree`, `shard`, and `tenant`. Trimming is decided per shard, so issue #3206 moved the emission from one per-tree total to one measurement per shard the pass actually scanned - **including a zero for a shard that reclaimed nothing**, which is what makes an absent series mean "this shard was not scanned on this silo" rather than "this shard reclaimed nothing". Read it per shard against `orleans.lattice.wal.compactions` (below): trimming that advances while that same shard's compactions stay flat is dead space stranding, and the tree-scoped sum cannot see it. **That is one of two failure shapes, and this row long described only the first (issue #3258).** It assumes trimming is happening at all. Where the WAL GC floor holder is refused admission the trim loop never runs, so no entry is ever marked dead, `wal.compaction.eval.dead_bytes` reports a measured **zero**, and compaction correctly declines - and that zero is **structurally pinned, not merely low**. The file WAL provider's per-shard dead-byte count is incremented at exactly two sites: inside the shard's trim, and during recovery, when it replays entries at or below a watermark rebuilt solely from trim-marker records whose only writer is that same trim. A shard that has never trimmed therefore reports zero dead at any retention, forever, and **no value of any compaction option can move it** (see issue #3207) - the tree that most needs reclamation presents as the one that least needs it. Measured on a live estate: a wedged tree reported 8 evaluation series summing 3,071,327,085 retained bytes against **0** dead bytes, while a healthy sibling sat at 4.1% dead and the estate's busiest tree at 65.4%. **A dead ratio of 0.0% is the best possible value and the wedged tree scores it**, so dead ratio, compaction count and reclaimed bytes must not be ranked, sorted, coloured or thresholded as health signals. `wal.entries_trimmed` reading a flat zero for a tree's whole lifetime is the lowest arm that is not inverted; `orleans.lattice.wal.gc.floor_holder_admission` is the arm that attributes it. |
| `orleans.lattice.wal.gc.blocked_leaf_reactivations` | `Counter<long>` | `{reactivation}` | Reactivations of a dormant leaf whose unusable durable materialiser pin was blocking its tree's WAL cursor floor (issue #2710 Limitation 2). Tagged `tree` and `outcome`, whose values are a **union of three kinds** rather than a single enumeration (issues #2938, #2692). Four are **lifecycle events**: `attempted` when the scheduler touched a blocking leaf, `healed` when a previously-touched consumer stopped blocking, `abandoned` when a consumer stayed blocked across every permitted attempt of a cycle, and `rearmed` when an abandoned consumer's attempt budget was restored after a backoff (issue #2783). Seven are **terminal outcomes** of a single touch, and every one of them but `admission_refused` partitions `attempted` exactly once: `completed` when the grain call returned, `unresolvable` when the reported blocker could not be resolved to a materialiser consumer so no call was ever issued, `faulted` when the call threw, `undelivered` when the touch was issued but the call did not return within its budget, so the leaf's state is unknown rather than unhealed (issue #2768), and `orphaned` when the leaf reported no bound tree id, proving its state was cleared after the pin was registered, so the pin outlived a reclaimed or purged leaf and the sweep retired it rather than reactivating anything (issue #3101), `latched_stale` when the drive was refused because the leaf is latched stale (issue #3451), which activation provably cannot heal (issue #3478), and `admission_refused` when the leaf's silo refused the drive admission to its WAL replay gate before anything was replayed (issue #3575). `latched_stale` is **terminal for the pin, not only for the touch**: the leaf is driven once and never again for the rest of the blocked episode, including through its sibling partition pins, it is neither abandoned nor re-armed, and its pin is **kept** in the cursor floor so no WAL it still needs is trimmed. The tree is instead reported once per GC pass on a warning log naming how many latched leaves block it, until an operator rebuilds them. The two refunded outcomes are `faulted` and `undelivered`, which is why a tree failing on those costs twice the attempts per abandonment that one failing on `completed` or `unresolvable` does. `admission_refused` is the one terminal outcome that costs the consumer nothing: the refusal is raised before the drive starts, so it says nothing about the leaf, and the touch is **neither charged nor refunded** - it can never lead to `abandoned`, and the consumer is retried after a jittered delay of one to one and a half minutes that doubles with each consecutive refusal, up to the fifteen-minute retry cooldown. It is also the one terminal outcome **outside `attempted`** (issue #3761): a refused try never reached the drive, so `attempted` counts every touch except a refused one, while `admission_refused` counts every refused try on its own. A pass holds no more touches in flight than this silo's GC starvation share and re-drives a refused touch once a sibling touch of the same pass frees a slot, so one consumer can be refused more than once in a pass; divide `attempted` by `orleans.lattice.wal.gc.blocked_consumers` for an honest per-blocker touch rate (issue #2878). Before it had an arm the refusal was counted as `faulted` and logged with a stack, and a consumer the sweep never managed to drive was abandoned with advice about its snapshot capture; on one deployment that was 89% of all touches, because the leaves' own coverage-lag timer drives were taking the gate's GC share first, and a timer drive can no longer take the slot the sweep needs. Seven are **drive verdicts** recording what the touch actually *achieved* - five added by issue #2692, a sixth by issue #3065 and a seventh by issue #3761: `drove_lifted` when the drive replayed the leaf's outstanding WAL, stamped its durable snapshot coverage, and left no partition checkpointed without coverage, so the block pin is genuinely gone (since issue #3599 it is also recorded, with no drive issued and no replay permit spent, when the sweep's permit-free bank step lifted the admitted floor-holding consumer's own durable pin offset, graded on the same offset axis as a drive); `drove_no_advance` when the drive ran and the pin survived it; `drove_memory_refused` when the replay was refused under heap pressure, which is a retry-worthy condition rather than a verdict on the leaf; `drove_not_driven` when the grain declined to drive at all, which on a leaf with no tree id means the blocking report and the grain disagree about what is there; and `drove_already_driving` when a drive was already in flight for that leaf, which is sweep contention rather than leaf starvation; and `drove_timed_out` when the drive exhausted its `StarvationDriveBudget` and abandoned its replay so that it would stop holding a per-silo replay permit (issue #3065); and `drove_admission_refused` when the leaf's silo refused the drive a replay permit before it replayed anything, the routine answer to a bounded background drive, which is counted here and under the `replay_permit_admission` source on `orleans.lattice.saturation.refusals` rather than raised as an exception (issue #3761). **`drove_timed_out` is structurally near-silent in exactly the case it describes** and must not be read as the abandonment signal: the scheduler's touch times out at the Orleans response-timeout default long before the drive's five-minute budget elapses, so the verdict is usually recorded by a grain whose caller has already given up and gone to `undelivered`. The arm exists because folding the verdict into `drove_no_advance` would be a misfiling - a timed-out drive made no statement whatever about the pin, whereas `drove_no_advance` asserts the pin survived a completed replay - and because the scheduler's mapping from drive verdict to tag value throws on an unmapped member rather than silently folding. Read `orleans.lattice.wal.replay.starvation_drive_abandonments` for the abandonment count; it is emitted grain-side and does not depend on the caller still being there. **`drove_no_advance` is the arm that makes the older `attempted`/`healed` pair honest.** Before #2692 the sweep issued a read-only touch and the fixture guarding it asserted only that a call had been issued, so a sweep that reached every blocking leaf and achieved nothing on all of them was indistinguishable from one that was working: `attempted` climbed, `healed` stayed flat, and nothing anywhere said which of the two had happened. Read `drove_lifted` over `drove_lifted + drove_no_advance` as the drive's true success rate, and note that it is a **stronger** statement than `healed`, which is only evidence the sweep is working rather than proof it caused the heal - ordinary traffic could have reactivated the leaf first, whereas a `drove_lifted` is the sweep's own call returning its own verdict. **Sustained `drove_no_advance` is the sharper alarm**: the leaf was reached, its replay ran to completion, and the pin still did not lift, so the residual cause is the snapshot-coverage half rather than reachability. **The `drove_lifted` verdict deliberately requires both halves.** Advancing the checkpoint alone does not lift the pin, because the durable pin for a partition is `min(checkpoint, covered)` and a starved leaf's coverage is `-1` until a snapshot is captured; a drive that claimed `drove_lifted` on the checkpoint alone would release WAL the leaf cannot actually reconstruct from, which is silent data loss that looks exactly like the fix working. The blocked population is **dormant** leaves and that is structural: `LatticeWalGc` skips any consumer present in the live cursor registry before it evaluates the pin, so a blocking consumer is by construction one whose leaf is not activated - which is why no leaf-local driver (the write path, the capture path, the periodic recheck, activation itself) can reach it, and why the retention path has to. All 18 outcome arms share one instrument so that a zero on `healed` is a **measured** zero rather than an unpublished series; that distinction is load-bearing, because a sweep that reactivates a leaf and moves on cannot otherwise tell "the pin lifted" from "the capture failed again", and a deployment can sit indefinitely in the second state. Read `attempted` as the cost the sweep imposes, `healed` over `attempted` as whether it is working, and any sustained `abandoned` as the alarm: the block is not clearable by activation alone, so the leaf's snapshot capture is failing for a separate reason. Distinguish that from `undelivered`, which is the arm that makes the two separable: `abandoned` means the leaf was reached and stayed blocked, whereas `undelivered` means it was never reached at all, so an unreachable leaf no longer reads as an unhealable one. Abandonment is a **pause, not a verdict** (issue #2783): the conditions that make a reactivation futile - memory pressure, a replay gate saturated by an ingest burst, a silo mid-recovery - are transient, so the budget is restored after a backoff that doubles per cycle to a ceiling, and each restoration emits `rearmed`. Read the pair together: `abandoned` and `rearmed` advancing in step is a hopeless tree retrying on an ever-slower schedule, which is the intended shape, whereas `abandoned` advancing while `rearmed` stays flat means the sweep has genuinely stopped. Every arm is **zero-primed once per tree per process**, on the tree's first GC pass and above every early return, so an absent series means the scheduler never evaluated that tree - a statement about the build, not about the events. A flat zero means it did evaluate and that arm's event never occurred, but that reading holds **only because the terminal outcomes and the drive verdicts are exhaustively armed**: priming walks each outcome enum rather than a hand-written list, and every member of both is mapped to an arm through a switch that throws on an unmapped member, so an outcome added later cannot ship unarmed. Before issue #2938 the condition did not hold - three of the then-four terminal outcomes had no arm at all, so their absence was a missing detector reported in exactly the same shape as a measured zero. Priming proves a detector exists for the arms that have one; it says nothing whatever about outcomes that have none, so it is the enum walk, and not the priming, that makes a zero here readable. Exhaustive arming is still only half of the condition: a primed arm whose recording path can never be reached is frozen at zero forever and is indistinguishable, on a scrape and on every enrolment, hygiene, ordering and doc-coverage gate, from an arm that is correct and merely quiet (issue #2942). Each terminal arm therefore carries a positive control that drives the scheduler into that outcome and observes the arm advance, so the zeros it reports elsewhere are earned rather than assumed; the same holds for each of the seven drive verdicts. Both controls are written as identity matrices - 7x7 over the terminal arms and 7x7 over the drive arms - so every off-diagonal zero is proven observable by the diagonal entry in the same matrix, and no zero in either is an unearned one. The leaf identity is carried on the paired warning log rather than as a tag, because the leaf population is unbounded and would be an unbounded metric dimension. Since issue #2815 that log is throttled to one line per blocked **episode** rather than one per change of reported blocker, so on a tree whose blocker churns faster than the minimum block age it names only the first blocker; the distinct-blocker count and the most recent blocker are carried on the escalation and on the end-of-episode summary. Read the count before concluding from a single named leaf that a single leaf was blocking. |
| `orleans.lattice.wal.gc.blocked_consumers` | `ObservableGauge<long>` | `{consumer}` | Latest uncapped count of distinct durable consumers blocking the cursor floor (#3161), tagged `tree` and `tenant`, including offset-plane population gaps once. Denominator for `orleans.lattice.wal.gc.blocked_leaf_reactivations`: together the population trend and heal/drive outcomes expose convergence; neither alone does. Counts consumer ids (partition pins), not leaves. Registry-present pins with real cursors, offset-covered zero pins, and proven-empty abstained partitions do not count. The scan continues beyond the eight-id report cap even when every partition is blocked; it reuses the unioned pin dictionary and memoises empty-WAL probes per partition, without changing trim entitlement. Zero-primed per tree per process before the first await; a completed census replaces the value, while an unreadable census records `-1` (unknown, never a healthy zero). Latest value persists between passes; consult scheduler phase age for stalled collection. Do not sum across silos, which can observe the same durable pins. |
| `orleans.lattice.wal.gc.blocking_pin_state` | `Counter<long>` | `{consumer}` | Durable-pin state of each **absent** consumer blocking a WAL GC pass (issue #3042). Tagged `tree`, `partition` and `status`: `checkpointed_uncovered` (the leaf durably checkpointed that partition but published an unusable pin because snapshot coverage is absent - repairable), `never_checkpointed` (the leaf holds live data it has never checkpointed, so there is no WAL offset it could honestly claim - correct by design, no repair exists), `no_durable_state` (the provider answered and reported nothing ever persisted for that leaf - a fourth state, not a flavour of either), and `unreadable` (the classifier could not answer: an unparseable consumer id, no storage provider on this silo, or a read that threw - a failure of the measurement, never rendered as a finding about the system), and `orphaned` (the leaf's durable state exists but carries no bound tree id, so the leaf was reclaimed or purged after the pin was registered - the pin has outlived its publisher and no amount of driving will lift it, issue #3105), and `checkpointed_coverage_unknown` (the leaf durably checkpointed that partition but its pin is **not** known to be unusable, so whether that checkpoint is covered was never determined and no claim is made, issue #3168). **`checkpointed_uncovered` infers the `uncovered` half rather than measuring it** - coverage is per-activation in-memory state no storage read can reach, and what licenses the inference is knowing independently that the pin is unusable, since the published pin is `min(checkpoint, covered)`. The blocked arm has that premise by construction; the floor-holder sample (issue #3158) does not, because it runs only when the cursor floor reports *usable* and selects by lowest durable **offset** rather than by usability. A usable pin sampled there is `checkpointed_coverage_unknown` and is **not** driven into the coverage repair, which on a live estate had been reactivating a healthy, fully covered, merely idle leaf every pass while that leaf answered `no_checkpointed_uncovered_partition` over four thousand times (issue #3168). It **is** driven for a different reason: a pin on this arm whose durable offset equals the tree's offset floor is reactivated for **liveness** (issue #3178), because a scanned-through projection checkpoint advances only during replay and therefore freezes when its leaf deactivates - so on a converged corpus, where nothing reactivates it, the lowest frozen pin holds the whole tree's WAL indefinitely while no leaf is blocked and no pass fails. Since issue #3310 a pin on this arm sitting *above* the floor is also driven, but only once a pin at the floor has been admitted on the same sweep and only within the tree's remedy candidate budget: those pins are the next floors, in the order they will become the floor, so driving them is prefetch rather than the waste issue #3168 measured. If the floor's own holder is inadmissible the level never drains, nothing above it is ever in the way, and nothing is driven. **`orphaned` did not merely go unreported before that fix - it was reported as `checkpointed_uncovered`**, because the checkpoint classifier read only the persisted checkpoint offset and a husk leaf retains the last offset it wrote. That is why a tree could report thousands of *repairable* pins while every repair attempt answered that there was no checkpointed-uncovered partition to repair. The checkpoint half is read through the same guarded accessor the leaf resolves its own pin from, so an unassigned born-0 checkpoint on partition 0 reads as `never_checkpointed` rather than being misreported as repairable (issues #2703, #3157). Classified by a direct storage-provider read that **never activates the leaf**, because the blocking population is exactly the population that cannot be activated. Charged once per consumer per blocked episode. **It is a cumulative tally of classification events, not a population** - a consumer classified again in a later episode increments it again, so the total only ever grows and cannot answer "how many leaves are blocked now". Read as a population on a live estate it climbed from 281 to 337 on a process that never restarted, which a population on a counter cannot do (issue #3175). Both inputs are capped as well - the blocked arm at the GC pass's report cap of eight blocking consumers, the floor-holder arm at the sweep's own classification cap - so even as a tally it is a **lower bound** on what was blocking, and nothing in the estate counts distinct blocked leaves per tree. All six `status` arms are zero-primed **once per tree per process** under the reserved partition value `none` - latched on the tree's first collection, not repeated per pass - and again per classified partition. That priming establishes the classifier is wired on this silo and nothing more: `Add(0)` is idempotent on a counter, so a primed series can never show the region ran on any particular pass. Use `orleans.lattice.wal.gc.pass.reach` for that (issue #3075). **Reading it: filter `partition!="none"` before aggregating** - the reachability series exists on every tree whether or not anything blocked, so an aggregate that forgets the filter counts every tree twice. A consumer id that cannot be parsed back to a leaf and partition is recorded under the other reserved partition value, `unknown`, and always on the `unreadable` arm. **A non-zero `unreadable` voids the other five `status` arms for that tree**, rather than sitting alongside them: if the classifier could not read, the classifications it did emit for that tree are not trustworthy either, so the six `status` arms are not a partition of one population whenever `unreadable` is non-zero. Diagnostic only: it never changes what a pass is allowed to trim. |
| `orleans.lattice.wal.gc.orphan_pin_sweep` | `Counter<long>` | `{pin}` | Outcome of each durable materialiser pin examined by the WAL GC's **bulk orphan sweep** (issue #3105). Tagged `tree` and `status`, and on its removal-decision arms also `cause` (issue #4246); the `status` arms **partition the examined population**, so `sum by (tree)` over one sweep is exactly the number of durable pins that tree holds - the only place that population is visible at all, because the cursor floor's blocking report is capped at `MaxReportedBlockingConsumers` and shows at most eight. `retired` is a pin whose leaf read as gone and which the sweep removed from every shard key it was found under; each removal also writes one audit log line naming the tree, the consumer id and the cause, bounded to the first 32 per tree per pass with the remainder reported as a suppressed count. `retire_failed` is such a pin whose removal did not complete on every key, so it stays registered and is re-examined next sweep. `deferred` is such a pin left in place because the pass's retirement budget was already spent. `refused_malformed_id` and `refused_ambiguous_partition` are such a pin the fail-safe gate declined to remove because its consumer id is not provably the leaf's own (issue #4238): the first for another tree's prefix, an id that is not a grain id, or a leaf key that is not a canonical guid; the second for a partition that cannot be read unambiguously against the registry-pinned count - out of range, non-canonical, missing on a partitioned tree, or present on a tree read as single-partition, which is where the #4238 misresolution lands. Those five removal-decision arms carry `cause`: `orphaned` when the leaf's record exists with no bound tree id (a reclaimed or purged leaf, the expected source of `retired`), and `no_durable_state` when there is no record at all. **The two are never folded (issue #4246)**: `no_durable_state` is benign when the leaf really is gone and is also exactly what a live leaf misresolved to a grain that never existed reads as, so before the split a misresolution and healthy reclaim exported byte-identical series. A non-zero `refused_*` arm, or a `retired` with `cause="no_durable_state"` on a tree where none was expected, is the series to investigate. `live` is a pin whose leaf still carries a bound tree id, which the sweep never touches - a live leaf's pin is retention the WAL GC is obliged to honour. `unresolved` is a consumer id that does not parse back to a leaf grain id, so no leaf could be read for it and no claim was made. `unreadable` is a leaf-state read that failed, and the sweep **fails closed** on it: an unreadable leaf is never retired, because retiring a pin whose leaf might still be live would authorise a trim over a prefix that leaf has not replayed. **`deferred` is the backlog signal, and it is why this is a counter rather than a gauge**: a sustained non-zero `deferred` rate means orphans remain and the sweep is still draining them, while `deferred` at zero with `retired` no longer advancing means the backlog is gone. That reads the operational question off an advancing series and carries none of the declaration-order hazards an observable gauge would. The absence of any such signal is what let issue #3105 run for days undetected: the only orphan series that existed was the reactivation sweep's monotonic `orphaned` arm, on which a 63-pin-per-hour drain against a 9,468-pin backlog is indistinguishable from steady healthy progress. Every arm is zero-primed per tree, and the removal-decision arms under both causes, so a zero here is a measured zero rather than an absence. |
| `orleans.lattice.wal.gc.drive_orphan_pin_retirement` | `Counter<long>` | `{pin}` | Every orphan-pin removal decision the WAL GC blocked-leaf **drive** takes when a leaf reports `NotDriven` (issue #4246) - the second route that deletes a durable materialiser pin, beside the bulk sweep (`orleans.lattice.wal.gc.orphan_pin_sweep`). Tagged `tree`, `status`, `cause` and `tenant`. The `status` arms partition every decision: `retired` removed the pin and wrote an audit log line naming the tree, the consumer id and the cause; `retire_failed` attempted the removal and it threw, so the pin stays registered; `refused_malformed_id` and `refused_ambiguous_partition` are the fail-safe gate declining because the consumer id is not provably the driven leaf's own (issue #4238), so the pin is left holding the floor. `cause` is always `not_driven`: an activated leaf cannot tell a record with no tree id from no record, so the drive's evidence is its own cause rather than either of the sweep's. The `orphaned` arm of `orleans.lattice.wal.gc.blocked_leaf_reactivations` records the leaf's verdict, not what was done about the pin, which is why this instrument exists. Every arm is zero-primed per tree on the tree's first GC pass. |
| `orleans.lattice.wal.gc.floor_holder_classification` | `Counter<long>` | `{pin}` | How much of a tree's durable materialiser pin population the WAL GC **floor-holder classification** examined on a sweep (issue #3158). Tagged `tree` and `status`, whose two arms **partition the enumerated population**, so `sum by (tree)` over one sweep is the tree's whole durable pin count and `classified / sum` is the coverage fraction outright. `classified` is a pin whose leaf state was read and whose result was recorded on `orleans.lattice.wal.gc.blocking_pin_state`; `unclassified` is a pin the sample did not reach. **This is the denominator that instrument never carried**, and its absence is the whole reason issue #3158 stayed invisible: five zero `status` arms on a tree nothing ever classified are byte-identical to five measured zeroes on a tree that was classified and found nothing, so a live scrape showing `blocking_pin_state` all-zero under partition `none` was read as "no pin is in a notable state" when it meant "no pin was ever looked at". That was on the byte-ceiling tree - the one tree the classifier exists to diagnose, and the one arm it could not reach, because the classifier's single call site sat on the floor-blocked heal path and its input was derived from a floor-blocked report an `over_ceiling` tree does not produce. **A very small `classified` fraction is the designed behaviour, not an alarm.** Each classification is a durable storage read and the population is unbounded - one live tree was measured at 52,224 pins on a single sweep - so the sample is capped per sweep by a constant independent of the population, and a tree holding tens of thousands of pins is expected to report a handful `classified` against the rest `unclassified`. The sample is not arbitrary, and **issue #3178 corrected which axis it is taken on**: the durable materialiser offset floor is a *minimum* over the pins that **reported** an offset, so the pins carrying the lowest offset are the ones holding it, and those are the ones sampled, in ascending offset order, with the frontier as the secondary key and ties broken on the consumer id so the sample is stable across sweeps rather than churning with dictionary order on a tree whose pins all sit at `HybridLogicalClock.Zero`. It previously ranked on the frontier alone - the floor's *other* axis, folded into the HLC cursor floor rather than into the offset floor - so on a tree whose every trim stop was `offset_floor`, the sample confidently described pins that were not the blocker. Pins reporting no offset constrain no offset floor at all and are sampled into a separate list by frontier, so the thousands of them a never-trimmed tree carries cannot evict the pins that do. The sample is also **deduplicated by leaf**: a leaf publishes one pin per WAL partition and they carry a byte-identical frontier, so an undeduplicated sample of eight on an eight-partition tree is one leaf described eight times. Both arms are zero-primed per tree at the top of the tree's collection, above every early return, so an absent series means the classification is not wired on this silo rather than that it found nothing. Diagnostic only: it never changes what a pass is allowed to trim. |
| `orleans.lattice.wal.gc.coverage_unknown_pin_offset` | `Counter<long>` | `{pin}` | Whether a floor-holding pin classified `checkpointed_coverage_unknown` carried a usable durable offset (issue #3199). Tagged `tree`, `partition` and `status`, whose two arms **partition this instrument's population exactly**, since a candidate's durable offset is either negative or it is not. `offset_usable` is a candidate whose offset is `>= 0`, which is the pin population the offset floor is a minimum over and the only one issue #3178's liveness drive can admit - it is driven when its offset equals that floor, or (since issue #3310) when it sits above a floor whose own holder was admitted on the same sweep, within the tree's remedy candidate budget. `offset_absent` is a candidate that reported no usable offset, which constrains no offset floor and can therefore never satisfy that gate. **That distinction is the whole of issue #3199**, because such a pin is excluded from the blocked arm (its frontier is above `HybridLogicalClock.Zero`), from the issue #3164 coverage repair (the frontier gate promoted its state to `checkpointed_coverage_unknown`), and from the issue #3178 liveness drive (it carries no offset) - and the durable pin store merges monotonic-max on **both** axes, so neither exclusion can be cleared by anything the leaf subsequently does. **`orleans.lattice.wal.gc.blocking_pin_state` cannot answer this and no aggregation of it can**: it records the classifier's verdict and nothing about the candidate that produced it, and both floor-holder sample lists - the offset-bearing one and the one holding pins that reported no offset - are recorded through the same call, so the two cases are byte-identical there. The bit is not derived here either; the sweep already computes it when it routes a candidate into one of those two lists, and until this instrument existed it was computed and then discarded. **The population is every observation of the `checkpointed_coverage_unknown` arm anywhere**, because the promotion that mints that state exists only on the floor-holder path while the blocked arm records the classifier's verdict unpromoted - so `sum by (tree)` here equals `blocking_pin_state{status="checkpointed_coverage_unknown"}` on the same tree, and a divergence is a defect in one of the two rather than a reading. Both arms are zero-primed per classified `(tree, partition)`, alongside the `blocking_pin_state` priming and independently of which state resolved, so `offset_absent` reading zero against a large `offset_usable` is a **measured absence** rather than silence. That discrimination is the entire value of the instrument: issue #3199 turns on whether the `offset_absent` slice is empty, and an unprimed zero could not answer it in either direction. It carries no per-tree reachability priming of its own - whether the floor-holder classifier is wired on this silo is already answered by `orleans.lattice.wal.gc.floor_holder_classification`, and duplicating that latch here would add a second thing to keep true without adding a discrimination. Diagnostic only: it never changes what a pass is allowed to trim. |
| `orleans.lattice.wal.gc.never_checkpointed_pin_offset` | `Counter<long>` | `{pin}` | Whether a floor-holding pin classified `never_checkpointed` carried a usable durable offset (issue #4198). Tagged `tree`, `partition`, `status` and `tenant`, whose two arms **partition this instrument's population exactly**, since a candidate's durable offset is either negative or it is not. **The state alone is ambiguous, and the ambiguity is load-bearing**: `never_checkpointed` is derived from the leaf's *persisted* checkpoint alone and constrains the published pin offset not at all, so one arm of `orleans.lattice.wal.gc.blocking_pin_state` covers two populations with opposite meanings. `offset_absent` is the blocking sentinel at offset `-1` - benign, ubiquitous, routed into the floor-holder sweep's unusable sample where it constrains no offset floor, and cleared as soon as the leaf checkpoints. `offset_usable` carries a non-negative published offset, so it sits in the offset-bearing sample and **can hold the tree's offset floor**, where it is refused admission to the issue #3178 liveness drive - and that refusal is **terminal for the whole tree**, because the floor is defined by its own holder, every other candidate is strictly above it, and issue #3310's prefetch is gated on the floor's holder having been admitted first. That is the permanent-wedge shape diagnosed in issue #3258. **`orleans.lattice.wal.gc.blocking_pin_state` cannot answer this and no aggregation of it can**: it records the classifier's verdict and nothing about the candidate that produced it, and both floor-holder sample lists are recorded through the same call, so the two readings are byte-identical there. The bit is not derived here either; the sweep already computes it when it routes a candidate into one of those lists, and until this instrument existed it was computed and then discarded. It mirrors `orleans.lattice.wal.gc.coverage_unknown_pin_offset` exactly and reuses its `offset_usable` / `offset_absent` vocabulary, so a reader who knows one needs nothing new to read the other. `sum by (tree)` here equals `blocking_pin_state{status="never_checkpointed"}` on the same tree only while every such observation came from the floor-holder classifier, and a divergence then is a defect in one of the two rather than a reading. The blocked arm also records `never_checkpointed` - for an absent consumer blocking a pass - and charges nothing here, so on a tree whose blocked arm has classified such a consumer the `blocking_pin_state` side is larger by exactly those classifications. Do not write the comparison naively either, because `blocking_pin_state` is also primed per collected tree at `partition="none"`, a label this instrument never mints, so sum over `partition` on both sides, and a tree that never reached the floor-holder classifier has no series here at all rather than a zero. Read a sustained non-zero `offset_usable` beside `orleans.lattice.wal.gc.floor_holder_admission`: the two together separate a latent holder sitting harmlessly above the floor from one actually holding it and blocking the tree. Both arms are zero-primed per classified `(tree, partition)`, independently of which state resolved, so `offset_usable` reading zero is a **measured absence** rather than silence - the entire value of the instrument, since the question is whether that slice is empty and an unprimed zero could not answer it in either direction. Diagnostic only: it never changes what a pass is allowed to trim, and never changes admission - the refusal it makes visible is correct, and driving a leaf with no proven durable checkpoint toward a durable claim is silent data loss. |
| `orleans.lattice.wal.gc.floor_holder_admission` | `Counter<long>` | `{sweep}` | Whether the candidate that **defines** a tree's durable materialiser offset floor was admitted for repair on a classifying sweep (issue #3258). Tagged `tree` and `status`, whose two arms are mutually exclusive and are charged at most once per sweep, because a tree has one offset floor per sweep however many pins sit on it. `admitted` is a sweep whose floor-holding candidate cleared the admission gate and was handed to the issue #3178 liveness drive; `blocked` is a sweep whose floor holder was refused. **It is the floor's own holder that decides whether any repair is possible at all.** The sample's offset list is ascending, so its head is the offset floor outright and every other candidate is strictly above it; the gate admits a `checkpointed_coverage_unknown` candidate at that floor, and since issue #3310 also above it but only once one AT the floor has been admitted on the same sweep, so if the floor's holder is refused then no candidate on the tree can ever be admitted, the tree is dropped from the repairable set, and the drive is never entered. That precondition is what preserves this property through the issue #3310 widening rather than voiding it. **No other instrument separates that from a tree with no repair to do**: `orleans.lattice.wal.gc.blocked_leaf_reactivations` reads zero on both, and on the live estate that ambiguity hid a tree at 26 stranded passes and 0 reclaimed beside siblings at 45 to 50 attempts. **State the retention correctly**: that tree's `orleans.lattice.wal.entries_trimmed` is zero lifetime, so its ~91 MB is permanently unreleasable and every future write adds to it permanently - but it is not necessarily growing at any moment, and was measured flat for 5.5 minutes of an 8-minute window. **Do not alert on bytes or on growth stopping**, which reads green on a wedged tree; alert on the `blocked` arm here and corroborate with `entries_trimmed` at zero. Both arms are zero-primed per tree on every sweep that reaches the classifier, **before** either is charged, which is what makes three readings distinguishable rather than two: an absent series means the classifier never ran on this tree, both arms present and static at zero mean it ran and found no pin constraining an offset floor, and a climbing `blocked` arm means it found one and refused it on every sweep. The resolved arm is charged **only when an offset floor exists**, since with no floor there is no holder to admit or refuse and charging either arm would report absence of work as a verdict. It reports admission and deliberately not outcome: folding the drive's result in would merge "never admitted" with "admitted but unhealed", and it is the former that has no witness today. Diagnostic only: it never changes what a pass is allowed to trim. |
| `orleans.lattice.wal.gc.floor_holder_offset_admission` | `Counter<long>` | `{candidate}` | How WIDE a tree's floor-holder admission was on the offset axis, per classifying sweep (issue #3310). Tagged `tree` and `status`. `at_floor` counts candidates admitted because their durable checkpoint offset equals the tree's offset floor - the population that was admissible before issue #3310. `above_floor` counts candidates admitted because they sit above a floor whose OWN holder was admitted on the same sweep - the population issue #3310 added, and **the only arm that distinguishes a working widening from an inert one**. Unlike `floor_holder_admission`, which is a per-sweep yes/no about the single candidate defining the floor and is blind to width by construction, this counts every admitted candidate, so it is the series that answers "how many levels can this tree's floor walk per sweep". **Why width is the thing to watch**: the admitted set used to be "every pin sitting on exactly one offset", so its size was a function of pin offset *distribution* and of nothing an operator could configure. Measured on one live silo, same binary, same sweep: a tree whose pins shared a single offset admitted **166** per sweep, while a tree with nine distinct offsets admitted **1 to 3** and drove 0.185 leaves per minute against a touch budget of 32 - 1.4% of the candidate budget issue #3279 had already granted it - and grew its WAL to **120.8%** of its ceiling with zero decreasing intervals over sixty minutes. Spread is the normal steady state of a tree under continuous ingest, whose leaves checkpoint at their own offsets rather than as a bulk-written cohort, so the gate starved exactly the trees that needed it most. **Read it against `floor_holder_admission`**: that series going `admitted` while `above_floor` here stays at zero on a tree with spread offsets means the widening is present but contributing nothing. Both arms are zero-primed per tree on every sweep that reaches the classifier, before either is charged, so an absent series (classifier never ran here) is distinguishable from a measured zero (it ran and admitted nothing on this axis). Admission is bounded by a per-sweep budget scaled from the tree's floor-holding pin population, so a wide tree cannot spend an unbounded sweep. Diagnostic only: admission grants no trim entitlement, and the published pin is still resolved as the minimum of checkpoint and covered offset. |
| `orleans.lattice.wal.gc.pass.reach` | `Counter<long>` | `{pass}` | How far each WAL GC scheduling pass got (issue #3075), tagged `stage`, plus the reserved `tree` value `_none_` and the platform `tenant` sentinel. The scheduler-wide half of the reachability layer for every instrument sited inside the region that stops executing: `orleans.lattice.wal.gc.passes` and `orleans.lattice.wal.gc.interval` are both recorded inside the per-tree collection, so any exit taken above them leaves both silent, and a series that exists and does not advance is byte-identical to one being measured as zero. **Every arm advances by one rather than priming a zero, and that is the point.** `Add(0)` is idempotent on a counter's exported value, so a primed series proves only that a region was reached *at least once* and can never show it was reached on this pass; an advancing arm gives a three-way reading instead of a two-way one - a non-zero rate proves the region is executing **now**, a zero rate on a present series proves it **stopped**, and an absent series proves the build is not deployed. Eight arms, one per region: `pass_entered` plus the seven terminating exits `registry_cancelled`, `registry_failed`, `registry_timed_out`, `loop_cancelled`, `no_due_tree`, `pass_completed_immediate` and `pass_completed_scheduled`. `registry_timed_out` is the scheduler's own enumeration budget firing and is kept apart from `registry_failed` for the reason the scheduler keeps the two `catch` arms apart: a fault is a property of the registry and a timeout is a property of the bound we chose, so merging them would let a decision of ours present as a finding about the system. It was **missing from this layer as first shipped** - the `catch` it accounts for was inserted ahead of the general fault arm by a later change, and an inserted exit is invisible to every detector that reasons about the arms already present. **Why this is a separate instrument from `orleans.lattice.wal.gc.tree.reach` rather than ten stages of one:** a pass spans the entire registry, so it belongs to no tenant and carries the platform sentinel, while a tree visit belongs to the tenant owning the tree exactly as `wal.gc.passes` and `wal.gc.interval` do. One instrument carrying both would split its own series across two attribution rules, so a tenant-scoped query would silently drop the pass-level arms - and those are precisely the arms that answer whether the scheduler ran at all. Deriving a tenant for a pass instead would satisfy the same constraint by inventing an attribution that does not exist. The reserved `tree` value is structurally required for the same reason: two of the exits are the `catch` arms of the tree-registry enumeration itself, where obtaining the tree list is the operation that failed, so no tree id exists to label them with and none ever can. **Reading it: do not filter this instrument by `tree` or `tenant`** - either clause hides every arm. `pass_entered` accounts for the seven exit arms, but it is taken before the work and an exit arm after it, so the correct relation is `0 <= pass_entered - sum(exit arms) <= passes in flight` (one, for a single scheduler loop); **assert equality only at quiescence and alert on the inequality**, or it flaps once per pass forever and gets muted, at which point the completeness guard is gone and nothing says so. That relation is what catches an early return added without an arm - a structural check over the exit *population* rather than a count of occurrences, which reported four exits where the source has nine. Diagnostic only: it never changes what a pass is allowed to trim. |
| `orleans.lattice.wal.gc.tree.reach` | `Counter<long>` | `{tree}` | How far each WAL GC scheduling pass got **for one tree** (issue #3075), tagged `tree`, `stage` and `tenant`. The per-tree half of the reachability layer, tenant-derived exactly like its siblings `orleans.lattice.wal.gc.passes` and `orleans.lattice.wal.gc.interval`; the scheduler-wide half, which belongs to no tenant, is the separate `orleans.lattice.wal.gc.pass.reach`. Two arms, and they are **a pair - neither half is interpretable alone**. `tree_seen` is recorded for every tree the pass enumerated, whether or not it was due; `tree_collected` is recorded above every exit of the per-tree collection, and is the arm that licenses reading a flat `wal.gc.interval` or `wal.gc.passes` **for that tree** as measured rather than as never-executed. Their difference is the set of trees skipped because the adaptive interval had not elapsed, which is *expected to be large and non-zero* in a healthy steady state - on any given pass most registered trees are not yet due - so a large gap is the normal reading and a zero gap on a many-tree silo is the anomaly. Read alone, `tree_collected` understates coverage and `tree_seen` overstates it; record and read both. **Both arms advance by one rather than priming a zero**, for the reason given on `wal.gc.pass.reach`: `Add(0)` is idempotent on a counter, so priming can establish that a region was reached at least once and nothing more. Diagnostic only: it never changes what a pass is allowed to trim. |
| `orleans.lattice.wal.gc.passes` | `Counter<long>` | `{pass}` | WAL garbage-collection passes driven by the per-silo scheduler. Emitted once per pass, unconditionally. Tagged `tree` and `outcome`: `reclaimed` (the pass trimmed at least one entry), `no_partitions` (no WAL partition's pinned provider resolved on this silo, so the pass visited none - see below), `blocked` (the pass reclaimed nothing because an unusable durable materialiser pin disabled the consumer-cursor branch for at least one WAL partition, so that partition cannot reclaim at all and its WAL is growing without bound), `no_consumer` (the pass reclaimed nothing because no consumer has reported a cursor, so there was no frontier to trim against), `idle` (the pass evaluated a usable cursor floor and found nothing beneath it, with the tree inside its byte ceiling or no ceiling configured), `over_ceiling` (the pass evaluated a usable cursor floor and found nothing beneath it, yet the tree's WAL occupancy - dead bytes not yet compacted included - is **still over** its configured `LatticeOptions.WalMaxRetainedBytes`; definite about that condition and advisory about its cause, which may be a safe trim frontier pinned below bytes the policy wants back or dead bytes awaiting compaction - `orleans.lattice.wal.gc.trim_stop` names where the scan stopped), `stranded` (the pass evaluated a usable cursor floor, trimmed nothing, and its trim scan stopped on WAL it had to retain, with no configured byte ceiling complaining about it - the "could not do anything" case, measured with no byte accounting and no opt-in), `unclassified` (the pass reached a cursor-floor state this build does not name), or `failed` (the pass threw). Reclaimed over total is the per-tree reclaim rate; `failed` isolates a wedged tree; `blocked` isolates a tree whose unusable pin disabled reclamation outright; `no_consumer` isolates a tree nobody is reading. **Only `reclaimed` is an affirmative reading. Every other arm means nothing was trimmed**, so `blocked = 0` on its own is not evidence of reclamation - it says one named predicate did not fire, which a tree parked forever on `no_consumer` also satisfies. `no_consumer` and `unclassified` were split out of `idle` by issue #2850 for exactly that reason: until then `idle` was a catch-all, and a tree whose WAL could not shrink for want of a frontier reported the same shape as a healthy quiet one. `over_ceiling` was split out by issue #3119 for the same reason one level further on: a tree breaching the operator's byte ceiling reports `Available` and trims nothing, which is byte for byte a quiet tree's reading, so the arm an operator reads as "nothing to do" was the arm a starved tree landed in. That tag change travels with a scheduling change - a breaching tree now holds the cadence floor instead of relaxing toward the ceiling, because when the safe trim frontier is pinned, pass frequency is the only lever the byte-pressure policy has left. `stranded` was split out by issue #3213 to finish that job on a deployment that configured nothing. `over_ceiling` can fire only where `WalMaxRetainedBytes` is set, and that option has no default, so on a stock silo a tree whose whole retained range sat behind a pinned trim frontier still reported `idle` - the arm that asserts health - for as long as it stayed stranded. The evidence this arm reads is the trim scan's own stop reason rather than a byte sample: the scan halts at the first entry it may not trim, so a pass that trimmed nothing and stopped on anything but an exhausted or an empty log left WAL behind by construction. That costs nothing extra and is never gated behind an opt-in, which is precisely what `over_ceiling` could not offer. It too travels with a scheduling change, and a different one: a stranded tree keeps the geometric backoff but loses its terminal interval, relaxing only as far as the fault-retry ceiling, because the stranded population is far larger than the blocked or the breaching one and pinning all of it to the cadence floor would reintroduce the poll storm the backoff exists to prevent. `over_ceiling` outranks `stranded` wherever both hold, so a breaching tree keeps the more specific diagnosis and the arms stay mutually exclusive: a pass records a single arm, so the arms partition invocations rather than merely covering them, and their sum is exactly the pass count rather than an over-count. `unclassified` is unreachable against today's `WalGcCursorFloorState` and is expected to read a permanent measured zero; it exists so that a floor state added later lands somewhere a reader can see rather than being absorbed by the arm that means "healthy and quiet". Note also that `reclaimed` outranks `blocked`: since issue #2849 the block is scoped to the WAL partitions an unusable pin actually covers, so a partially blocked tree can trim its healthy partitions and report `reclaimed` on the same pass. The scheduler's blocked-leaf remedy and its cadence floor read the report's cursor-floor state directly rather than this tag, so neither is weakened by that. Every outcome arm - `reclaimed`, `idle`, `over_ceiling`, `stranded`, `blocked`, `no_consumer`, `no_partitions`, `unclassified`, and `failed` - is primed at zero for every collected tree (issue #2774), so an absent arm means "this silo is not reporting" while a flat zero means "measured, never happened". That distinction is load-bearing on `reclaimed`, which backs a release acceptance criterion: unprimed, a system that reclaimed perfectly and one that never ran produced the same absent series, so a success was indistinguishable from a failure. Priming also removes a reader's need to know that the non-`failed` arms share one emission site, which was previously the only thing making a sibling arm's presence proof that the instrument was wired at all. `no_partitions` means no pinned WAL provider resolved on this silo (issue #2465), not an empty WAL. It outranks the other non-reclaiming arms and is zero-primed per tree. Inspect WAL placement and provider registration. `ShardsScanned` counts only resolved partitions visited for trimming or compaction; partial resolution keeps the existing outcome and does not imply every partition was examined. With no cursor and no TTL, `no_consumer` names the no-predicate return; `blocked` names an unusable durable pin, while `stranded` can describe a legitimate lagging consumer holding retained WAL. |
| `orleans.lattice.wal.gc.interval` | `Histogram<double>` | `s` | Adaptive WAL garbage-collection interval the scheduler selected for the tree after its most recent pass, inside the configured band `[LatticeOptions.WalGcMinInterval, LatticeOptions.WalGcInterval]`. A sustained reading at the floor is a tree whose log grows faster than one pass reclaims; a reading at the ceiling is a quiet tree. Tagged `tree`. |
| `orleans.lattice.wal.gc.scheduler_backoff` | `Histogram<double>` | `s` | Scheduler-wide WAL garbage-collection backoff currently in force, recorded once per scheduling pass (issue #3064). Distinct from `orleans.lattice.wal.gc.interval`, which is one **tree's** adaptive cadence; this is the whole scheduler's retry wait after a pass that collected nothing, and it is silo-scoped, so it carries `tenant` at the platform sentinel and no `tree`. Tagged `cause` with three arms: `scheduled` (the pass enumerated the registry and selected an ordinary cadence - no backoff is in force, and the floor is reported), `faulted` (the enumeration threw, so the pass learned nothing), and `empty` (the enumeration **succeeded** and reported no collectable tree - a correct observation of an idle silo, and deliberately not an alarm). **This instrument exists because a WAL GC sweep that has backed off and one that is dead emit byte-identical scrapes.** A pass that collects nothing writes no per-tree series at all, so every WAL GC series simply freezes in both cases; a live container was scraped five times over twenty minutes and recorded not one `wal.gc.interval` while leaf activation was demonstrably healthy, and the scrape could not settle which had happened. The two ladders have **disjoint ranges above the fault ceiling**: `faulted` relaxes only to the scheduler's reactivation block age (five minutes at stock defaults, clamped into the configured band) so the window in which a recovered silo still looks dead is bounded, while `empty` relaxes to the full `LatticeOptions.WalGcInterval`. A high reading is therefore self-identifying even before the tag is read. **Recorded unconditionally on every pass and every path** - including the first pass, and including a silo with no trees at all - which is load-bearing: primed on the fault path instead, an absent series would mean either "nothing has gone wrong" or "this build is not deployed", and this investigation confused exactly those two more than once. It cannot share `wal.gc.interval`'s site either, because that one is per-tree and so is silent on precisely the zero-tree silo that most needs a liveness witness. Read an absent series as a statement about the deployment, never about the system's health. Tagged `cause` and `tenant`. |
| `orleans.lattice.wal.gc.scheduler_consecutive_faults` | `Histogram<long>` | `{fault}` | Consecutive failed WAL garbage-collection registry enumerations, reset to zero by the first pass whose enumeration succeeded (issue #3064). A **current streak** rather than a lifetime total, which is why it is a histogram and not a counter - a counter cannot go back down. Recorded beside `orleans.lattice.wal.gc.scheduler_backoff` at the same unconditional per-pass site and carrying the same `cause` tag, so the same priming argument applies verbatim: a zero here is a measured zero, and an absent series is a statement about the build rather than about the registry. A streak that keeps returning to zero is transient fault absorption and needs no action; one that only climbs is a registry the scheduler cannot read at all, and it is operator-actionable - the enumeration is a full key-range scan across every shard of the tree registry, so it is the first thing to fail when leaf activation is saturated. Tagged `cause` and `tenant`. |
| `orleans.lattice.wal.gc.backlog_bytes` | `Histogram<long>` | `By` | WAL bytes the tree still occupies after a garbage-collection pass - its physical occupancy, dead but not yet compacted bytes included, falling back to the retained byte count on a provider without physical accounting - sampled from the pass's own report so it costs no extra I/O. Emitted whenever the pass took a byte sample and the WAL provider supports byte accounting. A pass samples when the byte-pressure policy (`LatticeOptions.WalMaxRetainedBytes`) or the durability hold (`LatticeOptions.WalDurabilityHoldCeilingBytes`, on by default) has a positive ceiling, so with the default options every pass on a provider that accounts bytes emits it. When it is not emitted, `orleans.lattice.wal.gc.backlog_bytes_unavailable` names the reason on the same pass, so "not measured" is an explicit positive signal rather than something a reader has to infer from the absence of samples (issue #2694). Tagged `tree`. |
| `orleans.lattice.wal.gc.backlog_bytes_unavailable` | `Counter<long>` | `{pass}` | WAL GC passes that could not sample `orleans.lattice.wal.gc.backlog_bytes` because no byte figure was available. Tagged `tree` and `reason`: `policy_disabled` when no byte-pressure ceiling (`LatticeOptions.WalMaxRetainedBytes`) is set, or `provider_unsupported` when one is set but the `IWalStorageProvider` reports no byte accounting. The reason is decided on the byte-pressure ceiling alone, while the byte sample is also taken for the durability hold (`LatticeOptions.WalDurabilityHoldCeilingBytes`, on by default): with the default options a provider that accounts bytes never advances this counter, and one that cannot is reported as `policy_disabled`, so that arm means byte sampling was switched off only when the durability hold is off as well. It exists because a silent backlog histogram is the same shape whether sampling is off, the provider cannot count, or the collector is wedged. This counter separates the first two from the third and carries the reason as a dimension rather than leaving it to be inferred (issue #2694). |
| `orleans.lattice.wal.gc.ceiling_unsatisfiable` | `Counter<long>` | `{pass}` | WAL GC passes that found the tree's configured `LatticeOptions.WalMaxRetainedBytes` **arithmetically unreachable** against the live set the pass just measured - below `LatticeOptions.WalMaxRetainedBytesWorkingSetMultiple` (2) times the tree's logical retained payload (issue #3242). **It names a configuration fault, not a lag, and that is why it is a separate instrument rather than another arm of `orleans.lattice.wal.gc.passes`.** Those arms partition invocations - exactly one is recorded per pass, so their sum can never over-count - and an unsatisfiable ceiling is not a pass outcome: it co-occurs with `reclaimed`, `over_ceiling` and `stranded` alike, so an arm would have had to take invocations away from whichever arm names them today, and an operator alerting on `stranded` would have watched it fall silent at the moment the condition worsened. It would also have been unreachable where it matters most, because the pass classifier consults the byte and backlog verdicts only when the cursor floor is `Available`, so a blocked tree would never have reached it. **The two conditions demand opposite responses:** `stranded` and `over_ceiling` say bytes could not be reclaimed *now*, which is answered by unblocking a consumer; this says no pass could ever bring the tree inside its ceiling, which is answered by raising the ceiling or shrinking the tree. **Why the multiple:** a log-structured provider reclaims dead bytes only by rewriting a segment, and rewrites once dead bytes reach a configured fraction `t` of total payload, so designed steady-state occupancy peaks at `live / (1 - t)` - twice the live set at the file provider's default 0.5 ratio - and the ceiling has been compared against *physical* occupancy since issue #3107. A ceiling below that multiple is breached by a perfectly healthy tree, and because `LatticeOptions.WalBytePressureReclaimTarget` (0.8) puts the disarm point below the natural floor of the same compaction cycle, it is breached **permanently**: the advisory byte-pressure alarm arms and can never clear, losing exactly the pathological-growth-versus-normal-size distinction it exists to draw. **It is evaluated against the logical total rather than the occupancy figure the ceiling actually bounds, and that is not a detail.** Compaction fires on a long period, one shard at a time, so occupancy is a sawtooth rather than a level: a check against it would read whatever phase of the cycle the pass landed in, condemning a reachable ceiling at the peak and clearing an unreachable one at the trough - both for the same tree under the same unchanged configuration, with the gap between the two quantities routinely a large fraction of the live set on a real deployment. The ceiling has to clear the *peak* of that sawtooth, and only the logical total predicts where the peak is. **The instrument is justified by the condition's silence rather than by any deployment currently exhibiting it:** before this series existed, an unsatisfiable ceiling was indistinguishable from a lagging consumer and could persist indefinitely while looking like a transient. Zero-primed per tree beside the pass-outcome arms, so a flat zero is a measurement rather than silence - but read that zero as "not provably unsatisfiable" rather than "comfortably sized", since it is also what a tree reads when no ceiling is configured, when the provider cannot account logical bytes, and when the tree is empty. Recorded once per pass, so the rate is readable against `orleans.lattice.wal.gc.passes`. Tagged `tree`. |
| `orleans.lattice.wal.gc.terminal_breach` | `Counter<long>` | `{pass}` | WAL GC passes on which a tree has been **over its byte ceiling, with an available cursor floor, and reclaimed nothing** - exactly the passes `orleans.lattice.wal.gc.passes` classifies `over_ceiling` - for 10 consecutive passes (issue #3149). `over_ceiling` on `orleans.lattice.wal.gc.passes` is by design a transient - the scheduler holds such a tree at the floor interval and asks again - and it reads the same on the tenth fruitless pass as on the first, so a retry loop that has stopped helping was indistinguishable from one about to succeed. This series is what says the retries have stopped helping. Recorded once per qualifying pass from the tenth onwards, so the rate is readable against `orleans.lattice.wal.gc.passes`; any completed pass that reclaims, falls inside the ceiling, or reports a cursor floor that is not available (the `blocked` or `no_consumer` arm) resets the run, while a pass that throws (the `failed` arm) neither advances nor resets it. An unavailable floor is excluded because the `blocked` and `no_consumer` arms already name that tree and their remedies differ. The run is held in memory, so a silo restart delays the signal by at most one threshold's worth of passes and never fabricates it. Read it with `orleans.lattice.wal.gc.floor_head_distance`, which says how much of each shard the floor is holding, and `orleans.lattice.wal.gc.trim_stop`, which names what holds it. Zero-primed per tree beside the pass-outcome arms. Tagged `tree` and `tenant`. |
| `orleans.lattice.wal.gc.offset_floor_population_gap` | `Counter<long>` | `{consumer}` | Durable-pin consumers a WAL GC pass found that the durable offset floor does not speak for, because they never reported an offset at all. The floor is a minimum over the consumers that *reported* an offset, not over the consumers that *owe* WAL entries (issue #2314), so a consumer present in the durable pin dictionary but in neither the covered set nor the abstained set leaves the floor too high. That is a different state from reporting the `-1` no-dependency sentinel, which is a real answer from a participating leaf. The pass fails closed by blocking the partitions the unreported consumer holds, so the effect is bounded over-retention that clears the moment it reports, never a lost WAL prefix. Distinct from `orleans.lattice.wal.gc.offset_floor_unavailable`, which counts the adjacent case where the pin-store read itself threw; this counter ticks when the read succeeded and the two planes disagree about the population. Expected to sit at zero, because `WalMaterialiserPinGrain.Merge` writes the pin and the offset in lockstep; a non-zero value means state predating the offsets plane, a swallowed birth seed, or a partially-upgraded population, and the amount added is the number of such consumers on that pass. Tagged `tree`. |
| `orleans.lattice.wal.gc.offset_floor_unavailable` | `Counter<long>` | `{pass}` | WAL GC passes that could not compute the durable leaf-materialiser offset floor because the pin store read (`GetPinOffsetsAsync`) threw. Since issue #3576 such a pass fails closed and trims nothing at all, TTL included, and reports `BlockedByUnusablePin`; before that it fell back to the HLC floor alone, which over-trims the low-HLC/high-offset reap class the offset floor retains. The condition was previously silent (issue #2314). A failed `GetPinsAsync` read fails the pass closed the same way and records `-1` on `orleans.lattice.wal.gc.blocked_consumers`; each skipped pass logs a warning naming the unreadable plane. The healthy "no offset floor" outcomes - a host that never wired the durable pin store, or a store reporting no offsets - do not reach the swallowing catch and do not increment this counter, which is what separates "no floor because unreachable" from "no floor because none needed". A transient tick is expected (the next pass retries; it also ticks during a rolling upgrade past an older pin grain); a sustained non-zero rate means the tree reclaims no WAL until the store recovers. Tagged `tree`. |
| `orleans.lattice.wal.gc.durability_hold_forced` | `Counter<long>` | (none) | WAL GC partition scans that trimmed a tree with **no durable materialiser offset floor** despite the durability hold (`LatticeOptions.WalDurabilityHoldCeilingBytes`) (issue #3300). Tagged `tree` and `reason`. This is not a tuning hint. A non-zero value states that the collector released WAL entries that nothing is known to have durably applied - the exact condition in which issue #3300 discarded eleven hours of writes while every series the collector published read healthy. Alert on it. **The two `reason` arms are not interchangeable and do not share a repair.** `ceiling_exhausted` means the hold engaged, retained up to its ceiling and then yielded; the fault is the floor that never arrived, and raising the ceiling suppresses the counter for longer without changing anything about it. `unmeasurable_footprint` means the hold never engaged at all, because the provider reports no retained-byte figure and a hold with nothing to bound it cannot be allowed to run by default - that is a storage-provider defect, and the ceiling is irrelevant to it. Collapsing the two would rebuild the one-arm-means-two-things conflation that `durability_unverified` was added to undo. `orleans.lattice.wal.gc.durable_floor_stall_seconds` is the series that says how long the underlying fault has been present, and the two are read together. Zero here is genuinely good news, because the hold is **on by default** (`LatticeOptions.DefaultWalDurabilityHoldCeilingBytes`, 256 MiB): a flat zero means the check ran and found nothing to force. That holds only while the ceiling is left positive - set it to `0` and the hold never engages, so zero reverts to meaning the check is switched off rather than that it passed. |
| `orleans.lattice.wal.gc.durable_floor_stall_seconds` | `Histogram<long>` | `s` | Seconds since a tree's durable leaf-materialiser offset floor last advanced, recorded once per garbage-collection pass and tagged `tree` and `status` (issue #3300). It exists because no counter already published could answer whether a tree is making anything durable: every one of them measures **volume**, and a volume that tracks how much work was attempted rises identically whether or not any of it landed. This measures **elapsed time without progress**, which is the only quantity that separates a slow tree from a stopped one. Three arms: `advanced` (the floor rose on this pass, recorded at a true zero), `stalled` (a floor exists and did not rise, recorded at its age since the last genuine advance) and `absent` (no durable floor exists at all, recorded at the age since this process first observed the tree). **The `absent` arm is deliberately never zero.** A tree with no floor has no advance timestamp to subtract from, and reporting zero there would put the state in which nothing is known to be durable on exactly the reading a perfectly healthy tree produces - the conflation that let issue #3300 run for eleven hours behind counters that all looked fine. Progress is judged against a **high-water mark**, not against the previous pass: the floor is a minimum over reporting leaves, so it can legitimately fall when a lagging leaf starts reporting, and treating that fall as movement would restart the stall clock on a tree making none. Losing a floor entirely does not reset it either. **Reading it: a sustained `absent` or `stalled` age is a tree whose materialiser is not making progress**, and the age is directly the sentence an operator needs - "this tree's durable floor has not advanced in N hours" - without a census, an archive diff, or a restart to find out what survived. A healthy tree sits on `advanced` at zero. |
| `orleans.lattice.wal.gc.trim_stop` | `Counter<long>` | `{scan}` | Why each WAL GC shard scan stopped (issues #3149, #3155, #3207), tagged `tree`, `shard` and `reason`. It exists because a pass that reclaims nothing is the single most consequential WAL GC reading and was, until this counter, unattributable. `orleans.lattice.wal.gc.passes` classifies such a pass `over_ceiling` whenever the tree is above its byte ceiling, and that classification is decided by the **consumer-cursor** floor alone - so a tree stranded by the **durable leaf-checkpoint offset floor** reports a healthy, available cursor floor, reclaims zero on every pass, and grows without bound, while `over_ceiling` is also the one outcome on which the blocking-pin diagnostic is never written. The arms separate the two innocent stops - `exhausted` (the scan consumed the whole shard - the healthy reading) and `empty` (the shard holds no entries at all, which is not the same as having just reclaimed everything and must not be averaged into it) - from the four that indict something, each naming a different subsystem: `offset_floor` (the first entry examined sat above the durable leaf-checkpoint offset floor, so **one lagging leaf is holding the whole tree**, since that floor is a minimum over reported checkpoints), `cursor_floor` (neither the consumer cursor nor the TTL ceiling accepted the entry, which indicts a lagging *consumer cursor*), `causal_frontier` (a reported per-origin stable frontier does not dominate the entry's vector clock, which indicts a *replication origin* rather than the consumer that reported it) and `block_pin` (a consumer's buffer pin holds the entry back, a deliberate hold rather than a lag, so a sustained run against a consumer that is no longer buffering is a *leaked pin*). Those four were the distinct causes of a zero-reclaim pass until the durability arms below joined them, and the last three of them were a single `not_eligible` arm until issue #3155 split them - which left a tree provably stranded with no way to say by what. Issue #3300 added three further arms on the durability axis: `durability_unverified` (the scan consumed a **non-empty** shard while **no** durable offset floor existed, so it released every entry without any check that the data had been applied anywhere - previously indistinguishable from `exhausted`, which is how a tree discarding live, never-checkpointed records published the arm documented as healthy), `durability_hold` (the same absent floor, but the durability hold (`LatticeOptions.WalDurabilityHoldCeilingBytes`, on by default at 256 MiB) still had budget remaining, so the shard was **retained** untouched instead), and `durable_offset_refusal` (the scan stopped at an entry the consumer cursor would have admitted, because the durable leaf-checkpoint offset floor overruled it and no configured retention TTL admitted it either - the data is correctly retained pending durability, and a sustained run with nothing reclaimed indicts a leaf whose durable checkpoint or snapshot coverage has stopped advancing). The first two are a deliberate pair: both describe a tree with no durable floor, and they differ in whether the data survived. Every arm - `exhausted`, `empty`, `offset_floor`, `cursor_floor`, `causal_frontier`, `block_pin`, `durability_unverified`, `durability_hold` and `durable_offset_refusal` - is zero-primed above every early return, so an absent series means WAL GC is not running for this tree on this silo rather than that no scan ever stopped, and one measurement is recorded **per shard per pass**, carrying the `shard` it stopped on, so a single stranded partition stays visible instead of being absorbed by its healthy siblings. **The `shard` dimension is what separates a slow tree from a wedged one (issue #3207).** A stop is not on its own an indictment: a healthy shard releases every entry it may and then stops at the first one it must retain, publishing `offset_floor` exactly like a shard that has never released an entry, so at tree scope the two readings are identical. Joined per shard against `orleans.lattice.wal.entries_trimmed` they separate exactly: **a stop arm advancing while that same shard's entries-trimmed stays flat is a shard that is asked on every pass and releases nothing.** That join is the only signal available for such a shard, because dead bytes rise only as a consequence of a trim - no write path marks anything dead - so a shard that has never trimmed reports zero dead bytes, a zero dead ratio and zero compactions, which is the best score available on every arm derived from dead-byte accounting. **Reading it: a sustained `offset_floor` rate alongside `over_ceiling` passes is unbounded WAL growth with a named cause** - find the leaf whose checkpoint is not advancing; a sustained `cursor_floor` rate is a consumer that is not acknowledging, while a sustained `causal_frontier` or `block_pin` rate is not. |
| `orleans.lattice.wal.gc.floor_head_distance` | `Histogram<long>` | `{offset}` | The floor-to-head retained distance of each WAL GC shard scan (issue #3149): the offsets from the first entry the scan had to retain through the shard's head, inclusive. `orleans.lattice.wal.gc.trim_stop` says a scan stopped and `orleans.lattice.wal.entries_trimmed` says what it released first, but neither says how much of the shard the floor is holding - the quantity that separates a floor holding a handful of recent writes from one stranded far behind the head. Measured in offsets rather than bytes because byte accounting is per tree, not per shard. Recorded on **every** scanned shard, at zero when the scan did not stop on a retained entry (it released everything it was offered, or the shard was empty), so a measured zero means the floor holds nothing rather than that the shard is not reporting. A stopped scan costs one extra head read; an exhausted or empty one costs none. Tagged `tree`, `shard` and `tenant` - the same set as `orleans.lattice.wal.entries_trimmed`, so the retained side of a scan joins its released side shard for shard. |
| `orleans.lattice.wal.gc.durability_hold_engaged` | `Counter<long>` | (none) | WAL GC passes on which the durability hold **engaged and retained the scan** (issue #3300). Tagged `tree` and `reason`. The mirror of `durability_hold_forced`: that one counts the hold yielding, this one counts it working. It is emitted only when every cursor admitting the trim is a leaf materialiser the durable offset floor does not speak for - the state in which the collector would otherwise release records that nothing outside this process has attested to. A tree with a replication shipper, view maintainer, log subscriber or backup capture reporting a cursor never reaches it, because those cursors survive the process. **The three `reason` arms need different operator responses, which is the whole reason this instrument exists separately from the `durability_hold` trim-stop arm.** `never_pinned` means no durable materialiser offset floor has ever been observed for this tree in this process: it is stalled, it will hold until the ceiling forces it onto `durability_hold_forced`, and it needs someone to wire or repair a materialiser. `pin_regressed` means a floor was observed earlier and is absent now - a rolling upgrade or leaf churn - and it clears itself without intervention as the leaves re-pin. Reporting both on one arm would tell an operator mid-upgrade that they had an outage. `cursor_unreadable` means the consumer-cursor registry could not be read on this pass (issue #3366), so the hold engaged without establishing who is watching at all. **It is not a claim about the pin history, which this pass never measured**, and that is precisely why it cannot be folded into either of the other two arms: both of those assert a fact about pins that an unread registry does not supply, so reporting either would put a measurement this pass does not have onto a series an operator reads as one. It indicts the *registry read path* in this process, not the tree's materialiser, so a sustained run is a defect to fix here rather than a leaf to repair elsewhere. Before issue #3366 this pass classified the cursor `durable` instead, which disabled the hold entirely and released the scan with no durability evidence, and did so silently - `durability_hold_forced` is itself gated on the hold being configured, so the counter that would have reported the release was switched off by the same fault that caused it. Read alongside `orleans.lattice.wal.gc.durable_floor_stall_seconds`, which says how long the condition has persisted. |
| `orleans.lattice.wal.gc.scheduler.phase_age` | `ObservableGauge<double>` | `s` | How long the silo's WAL GC scheduler loop has been in its current phase (issue #3060). Tagged `phase` and, while collecting, `tree`. **This is the spine of the scheduler liveness set, and it is a gauge rather than a counter deliberately**: a loop that has stopped is a *location* problem, not an *event* problem, and no counter can say where a loop is parked - only that it has not moved, which a quiet healthy silo also reports. Eleven phase arms: `unstarted` (the hosted service was constructed but `ExecuteAsync` has not been entered), `disabled` (`WalGcInterval <= 0`, a correct exit), `starting` (inside the startup delay), `enumerating` (awaiting the tree registry), the four-way collecting split `collecting.priming`, `collecting.reconciling`, `collecting.gc_run` and `collecting.healing`, then `pruning`, `waiting` (parked on the inter-pass delay) and `stopped`. **Exactly one series is emitted at a time** - the phase the loop is actually in - because a loop is in one phase, and publishing a stale age for the ten it is not in would leave a reader to pick the live number out of eleven climbing ones. That is also why the phase dimension is **not** zero-primed: priming it would manufacture the ten fictions it exists to avoid. It needs no priming, because an `ObservableGauge` reports from its declaration site on every scrape whether or not the scheduler ever ran, which makes the presence of this series the **build-version witness** for the whole set: if `phase_age` is absent the build predates these instruments, and if it is present but the five counters and histograms are not, that is a wiring fault rather than an idle silo. **Reading it: an age that exceeds the cadence the scheduler selected (`orleans.lattice.wal.gc.scheduler.wait`) is a stalled sweep, and the `phase` tag names where it stalled.** An age climbing without bound on `collecting.*` names the tree in the `tree` tag; on `enumerating` the registry is not answering; on `waiting` the delay itself did not return. |
| `orleans.lattice.wal.gc.scheduler.passes_started` | `Counter<long>` | `{pass}` | Passes the silo's WAL GC scheduler loop has begun (issue #3060). Incremented as a **heartbeat above the registry enumeration**, so it advances even on a pass that dies at its first await - which is the whole point, since the pre-existing `orleans.lattice.wal.gc.passes` is emitted per tree *after* enumeration succeeds and is therefore silent on exactly the failures that stop the loop. Untagged by tree: the loop is one per silo, and a pass that failed before enumerating has no tree ids to tag with. **Read it against `phase_age`**: a total that keeps climbing while every per-tree `wal.gc.*` series is frozen is a loop that is alive and failing, which needs a different remedy from a loop that has stopped, and before this counter existed those two states produced byte-identical scrapes. Zero-primed as the first statement of `ExecuteAsync`, above the disabled return and above the startup delay, so a zero here is a measured zero and an absent series means the build lacks the instrument. |
| `orleans.lattice.wal.gc.scheduler.pass_duration` | `Histogram<double>` | `s` | How long one WAL GC scheduling pass actually took (issue #3060), recorded in a `finally` so a pass that faults, is cancelled, or is abandoned at a bound is measured too - that population's duration is itself the diagnosis, and siting the record on the success path would discard exactly the sample worth having. **This is the counterpart to `orleans.lattice.wal.gc.scheduler.wait`, and the pair is only informative read together**: `wait` is sited on the scheduler's *decision* and therefore cannot observe anything that happens after it, so a silo whose per-tree touches all end at the Orleans default 30s response timeout selects the interval floor on every pass while the wall period is twice that, and `wait` reads a flat, healthy-looking floor throughout. Divergence between the two is the signature of a loop whose passes are being stretched by something the cadence policy never saw. Carries no tree dimension: a pass is a property of the loop rather than of any tree it visited. Deliberately **not** zero-primed, because the empty state of a duration distribution is undefined rather than zero and a primed `0 s` would corrupt every quantile; its liveness is anchored instead to `orleans.lattice.wal.gc.scheduler.passes_started`, with exactly one observation per started pass asserted by fixture. |
| `orleans.lattice.wal.gc.scheduler.enumerations` | `Counter<long>` | `{enumeration}` | Outcome of the tree-registry enumeration that opens each WAL GC scheduler pass (issue #3060), tagged `outcome`. `succeeded` saw at least one non-blank tree id. `faulted` is the registry throwing, which before this instrument existed was swallowed into a debug log and left no trace on any scrape at any level. `cancelled` is an orderly silo shutdown and is not a fault. `empty` and `all_blank` are separated because a registry that answers with ids that are all blank *reports success, returns content, and collects nothing*, so on every other series it presents exactly as an idle silo. `timed_out` is a budget-shaped failure, though not necessarily this scheduler's own bound: both awaits are wrapped by `Task.WaitAsync`, which raises `TimeoutException` when the budget expires and *also* propagates one raised inside the operation - an Orleans response timeout at the 30s default, for instance. The two are indistinguishable by exception type, so this arm deliberately does not split them; the abandonment log line carries the measured elapsed and an explicit attribution, and that is where a reader who needs the cause should look. Reading this arm as "our bound fired" is how a pre-registered diagnostic reached a verdict that had to be withdrawn. It is kept apart from `faulted` because a fault is a property of the registry and a timeout is a property of a bound, so that **a limit of ours is never reported as a failure of the subject**. All six outcome arms are zero-primed by walking the enum rather than a hand-written list, so an arm added later cannot ship unarmed, and each arm carries a positive control that drives the scheduler into that outcome and observes it advance - so the zeros the others report are earned rather than assumed. |
| `orleans.lattice.wal.gc.scheduler.terminations` | `Counter<long>` | `{termination}` | Why the silo's WAL GC scheduler loop stopped (issue #3060), tagged `reason`. **A `BackgroundService` that returns from `ExecuteAsync` never runs again for the life of the process and the host reports nothing**, so every arm here is a permanent end to silo-wide WAL collection and any non-zero reading is terminal until the silo restarts. `disabled` is `WalGcInterval <= 0`: a correct exit which, on every other instrument in the system, is indistinguishable from a wedge. `cancelled` is an orderly host shutdown. `faulted` is the loop body throwing, recorded and then rethrown so the host's configured exception behaviour is unchanged. All three reason arms are zero-primed by walking the enum, so a zero is a reading and not a missing wire. |
| `orleans.lattice.wal.gc.scheduler.wait` | `Histogram<double>` | `s` | How long the silo's WAL GC scheduler **chose** to wait before its next pass (issue #3060), observed once per pass at the delay site so it is recorded even when the pass reached no tree. The adaptive cadence returns to the configured minimum only after a fully successful pass, so every other exit relaxes geometrically toward the ceiling, and that climbing ladder is the one positive signature a loop which is alive and retrying has - every other series such a loop touches is frozen exactly as a stopped loop leaves them. **This is the selected wait, not the elapsed pass.** It is sited on the decision and so cannot observe a failure that happens after the decision; `orleans.lattice.wal.gc.scheduler.pass_duration` is the counterpart that can, and selected-versus-actual divergence is the reading worth having. Deliberately not zero-primed (an empty duration distribution is undefined, not zero); its count is anchored to `orleans.lattice.wal.gc.scheduler.passes_started`. |
| `orleans.lattice.storage.wal_bytes` | `ObservableGauge<long>` | `By` | Retained WAL bytes for the tree. Observed lazily from the per-tree storage-usage aggregator's last-known report (coalesced behind `StorageUsageCacheTtl`); a tree on a provider that does not support byte accounting reports no data rather than `0`. This and the other byte and depth `orleans.lattice.storage.*` gauges are withheld for a tree whose last report was partial (a WAL surface reported no byte accounting or did not answer, or a shard did not answer); `orleans.lattice.storage.policy.over_threshold` is tracked apart from them and set only from a complete WAL measurement. A tree's series stops being reported once it has not been refreshed within the storage sink's staleness horizon - four times the slowest active storage-usage poll interval, never less than 60 s (60 s at the default 15 s `StorageUsagePollInterval`) - so when one of the tree's aggregators migrates to another silo the old silo stops reporting it once that horizon passes. Because a tree's aggregators are placed independently, more than one silo can still export the same tree at once; see [The `tree` dimension across aliasing](tag-conventions.md#the-tree-dimension-across-aliasing). Tagged `tree`. |
| `orleans.lattice.storage.snapshot_bytes` | `ObservableGauge<long>` | `By` | Snapshot blob bytes for the tree. **Reports no data until the tree has had a deep storage-usage publish**, which only a full storage-usage report makes: an `ILattice.GetStorageUsageAsync` call, an `ILatticeAdmin.GetTotalStorageUsageAsync` or `RefreshStorageUsageAsync` roll-up, the deep poll when `StorageUsageDeepPollInterval` is positive (it is off by default), the admission write guard's refresh when a cap or advisory ceiling is configured, or the tenancy package's usage metering; the continuous background WAL poller does not measure this surface. Previously the poller seeded it to `0` without measuring it, so an unmeasured tree was exported as a measured zero and read as "no snapshots exist" while 137 MB of snapshot state was on disk (issues #2693, #2692). Pair with `orleans.lattice.storage.usage_deep_published` to tell the two apart. Tagged `tree`. |
| `orleans.lattice.storage.leaf_state_bytes` | `ObservableGauge<long>` | `By` | Summed leaf/shard-root state bytes for the tree. Carries the same deep-publish precondition as `snapshot_bytes` above: no data until a deep publish has run for the tree, never a synthesised `0` (issue #2693). Tagged `tree`. |
| `orleans.lattice.storage.total_bytes` | `ObservableGauge<long>` | `By` | Sum of the three storage surfaces for the tree, with a WAL term that depends on which path published last. A deep publish sums the WAL's **physical** occupancy with `snapshot_bytes` and `leaf_state_bytes`; that WAL term is not `wal_bytes`: it also counts dead bytes a log-structured provider has trimmed but not yet compacted away, so the deep total can exceed `wal_bytes` + `snapshot_bytes` + `leaf_state_bytes` (issue #3107). Each background WAL-only poll in between recomputes it as the retained `wal_bytes` figure plus the last deep `snapshot_bytes` and `leaf_state_bytes`, so on a provider that carries dead bytes the gauge can step down at a poll and back up at the next deep publish with nothing changing on disk. Carries the same deep-publish precondition as the two surfaces above, and it is the one where a synthesised zero was worst: it presents as a total while silently omitting two of its three terms, so a WAL-only publish made it read as the WAL size alone (issue #2693). Tagged `tree`. |
| `orleans.lattice.storage.usage_deep_published` | `ObservableGauge<long>` | `1` | `1` once a report from the deep path that measures snapshot and leaf-state bytes has landed for the tree in this silo's sink - later WAL-only polls on the same silo keep it at `1` - and `0` when only the cheap WAL-only poller has reported. It is the companion signal that makes the three suppressed gauges above readable: no series at all means the tree has not been observed, `0` means observed but not deeply measured, and `1` means a real measurement was taken and any zero on the deep surfaces is a genuine zero (issue #2693). Tagged `tree`. |
| `orleans.lattice.storage.policy.over_threshold` | `ObservableGauge<long>` | `1` | `1` when the tree's WAL occupancy currently breaches the advisory ceiling (`LatticeOptions.WalMaxRetainedBytes`), else `0`. Occupancy is the physical figure - dead bytes not yet compacted included - not the retained figure `orleans.lattice.storage.wal_bytes` reports, so the gauge can read `1` while `wal_bytes` sits below the ceiling (issue #3107). Advisory only - the byte-pressure policy never trims past the safe consumer frontier, so the gauge stays at `1` while a lagging consumer pins the bytes. A `1` asserts no cause, though: because occupancy counts dead bytes, a tree that trimmed everything it was entitled to and is only waiting on compaction also reads `1` with nothing pinning it (issue #3204). Tagged `tree`. |
| `orleans.lattice.storage.policy.trim_triggered` | `Counter<long>` | `{trim}` | Incremented once per WAL GC pass on which the advisory byte-pressure policy is armed: the pass's pre-trim WAL occupancy exceeds the ceiling (`LatticeOptions.WalMaxRetainedBytes`), or sits between `LatticeOptions.WalBytePressureReclaimTarget` of the ceiling and the ceiling while still armed from an earlier breach. The pass trims to the same safe frontier either way, so the counter records that the policy considered the tree under byte pressure, not that bytes were released. Tagged `tree` and `reason` (`byte_pressure`). Not emitted when the policy is disabled or the WAL provider does not support byte accounting. |
| `orleans.lattice.storage.policy.bytes_reclaimed` | `Counter<long>` | `By` | WAL bytes freed by a garbage-collection pass on a tree with an advisory ceiling (`LatticeOptions.WalMaxRetainedBytes`) configured: the pass's pre-trim minus post-trim WAL occupancy. Recorded on every such pass whose occupancy fell, whether or not the byte-pressure policy triggered on it (see `orleans.lattice.storage.policy.trim_triggered`). Zero-reclaim passes do not emit, and a missing sample names no cause: the floor may be pinned below the bytes, or the pass may have trimmed entries whose dead bytes still count toward occupancy until compaction removes them (issue #3204). Tagged `tree`. |
| `orleans.lattice.storage.wal.uncompressed_bytes` | `Counter<long>` | `By` | Pre-compression encoded WAL payload bytes the Azure Table WAL provider committed for the tree, summed once per append batch after the batch's phase-1 rows land. Emitted whether or not row compression is on (it is on by default: `AzureTableWalStorageOptions.Compression` defaults to `Zstd`); with compression off it equals `stored_bytes`. No other WAL provider emits it. Tagged `tree` and `tenant`. |
| `orleans.lattice.storage.wal.stored_bytes` | `Counter<long>` | `By` | Post-compression WAL payload bytes actually stored for the tree, summed once per append batch. Dividing this by `uncompressed_bytes` gives the realised compression ratio; `1 -` that ratio is the savings. Tagged `tree` and `tenant`. |
| `orleans.lattice.storage.wal.compression_skipped` | `Counter<long>` | `{row}` | WAL rows the Azure Table WAL provider stored verbatim instead of compressed. Tagged `tree`, `tenant` and `reason`: `below_threshold` (payload shorter than `CompressionMinPayloadBytes`, default 256 bytes), `inflation_guard` (compressing did not shrink the payload), or `disabled` (no compressor is active because `AzureTableWalStorageOptions.Compression` is `None`, so every row lands here). A reason is emitted only for a batch in which it is non-zero and is never zero-primed, so an absent `reason` series means no row has been skipped for that reason on this silo. |

Previous: [Instrument catalog: Snapshot cursors (sourced from SnapshotLeafGrain / snapshot-cursor open path)](instrument-catalog-3.md). Next: [Instrument catalog: WAL compaction (sourced from FileWalShard) to Foreground read envelopes (sourced from LatticeGrain)](instrument-catalog-5.md). Contents: [Metrics](../metrics.md).
