Table of Contents

orleans.lattice meter

This page documents Orleans.Lattice.Dashboards 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at orleans-lattice-meter.md, and llms.txt lists every page.

Part of Metric-to-panel coverage map.

Instrument Type Tags Dashboard Panel(s)
orleans.lattice.build.info observable gauge ({build}) version, sha, tenant Overview Deployed build (version / commit)
orleans.lattice.shard.reads counter ({op}, per-operation) tree, shard, tenant Overview Cluster throughput (ops/s), Active shards (per tree)
orleans.lattice.shard.writes counter ({op}, per-operation) tree, shard, tenant Overview Cluster throughput (ops/s), Per-tree write throughput (operations/s and records/s)
orleans.lattice.shard.records_written counter ({record}, per-record) tree, shard, tenant Overview Per-tree write throughput (operations/s and records/s)
orleans.lattice.shard.splits_committed counter ({split}) tree, shard, tenant Overview Splits committed
orleans.lattice.shard.consolidations_committed counter ({consolidation}) tree, shard, tenant Overview Consolidations committed
orleans.lattice.shard.healing.backlog histogram ({shard}) tree, tenant Overview Healing backlog (shards above base)
orleans.lattice.shard.healing.decisions counter ({decision}) tree, decision, tenant Overview Healing decisions by outcome
orleans.lattice.leaf.write.duration histogram (ms) tree, kind, tenant Overview, CommitPath Leaf write duration percentiles (Overview "Leaf write duration (p50/p95/p99 ms)", CommitPath "Storage-provider write duration (leaf.write.duration p95)"). Both panels aggregate over kind, so they blend the untagged state-row writes with the compact, merge and backstop WAL appends; add kind="" to a query to read the state-row writes alone
orleans.lattice.leaf.scan.duration histogram (ms) tree, operation, tenant Overview Leaf scan duration p95 by operation
orleans.lattice.leaf.compaction.duration histogram (ms) tree, trigger, tenant Overview, CommitPath Compaction duration p95
orleans.lattice.leaf.tombstones.created counter ({tombstone}) tree, tenant Overview, CommitPath Tombstone churn on Overview; Tombstones reaped vs expired (TTL churn) on CommitPath
orleans.lattice.leaf.tombstones.reaped counter ({tombstone}) tree, trigger, tenant Overview, CommitPath Tombstone churn on Overview; Tombstones reaped vs expired (TTL churn) on CommitPath
orleans.lattice.leaf.tombstones.expired counter ({tombstone}) tree, trigger, tenant Overview, CommitPath Tombstone churn on Overview; Tombstones reaped vs expired (TTL churn) on CommitPath
orleans.lattice.compaction.pass.duration histogram (ms) tree, trigger, tenant Overview Compaction pass duration p95 by trigger
orleans.lattice.compaction.leaves.visited counter ({leaf}) tree, outcome, trigger, path, tenant Overview Compaction leaves visited (rate, by outcome)
orleans.lattice.compaction.shard.retries counter ({retry}) tree, tenant Overview Compaction shard retries / skips
orleans.lattice.compaction.shard.skipped counter ({shard}) tree, tenant Overview Compaction shard retries / skips
orleans.lattice.leaf.tombstone.ratio histogram (1) tree, tenant Overview Leaf tombstone ratio p95
orleans.lattice.leaf.splits counter ({split}) tree, tenant Overview Splits committed
orleans.lattice.leaf.split.completion.in_flight observable gauge ({completion}) tree, tenant Overview Leaf-split completions currently suspended inside a leaf split's completion phase (issue #2967). The divided/faulted outcome counters partition only TERMINATED divisions, so a division still in the completion body at scrape time belongs to none of them, and because the catch that records faulted is unconditional a division carrying neither divided nor faulted did not throw - it is genuinely in flight, a state nothing named before this gauge. Read beside the oldest-completion-age panel: a non-zero count is only actionable once that age is climbing. No per-tree pre-mint, so a tree with nothing in flight emits no series and No data is the healthy steady state
orleans.lattice.leaf.split.completion.oldest_age observable gauge (s) tree, tenant Overview Age in seconds of the oldest leaf-split completion suspended inside a leaf split's completion phase (issue #2967). Separates a division momentarily in flight (age near zero, falling) from one that is stuck (age climbing without bound), which the bare in-flight count cannot. A measurement, not a threshold: the code encodes no abandonment age, so alert on a sustained climb read against real completion durations. A registry-backed gauge rather than a completion-duration histogram on purpose - a histogram only records on completion, so a division that never completes would never appear in one. The panel queries orleans_lattice_leaf_split_completion_oldest_age_seconds, the .AddPrometheusExporter() family for its s unit, matching orleans.lattice.backup.inventory.oldest_age; the repository-context container publishes the bare name instead (see How an instrument name becomes a PromQL series name)
orleans.lattice.leaf.bisect_refusals counter ({refusal}) tree, reason, detach_seam, tenant Overview Leaf division fast path forfeited (issue #2787). Read reason and detach_seam together (issue #2856): no_snapshot_attached means no frame is attached and the fallback materialises nothing - seam none is a leaf that never had a frame, any other seam names the earlier (typically read-path) surface that already paid for the whole leaf. frame_key_unreadable and no_key_sorts_below_pivot refuse with a frame still attached, so the split itself materialises the entire leaf. too_few_rows also refuses with a frame still attached, but that frame holds fewer than two rows, so the fallback materialises at most one row from it. Seam values underlying_rows_accessor and range_hydration_completed are never produced by a deployed process (issue #2865), so a missing line for either is expected and is not evidence about leaf behaviour
orleans.lattice.leaf.byte.overflow counter ({leaf}) tree, outcome, tenant Overview Splits committed. Pre-minted at zero for both outcomes when a leaf reaches the byte-bound check during snapshot capture, with the tag set a real emission carries, so a missing series means the build did not land rather than that no leaf overflowed (issue #2756). The panel's byte-overflow target nonetheless still appends or vector(0), so on the chart that absence draws as a zero line
orleans.lattice.leaf.split_attempts counter ({attempt}) tree, outcome, failure_class, tenant Overview Leaf divisions sought; failure_class (unaffordable, timeout, other) is carried by the faulted arm only. Read beside the bisect-refusal panel, which is uninterpretable alone: a zero refusal rate means either no division forfeited its fast path or none was ever attempted, and only this counter separates them. All six outcomes are pre-minted at zero when a leaf reaches the byte-bound check during snapshot capture, so the panel omits or vector(0) and a flat zero is a measured zero (issue #2756). recovered is a division completed by the recovery path after its intent was stranded by an earlier fault, a deactivation, or a crash; it is never also counted as divided nor as a second orleans.lattice.leaf.splits increment (issue #2860). no_admissible_pivot means the division was declined because no row fell strictly inside the leaf's own declared range, so it owned nothing it was entitled to divide (issue #3117); it is self-correcting rather than a wedge - the out-of-span rows drain to their real custodians and the leaf is then under threshold or divisible - but a sustained non-zero rate means rows are not draining and should be read beside the span-admission forward metrics
orleans.lattice.leaf.commit.duration histogram (ms) tree, step, tenant CommitPath Commit-step latency p50/p95/p99
orleans.lattice.cache.hits counter ({hit}) tree, tenant Overview Cache hit ratio
orleans.lattice.cache.misses counter ({miss}) tree, tenant Overview Cache hit ratio
orleans.lattice.atomic_write.completed counter ({saga}) tree, outcome, tenant Overview, AtomicWrites Atomic write outcomes (rate); per-tree committed throughput; saga failure rate (failed + compensated / total); range-window non-committed saga count. compensated is an abort with no captured failure message - in the current code a cross-tree participant its coordinator finalized with abort - and neither abort arm issues rollback writes. The failure-rate panels also chart a shutdown_refused series, but no code path produces that outcome, so it returns no data: a shutdown refusal throws LatticeShuttingDownException without a terminal transition, and a guard-rejected batch records no sample, so neither is in these rates
orleans.lattice.atomic_write.duration histogram (ms) tree, outcome, tenant Overview, AtomicWrites Saga duration p50/p95/p99; saga duration p95 by outcome
orleans.lattice.atomic_write.batch_size histogram ({entry}) tree, outcome, tenant Overview, AtomicWrites Batch size p50/p95/p99; batch size p95 by outcome
orleans.lattice.coordinator.completed counter ({operation}) tree, kind, tenant Overview Coordinator completions
orleans.lattice.coordinator.phase_tick.failures counter ({failure}) tree, kind, tenant Overview Coordinator phase-tick failures (rate) - zero-primed per coordinator, so a flat zero is a reading that ticks are succeeding; any non-zero value is discarded phase-loop work and is operator-actionable
orleans.lattice.coordinator.phase_tick.consecutive_failures observable gauge ({failure}) tree, kind, tenant Overview Consecutive failed coordinator phase ticks, as the max over the activations sharing a tag set - every live coordinator reports, so 0 is a reading that the last tick succeeded; a value that keeps returning to zero is transient fault absorption, while one that only climbs is a wedged phase machine and is operator-actionable
orleans.lattice.tree.lifecycle counter ({event}) tree, kind, tenant Overview Tree lifecycle events (annotation + stat). One event per logical delete, recover or purge under the logical tree id, including on an aliased tree; a resize's retirement of its old copy records nothing, so a live resized tree never appears as deleted
orleans.lattice.events.published counter ({event}) tree, kind, tenant Overview Events published
orleans.lattice.events.dropped counter ({event}) tree, reason, tenant Overview Events dropped
orleans.lattice.config.changed counter ({change}) tree, config, tenant Overview Runtime config changes
orleans.lattice.observer.duration histogram (ms, per-observer per-mutation sample) tree, observer, tenant CommitPath Mutation-observer inline latency p95 (ms) by observer
orleans.lattice.storage.wal_bytes observable gauge (By) tree, tenant Overview Storage footprint by tree
orleans.lattice.storage.snapshot_bytes observable gauge (By) tree, tenant Overview Storage footprint by tree
orleans.lattice.storage.leaf_state_bytes observable gauge (By) tree, tenant Overview Storage footprint by tree
orleans.lattice.storage.total_bytes observable gauge (By) tree, tenant Overview Cluster total stored bytes. The per-tree total's WAL term depends on which path published last: a deep publish adds the WAL's physical occupancy - dead bytes not yet compacted included, so it can exceed orleans.lattice.storage.wal_bytes - to the snapshot and leaf-state bytes, while each background WAL-only poll in between recomputes it from the retained wal_bytes figure plus the last deep snapshot and leaf-state bytes, so on a provider that carries dead bytes the total can step down at a poll and back up at the next deep publish with nothing changing on disk (issue #3107). No data for a tree until its first deep publish (issue #2693)
orleans.lattice.storage.usage_deep_published observable gauge (1, 0/1) tree, tenant Overview Storage usage measurement depth by tree
orleans.lattice.storage.policy.over_threshold observable gauge (1, 0/1) tree, tenant Overview Trees over advisory threshold. 1 while the tree's WAL occupancy breaches WalMaxRetainedBytes: occupancy is the physical figure, dead bytes not yet compacted included, not the retained figure orleans.lattice.storage.wal_bytes reports, so it can read 1 while wal_bytes sits below the ceiling (issue #3107). Advisory only: the byte-pressure policy never trims past the safe consumer frontier, so it stays at 1 while a lagging consumer pins the bytes. A 1 asserts no cause, though: because occupancy counts dead bytes, a tree that trimmed everything it was entitled to and is only waiting on compaction also reads 1 with nothing pinning it (issue #3204)
orleans.lattice.storage.policy.trim_triggered counter ({trim}) tree, reason, tenant Overview Byte-pressure trim activity. One increment per WAL GC pass on which the byte-pressure policy is armed: the pre-trim WAL occupancy (physical, dead bytes not yet compacted included) exceeds WalMaxRetainedBytes, or sits above WalBytePressureReclaimTarget of it while still armed from an earlier breach. The pass trims to the same safe frontier either way, so it records that the policy considered the tree under byte pressure, not that bytes were released
orleans.lattice.storage.policy.bytes_reclaimed counter (By) tree, tenant Overview Byte-pressure trim activity. The pass's pre-trim minus post-trim WAL occupancy, recorded on every WAL GC pass of a tree with WalMaxRetainedBytes configured whose occupancy fell, whether or not the byte-pressure policy triggered on it; zero-reclaim passes do not emit
orleans.lattice.wal.gc.passes counter ({pass}) tree, outcome, tenant Replication, CommitPath WAL GC pass rate by outcome; WAL GC blocked passes by tree. Also charted on CommitPath alongside leaf.snapshot.coverage_repairs, where the reclaimed arm is the counterpart a zero-coverage repair is meant to unblock (issue #2692). All nine outcome arms are primed at zero per tree (issues #2774, #2850, #3119, #3213), so the by-outcome panel carries no or vector(0) compensation: a missing line there means the build did not land, not that the system is healthy. Only reclaimed is affirmative; idle means a usable floor was evaluated, nothing was eligible, and the tree is inside its byte ceiling, over_ceiling means the same floor evaluation on a tree whose WAL occupancy - dead bytes not yet compacted included - is still over WalMaxRetainedBytes, stranded means the same floor evaluation on a tree whose trim scan stopped on WAL it could not reclaim while no configured ceiling complains about it (issue #3213, the arm a silo with no WalMaxRetainedBytes had no way to report), and no_consumer means there was no floor to evaluate. The arms are mutually exclusive - a pass records a single arm, over_ceiling outranking stranded where both hold - so they partition invocations and their sum is the pass count rather than an over-count. no_partitions means no pinned WAL provider resolved on this silo (issue #2465), not an empty WAL. It outranks the other non-reclaiming arms and is zero-primed per tree. Inspect WAL placement and provider registration. ShardsScanned counts only resolved partitions visited for trimming or compaction; partial resolution keeps the existing outcome and does not imply every partition was examined. With no cursor and no TTL, no_consumer names the no-predicate return; blocked names an unusable durable pin, while stranded can describe a legitimate lagging consumer holding retained WAL.
orleans.lattice.wal.gc.blocked_leaf_reactivations counter ({reactivation}) tree, outcome, tenant Replication, CommitPath (160, 161) On Replication: WAL GC blocked-leaf reactivations by outcome; WAL GC blocked-leaf sweep abandon/re-arm totals over range; WAL GC starved-leaf checkpoint drive by verdict over range. Reactivations of a dormant leaf whose unusable durable pin was blocking its tree's cursor floor, by four lifecycle arms (attempted/healed/abandoned/rearmed) and seven terminal outcome arms, every one but admission_refused partitioning attempted exactly once (issue #3761) (completed/unresolvable/faulted/undelivered/orphaned/latched_stale/admission_refused, orphaned retiring a pin that outlived its reclaimed leaf per issue #3101, latched_stale recording, once per leaf, a drive refused by the issue #3451 stale latch, which is terminal for that pin while the pin stays in the floor per issue #3478, and admission_refused recording a drive the leaf's silo refused admission to its WAL replay gate before replaying anything, which is neither charged against the attempt budget nor refunded to it and so never leads to abandoned, per issue #3575, and is counted per refused try outside attempted, since a pass keeps no more touches in flight than its GC starvation share and re-drives a refused one once a sibling frees a slot, per issue #3761), and seven drive verdicts recording what the touch achieved (drove_lifted/drove_no_advance/drove_memory_refused/drove_not_driven/drove_already_driving from issue #2692, drove_timed_out from issue #3065 - the drive exhausted its StarvationDriveBudget and abandoned its replay rather than keep holding a per-silo replay permit - and drove_admission_refused from issue #3761 - the leaf's silo refused the drive a replay permit before it replayed anything, counted rather than thrown; drove_lifted is also recorded, with no drive issued and no replay permit spent, when the sweep's permit-free bank step alone lifted the consumer's durable pin offset to within a small tolerance of its WAL partition head, issues #3599 and #3649) (issue #2710 Limitation 2, re-arm per issue #2783, undelivered per issue #2768, terminal arms armed exhaustively per issue #2938). Chart beside wal.gc.passes{outcome="blocked"}, which is the series this sweep exists to drive to zero. All 18 outcome arms share one instrument and every one is zero-primed once per tree per process, latched on the tree's first collection rather than repeated per pass - the terminal arms and the drive verdicts by walking their respective outcome enums, so one added later cannot ship unarmed (issue #2938, issue #2692), and each terminal arm and each drive verdict additionally carries a positive control proving its recording path is reachable rather than frozen at zero (issue #2942) - so a zero on healed is a measured zero rather than an unpublished series: a sweep that touches leaves and achieves nothing reads identically to one that was never reached, and an absent series means the build carrying the sweep never landed. abandoned is the alarm arm - the leaf stayed blocked across every permitted attempt of a cycle, so activation alone cannot clear it and its snapshot capture is failing for a separate reason - and rearmed is its counterweight, marking each restoration of the budget after an escalating backoff. abandoned rising while rearmed stays flat is the signature of a sweep that has stopped. undelivered separates a leaf that was never reached, because the touch did not return within its budget, from one that was reached and stayed blocked, so an unreachable leaf no longer reads as an unhealable one. Deliberately not tagged with the leaf identity, which is unbounded; the blocking consumer is named on the paired warning log (issue #2464), which since issue #2815 emits one line per blocked episode rather than one per change of reported blocker and carries a distinct-blocker count, so a single named consumer is not evidence that only one was blocking. On CommitPath its rate summed over every outcome arm is the WAL GC drive-completion series that panels 160 and 161 read against capture-driver minting (issue #3185)
orleans.lattice.wal.gc.blocked_consumers observable gauge ({consumer}) tree, tenant CommitPath (3161) WAL GC blocked-consumer population (uncapped). Denominator for blocked-leaf reactivation outcomes; read both to judge convergence. Zero-primed on the first pass, -1 is unknown, not healthy. Consumer ids count partition pins, not leaves. Keep silo series separate rather than summing duplicated durable populations.
orleans.lattice.wal.gc.blocking_pin_state counter ({consumer}) tree, partition, status, tenant Replication Durable-pin state of each absent consumer blocking a WAL GC pass, so that two opposite conditions the leaf's durable-pin resolution reports identically - both as the same zero-frontier block pin at offset -1 - can be told apart after the fact (issue #3042). Six status arms: checkpointed_uncovered (durably checkpointed but coverage absent - repairable), never_checkpointed (holds live data never checkpointed, so the block is correct by design and no repair exists or should), no_durable_state (the provider answered and nothing was ever persisted - a fourth state, deliberately not folded into never_checkpointed, which would assert a property of the leaf on the strength of an absence), and unreadable (a failure of the measurement - unparseable consumer id, no storage provider on this silo, or a read that threw), and orphaned (the durable state exists but carries no bound tree id, so the leaf was reclaimed or purged after the pin was registered and the pin has outlived its publisher, issue #3105 - before that fix such a pin was reported as checkpointed_uncovered, because the checkpoint classifier never read the tree id and a husk leaf retains its last offset, which is how a tree could report thousands of repairable pins that every repair attempt found nothing to repair), and checkpointed_coverage_unknown (durably checkpointed, but the pin is not known to be unusable, so whether the checkpoint is covered was never determined and no claim is made - the uncovered half of checkpointed_uncovered is inferred from the pin being independently known unusable, a premise the blocked arm has by construction and the floor-holder sample of issue #3158 does not, since it runs only when the cursor floor reports usable and selects by lowest durable offset rather than by usability; excluded from the coverage repair, issue #3168, but a pin on this arm whose offset equals the tree offset floor is driven for liveness, issue #3178 - a scanned-through checkpoint advances only during replay, so it freezes when its leaf deactivates and holds the floor indefinitely on a converged corpus; a pin on this arm above the floor is still not driven). Classified by a direct storage-provider read of the leaf's persisted state that never activates the leaf, following the precedent of the leaf cursor-reporting path, which falls back to reading and writing durable pin state through the storage provider directly when the pin store cannot be reached through a grain, because the blocking population is exactly the population that cannot be activated - an instrument requiring a live leaf would measure the wrong set. Chart beside wal.gc.passes{outcome="blocked"}: this names why each of those blocks is held. Filter partition!="none" before aggregating - all six status arms are zero-primed once per tree per process under that reserved partition value (latched on the tree's first collection, not repeated per pass; Add(0) is idempotent on a counter, so this proves the classifier is wired here and never that it ran on a given pass - orleans.lattice.wal.gc.pass.reach is the advancing layer for that, issue #3075), which is the reachability claim a per-partition prime structurally cannot make (a partition is unknown until a blocker has been parsed, so a tree that never blocked would otherwise carry no series of any shape and "not deployed" would read identically to "nothing to classify"). A non-zero unreadable voids the other five status arms for that tree rather than sitting alongside them, so the six status arms are a partition of one population only while it is zero. Charged once per consumer per blocked episode, latched before the read, so a tree blocked indefinitely costs one durable read rather than one per cadence tick. It is a cumulative tally of classification events, not a population - a consumer classified again in a later episode increments it again, so the total only ever grows and cannot answer "how many leaves are blocked now"; read that way on a live estate it climbed from 281 to 337 on a process that never restarted (issue #3175). Both inputs are capped too - the blocked arm at the fixed number of blocking consumer ids a single WAL GC pass reports, the floor-holder arm at the sweep's own classification cap - so even as a tally it is a lower bound. Do not chart it as a gauge of outstanding blockers; nothing in the estate counts distinct blocked leaves per tree. Diagnostic only: it never changes what a pass is allowed to trim
orleans.lattice.wal.gc.orphan_pin_sweep counter ({pin}) tree, status, cause, tenant Replication Outcome of each durable materialiser pin examined by the WAL GC bulk orphan sweep (issue #3105). The eight status arms partition the examined population, so sum by (tree) over one sweep is the tree's whole durable pin count - the only place it is visible, because the cursor floor's blocking report is capped at eight consumers. retired (leaf read as gone; removed from every shard key it was found under), retire_failed (the same, removal did not complete), deferred (the same, left in place because the pass's retirement budget was spent), refused_malformed_id / refused_ambiguous_partition (the same, but the fail-safe gate declined because the consumer id is not provably the leaf's own, issue #4238), with those five decision arms split by cause - orphaned (record with no tree id) or no_durable_state (no record, the #4238 signature when it is not benign) - so a misresolution is no longer indistinguishable from healthy reclaim (issue #4246), live (leaf still carries a bound tree id, never touched), unresolved (consumer id does not parse back to a leaf grain id), unreadable (leaf-state read failed; the sweep fails closed and never retires on it). Chart deferred beside wal.gc.passes{outcome="blocked"} and wal.gc.blocking_pin_state{status="orphaned"}: deferred is the backlog signal, deliberately a counter arm rather than a gauge, so a sustained non-zero rate means orphans remain and zero with retired flat means the backlog has drained. Its absence is what let issue #3105 run undetected - the only orphan series was a monotonic counter on which a 63-per-hour drain against a 9,468-pin backlog looked like healthy progress. Every arm is zero-primed per tree
orleans.lattice.wal.gc.drive_orphan_pin_retirement counter ({pin}) tree, status, cause, tenant Replication Every orphan-pin removal decision the WAL GC blocked-leaf drive takes on a NotDriven verdict (issue #4246), charted as series B of the bulk orphan-pin sweep panel. retired (removed), retire_failed (the removal threw), refused_malformed_id / refused_ambiguous_partition (the fail-safe gate declined, issue #4238); cause is always not_driven. Zero-primed per tree
orleans.lattice.wal.gc.floor_holder_classification counter ({pin}) tree, status, tenant Replication WAL GC floor-holder classification coverage. How much of a tree's durable materialiser pin population the floor-holder classification examined on a sweep (issue #3158). The two status arms partition the enumerated population, so sum by (tree) over one sweep is the tree's whole durable pin count and classified / sum is the coverage fraction outright. classified is a pin whose leaf state was read and recorded on orleans.lattice.wal.gc.blocking_pin_state; unclassified is a pin the sample did not reach. Chart it directly beneath blocking_pin_state, because it is the denominator that instrument never carried: five zero status arms on a tree nothing ever classified are byte-identical to five measured zeroes on a tree that was classified and found nothing, and on the live estate that ambiguity read as "no pin is in a notable state" for 1.07 GiB of WAL that meant "no pin was ever looked at". A tiny classified fraction is designed, not an alarm - each classification is a durable storage read and one live tree was measured at 52,224 pins on a single sweep, so the sample is capped per sweep by a constant independent of the population. The sample is the lowest-offset pins, which are the ones actually holding the floor (it is a minimum over the pins that reported an offset), in ascending offset order with the frontier as the secondary key and ties broken on the consumer id so the reported holders are stable across scrapes. Issue #3178 corrected this from ascending frontier, which ranks on the floor other axis and therefore described non-blocking pins on a tree stopping at offset_floor; pins reporting no offset are sampled into a separate list so they cannot evict the pins that constrain a floor, and the sample is deduplicated by leaf so a budget of eight is eight leaves rather than eight partitions of one. Both arms are zero-primed per tree above every early return
orleans.lattice.wal.gc.floor_holder_admission counter ({sweep}) tree, status, tenant Replication WAL GC floor-holder admission (issue #3258). Whether the candidate that defines a tree's durable materialiser offset floor cleared the admission gate on a classifying sweep. Charged at most once per sweep, because a tree has one offset floor per sweep however many pins sit on it. Chart it beside floor_holder_classification and blocked_leaf_reactivations: classification says how much of the pin population was examined, this says whether the one candidate that matters was admitted, and reactivation says what happened next. The floor-holder sample is ascending by offset, so its head is the offset floor outright and every other candidate is strictly above it; the gate admits a checkpointed_coverage_unknown candidate at that floor, and since issue #3310 also above it but only once one AT the floor has been admitted on the same sweep, so a refused floor holder means nothing on the tree can ever be admitted and the drive is never entered - permanently, since the durable pin store merges monotonic-max. No other series separates that from a tree with no repair to do: blocked_leaf_reactivations reads zero on both, and a healthy tree retaining live data legitimately reports stranded passes. On the live estate this hid a tree at 26 stranded passes, 0 reclaimed and 0 reactivation attempts beside siblings at 45 to 50. That tree's orleans.lattice.wal.entries_trimmed is zero lifetime, so its ~91 MB is permanently unreleasable - but it is not necessarily growing at any moment, so alert on the blocked arm rather than on bytes, which reads green on a wedged tree that has gone flat. Both arms are zero-primed per tree on every sweep that reaches the classifier, before either is charged, so three readings are distinguishable: absent series means the classifier never ran here, both arms static at zero mean it ran and no pin constrained an offset floor, and a climbing blocked arm is the wedge. The resolved arm is charged only when an offset floor exists. Reports admission, deliberately not outcome
orleans.lattice.wal.gc.floor_holder_offset_admission counter ({candidate}) tree, status, tenant Replication WAL GC floor-holder admission WIDTH on the offset axis (issue #3310). How many candidates a classifying sweep admitted, split by why. at_floor is the population admissible before issue #3310 - checkpoint offset equal to the tree's offset floor. above_floor is the population it added - candidates above a floor whose own holder was admitted on the same sweep. Chart it directly beneath floor_holder_admission, which is a per-sweep yes/no about the single floor-defining candidate and is blind to width by construction; this is the series that says how many levels the floor can walk per sweep. above_floor is the acceptance signal: on a tree with spread pin offsets it sitting at zero while floor_holder_admission reads admitted means the widening is present and inert. Why it matters: the admitted set used to be "every pin on exactly one offset", so width tracked pin offset distribution rather than any budget. On one live silo, same binary and sweep, a single-offset tree admitted 166 per sweep while a nine-offset tree admitted 1 to 3, drove 0.185 leaves/minute against a 32 touch budget, and reached 120.8% of its WAL ceiling with zero decreasing intervals over sixty minutes. Spread is normal for continuous ingest, so the gate starved the trees that needed it most and no configuration could widen it. Alert on above_floor flat at zero while wal_gc_trim_stop_total{reason="offset_floor"} climbs - that pair is the wedge this instrument exists to expose; do not alert on bytes, which read green on a tree that has gone flat. Both arms zero-primed per tree per classifying sweep before either is charged, so an absent series is distinguishable from a measured zero. Bounded by a per-sweep budget scaled from the tree's floor-holding pin population. Reports admission, deliberately not outcome
orleans.lattice.wal.gc.coverage_unknown_pin_offset counter ({pin}) tree, partition, status, tenant Replication WAL GC coverage-unknown floor holders by durable offset availability. Splits the checkpointed_coverage_unknown arm of blocking_pin_state by whether the candidate carried a usable durable offset (issue #3199), which that instrument cannot express: it records the classifier's verdict and nothing about the candidate, and both floor-holder sample lists are recorded through the same call. offset_usable (offset >= 0) is the population the offset floor is a minimum over and the only one issue #3178's liveness drive can admit; offset_absent constrains no floor and can never satisfy that gate, so it is the triple exclusion of issue #3199 actually occurring. Both arms zero-primed per classified (tree, partition), so an empty offset_absent slice is a measured absence - the property that makes the reading conclusive either way. sum by (tree) here equals blocking_pin_state{status="checkpointed_coverage_unknown"}; a divergence is a defect in one of the two
orleans.lattice.wal.gc.never_checkpointed_pin_offset counter ({pin}) tree, partition, status, tenant Replication WAL GC never-checkpointed floor holders by durable offset availability. Splits the never_checkpointed arm of blocking_pin_state by whether the candidate carried a usable durable offset (issue #4198), which that instrument cannot express: it records the classifier's verdict and nothing about the candidate, and both floor-holder sample lists are recorded through the same call. offset_absent (offset -1) is the benign blocking sentinel, which constrains no offset floor and clears once the leaf checkpoints; offset_usable (offset >= 0) sits in the offset-bearing sample and can hold the tree's offset floor, where its refusal is terminal for the whole tree - the permanent-wedge shape of issue #3258, since issue #3310's prefetch is gated on the floor's own holder having been admitted. Mirrors coverage_unknown_pin_offset and reuses its vocabulary. Both arms zero-primed per classified (tree, partition), so an empty offset_usable slice is a measured absence rather than silence. sum by (tree) here equals blocking_pin_state{status="never_checkpointed"} summed over partition; a divergence is a defect in one of the two. Read beside floor_holder_admission to separate a latent holder from one actually blocking the tree
orleans.lattice.wal.gc.pass.reach counter ({pass}) stage, tree (reserved _none_), tenant (platform sentinel) Replication Which region of a WAL GC pass was reached, so that a flat wal.gc.passes or wal.gc.interval series can be read as measured rather than as never-executed (issue #3075). Both of those instruments are recorded inside the per-tree collection, which a pass reaches only after clearing several earlier exits; the standing remedy of siting an instrument outside the failing region is unavailable because the quantity being counted only exists inside it. Eight stage arms: pass_entered and the seven terminating exits registry_cancelled, registry_failed, registry_timed_out, loop_cancelled, no_due_tree, pass_completed_immediate, pass_completed_scheduled. registry_timed_out is our own enumeration budget firing, kept apart from registry_failed so that a bound we chose cannot present as a wedged registry. This is a separate instrument from orleans.lattice.wal.gc.tree.reach because the two halves have different attribution subjects and an instrument may only have one: a pass spans the whole registry and belongs to no tenant, while a tree visit belongs to the tenant owning the tree exactly as wal.gc.passes does. Carrying both under one name would split its series across two attribution rules, so a tenant-scoped query would return the per-tree arms and silently drop the pass-level ones - the very arms that answer whether the scheduler is running. The reserved tree value is structurally required rather than tidy: two of the exits are the catch arms of the registry enumeration itself, where obtaining the tree list is the operation that failed, so no tree id exists and none ever can. It takes the underscore-delimited form (matching _platform_) because a tree id is a caller-supplied string and a tree could legitimately be named none. Do not filter this panel by tree or tenant: either clause hides every arm. pass_entered is the denominator; at quiescence the seven exit arms sum to it, and the completeness of that sum is the guard against an early return added later without an arm. While a pass is in flight the entry arm legitimately leads the exits by the number of passes running, so the production relation is 0 <= entered - sum(exits) <= passes in flight - asserted as an equality outside quiescence it flaps once per pass and gets muted, at which point the guard is gone and nothing says so. pass_completed_immediate means the next tree was already due when the pass ended, i.e. the scheduler is running flat out, which is an operational signal in its own right and is why it is not folded into pass_completed_scheduled. Every arm advances by one rather than priming a zero, which is the whole design and the one way it departs from the partition="none" layer it generalises: Add(0) is idempotent on a counter, so a series primed once and a series primed ten thousand times are byte-identical and no sample count exists to separate them - priming can establish that a region was reached at least once and nothing more, which makes a priming layer an instance of the very defect class it was prescribed to cure. Read rates here, not only totals. Diagnostic only: it never changes what a pass is allowed to trim, and it does not change what any existing WAL GC instrument counts
orleans.lattice.wal.gc.tree.reach counter ({tree}) tree, stage, tenant Replication The per-tree half of the WAL GC reachability layer (issue #3075), tenant-derived exactly like its siblings wal.gc.passes and wal.gc.interval. Two stage arms, tree_seen and tree_collected, and they are a pair - neither half is interpretable alone. tree_seen fires for every tree the pass enumerated, due or not; tree_collected fires above every exit of the per-tree collection and is the arm that licenses reading a flat wal.gc.interval or wal.gc.passes for that tree as measured rather than never-executed. Their difference is the set skipped by the not-yet-due check, which is the healthy majority on any pass rather than a fault - so a large gap is expected, and a zero gap on a many-tree silo is the anomaly. Chart the two arms together and never sum them with the pass-level instrument, whose arms are scheduler-wide and count a different population. Both arms advance rather than priming a zero, for the reason given on wal.gc.pass.reach. Diagnostic only: it never changes what a pass is allowed to trim, and it does not change what any existing WAL GC instrument counts
orleans.lattice.wal.gc.interval histogram (s) tree, tenant Replication WAL GC adaptive interval
orleans.lattice.wal.gc.scheduler_backoff histogram (s) cause, tenant Replication WAL GC scheduler-wide backoff by cause (p95, s) - issue #3064. Scheduler-wide WAL GC backoff currently in force, by cause (issue #3064). Not the same quantity as wal.gc.interval beside it: that is one tree's adaptive cadence, this is the whole scheduler's retry wait, so it is silo-scoped and carries the platform tenant sentinel and no tree. Chart it beside wal.gc.passes as the answer to "is WAL GC asleep or dead?" - a pass that collects nothing writes no per-tree series, so every other WAL GC line freezes identically in both cases and no amount of staring at them separates the two. cause=scheduled reports the floor and means no backoff is in force; cause=faulted is the registry enumeration throwing and is capped at the reactivation block age so recovery is bounded; cause=empty is a successful enumeration of an idle silo and relaxes to the configured WalGcInterval. The two ladders have disjoint ranges above the fault ceiling, so a high reading is self-identifying before the tag is read. Recorded unconditionally on every pass, including the first and including a zero-tree silo, so the panel carries no or vector(0) compensation and a missing line means the build did not land - deliberately not primed on the fault path, where an absent series would have meant either "no fault" or "not deployed"
orleans.lattice.wal.gc.scheduler_consecutive_faults histogram ({fault}) cause, tenant Replication WAL GC consecutive registry-enumeration faults (p95) - issue #3064. Consecutive failed WAL GC registry enumerations, reset to zero by the first success (issue #3064). A current streak, not a lifetime total, which is why it is a histogram - a counter cannot go back down. Same unconditional per-pass site and same cause tag as wal.gc.scheduler_backoff, so a zero is a measured zero and an absent series names the build. A streak that keeps returning to zero is transient fault absorption; one that only climbs is a registry the scheduler cannot read, and is the operator-actionable alarm
orleans.lattice.wal.gc.backlog_bytes histogram (By) tree, tenant Replication WAL GC backlog after pass. The WAL bytes the tree still occupies after the pass - its physical occupancy, dead but not yet compacted bytes included, falling back to the retained byte count on a provider without physical accounting - taken from the pass's own byte sample at no extra I/O. Recorded on every completed pass that took a byte sample on a provider that accounts bytes: a pass samples when WalMaxRetainedBytes or the durability hold (WalDurabilityHoldCeilingBytes, on by default) has a positive ceiling, so with the default options every pass on such a provider records it. On any other completed pass orleans.lattice.wal.gc.backlog_bytes_unavailable names the reason instead (issue #2694)
orleans.lattice.wal.gc.backlog_bytes_unavailable counter ({pass}) tree, reason, tenant Replication WAL GC backlog bytes unavailable by reason. A completed pass that recorded no orleans.lattice.wal.gc.backlog_bytes sample: policy_disabled when WalMaxRetainedBytes is unset, provider_unsupported when it is set but the WAL provider reports no byte accounting. The reason is decided on that byte-pressure ceiling alone, while the byte sample is also taken for the default-on durability hold, so with the default options a provider that accounts bytes never advances this counter and one that cannot is reported as policy_disabled; that arm means byte sampling was switched off only when the durability hold is off as well (issue #2694)
orleans.lattice.wal.gc.ceiling_unsatisfiable counter ({pass}) tree, tenant Replication WAL GC ceiling unsatisfiable - the tree's configured WalMaxRetainedBytes is below WalMaxRetainedBytesWorkingSetMultiple (2) times its measured logical working set, so it is arithmetically unreachable by a healthy tree (issue #3242). Chart it directly beside the pass-outcome panel, because it is the discriminator that panel cannot carry. stranded and over_ceiling both say bytes could not be reclaimed, which reads as a lagging consumer; this says the ceiling itself cannot be met at this size, whose remedy is to raise the ceiling rather than chase consumers. Deliberately not an arm of wal.gc.passes: those arms partition invocations and this condition co-occurs with reclaimed, over_ceiling and stranded alike, so an arm would have stolen invocations from whichever arm names them and an alert on stranded would have gone silent as the condition worsened. Every pass on an affected tree agrees, because this is a standing property of the configuration rather than an event, so the rate tracks the tree's pass rate. Zero-primed per tree, so the panel carries no or vector(0) compensation and an absent series means WAL GC is not running for that tree on the selected silo. A measured zero means "not provably unsatisfiable" - it also reads zero with no ceiling configured, with a provider that cannot account logical bytes, and on an empty tree
orleans.lattice.wal.gc.terminal_breach counter ({pass}) tree, tenant Replication WAL GC terminal breach - the tree has been over its byte ceiling, with an available cursor floor, reclaiming nothing, for 10 consecutive passes (issue #3149). Chart it beside the pass-outcome panel: over_ceiling is a transient the scheduler keeps retrying and reads the same on every pass, so this is the only series that says the retries have stopped helping. Counted once per qualifying pass from the tenth onwards, so the rate tracks the tree's pass rate while the breach persists; any pass that reclaims, falls under the ceiling, or lacks an available cursor floor (blocked, or no cursor reported) resets the run. Zero-primed per tree, so the panel carries no or vector(0) compensation and an absent series means WAL GC is not running for that tree on the selected silo. Read it with the floor-to-head distance panel, which sizes what the floor is holding.
orleans.lattice.wal.gc.offset_floor_population_gap counter ({consumer}) tree, tenant Replication WAL GC offset floor population gap (pinned consumer never reported an offset)
orleans.lattice.wal.gc.offset_floor_unavailable counter ({pass}) tree, tenant Replication WAL GC offset floor unavailable (pin store unreachable)
orleans.lattice.wal.gc.durability_hold_forced counter tree, tenant, reason Replication WAL GC trimmed a tree with no durable materialiser offset floor despite the durability hold (LatticeOptions.WalDurabilityHoldCeilingBytes, on by default at 256 MiB) (issue #3300). Not a tuning hint: a non-zero value states that WAL entries nothing is known to have durably applied were discarded. Alert on any increase. reason says which way the hold failed and the two arms need different repairs: ceiling_exhausted means the hold engaged, retained to its ceiling and yielded - fix the materialiser, since raising the ceiling only delays it; unmeasurable_footprint means the hold never engaged because the provider reports no retained bytes, leaving nothing to bound it - fix the provider, since the ceiling is irrelevant to this arm. Pair the panel with orleans.lattice.wal.gc.durable_floor_stall_seconds, which says how long the underlying cause has been present. Zero is good news because the hold is on by default; it reverts to meaning the check is off only if the ceiling is explicitly set to 0.
orleans.lattice.wal.gc.durable_floor_stall_seconds histogram (s) tree, status, tenant Replication Seconds since a tree's durable materialiser offset floor last advanced (issue #3300). The panel that answers "is anything actually being made durable for this tree", which no volume counter can: a count of checkpoint attempts rises just as happily when none of them land. Three status arms: advanced (the floor rose on this pass, recorded at zero), stalled (a floor exists and did not rise, recorded at its age) and absent (no floor exists at all, recorded at the age since this process first saw the tree - deliberately never zero, because reporting the worst state on the healthiest reading is what hid issue #3300 for eleven hours). Progress is measured against a high-water mark, so a floor falling back - legitimate, since it is a minimum over reporting leaves - does not reset the clock. Alert on a sustained absent or stalled age; a healthy tree sits on advanced at zero.
orleans.lattice.wal.gc.trim_stop counter ({scan}) tree, shard, reason, tenant Replication WAL GC trim stops by reason - why each shard scan stopped. The panel that makes a zero-reclaim pass attributable: offset_floor, cursor_floor, causal_frontier and block_pin are the distinct causes of reclaiming nothing and each indicts a different subsystem (a lagging leaf checkpoint, a lagging consumer cursor, a replication origin whose stable frontier has not advanced, and a buffering receiver holding a pin), while exhausted and empty are the innocent stops. The last three were a single not_eligible arm until issue #3155 split them. Issue #3300 added three durability-axis arms beside them: durability_unverified (a non-empty shard released with no durable offset floor at all, previously indistinguishable from exhausted), durability_hold (the same absent floor, but the shard retained under the durability hold, LatticeOptions.WalDurabilityHoldCeilingBytes), and durable_offset_refusal (the durable offset floor overruled an entry the consumer cursor would have admitted, so a sustained run with nothing reclaimed names a leaf whose durable checkpoint or snapshot coverage has stopped advancing). Chart against the pass-outcome panel: over_ceiling is decided by the cursor floor alone, so a tree stranded by the offset floor reports a healthy floor there and only this panel names the cause. Break out by shard and join against orleans.lattice.wal.entries_trimmed to tell a shard that stops after releasing its eligible prefix from one that has never released anything (issue #3207): the stop arm advancing while that shard's entries-trimmed stays flat is the wedged reading, and it is the only one available, because a shard that has never trimmed reports zero dead bytes on every compaction arm.
orleans.lattice.wal.gc.floor_head_distance histogram ({offset}) tree, shard, tenant Replication WAL GC floor-to-head distance per shard (issue #3149) - how many offsets each shard's effective trim floor is holding between its first retained entry and its head. Zero on a shard whose scan released everything or found nothing, so a large value on one shard beside zeros on its siblings names the stranded shard, and a large value that keeps growing alongside the terminal-breach panel is unbounded retention with a measured size. Recorded on every scanned shard, zeros included, so the query needs no or vector(0) fallback.
orleans.lattice.wal.gc.durability_hold_engaged counter tree, tenant, reason Replication WAL GC passes on which the durability hold engaged and retained the scan (issue #3300). The mirror of durability_hold_forced: that counts the hold yielding, this counts it working, so a non-zero value here is the safeguard doing its job and is not an alert on its own. It fires only where every cursor admitting the trim is a leaf materialiser with no durable offset coverage; a tree with a shipper, view maintainer, log subscriber or backup capture reporting a cursor never reaches it. reason decides whether anyone needs to act: never_pinned is a stalled tree that will hold until its ceiling forces it onto durability_hold_forced and needs a materialiser wired or repaired, while pin_regressed is a bounded transient (rolling upgrade, leaf churn) that clears itself as the leaves re-pin. cursor_unreadable is neither: the cursor registry did not answer, so the pin history was never measured on that pass (issue #3366), and it points at the registry read path in the collector rather than at the tree. Alert on sustained never_pinned and on sustained cursor_unreadable, not on pin_regressed. Pair with orleans.lattice.wal.gc.durable_floor_stall_seconds for how long the condition has persisted.
orleans.lattice.wal.gc.scheduler.phase_age observable gauge (s) phase, tree, tenant Replication WAL GC scheduler phase age - where the silo-wide sweep loop is parked and for how long. The spine of the liveness set: an age exceeding the selected cadence is a stalled sweep, and the phase arm names where it stalled. Present from process start, so its absence dates the build rather than describing the scheduler.
orleans.lattice.wal.gc.scheduler.passes_started counter ({pass}) tenant Replication WAL GC scheduler pass heartbeat, counted above the registry enumeration. Chart against phase_age: climbing here while every per-tree wal.gc.* series is frozen is a loop alive and failing, not a stopped one.
orleans.lattice.wal.gc.scheduler.pass_duration histogram (s) tenant Replication WAL GC scheduler actual pass duration. Chart alongside scheduler.wait: divergence between selected cadence and elapsed pass is the signature of passes stretched by something the cadence policy never observed.
orleans.lattice.wal.gc.scheduler.enumerations counter ({enumeration}) outcome, tenant Replication WAL GC scheduler tree-registry enumeration outcome by arm. all_blank and empty are charted separately from succeeded because a blank-id registry reports success and collects nothing; timed_out is charted apart from faulted so a limit of ours never reads as a failed subject; note it does not distinguish this scheduler's own bound from a timeout raised inside the awaited call, which only the abandonment log line's measured elapsed separates.
orleans.lattice.wal.gc.scheduler.terminations counter ({termination}) reason, tenant Replication WAL GC scheduler loop termination by reason. Any non-zero value is terminal for the life of the process, so this belongs on an alert rather than only on a chart.
orleans.lattice.wal.gc.scheduler.wait histogram (s) tenant Replication WAL GC scheduler selected inter-pass wait. A reading climbing geometrically toward the configured ceiling is a loop retrying after failed passes, which is the one positive signature an alive-but-failing loop emits.
orleans.lattice.admission.live_keys observable gauge ({key}) tree, tenant Overview Admission - live keys by tree
orleans.lattice.admission.estimated_bytes observable gauge (By) tree, tenant Overview Admission - estimated bytes by tree
orleans.lattice.admission.over_advisory observable gauge (1, 0/1) tree, tenant Overview Admission - trees over advisory ceiling
orleans.lattice.admission.would_reject counter ({write}) tree, dimension, tenant Overview Admission - would-reject rate (advisory dry-run)
orleans.lattice.admission.utilization observable gauge (1, ratio) tree, dimension, tenant Overview Admission - utilization by dimension
orleans.lattice.admission.rejected counter ({write}) tree, dimension, tenant Overview Admission - rejected write rate (enforced)
orleans.lattice.lock.acquired counter ({acquire}) outcome, tenant Overview Distributed lock - acquire / release / reclaim rate
orleans.lattice.lock.released counter ({release}) tenant Overview Distributed lock - acquire / release / reclaim rate
orleans.lattice.lock.lease_reclaimed counter ({lease}) tenant Overview Distributed lock - acquire / release / reclaim rate
orleans.lattice.lock.acquire.wait histogram (ms) tenant Overview Distributed lock - acquire wait latency
orleans.lattice.atomic_action.completed counter ({saga}) outcome, tenant Overview Atomic action - saga and step rate
orleans.lattice.atomic_action.step counter ({step}) phase, outcome, tenant Overview Atomic action - saga and step rate
orleans.lattice.atomic_action.duration histogram (ms) outcome, tenant Overview Atomic action - saga duration
orleans.lattice.grain.call.outstanding_depth histogram ({call}) grain_type, tenant Overview Outstanding calls per target activation, by grain type
orleans.lattice.grain.call.duration histogram (ms) grain_type, outcome, tenant Overview Grain-call duration by grain type and outcome
orleans.lattice.shard_root.forward.timeouts counter ({timeout}) tree, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.shard_root.scan_page.chain_regressions counter ({regression}) tree, shard, outcome, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.shard_root.scan_page.stalls counter ({stall}) tree, shard, phase, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.shard_root.scan_page.ceiling_outcomes counter ({fire}) tree, shard, outcome, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.shard_root.scan_page.leaf_read_outcomes counter ({read}) tree, shard, outcome, tenant CommitPath Scan-page leaf-read coalescing (issued, joined, served)
orleans.lattice.shard_root.scan_page.zero_progress_stalls counter ({stall}) tree, shard, outcome, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.scan.stall_resumptions counter ({resumption}) tree, phase, outcome, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.scan.stall_futility_outcomes counter ({outcome}) tree, phase, outcome, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.shard_root.flush.retries_suspended counter ({suspension}) tree, shard, kind, tenant CommitPath Shard-root wedge guards (forward timeouts, scan-page stalls, scan resumptions, and flush suspensions)
orleans.lattice.saturation.refusals counter ({refusal}) tree, source, arm, tenant CommitPath (3761, 3921) LatticeSaturatedException refusals by source (rate): each refusing seam records its own refusal, so a writer-side refusal that surfaces through an atomic-write saga is counted under both wal_admission and atomic_write_saga, and a saga refused because its quiesce budget elapsed is counted twice under atomic_write_saga; background starvation drives refused a replay permit are also counted under replay_permit_admission (counted, not raised; issue #3761). A replay_permit_admission refusal also carries arm (wait_exceeded, no_progress or gc_share), charted on its own panel because the first two have opposite remedies (issue #3921)
orleans.lattice.wal.writer.append.admission_saturation_refusals counter ({refusal}) tree, partition, tenant CommitPath WAL writer admission & dispatch (rate)
orleans.lattice.wal.writer.append.admission_timeouts counter ({timeout}) tree, partition, tenant CommitPath WAL writer admission & dispatch (rate)
orleans.lattice.wal.writer.append.dispatched counter ({dispatch}) tree, partition, tenant CommitPath WAL writer admission & dispatch (rate)
orleans.lattice.wal.writer.append.drain.releases counter ({release}) tree, partition, tenant CommitPath WAL writer admission & dispatch (rate)
orleans.lattice.wal.writer.append.admission_wait histogram (ms) tree, partition, tenant CommitPath WAL writer admission wait p50/p95/p99
orleans.lattice.wal.writer.partition.pending_appends histogram ({dispatch}) tree, partition, tenant CommitPath WAL writer partition pending appends
orleans.lattice.wal.shard.pending_segments histogram ({segment}) tree, shard, tenant CommitPath WAL shard backlog
orleans.lattice.wal.shard.deactivate.in_flight histogram ({slot}) tree, shard, tenant CommitPath WAL shard backlog
orleans.lattice.wal.shard.drain.budget.force_faulted_slots histogram ({slot}) tree, shard, tenant CommitPath WAL shard backlog
orleans.lattice.wal.shard.drain.budget.expirations counter ({expiration}) tree, shard, tenant CommitPath WAL shard drain budget & flush calls
orleans.lattice.wal.shard.start_flush.calls counter ({call}) tree, shard, tenant CommitPath WAL shard drain budget & flush calls
orleans.lattice.wal.append_dispatch.timeouts counter ({timeout}) tree, shard, tenant CommitPath WAL flush / dispatch timeouts
orleans.lattice.wal.flush.preflight.timeouts counter ({timeout}) tree, shard, tenant CommitPath WAL flush / dispatch timeouts
orleans.lattice.provider.phase2.commit.timeouts counter ({commit}) tree, shard, tenant CommitPath Provider commit timeouts & retry short-circuits
orleans.lattice.provider.retry.short_circuited counter ({attempt}) status, tenant CommitPath Provider commit timeouts & retry short-circuits
orleans.lattice.provider.phase1.transient_retries counter ({attempt}) tree, shard, phase, tenant CommitPath Storage-provider retries (attempts vs exhausted vs idempotent-replays vs phase1-transient)
orleans.lattice.shard.digest_reads counter ({op}) tree, shard, tenant CommitPath Digest reads & publish timeouts
orleans.lattice.internal.digest_publish.timeouts counter ({timeout}) tree, tenant CommitPath Digest reads & publish timeouts
orleans.lattice.shard_root.reshard.initiated counter ({reshard}) tree, tenant CommitPath Reshard activity
orleans.lattice.shard_root.reshard.completed counter ({reshard}) tree, tenant CommitPath Reshard activity
orleans.lattice.shard_root.reshard.rejected counter ({rejection}) tree, reason, tenant CommitPath Reshard activity
orleans.lattice.shard_root.activation_ready.timeouts counter ({timeout}) tree, tenant CommitPath Reshard activity
orleans.lattice.shard_root.reshard.in_flight histogram ({reshard}) tree, tenant CommitPath Reshard runs in flight
orleans.lattice.materialiser.pin.durable_writes counter ({write}) tree, outcome, tenant CommitPath Leaf-materialiser durable pin path (issue #1030)
orleans.lattice.materialiser.pin.advances counter ({report}) tree, outcome, tenant CommitPath Leaf-materialiser pin advancement by outcome (issue #2694)
orleans.lattice.leaf.activation_replays counter ({replay}) tree, activation_temperature, tenant CommitPath Leaf-materialiser durable pin path (issue #1030); the cold/warm arms give the activation-temperature ratio (issue #2148)
orleans.lattice.wal.replay.permit_adaptations counter ({permit}) outcome, trigger (withheld arm only), tenant CommitPath Memory-adaptive backpressure on the per-silo WAL replay concurrency gate, by withheld/restored (issues #2781, #2862). Withholding has two triggers - proactively when a replay finishes with managed heap occupancy at or above 75% of the GC hard limit, and reactively when a replay fails for memory pressure - while restoring requires a clean replay and occupancy back under 60%, so the arms trace a hysteresis band. The withheld arm carries trigger (occupancy/fault) to separate those two producers (issue #2883); before it did, withheld was a sum no scrape could attribute, and run 12's withheld = 6 was published as evidence for the reactive trigger when it was equally consistent with that trigger firing zero times. Summing over trigger recovers the untagged total, so pre-tag comparisons stay valid. The restored arm carries no trigger on purpose: withheld permits are fungible, so withheld{trigger=X} - restored is meaningless and only the total withheld - restored is a valid level. Charted twice: as a rate, to show how hard the gate is oscillating, and as the level withheld - restored, which is the number of permits currently out of circulation and so the actual reduction in leaf-activation concurrency. Both arms share one instrument so a zero on restored is a measured zero rather than an unpublished series, and all three series - each trigger on the withheld arm, plus restored - are primed at zero when the gate is sized, so an absent series means the build did not land. Deliberately carries no tree tag - the gate is process-wide, and a per-tree tag would invite summing arms that share one resource. Chart beside leaf.activation.failures{reason="canceled_awaiting_permit"}, which is the population this reduction exists to stop starving
orleans.lattice.wal.replay.permits_withheld observable gauge ({permit}) tenant CommitPath Live count of replay permits currently withheld from the per-silo WAL replay concurrency gate by memory backpressure (issue #2784). The CommitPath panel that charts adaptation already renders this level as withheld - restored from the counter arms; this series publishes it directly, which is not cosmetic. Both counter arms are process-lifetime, so a restart resets them together and the derived difference silently returns to zero with nothing in the series marking the discontinuity - adaptation a restart discarded and adaptation that never happened render identically. The gauge supplies that boundary: the last sample before the gap is the discarded level, and the step to zero is the discard. No tree tag, matching the sibling counter, because the gate is a process-wide static and a per-tree series would repeat one figure under many names. Not a replacement for the counters - a sampled gauge misses a withhold and restore that fall between scrapes, so the counters remain the record of how often the mechanism fired
orleans.lattice.wal.replay.permit_ceiling observable gauge ({permit}) tenant CommitPath Ceiling the per-silo WAL replay concurrency gate was sized to, or 0 before the first activation sizes it (issue #3047). Charted alongside permits_available and permits_queued on the replay-gate occupancy panel, because none of the three is interpretable alone: this one is the denominator for both the withheld level and the available level, and a SemaphoreSlim does not expose its own maximum, so without it a reader cannot tell one permit withheld from sixteen (noise) from one withheld from two (half the silo). Zero is an unambiguous not-yet-sized sentinel, never a sized gate
orleans.lattice.wal.replay.permits_available observable gauge ({permit}) tenant CommitPath Permits currently available on the per-silo WAL replay concurrency gate (issue #3047). Charted against permit_ceiling on the same panel deliberately: a zero here is ambiguous between an unsized gate and a saturated one, and only the ceiling separates them. Saturation is the reading the panel exists to show. It does not chart queue depth - the underlying count saturates at zero and cannot distinguish one waiter from a thousand - for which the companion series on the same panel is permits_queued
orleans.lattice.wal.replay.permits_queued observable gauge ({activation}) tenant CommitPath Activations currently blocked waiting for a permit on the per-silo WAL replay concurrency gate (issue #3047). The only series on the panel that can show a gate admitting nothing: the queue-wait histogram and the replay counters all record at a terminal outcome downstream of the wait, so an activation that never acquires and is never cancelled is invisible to every one of them, and a permanently saturated gate renders identically to a tree that never activated a leaf. Read with the other two on the panel - ceiling is the width of the door, availability is whether it is open, this is the queue outside it
orleans.lattice.wal.replay.permits_served counter ({permit}) tenant CommitPath (3923) The per-silo WAL replay concurrency gate's service rate (rate): one increment per permit hold that ended, returned or withheld (issue #3921). Zero-primed when the gate is sized, so a flat zero while activations are queued is a measured stall, the regime the no_progress refusal arm reports
orleans.lattice.wal.replay.permit_hold histogram (ms) tree, tenant CommitPath (3922) Mean time a replay held a permit on the per-silo WAL replay concurrency gate (issue #3921). Read against the service rate: hold time rising with the ceiling while the service rate stays flat is a store-bound gate, where raising WalMaterialiserMaxConcurrentReplays makes the no_progress refusal arm fire harder. Not primed, so absent means no hold ended
orleans.lattice.wal.replay.permit_queue_wait histogram (ms) tree, tenant, outcome CommitPath Time spent queued on the per-silo WAL replay concurrency gate, by acquired/canceled (issue #2873). The discriminator for leaf.activation.failures{reason="canceled_awaiting_permit"}, which counts where an activation was cancelled and cannot say why: a canceled arm whose wait approaches the request budget is genuine permit starvation, while a canceled arm at or near zero is upstream budget exhaustion arriving already doomed. The counter is identical in both cases, so the duration is the only separator. Chart it as rate(_sum) / rate(_count), never as a quantile: this container's exposition renders a histogram as a Prometheus summary with _sum and _count and no _bucket, so histogram_quantile returns nothing and the or vector(0) repair would manufacture a literal zero indistinguishable from the genuine near-zero wait that is one of the two readings. The mean suffices because the regimes differ by orders of magnitude. Since issue #3044 the instrument declares explicit InstrumentAdvice bucket boundaries in-process, and that does NOT change the guidance above for this container. Advice is a .NET-level hint that only an exporter which reads it can honour; the deployed exposition renders every one of its 98 histogram families as a bucketless summary (no _bucket, and no quantile= either), so the boundaries are real in the process and absent on the wire. The advice is therefore necessary and not sufficient - it makes buckets available to an exporter that honours them, such as the OpenTelemetry AddPrometheusExporter used by samples/McpTelemetry, and changes nothing under the current adapter. Read the mean here until a scrape actually shows _bucket series. Carries a tree tag although its permit_adaptations sibling deliberately does not: a ceiling is process-wide, a wait is one caller's own and is attributable. Not primed at zero, because a primed histogram sample is a 0 ms datum reading as "instant" rather than a neutral marker, so an absent series is uninterpretable and must be corroborated against the primed permit_adaptations arms
orleans.lattice.wal.replay.permit_waits_in_flight observable gauge ({activation}) tree, tenant CommitPath Activations currently queued on the per-silo WAL replay concurrency gate (issue #3044). Exists because permit_queue_wait is structurally blind to a wait that never terminates: both of its arms are terminal - acquired after the semaphore is entered, canceled from the catch around the wait - so an activation parked on a saturated gate forever writes to neither, and the instrument goes quiet exactly when the gate is worst. No amount of zero-priming reaches that, because a primed canceled arm reading zero asserts only that no cancellation completed; the missing observable is a level, not a terminal event. Primed at zero per tree once that tree has queued at least once in this process, so a flat zero is a measured zero. The priming has a boundary that is part of the reading: it covers trees this process has observed queueing, not every tree in the estate, so a zero means "has queued before, is not queued now" and never "cannot queue". Carries tree for the same reason permit_queue_wait does and permit_adaptations does not - a ceiling is process-wide, a wait is one caller's own. Useless alone: read with permit_wait.oldest_age, since a count of three cannot separate three healthy 200ms waits from one parked for twelve minutes
orleans.lattice.wal.replay.permit_wait.oldest_age observable gauge (s) tree, tenant CommitPath Age of the oldest activation currently queued on the per-silo WAL replay concurrency gate (issue #3044). The half of the pair that separates busy from stuck: a steady count with this age oscillating near zero is a gate serving a queue promptly, while the same count with this age climbing without bound is an activation that will never be admitted - the shape that closes the loop #2691 measured, since a leaf that cannot get a permit cannot replay, cannot bank a snapshot, cannot advance the durable materialiser pin, so WAL GC finds the pin unusable and reclaims nothing. No threshold is encoded; "stuck" is read off the slope against real data, because the honest threshold differs per host and per WAL size. Deliberately NOT primed, unlike its sibling count, and the asymmetry is the design: the age of the oldest waiter when there is no waiter is undefined rather than zero, and a fabricated zero would publish the healthiest possible value for the emptiest possible state - making an idle gate indistinguishable from one whose waits are all instant, which is one of the two conclusions the pair exists to choose between. An absent series here is therefore correct and means nothing is queued; corroborate against the primed count, which does distinguish a measured zero from a build that did not land
orleans.lattice.wal.replay.slice_narrowings counter ({narrowing}) tree, partition, tenant CommitPath Activation-time replay slice-width narrowings forced by memory pressure on a commit-log read (issue #2867). The second of the two factors that set peak replay memory, and the only one no operator can configure. Peak memory is the product of how many replays run at once and how much each buffers: the first factor is the WalMaterialiserMaxConcurrentReplays ceiling, surfaced on the container overlay and already charted from both sides by wal.replay.permit_adaptations and wal.replay.permit_queue_wait; the second is the per-replay slice width, a private constant of 256 with no option behind it, so the reactive #2742 narrowing is the only thing that ever moves it. Chart beside leaf.activation.failures on the same tree, because together they separate two stories a failure count alone cannot: narrowings climbing with failures means the width is too coarse for the host and the narrowing is not keeping up, while failures climbing against a flat zero here means the allocation that failed was not the slice read at all and width is not the lever. Carries a tree tag for the same reason permit_queue_wait does and permit_adaptations does not - a ceiling is process-wide, but a buffer width is one replay's own and is attributable, which is what lets one tree be shown buffering harder than its siblings under identical cycling. Primed at zero per (tree, partition) when a partition replay begins: the healthy steady state is never to narrow, so without the prime the common case would be indistinguishable on the scrape from a build that cannot narrow, and an absent series here does mean the build did not land
orleans.lattice.wal.replay.starvation_drive_abandonments counter tree, tenant Replication WAL GC starvation drives abandoned after exhausting their StarvationDriveBudget, recorded grain-side as the drive gives up (issue #3065). The sole discriminator between a bounded-but-slow drive and one that would have parked forever. Bounding the drive necessarily destroys the diagnostic that found the defect: the frozen signature was attempted climbing with undelivered just behind it and drove_already_driving collecting the successors that bounced off the in-flight latch, and a merely-slow drive now reproduces that shape byte for byte, because the scheduler's touch still times out at the Orleans response-timeout default long before the drive's own budget elapses. None of the scheduler's arms can separate the two any more, and the drove_timed_out verdict arm is not a substitute, because it is recorded by a grain whose caller has usually already given up. Chart beside the starved-leaf drive panel: advancing means drives are hitting the budget and the remedy is upstream of Lattice in the host-supplied commit-log read, while flat-at-zero means a frozen-looking reactivation panel is ordinary sweep contention. Zero-primed per tree beside the attempted arm on the scheduler rather than at the grain's drive entry point, so an absent series means the build carrying the fix never landed and nothing else - priming at the drive would have made absence mean "no drive has ever run" as well, which is the ambiguity the prime exists to remove. Tagged tree and, like every instrument on this meter, the derived tenant label: the grain-side recording and the scheduler's zero-prime both carry it
orleans.lattice.leaf.activation_replays_over_budget counter ({replay}) tree, partition, tenant CommitPath Leaf-materialiser durable pin path: per-leaf post-filter replay cost over budget against an intact WAL (issues #1738, #2149)
orleans.lattice.leaf.activation_stalled_replays counter ({replay}) tree, partition, tenant CommitPath Leaf-materialiser durable pin path: leaf replay re-entered from a checkpoint that did not advance; fault arm, alert on persistence not appearance (issue #2285)
orleans.lattice.leaf.activation_cursor_publish_failures counter ({failure}) tree, tenant CommitPath Leaf-materialiser durable pin path (issue #1030)
orleans.lattice.leaf.deactivation.checkpoint_delta histogram ({offset}) tree, deactivation_reason, activation_temperature, tenant CommitPath Checkpoint offsets banked by an activation during graceful deactivation (issue #2280). LOWER BOUND, not a census: crash teardowns bypass the hook and a failed activation never reaches it. A zero on the cold arm is arithmetically forced, not symptomatic
orleans.lattice.leaf.activation.failures counter ({activation}) tree, activation_temperature, reason, tenant CommitPath Failures of a leaf's activation-time WAL replay (issue #2280), by reason: canceled (a replay in progress was cancelled), canceled_awaiting_permit (cancelled while still queued for the per-silo replay permit, before any replay began), canceled_resolving_options and canceled_rehydrating_snapshot (cancelled while resolving the tree's options or rehydrating the snapshot, both before the permit is requested, issue #2770), refused_replay_admission (refused a place in the replay permit queue by admission control, issues #3284 and #3290) or faulted (any other failure). Since issue #2871 the replay runs behind a barrier the activation arms and does not await, so the activation itself no longer fails: the replay's failure path counts the failure here and rethrows it to the data operation waiting on the barrier, and the barrier counts it again on orleans.lattice.leaf.replay_barrier_outcomes. Read it beside the deactivation histogram above
orleans.lattice.leaf.activation.cold_replay_loop counter ({activation}) tree, tenant CommitPath Cold activations cancelled at or past the consecutive-cancellation escalation threshold: the self-reinforcing cold WAL replay loop (issue #2280). A DEFECT signal, where leaf.activation.failures above is a COST signal - that counter is an aggregate and cannot separate one leaf cancelled five times from five leaves cancelled once. The count is consecutive and resets on any successful activation, so it measures leaf health rather than process age, and the threshold sits one above the highest value in the field measurement, so zero is the expected reading. Not tagged by leaf; identity, the consecutive count and the mid-replay/queued-for-permit/resolving-options split are on the paired warning
orleans.lattice.leaf.replay_barrier_outcomes counter ({replay}) tree, outcome, tenant CommitPath Terminal states of a leaf's deferred WAL replay, by completed/faulted/canceled (issue #2871). Since the replay no longer runs on the activation path, the activation succeeds and only its data operations fail. The replay's own failure path also counts the failure on leaf.activation.failures above, under a reason naming where it landed, so the two count failed replays from different sides: read them together and do not substitute one for the other. Zero-primed per tree at the arming site, so a flat zero on faulted is a measured zero rather than an absent series
orleans.lattice.leaf.unresolved_prepare_ledger_beyond_cap counter ({prepare}) tree, partition, tenant (not charted) Resident unresolved prepares recorded into the leaf's replay-work ledger once that ledger holds at least MaxDurableUnresolvedReplayWork entries (issue #2183). The threshold is inclusive (issue #2756): the prepare that fills the ledger to the cap is counted as well as every prepare recorded past it, because the cap is the resting size at which the ledger starts refusing to record deferred terminals. Persist risk: Azure Table grain storage rejects writes above its ~960 KB grain-state limit (see Storage provider per-row limits), bounding persisted row growth. Read risk: the larger SQLite local limit permits growth that can exhaust memory or the read budget during activation, before grain-level repair can run. Alert on every profile. Remains alert-only rather than a bundled trend panel; a normally quiet series is not evidence that growth is harmless
orleans.lattice.leaf.deferred_terminals_dropped_at_cap counter ({terminal}) tree, partition, tenant CommitPath Durable ledger records refused in pass 1 when capped or disabled; the terminals themselves are retained in memory and drained by pass 2 (issue #3190). Interruption can require re-reading unbanked work. Range deletes can saturate the ledger without sagas. This aggregate is neither per-leaf occupancy nor evidence of a stalled WAL floor; the uncapped prepare counter is not a substitute. Zero-primed on replay entry; no or vector(0) fallback. Shipped metric name and query unchanged.
orleans.lattice.leaf.span_fail_open_commits counter ({key}) tree, reason, origin, tenant CommitPath Keys a leaf committed locally although its declared span excludes them, because span admission found no neighbour to forward them to (issue #2125). reason is no_sibling or self_reference; origin separates cross_shard_migration (the deliberate graft) from client_write / merge (an accidental fall-back). Not pre-minted: the panel omits or vector(0), a healthy estate is flat at absent, and a missing line does not prove the build landed
orleans.lattice.leaf.snapshot.captures counter ({capture}) tree, outcome, tenant CommitPath Leaf-snapshot capture attempts by succeeded/failed/abandoned (issue #2696). Before this the capture path carried no instrument at all, so "every capture is failing" and "no capture has ever run" were the same reading - both left orleans.lattice.storage.snapshot_bytes at zero. Chart attempts as the sum across outcome; keep abandoned separate, because a graceful fleet shutdown abandons captures by design and would otherwise read as a provider outage
orleans.lattice.leaf.snapshot.capture.declines counter ({decline}) tree, reason, tenant CommitPath Capture invocations that declined before the attempt boundary, by no_tree_id/not_eligible/already_in_flight/no_coverage_claim (issues #2696, #2725). Kept off the outcome tag of the attempt counter so attempts and durations stay exactly co-populated. Chart it beside the attempt counter: a declined capture is a third state, distinct from a failed one and from an idle deployment, and is the most likely way a self-heal silently never runs. already_in_flight dominating is contention against the shared snapshot provider; not_eligible dominating is the starved-leaf population; no_tree_id above zero is a bug and carries no tree tag; no_coverage_claim is the benign unclaimable-rows population - live rows, never checkpointed, so a blob would claim no coverage and be refused by the load gate
orleans.lattice.leaf.snapshot.driver.declines counter ({decline}) tree, reason, tenant CommitPath (160, 161) Capture drivers that declined to drive a capture at all, by deactivate_unproven_coverage_stale/deactivate_unproven_coverage_current/deactivate_unproven_unclassified/recheck_cadence_not_reached/recheck_capture_in_flight/recheck_coverage_current (issue #3185), the coverage-lag timer's routing arms recheck_no_durable_checkpoint/recheck_checkpoint_stalled (issues #3300, #3389), and recheck_drive_refused/recheck_drive_deferred (issue #3575), which qualify a routing arm recorded on the same tick - the routed starvation drive was refused admission to the per-silo WAL replay gate, or skipped while the leaf backs off after a refusal - and so must not be summed with it. Distinct from orleans.lattice.leaf.snapshot.capture.declines above and must not be summed with it: those declines are raised inside the leaf's snapshot-capture routine itself and so partition its invocations exactly, whereas these are raised in a caller that never reached it, so folding them together counts a strict superset and corrupts the decline-versus-attempt ratio that counter exists to report. Chart deactivate_unproven_coverage_stale on its own - it is the minting rate of the frozen floor-holder population that blocks WAL reclamation, and it is the arm that matters operationally. Plot it against rate(orleans.lattice.wal.gc.blocked_leaf_reactivations): sustained above the drive's completion rate the frozen population grows without bound and no drive-side tuning converges, because the mismatch is asymptotic rather than a matter of constants; below it the drive is merely slow and tuning applies. The decline itself is correct behaviour and not a fault - a leaf that neither advanced a checkpoint over cache-resident applies nor cold-rebuilt its cache from the WAL start cannot honestly stamp coverage - so what is charted is a rate, not an error. Not zero-primed, unlike most counters on this map: a series appears on first decline, so an absent line does NOT establish that the build failed to land. Panel 160 is the by-arm breakdown and foregrounds the stale arm rather than plotting every arm together, with the two routing arms and their two qualifiers on targets of their own, because deactivate_unproven_coverage_current dominates on a healthy estate and would otherwise bury the one arm that carries signal. Panel 161 states the comparison as a single number, rate(stale) / rate(blocked_leaf_reactivations), with the divergence threshold at 1; read it as a level and not as a trend, since it is already a ratio of two rates. On a low-traffic tree both rates fall to zero, because the drive is reactivation-driven, so the quotient renders as 0 and reads as health on a tree that may be permanently stranded - a near-zero value there means no drive activity, not convergence, and strandedness on that population is visible only on orleans.lattice.wal.gc.passes{outcome="stranded"}, which advances with zero reclaim independently of tree size. Do not corroborate it against WAL byte totals: compaction reclaims DEAD bytes on a sawtooth whose envelope can fall while this ratio sits above 1, because the floor stall holds RETAINED bytes that no compaction arm may discard
orleans.lattice.leaf.snapshot.capture.duration histogram (ms) tree, outcome, tenant CommitPath Duration of a capture attempt, same boundary and tags as orleans.lattice.leaf.snapshot.captures so the two share a population (issue #2696). The activation itself no longer awaits a capture (issue #2871): the activation-time capture runs inside the deferred WAL replay the activation starts, which the leaf's data operations wait on, and captures also run when replay banks progress, on the periodic snapshot recheck, during zero-coverage repair and warm-cache rescue, and at graceful deactivation, so read it as capture cost across all of those paths rather than as activation latency. Exported with explicit buckets like every other ms histogram on this meter, so chart histogram_quantile over _bucket; an interval mean over delta _sum / delta _count is a useful companion but is not the only available form
orleans.lattice.leaf.snapshot.capture.concurrency_peak observable gauge ({capture}) tenant (platform sentinel) CommitPath Monotone high-water mark of how many leaf-snapshot captures ran concurrently on one silo (issue #2696), the cross-leaf quantity PR #2723 deferred and that no per-leaf series can show: the single-flight guard is per activation, so a thousand different leaves each passing their own guard is invisible to every other capture instrument. Chart it as a plain max over the raw series and not as a rate or a quantile. It never falls, which is the point - a spike is reported by the scrape that sees it and by every scrape after, so an instantaneous gauge panel that would have missed the transient cannot miss this one. Do not expect it to return to zero when load subsides; it is a peak for the life of the process and a restart is what clears it. Carries no tree tag by design - the maximum spans every leaf on the silo and the contended snapshot provider is silo-wide, so a per-tree split would report several numbers none of which is the depth the provider saw
orleans.lattice.leaf.deactivation.barrier.failures counter ({failure}) tree, reason, tenant CommitPath (163) Leaf deactivation barriers that threw and were contained, by digest_publish/checkpoint_flush/snapshot_capture/frontier_pin (issue #3366). A teardown runs those four barriers in order; they previously shared one try/catch, so the first to throw cancelled every barrier after it and the teardown reported nothing at all - which is how a faulting checkpoint flush silently suppressed the snapshot capture that would have preserved the leaf's rows, and so how a graceful restart destroyed recently-written durable entries. Each barrier now fails independently and is attributed here. Read reason as the remedy selector, because the arms are not interchangeable: checkpoint_flush is the arm implicated in #3366 and means durable coverage is behind in-memory state; snapshot_capture is strictly worse, because the blob carrying the uncheckpointed tail was not written and those rows do not survive the process; digest_publish and frontier_pin are internally best-effort and realistically cannot throw, so a non-zero value on either is a defect report rather than a rate to tune. Do not read a non-zero value as a data loss on its own - it establishes that a barrier faulted and that the barriers after it still ran, which is the distinction that was unavailable before. Chart it against orleans.lattice.leaf.snapshot.capture.declines{reason="no_tree_id"} on the capture-declines panel: the two together partition the silent-teardown population, so an absent snapshot is attributable to a declining capture, to a barrier that threw, or to the hook never running at all. Not zero-primed, matching its capture.declines and driver.declines siblings: a series appears on first fault, so an absent line does NOT establish that the build lacks the instrument. Confirm deployment from the image digest. A healthy estate is flat at absent, so any line here is worth reading
orleans.lattice.leaf.deactivation.barrier.duration histogram (ms) tree, reason, tenant CommitPath (166) Wall-clock duration of each leaf deactivation barrier, by digest_publish/checkpoint_flush/snapshot_capture/frontier_pin (issue #3628). Panel 163 counts barriers that did not complete; this one decomposes a drain by barrier, so a rise in per-activation drain cost is attributable. Recorded whatever the barrier's outcome, so read it beside 163 to tell a slow barrier from a failing one. Charted as an interval mean, delta _sum / delta _count, by tree and reason. Absent means no graceful deactivation ran in the window, not that barriers are free
orleans.lattice.leaf.deactivation.barrier.elided counter ({barrier}) tree, reason, tenant CommitPath (167) Leaf deactivation barriers skipped because the pin store had already acknowledged, in the same deactivation, a pin dominating the whole batch the barrier would publish, by frontier_pin (issue #3643). The teardown tail publishes the frontier pin and the frontier_pin barrier re-resolves the same batch; when every partition is dominated on both frontier and offset by the tail's acknowledged pin, the pin-store round trip is skipped. Read beside panel 166, whose frontier_pin duration is still recorded for an elided barrier: the ratio is the fraction of drains whose pin publish was redundant. Not zero-primed; absent means no barrier has elided yet.
orleans.lattice.leaf.checkpoint.flush.tail.failures counter ({failure}) tree, reason, tenant CommitPath (164) Post-flush notification steps that threw after the checkpoint flush itself succeeded, by cursor_report/inline_digest_publish/snapshot_recheck (issue #3393). Read it beside the barrier counter above, not instead of it: that counter covers the flush BODY and stands down once the durable write lands, while the tail runs three notifications after it, each contained separately and deliberately so, because under issue #2220 a failed notification tearing down the activation produced a replay loop. This instrument does not change that containment; it makes the fault countable. The consequence for reading the barrier counter is the point - a checkpoint_flush count of zero there does not establish that the checkpoint path is healthy. On the drain captured under #3393 the barrier read zero for checkpoint_flush while 1,699 dirty digests died in this tail, so barrier counts are a floor on the damage, not a total. inline_digest_publish is the arm seen in production and means a leaf's parent was never told its digest changed, so the upward digest chain is stale until the next mutation re-drives it - on a quiescing tree, possibly never. snapshot_recheck means the post-flush re-evaluation of snapshot need did not run, so a block pin can retain the tree's shared WAL, making it a candidate contributor to unbounded WAL growth (issue #3094). cursor_report carries a specific warning: the reporter call is also wrapped in its own per-partition try/catch inside the leaf's cursor-report step, so only the option lookup and consumer-id construction around it can reach this counter - a zero on cursor_report is therefore not evidence that cursor reporting is healthy, because it is double-contained and its ordinary failures are invisible here by construction. Not zero-primed, matching its barrier sibling: a series appears on first fault, so an absent line does NOT establish that the build lacks the instrument. Confirm deployment from the image digest. A healthy estate is flat at absent, so any line here is worth reading
orleans.lattice.leaf.snapshot.capture.concurrent_entries counter ({capture}) tree, tenant CommitPath Captures admitted across the attempt boundary while another capture was already in flight (issue #2696). Chart it as a rate beside the peak gauge: the peak says how deep it got, this says how often it happens, and a peak alone cannot separate one transient spike at startup from sustained contention. Do not confuse it with the already_in_flight reason on orleans.lattice.leaf.snapshot.capture.declines - that counts captures the per-leaf guard rejected, this counts captures that were admitted alongside a different leaf, so the two populations are disjoint and only this one is real concurrent load. Zero-primed per tree by adding 0 on uncontended attempts, so a flat zero here means measured absence of contention rather than a tree that never captured
orleans.lattice.leaf.snapshot.load_failures counter ({load}) tree, reason, tenant CommitPath Snapshot load failures by resource_exhausted/unclassified/contiguity_exhausted (issues #2364, #2404, #2844). Observed OOMs decline activation; contiguity_exhausted specifically requires sole-occupant admission and calls for dividing the leaf or lowering MaxLeafBytes. unclassified retains cold-replay fallback and does not rule out an OOM hidden by snapshot-grain activation; correlate logs and memory headroom. It replaces historical faulted: reason-filtered alerts must include both during rollout. The panel groups all reasons, including historical values, without a reason selector. Read with activation failures; this counter does not imply activation succeeded.
orleans.lattice.leaf.snapshot.hydration_admissions counter ({hydration}) tree, outcome, tenant CommitPath Snapshot hydrations passing the per-silo byte-budgeted admission gate, by immediate/queued/sole_occupancy (issues #2765, #2844). One hydration costs several multiples of the stored frame in peak heap - about 2x for a current binary LGB1 blob, roughly 10-13x for a legacy JSON blob written before issue #2516, which still decodes as JSON - and the gate prices every claim at 5x because it cannot tell which before the read (issue #2858); an unbounded cold-start fan-out crossed the .NET heap hard limit the runtime sizes from the cgroup. A rising queued rate is the gate serialising oversized hydrations rather than letting them exhaust the heap together, and should decay as leaf division shrinks the corpus; sole_occupancy is a different predicate - a hydration serialised because its largest contiguous allocation was too big to attempt alongside anything else - and does NOT decay with a larger memory grant, which raises every byte-denominated bound and admits more of exactly these claims; all three outcome arms are pre-minted at zero per tree, so absent means the build did not land rather than nothing happening
orleans.lattice.leaf.snapshot.segment_reads counter ({read}) tree, outcome, tenant CommitPath Individual snapshot segment frames read during a segmented hydration, by loaded/missing/failed (issue #2914). A snapshot above LeafSnapshotSegmentBytes is stored as row-aligned segments in separate grain-state rows and folded one at a time, because the contiguous allocation that failed in issue #2844 happens in the storage provider's column read before any lattice code runs, so the column is what has to be bounded. Chart missing and failed separately and alert on either: both fail the hydration closed and force a full-window WAL replay, but missing is a durability signal (the manifest references a segment that is not there) while failed is an I/O signal. All three outcome arms are pre-minted at zero per tree, but only when the tree's first segmented hydration is observed, so an absent series is ambiguous between a build without the instrument and a tree whose snapshots have never exceeded the segment window - which is why the panel omits or vector(0)
orleans.lattice.leaf.snapshot.segmented_hydrations counter ({hydration}) tree, tenant CommitPath Activation-time hydrations that folded a segmented snapshot to completion (issue #2914). This is the population for which the contiguity bound was actually exercised - leaves whose snapshot fits inline are never counted - so chart it against the sole_occupancy arm of orleans.lattice.leaf.snapshot.hydration_admissions, which should fall as this rises, since a segmented snapshot no longer presents an oversized contiguous claim to the admission gate. Counted only on a complete fold, so it is not an attempt count; partial folds appear on the missing/failed arms of orleans.lattice.leaf.snapshot.segment_reads instead. Pre-minted at zero per tree
orleans.lattice.leaf.snapshot.segment_peak_bytes observable gauge (By) tree, tenant CommitPath High-water mark of the largest single contiguous segment frame this process has materialised while hydrating a segmented snapshot (issue #2914). This is the panel that answers whether the bound held: it must stay at or below the configured LeafSnapshotSegmentBytes window no matter how large the snapshot, so alert on it exceeding that window. Deliberately a monotonic high-water gauge and NOT a histogram, because on the repository-context container's exposition a histogram exports as a summary with only _sum and _count - no buckets, no quantiles - so there the only available reading would be a mean, which dilutes the one large allocation that matters into an average of many small ones, and the peak is what fails. Read it as "the worst this host has seen since it started", not as a current level; it resets on restart by design. Pre-minted at zero per tree
orleans.lattice.leaf.residency.sheds counter ({activation}) tree, kind, tenant CommitPath Leaf activations deactivated by the per-silo resident leaf working set to hold the hydrated population under its derived byte budget, by banked/unbanked (issue #2767). The hydration admission gate bounds concurrency only, while a hydrated leaf's frame is retained for the activation's lifetime, so the steady state grows with the number of live activations and was unbounded. A steady banked rate is the bound working cheaply; a sustained unbanked rate means it has run out of snapshot-backed candidates and is buying headroom with whole-window replays, which points at snapshot coverage rather than at this bound. Both arms are pre-minted at zero per tree, so absent means the build did not land rather than nothing happening
orleans.lattice.leaf.residency.budget_bytes observable gauge (By) tenant CommitPath Resolved byte budget the per-silo resident leaf working set enforces (issue #2788). Read against the container memory grant: a budget of the same order as the whole grant means the process dies before the threshold can be crossed, so the bound never engages and the shed counter's zero says nothing. Originally derived from GCMemoryInfo.TotalAvailableMemoryBytes alone, which reports host physical memory rather than zero when no heap hard limit is set; now the smaller of the heap hard limit and the cgroup memory limit. A flat 1073741824 is the both-unknown fallback and is a detection defect, not a tuned value
orleans.lattice.leaf.residency.resident_bytes observable gauge (By) tenant CommitPath Bytes accounted to live, un-shed leaf registrations (issue #2788). Stack against the budget gauge; the ratio is the headroom. Resident near budget with a steady shed rate is the bound working; resident at a small fraction of budget while memory is exhausting means the budget is too large to bind, not that leaf residency is cheap
orleans.lattice.leaf.residency.registrations observable gauge ({registration}) tenant CommitPath Leaf registrations currently held by the per-silo resident leaf working set (issue #2788). The instrument that separates an empty ledger from a populated one correctly under budget, which the shed counter reads as zero for both because a counter reports events and both are the absence of one. Registered eagerly from AddLattice, so present-and-zero is a measured zero and absent means the build did not land
orleans.lattice.leaf.snapshot.coverage_repairs counter ({evaluation}) tree, outcome, tenant CommitPath Zero-coverage repair evaluations, by repaired/unsatisfied/exhausted/rearmed/backing_off/capture_in_flight/no_checkpointed_uncovered_partition (issues #2692, #2940). A leaf holding a checkpointed partition with no durable snapshot coverage publishes the Zero block pin, and one such leaf disables cursor-based WAL trimming for its whole tree. All seven outcome arms are zero-primed per tree, so an absent series means the path never ran for that tree (or the build did not land) and the CommitPath panel carries no or vector(0) compensation. The two rejection arms are the diagnostic: no_checkpointed_uncovered_partition routes a remedy to the pin/guard seam, unsatisfied routes it to the capture seam, and before #2940 both exits were silent. Expect no_checkpointed_uncovered_partition to dominate the panel by rate on a healthy estate - it is the denominator, not a fault. rearmed is a lifecycle transition, not a terminal outcome: only the other six partition an invocation, and a re-arming invocation records rearmed AND a terminal arm, so do not sum all seven as an invocation tally. backing_off is the suppressed-by-backoff arm, and its rate is the direct cost of the coverage-lag timer spinning against leaves it has already abandoned - the number to read when tuning the recheck cadence against the backoff ceiling. Read it paired with exhausted: an exhausted count with no matching rearmed is a leaf still abandoned, the two moving together is the backoff working. Pair with wal_gc_passes_total{outcome="reclaimed"}: repaired falling to zero while reclaimed passes rise is the drain completing
orleans.lattice.materialiser.drain_lag histogram (ms) tree, tenant CommitPath Leaf-materialiser drain lag p50/p95 (issue #1030 back-pressure)
orleans.lattice.materialiser.lagging_consumers histogram ({consumer}) tree, tenant CommitPath Consumers individually past the drain-lag threshold on a tree already over it (issue #2444). Decomposes the materialiser.drain_lag minimum above, which reads identically for one dormant consumer and for many falling behind. Emitted only for over-threshold trees, so a healthy estate reports nothing and an absent series is the expected reading; the panel therefore omits the or vector(0) fallback its neighbours use, which would draw that healthy absence as a literal zero. Triageable, not diagnosable: it does not name the contributor (issue #2505)
orleans.lattice.materialiser.pin.durable_write_latency histogram (ms) tree, tenant CommitPath Durable pin-write latency: the only materialiser instrument that observes the retention floor rather than in-memory progress (issue #2015)
orleans.lattice.materialiser.pin.reports_shed counter ({report}) tree, pin_shard, tenant CommitPath Steady-state pin reports dropped to protect a pin store that is not keeping up (issue #2014). Safe for durability and NOT for boundedness: this is the only write path carrying an advancing checkpoint offset, so a sustained shed freezes the WAL GC offset floor and retention grows unbounded (issue #3310). Read alongside pin.shed_stall_seconds, which is what tells a burst apart from a latch. pin_shard is the durable-pin routing shard and is NOT the WAL shard tag - different hash, different key, both defaulting to 8, so a join between them is well-formed and meaningless
orleans.lattice.materialiser.pin.shed_forced counter ({report}) tree, pin_shard, tenant CommitPath Pin reports forced through a shed window older than WalMaterialiserPinShedCeiling (issue #3310). Structurally zero while that option is unset, which is the library default, so an absent or flat-zero series means disarmed and not healthy - confirm which before reading it as good news. Every increment is one deliberate override of the issue #2014 back-pressure and is an alarm, not routine. Cannot overstate durability: the offset is clamped to min(checkpoint, covered) in the leaf before it is written
orleans.lattice.materialiser.pin.shed_stall_seconds observable gauge (s) tree, pin_shard, tenant CommitPath Age of each pin shard's current unbroken shed run (issue #3310). Emitted whether or not the ceiling is armed, which is what makes the default-disarmed build observable rather than silent. Series are omitted for shards that are not shedding, so a healthy estate publishes nothing and the panel must not compensate with or vector(0) - that would draw a measured absence as a literal zero. A value that only ever climbs is the latch; reports_shed alone cannot distinguish it from a healthy burst
orleans.lattice.snapshot.replay.entries counter ({entry}) tree, shard, tenant Overview Snapshot replay throughput
orleans.lattice.snapshot.replay.duration histogram (ms) tree, shard, tenant Overview Snapshot replay duration p50/p95/p99
orleans.lattice.snapshot.pins observable gauge ({pin}) tree, tenant Overview Snapshot pins (current). Counts snapshot cursors registered in the WAL cursor registry; those registrations carry a zero cursor and no floor, so they do not hold back WAL trimming - snapshot pages are served from per-shard baselines frozen at open
orleans.lattice.split.retroactive_forward.entries counter ({entry}) tree, shard, tenant Overview Retroactive split-forward throughput
orleans.lattice.split.retroactive_forward.duration histogram (ms) tree, shard, tenant Overview Retroactive split-forward duration p50/p95/p99
orleans.lattice.split.in_flight histogram ({split}) tree, tenant Overview Autonomic split admission (cluster gate)
orleans.lattice.split.candidates_suppressed counter ({shard}) tree, tenant Overview Autonomic split admission (cluster gate)
orleans.lattice.split.admission.deferred counter ({shard}) tree, reason, tenant Overview Autonomic split admission (cluster gate)
orleans.lattice.compaction.shard.dirty_leaves histogram ({leaf}) tree, tenant Overview Compaction dirty leaves per pass
orleans.lattice.compress.dictionary.training_runs counter ({run}) outcome, tenant Overview Auto-trained dictionary - training runs by outcome
orleans.lattice.compress.dictionary.active_version observable gauge ({version}) tenant Overview Auto-trained dictionary - active version
orleans.lattice.compress.dictionary.reservoir_fill observable gauge (1) kind, tenant Overview Auto-trained dictionary - reservoir fill
orleans.lattice.compress.dictionary.trained_bytes_in counter (By) tenant Overview Auto-trained dictionary - trained vs baseline compression ratio
orleans.lattice.compress.dictionary.trained_bytes_out counter (By) tenant Overview Auto-trained dictionary - trained vs baseline compression ratio
orleans.lattice.view.apply_lag histogram ({entry}) view, tenant MaterialisedViews Apply lag (entries) p50/p95/p99, Apply lag p95 by view
orleans.lattice.view.backlog_depth histogram ({entry}) view, tenant MaterialisedViews Drain backlog depth (entries) p50/p95/p99
orleans.lattice.view.applied counter ({write}) view, tenant MaterialisedViews View writes applied (rate)
orleans.lattice.view.aggregation_applied counter ({contribution}) view, tenant MaterialisedViews Aggregation contributions applied (rate)
orleans.lattice.view.aggregation_rejected counter ({contribution}) view, tenant MaterialisedViews Aggregation reserved-key rejections (rate)
orleans.lattice.view.lag_budget_eviction counter ({eviction}) view, tenant MaterialisedViews Lag-budget evictions (rate)
orleans.lattice.view.key_collisions counter ({collision}) view, tenant MaterialisedViews Re-key collisions (rate). Counts colliding view keys, once per drain batch - not the source keys behind them
orleans.lattice.view.atomic_staging_backstop counter ({rebuild}) view, tenant MaterialisedViews Atomic-staging backstop fall-backs (rate)
orleans.lattice.view.cross_tree_joint_violation counter ({degradation}) view, tenant MaterialisedViews Cross-tree joint-atomicity violations (rate)
orleans.lattice.view.source_backpressure counter ({pass}) view, state, tenant MaterialisedViews Source back-pressure self-throttle (rate). Counts throttled drain passes of every trigger, read-your-writes barrier drains included; only background timer ticks are also deferred
orleans.lattice.get.duration histogram (ms) tree, tenant Overview GetAsync / GetManyAsync envelope p50 (ms); GetAsync / GetManyAsync envelope p95 / p99 (ms)
orleans.lattice.get.stage.duration histogram (ms) tree, stage, tenant Overview GetAsync stage breakdown p95 (ms)
orleans.lattice.shard_root.optimistic_read.outcomes counter ({read}) tree, outcome, tenant Overview Shard-root optimistic point reads by outcome (reads/s); includes leaf-generation retries and validated absences
orleans.lattice.get_many.duration histogram (ms) tree, tenant Overview GetAsync / GetManyAsync envelope p50 (ms); GetAsync / GetManyAsync envelope p95 / p99 (ms)
orleans.lattice.get_many.stage.duration histogram (ms) tree, stage, tenant Overview GetManyAsync stage breakdown p95 (ms)
orleans.lattice.exists.duration histogram (ms) tree, tenant Overview ExistsAsync / GetWithVersionAsync envelope p95 (ms)
orleans.lattice.get_with_version.duration histogram (ms) tree, tenant Overview ExistsAsync / GetWithVersionAsync envelope p95 (ms)
orleans.lattice.set.duration histogram (ms) tree, tenant CommitPath SetAsync / SetManyAsync envelope p50 (ms); SetAsync / SetManyAsync envelope p95 (ms)
orleans.lattice.set.stage.duration histogram (ms) tree, stage, tenant CommitPath SetAsync stage breakdown p95 (ms)
orleans.lattice.set_many.duration histogram (ms) tree, tenant CommitPath SetAsync / SetManyAsync envelope p50 (ms); SetAsync / SetManyAsync envelope p95 (ms)
orleans.lattice.set_many.stage.duration histogram (ms) tree, stage, tenant CommitPath SetManyAsync stage breakdown p95 (ms)
orleans.lattice.set_many_where_predicate.duration histogram (ms) tree, tenant CommitPath SetAsync / SetManyAsync envelope p50 (ms); SetAsync / SetManyAsync envelope p95 (ms)
orleans.lattice.shard_root.set_many.leaf_rpc.duration histogram (ms) tree, operation, tenant CommitPath ShardRoot.SetMany sub-attribution p95 (ms)
orleans.lattice.shard_root.set_many.local_apply.duration histogram (ms) tree, operation, tenant CommitPath ShardRoot.SetMany sub-attribution p95 (ms)
orleans.lattice.shard_root.set_many.shadow_forward.duration histogram (ms) tree, tenant CommitPath ShardRoot.SetMany sub-attribution p95 (ms)
orleans.lattice.warmup.invocations counter ({call}) tree, tenant CommitPath WarmUpAsync - invocations and duration
orleans.lattice.warmup.duration histogram (ms) tree, shard_count, tenant CommitPath WarmUpAsync - invocations and duration
orleans.lattice.warmup.leaf_cache.prewarmed counter ({leaf}) tree, shard, tenant CommitPath Leaf-cache pre-warm (on by default) - leaves primed, fan-out cost, model size
orleans.lattice.warmup.leaf_cache.duration histogram (ms) tree, shard, tenant CommitPath Leaf-cache pre-warm (on by default) - leaves primed, fan-out cost, model size
orleans.lattice.leaf_access.model.leaves histogram ({leaf}) tree, shard, tenant CommitPath Leaf-cache pre-warm (on by default) - leaves primed, fan-out cost, model size
orleans.lattice.registry.call.duration histogram (ms) operation, tenant Overview Tree-registry singleton - service time and fan-in width
orleans.lattice.registry.call.in_flight histogram ({call}) operation, tenant Overview Tree-registry singleton - service time and fan-in width
orleans.lattice.registry.caller.duration histogram (ms) method, outcome, tenant Overview Tree-registry calls - caller-observed duration and outcome
orleans.lattice.registry.admission.wait histogram (ms) tenant Overview Tree-registry fan-in gate - caller-side admission wait (the only signal that sees a stall relocated out of the registry)
orleans.lattice.registry.admission.in_flight histogram ({dispatch}) tenant (not charted) Fan-in permits this silo's gate held at dispatch, counting the dispatch, so a gate at its ceiling reads as its full share of the cluster-wide bound of 16 (divided by the live silo count and floored at one, so 16 on a single silo) rather than one below it
orleans.lattice.registry.admission.batch.size histogram ({tree}) tenant (not charted) Distinct tree ids per gated round trip - the share above 1 is the share the bound actually coalesced
orleans.lattice.registry.admission.queue.depth histogram ({tree}) tenant (not charted) Offered fan-in at enqueue - distinguishes "the bound had room" from "nothing asked for it"
orleans.lattice.leaf.commit.in_flight histogram ({commit}) tree, tenant CommitPath Leaf commit concurrency (in-flight) p95
orleans.lattice.leaf.digest.publishes counter ({publish}) tree, path, tenant CommitPath Digest publish path attribution (ops/s) - coalescing efficacy
orleans.lattice.provider.commit.duration histogram (ms) tree, shard, phase, pipeline_phase2, tenant CommitPath Storage-provider phase-2 commit p95 (ms) + batch size
orleans.lattice.provider.phase2.batch_size histogram ({commit}) tree, shard, tenant CommitPath Storage-provider phase-2 commit p95 (ms) + batch size
orleans.lattice.provider.retry.attempts counter ({attempt}) status, tenant CommitPath Storage-provider retries (ops/s) - attempts vs exhausted vs idempotent-replays vs phase1-transient. The instrument is tagged by status only, to bound cardinality, so the attempts target applies no tree matcher: its line is silo-wide and the dashboard's tree selector does not narrow it
orleans.lattice.provider.retry.exhausted counter ({call}) tree, shard, phase, status, tenant CommitPath Storage-provider retries (ops/s) - attempts vs exhausted vs idempotent-replays vs phase1-transient
orleans.lattice.provider.idempotent_replays counter ({call}) tree, shard, phase, tenant CommitPath Storage-provider retries (ops/s) - attempts vs exhausted vs idempotent-replays vs phase1-transient
orleans.lattice.wal.append.turn_wait histogram (ms) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL append latency p95 (ms) - turn-wait / provider / dispatch
orleans.lattice.wal.append.provider.duration histogram (ms) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL append latency p95 (ms) - turn-wait / provider / dispatch
orleans.lattice.wal.append.in_flight histogram ({flush}) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL pipeline depth p95 - in-flight flushes / queue depth
orleans.lattice.wal.append.queue_depth histogram ({entry}) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL pipeline depth p95 - in-flight flushes / queue depth
orleans.lattice.wal.append.batch_entries histogram ({entry}) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL batch shape p95 - entries / bytes / dispatch-entries
orleans.lattice.wal.append.batch_bytes histogram (By) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL batch shape p95 - entries / bytes / dispatch-entries
orleans.lattice.wal.shard.dispatch.duration histogram (ms) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath WAL append latency p95 (ms) - turn-wait / provider / dispatch
orleans.lattice.wal.shard.dispatch.entries histogram ({entry}) tree, shard, wal_partitions, wal_max_pending_batches, tenant CommitPath, Replication WAL batch shape p95 - entries / bytes / dispatch-entries; Log-tailing producer: leaf WAL append vs ship rate (ops/s)
orleans.lattice.wal.saturation.state observable gauge ({state}, 0/1/2) tree, tenant Overview WAL saturation regime - % time non-Healthy (1h); WAL saturation regime - current state per tree
orleans.lattice.wal.saturation.transitions counter ({transition}) tree, shard, partition, state, previous_state, cause, tenant Overview WAL saturation regime - per-partition attribution (heat-map); WAL saturation regime - transition rate by direction (ops/s). cause (dispatch_timeouts, provider_failures, flush_latency, admission_depth, materialiser_drain_lag, materialiser_pin_latency, or none on a return to healthy) names the input the transition was attributed to and is not charted. partition and shard are recorded whatever the cause - the partition with the highest admission-depth ratio, and the first shard whose dispatch-timeout, provider-failure, flush-latency or pin-write-latency input advanced that window - so the heat-map shows where admission depth was highest when a tree saturated, not what saturated it
orleans.lattice.storage.wal.uncompressed_bytes counter (By) tree, tenant Overview WAL compression savings ratio by tree
orleans.lattice.storage.wal.stored_bytes counter (By) tree, tenant Overview WAL compression savings ratio by tree
orleans.lattice.storage.wal.compression_skipped counter ({row}) tree, reason, tenant Overview WAL compression skips by reason
orleans.lattice.saga.prepare.duration histogram (ms) tree, wal_partitions, tenant AtomicWrites Saga phase durations p95 (ms) - prepare / decision / broadcast / checkpoint / reminder
orleans.lattice.saga.terminal_decision.duration histogram (ms) tree, wal_partitions, tenant AtomicWrites Saga phase durations p95 (ms) - prepare / decision / broadcast / checkpoint / reminder
orleans.lattice.saga.broadcast.duration histogram (ms) tree, wal_partitions, tenant AtomicWrites Saga phase durations p95 (ms) - prepare / decision / broadcast / checkpoint / reminder
orleans.lattice.saga.broadcast.shard.duration histogram (ms) tree, shard, tenant AtomicWrites Saga broadcast sub-attribution p95 (ms) - per-shard / per-leaf / per-shard-stage
orleans.lattice.saga.broadcast.leaf.duration histogram (ms) tree, shard, tenant AtomicWrites Saga broadcast sub-attribution p95 (ms) - per-shard / per-leaf / per-shard-stage
orleans.lattice.saga.broadcast.shard.stage.duration histogram (ms) tree, shard, stage, tenant AtomicWrites Saga broadcast sub-attribution p95 (ms) - per-shard / per-leaf / per-shard-stage
orleans.lattice.saga.checkpoint.duration histogram (ms) tree, wal_partitions, phase, tenant AtomicWrites Saga phase durations p95 (ms) - prepare / decision / broadcast / checkpoint / reminder
orleans.lattice.saga.reminder.duration histogram (ms) tree, wal_partitions, phase, tenant AtomicWrites Saga phase durations p95 (ms) - prepare / decision / broadcast / checkpoint / reminder
orleans.lattice.saga.perkey.duration histogram (ms) tree, wal_partitions, tenant AtomicWrites Per-key saga work - p95 per-key duration (ms)
orleans.lattice.saga.fanout.size histogram ({entry}) tree, wal_partitions, tenant AtomicWrites Saga fan-out size (entries per saga)
orleans.lattice.atomic_write.cross_tree.completed counter ({saga}) outcome, tree_count, tenant AtomicWrites Cross-tree atomic write outcomes (rate); Cross-tree failure rate (%)
orleans.lattice.atomic_write.cross_tree.duration histogram (ms) outcome, tenant AtomicWrites Cross-tree coordinator duration (p50/p95/p99 ms)
orleans.lattice.atomic_write.cross_tree.participants histogram ({tree}) outcome, tenant AtomicWrites Cross-tree participant fan-out (trees per saga)
orleans.lattice.tx_registry.writes counter ({write}) tree, outcome, tenant AtomicWrites Tx registry writes (rate) by outcome
orleans.lattice.tx_registry.write.mutations histogram ({mutation}) tree, outcome, tenant AtomicWrites Tx registry group-commit coalescing factor (mutations per write)
orleans.lattice.tx_registry.write.duration histogram (ms) tree, outcome, tenant AtomicWrites Tx registry write duration (mean ms) by outcome
orleans.lattice.grainindex.grains_enrolled counter ({grain}) index, path, tenant GrainIndex Grains enrolled per second, by route
orleans.lattice.grainindex.entries up-down counter ({entry}) index, tenant GrainIndex Index entries held
orleans.lattice.grainindex.write_failures counter ({failure}) index, path, tenant GrainIndex Index write failures per second, by route
orleans.lattice.grainindex.projection.duration histogram (ms) index, tenant GrainIndex Projection latency percentiles
orleans.lattice.grainindex.backfill.processed observable gauge ({grain}) index, tenant GrainIndex Backfill progress (processed vs total)
orleans.lattice.grainindex.backfill.total observable gauge ({grain}) index, tenant GrainIndex Backfill progress (processed vs total)
orleans.lattice.grainindex.backfill.percent_complete observable gauge (%) index, tenant GrainIndex Backfill percent complete
orleans.lattice.grainindex.backfill.state observable gauge ({state}) index, tenant GrainIndex Backfill state
orleans.lattice.tag_index.reconcile.sweeps counter ({sweep}) index, outcome, tenant (not charted) Background tag-index reconciliation sweeps, by outcome (clean, repaired, probe_only)
orleans.lattice.tag_index.reconcile.trees.probed counter ({tree}) index, tenant (not charted) Covered trees whose digest fingerprint a sweep probed
orleans.lattice.tag_index.reconcile.trees.mismatched counter ({tree}) index, tenant (not charted) Covered trees a sweep found divergent from their digest baseline
orleans.lattice.tag_index.reconcile.orphan_rows.removed counter ({row}) index, tenant (not charted) Orphan membership rows removed by background reconciliation
orleans.lattice.tag_index.reconcile.duration histogram (ms) index, tenant (not charted) Wall-clock duration of a background reconciliation sweep

The orleans.lattice.tag_index.reconcile.* family is emitted by the background tag-index reconciliation sweep and is documented in Metrics. No bundled dashboard charts it yet; scrape it directly, or add panels and update the rows above.