---
title: "Instrument catalog: Storage-provider commit pipeline to Grain-call observation (opt-in) - Metrics"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/metrics/instrument-catalog-7.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/metrics.md?plain=1#L608-L815"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Instrument catalog: Storage-provider commit pipeline to Grain-call observation (opt-in)

Part of [Instrument catalog](instrument-catalog.md), in [Metrics](../metrics.md).

## Storage-provider commit pipeline

The Azure Table WAL provider (and any other WAL provider that opts in)
emits per-commit-phase histograms so the per-batch partition transaction
and the manifest partition transaction can be observed independently.
See [Write-Ahead Log](../wal.md) and
[WAL Storage Providers](../wal-storage-providers.md) for the underlying
two-phase commit shape.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.provider.commit.duration` | `Histogram<double>` | `ms` | Wall-clock duration of one storage-provider commit transaction. Tagged `tree`, `shard`, `phase=phase1` (per-batch partition transaction) or `phase=phase2` (manifest partition transaction), and `pipeline_phase2` reflecting `AzureTableWalStorageOptions.PipelinePhaseTwoCommits`. The phase-2 measurement covers a single coalesced commit transaction, not the per-shard worker's whole drain loop. |
| `orleans.lattice.provider.phase2.batch_size` | `Histogram<int>` | `{commit}` | Number of coalesced phase-2 commits the per-shard provider worker bundled into a single transaction. Tagged `tree`, `shard`. A distribution concentrated near 1 means the worker is never catching up against backed-up arrivals; values closer to the per-transaction cap (~49 commits) indicate the worker is the shard's effective rate limiter. |
| `orleans.lattice.provider.retry.attempts` | `Counter<long>` | `{attempt}` | Individual retry attempts the storage-provider SDK performs on a WAL append or commit, incremented once per attempt regardless of whether the retry ultimately succeeds. Tagged `status` (the HTTP status of the response that triggered the retry, e.g. `503` / `429`, or `0` for a transport-level failure with no HTTP exchange); per-tree / per-shard correlation is intentionally omitted to keep cardinality low - use `provider.retry.exhausted` for shard-attributed failures. A non-zero rate is normal; sustained climbing rates suggest provider-side throttling that the local retry policy is masking. |
| `orleans.lattice.provider.retry.exhausted` | `Counter<long>` | `{call}` | Calls that exhausted the storage-provider retry budget and surfaced the underlying fault. Tagged `tree`, `shard`, `phase` (`phase1` per-batch partition transaction or `phase2` manifest partition transaction), and `status` (the HTTP status string the SDK observed, or `unknown`). **Any non-zero rate is alert-worthy** - the WAL append or commit failed past the SDK's retry envelope. |
| `orleans.lattice.provider.retry.short_circuited` | `Counter<long>` | `{attempt}` | Storage-provider SDK retry attempts abandoned early by the Azure Tables provider's saturation-aware retry policy (`AzureTableWalStorageOptions.HonorSaturationSignal = true`, the default). Tagged `status` (the synthetic-response HTTP status that signalled the SDK to exit the retry chain; `503` matches the policy's default short-circuit shape). Increments once per abandoned retry, including retries the policy short-circuits during the `AzureTableWalStorageOptions.SaturationShortCircuitCooldown` sticky window after the last observed `Saturated` tick. A non-zero rate confirms the policy is firing under storage-side back-pressure; a sustained climbing rate alongside `provider.retry.exhausted` indicates the storage account is shedding load faster than the SDK retry budget would otherwise surface to the writer. Always zero when `HonorSaturationSignal = false` (the historical unguarded SDK retry behaviour) or when the host supplies a pre-built `TableServiceClient` via `ServiceClient` (the host owns the pipeline in that mode). |
| `orleans.lattice.provider.idempotent_replays` | `Counter<long>` | `{call}` | Phase-one WAL batch commits whose `409 EntityAlreadyExists` conflict was proven to be an idempotent replay of an already-durable write and resolved as a success. Tagged `tree`, `shard`, and `phase` (always `phase1` - idempotent replays are only detected on the phase-one batch commit). Increments when the Azure Tables SDK retry pipeline resends a phase-one batch whose first attempt committed server-side but whose response was lost (a socket / read timeout under CPU or network pressure); the provider reads the resident rows back, confirms they are byte-identical to the batch it tried to write, and treats the conflict as success instead of failing the whole batch. A non-zero rate is **benign** (no data was lost) but is a leading indicator that the cluster is dropping phase-one responses under load - watch it alongside `provider.retry.exhausted`, not instead of it. A genuine offset collision (resident rows present but differing) is never counted here; it still surfaces as a hard failure on `provider.retry.exhausted`. |
| `orleans.lattice.provider.phase1.transient_retries` | `Counter<long>` | `{attempt}` | Phase-1 commit attempts the provider re-issued in place after a transient fault, resubmitting the byte-identical batch at the same offsets rather than faulting the calling WAL shard. Tagged `tree`, `shard`, and `phase` (always `phase1`). |
| `orleans.lattice.provider.phase2.commit.timeouts` | `Counter<long>` | `{commit}` | Phase-2 manifest commits abandoned by the per-shard worker after exceeding `AzureTableWalStorageOptions.PhaseTwoCommitTimeout`. Tagged `tree`, `shard`. Zero unless a commit's Azure Tables transaction stopped making progress (hung socket, server-side partition stall, or an SDK retry loop running past the deadline); **any non-zero rate is alert-worthy** - it is the direct signal that the per-shard phase-2 drain loop would otherwise have wedged. `PhaseTwoCommitTimeout` defaults to 12 seconds; set it to `null` and the commit is left unbounded, so this counter never increments. |

## Saga / coordinator / lifecycle

Long-running maintenance and atomicity primitives each emit a completion
signal so operators can alert on stalls, compensation spikes, or missing
coordinator progress independently of whether the event stream is enabled.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.atomic_write.completed` | `Counter<long>` | `{saga}` | Terminal transition of a `SetManyAtomicAsync` saga - one increment per single-tree saga, and one per participating tree of a cross-tree write. Tagged `outcome=committed` (the commit decision was recorded and the commit terminals broadcast, so every write became visible), `failed` (a batch still failed after its retries, so the saga captured the failure message, recorded an abort decision in the transaction registry and broadcast abort terminals), or `compensated` (an abort with no captured failure message - in the current code, a participating tree of a cross-tree write that its coordinator finalized with abort because another participant did not prepare). Neither abort arm issues per-key rollback writes: prepared writes stay invisible in each leaf's pending-transaction buckets until the abort terminal drops them, so there is nothing to undo. The outcome mapping also names `shutdown_refused`, but no code path produces it, because the failure-message prefix it tests for is never written: a shutdown refusal throws `LatticeShuttingDownException` to the caller and leaves the saga at its persisted position to resume on the next activation, recording no terminal transition, so count shutdown refusals at the caller instead (see [Atomic Writes](../atomic-writes.md)). No sample is recorded either for a saturation refusal, which throws `LatticeSaturatedException` and leaves the saga to resume later from its persisted position (a transaction-registry capacity refusal, `LatticeSaturationSource.TxRegistryCapacity`, is raised before the saga persists anything, so a retry under the same operation id starts cleanly), or for a batch whose guard predicate fails (a `SetManyAtomicWhereAsync` call returning `AtomicWriteOutcome.PreconditionFailed`, or a cross-tree participant voting precondition-failed), which commits nothing and records no terminal transition. |
| `orleans.lattice.atomic_write.duration` | `Histogram<double>` | `ms` | End-to-end `SetManyAtomicAsync` saga duration, captured from the first `Prepare` to `Completed` and persisted across reminder-driven recovery so the recorded ms reflects true wall-clock cost (including any time the saga was suspended across silo restarts). Tagged with the same `outcome` values as `atomic_write.completed`; emitted alongside it on every terminal transition, except for a saga whose persisted state predates the start timestamp (it decodes as zero), for which the duration is suppressed rather than recorded as time since year one. Pair with `atomic_write.completed` to derive sustained atomic-write throughput and SLO percentiles. |
| `orleans.lattice.atomic_write.batch_size` | `Histogram<int>` | `{entry}` | Entry count of each `SetManyAtomicAsync` saga at terminal transition. Tagged with the same `outcome` values as `atomic_write.completed`; emitted alongside it. Lets operators correlate p99 saga duration with batch size and detect distribution shifts in caller batch sizing. |
| `orleans.lattice.atomic_write.cross_tree.completed` | `Counter<long>` | `{saga}` | Terminal transition of a cross-tree atomic-write saga (a `SetManyAtomicAsync` spanning more than one tree). Tagged `outcome=committed` or `precondition_failed` - those two are the whole domain, because the coordinator has no compensation arm of its own, and `precondition_failed` covers every abort, including one caused by a participant that failed to prepare rather than by a precondition miss (each participant it finalizes with abort reports `compensated` on `orleans.lattice.atomic_write.completed`) - and `tree_count` (number of participating trees). It carries no `tree` tag and records under the `_platform_` tenant, and an empty cross-tree batch commits vacuously without recording anything. A panel filtering this counter for `failed` or `compensated` selects zero series forever; `DashboardPanelTagDomainTests` now fails the build on exactly that mistake. |
| `orleans.lattice.atomic_write.cross_tree.duration` | `Histogram<double>` | `ms` | End-to-end cross-tree atomic-write coordinator duration. Tagged `outcome`. |
| `orleans.lattice.atomic_write.cross_tree.participants` | `Histogram<int>` | `{tree}` | Participating-tree count of each cross-tree atomic-write saga. Tagged `outcome`. |
| `orleans.lattice.saga.prepare.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside one run of the saga's staging (execute) phase: the batched `ILattice.SetManyAsync` dispatch that stages every per-key prepared mutation across all touched shards, together with the saga checkpoint persists and batch retries that phase issues, recorded whether the phase completes or throws. Tagged `tree` and `wal_partitions`. Pair with `set_many.duration` and `saga.checkpoint.duration` to confirm the saga prepare is dominated by the foreground multi-key write rather than saga-internal bookkeeping. |
| `orleans.lattice.saga.terminal_decision.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside one saga `TerminalDecision` phase (decide whether to commit or compensate after every prepared mutation has acknowledged). Tagged `tree` and `wal_partitions`. Sub-millisecond except under failure paths. |
| `orleans.lattice.saga.broadcast.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside one saga `Broadcast` phase (the `Task.WhenAll` across affected shards that flips every prepared mutation to its terminal state). Tagged `tree` and `wal_partitions`. The c2-xxiii batched-WAL terminal lift collapsed this phase from ~880ms p50 to ~96ms p50 at the `500:10` rung. |
| `orleans.lattice.saga.broadcast.shard.duration` | `Histogram<double>` | `ms` | Per-shard contribution inside the saga broadcast: wall-clock ms inside one `ShardRootGrain.AppendTxTerminalAsync` call. Tagged `tree` and `shard`. The gap between `saga.broadcast.duration` and this is the max-of-N parallel tail (Orleans scheduling + the slowest shard). |
| `orleans.lattice.saga.broadcast.shard.stage.duration` | `Histogram<double>` | `ms` | Per-sub-stage wall-clock ms inside one `ShardRootGrain.AppendTxTerminalAsync` call. Tagged `tree`, `shard`, and `stage=resolve` (affected-leaves resolution), `hlc` (terminal HLC compute via `GetClockAsync` fan-out), `wal` (the shard root's WAL terminal append; collapsed to ~0ms after c2-xxiii batched the WAL terminal at the saga layer), or `fanout` (the per-leaf `ApplyTxTerminalAsync` fan-out). |
| `orleans.lattice.saga.broadcast.leaf.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside a single per-leaf `IBPlusLeafGrain.ApplyTxTerminalAsync` RPC dispatched from the shard's terminal broadcast (step 4). Tagged `tree` and `shard`. |
| `orleans.lattice.saga.checkpoint.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside one persist of the saga's state (a single `WriteStateAsync`), recorded at every site that persists it rather than only at the terminal decision, and on the failure path too. Tagged `tree`, `wal_partitions` and `phase`, the call site: `prepare`, `prepared`, `execute-batch-commit`, `execute-batch-retry`, `execute-to-compensate`, `broadcast-touched-init`, `broadcast-touched-expand`, `broadcast-touched-drift`, `broadcast-touched-late` or `complete`. |
| `orleans.lattice.saga.reminder.duration` | `Histogram<double>` | `ms` | Wall-clock ms inside the Orleans reminder-registry calls that manage the saga's keepalive reminder, its crash-recovery anchor. Tagged `tree`, `wal_partitions` and `phase`: `register` (registering or refreshing the reminder, including the bounded wait while the silo's reminder service is still initialising), `unregister-get` (looking the reminder up to tear it down once the saga no longer needs it) or `unregister-drop` (removing it). |
| `orleans.lattice.saga.fanout.size` | `Histogram<int>` | `{entry}` | Entry count per atomic-write saga, observed at execute-phase entry. Tagged `tree` and `wal_partitions`. The distribution tells operators whether the workload is dominated by small single-entry sagas or large multi-entry batches. |
| `orleans.lattice.saga.perkey.duration` | `Histogram<double>` | `ms` | Amortised per-key cost of a saga's prepare dispatch. Each staging batch is sent as one `ILattice.SetManyAsync`, so per-key timing is not observable: the batch's wall-clock time divided by the number of keys it carried is recorded once per key, so every sample from one dispatch is identical, and a batch that faulted is recorded as well. Tagged `tree` and `wal_partitions`. Pair with `atomic_write.batch_size` to derive amortised per-key atomic-write cost across the saga envelope. |
| `orleans.lattice.tx_registry.writes` | `Counter<long>` | `{write}` | Whole-state writes issued by the saga decision registry. Tagged `tree` and `outcome=ok` (the write completed durably) or `fault` (the write threw; every mutation it carried, and any queued behind it, was rolled back and failed to its caller, which retries). A registry shard raises the tree's shard high-water mark before its first write; if that raise fails, the pending mutations are rolled back and failed the same way and one `fault` is recorded although no state write was issued. Every atomic saga records its participants, decision and cleanup through its registry, so the registry's write rate bounds saga throughput. At the default [`TxRegistryShardCount`](../configuration/options-reference-3.md#txregistryshardcount) of `1` that is one registry activation, and one persisted row, per tree; a higher count mints new sagas across that many registry shards, each its own activation with its own row, beside the legacy per-tree registry that sagas minted earlier keep routing to. Every shard records under the logical tree id, so a tree's series sums all of its registries. Each registry group-commits: at most one write is in flight, and mutations that arrive meanwhile join the next write ([#3475](https://github.com/NSTA1/Orleans.Lattice/issues/3475)). A write rate that stays flat while `tx_registry.write.mutations` climbs is coalescing at work. **Any sustained `fault` rate is alert-worthy.** |
| `orleans.lattice.tx_registry.write.mutations` | `Histogram<int>` | `{mutation}` | Registry mutations carried by one saga decision registry state write: the group-commit coalescing factor. Tagged `tree` and `outcome` (as `tx_registry.writes`). Each mutation is one caller's call, so on the `ok` arm `1` means the write acknowledged a single caller and `N` means one durable write acknowledged `N` callers; on the `fault` arm it counts the mutations the failed write carried, not the queued ones that fail with it. A p50 stuck at `1` under heavy atomic load means there was no concurrency to coalesce. |
| `orleans.lattice.tx_registry.write.duration` | `Histogram<double>` | `ms` | Duration of one saga decision registry whole-state write, measured around the storage call (for a `fault` caused by a failed shard high-water raise, around that raise instead). Tagged `tree` and `outcome` (as `tx_registry.writes`). Each registry serialises its writes, so the reciprocal of this duration is the most writes per second one registry can issue - a whole tree's ceiling at the default `TxRegistryShardCount` of `1`, and one shard's share of it above that. It grows with the registry row, which carries a tombstone per completed saga for `TxDecisionRetention`; see `TxRegistryAdmissionBudgetBytes` in [Configuration](../configuration/options-reference-3.md#txregistryadmissionbudgetbytes). |
| `orleans.lattice.coordinator.completed` | `Counter<long>` | `{operation}` | Successful completion of a long-running coordinator. Tagged `tree` and `kind=snapshot`, `resize`, `reshard`, `merge`, or `compaction`. |
| `orleans.lattice.coordinator.phase_tick.failures` | `Counter<long>` | `{failure}` | A coordinator phase-timer tick whose phase step threw and was swallowed. That tick advanced the phase machine by nothing, so the step it would have taken is retried from the start and any work it had accumulated is discarded. Tagged `kind` (the coordinator's keepalive reminder name, for example `resize-keepalive` or `reshard-keepalive` - not the bare kind `coordinator.completed` carries), `tree`, and the tenant label. **Zero-primed when a coordinator arms its phase timer**, so a flat zero on a live series is a reading - this coordinator has swallowed nothing - rather than the absence a counter reports before its first `Add`. A zero does **not** mean the coordinator is *progressing*: a tick that returns normally without advancing is a success here, so denominate against `coordinator.completed` for that question. The first two consecutive failures on one activation log at warning and the third onwards at error, because a run of them is a phase loop that has stopped advancing rather than a transient the pump absorbs. |
| `orleans.lattice.coordinator.phase_tick.consecutive_failures` | `ObservableGauge<long>` | `{failure}` | Length of the *current run* of consecutive failed coordinator phase ticks, tagged identically to `coordinator.phase_tick.failures` so the two series join. The counter beside it cannot express consecutiveness: a coordinator that fails one tick in a thousand and one that has failed every tick since process start both present as a rising total, yet the first is a transient the pump absorbs by design and the second is a phase machine that has stopped advancing. **Every live coordinator reports, enrolling at `0` when it arms its phase timer**, so `0` means the last tick succeeded rather than no data - and, exactly as for the counter, `0` does **not** mean the coordinator is *advancing*, because a tick that returns without moving the phase machine forward is a success by this measure. Reported as the **maximum** over the activations sharing a tag set, which is coarser than the activation because a coordinator with a composite key deliberately reports under the subject alone; `max` is the correct reduction because a sum would invent a run no activation experienced and a last-writer-wins would let a healthy sibling hide a wedged one. A value that keeps returning to zero is a coordinator absorbing transients; a value that only climbs is a wedge, and its magnitude is how many ticks of work have been discarded back to back. |
| `orleans.lattice.tree.lifecycle` | `Counter<long>` | `{event}` | Tree-lifecycle transition, recorded by the tree's soft-delete and purge manager. Tagged `tree` and `kind=deleted`, `recovered`, or `purged`. Emitted **unconditionally** - regardless of the tree's `PublishEvents` setting. Exactly one increment per logical delete, recover or purge, tagged with the logical tree id, including on an aliased tree; the work a logical operation delegates to the backing physical tree records nothing of its own. A resize's retirement of its old physical copy, and a resize undo's recovery or discard of a copy, are physical maintenance and record nothing. |
| `orleans.lattice.warmup.invocations` | `Counter<long>` | `{call}` | One increment per successful `ILattice.WarmUpAsync` call. Tagged `tree`. Operators alerting on cold-start health expect to see exactly one increment per silo startup per warmed tree. |
| `orleans.lattice.warmup.duration` | `Histogram<double>` | `ms` | End-to-end duration of `ILattice.WarmUpAsync` - the wall-clock cost of pre-activating every physical shard root via a bounded-concurrency read-only probe. Tagged `tree` and `shard_count` (the per-tree physical-shard-root probe fan-out). The p99 is the primary warm-start latency signal; sustained increases are a leading indicator of placement-directory or grain-storage cold-touch cost growth. |
| `orleans.lattice.warmup.leaf_cache.prewarmed` | `Counter<long>` | `{leaf}` | Leaf caches successfully primed by a shard root's post-restart pre-warm. Tagged `tree`, `shard`, and the tenant label. Flat at zero only where `LatticeOptions.LeafCachePreWarmCount` has been set to `0`; the default is `8`, so the feature is on unless it is explicitly turned off. Individual priming failures are swallowed by design, so a value materially below the configured count is the only signal that they occurred. |
| `orleans.lattice.warmup.leaf_cache.duration` | `Histogram<double>` | `ms` | Wall-clock duration of one shard root's leaf-cache pre-warm fan-out. One observation per shard per warm-up when the feature is enabled and the access model ranked at least one leaf. Tagged `tree`, `shard`, and the tenant label. Read against `warmup.duration` to attribute how much of a tree's warm-start cost is leaf priming. |
| `orleans.lattice.leaf_access.model.leaves` | `Histogram<int>` | `{leaf}` | Leaves resident in a shard root's leaf-access frequency histogram, observed each time the model is persisted. Tagged `tree`, `shard`, and the tenant label. Bounded above by the model's tracked-leaf cap, so a distribution pinned at that cap means the shard's read set is wider than the model can represent and the pre-warm ranking is drawn from a pruned view. |

## Distributed lock (sourced from `LatticeLockGrain`)

The FIFO-fair distributed lock / lease grain behind `ILatticeLockGrain` (see
[Distributed lock](../distributed-lock.md)). One activation per lock name serialises
all contending callers, so these instruments carry no per-key tag; scope them by
`cluster` (and `silo`) in the dashboards. Charted on the Overview dashboard's
"Distributed lock" row.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.lock.acquired` | `Counter<long>` | `{acquire}` | One increment per acquire attempt that reached a terminal outcome. Tagged `outcome=granted` (the caller was handed the lease), `timeout` (a blocking `AcquireAsync` whose wait-timeout elapsed, or a non-blocking `MaxWait = 0` acquire on a held lock), or `unavailable` (a `TryAcquireAsync` that returned `null` because the lock was held). Pair `granted` with `released` to see live lock churn. |
| `orleans.lattice.lock.released` | `Counter<long>` | `{release}` | One increment per explicit `ReleaseAsync` by the current holder (a stale-token release is a silent no-op and is not counted). |
| `orleans.lattice.lock.lease_reclaimed` | `Counter<long>` | `{lease}` | One increment each time an expired lease was reclaimed because the holder neither renewed nor released before expiry. A sustained non-zero rate means holders are crashing or failing to renew within their lease duration; downstream resources should be relying on the fencing token to reject the reclaimed holder. |
| `orleans.lattice.lock.acquire.wait` | `Histogram<double>` | `ms` | Time a granted acquire spent parked in the FIFO queue before it was granted (recorded once per grant, so an uncontended acquire records ~0). The p95/p99 tail is the primary lock-contention signal. |

## Atomic action / saga (sourced from `AtomicActionGrain`)

The generic atomic-action (saga / TCC) coordinator behind `IAtomicActionGrain` (see
[Atomic action](../atomic-action.md)). One activation per operation id, so these
instruments carry no per-key tag; scope them by `cluster` (and `silo`) in the
dashboards. Charted on the Overview dashboard's "Atomic action (saga / TCC)" row.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.atomic_action.completed` | `Counter<long>` | `{saga}` | One increment per terminal transition of a saga. Tagged `outcome=committed` (every forward step committed), `compensated` (a forward step faulted and every committed step was rolled back in reverse order), or `compensation_failed` (a compensating effect itself faulted, so the saga parked for operator intervention). A rising `compensation_failed` rate means a caller's compensation contract is being violated. |
| `orleans.lattice.atomic_action.step` | `Counter<long>` | `{step}` | One increment per step effect the saga runs. Tagged `phase=forward` \| `compensate` and `outcome=ok` \| `fault`. The `compensate` series is non-zero only when a saga rolls back; a `compensate,fault` point is the leading indicator of a park. |
| `orleans.lattice.atomic_action.duration` | `Histogram<double>` | `ms` | End-to-end saga duration measured from saga start (persisted, so it includes time suspended across silo restarts) to the terminal transition. Tagged by `outcome` so rollback-path latency is separable from happy-path latency. |

## Events

These counters are emitted only when event publication is enabled on at least
one tree. They let operators detect a misconfigured stream provider or a
failing downstream queue before it starts consuming silo resources.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.events.published` | `Counter<long>` | `{event}` | `LatticeTreeEvent` instances successfully dispatched to the configured stream provider. Tagged `tree` and `kind` = the `LatticeTreeEventKind` name (e.g. `Set`, `SnapshotCompleted`). |
| `orleans.lattice.events.dropped` | `Counter<long>` | `{event}` | Events dropped by the publisher. Tagged `tree` and `reason=missing_provider` (no stream provider by the configured name is registered on this silo) or `publish_error` (the stream provider threw during dispatch). A non-zero rate on `missing_provider` means [`LatticeOptions.PublishEvents`](../configuration/options-reference-2.md#publishevents) is `true` but the corresponding `AddMemoryStreams` / `AddEventHubStreams` call is missing on the silo. |

## Configuration

Runtime overrides applied through `ILattice` that mutate per-tree behaviour
emit a lightweight change counter so operators can audit policy changes on
the same pipeline as the traffic they affect.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.config.changed` | `Counter<long>` | `{change}` | A per-tree configuration change was applied. Tagged `tree` and `config` = the configuration dimension: `publish_events` (from `ILattice.SetPublishEventsEnabledAsync`) or `history_retention` (from `ILattice.SetHistoryRetentionAsync`). Both arms are zero-primed whenever the tree's entry point activates (never for a system tree), so a flat zero is a measured zero. |

## Mutation observers

Registered [`IMutationObserver`](../api/mutation-observers.md) callbacks run *inline*
on the grain write path, so every millisecond an observer spends is a millisecond
added to the caller's write latency. This histogram attributes that cost to the
specific observer that incurred it, on the same pipeline as the traffic it slows
down - which is how a misbehaving observer (synchronous I/O, a blocking call, a
chatty downstream) is identified rather than merely suspected.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.observer.duration` | `Histogram<double>` | `ms` | Wall-clock time one registered `IMutationObserver` spent inside a single `OnMutationAsync` callback. Tagged `observer` = the observer's CLR type name and `tree` = the mutated tree. Recorded on the faulting path too, so an observer that throws slowly is still visible; the dispatcher continues to suppress the exception. The sample spans only the callback - the dispatcher's own swallow-and-log work is excluded, so a slow log sink is never billed to the observer. Zero-cost when no observer is registered (the dispatcher returns before any timing work) and elided when no metrics listener is attached. |

The instrument takes one sample per observer per published mutation, so N
registered observers contribute N samples per mutation. `sum by (observer)` over
the `_sum` series ranks observers by total latency contributed, while
`histogram_quantile` over the `_bucket` series exposes a single slow observer that
a mean would hide.

## Materialised views

The view maintainer publishes on the core `orleans.lattice` meter (views need a
WAL-backed lattice, not the replication package), each instrument tagged with the
view name. See [Materialised views](../materialised-views.md).

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.view.apply_lag` | `Histogram<long>` | `{entry}` | Apply lag (committed-but-unapplied source entries) sampled at the end of each drain pass. |
| `orleans.lattice.view.backlog_depth` | `Histogram<long>` | `{entry}` | WAL entries read in the drain pass. |
| `orleans.lattice.view.applied` | `Counter<long>` | `{write}` | View writes applied to the view tree. |
| `orleans.lattice.view.key_collisions` | `Counter<long>` | `{collision}` | View keys that two or more distinct source keys re-mapped to within one drain batch of a filter / re-project view (an injectivity violation in its key re-map), counted once per colliding view key per batch rather than once per source key. A write that carries no source key is ignored, and one source key rewriting the same view key is not a collision. Each colliding view key resolves by source-HLC last-writer-wins, and the maintainer logs a warning naming an example key. |
| `orleans.lattice.view.aggregation_applied` | `Counter<long>` | `{contribution}` | Aggregation contributions folded into the view (count / sum / min / max / set-union / fold). |
| `orleans.lattice.view.aggregation_rejected` | `Counter<long>` | `{contribution}` | Aggregation contributions dropped because the group-key selector produced an empty key or one under the reserved NUL (`\u0000`) prefix; deterministic on the key, so clusters stay convergent. A non-zero value flags a selector emitting reserved keys. |
| `orleans.lattice.view.atomic_staging_backstop` | `Counter<long>` | `{rebuild}` | Times the bounded-buffer / retention backstop abandoned atomic staging and forced a rebuild. |
| `orleans.lattice.view.cross_tree_joint_violation` | `Counter<long>` | `{degradation}` | Cross-tree view batches that degraded to per-tree atomicity because a participant view did not become ready within `CrossTreeReadinessTimeout`. |
| `orleans.lattice.view.lag_budget_eviction` | `Counter<long>` | `{eviction}` | Views force-evicted (WAL unpinned and rebuilt) for exceeding their `MaxLagBudget`. |
| `orleans.lattice.view.source_backpressure` | `Counter<long>` | `{pass}` | View drain passes that throttled themselves because the source tree was under WAL saturation back-pressure - any drain pass, whatever triggered it: the background timer tick, the keepalive reminder, activation, or a read-your-writes `ILatticeView.WaitForSourceHlcAsync` barrier, on filter / re-project and aggregation views alike. Every such pass runs a scaled-down batch; only a background timer tick also defers the next tick, and a deferred tick drains nothing and records nothing. Tagged `view` and `state` (the observed source regime, `throttled` / `saturated`); never recorded on a healthy source or when `ObeySourceBackpressure` is disabled. |

## Tag index reconciliation

Background tag-index reconciliation publishes on the core `orleans.lattice` meter,
each instrument tagged with the `index` name. See [Tag indexes](../api/tag-indexes.md).

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.tag_index.reconcile.sweeps` | `Counter<long>` | `{sweep}` | Background tag-index reconciliation sweeps, tagged by outcome (`clean`, `repaired`, `probe_only`). |
| `orleans.lattice.tag_index.reconcile.trees.probed` | `Counter<long>` | `{tree}` | Covered trees whose digest fingerprint a reconciliation sweep probed. |
| `orleans.lattice.tag_index.reconcile.trees.mismatched` | `Counter<long>` | `{tree}` | Covered trees a reconciliation sweep found divergent from their digest baseline. |
| `orleans.lattice.tag_index.reconcile.orphan_rows.removed` | `Counter<long>` | `{row}` | Orphan membership rows removed by background tag-index reconciliation. |
| `orleans.lattice.tag_index.reconcile.duration` | `Histogram<double>` | `ms` | Wall-clock duration of a background tag-index reconciliation sweep. |

## Grain-call observation (opt-in)

Registered by `ISiloBuilder.AddLatticeGrainCallObservation()`, which installs
`IOutgoingGrainCallFilter` on the silo. It is **opt-in** because it observes
every outgoing grain call the silo makes - not only calls into Lattice grains -
so a host chooses to pay for it. Both instruments are tagged `grain_type` (the
target's Orleans grain type name, low-cardinality) and carry the platform tenant
sentinel: a target activation's queue is the aggregate of every caller's arrivals
and no single tenant owns it, so the series is deliberately not tenant-scoped and
is invisible to a tenant-scoped telemetry query.

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.grain.call.outstanding_depth` | `Histogram<int>` | `{call}` | Calls this silo had already issued to the same target activation and not yet seen complete, sampled **at dispatch** on every outgoing call. On a non-reentrant target this is the depth the new call queues behind. Tagged `grain_type`. |
| `orleans.lattice.grain.call.duration` | `Histogram<double>` | `ms` | Wall-clock duration of an outgoing grain call, from dispatch to completion or fault. Tagged `grain_type` and `outcome` (`completed` / `faulted`). |

### Why this exists, and what it is not

Orleans' only built-in queue signal is the `NonReentrancyQueueSize=` clause of
its `Response did not arrive on time` timeout diagnostic. That clause is
**censored twice over**: the runtime emits it only for a request already
approaching the 30-second response deadline, and the clause reports the
*emitting* request's own wait - so a grain type whose calls queue deeply, but
which does not itself trip the timeout, contributes **no rows at all**. On one
real investigation the grain type carrying the deepest queues in the system
contributed 0 of the diagnostic's 154 samples, and two independent extractions
from those samples agreed - consistently and wrongly - that nothing was queueing.
Agreement between two extractions that apply the same selection predicate
validates the arithmetic, not the sampling frame.

`orleans.lattice.grain.call.outstanding_depth` removes both conditions: it is
recorded at dispatch, before the call is awaited, so no timeout, fault, or
threshold gates the emission, and it describes the target's contention rather
than the emitter's luck.

Three limits, all of which understate rather than invent contention:

- **Per-silo.** Only calls issued from *this* silo are counted, so cluster-wide
  contention on a shared activation is under-stated. Treat the value as a floor.
- **Reentrancy changes the meaning.** On a `[Reentrant]` or `[AlwaysInterleave]`
  target the outstanding calls interleave rather than queue, so a high value
  there means pipelining, not contention.
- **Dispatch, not admission.** The window includes network transit and the
  response hop, not only scheduler queueing.

The companion duration histogram is split by `outcome` deliberately: the
`faulted` series carries the 30-second message timeouts, which would otherwise
pile up at the deadline and dominate any quantile taken over the combined
population - reproducing, inside the new channel, exactly the censoring that
makes the log diagnostic unusable.

Previous: [Instrument catalog: Shard-root SetManyAsync decomposition (sourced from ShardRootGrain) to WAL append pipeline (sourced from WalShardGrain and WalCommitLogWriter)](instrument-catalog-6.md). Next: [Bundled Grafana dashboards](bundled-grafana-dashboards.md). Contents: [Metrics](../metrics.md).
