Table of Contents

Instrument catalog: Storage-provider commit pipeline to Grain-call observation (opt-in)

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at instrument-catalog-7.md, and llms.txt lists every page.

Part of Instrument catalog, in Metrics.

Storage-provider commit pipeline

The Azure Table WAL provider (and any other WAL provider that opts in) emits per-commit-phase histograms so the per-batch partition transaction and the manifest partition transaction can be observed independently. See Write-Ahead Log and WAL Storage Providers for the underlying two-phase commit shape.

Name Kind Unit Description
orleans.lattice.provider.commit.duration Histogram<double> ms Wall-clock duration of one storage-provider commit transaction. Tagged tree, shard, phase=phase1 (per-batch partition transaction) or phase=phase2 (manifest partition transaction), and pipeline_phase2 reflecting AzureTableWalStorageOptions.PipelinePhaseTwoCommits. The phase-2 measurement covers a single coalesced commit transaction, not the per-shard worker's whole drain loop.
orleans.lattice.provider.phase2.batch_size Histogram<int> {commit} Number of coalesced phase-2 commits the per-shard provider worker bundled into a single transaction. Tagged tree, shard. A distribution concentrated near 1 means the worker is never catching up against backed-up arrivals; values closer to the per-transaction cap (~49 commits) indicate the worker is the shard's effective rate limiter.
orleans.lattice.provider.retry.attempts Counter<long> {attempt} Individual retry attempts the storage-provider SDK performs on a WAL append or commit, incremented once per attempt regardless of whether the retry ultimately succeeds. Tagged status (the HTTP status of the response that triggered the retry, e.g. 503 / 429, or 0 for a transport-level failure with no HTTP exchange); per-tree / per-shard correlation is intentionally omitted to keep cardinality low - use provider.retry.exhausted for shard-attributed failures. A non-zero rate is normal; sustained climbing rates suggest provider-side throttling that the local retry policy is masking.
orleans.lattice.provider.retry.exhausted Counter<long> {call} Calls that exhausted the storage-provider retry budget and surfaced the underlying fault. Tagged tree, shard, phase (phase1 per-batch partition transaction or phase2 manifest partition transaction), and status (the HTTP status string the SDK observed, or unknown). Any non-zero rate is alert-worthy - the WAL append or commit failed past the SDK's retry envelope.
orleans.lattice.provider.retry.short_circuited Counter<long> {attempt} Storage-provider SDK retry attempts abandoned early by the Azure Tables provider's saturation-aware retry policy (AzureTableWalStorageOptions.HonorSaturationSignal = true, the default). Tagged status (the synthetic-response HTTP status that signalled the SDK to exit the retry chain; 503 matches the policy's default short-circuit shape). Increments once per abandoned retry, including retries the policy short-circuits during the AzureTableWalStorageOptions.SaturationShortCircuitCooldown sticky window after the last observed Saturated tick. A non-zero rate confirms the policy is firing under storage-side back-pressure; a sustained climbing rate alongside provider.retry.exhausted indicates the storage account is shedding load faster than the SDK retry budget would otherwise surface to the writer. Always zero when HonorSaturationSignal = false (the historical unguarded SDK retry behaviour) or when the host supplies a pre-built TableServiceClient via ServiceClient (the host owns the pipeline in that mode).
orleans.lattice.provider.idempotent_replays Counter<long> {call} Phase-one WAL batch commits whose 409 EntityAlreadyExists conflict was proven to be an idempotent replay of an already-durable write and resolved as a success. Tagged tree, shard, and phase (always phase1 - idempotent replays are only detected on the phase-one batch commit). Increments when the Azure Tables SDK retry pipeline resends a phase-one batch whose first attempt committed server-side but whose response was lost (a socket / read timeout under CPU or network pressure); the provider reads the resident rows back, confirms they are byte-identical to the batch it tried to write, and treats the conflict as success instead of failing the whole batch. A non-zero rate is benign (no data was lost) but is a leading indicator that the cluster is dropping phase-one responses under load - watch it alongside provider.retry.exhausted, not instead of it. A genuine offset collision (resident rows present but differing) is never counted here; it still surfaces as a hard failure on provider.retry.exhausted.
orleans.lattice.provider.phase1.transient_retries Counter<long> {attempt} Phase-1 commit attempts the provider re-issued in place after a transient fault, resubmitting the byte-identical batch at the same offsets rather than faulting the calling WAL shard. Tagged tree, shard, and phase (always phase1).
orleans.lattice.provider.phase2.commit.timeouts Counter<long> {commit} Phase-2 manifest commits abandoned by the per-shard worker after exceeding AzureTableWalStorageOptions.PhaseTwoCommitTimeout. Tagged tree, shard. Zero unless a commit's Azure Tables transaction stopped making progress (hung socket, server-side partition stall, or an SDK retry loop running past the deadline); any non-zero rate is alert-worthy - it is the direct signal that the per-shard phase-2 drain loop would otherwise have wedged. PhaseTwoCommitTimeout defaults to 12 seconds; set it to null and the commit is left unbounded, so this counter never increments.

Saga / coordinator / lifecycle

Long-running maintenance and atomicity primitives each emit a completion signal so operators can alert on stalls, compensation spikes, or missing coordinator progress independently of whether the event stream is enabled.

Name Kind Unit Description
orleans.lattice.atomic_write.completed Counter<long> {saga} Terminal transition of a SetManyAtomicAsync saga - one increment per single-tree saga, and one per participating tree of a cross-tree write. Tagged outcome=committed (the commit decision was recorded and the commit terminals broadcast, so every write became visible), failed (a batch still failed after its retries, so the saga captured the failure message, recorded an abort decision in the transaction registry and broadcast abort terminals), or compensated (an abort with no captured failure message - in the current code, a participating tree of a cross-tree write that its coordinator finalized with abort because another participant did not prepare). Neither abort arm issues per-key rollback writes: prepared writes stay invisible in each leaf's pending-transaction buckets until the abort terminal drops them, so there is nothing to undo. The outcome mapping also names shutdown_refused, but no code path produces it, because the failure-message prefix it tests for is never written: a shutdown refusal throws LatticeShuttingDownException to the caller and leaves the saga at its persisted position to resume on the next activation, recording no terminal transition, so count shutdown refusals at the caller instead (see Atomic Writes). No sample is recorded either for a saturation refusal, which throws LatticeSaturatedException and leaves the saga to resume later from its persisted position (a transaction-registry capacity refusal, LatticeSaturationSource.TxRegistryCapacity, is raised before the saga persists anything, so a retry under the same operation id starts cleanly), or for a batch whose guard predicate fails (a SetManyAtomicWhereAsync call returning AtomicWriteOutcome.PreconditionFailed, or a cross-tree participant voting precondition-failed), which commits nothing and records no terminal transition.
orleans.lattice.atomic_write.duration Histogram<double> ms End-to-end SetManyAtomicAsync saga duration, captured from the first Prepare to Completed and persisted across reminder-driven recovery so the recorded ms reflects true wall-clock cost (including any time the saga was suspended across silo restarts). Tagged with the same outcome values as atomic_write.completed; emitted alongside it on every terminal transition, except for a saga whose persisted state predates the start timestamp (it decodes as zero), for which the duration is suppressed rather than recorded as time since year one. Pair with atomic_write.completed to derive sustained atomic-write throughput and SLO percentiles.
orleans.lattice.atomic_write.batch_size Histogram<int> {entry} Entry count of each SetManyAtomicAsync saga at terminal transition. Tagged with the same outcome values as atomic_write.completed; emitted alongside it. Lets operators correlate p99 saga duration with batch size and detect distribution shifts in caller batch sizing.
orleans.lattice.atomic_write.cross_tree.completed Counter<long> {saga} Terminal transition of a cross-tree atomic-write saga (a SetManyAtomicAsync spanning more than one tree). Tagged outcome=committed or precondition_failed - those two are the whole domain, because the coordinator has no compensation arm of its own, and precondition_failed covers every abort, including one caused by a participant that failed to prepare rather than by a precondition miss (each participant it finalizes with abort reports compensated on orleans.lattice.atomic_write.completed) - and tree_count (number of participating trees). It carries no tree tag and records under the _platform_ tenant, and an empty cross-tree batch commits vacuously without recording anything. A panel filtering this counter for failed or compensated selects zero series forever; DashboardPanelTagDomainTests now fails the build on exactly that mistake.
orleans.lattice.atomic_write.cross_tree.duration Histogram<double> ms End-to-end cross-tree atomic-write coordinator duration. Tagged outcome.
orleans.lattice.atomic_write.cross_tree.participants Histogram<int> {tree} Participating-tree count of each cross-tree atomic-write saga. Tagged outcome.
orleans.lattice.saga.prepare.duration Histogram<double> ms Wall-clock ms inside one run of the saga's staging (execute) phase: the batched ILattice.SetManyAsync dispatch that stages every per-key prepared mutation across all touched shards, together with the saga checkpoint persists and batch retries that phase issues, recorded whether the phase completes or throws. Tagged tree and wal_partitions. Pair with set_many.duration and saga.checkpoint.duration to confirm the saga prepare is dominated by the foreground multi-key write rather than saga-internal bookkeeping.
orleans.lattice.saga.terminal_decision.duration Histogram<double> ms Wall-clock ms inside one saga TerminalDecision phase (decide whether to commit or compensate after every prepared mutation has acknowledged). Tagged tree and wal_partitions. Sub-millisecond except under failure paths.
orleans.lattice.saga.broadcast.duration Histogram<double> ms Wall-clock ms inside one saga Broadcast phase (the Task.WhenAll across affected shards that flips every prepared mutation to its terminal state). Tagged tree and wal_partitions. The c2-xxiii batched-WAL terminal lift collapsed this phase from ~880ms p50 to ~96ms p50 at the 500:10 rung.
orleans.lattice.saga.broadcast.shard.duration Histogram<double> ms Per-shard contribution inside the saga broadcast: wall-clock ms inside one ShardRootGrain.AppendTxTerminalAsync call. Tagged tree and shard. The gap between saga.broadcast.duration and this is the max-of-N parallel tail (Orleans scheduling + the slowest shard).
orleans.lattice.saga.broadcast.shard.stage.duration Histogram<double> ms Per-sub-stage wall-clock ms inside one ShardRootGrain.AppendTxTerminalAsync call. Tagged tree, shard, and stage=resolve (affected-leaves resolution), hlc (terminal HLC compute via GetClockAsync fan-out), wal (the shard root's WAL terminal append; collapsed to ~0ms after c2-xxiii batched the WAL terminal at the saga layer), or fanout (the per-leaf ApplyTxTerminalAsync fan-out).
orleans.lattice.saga.broadcast.leaf.duration Histogram<double> ms Wall-clock ms inside a single per-leaf IBPlusLeafGrain.ApplyTxTerminalAsync RPC dispatched from the shard's terminal broadcast (step 4). Tagged tree and shard.
orleans.lattice.saga.checkpoint.duration Histogram<double> ms Wall-clock ms inside one persist of the saga's state (a single WriteStateAsync), recorded at every site that persists it rather than only at the terminal decision, and on the failure path too. Tagged tree, wal_partitions and phase, the call site: prepare, prepared, execute-batch-commit, execute-batch-retry, execute-to-compensate, broadcast-touched-init, broadcast-touched-expand, broadcast-touched-drift, broadcast-touched-late or complete.
orleans.lattice.saga.reminder.duration Histogram<double> ms Wall-clock ms inside the Orleans reminder-registry calls that manage the saga's keepalive reminder, its crash-recovery anchor. Tagged tree, wal_partitions and phase: register (registering or refreshing the reminder, including the bounded wait while the silo's reminder service is still initialising), unregister-get (looking the reminder up to tear it down once the saga no longer needs it) or unregister-drop (removing it).
orleans.lattice.saga.fanout.size Histogram<int> {entry} Entry count per atomic-write saga, observed at execute-phase entry. Tagged tree and wal_partitions. The distribution tells operators whether the workload is dominated by small single-entry sagas or large multi-entry batches.
orleans.lattice.saga.perkey.duration Histogram<double> ms Amortised per-key cost of a saga's prepare dispatch. Each staging batch is sent as one ILattice.SetManyAsync, so per-key timing is not observable: the batch's wall-clock time divided by the number of keys it carried is recorded once per key, so every sample from one dispatch is identical, and a batch that faulted is recorded as well. Tagged tree and wal_partitions. Pair with atomic_write.batch_size to derive amortised per-key atomic-write cost across the saga envelope.
orleans.lattice.tx_registry.writes Counter<long> {write} Whole-state writes issued by the saga decision registry. Tagged tree and outcome=ok (the write completed durably) or fault (the write threw; every mutation it carried, and any queued behind it, was rolled back and failed to its caller, which retries). A registry shard raises the tree's shard high-water mark before its first write; if that raise fails, the pending mutations are rolled back and failed the same way and one fault is recorded although no state write was issued. Every atomic saga records its participants, decision and cleanup through its registry, so the registry's write rate bounds saga throughput. At the default TxRegistryShardCount of 1 that is one registry activation, and one persisted row, per tree; a higher count mints new sagas across that many registry shards, each its own activation with its own row, beside the legacy per-tree registry that sagas minted earlier keep routing to. Every shard records under the logical tree id, so a tree's series sums all of its registries. Each registry group-commits: at most one write is in flight, and mutations that arrive meanwhile join the next write (#3475). A write rate that stays flat while tx_registry.write.mutations climbs is coalescing at work. Any sustained fault rate is alert-worthy.
orleans.lattice.tx_registry.write.mutations Histogram<int> {mutation} Registry mutations carried by one saga decision registry state write: the group-commit coalescing factor. Tagged tree and outcome (as tx_registry.writes). Each mutation is one caller's call, so on the ok arm 1 means the write acknowledged a single caller and N means one durable write acknowledged N callers; on the fault arm it counts the mutations the failed write carried, not the queued ones that fail with it. A p50 stuck at 1 under heavy atomic load means there was no concurrency to coalesce.
orleans.lattice.tx_registry.write.duration Histogram<double> ms Duration of one saga decision registry whole-state write, measured around the storage call (for a fault caused by a failed shard high-water raise, around that raise instead). Tagged tree and outcome (as tx_registry.writes). Each registry serialises its writes, so the reciprocal of this duration is the most writes per second one registry can issue - a whole tree's ceiling at the default TxRegistryShardCount of 1, and one shard's share of it above that. It grows with the registry row, which carries a tombstone per completed saga for TxDecisionRetention; see TxRegistryAdmissionBudgetBytes in Configuration.
orleans.lattice.coordinator.completed Counter<long> {operation} Successful completion of a long-running coordinator. Tagged tree and kind=snapshot, resize, reshard, merge, or compaction.
orleans.lattice.coordinator.phase_tick.failures Counter<long> {failure} A coordinator phase-timer tick whose phase step threw and was swallowed. That tick advanced the phase machine by nothing, so the step it would have taken is retried from the start and any work it had accumulated is discarded. Tagged kind (the coordinator's keepalive reminder name, for example resize-keepalive or reshard-keepalive - not the bare kind coordinator.completed carries), tree, and the tenant label. Zero-primed when a coordinator arms its phase timer, so a flat zero on a live series is a reading - this coordinator has swallowed nothing - rather than the absence a counter reports before its first Add. A zero does not mean the coordinator is progressing: a tick that returns normally without advancing is a success here, so denominate against coordinator.completed for that question. The first two consecutive failures on one activation log at warning and the third onwards at error, because a run of them is a phase loop that has stopped advancing rather than a transient the pump absorbs.
orleans.lattice.coordinator.phase_tick.consecutive_failures ObservableGauge<long> {failure} Length of the current run of consecutive failed coordinator phase ticks, tagged identically to coordinator.phase_tick.failures so the two series join. The counter beside it cannot express consecutiveness: a coordinator that fails one tick in a thousand and one that has failed every tick since process start both present as a rising total, yet the first is a transient the pump absorbs by design and the second is a phase machine that has stopped advancing. Every live coordinator reports, enrolling at 0 when it arms its phase timer, so 0 means the last tick succeeded rather than no data - and, exactly as for the counter, 0 does not mean the coordinator is advancing, because a tick that returns without moving the phase machine forward is a success by this measure. Reported as the maximum over the activations sharing a tag set, which is coarser than the activation because a coordinator with a composite key deliberately reports under the subject alone; max is the correct reduction because a sum would invent a run no activation experienced and a last-writer-wins would let a healthy sibling hide a wedged one. A value that keeps returning to zero is a coordinator absorbing transients; a value that only climbs is a wedge, and its magnitude is how many ticks of work have been discarded back to back.
orleans.lattice.tree.lifecycle Counter<long> {event} Tree-lifecycle transition, recorded by the tree's soft-delete and purge manager. Tagged tree and kind=deleted, recovered, or purged. Emitted unconditionally - regardless of the tree's PublishEvents setting. Exactly one increment per logical delete, recover or purge, tagged with the logical tree id, including on an aliased tree; the work a logical operation delegates to the backing physical tree records nothing of its own. A resize's retirement of its old physical copy, and a resize undo's recovery or discard of a copy, are physical maintenance and record nothing.
orleans.lattice.warmup.invocations Counter<long> {call} One increment per successful ILattice.WarmUpAsync call. Tagged tree. Operators alerting on cold-start health expect to see exactly one increment per silo startup per warmed tree.
orleans.lattice.warmup.duration Histogram<double> ms End-to-end duration of ILattice.WarmUpAsync - the wall-clock cost of pre-activating every physical shard root via a bounded-concurrency read-only probe. Tagged tree and shard_count (the per-tree physical-shard-root probe fan-out). The p99 is the primary warm-start latency signal; sustained increases are a leading indicator of placement-directory or grain-storage cold-touch cost growth.
orleans.lattice.warmup.leaf_cache.prewarmed Counter<long> {leaf} Leaf caches successfully primed by a shard root's post-restart pre-warm. Tagged tree, shard, and the tenant label. Flat at zero only where LatticeOptions.LeafCachePreWarmCount has been set to 0; the default is 8, so the feature is on unless it is explicitly turned off. Individual priming failures are swallowed by design, so a value materially below the configured count is the only signal that they occurred.
orleans.lattice.warmup.leaf_cache.duration Histogram<double> ms Wall-clock duration of one shard root's leaf-cache pre-warm fan-out. One observation per shard per warm-up when the feature is enabled and the access model ranked at least one leaf. Tagged tree, shard, and the tenant label. Read against warmup.duration to attribute how much of a tree's warm-start cost is leaf priming.
orleans.lattice.leaf_access.model.leaves Histogram<int> {leaf} Leaves resident in a shard root's leaf-access frequency histogram, observed each time the model is persisted. Tagged tree, shard, and the tenant label. Bounded above by the model's tracked-leaf cap, so a distribution pinned at that cap means the shard's read set is wider than the model can represent and the pre-warm ranking is drawn from a pruned view.

Distributed lock (sourced from LatticeLockGrain)

The FIFO-fair distributed lock / lease grain behind ILatticeLockGrain (see Distributed lock). One activation per lock name serialises all contending callers, so these instruments carry no per-key tag; scope them by cluster (and silo) in the dashboards. Charted on the Overview dashboard's "Distributed lock" row.

Name Kind Unit Description
orleans.lattice.lock.acquired Counter<long> {acquire} One increment per acquire attempt that reached a terminal outcome. Tagged outcome=granted (the caller was handed the lease), timeout (a blocking AcquireAsync whose wait-timeout elapsed, or a non-blocking MaxWait = 0 acquire on a held lock), or unavailable (a TryAcquireAsync that returned null because the lock was held). Pair granted with released to see live lock churn.
orleans.lattice.lock.released Counter<long> {release} One increment per explicit ReleaseAsync by the current holder (a stale-token release is a silent no-op and is not counted).
orleans.lattice.lock.lease_reclaimed Counter<long> {lease} One increment each time an expired lease was reclaimed because the holder neither renewed nor released before expiry. A sustained non-zero rate means holders are crashing or failing to renew within their lease duration; downstream resources should be relying on the fencing token to reject the reclaimed holder.
orleans.lattice.lock.acquire.wait Histogram<double> ms Time a granted acquire spent parked in the FIFO queue before it was granted (recorded once per grant, so an uncontended acquire records ~0). The p95/p99 tail is the primary lock-contention signal.

Atomic action / saga (sourced from AtomicActionGrain)

The generic atomic-action (saga / TCC) coordinator behind IAtomicActionGrain (see Atomic action). One activation per operation id, so these instruments carry no per-key tag; scope them by cluster (and silo) in the dashboards. Charted on the Overview dashboard's "Atomic action (saga / TCC)" row.

Name Kind Unit Description
orleans.lattice.atomic_action.completed Counter<long> {saga} One increment per terminal transition of a saga. Tagged outcome=committed (every forward step committed), compensated (a forward step faulted and every committed step was rolled back in reverse order), or compensation_failed (a compensating effect itself faulted, so the saga parked for operator intervention). A rising compensation_failed rate means a caller's compensation contract is being violated.
orleans.lattice.atomic_action.step Counter<long> {step} One increment per step effect the saga runs. Tagged phase=forward | compensate and outcome=ok | fault. The compensate series is non-zero only when a saga rolls back; a compensate,fault point is the leading indicator of a park.
orleans.lattice.atomic_action.duration Histogram<double> ms End-to-end saga duration measured from saga start (persisted, so it includes time suspended across silo restarts) to the terminal transition. Tagged by outcome so rollback-path latency is separable from happy-path latency.

Events

These counters are emitted only when event publication is enabled on at least one tree. They let operators detect a misconfigured stream provider or a failing downstream queue before it starts consuming silo resources.

Name Kind Unit Description
orleans.lattice.events.published Counter<long> {event} LatticeTreeEvent instances successfully dispatched to the configured stream provider. Tagged tree and kind = the LatticeTreeEventKind name (e.g. Set, SnapshotCompleted).
orleans.lattice.events.dropped Counter<long> {event} Events dropped by the publisher. Tagged tree and reason=missing_provider (no stream provider by the configured name is registered on this silo) or publish_error (the stream provider threw during dispatch). A non-zero rate on missing_provider means LatticeOptions.PublishEvents is true but the corresponding AddMemoryStreams / AddEventHubStreams call is missing on the silo.

Configuration

Runtime overrides applied through ILattice that mutate per-tree behaviour emit a lightweight change counter so operators can audit policy changes on the same pipeline as the traffic they affect.

Name Kind Unit Description
orleans.lattice.config.changed Counter<long> {change} A per-tree configuration change was applied. Tagged tree and config = the configuration dimension: publish_events (from ILattice.SetPublishEventsEnabledAsync) or history_retention (from ILattice.SetHistoryRetentionAsync). Both arms are zero-primed whenever the tree's entry point activates (never for a system tree), so a flat zero is a measured zero.

Mutation observers

Registered IMutationObserver callbacks run inline on the grain write path, so every millisecond an observer spends is a millisecond added to the caller's write latency. This histogram attributes that cost to the specific observer that incurred it, on the same pipeline as the traffic it slows down - which is how a misbehaving observer (synchronous I/O, a blocking call, a chatty downstream) is identified rather than merely suspected.

Name Kind Unit Description
orleans.lattice.observer.duration Histogram<double> ms Wall-clock time one registered IMutationObserver spent inside a single OnMutationAsync callback. Tagged observer = the observer's CLR type name and tree = the mutated tree. Recorded on the faulting path too, so an observer that throws slowly is still visible; the dispatcher continues to suppress the exception. The sample spans only the callback - the dispatcher's own swallow-and-log work is excluded, so a slow log sink is never billed to the observer. Zero-cost when no observer is registered (the dispatcher returns before any timing work) and elided when no metrics listener is attached.

The instrument takes one sample per observer per published mutation, so N registered observers contribute N samples per mutation. sum by (observer) over the _sum series ranks observers by total latency contributed, while histogram_quantile over the _bucket series exposes a single slow observer that a mean would hide.

Materialised views

The view maintainer publishes on the core orleans.lattice meter (views need a WAL-backed lattice, not the replication package), each instrument tagged with the view name. See Materialised views.

Name Kind Unit Description
orleans.lattice.view.apply_lag Histogram<long> {entry} Apply lag (committed-but-unapplied source entries) sampled at the end of each drain pass.
orleans.lattice.view.backlog_depth Histogram<long> {entry} WAL entries read in the drain pass.
orleans.lattice.view.applied Counter<long> {write} View writes applied to the view tree.
orleans.lattice.view.key_collisions Counter<long> {collision} View keys that two or more distinct source keys re-mapped to within one drain batch of a filter / re-project view (an injectivity violation in its key re-map), counted once per colliding view key per batch rather than once per source key. A write that carries no source key is ignored, and one source key rewriting the same view key is not a collision. Each colliding view key resolves by source-HLC last-writer-wins, and the maintainer logs a warning naming an example key.
orleans.lattice.view.aggregation_applied Counter<long> {contribution} Aggregation contributions folded into the view (count / sum / min / max / set-union / fold).
orleans.lattice.view.aggregation_rejected Counter<long> {contribution} Aggregation contributions dropped because the group-key selector produced an empty key or one under the reserved NUL (\u0000) prefix; deterministic on the key, so clusters stay convergent. A non-zero value flags a selector emitting reserved keys.
orleans.lattice.view.atomic_staging_backstop Counter<long> {rebuild} Times the bounded-buffer / retention backstop abandoned atomic staging and forced a rebuild.
orleans.lattice.view.cross_tree_joint_violation Counter<long> {degradation} Cross-tree view batches that degraded to per-tree atomicity because a participant view did not become ready within CrossTreeReadinessTimeout.
orleans.lattice.view.lag_budget_eviction Counter<long> {eviction} Views force-evicted (WAL unpinned and rebuilt) for exceeding their MaxLagBudget.
orleans.lattice.view.source_backpressure Counter<long> {pass} View drain passes that throttled themselves because the source tree was under WAL saturation back-pressure - any drain pass, whatever triggered it: the background timer tick, the keepalive reminder, activation, or a read-your-writes ILatticeView.WaitForSourceHlcAsync barrier, on filter / re-project and aggregation views alike. Every such pass runs a scaled-down batch; only a background timer tick also defers the next tick, and a deferred tick drains nothing and records nothing. Tagged view and state (the observed source regime, throttled / saturated); never recorded on a healthy source or when ObeySourceBackpressure is disabled.

Tag index reconciliation

Background tag-index reconciliation publishes on the core orleans.lattice meter, each instrument tagged with the index name. See Tag indexes.

Name Kind Unit Description
orleans.lattice.tag_index.reconcile.sweeps Counter<long> {sweep} Background tag-index reconciliation sweeps, tagged by outcome (clean, repaired, probe_only).
orleans.lattice.tag_index.reconcile.trees.probed Counter<long> {tree} Covered trees whose digest fingerprint a reconciliation sweep probed.
orleans.lattice.tag_index.reconcile.trees.mismatched Counter<long> {tree} Covered trees a reconciliation sweep found divergent from their digest baseline.
orleans.lattice.tag_index.reconcile.orphan_rows.removed Counter<long> {row} Orphan membership rows removed by background tag-index reconciliation.
orleans.lattice.tag_index.reconcile.duration Histogram<double> ms Wall-clock duration of a background tag-index reconciliation sweep.

Grain-call observation (opt-in)

Registered by ISiloBuilder.AddLatticeGrainCallObservation(), which installs IOutgoingGrainCallFilter on the silo. It is opt-in because it observes every outgoing grain call the silo makes - not only calls into Lattice grains - so a host chooses to pay for it. Both instruments are tagged grain_type (the target's Orleans grain type name, low-cardinality) and carry the platform tenant sentinel: a target activation's queue is the aggregate of every caller's arrivals and no single tenant owns it, so the series is deliberately not tenant-scoped and is invisible to a tenant-scoped telemetry query.

Name Kind Unit Description
orleans.lattice.grain.call.outstanding_depth Histogram<int> {call} Calls this silo had already issued to the same target activation and not yet seen complete, sampled at dispatch on every outgoing call. On a non-reentrant target this is the depth the new call queues behind. Tagged grain_type.
orleans.lattice.grain.call.duration Histogram<double> ms Wall-clock duration of an outgoing grain call, from dispatch to completion or fault. Tagged grain_type and outcome (completed / faulted).

Why this exists, and what it is not

Orleans' only built-in queue signal is the NonReentrancyQueueSize= clause of its Response did not arrive on time timeout diagnostic. That clause is censored twice over: the runtime emits it only for a request already approaching the 30-second response deadline, and the clause reports the emitting request's own wait - so a grain type whose calls queue deeply, but which does not itself trip the timeout, contributes no rows at all. On one real investigation the grain type carrying the deepest queues in the system contributed 0 of the diagnostic's 154 samples, and two independent extractions from those samples agreed - consistently and wrongly - that nothing was queueing. Agreement between two extractions that apply the same selection predicate validates the arithmetic, not the sampling frame.

orleans.lattice.grain.call.outstanding_depth removes both conditions: it is recorded at dispatch, before the call is awaited, so no timeout, fault, or threshold gates the emission, and it describes the target's contention rather than the emitter's luck.

Three limits, all of which understate rather than invent contention:

  • Per-silo. Only calls issued from this silo are counted, so cluster-wide contention on a shared activation is under-stated. Treat the value as a floor.
  • Reentrancy changes the meaning. On a [Reentrant] or [AlwaysInterleave] target the outstanding calls interleave rather than queue, so a high value there means pipelining, not contention.
  • Dispatch, not admission. The window includes network transit and the response hop, not only scheduler queueing.

The companion duration histogram is split by outcome deliberately: the faulted series carries the 30-second message timeouts, which would otherwise pile up at the deadline and dominate any quantile taken over the combined population - reproducing, inside the new channel, exactly the censoring that makes the log diagnostic unusable.