---
title: "Instrument catalog: Process identity to Shard-level and tree registry - Metrics"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/metrics/instrument-catalog-1.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/metrics.md?plain=1#L174-L208"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Instrument catalog: Process identity to Shard-level and tree registry

Part of [Instrument catalog](instrument-catalog.md), in [Metrics](../metrics.md).

## Process identity

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.build.info` | `ObservableGauge<long>` | `{build}` | Always `1`, tagged with the `version` and the full 40-character `sha` the running process was built from. **This is the only instrument that reports a property of the process rather than of a feature**, and it exists because "is the build under test the build that is actually running" cannot be answered from any feature metric: every other instrument in this catalog stays silent until its feature is exercised, so a missing series is ambiguous between "the image did not deploy" and "the code deployed but that path was never reached" - opposite conclusions drawn from byte-identical evidence. It is emitted unconditionally on every collection from process start, with no registry to populate and no work to wait for, so it has no empty state to prime: an absent series means the process is not running or is not exporting, and means nothing else. The value carries no information and is always `1`; all of the content is in the tags, which is the conventional info-metric shape and keeps the series safe to join against in a query. The sha is sourced from the `+<sha>` build-metadata suffix the SDK appends to `AssemblyInformationalVersionAttribute` from `SourceRevisionId`, so it comes **from the build, never from a runtime `git` call** - a deployment container is distroless and has no git, which is precisely where the value matters most - and it identifies the **image** rather than the checkout the process happens to be running beside. Carries `version` and `sha` only: cardinality on an info gauge multiplies across every other series a reader joins it against. A placeholder sha would be worse than an absent one, because it reads as a working detector while identifying nothing, so `BuildInfoMetricTests` fails outright on an empty or sentinel value rather than tolerating it. |

## Shard-level and tree registry

| Name | Kind | Unit | Description |
|---|---|---|---|
| `orleans.lattice.registry.call.duration` | `Histogram<double>` | `ms` | Service time of one tree-registry call, measured **inside** the registry singleton's grain body. Tagged `operation` = `exists`, `get_entry`, `resolve`, `get_shard_map`, `get_entries`, or `get_all_tree_ids` for the reads, and `register` or `unregister` for the two mutators the census also covers; `get_all_tree_ids`, `register` and `unregister` are not `[AlwaysInterleave]`, so a short service time on those arms beside a low admitted count does not by itself mean calls went unserved. The registry is a cluster singleton every per-tree background service addresses, so a cold start fans a whole estate onto one activation and callers see only a response-deadline `TimeoutException` - the same observation whether the call was served slowly or never served at all, which have opposite remedies. Because only admitted calls are recorded here, reading the tail **and the count** against the caller-side timeout population separates them: a short tail with a count far below the offered load means the calls never reached the body (blocked upstream, in activation or in the turn queue, neither of which an `[AlwaysInterleave]` attribute can admit past), while a tail approaching the caller deadline with a matching count means they were admitted and the time went on the awaited hop to the backing `_lattice_trees` tree. Deliberately not derived from `orleans-storage-read-latency`: that instrument's `state_name` tag comes from a grain's `[PersistentState]` declaration and the registry grain declares none, so no `state_name` series for it can exist. |
| `orleans.lattice.registry.call.in_flight` | `Histogram<int>` | `{call}` | Concurrent tree-registry calls in flight on the registry singleton at the moment a new call is admitted, tagged `operation` as above. The count is global across every `operation` arm, so it is the fan-in width, not evidence of whether a given member interleaves. The recorded value excludes the arriving call, so it is `0` on the first concurrent call and `1` on the second - the same convention as `orleans.lattice.leaf.commit.in_flight`. This is the fan-in **width** the singleton carries, which the duration histogram alone cannot supply: a long tail over a flat-zero width is a slow backing hop, whereas a long tail whose width climbs with estate size is the (trees x per-tree background services) scaling law saturating one activation. |
| `orleans.lattice.registry.caller.duration` | `Histogram<double>` | `ms` | Duration of one `ILatticeRegistry` call as observed by its **caller**, from dispatch to response, timeout, or fault (issue #3088). Tagged `method` (every interface member, not only the reads the census wraps) and `outcome` = `completed`, `timeout` (a `TimeoutException` - the response deadline), or `faulted`. Recorded by the caller-side decorator every production caller obtains the registry through - **not opt-in**, yet charged to no other grain call, because it is not an outgoing grain call filter - for every registry call the process dispatches, whatever becomes of it. That is what `registry.call.duration` cannot do: it sees only calls the grain admitted, so an unreachable registry is silent there. Read the three states directly: **fine** is `completed` with a short tail, **slow** is `completed` with a long tail, **unreachable** is a non-zero `timeout` rate - and `timeout` on a member here with no admitted calls for it in `registry.call.duration` means the calls were never served. Replaces the Orleans `Diagnostics: [... CurrentlyExecuting=...]` timeout-log field, which Orleans stops emitting silo-wide under saturation and whose absence reads as an idle registry. Per-process; a `timeout` sample is pinned at the deadline, which is why it is its own arm. Platform-tenant series. |
| `orleans.lattice.registry.admission.wait` | `Histogram<double>` | `ms` | Time one tree id spent waiting for a fan-in permit before its batched registry read was dispatched, measured **caller-side** in `RegistryFanInGate`. Recorded unconditionally on every dequeue, including a zero wait, so the population is every gated read rather than only the slow ones. This exists because bounding fan-in does not remove work, it queues it: once concurrent reads are capped, a saturated registry stops producing registry-side symptoms and starts producing caller-side waiting instead. Every other registry signal - both instruments above, and `orleans_app_requests_timedout_total{grain_type="latticeregistry"}` - is scoped to the registry grain, so a stall that relocates into the admission queue drives all of them toward zero and reads as a clean recovery. This is the only instrument that observes the relocated wait, which makes it the one signal that distinguishes "the bound is working" from "the bound is now the constraint": a rising upper percentile here **without** a matching rise in `orleans.lattice.registry.call.duration` means the gate, not the registry, is what callers are waiting on. Admission is deliberately uncapped and untimed - shedding a cold-start options resolution would turn latency into a correctness failure, since the caller cannot proceed without its entry - so unbounded waiting is a real possible state and this histogram is the only thing that would show it. |
| `orleans.lattice.registry.admission.in_flight` | `Histogram<int>` | `{dispatch}` | Fan-in permits held by **this silo's** registry fan-in gate at the moment it dispatches a gated round trip - the width this silo's share of the gate's bound actually caps. The bound is 16 concurrent reads across the whole cluster, divided by the live silo count and floored at one per silo, so a single silo may hold all 16 and each of four silos at most 4. It is not a duplicate of `orleans.lattice.registry.call.in_flight`, and confusing the two is the specific error it exists to prevent: that instrument counts calls executing inside the registry singleton's body, summed over every caller in the cluster and including callers that never pass through a gate at all, so it has a different population and a different ceiling. Comparing it against that bound compares two quantities that were never the same number, and a low reading there is not evidence that the bound has headroom. **The recorded value includes the dispatch being recorded**, so it runs from `1` to this silo's share and a recorded value equal to the share means the ceiling was reached. That deliberately inverts the exclude-the-arrival convention of `orleans.lattice.registry.call.in_flight` and `orleans.lattice.leaf.commit.in_flight`: under that convention a fully saturated gate would top out one below its own bound, so saturation would be indistinguishable from headroom by inspection - an off-by-one that makes the bound's binding state unobservable, which is precisely the failure this instrument removes. Read the distribution, never the window mean: permit occupancy is bursty, so a mean is dominated by idle time and measures something other than the peak. The actionable reading is the fraction of dispatches at the ceiling: on a single silo, whose share is the whole 16, that is `_count` minus `_bucket{le="15"}`, against `_count` (buckets are cumulative, so `_bucket{le="16"}` holds every dispatch and always equals `_count`), and on a larger cluster the same reading taken one below that silo's share. That needs a bucket boundary one below the ceiling. The instrument declares no bucket advice, so under the OpenTelemetry SDK's default boundaries (`0, 5, 10, 25, ...`) every width from 11 to 16 lands in the one `le="25"` bucket; configure an explicit-bucket view for it before reading the ceiling share. |
| `orleans.lattice.registry.admission.batch.size` | `Histogram<int>` | `{tree}` | Distinct tree ids carried by one gated registry round trip, recorded once per dispatch. A value of `1` is the single-key registry read a silo issues when nothing is queued behind the bound - byte-for-byte the traffic it had before the gate existed - and `2` or more is a batched multi-key registry read, so the share above one is the share of reads the bound actually coalesced. That share is the only direct evidence the batching half of the gate ran at all; a population sitting almost entirely at `1` is a statement about the offered load, not about the batching. Read beside `orleans.lattice.registry.admission.in_flight`: a batch larger than one forms from ids already queued when a dispatch is taken, which happens chiefly once the permits are saturated but also when arrivals on concurrent threads enqueue before the first of them dispatches, so a batch-size distribution above one does not on its own show the bound binding - a width distribution reaching this silo's share does. |
| `orleans.lattice.registry.admission.queue.depth` | `Histogram<int>` | `{tree}` | Distinct tree ids already waiting for a fan-in permit when another arrives, counting the arrival, recorded at enqueue. This is the **offered** fan-in, and it is the signal that separates "the bound had room" from "nothing asked for it". Admission dispatches synchronously on the arriving thread whenever a permit is free, so when the gate is not binding every other gate instrument reports its structural floor - a wait of microseconds, a width of `1`, a batch of `1` - and those floors are visually identical to a comfortable bound. They are not a weak measurement of headroom; they are what absent demand looks like. A depth that stays at `1` says the demand never arrived, and until it rises above `1` no reading from the other three is evidence about the bound at all. This is the instrument to check first when a fan-in experiment comes back green. |
| `orleans.lattice.shard.reads` | `Counter<long>` | `{op}` | Read **operations** served by a shard root (`GetAsync`, `ExistsAsync`, scan, count, etc.). One increment per operation, not per record returned. Tagged `tree` and `shard`. |
| `orleans.lattice.shard.writes` | `Counter<long>` | `{op}` | Write **operations** served by a shard root (`SetAsync`, `DeleteAsync`, `MergeManyAsync`, `BulkLoadAsync`, etc.). One increment per operation: a batched or bulk call (`SetManyAsync`, `MergeManyAsync`, `DeleteRangeAsync`, `SetManyWherePredicateAsync`, `BulkLoadAsync`, `BulkLoadRawAsync`, `BulkAppendAsync`) counts **once regardless of entry count**, so a 5000-record import advances this by only the number of bulk operations. Use `orleans.lattice.shard.records_written` for the record rate. Tagged `tree` and `shard`. |
| `orleans.lattice.shard.records_written` | `Counter<long>` | `{record}` | Individual **records** written by a shard root - the per-record companion to `orleans.lattice.shard.writes`. Incremented by 1 on a single-key write, by the batch size on `SetManyAsync` / `MergeManyAsync` / bulk load, and by the affected count on `DeleteRangeAsync` / `SetManyWherePredicateAsync` (published once the operation succeeds). Plot both: their ratio is the effective batch size. Tagged `tree` and `shard`. |
| `orleans.lattice.shard.digest_reads` | `Counter<long>` | `{op}` | Projection-digest reads served by a shard root - one increment per whole-shard digest read (`ILattice.GetLeafProjectionDigestAsync`) and one per key-range digest read (`ILattice.GetLeafProjectionDigestForRangeAsync`), each served by the one shard the call names. A whole-tree poll of `ILattice.GetLeafProjectionDigestAsync` produces exactly `shardCount` increments; beyond the key-range reads, a higher rate signals a regression that fell back to walking every leaf. Tagged `tree` and `shard`. |
| `orleans.lattice.shard.splits_committed` | `Counter<long>` | `{split}` | Adaptive shard-split commits - fired once per successful `ShardMap` swap, when the split coordinator finalises. Tagged `tree` and `shard` (the source shard that split). |
| `orleans.lattice.shard.consolidations_committed` | `Counter<long>` | `{consolidation}` | Online shard-consolidation commits - fired once per fold that retires a donor shard from the `ShardMap`, whether automatic healing or a shrinking `ILattice.ReshardAsync` started the fold. It is recorded at the fold's final step, and only after that step's terminal state write succeeds, so an increment is always a durably committed fold rather than an attempt. The same step releases the donor's storage, except where the routing map cannot be read or still routes a slot to the donor, or the donor refuses retirement: the fold then commits as a routing-only retirement that leaves the donor's storage in place, and still counts here. The exact inverse of `orleans.lattice.shard.splits_committed`; plot both, because a sustained gap between them is a tree whose physical shard count is still climbing. Tagged `tree` and `shard` (the retired donor). |
| `orleans.lattice.shard.healing.backlog` | `Histogram<int>` | `{shard}` | Physical shards a tree carries **above** its configured base shard count, sampled once per `ShardHealingOrchestratorGrain` sweep - the automatic over-split healing work outstanding. Tagged `tree`. A tree is healed exactly when this reaches zero, so "trees healed" is read straight off this instrument (`count` of series at zero) rather than from a second counter that could disagree with it. Pair it with `orleans.lattice.shard.consolidations_committed` for the reclaim rate: backlog answers how much damage is left, that counter answers whether it is going down. |
| `orleans.lattice.shard.healing.decisions` | `Counter<long>` | `{decision}` | Automatic over-split healing sweeps by outcome - exactly one increment per tree per sweep. Tagged `tree` and `decision` (`admitted`, `disabled`, `admission_closed`, `not_over_split`, `skewed_load`, `split_in_flight`, `tree_maintenance`, `cooldown`, `backpressure`, `at_capacity`, `no_foldable_pair`; a `not_observed` arm is also mapped, for the persisted default before any sweep, but a sweep never publishes it). The series currently advancing for a tree **is** that tree's current healing decision, which is what distinguishes a tree that needs no healing (`not_over_split`) from one that needs healing and is being held back (`skewed_load`, `backpressure`, `cooldown`) from one where the mechanism is switched off (`disabled`, `admission_closed`). |
| `orleans.lattice.split.retroactive_forward.entries` | `Counter<long>` | `{entry}` | Pending prepared mutations retroactively shadow-forwarded from a source shard's leaf chain into the destination shard's `_pendingTx` buckets at the start of an adaptive split's `BeginShadowWrite` phase. Tagged `tree` and `shard` (source). |
| `orleans.lattice.split.retroactive_forward.duration` | `Histogram<double>` | `ms` | Wall-clock duration of the retroactive prepared-mutation sweep before the split coordinator transitions to the `Drain` phase. Tagged `tree` and `shard` (source). |
| `orleans.lattice.split.in_flight` | `Histogram<long>` | `{split}` | Per-tree count of shard migrations in flight: shards that are the source of an unfinished split or the donor of a consolidation fold, both of which carry the same durable per-shard migration record, so an automatic-healing fold's donor is counted alongside the adaptive splits. Sampled once per hot-shard monitor pass that reaches its shard poll - a pass records nothing while the tree's `AutoSplitEnabled` is off, before the tree reaches `AutoSplitMinTreeAge`, or while a resize, reshard, merge or snapshot of the tree is in flight, so the splits and folds a reshard drives are not sampled while it runs. Every migration counted here occupies one of the tree's `MaxConcurrentAutoSplits` slots - a healing fold in flight leaves one fewer for an autonomic split - and the same count, plus the splits the pass triggers, is the footprint the tree reports to the cluster-wide split gate and that `ILatticeAdmin.GetSplitActivityAsync` reads. Tagged `tree`. Emitted on every such pass **regardless** of whether the cluster-wide split gate (`MaxClusterConcurrentAutoSplits`) is enabled; compute the cluster aggregate as a `sum` across the `tree` tag to size the ceiling and decide whether the gate is needed. |
| `orleans.lattice.split.candidates_suppressed` | `Counter<long>` | `{shard}` | Hot, admitted split candidates a monitor pass found but could not start, because the per-tree cap (`MaxConcurrentAutoSplits`) had fewer free slots than candidates or the cluster-wide gate (`MaxClusterConcurrentAutoSplits`) withheld a slot. Both are occupied by every migration `orleans.lattice.split.in_flight` counts, so an in-flight healing fold holds a slot an autonomic split would otherwise take. A pass that starts with the per-tree cap already fully occupied evaluates no candidates and records nothing here. Tagged `tree`. Emitted regardless of whether the gate is enabled; chronically non-zero across many trees while `split.in_flight` is high signals aggregate drain pressure the per-tree cap alone cannot see. |
| `orleans.lattice.split.admission.deferred` | `Counter<long>` | `{shard}` | Hot shards whose autonomic split was held back this pass, counted by the clause that held it back. Tagged `tree` and `reason`: `cluster_cap` (the **cluster-wide** admission gate had no headroom left under `MaxClusterConcurrentAutoSplits`, against which every tree's in-flight migrations count, healing folds included), `shard_ceiling` (the tree has no split headroom left under `MaxPhysicalShardsPerTree`), `uniform_load` (the shard is over the ops/sec threshold but its tree's load skew is below `HotShardMinSkewRatio`, so a split would relieve nothing - the signature of a bulk ingest rather than a hot spot), or `low_occupancy` (a hot, skewed shard holding fewer live entries than `HotShardMinShardEntries`, so a split would redistribute nothing). On `cluster_cap`, flat-zero means the ceiling never binds (pure overhead - raise it or leave the gate off), and sustained non-zero alongside rising hot-shard latency means the ceiling is set too low and is starving legitimate elasticity; sustained `shard_ceiling` means the per-tree ceiling is binding and should be reviewed alongside the tree's shard count. |

To derive ops/sec, compute the rate of `shard.reads + shard.writes` at the
collector; the same underlying counters back the internal hotness monitor that
drives autonomic splitting.

Previous: [Instrument catalog](instrument-catalog.md). Next: [Instrument catalog: Leaf-level (sourced from BPlusLeafGrain)](instrument-catalog-2.md). Contents: [Metrics](../metrics.md).
