Table of Contents

Instrument catalog: Process identity to Shard-level and tree registry

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at instrument-catalog-1.md, and llms.txt lists every page.

Part of Instrument catalog, in Metrics.

Process identity

Name Kind Unit Description
orleans.lattice.build.info ObservableGauge<long> {build} Always 1, tagged with the version and the full 40-character sha the running process was built from. This is the only instrument that reports a property of the process rather than of a feature, and it exists because "is the build under test the build that is actually running" cannot be answered from any feature metric: every other instrument in this catalog stays silent until its feature is exercised, so a missing series is ambiguous between "the image did not deploy" and "the code deployed but that path was never reached" - opposite conclusions drawn from byte-identical evidence. It is emitted unconditionally on every collection from process start, with no registry to populate and no work to wait for, so it has no empty state to prime: an absent series means the process is not running or is not exporting, and means nothing else. The value carries no information and is always 1; all of the content is in the tags, which is the conventional info-metric shape and keeps the series safe to join against in a query. The sha is sourced from the +<sha> build-metadata suffix the SDK appends to AssemblyInformationalVersionAttribute from SourceRevisionId, so it comes from the build, never from a runtime git call - a deployment container is distroless and has no git, which is precisely where the value matters most - and it identifies the image rather than the checkout the process happens to be running beside. Carries version and sha only: cardinality on an info gauge multiplies across every other series a reader joins it against. A placeholder sha would be worse than an absent one, because it reads as a working detector while identifying nothing, so BuildInfoMetricTests fails outright on an empty or sentinel value rather than tolerating it.

Shard-level and tree registry

Name Kind Unit Description
orleans.lattice.registry.call.duration Histogram<double> ms Service time of one tree-registry call, measured inside the registry singleton's grain body. Tagged operation = exists, get_entry, resolve, get_shard_map, get_entries, or get_all_tree_ids for the reads, and register or unregister for the two mutators the census also covers; get_all_tree_ids, register and unregister are not [AlwaysInterleave], so a short service time on those arms beside a low admitted count does not by itself mean calls went unserved. The registry is a cluster singleton every per-tree background service addresses, so a cold start fans a whole estate onto one activation and callers see only a response-deadline TimeoutException - the same observation whether the call was served slowly or never served at all, which have opposite remedies. Because only admitted calls are recorded here, reading the tail and the count against the caller-side timeout population separates them: a short tail with a count far below the offered load means the calls never reached the body (blocked upstream, in activation or in the turn queue, neither of which an [AlwaysInterleave] attribute can admit past), while a tail approaching the caller deadline with a matching count means they were admitted and the time went on the awaited hop to the backing _lattice_trees tree. Deliberately not derived from orleans-storage-read-latency: that instrument's state_name tag comes from a grain's [PersistentState] declaration and the registry grain declares none, so no state_name series for it can exist.
orleans.lattice.registry.call.in_flight Histogram<int> {call} Concurrent tree-registry calls in flight on the registry singleton at the moment a new call is admitted, tagged operation as above. The count is global across every operation arm, so it is the fan-in width, not evidence of whether a given member interleaves. The recorded value excludes the arriving call, so it is 0 on the first concurrent call and 1 on the second - the same convention as orleans.lattice.leaf.commit.in_flight. This is the fan-in width the singleton carries, which the duration histogram alone cannot supply: a long tail over a flat-zero width is a slow backing hop, whereas a long tail whose width climbs with estate size is the (trees x per-tree background services) scaling law saturating one activation.
orleans.lattice.registry.caller.duration Histogram<double> ms Duration of one ILatticeRegistry call as observed by its caller, from dispatch to response, timeout, or fault (issue #3088). Tagged method (every interface member, not only the reads the census wraps) and outcome = completed, timeout (a TimeoutException - the response deadline), or faulted. Recorded by the caller-side decorator every production caller obtains the registry through - not opt-in, yet charged to no other grain call, because it is not an outgoing grain call filter - for every registry call the process dispatches, whatever becomes of it. That is what registry.call.duration cannot do: it sees only calls the grain admitted, so an unreachable registry is silent there. Read the three states directly: fine is completed with a short tail, slow is completed with a long tail, unreachable is a non-zero timeout rate - and timeout on a member here with no admitted calls for it in registry.call.duration means the calls were never served. Replaces the Orleans Diagnostics: [... CurrentlyExecuting=...] timeout-log field, which Orleans stops emitting silo-wide under saturation and whose absence reads as an idle registry. Per-process; a timeout sample is pinned at the deadline, which is why it is its own arm. Platform-tenant series.
orleans.lattice.registry.admission.wait Histogram<double> ms Time one tree id spent waiting for a fan-in permit before its batched registry read was dispatched, measured caller-side in RegistryFanInGate. Recorded unconditionally on every dequeue, including a zero wait, so the population is every gated read rather than only the slow ones. This exists because bounding fan-in does not remove work, it queues it: once concurrent reads are capped, a saturated registry stops producing registry-side symptoms and starts producing caller-side waiting instead. Every other registry signal - both instruments above, and orleans_app_requests_timedout_total{grain_type="latticeregistry"} - is scoped to the registry grain, so a stall that relocates into the admission queue drives all of them toward zero and reads as a clean recovery. This is the only instrument that observes the relocated wait, which makes it the one signal that distinguishes "the bound is working" from "the bound is now the constraint": a rising upper percentile here without a matching rise in orleans.lattice.registry.call.duration means the gate, not the registry, is what callers are waiting on. Admission is deliberately uncapped and untimed - shedding a cold-start options resolution would turn latency into a correctness failure, since the caller cannot proceed without its entry - so unbounded waiting is a real possible state and this histogram is the only thing that would show it.
orleans.lattice.registry.admission.in_flight Histogram<int> {dispatch} Fan-in permits held by this silo's registry fan-in gate at the moment it dispatches a gated round trip - the width this silo's share of the gate's bound actually caps. The bound is 16 concurrent reads across the whole cluster, divided by the live silo count and floored at one per silo, so a single silo may hold all 16 and each of four silos at most 4. It is not a duplicate of orleans.lattice.registry.call.in_flight, and confusing the two is the specific error it exists to prevent: that instrument counts calls executing inside the registry singleton's body, summed over every caller in the cluster and including callers that never pass through a gate at all, so it has a different population and a different ceiling. Comparing it against that bound compares two quantities that were never the same number, and a low reading there is not evidence that the bound has headroom. The recorded value includes the dispatch being recorded, so it runs from 1 to this silo's share and a recorded value equal to the share means the ceiling was reached. That deliberately inverts the exclude-the-arrival convention of orleans.lattice.registry.call.in_flight and orleans.lattice.leaf.commit.in_flight: under that convention a fully saturated gate would top out one below its own bound, so saturation would be indistinguishable from headroom by inspection - an off-by-one that makes the bound's binding state unobservable, which is precisely the failure this instrument removes. Read the distribution, never the window mean: permit occupancy is bursty, so a mean is dominated by idle time and measures something other than the peak. The actionable reading is the fraction of dispatches at the ceiling: on a single silo, whose share is the whole 16, that is _count minus _bucket{le="15"}, against _count (buckets are cumulative, so _bucket{le="16"} holds every dispatch and always equals _count), and on a larger cluster the same reading taken one below that silo's share. That needs a bucket boundary one below the ceiling. The instrument declares no bucket advice, so under the OpenTelemetry SDK's default boundaries (0, 5, 10, 25, ...) every width from 11 to 16 lands in the one le="25" bucket; configure an explicit-bucket view for it before reading the ceiling share.
orleans.lattice.registry.admission.batch.size Histogram<int> {tree} Distinct tree ids carried by one gated registry round trip, recorded once per dispatch. A value of 1 is the single-key registry read a silo issues when nothing is queued behind the bound - byte-for-byte the traffic it had before the gate existed - and 2 or more is a batched multi-key registry read, so the share above one is the share of reads the bound actually coalesced. That share is the only direct evidence the batching half of the gate ran at all; a population sitting almost entirely at 1 is a statement about the offered load, not about the batching. Read beside orleans.lattice.registry.admission.in_flight: a batch larger than one forms from ids already queued when a dispatch is taken, which happens chiefly once the permits are saturated but also when arrivals on concurrent threads enqueue before the first of them dispatches, so a batch-size distribution above one does not on its own show the bound binding - a width distribution reaching this silo's share does.
orleans.lattice.registry.admission.queue.depth Histogram<int> {tree} Distinct tree ids already waiting for a fan-in permit when another arrives, counting the arrival, recorded at enqueue. This is the offered fan-in, and it is the signal that separates "the bound had room" from "nothing asked for it". Admission dispatches synchronously on the arriving thread whenever a permit is free, so when the gate is not binding every other gate instrument reports its structural floor - a wait of microseconds, a width of 1, a batch of 1 - and those floors are visually identical to a comfortable bound. They are not a weak measurement of headroom; they are what absent demand looks like. A depth that stays at 1 says the demand never arrived, and until it rises above 1 no reading from the other three is evidence about the bound at all. This is the instrument to check first when a fan-in experiment comes back green.
orleans.lattice.shard.reads Counter<long> {op} Read operations served by a shard root (GetAsync, ExistsAsync, scan, count, etc.). One increment per operation, not per record returned. Tagged tree and shard.
orleans.lattice.shard.writes Counter<long> {op} Write operations served by a shard root (SetAsync, DeleteAsync, MergeManyAsync, BulkLoadAsync, etc.). One increment per operation: a batched or bulk call (SetManyAsync, MergeManyAsync, DeleteRangeAsync, SetManyWherePredicateAsync, BulkLoadAsync, BulkLoadRawAsync, BulkAppendAsync) counts once regardless of entry count, so a 5000-record import advances this by only the number of bulk operations. Use orleans.lattice.shard.records_written for the record rate. Tagged tree and shard.
orleans.lattice.shard.records_written Counter<long> {record} Individual records written by a shard root - the per-record companion to orleans.lattice.shard.writes. Incremented by 1 on a single-key write, by the batch size on SetManyAsync / MergeManyAsync / bulk load, and by the affected count on DeleteRangeAsync / SetManyWherePredicateAsync (published once the operation succeeds). Plot both: their ratio is the effective batch size. Tagged tree and shard.
orleans.lattice.shard.digest_reads Counter<long> {op} Projection-digest reads served by a shard root - one increment per whole-shard digest read (ILattice.GetLeafProjectionDigestAsync) and one per key-range digest read (ILattice.GetLeafProjectionDigestForRangeAsync), each served by the one shard the call names. A whole-tree poll of ILattice.GetLeafProjectionDigestAsync produces exactly shardCount increments; beyond the key-range reads, a higher rate signals a regression that fell back to walking every leaf. Tagged tree and shard.
orleans.lattice.shard.splits_committed Counter<long> {split} Adaptive shard-split commits - fired once per successful ShardMap swap, when the split coordinator finalises. Tagged tree and shard (the source shard that split).
orleans.lattice.shard.consolidations_committed Counter<long> {consolidation} Online shard-consolidation commits - fired once per fold that retires a donor shard from the ShardMap, whether automatic healing or a shrinking ILattice.ReshardAsync started the fold. It is recorded at the fold's final step, and only after that step's terminal state write succeeds, so an increment is always a durably committed fold rather than an attempt. The same step releases the donor's storage, except where the routing map cannot be read or still routes a slot to the donor, or the donor refuses retirement: the fold then commits as a routing-only retirement that leaves the donor's storage in place, and still counts here. The exact inverse of orleans.lattice.shard.splits_committed; plot both, because a sustained gap between them is a tree whose physical shard count is still climbing. Tagged tree and shard (the retired donor).
orleans.lattice.shard.healing.backlog Histogram<int> {shard} Physical shards a tree carries above its configured base shard count, sampled once per ShardHealingOrchestratorGrain sweep - the automatic over-split healing work outstanding. Tagged tree. A tree is healed exactly when this reaches zero, so "trees healed" is read straight off this instrument (count of series at zero) rather than from a second counter that could disagree with it. Pair it with orleans.lattice.shard.consolidations_committed for the reclaim rate: backlog answers how much damage is left, that counter answers whether it is going down.
orleans.lattice.shard.healing.decisions Counter<long> {decision} Automatic over-split healing sweeps by outcome - exactly one increment per tree per sweep. Tagged tree and decision (admitted, disabled, admission_closed, not_over_split, skewed_load, split_in_flight, tree_maintenance, cooldown, backpressure, at_capacity, no_foldable_pair; a not_observed arm is also mapped, for the persisted default before any sweep, but a sweep never publishes it). The series currently advancing for a tree is that tree's current healing decision, which is what distinguishes a tree that needs no healing (not_over_split) from one that needs healing and is being held back (skewed_load, backpressure, cooldown) from one where the mechanism is switched off (disabled, admission_closed).
orleans.lattice.split.retroactive_forward.entries Counter<long> {entry} Pending prepared mutations retroactively shadow-forwarded from a source shard's leaf chain into the destination shard's _pendingTx buckets at the start of an adaptive split's BeginShadowWrite phase. Tagged tree and shard (source).
orleans.lattice.split.retroactive_forward.duration Histogram<double> ms Wall-clock duration of the retroactive prepared-mutation sweep before the split coordinator transitions to the Drain phase. Tagged tree and shard (source).
orleans.lattice.split.in_flight Histogram<long> {split} Per-tree count of shard migrations in flight: shards that are the source of an unfinished split or the donor of a consolidation fold, both of which carry the same durable per-shard migration record, so an automatic-healing fold's donor is counted alongside the adaptive splits. Sampled once per hot-shard monitor pass that reaches its shard poll - a pass records nothing while the tree's AutoSplitEnabled is off, before the tree reaches AutoSplitMinTreeAge, or while a resize, reshard, merge or snapshot of the tree is in flight, so the splits and folds a reshard drives are not sampled while it runs. Every migration counted here occupies one of the tree's MaxConcurrentAutoSplits slots - a healing fold in flight leaves one fewer for an autonomic split - and the same count, plus the splits the pass triggers, is the footprint the tree reports to the cluster-wide split gate and that ILatticeAdmin.GetSplitActivityAsync reads. Tagged tree. Emitted on every such pass regardless of whether the cluster-wide split gate (MaxClusterConcurrentAutoSplits) is enabled; compute the cluster aggregate as a sum across the tree tag to size the ceiling and decide whether the gate is needed.
orleans.lattice.split.candidates_suppressed Counter<long> {shard} Hot, admitted split candidates a monitor pass found but could not start, because the per-tree cap (MaxConcurrentAutoSplits) had fewer free slots than candidates or the cluster-wide gate (MaxClusterConcurrentAutoSplits) withheld a slot. Both are occupied by every migration orleans.lattice.split.in_flight counts, so an in-flight healing fold holds a slot an autonomic split would otherwise take. A pass that starts with the per-tree cap already fully occupied evaluates no candidates and records nothing here. Tagged tree. Emitted regardless of whether the gate is enabled; chronically non-zero across many trees while split.in_flight is high signals aggregate drain pressure the per-tree cap alone cannot see.
orleans.lattice.split.admission.deferred Counter<long> {shard} Hot shards whose autonomic split was held back this pass, counted by the clause that held it back. Tagged tree and reason: cluster_cap (the cluster-wide admission gate had no headroom left under MaxClusterConcurrentAutoSplits, against which every tree's in-flight migrations count, healing folds included), shard_ceiling (the tree has no split headroom left under MaxPhysicalShardsPerTree), uniform_load (the shard is over the ops/sec threshold but its tree's load skew is below HotShardMinSkewRatio, so a split would relieve nothing - the signature of a bulk ingest rather than a hot spot), or low_occupancy (a hot, skewed shard holding fewer live entries than HotShardMinShardEntries, so a split would redistribute nothing). On cluster_cap, flat-zero means the ceiling never binds (pure overhead - raise it or leave the gate off), and sustained non-zero alongside rising hot-shard latency means the ceiling is set too low and is starving legitimate elasticity; sustained shard_ceiling means the per-tree ceiling is binding and should be reviewed alongside the tree's shard count.

To derive ops/sec, compute the rate of shard.reads + shard.writes at the collector; the same underlying counters back the internal hotness monitor that drives autonomic splitting.