Table of Contents

Diagnostics

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at diagnostics.md, and llms.txt lists every page.

ILattice.DiagnoseAsync returns a point-in-time health snapshot of a tree. It is an admin-rate API intended for dashboards, health probes, and post-mortem investigation - not for hot-path application logic.

var report = await tree.DiagnoseAsync(deep: true, cancellationToken);
Console.WriteLine($"Tree {report.TreeId}: {report.TotalLiveKeys} live / {report.TotalTombstones} tombstones across {report.ShardCount} shards.");
foreach (var shard in report.Shards)
{
    Console.WriteLine($"  shard {shard.ShardIndex}: depth={shard.Depth}, live={shard.LiveKeys}, ops/s={shard.OpsPerSecond:F1}");
}

Report shape

TreeDiagnosticReport aggregates per-shard structural and runtime metrics plus a bounded ring buffer of recent split events:

Field Description
TreeId Logical tree identifier.
ShardCount Number of physical shards currently owning virtual slots.
VirtualShardCount Size of the tree's virtual slot space: 4096 by default, or the slot count an installed app's manifest declared for a tree it created, while that tree keeps the map it was created with.
TotalLiveKeys Sum of live keys across all shards.
TotalTombstones Sum of tombstones across all shards (always 0 when deep: false).
Shards Per-shard reports, ordered by ShardIndex.
RecentSplits Most recent split commits, adaptive or driven by an online reshard that grows the tree, as RecentSplit entries (oldest first, capped at 32). Folds are not recorded.
SampledAt UTC timestamp when the report was assembled.
Deep Whether the report includes tombstone counts.

Each ShardDiagnosticReport carries structural, volume, and hotness fields:

Field Description
ShardIndex Zero-based physical shard index.
Depth B+ tree depth - 1 when the root is a leaf, 2 with one internal level, etc. 0 when the shard has no root yet.
RootIsLeaf Whether the shard's root is currently a leaf; false when the shard has no root yet.
LiveKeys / Tombstones Live and tombstoned entry counts. Tombstones is 0 unless deep: true.
TombstoneRatio Tombstones / (LiveKeys + Tombstones); 0 when the shard is empty or deep: false.
OpsPerSecond (Reads + Writes) / HotnessWindow.TotalSeconds since shard activation.
Reads / Writes Volatile counters; reset on shard-grain deactivation.
HotnessWindow Wall-clock duration over which Reads and Writes accumulated - the time since the shard activated. Exactly TimeSpan.Zero only on the placeholder entry of a shard whose fan-out failed (SampleFailed).
SplitInProgress Whether the shard is the source of an in-flight slot migration: a split (adaptive, or driven by an online reshard that grows the tree) or a fold that hands its slots to an adjacent shard (an online reshard that shrinks the tree, or automatic shard healing).
BulkOperationPending Whether a bulk-load graft is pending on this shard.
SampleFailed Whether the fan-out failed to sample this shard (the shard grain faulted or timed out). When true the entry is a placeholder: only ShardIndex is meaningful, and its zero counts mean not measured, not empty.

Shallow vs deep

The deep parameter controls whether tombstones are counted. Both modes fan out to every physical shard, and in both each shard walks its whole leaf chain - in work-bounded batches that each visit a bounded number of leaves and then release the shard, so a large shard answers over several calls instead of being held for the whole walk. Cost is therefore proportional to the number of leaves in either mode:

  • deep: false (default) - each leaf reports its live-key count only, so Tombstones and TombstoneRatio stay 0.
  • deep: true - each leaf also counts its tombstones, populating Tombstones and TombstoneRatio.

Use shallow reports for routine health probing; reach for deep reports when you suspect tombstone bloat or are diagnosing a compaction issue.

Caching

Repeat callers within LatticeOptions.DiagnosticsCacheTtl (default: 5 seconds) share the same snapshot - identical SampledAt timestamps are returned and no fan-out is performed. Shallow and deep reports are cached independently, so a shallow-then-deep sequence will always produce two distinct samples.

Recording a split commit - adaptive or reshard-driven - invalidates both cached reports, so the first DiagnoseAsync call after the record lands returns a fresh report rather than a stale pre-split view. The record is sent best-effort (see Recent splits); if it is lost, the cached report simply lives out its TTL. The tree-administration facade's BeginBulkLoadAsync also drops both cached reports before its deep emptiness probe, so that probe is always sampled fresh; see Bulk Loading.

Set DiagnosticsCacheTtl = TimeSpan.Zero to disable caching entirely - every call assembles a new report. See Configuration for details.

Recent splits

RecentSplits is a bounded (32-entry) ring buffer of the most recent split commits observed by the diagnostics grain - adaptive splits and the per-shard splits a growing online reshard drives alike. The folds that a shrinking reshard or automatic shard healing commits are not recorded; orleans.lattice.shard.consolidations_committed counts them. Each RecentSplit entry carries the source ShardIndex and the UTC commit time (AtUtc). The buffer lives in memory for the lifetime of the diagnostics grain's activation - it is not persisted, so it starts empty again if that activation is collected - and is useful for correlating shard-count changes with recent traffic bursts.

Splits are pushed to the diagnostics grain on a best-effort, fire-and-forget basis from the split-coordinator's commit path, so a split may occasionally be missing from the buffer if the push RPC fails - the split itself still completes and is reflected in ShardCount/Shards on the next full report.

Cancellation

DiagnoseAsync honours its CancellationToken cooperatively. A pre-cancelled token throws OperationCanceledException immediately. Cancellation during the fan-out is observed between a shard's bounded batches and once the fan-out completes; shard calls already in flight are not aborted.

Operational notes

  • DiagnoseAsync is safe to call at any time, including during an ongoing resize or reshard - per-shard reports whose fan-out fails are returned as empty entries rather than failing the whole report. Such an entry carries only its ShardIndex and SampleFailed = true, which is what tells it apart from a genuinely empty shard; its HotnessWindow also reads exactly TimeSpan.Zero, where a genuinely empty shard reports the time since it activated.
  • DiagnoseAsync is authorized as a whole-tree read. With the authorization add-on registered, the caller needs read access to the entire tree: a denial, or a grant scoped to a key prefix or single keys, throws LatticeAuthorizationDeniedException instead of returning a narrowed report, because per-shard counts would still disclose the keys they counted. With no authorization add-on registered every caller is allowed, at no cost. See Security posture.
  • Reserved system trees (ids starting with _lattice_) cannot be diagnosed through ILattice: the call throws LatticeReservedTreeNamespaceException.
  • The returned DTOs are readonly record struct types with [Immutable] serialization - they are cheap to copy and log.
  • Do not call DiagnoseAsync on a hot path. For routing or capacity decisions inside the data plane, use the dedicated public APIs (CountAsync, CountPerShardAsync) rather than decoding a diagnostics report.