Diagnostics
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at diagnostics.md, and llms.txt lists every page.ILattice.DiagnoseAsync returns a point-in-time health snapshot of a tree. It is an admin-rate API intended for dashboards, health probes, and post-mortem investigation - not for hot-path application logic.
var report = await tree.DiagnoseAsync(deep: true, cancellationToken);
Console.WriteLine($"Tree {report.TreeId}: {report.TotalLiveKeys} live / {report.TotalTombstones} tombstones across {report.ShardCount} shards.");
foreach (var shard in report.Shards)
{
Console.WriteLine($" shard {shard.ShardIndex}: depth={shard.Depth}, live={shard.LiveKeys}, ops/s={shard.OpsPerSecond:F1}");
}
Report shape
TreeDiagnosticReport aggregates per-shard structural and runtime metrics plus a bounded ring buffer of recent split events:
| Field | Description |
|---|---|
TreeId |
Logical tree identifier. |
ShardCount |
Number of physical shards currently owning virtual slots. |
VirtualShardCount |
Size of the tree's virtual slot space: 4096 by default, or the slot count an installed app's manifest declared for a tree it created, while that tree keeps the map it was created with. |
TotalLiveKeys |
Sum of live keys across all shards. |
TotalTombstones |
Sum of tombstones across all shards (always 0 when deep: false). |
Shards |
Per-shard reports, ordered by ShardIndex. |
RecentSplits |
Most recent split commits, adaptive or driven by an online reshard that grows the tree, as RecentSplit entries (oldest first, capped at 32). Folds are not recorded. |
SampledAt |
UTC timestamp when the report was assembled. |
Deep |
Whether the report includes tombstone counts. |
Each ShardDiagnosticReport carries structural, volume, and hotness fields:
| Field | Description |
|---|---|
ShardIndex |
Zero-based physical shard index. |
Depth |
B+ tree depth - 1 when the root is a leaf, 2 with one internal level, etc. 0 when the shard has no root yet. |
RootIsLeaf |
Whether the shard's root is currently a leaf; false when the shard has no root yet. |
LiveKeys / Tombstones |
Live and tombstoned entry counts. Tombstones is 0 unless deep: true. |
TombstoneRatio |
Tombstones / (LiveKeys + Tombstones); 0 when the shard is empty or deep: false. |
OpsPerSecond |
(Reads + Writes) / HotnessWindow.TotalSeconds since shard activation. |
Reads / Writes |
Volatile counters; reset on shard-grain deactivation. |
HotnessWindow |
Wall-clock duration over which Reads and Writes accumulated - the time since the shard activated. Exactly TimeSpan.Zero only on the placeholder entry of a shard whose fan-out failed (SampleFailed). |
SplitInProgress |
Whether the shard is the source of an in-flight slot migration: a split (adaptive, or driven by an online reshard that grows the tree) or a fold that hands its slots to an adjacent shard (an online reshard that shrinks the tree, or automatic shard healing). |
BulkOperationPending |
Whether a bulk-load graft is pending on this shard. |
SampleFailed |
Whether the fan-out failed to sample this shard (the shard grain faulted or timed out). When true the entry is a placeholder: only ShardIndex is meaningful, and its zero counts mean not measured, not empty. |
Shallow vs deep
The deep parameter controls whether tombstones are counted. Both modes fan out to every physical shard, and in both each shard walks its whole leaf chain - in work-bounded batches that each visit a bounded number of leaves and then release the shard, so a large shard answers over several calls instead of being held for the whole walk. Cost is therefore proportional to the number of leaves in either mode:
deep: false(default) - each leaf reports its live-key count only, soTombstonesandTombstoneRatiostay0.deep: true- each leaf also counts its tombstones, populatingTombstonesandTombstoneRatio.
Use shallow reports for routine health probing; reach for deep reports when you suspect tombstone bloat or are diagnosing a compaction issue.
Caching
Repeat callers within LatticeOptions.DiagnosticsCacheTtl (default: 5 seconds) share the same snapshot - identical SampledAt timestamps are returned and no fan-out is performed. Shallow and deep reports are cached independently, so a shallow-then-deep sequence will always produce two distinct samples.
Recording a split commit - adaptive or reshard-driven - invalidates both cached reports, so the first DiagnoseAsync call after the record lands returns a fresh report rather than a stale pre-split view. The record is sent best-effort (see Recent splits); if it is lost, the cached report simply lives out its TTL. The tree-administration facade's BeginBulkLoadAsync also drops both cached reports before its deep emptiness probe, so that probe is always sampled fresh; see Bulk Loading.
Set DiagnosticsCacheTtl = TimeSpan.Zero to disable caching entirely - every call assembles a new report. See Configuration for details.
Recent splits
RecentSplits is a bounded (32-entry) ring buffer of the most recent split commits observed by the diagnostics grain - adaptive splits and the per-shard splits a growing online reshard drives alike. The folds that a shrinking reshard or automatic shard healing commits are not recorded; orleans.lattice.shard.consolidations_committed counts them. Each RecentSplit entry carries the source ShardIndex and the UTC commit time (AtUtc). The buffer lives in memory for the lifetime of the diagnostics grain's activation - it is not persisted, so it starts empty again if that activation is collected - and is useful for correlating shard-count changes with recent traffic bursts.
Splits are pushed to the diagnostics grain on a best-effort, fire-and-forget basis from the split-coordinator's commit path, so a split may occasionally be missing from the buffer if the push RPC fails - the split itself still completes and is reflected in ShardCount/Shards on the next full report.
Cancellation
DiagnoseAsync honours its CancellationToken cooperatively. A pre-cancelled token throws OperationCanceledException immediately. Cancellation during the fan-out is observed between a shard's bounded batches and once the fan-out completes; shard calls already in flight are not aborted.
Operational notes
DiagnoseAsyncis safe to call at any time, including during an ongoing resize or reshard - per-shard reports whose fan-out fails are returned as empty entries rather than failing the whole report. Such an entry carries only itsShardIndexandSampleFailed = true, which is what tells it apart from a genuinely empty shard; itsHotnessWindowalso reads exactlyTimeSpan.Zero, where a genuinely empty shard reports the time since it activated.DiagnoseAsyncis authorized as a whole-tree read. With the authorization add-on registered, the caller needs read access to the entire tree: a denial, or a grant scoped to a key prefix or single keys, throwsLatticeAuthorizationDeniedExceptioninstead of returning a narrowed report, because per-shard counts would still disclose the keys they counted. With no authorization add-on registered every caller is allowed, at no cost. See Security posture.- Reserved system trees (ids starting with
_lattice_) cannot be diagnosed throughILattice: the call throwsLatticeReservedTreeNamespaceException. - The returned DTOs are
readonly record structtypes with[Immutable]serialization - they are cheap to copy and log. - Do not call
DiagnoseAsyncon a hot path. For routing or capacity decisions inside the data plane, use the dedicated public APIs (CountAsync,CountPerShardAsync) rather than decoding a diagnostics report.