Table of Contents

Troubleshooting

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at troubleshooting.md, and llms.txt lists every page.

A symptom-driven guide to the problems Lattice trees actually hit in production: storage-provider failures on write, slowdowns caused by an in-flight shard split, scans that take far longer than expected, and reads that look stale. Each section is laid out as symptom -> likely cause -> how to confirm -> how to fix.

The centre of the guide is Reading a DiagnoseAsync report. Nearly every investigation starts by taking a report and following the field that looks wrong into one of the sections below.

This guide deliberately does not restate the reference material it points at. Field-by-field definitions live in Diagnostics, the instrument catalog lives in Metrics, and every option named here is described in Configuration.


Take a diagnostic report first

ILattice.DiagnoseAsync returns a TreeDiagnosticReport: a per-tree health snapshot assembled by fanning out to every physical shard.

var report = await tree.DiagnoseAsync(deep: true, cancellationToken);

Console.WriteLine(
    $"tree={report.TreeId} shards={report.ShardCount}/{report.VirtualShardCount} " +
    $"live={report.TotalLiveKeys} tombstones={report.TotalTombstones} " +
    $"deep={report.Deep} sampledAt={report.SampledAt:O}");

foreach (var shard in report.Shards)
{
    Console.WriteLine(
        $"  shard {shard.ShardIndex}: depth={shard.Depth} rootIsLeaf={shard.RootIsLeaf} " +
        $"live={shard.LiveKeys} tombstones={shard.Tombstones} " +
        $"ratio={shard.TombstoneRatio:P1} ops/s={shard.OpsPerSecond:F1} " +
        $"reads={shard.Reads} writes={shard.Writes} window={shard.HotnessWindow} " +
        $"split={shard.SplitInProgress} bulk={shard.BulkOperationPending}");
}

foreach (var split in report.RecentSplits)
{
    Console.WriteLine($"  recent split: shard {split.ShardIndex} at {split.AtUtc:O}");
}

deep: true is the mode that produces tombstone counts; see the traps below before you draw a conclusion from a shallow report. Neither mode is asymptotically cheaper: both walk every shard's whole leaf chain, and deep: true only reads each leaf's tombstone count alongside its live count (see Diagnostics). What makes a dashboard poll affordable is the per-mode report cache described in the traps.

On a cluster that registers the authorization add-on, DiagnoseAsync - and GetStorageUsageAsync, used further down - is authorized as a read of the whole tree. Run the investigation as a caller whose grant covers the whole tree: a grant scoped to a key prefix or single keys is refused with LatticeAuthorizationDeniedException rather than narrowed, because the report's per-shard counts would still disclose the keys they counted. See Diagnostics.


Reading a DiagnoseAsync report

Tree-level fields

Field What a healthy value looks like What an unhealthy value points at
ShardCount The tree's physical shard count. It rises by one with each committed split - adaptive, or driven by an online reshard that grows the tree - and falls by one with each committed fold, which an online reshard that shrinks the tree drives, and automatic shard healing drives when it folds an over-split tree back towards its registry-pinned base shard count. A value you did not expect means you are looking at a different tree than you think, or splits, folds or healing changed the topology: check RecentSplits and orleans.lattice.shard.splits_committed for splits, and orleans.lattice.shard.consolidations_committed for folds, which RecentSplits never records. See Shard splitting, Online reshard, and ShardHealingEnabled.
VirtualShardCount 4096 by default - a compile-time constant - or the slot count an installed app's manifest declared for a tree it created (see Virtual shard space). It is not a runtime option: persisted shard maps reference virtual slots by index. Nothing to act on - it is the routing map's resolution, not a load signal. See Tree sizing.
TotalLiveKeys Tracks your expected working-set size. Growth that outruns your model points at Slow scans and admission headroom. Compare against LatticeOptions.MaxLiveKeys if you have set one.
TotalTombstones Small relative to TotalLiveKeys. A large share of the total points at tombstone bloat. Only populated when deep: true.
SampledAt Within DiagnosticsCacheTtl (default 5 s) of now. Older than the TTL means you are reading a cached report; see the traps.
Deep Echoes the argument you passed. If it is false, ignore every tombstone field in the report.
Shards One entry per physical shard, ordered by ShardIndex. See the per-shard table below.
RecentSplits Empty on a stable tree; a short list after adaptive splits or an online reshard that grows the tree. A fold is never recorded here. A steady stream of entries points at Concurrent split activity.

Per-shard fields

Field What a healthy value looks like What an unhealthy value points at
Depth 0 for a shard with no root yet, 1 while the root is still a leaf, and small (single digits) once the shard has internal levels. Depth climbing across shards means the shards hold far more keys than the tree was sized for. See Tree sizing and Tree structure.
RootIsLeaf true while the shard's root is still a single leaf, false once the shard has grown past one leaf. A shard with no root yet (Depth = 0) also reports false. Not a fault on its own; read it alongside Depth.
LiveKeys In proportion to the virtual slots each shard owns: roughly even on a tree that has never split. A split moves half of the source shard's slots to a new shard, so each of the two then holds about half the keys of an unsplit peer. Keys are placed by hashing the whole key into a virtual slot, so key design cannot concentrate keys on one shard. Uneven counts that track slot ownership are the footprint of past splits and healing, not a fault; see Shard splitting.
Tombstones / TombstoneRatio Low. 0.0 on a tree that never deletes. A high ratio means deleted rows are still being walked on every scan. See Slow scans and Tombstone compaction.
OpsPerSecond Comparable across shards. One shard far above its peers is a hot shard - the exact condition adaptive splitting exists to relieve. See Concurrent split activity.
Reads / Writes The raw counters OpsPerSecond is derived from, over HotnessWindow: (Reads + Writes) / HotnessWindow.TotalSeconds. A read-heavy shard and a write-heavy shard need different remedies; the split between the two counters tells you which you have.
HotnessWindow The window the counters cover: the time since the shard activated, so it restarts when the shard deactivates. Positive on every shard that answered the fan-out. Exactly TimeSpan.Zero marks the placeholder entry of a shard whose diagnostics call failed (OpsPerSecond then reads 0.0); see the traps. A very short window means the shard activated recently, so its counters and OpsPerSecond cover only that span.
SplitInProgress false on a stable tree. true means this shard is the source of an in-flight slot migration: a split (adaptive, or driven by an online reshard that grows the tree), or a fold that hands this shard's slots to an adjacent shard (an online reshard that shrinks the tree, or automatic shard healing). Normal if transient; see Concurrent split activity.
BulkOperationPending false on a stable tree. true means the shard has recorded a bulk-load graft it has not finished linking in: normal while a chunk of an append-based bulk load grafts, and cleared by the shard's next read or write if the graft was interrupted. While any shard reports it, autonomic splitting is suspended for the whole tree. See Bulk loading.
SampleFailed false on every shard. true marks the placeholder entry of a shard whose diagnostics call failed: its counts are unmeasured, not zero, and the tree-administration facade's streamed bulk load (ILatticeTreeAdmin.BeginBulkLoadAsync) refuses to begin until every shard answers. See the traps.

A worked reading

A tree with a hot shard mid-split, and a second shard that needs compaction:

tree=orders shards=4/4096 live=812433 tombstones=196022 deep=True sampledAt=2025-...
  shard 0: depth=3 live=201110 tombstones=1204  ratio=0.6%  ops/s=48.2  split=False bulk=False
  shard 1: depth=3 live=198740 tombstones=1190  ratio=0.6%  ops/s=51.7  split=False bulk=False
  shard 2: depth=4 live=210301 tombstones=192455 ratio=47.8% ops/s=44.9  split=False bulk=False
  shard 3: depth=3 live=202282 tombstones=1173  ratio=0.6%  ops/s=502.4 split=True  bulk=False

Read it in this order:

  1. LiveKeys is even across every shard. Key distribution is fine, so whatever is wrong is not a hashing problem.
  2. Shard 3 carries roughly ten times its peers' OpsPerSecond. That is a hot shard: even load by key count, very uneven load by request rate. Its SplitInProgress is true, so the autonomic splitter has already noticed and is acting. The split is not in RecentSplits yet: a split is recorded there only when it finalises, after the flag clears. Nothing to do unless the flag stays set - go to Concurrent split activity.
  3. Shard 2 has a TombstoneRatio of 47.8 percent while its peers sit under one percent. Nearly half of what a scan of that shard walks is deleted rows. That is the compaction signal - go to Slow scans.
  4. Shard 2's Depth is one greater than its peers', which is consistent with the tombstone bloat rather than a separate problem: the shard is holding more physical entries than its peers for the same live-key count.

Traps when reading a report

These are the report behaviours that most often lead to a wrong conclusion.

  • A shallow report always reports zero tombstones. The shallow path never asks leaves for tombstone counts, so Tombstones and TombstoneRatio come back 0 (and TotalTombstones with them) regardless of what is on disk. Check report.Deep before believing a zero. Only deep: true produces tombstone figures.

  • An all-zero shard entry can mean the fan-out failed. When a shard's diagnostics call throws, the aggregator logs a warning (Diagnostics fan-out failed for shard {ShardIndex} in tree {TreeId}) and substitutes an entry carrying only the shard index and SampleFailed = true; every other field is its default. A genuinely empty shard also reports Depth = 0 and LiveKeys = 0, but its SampleFailed is false and its HotnessWindow is positive, because a shard that answers always reports the time since it activated. An entry with SampleFailed set is therefore the placeholder for a shard that did not answer, and the silo log carries the exception. If one shard reads as empty on a tree you know holds data, check SampleFailed and the log before concluding the data is gone.

  • Reports are cached per mode. Shallow and deep results are cached independently for DiagnosticsCacheTtl (default 5 s), so a shallow poll never refreshes the deep report and vice versa. SampledAt tells you how old the report actually is. To make a report unconditionally fresh during an investigation, set the TTL to zero for that tree:

    siloBuilder.ConfigureLattice("orders", options =>
    {
        options.DiagnosticsCacheTtl = TimeSpan.Zero;
    });
    

    A zero TTL means every call fans out to every shard. Use it for triage, not for a steady-state dashboard poll.

  • RecentSplits is activation state, not history. It is a bounded ring buffer (32 entries) on the per-tree stats grain. It is emptied if that grain deactivates, and it is trimmed once 32 splits have accumulated. An empty RecentSplits does not prove no split happened; the orleans.lattice.shard.splits_committed counter in Metrics is the record that survives a deactivation.

  • A split commit invalidates both cached reports. Recording a split clears the shallow and deep caches, so the first report after the split is recorded is freshly fanned out. The split coordinator sends that record best-effort and does not wait for it, so if it is lost the cached report simply lives out its TTL (see Diagnostics).

What the report does not tell you

DiagnoseAsync is a structural and hotness snapshot. It carries no storage bytes, no latency percentiles, and no cache hit ratio. For those:

Question Where to look
How many bytes is this tree holding? ILattice.GetStorageUsageAsync; see Tree storage.
Is orleans.lattice.storage.leaf_state_bytes / storage.snapshot_bytes really zero, or just unmeasured? orleans.lattice.storage.usage_deep_published: no series means the tree is unobserved, 0 means only the WAL-only poller has run so those gauges report no data, 1 means a deep publish ran and a zero is a real zero. See Tree storage.
How slow are reads and writes? The orleans.lattice.get.duration / set.duration histograms in Metrics.
Is the read cache helping? orleans.lattice.cache.hits and orleans.lattice.cache.misses.
Is the WAL backing up? WAL saturation signal and WAL tuning.
Is compaction keeping up? orleans.lattice.compaction.*; see Tombstone compaction.

Storage-provider exceptions on write

Symptom

A write fails with an exception thrown by a storage provider rather than by Lattice: the WAL provider's append when the mutation's WAL row is too large, or the Orleans grain-storage provider's WriteStateAsync when a grain-state row is. The failure is usually reproducible for a particular key or leaf and unaffected by retry.

Likely cause

Lattice persists a tree across several storage rows, and each one is bounded independently by the provider's per-row or per-blob limit:

  • The WAL row, written once per mutation. Its size grows with the key and value you wrote, plus causal metadata. On the Azure Table WAL provider the whole encoded entry is stored in one binary property, which Azure Table caps at 64 KiB.
  • The leaf snapshot blob, capturing every entry a leaf holds - tombstones included, until compaction reaps them. Its size grows with the entries the leaf holds, which are bounded by that tree's pinned MaxLeafKeys and by LatticeOptions.MaxLeafBytes (default 64 MiB). A payload larger than LeafSnapshotSegmentBytes (default 4 MiB) is written as row-aligned segments of at most that size, so a single snapshot row exceeds the segment size only for an indivisible entry larger than it.
  • The leaf state row, which no longer scales with MaxLeafKeys.

An oversized single value pushes the WAL row over the limit; a snapshot segment larger than the provider's per-row limit pushes a snapshot row over it. Tree storage carries the per-provider limit table, the row-size formulas, and the sizing arithmetic for choosing MaxLeafKeys against a given provider.

Two things this is not:

  • It is not LatticeQuotaExceededException. That is Lattice's own opt-in admission control (LatticeOptions.MaxLiveKeys / MaxEstimatedBytes, both null by default) refusing a write because the tree hit a configured ceiling, and it carries TreeId, Dimension, Current, and Limit so you can act on it without parsing a message.
  • It is not Orleans's InconsistentStateException, which signals an etag conflict - a concurrent writer to the same state row - not a size problem. An atomic write reports that conflict as a LatticeStateWriteFailedException with Conflict set instead (see below), and is safe to retry with the same operation id.

How to confirm

  1. Read the provider exception itself. It names the limit it enforced. On the ordinary write path Lattice does not translate a provider size failure into a Lattice exception type, so the provider's own error is the primary evidence. An atomic write is the exception: when the atomic-write saga or the cross-tree coordinator fails to persist its own state with a fault raised by a storage provider, the caller receives LatticeStateWriteFailedException instead, because the provider's exception type need not be loadable on the client (a fault of a type every client can load, such as TimeoutException, propagates unchanged). Its message and FaultType summarise the provider fault, GrainType and GrainKey name the grain whose write failed, and Conflict is false for a failure that is not an optimistic-concurrency conflict. See Atomic writes.

  2. Take a storage-usage report and compare the surfaces against the provider's limit from Tree storage:

    var usage = await tree.GetStorageUsageAsync(cancellationToken);
    
    Console.WriteLine(
        $"tree={usage.TreeId} wal={usage.WalRetainedBytes} snapshot={usage.SnapshotBytes} " +
        $"leafState={usage.LeafStateBytes} total={usage.TotalBytes} partial={usage.Partial}");
    

    Partial set to true means at least one surface could not be sampled, so the totals are a floor rather than an exact figure.

  3. Check the silo log for the leaf snapshot warning. A snapshot capture driven by the activation advisory is best-effort: it is caught and logged (Proactive snapshot capture for leaf {GrainId} failed; will retry on next periodic recheck or reactivation.) rather than surfaced to the caller. A snapshot that is permanently too large to write therefore shows up as a repeating warning and no snapshot coverage, not as a failed request.

How to fix

  • Bound what callers can write, at the edge. Both guards are opt-in and both default to null; on SetAsync, SetIfVersionAsync, GetOrSetAsync, SetManyAsync, SetManyWherePredicateAsync, the single-tree and cross-tree atomic batches (checked before anything is staged) and the CRDT delta paths each throws an ArgumentException before the write reaches storage, which is a far better failure than a provider error deep in the persistence path. The bulk-load paths and MergeAsync do not check them:

    siloBuilder.ConfigureLattice("orders", options =>
    {
        options.MaxKeyLength = 512;
        options.MaxValueSizeBytes = 256 * 1024;
    });
    
  • Keep large payloads out of the tree. Store the blob in blob storage and put its identifier in the tree. This is the only fix that scales - a value large enough to threaten a WAL row will threaten the next provider too.

  • Lower the tree's MaxLeafKeys so each snapshot blob covers fewer entries. MaxLeafKeys is not a LatticeOptions property: it is pinned per tree in the registry (defaulting to 128) and changed only through ILattice.ResizeAsync. See Tree sizing for the resize procedure and Tree storage for how to pick the value.

  • Lower LeafSnapshotSegmentBytes (default 4 MiB, floor 64 KiB) below the grain-storage provider's per-row limit, so every snapshot row fits whatever MaxLeafKeys is. Set it on the silo-wide options: the snapshot storage grain does not see a per-tree override. See Tree storage.

  • Move the WAL to a higher-capacity backend. The WAL has its own storage seam (IWalStorageProvider), so it can be pointed at a backend with a larger per-row limit independently of the grain-storage provider the rest of the tree uses; see WAL storage providers. Note this does not help the snapshot blob: leaf snapshots are persisted through the same named grain-storage provider as the rest of the tree, so the levers there are LeafSnapshotSegmentBytes, MaxLeafKeys, MaxLeafBytes, the value sizes you write, and swapping that provider.

  • Admission control caps total growth, not row size. MaxLiveKeys and MaxEstimatedBytes make the tree refuse the write calls that check them - SetAsync, SetIfVersionAsync, GetOrSetAsync, SetManyAsync, SetManyWherePredicateAsync, the single-tree and cross-tree atomic batches, ApplyCrdtDeltaAsync and ApplyCrdtDeltaManyAsync - with a typed, actionable LatticeQuotaExceededException once the whole tree reaches a ceiling, and the non-enforcing AdmissionAdvisoryLiveKeys / AdmissionAdvisoryBytes dry-run ceilings help you size them; see Configuration. They bound the tree's total footprint, not any single row, and the bulk-load paths and MergeAsync do not check them, so they complement the fixes above rather than replace them.


Concurrent split activity

Symptom

One or more shards report SplitInProgress = true. Latency on the affected shard is elevated, and RecentSplits shows entries appearing regularly.

Likely cause

SplitInProgress means the shard is the source of an in-flight adaptive split: Lattice has detected a hot shard and is moving part of its virtual-shard range to a new physical shard. An online reshard (ILattice.ReshardAsync) that grows the tree drives the same per-shard split, so while one runs the flag and RecentSplits also move on the shards it divides, hot or not. A fold - an online reshard that shrinks the tree, or automatic shard healing - sets the same flag on the shard whose slots it is handing to an adjacent shard, but is never recorded in RecentSplits; orleans.lattice.shard.consolidations_committed counts folds. The autonomic splitter watches per-shard throughput and triggers when a shard's observed operations per second exceed HotShardOpsPerSecondThreshold (default 200). That figure is computed as (reads + writes) / window.TotalSeconds - the same quantity the report surfaces as OpsPerSecond, so the report shows you exactly what the splitter is reacting to. AutoSplitEnabled defaults to true.

BulkOperationPending is the analogous flag for a pending bulk graft. An append-based bulk load records each chunk's graft on the shard before linking it in and clears it once the graft completes, so the flag is normally set only while a chunk is grafting; an interrupted graft is finished by the next read or write that reaches the shard.

A split in progress is normal operation, not a fault. Callers do not see topology exceptions during one: routing staleness is caught and retried inside Lattice, so a read or write that races the topology change is retried against the new owner transparently. See Shard splitting for the phase-by-phase description of what the split actually does.

How to confirm

Poll the flags and correlate them against split history:

var report = await tree.DiagnoseAsync(deep: false, cancellationToken);

foreach (var shard in report.Shards.Where(s => s.SplitInProgress || s.BulkOperationPending))
{
    Console.WriteLine(
        $"shard {shard.ShardIndex}: split={shard.SplitInProgress} " +
        $"bulk={shard.BulkOperationPending} ops/s={shard.OpsPerSecond:F1}");
}

var totalOps = report.Shards.Sum(s => s.OpsPerSecond);
var hottest = report.Shards.OrderByDescending(s => s.OpsPerSecond).First();
var share = totalOps > 0 ? hottest.OpsPerSecond / totalOps : 0;

Console.WriteLine($"hottest shard {hottest.ShardIndex} carries {share:P0} of observed ops/s");

Then decide which case you are in:

Observation Reading
Flag set, clears within a few report intervals, RecentSplits gains one entry Normal. The split committed.
Flag set, clears, ShardCount one lower, RecentSplits unchanged Normal. A fold committed; orleans.lattice.shard.consolidations_committed counts it.
Flag set on several shards at once Also normal if it is bounded: MaxConcurrentAutoSplits (default 2) caps in-flight autonomic splits per tree, MaxClusterConcurrentAutoSplits (default null, disabled) adds a cluster-wide ceiling on top of it, MaxConcurrentMigrations (default 4) bounds the splits or folds an online reshard runs at once, and MaxConcurrentShardConsolidations (default 1) bounds the folds automatic shard healing runs on a tree.
Flag set on the same shard across many reports, no new RecentSplits entry, no throughput recovery Stuck. Treat as a fault.
Flag never set even though one shard is obviously hot The candidate is being suppressed.

For the last two cases, the metrics tell you which: orleans.lattice.split.in_flight shows how many of the tree's shards carry the flag on each monitor pass that polls the shards (a pass held off by AutoSplitEnabled, by the minimum tree age, or by an in-flight resize, reshard, merge or snapshot records nothing) - the donor of an in-flight healing fold included, so a fold occupies one of the MaxConcurrentAutoSplits slots too; orleans.lattice.split.candidates_suppressed counts hot, eligible shards a monitor pass found but could not start, because the per-tree cap (MaxConcurrentAutoSplits) had fewer free slots than candidates or the cluster-wide gate (MaxClusterConcurrentAutoSplits) withheld a slot; orleans.lattice.split.admission.deferred counts hot shards the admission policy held back, by reason: cluster_cap (the cluster-wide gate), uniform_load (the whole tree is hot - typically a bulk ingest - so a split would relieve nothing; governed by HotShardMinSkewRatio), low_occupancy (too few live entries to redistribute; HotShardMinShardEntries), or shard_ceiling (the tree's physical shard ceiling; MaxPhysicalShardsPerTree); and orleans.lattice.shard.splits_committed counts the splits that completed. A pass that starts with the per-tree cap already fully occupied evaluates no candidates at all, so while split.in_flight sits at MaxConcurrentAutoSplits a hot shard waiting its turn is counted on neither counter. Nor is a shard held off by the minimum tree age, by the per-shard cooldown, or because it owns fewer than two virtual slots, or any shard while a pass is suspended outright because a resize, reshard, merge or snapshot is in flight or a shard holds a pending bulk graft. See Metrics.

How to fix

  • A transient flag needs no action. Let it complete.
  • Splits never triggering on an obviously hot shard: check the suppression rules in Shard splitting. The common causes are AutoSplitMinTreeAge (default 60 s) holding off splits on a young tree, HotShardSplitCooldown (default 2 minutes) rate-limiting repeat splits on the same shard, and AutoSplitEnabled having been turned off. The admission policy can also hold a hot shard back - read the reason on orleans.lattice.split.admission.deferred: uniform_load means the whole tree is hot, so splitting would not help.
  • Splits triggering too eagerly on a bursty workload: raise HotShardOpsPerSecondThreshold. The rate the splitter compares against it is averaged over the shard's whole activation (the same HotnessWindow the report shows), not over one sampling interval, so a short burst barely moves it on a long-lived shard but dominates it on a recently activated one. Lengthening HotShardSampleInterval (default 30 s) only makes the monitor poll less often; it does not change the window the rate is averaged over.
  • The split itself is too disruptive: lower SplitDrainBatchSize (default 1024) to make each drain step smaller, at the cost of a longer overall split.
  • Uneven LiveKeys needs no fix. Splitting redistributes virtual slots to relieve request load, and each split itself leaves the source shard and its new sibling holding about half the keys of an unsplit peer. Placement hashes the whole key, so key design cannot concentrate keys on one shard: a LiveKeys column that is uneven in proportion to slot ownership is the footprint of past splits.

Slow scans

Symptom

A range scan or full enumeration takes much longer than the key count suggests it should, or its latency degrades over time on a tree whose live-key count is flat.

Likely cause

Four things dominate scan cost:

  1. Tombstone bloat. A deleted key leaves a tombstone that a scan still walks. A shard at 50 percent TombstoneRatio does twice the work per live result. This is the classic case of a scan getting slower while TotalLiveKeys stays flat.
  2. Shard fan-out. Keys are routed to shards by hash (a virtual slot derived from XxHash32 of the key), so a scan fans out to every physical shard no matter how narrow the range you ask for. The scan's wall-clock is therefore bounded by its slowest shard - one hot or bloated shard slows every scan of the tree.
  3. Leaf-chain walk depth. Within a shard, a scan walks the leaf chain one leaf at a time. More entries per shard means more leaves to hop.
  4. Page size and round trips. KeysPageSize (default 512) sets how many keys each shard returns per page; a small page size on a large scan multiplies round trips.

How to confirm

  • Take a deep report and read TombstoneRatio per shard. This is the fastest discriminator: a high ratio on the shards you scan is the answer. Remember that a shallow report reports zero tombstones unconditionally.
  • Watch orleans.lattice.leaf.scan.duration to see whether the time is going into per-leaf work, and orleans.lattice.leaf.tombstone.ratio for the distribution of per-leaf ratios within each tree (the histogram is tagged by tree, not by leaf, and is sampled only when a compaction pass visits a leaf - by default once per TombstoneGracePeriod, or on demand through CompactShardAsync - so it can trail a deep report by up to that period). See Metrics.
  • Compare per-shard figures in the report. Because every scan fans out to every shard, the slowest shard sets the pace. Look for the shard whose TombstoneRatio or OpsPerSecond is the outlier and treat that shard as the scan's bottleneck.

How to fix

  • Compact the bloated shards. MinTombstoneRatioForCompaction (default 0.0, meaning ratio-triggered compaction is off) and MaxLeafEntriesBeforeForcedCompaction (default 0, off) are the policy triggers; CompactionTriggerCooldown (default 5 minutes) rate-limits them. TombstoneGracePeriod (default 24 hours) is how long a tombstone must age before it can be reaped. For a one-off, drive a shard directly:

    var accepted = await tree.CompactShardAsync(shardIndex: 2, cancellationToken);
    
    Console.WriteLine(accepted
        ? "compaction pass accepted"
        : "request declined (compaction disabled for this tree, or a pass already in flight)");
    

    See Tombstone compaction for the full policy and its telemetry.

  • Bound the range. Prefer the startInclusive / endExclusive overloads over an unbounded enumeration. A bound does not reduce the number of shards visited - routing is by hash, so every shard is asked - but it cuts the work each shard does and keeps the per-scan dedup set the tree grain holds small, since that set grows with the number of distinct keys the scan has yielded.

  • Use the resilient streaming API for scans. ScanKeysAsync and ScanEntriesAsync transparently reconnect and resume when an enumeration is aborted mid-flight, with a default budget of 8 reconnect attempts (overridable per call via maxAttempts). The raw ILattice.KeysAsync and EntriesAsync streams surface the abort instead, and because the tree grain is a stateless worker the abort rate rises with concurrency on the tree rather than with scan length, so prefer the resilient pair for scans of any length.

  • Turn on prefetch for scans you will consume in full. PrefetchKeysScan and PrefetchEntriesScan both default to false; a per-call prefetch: true argument overrides them for a single scan. Prefetch fetches each shard's next page in parallel while the current page is being consumed, hiding per-shard grain-call latency. Prefetched pages are held in memory until consumed, so a caller that aborts early (a Take(n), say) pays for pages it never reads.

  • Raise KeysPageSize (default 512) if you are paging a large result set and the round-trip count dominates.

If a scan does not just run slowly but fails, that is a different problem. Strongly-consistent scans (CountAsync, CountPerShardAsync, KeysAsync, EntriesAsync) reconcile against shard-map changes that land mid-scan, bounded by MaxScanRetries (default 3); exhausting the budget throws an InvalidOperationException whose message names the operation and tells you to raise MaxScanRetries or reduce concurrent split activity. If you see it, read Concurrent split activity first - the scan is a symptom of the topology churn, not the cause. GetManyAsync spends the same budget when a shard-map change or a concurrently committing atomic-write saga races its batched read; its exhaustion message tells you to reduce the concurrent saga rate instead. A multi-key read whose result depended on a pending atomic write while the transaction registry could not be reached throws LatticeTransactionOutcomeUnavailableException (a TimeoutException) instead of either message: a transient condition, so retry after a back-off; see Atomic writes. See Consistency for the enumeration guarantees.


Stale reads and cache behaviour

Symptom

A read returns a value you believe was already overwritten or deleted, or two reads issued close together disagree.

Likely cause

Lattice serves ExistsAsync and GetManyAsync - and any GetAsync the shard root serves serially - through a per-silo read-through cache. With OptimisticShardRootPointReads on (the default), a GetAsync whose optimistic read validates reads the primary leaf instead and never consults the cache. Whether the cache can return a stale value is entirely determined by CacheTtl:

  • CacheTtl = TimeSpan.Zero (the default). Every read confirms freshness before answering. When the primary leaf is activated on the same silo and its revision has not moved since the cache last refreshed, that local check is enough; otherwise the cache performs a delta refresh against the primary. The delivery-cursor comparison behind that refresh is cheap, and either way the default configuration does not serve stale values.
  • CacheTtl set to a non-zero value. When the primary leaf is on another silo, the cache may answer from its local dictionary without contacting the primary for up to that interval. This is the trade you opted into: lower read latency, staleness bounded by the TTL. A same-silo primary whose revision has moved is still refreshed immediately.

So on a default-configured tree, a surprising read is almost never the read cache. See Caching for the refresh protocol and Consistency for the guarantees each read path gives.

How to confirm

Compare a cacheable read against a read that cannot be served from cache. GetWithVersionAsync bypasses the cache deliberately, because compare-and-swap callers need the authoritative version:

var cached = await tree.GetAsync("orders/42", cancellationToken);
var authoritative = await tree.GetWithVersionAsync("orders/42", cancellationToken);

Console.WriteLine($"cached-path bytes: {cached?.Length ?? -1}");
Console.WriteLine($"authoritative bytes: {authoritative.Value?.Length ?? -1}");
  • They agree. The cache is not involved; the value really is what the tree holds. Look at the writer instead.
  • They disagree and CacheTtl is non-zero. Expected staleness, bounded by the TTL you configured.
  • They disagree and CacheTtl is TimeSpan.Zero. That is not ordinary cache staleness. Capture both results and treat it as a consistency issue.

Two secondary checks:

  • orleans.lattice.cache.hits and orleans.lattice.cache.misses in Metrics tell you whether the cache is answering at all.
  • If you have set MaxCacheValueBytes, the cache evicts value payloads only - least recently used first, once the payload bytes it holds for a leaf exceed the budget - while each row's metadata envelope stays resident; a read landing on an evicted payload transparently fetches from the primary, and a single-key GetAsync counts it as a miss while a GetManyAsync batch leaves it out of both counters. That costs an RPC but cannot return a stale value.

How to fix

  • Set CacheTtl to TimeSpan.Zero for a tree whose reads must always reflect the latest committed write:

    siloBuilder.ConfigureLattice("orders", options =>
    {
        options.CacheTtl = TimeSpan.Zero;
    });
    
  • Use GetWithVersionAsync for read-modify-write. It bypasses the cache and returns the version a conditional write needs. Never build a compare-and-swap on top of GetAsync.

  • Do not confuse the two caches. CacheTtl governs value reads; DiagnosticsCacheTtl (default 5 s) governs DiagnoseAsync reports only. A stale report is a diagnostics-cache artefact and says nothing about read freshness - which is exactly why a caller that needs an authoritative live count should use CountAsync rather than reading TotalLiveKeys out of a possibly-cached report.

  • After a resize or reshard, no cache flush is needed. A resize rebuilds the tree into a new physical tree with different leaf grain identities, so reads land on fresh cache activations. A reshard that grows the tree moves virtual slots onto new physical shards, so moved keys likewise read through new leaves and their fresh caches; one that shrinks it folds a shard's keys into an adjacent shard's existing leaves, whose caches pick the moved keys up on their next refresh like any other write. Keys that stay put keep reading through the same leaves, whose caches keep refreshing against them. See Tree sizing and Online reshard.


Symptom index

Symptom Section
Provider exception from a WAL append or WriteStateAsync Storage-provider exceptions on write
LatticeQuotaExceededException on write Storage-provider exceptions on write (admission control, not a provider limit)
LatticeStateWriteFailedException from an atomic write Storage-provider exceptions on write - FaultType names the provider fault; with Conflict set, retry with the same operation id
Repeating "Proactive snapshot capture ... failed" warning Storage-provider exceptions on write
One shard far hotter than its peers Concurrent split activity
SplitInProgress stuck on for a long time Concurrent split activity
BulkOperationPending stuck on Concurrent split activity and Bulk loading
Scan latency climbing while live keys stay flat Slow scans
InvalidOperationException naming MaxScanRetries Slow scans, then Concurrent split activity; from GetManyAsync, concurrent atomic writes
LatticeTransactionOutcomeUnavailableException on a read Atomic writes - the transaction registry was unreachable for a key under a pending atomic write: transient, retry after a back-off
LatticeSaturatedException on a read or write WAL saturation signal - back-pressure, not a fault: back off and retry. SaturationSource names the seam that refused, and the source tag on orleans.lattice.saturation.refusals counts refusals by seam; a replay_permit_admission refusal also carries an arm tag naming which part of the replay admission check refused
LatticeTreeOwnershipDeniedException from an alias change Ownership-bounded aliasing - the registered ITreeOwnershipGuard refused the alias before anything was written; Reason carries the guard's explanation
Leaf Error that it cannot advance its durable projection checkpoint, or LeafProjectionStaleException A live leaf whose projection has gone stale - data at risk: capture a backup before the activation is recycled
High TombstoneRatio Slow scans and Tombstone compaction
Read returns an overwritten value Stale reads and cache behaviour
One shard reports all zeros on a tree that holds data Traps when reading a report
Report is older than expected Traps when reading a report
Depth growing across shards Tree sizing

See also

  • Diagnostics - the field-by-field definition of TreeDiagnosticReport and ShardDiagnosticReport.
  • Metrics - the full instrument catalog behind every metric named here.
  • Configuration - every LatticeOptions member, with defaults and tuning guidance.
  • Tree storage - per-provider size limits and the row-size arithmetic.
  • Tree sizing - choosing and changing MaxLeafKeys.
  • Shard splitting - the split protocol, suppression rules, and tunables.
  • Tombstone compaction - compaction policy, operator API, and telemetry.
  • Caching - the read-through cache and its refresh protocol.
  • Consistency - the guarantees each read and scan path gives.