Table of Contents

Operator tooling: orphaned-leaf repair

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at operator-tooling-orphaned-leaf-repair.md, and llms.txt lists every page.

Part of Lattice Public API Reference.

VerifiedKeyCount always means the verified prefix length before the first missing key or routing contradiction. The default audit and repair do not test later keys. Do not subtract this prefix from KeyCount to estimate damage. The opt-in SurveyOrphanedLeavesAsync preserves that field and the first failure in UnverifiedKey, and adds SurveyVerifiedKeyCount, SurveyMissingKeyCount and SurveyRoutingContradictionKeyCount on each finding. Together they account for every enumerated key; null means the leaf was not surveyed (for example, blocking state or the key limit), not zero. A missing routed copy is not proof of data loss. This is a live observation, not a consistent tree snapshot.

Reports include DryRun, Survey, LeavesWalked, OrphanedLeafCount, RepairableCount, RepairedCount and RefusedCount, with positions (shard, leaf id, LowKeyInclusive / HighKeyExclusive bounds, key counts) and OrphanedLeafDisposition values in Findings. The nullable report SurveyMissingKeyCount totals this batch only, and is unknown if any reached region or orphan could not be surveyed. Collect every batch using the same survey verb and inspect Gaps; IsComplete alone does not make a complete census. Findings marked Repairable identify the repair's candidate scope; repair independently rechecks safety and still refuses every unverified leaf.

The public ILattice and ILatticeTreeAdmin survey methods have default implementations for compatibility with older external implementations. Their default throws NotSupportedException: unsupported survey must not be mistaken for a zero-damage census. Lattice's grain and tree-admin facade implement the survey; older remote servers may not support it.

An orphaned leaf is a leaf that is present in a shard's doubly linked sibling chain but is not reachable by descent from the shard root - no routing entry points at it. Such a leaf can only arise from an interrupted split that spliced a new sibling into the chain before publishing its routing entry. The immediate cause of that window is fixed, but a tree that already acquired an orphan before the fix has no self-healing path: empty-leaf reclaim rejects a candidate on LiveRowCount != 0 before it ever evaluates descent reachability, and an orphan is rarely empty - activation-time WAL replay admits records by (ShardIndex, LowKeyInclusive, HighKeyExclusive) and never by leaf identity, so an orphan sharing bounds with a live leaf materialises a shadow copy of that range. The orphan's projection checkpoint then pins the WAL trim floor indefinitely, and because compaction is downstream of trim, the WAL grows without bound.

Three ILattice methods provide the operator path.

Method Description
InspectOrphanedLeavesAsync(string?, CancellationToken) Dry run. Walks the sibling chain of every physical shard, reports each leaf that is not reachable by descent, and evaluates the same safety verification the repair uses - so a leaf reported Repairable here is one RepairOrphanedLeavesAsync would unsplice. Mutates nothing. Requires LatticeOperation.Read.
SurveyOrphanedLeavesAsync(string?, CancellationToken) Opt-in read-only census of all key outcomes per orphan, bounded at 100,000 keys per leaf. Costs O(KeyCount) descents/owner reads rather than O(first miss). Requires LatticeOperation.Read.
RepairOrphanedLeavesAsync(string?, CancellationToken) Performs the repair. For each descent-unreachable leaf it verifies every key the leaf holds is also held by the descent-reachable leaf that key routes to; only then does it unsplice the leaf (relinking its neighbours) and retire its projection state, releasing the pin. Requires LatticeOperation.Admin.

The leading string? on all three is the resume token from the previous call's report. Pass null (the default) to start a new pass; pass report.ResumeFrom back unaltered to continue one. See Driving a pass to completion below - it is not optional detail, because one call is one bounded batch and a caller that ignores the token only ever sees the first one.

The verification in step two fails closed. If the orphan holds any key that the routed live leaf does not hold, the leaf is left exactly as it was and reported RefusedUnverifiedKeys with the offending key (a leaf the repair does unsplice is reported Repaired, and one the dry run would unsplice Repairable); unsplicing it would lose data. Key values are deliberately not compared, because an orphan and a live leaf replay the same WAL records on independent horizons and a benign version difference must not produce a spurious refusal. Other refusal dispositions - RefusedKeyCountExceeded, RefusedBlockingState, RefusedChainRace, RefusedRoutingContradiction - are likewise no-ops on the tree.

The unsplice deliberately does not widen the surviving predecessor to absorb the orphan's key range, which is where it diverges from empty-leaf reclaim. A reclaimed leaf is routed, so its range must be re-homed; an orphan is not routed, so the routed leaves already tile the keyspace and widening a neighbour over the orphan's bounds would create an overlap - and therefore a second materialisation of that range.

Driving a pass to completion

One call to either verb is one bounded batch. It returns when its work budget is spent, leaving ResumeFrom non-null and IsComplete false; pass that token back unaltered to continue from exactly where it stopped. The bound is deliberate: an earlier revision drove the whole fan-out inside a single call, and on an ordinary tree that had accumulated 236 orphans it ran past the Orleans client response deadline and surfaced a TimeoutException to the caller - after the grain had already completed every repair. The operation reported failure having entirely succeeded.

Read VerdictComplete too, and read it before Findings

IsComplete and VerdictComplete answer different questions. IsComplete says how far the pass got; VerdictComplete says whether it could judge what it reached.

A report carries Gaps: every region of the tree the pass could not establish a verdict over. A shard that declined because it was mid-split (ShardSplitInProgress) or because another orphaned-leaf pass already held it (ShardPassAlreadyRunning), a sibling chain severed part-way across the keyspace (ChainTruncated when the walk re-entered past the break, ChainTruncatedUnrecoverable when it could not), a leaf whose declared bounds make reachability undecidable (LeafBoundsUndecidable), an entry leaf - the chain head, a resume position, or a re-entry past a break - that is itself unreachable by descent (EntryLeafUnreachable), and a shard-level budget exhausted with no position to resume from (WalkBudgetExhaustedWithoutResumePosition) are all reported here, each as an OrphanedLeafAuditGap carrying its OrphanedLeafAuditGapReason, rather than silently folded into an empty findings list.

This matters because the enumeration walks each shard's sibling chain from its head, and the chain is the same structure an orphan damages. Before issue 3301 a pointer severed mid-keyspace ended the walk, the drive read the resulting null resume position as completion, and the shard was reported examined and clean - while range scans, which enter the chain by descending on their own lower bound, kept reaching the segment past the break and reporting orphans in it. Both were telling the truth about the same shard.

VerdictComplete is true only when Gaps is empty. An empty Findings list is a clean bill of health only when IsComplete and VerdictComplete are both true. When either is false the answer is "this could not be established", not "there is nothing here", and it does not rule an orphan out as the cause of an unbounded WAL.

// Dry run first: see what would be repaired, and why anything is refused.
// One call is one bounded batch, so drive it until IsComplete.
var findings = new List<OrphanedLeafFinding>();
var gaps = new List<OrphanedLeafAuditGap>();
string? cursor = null;
int refused = 0;
do
{
    OrphanedLeafRepairReport batch =
        await tree.InspectOrphanedLeavesAsync(cursor, cancellationToken);
    findings.AddRange(batch.Findings);
    gaps.AddRange(batch.Gaps);
    refused += batch.RefusedCount;
    cursor = batch.ResumeFrom;
}
while (cursor is not null);

foreach (OrphanedLeafAuditGap gap in gaps)
{
    // Read this BEFORE the findings: a region that could not be judged
    // contributes no findings by construction.
    _ = (gap.ShardIndex, gap.Reason, gap.LeafId, gap.KeyHint);
}

foreach (OrphanedLeafFinding finding in findings)
{
    if (finding.IsRefusal)
    {
        // Investigate before escalating - a refusal means repair is unsafe.
        _ = (finding.ShardIndex, finding.LeafId, finding.Disposition, finding.UnverifiedKey);
    }
}

if (gaps.Count == 0 && findings.Count == 0)
{
    // The only reading that actually rules the defect out: the pass was
    // driven to completion AND it could judge everything it reached.
}

if (findings.Count > 0 && refused == 0)
{
    cursor = null;
    do
    {
        OrphanedLeafRepairReport batch =
            await tree.RepairOrphanedLeavesAsync(cursor, cancellationToken);
        _ = (batch.LeavesWalked, batch.RepairedCount);
        cursor = batch.ResumeFrom;
    }
    while (cursor is not null);

    // Then RE-AUDIT. The repair's own return is not the source of truth.
}

Both calls are safe to run against a live tree under load, and both are idempotent: re-running cannot double-repair, because a leaf already unspliced is gone from the chain and a refused one is refused again on the same evidence. A pass restarted from null re-establishes the truth from scratch rather than compounding anything.

These consequences are worth stating plainly, because each is a way to misread or misuse a correct report.

  • An empty Findings on a partial batch is not a clean tree. It says only that the part of the tree this batch reached was clean. The clean bill of health requires IsComplete.
  • An empty Findings with a non-empty Gaps is not a clean tree either. A region the pass could not judge contributes zero findings by construction, so the clean bill of health requires VerdictComplete as well.
  • If you see a timeout or any transport error, the return value is not authoritative, and its absence is not evidence that nothing happened. The reply may have been lost after the work landed. Do not guess and do not simply retry: run the audit - which mutates nothing - and let it establish the true state. The safe loop is audit, repair to completion, then RE-AUDIT.
  • A retry is not free even though it is safe. A second repair pass started while the first is still running mutates the same leaf chains under compare-and-swap. It cannot corrupt the tree, but it wastes the budget re-verifying work the other pass is doing. Drive one pass to completion rather than starting a second.

The work budget is wall-clock, not a leaf count - a call stops starting new shard batches once LatticeOptions.BackgroundDrainMaxDuration (10 s by default) has elapsed, checked between batches, while each shard batch runs under the same wall-clock net plus its own 4,096-leaf cap - because the cost of this pass is dominated by per-key verification rather than by leaves traversed: in the incident above the tree that blew the deadline had walked fewer leaves (2121) than the tree that returned comfortably (2443), but held roughly 83 keys per orphaned leaf, each verified by an individual descent. Any leaf cap that admitted the second tree would have admitted the first.

One residual bound is worth knowing: a single leaf's key verification is atomic and cannot be split, since a partial verification proves nothing about safety. A pathological leaf approaching the 100,000-key verification ceiling can therefore still overrun on its own. The guarantee is bounded work per call, not an absolute wall-clock cap.