Operator tooling: orphaned-leaf repair
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at operator-tooling-orphaned-leaf-repair.md, and llms.txt lists every page.Part of Lattice Public API Reference.
VerifiedKeyCount always means the verified prefix length before the first
missing key or routing contradiction. The default audit and repair do not test
later keys. Do not subtract this prefix from KeyCount to estimate damage.
The opt-in SurveyOrphanedLeavesAsync preserves that field and the first failure
in UnverifiedKey, and adds SurveyVerifiedKeyCount, SurveyMissingKeyCount
and SurveyRoutingContradictionKeyCount on each finding. Together they account
for every enumerated key; null means the leaf was not surveyed (for example,
blocking state or the key limit), not zero. A missing routed copy is not proof
of data loss. This is a live observation, not a consistent tree snapshot.
Reports include DryRun, Survey, LeavesWalked, OrphanedLeafCount, RepairableCount,
RepairedCount and RefusedCount, with positions (shard, leaf id, LowKeyInclusive /
HighKeyExclusive bounds, key counts) and OrphanedLeafDisposition values in Findings. The nullable
report SurveyMissingKeyCount totals this batch only, and is unknown if any
reached region or orphan could not be surveyed. Collect every batch using the
same survey verb and inspect Gaps; IsComplete alone does not make a complete
census. Findings marked Repairable identify the repair's candidate scope;
repair independently rechecks safety and still refuses every unverified leaf.
The public ILattice and ILatticeTreeAdmin survey methods have default
implementations for compatibility with older external implementations. Their
default throws NotSupportedException: unsupported survey must not be mistaken
for a zero-damage census. Lattice's grain and tree-admin facade implement the
survey; older remote servers may not support it.
An orphaned leaf is a leaf that is present in a shard's doubly
linked sibling chain but is not reachable by descent from the shard
root - no routing entry points at it. Such a leaf can only arise from
an interrupted split that spliced a new sibling into the chain before
publishing its routing entry. The immediate cause of that window is
fixed, but a tree that already acquired an orphan before the fix has
no self-healing path: empty-leaf reclaim rejects a candidate on
LiveRowCount != 0 before it ever evaluates descent reachability, and
an orphan is rarely empty - activation-time WAL replay admits records
by (ShardIndex, LowKeyInclusive, HighKeyExclusive) and never by leaf
identity, so an orphan sharing bounds with a live leaf materialises a
shadow copy of that range. The orphan's projection checkpoint then
pins the WAL trim floor indefinitely, and because compaction is
downstream of trim, the WAL grows without bound.
Three ILattice methods provide the operator path.
| Method | Description |
|---|---|
InspectOrphanedLeavesAsync(string?, CancellationToken) |
Dry run. Walks the sibling chain of every physical shard, reports each leaf that is not reachable by descent, and evaluates the same safety verification the repair uses - so a leaf reported Repairable here is one RepairOrphanedLeavesAsync would unsplice. Mutates nothing. Requires LatticeOperation.Read. |
SurveyOrphanedLeavesAsync(string?, CancellationToken) |
Opt-in read-only census of all key outcomes per orphan, bounded at 100,000 keys per leaf. Costs O(KeyCount) descents/owner reads rather than O(first miss). Requires LatticeOperation.Read. |
RepairOrphanedLeavesAsync(string?, CancellationToken) |
Performs the repair. For each descent-unreachable leaf it verifies every key the leaf holds is also held by the descent-reachable leaf that key routes to; only then does it unsplice the leaf (relinking its neighbours) and retire its projection state, releasing the pin. Requires LatticeOperation.Admin. |
The leading string? on all three is the resume token from the previous
call's report. Pass null (the default) to start a new pass; pass
report.ResumeFrom back unaltered to continue one. See
Driving a pass to completion below -
it is not optional detail, because one call is one bounded batch and a
caller that ignores the token only ever sees the first one.
The verification in step two fails closed. If the orphan holds any
key that the routed live leaf does not hold, the leaf is left exactly
as it was and reported RefusedUnverifiedKeys with the offending key (a leaf the repair does unsplice is reported Repaired, and one the dry run would unsplice Repairable);
unsplicing it would lose data. Key values are deliberately not
compared, because an orphan and a live leaf replay the same WAL
records on independent horizons and a benign version difference must
not produce a spurious refusal. Other refusal dispositions -
RefusedKeyCountExceeded, RefusedBlockingState, RefusedChainRace,
RefusedRoutingContradiction - are likewise no-ops on the tree.
The unsplice deliberately does not widen the surviving predecessor to absorb the orphan's key range, which is where it diverges from empty-leaf reclaim. A reclaimed leaf is routed, so its range must be re-homed; an orphan is not routed, so the routed leaves already tile the keyspace and widening a neighbour over the orphan's bounds would create an overlap - and therefore a second materialisation of that range.
Driving a pass to completion
One call to either verb is one bounded batch. It returns when its
work budget is spent, leaving ResumeFrom non-null and IsComplete
false; pass that token back unaltered to continue from exactly where it
stopped. The bound is deliberate: an earlier revision drove the whole
fan-out inside a single call, and on an ordinary tree that had
accumulated 236 orphans it ran past the Orleans client response deadline
and surfaced a TimeoutException to the caller - after the grain had
already completed every repair. The operation reported failure having
entirely succeeded.
Read VerdictComplete too, and read it before Findings
IsComplete and VerdictComplete answer different questions.
IsComplete says how far the pass got; VerdictComplete says whether
it could judge what it reached.
A report carries Gaps: every region of the tree the pass could not
establish a verdict over. A shard that declined because it was mid-split
(ShardSplitInProgress) or because another orphaned-leaf pass already
held it (ShardPassAlreadyRunning), a sibling chain severed part-way
across the keyspace (ChainTruncated when the walk re-entered past the
break, ChainTruncatedUnrecoverable when it could not), a leaf whose
declared bounds make reachability undecidable (LeafBoundsUndecidable),
an entry leaf - the chain head, a resume position, or a re-entry past a
break - that is itself unreachable by descent (EntryLeafUnreachable),
and a shard-level budget exhausted with no position to resume from
(WalkBudgetExhaustedWithoutResumePosition) are all reported here, each
as an OrphanedLeafAuditGap carrying its OrphanedLeafAuditGapReason,
rather than silently folded into an empty findings list.
This matters because the enumeration walks each shard's sibling chain from its head, and the chain is the same structure an orphan damages. Before issue 3301 a pointer severed mid-keyspace ended the walk, the drive read the resulting null resume position as completion, and the shard was reported examined and clean - while range scans, which enter the chain by descending on their own lower bound, kept reaching the segment past the break and reporting orphans in it. Both were telling the truth about the same shard.
VerdictComplete is true only when Gaps is empty. An empty
Findings list is a clean bill of health only when IsComplete and
VerdictComplete are both true. When either is false the answer is
"this could not be established", not "there is nothing here", and it
does not rule an orphan out as the cause of an unbounded WAL.
// Dry run first: see what would be repaired, and why anything is refused.
// One call is one bounded batch, so drive it until IsComplete.
var findings = new List<OrphanedLeafFinding>();
var gaps = new List<OrphanedLeafAuditGap>();
string? cursor = null;
int refused = 0;
do
{
OrphanedLeafRepairReport batch =
await tree.InspectOrphanedLeavesAsync(cursor, cancellationToken);
findings.AddRange(batch.Findings);
gaps.AddRange(batch.Gaps);
refused += batch.RefusedCount;
cursor = batch.ResumeFrom;
}
while (cursor is not null);
foreach (OrphanedLeafAuditGap gap in gaps)
{
// Read this BEFORE the findings: a region that could not be judged
// contributes no findings by construction.
_ = (gap.ShardIndex, gap.Reason, gap.LeafId, gap.KeyHint);
}
foreach (OrphanedLeafFinding finding in findings)
{
if (finding.IsRefusal)
{
// Investigate before escalating - a refusal means repair is unsafe.
_ = (finding.ShardIndex, finding.LeafId, finding.Disposition, finding.UnverifiedKey);
}
}
if (gaps.Count == 0 && findings.Count == 0)
{
// The only reading that actually rules the defect out: the pass was
// driven to completion AND it could judge everything it reached.
}
if (findings.Count > 0 && refused == 0)
{
cursor = null;
do
{
OrphanedLeafRepairReport batch =
await tree.RepairOrphanedLeavesAsync(cursor, cancellationToken);
_ = (batch.LeavesWalked, batch.RepairedCount);
cursor = batch.ResumeFrom;
}
while (cursor is not null);
// Then RE-AUDIT. The repair's own return is not the source of truth.
}
Both calls are safe to run against a live tree under load, and both are
idempotent: re-running cannot double-repair, because a leaf already
unspliced is gone from the chain and a refused one is refused again on
the same evidence. A pass restarted from null re-establishes the truth
from scratch rather than compounding anything.
These consequences are worth stating plainly, because each is a way to misread or misuse a correct report.
- An empty
Findingson a partial batch is not a clean tree. It says only that the part of the tree this batch reached was clean. The clean bill of health requiresIsComplete. - An empty
Findingswith a non-emptyGapsis not a clean tree either. A region the pass could not judge contributes zero findings by construction, so the clean bill of health requiresVerdictCompleteas well. - If you see a timeout or any transport error, the return value is not authoritative, and its absence is not evidence that nothing happened. The reply may have been lost after the work landed. Do not guess and do not simply retry: run the audit - which mutates nothing - and let it establish the true state. The safe loop is audit, repair to completion, then RE-AUDIT.
- A retry is not free even though it is safe. A second repair pass started while the first is still running mutates the same leaf chains under compare-and-swap. It cannot corrupt the tree, but it wastes the budget re-verifying work the other pass is doing. Drive one pass to completion rather than starting a second.
The work budget is wall-clock, not a leaf count - a call stops starting
new shard batches once LatticeOptions.BackgroundDrainMaxDuration (10 s
by default) has elapsed, checked between batches, while each shard batch
runs under the same wall-clock net plus its own 4,096-leaf cap - because
the cost of this pass is dominated by per-key verification rather than
by leaves traversed: in the incident above the tree that blew the deadline had
walked fewer leaves (2121) than the tree that returned comfortably
(2443), but held roughly 83 keys per orphaned leaf, each verified by an
individual descent. Any leaf cap that admitted the second tree would
have admitted the first.
One residual bound is worth knowing: a single leaf's key verification is atomic and cannot be split, since a partial verification proves nothing about safety. A pathological leaf approaching the 100,000-key verification ceiling can therefore still overrun on its own. The guarantee is bounded work per call, not an absolute wall-clock cap.