The agent-operated backlog
This page documents Orleans.Lattice.Api.Mcp.RepoContext, which is unreleased, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at backlog.md, and llms.txt lists every page.An agent-operated backlog lets a fleet of agents pick up, execute, and close out work on this repository with the product owner steering rather than dispatching. Work items live as repository-context memory entries, mirrored to GitHub issues for human oversight; scheduled or on-demand worker agents drain them; a project manager agent curates them and is the human's single point of contact.
This page is the operator's view: what the mechanism is, what it guarantees, and where each part is defined. It does not restate the data model or the agent protocols, which live in a repository-neutral template this repository consumes unmodified, plus a small set of repository-specific instruction files, and would drift if copied here.
| Concern | Defined in |
|---|---|
| Item schema, relation vocabulary, ready-set algorithm, grouping model, mirroring, gating | samples/AgentBacklog/template/backlog-protocol.md; the two rules that bind every agent touching memory (never prune the backlog relations, never set a TTL on an item) are also restated in .github/instructions/repocontext.instructions.md, section The agent-operated backlog |
| Worker behaviour | samples/AgentBacklog/template/backlog-worker.base.md, with this repository's bindings in .github/agents/backlog-worker.agent.md |
| Project-manager behaviour | samples/AgentBacklog/template/backlog-pm.base.md, with this repository's bindings in .github/agents/backlog-pm.agent.md |
| Epic branch convention and CI trigger | .github/copilot-instructions.md |
| Claim and lease tools | Tools |
Why a claim needs a fence
The backlog is a shared, concurrently-drained queue held in a CRDT store, and
that combination has a specific hazard. Repository-context memory records merge
without coordination: scalar fields (title, body, author, provenance) are
last-writer-wins registers, so two concurrent writers silently lose one write,
while tags and links are add-wins sets, so two concurrent writers both
survive. Nothing in the surface offers compare-and-swap: repocontext_update
preconditions on record existence only, never on value.
So an edge asserting "this run owns this item" is an audit record of who tried, not a lock. Two workers can both assert it and both believe they won.
Worse, the natural fix - a boolean "claimed" flag - fails in exactly the case that matters. A scheduled agent session can be stopped at any instant, and in practice a meaningful fraction of them are. A flag set by a killed session stays set forever, and the item is stranded with no way to tell a live owner from a dead one.
The claim surface addresses both by wrapping the cluster-wide distributed lock rather than reimplementing mutual exclusion:
- Exclusion and fairness come from the lock, which is FIFO-fair across the cluster.
- Liveness comes from the lock's bounded, expiry-reclaimed lease. A claim's
lease is never a flag; it always expires. A worker that dies mid-item frees the
item by doing nothing - the lock reclaims the lapsed lease, so the next claim is
granted - and
Claimed -> Readyon lease expiry is the normal path, not an exception. Until that next claim or a release, the record still reports the dead worker's claim (claimed: truewithisHeld: falseinrepocontext_claim_status) and keeps refusing unfenced writes. - Safety after a handover comes from the lock's monotonic fencing token, which strictly increases and is never reused across activations or crashes.
The fence is enforced, not advertised
A fencing token that only the well-behaved consult is decoration. The
repository-context store therefore checks the token on the write path itself:
repocontext_remember, repocontext_update, and repocontext_forget each take
an optional fencing token, and each is refused when the token does not admit the
write. Refusal raises RepoContextClaimConflictException.
The admission rules, in order:
| Record state | Token presented | Outcome |
|---|---|---|
| Never claimed | any, or none | Accepted. Every pre-existing caller is unchanged. |
| Claim live | none | Refused - a live claim excludes unfenced writes. |
| Claim released | none | Accepted - a released record readmits unfenced writes. |
| Claim live or released | below the record's high-water mark | Refused - a superseded holder can never write. |
| Claim live or released | above the high-water mark, but not the token the lock currently holds | Refused - a token ahead of the record's stamp is honoured only when the lock confirms it issued it. |
| Claim released | at or above the high-water mark | Refused - re-claim first. |
| Claim live, token current, different region | current token | Refused - claims are region-scoped. |
| Claim live, token current, same region | current token | Accepted. |
A claim is live from the moment it is stamped until it is released; the write path never reads the lease's expiry. A holder whose lease has lapsed therefore still passes the check under its token until another claim is granted, which stamps a strictly higher token and puts the lapsed holder below the high-water mark.
Two consequences are worth stating plainly, because collapsing them is how "fenced" gets misread as "true":
- The fence guarantees authorship against a stale writer: a superseded holder
cannot overwrite the resume block in an item's
body, and neither can a write presenting no token while a claim is live. That is what makes an LWW register safe here - not convention, but enforcement on the write path itself. It checks the token, not the owner: the current token is readable throughrepocontext_claim_status, and release readmits unfenced writes, so it excludes a stale or tokenless writer, not one that presents the current token or writes after the claim is released. - The fence guarantees nothing about content. A resume note is the last attempt's own account of itself. A resuming worker re-decides from it and never continues blindly.
The high-water mark is stored as a bounded register keyed by the token itself, so
it behaves as a join-semilattice maximum: no lower token can displace it, whether
by a direct write or by a concurrent replica merge. That is how a trustworthy
high-water mark is obtained without compare-and-swap. Claim state sits outside
the settable-scalar allow-list, so repocontext_update cannot forge it.
Claims apply to memory records only, which is the family the write-path check guards. A token presented against any other record family is rejected rather than silently ignored.
Reading a claim decision
The claim tools report contention rather than throwing, so a worker branches on the result instead of catching an exception:
using Orleans.Lattice.Api.Mcp.RepoContext;
static string Describe(RepoContextClaimResult claim) =>
claim switch
{
{ Granted: true, FencingToken: { } token } =>
$"claimed {claim.Key} with fencing token {token}, lease to {claim.LeaseExpiresAtUtc}",
// Losing a race is an ordinary outcome, not a fault: another worker holds
// the lock, or the wait elapsed, or the record does not exist.
_ => $"not claimed ({claim.Reason}); pick another item",
};
repocontext_renew_claim returns the same shape, and a reason of superseded is
the authoritative signal that this run's lease is gone: it lapsed and the lock
reclaimed it, whether or not another worker has claimed the item since, and the
next claim by anyone fences this run's token out. It must abandon immediately
without writing anything further.
Always pass leaseSeconds explicitly on a renew. Omitting it does not
preserve the lease being held - it requests the cluster's configured default,
which is deliberately short (LatticeOptions.DefaultLockLeaseDuration, 30 seconds
unless the host overrides it), so renewing a long claim without a length cuts it
down. Both outcomes are granted: true, so the reduction is reported separately
as leaseShortened: true with the prior expiry in previousLeaseExpiresAtUtc.
Without that signal the reduction is invisible until the next renew, which
returns superseded - by which point the worker has lost its lease mid-task
while believing it held the claim. Treat a shortening renew as a prompt to renew
again with an explicit length, not as success.
An explicit length is still clamped: every granted or renewed lease is capped at
LatticeOptions.MaxLockLeaseDuration, 5 minutes unless the host overrides it. The
repository-context container raises that ceiling to 30 minutes, through
LATTICE_MAX_LOCK_LEASE_SECONDS (accepted range 30-7200 seconds), so a claim can
span a build-and-test cycle. Act on the returned leaseSeconds and
leaseExpiresAtUtc, never on the length requested.
using Orleans.Lattice.Api.Mcp.RepoContext;
static string AfterRenew(RepoContextClaimResult renewed) =>
renewed switch
{
// Not a failure: the renew succeeded, but it moved the expiry earlier.
// Renew again with an explicit leaseSeconds rather than carrying on.
{ Granted: true, LeaseShortened: true } =>
$"lease SHORTENED from {renewed.PreviousLeaseExpiresAtUtc} "
+ $"to {renewed.LeaseExpiresAtUtc}; renew again with an explicit length",
{ Granted: true } => $"lease held to {renewed.LeaseExpiresAtUtc}",
_ => $"lost the item ({renewed.Reason}); stop writing",
};
repocontext_claim_status is advisory only. Its authoritative property is
hard-wired to false so that no call site can project an authoritative status
from a read. Use it to observe and report; never to gate a decision. Only a
granted claim or a renew verdict is authoritative.
Grouping, and why an epic gets one branch
Items are grouped: an epic and the sub-items joined to it by partOf edges. A
grouping runs in up to three phases - an optional research phase that may never
nest inside itself, an implementation phase that fans out behind a small seam
item, and an integration phase.
The integration phase exists because maximum parallelism concentrates risk at the join. N pull requests, each green in isolation against a different base, none ever tested against the others, can satisfy every sub-item's acceptance criteria while failing the epic's. Every grouping therefore terminates in exactly one integration item that is blocked by the whole fan-out, runs the full cross-package suite rather than per-item targeted runs, verifies against the epic's criteria, and holds an exclusive claim with the grouping's other workers quiesced. A grouping is not complete until it completes, however green its sub-items are.
Groupings also share one branch. Because main requires a status check with
"up to date before merging", every merge into main invalidates every other open
pull request, which must then update and re-run the whole suite: N concurrent
sub-items cost quadratic CI. So an epic gets <type>/epic/<slug>, sub-items branch
off it as <type>/epic/<slug>-<item> (a hyphen, not a slash: git cannot hold a branch
X and a branch X/anything at once), and the epic reaches main as a single
fully-gated pull request. CI runs on epic-targeted pull requests, and an epic
branch deliberately carries no protection - CI running is what gives feedback,
while a strict required check is what serialises. The full rules, including who
keeps the branch current with main, are in .github/copilot-instructions.md.
Oversight stays with the human
Two stores, one source of truth each, and neither copies the other's content: GitHub owns identity, specification, priority, audit trail, and notification; repository-context memory owns the dependency graph, code anchors, claims, and resume pointers.
Every item is mirrored to a GitHub issue at creation and takes its id from that
issue, so nothing can be enqueued invisibly. An item an agent proposed
additionally opens carrying the existing needs-specification label and stays out
of the ready set until a human removes it - no agent removes that label, not even
after writing the item's specification itself. A human can reprioritise or respecify without an agent in the
loop, because the thing they edit is the thing that is authoritative.
Linking an item to the code it concerns (anchoredTo) captures those targets'
content digests, so an item whose code has drifted is reported stale and
re-validated before a run is spent on it. That is the one capability the issue
tracker cannot provide, and it doubles as the mitigation for items that would
otherwise fail repeatedly against a spec the code has outgrown.
See also
- Tools - the claim and lease tool contracts.
- Record model - record families and the CRDT store-of-record model.
- Memory and TTL - topics and per-entry expiry. Note that a backlog item never carries a TTL: expiry is silent, so a lapsed item starves its dependents with no event to explain it.
- Distributed lock - the fencing and lease primitive the claim surface wraps.
- Agent backlog sample - a runnable walkthrough of claiming, fencing, and release.
- Adopting the backlog - the copyable base protocol and agent definitions, plus the GitHub-side setup no file copy can do. This repository consumes that template unmodified.