The agent-operated backlog protocol
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at backlog-protocol.md, and llms.txt lists every page.This is the generic base for an agent-operated backlog: a durable work queue
that lives in repocontext memory, is drained concurrently by agent sessions
under fenced claims, and is mirrored to GitHub issues for human oversight.
It is written to be repository-neutral and is the single source of truth for the protocol. The repository that hosts it consumes it unmodified, so it cannot rot into a stale copy: if this document is wrong, that repository's own backlog agents are wrong with it.
The backlog topic is a specialisation of ordinary repocontext memory:
ordinary memory entries, an extended relation vocabulary, and rules that make
the graph safe for several agents to drain concurrently. Read it before
authoring, claiming, or completing a backlog item. The agent definitions that
implement it are backlog-pm.base.md and
backlog-worker.base.md; this document defines the
data they operate on. Background on the fencing mechanism is in
The agent-operated backlog.
Bindings
Everything repository-specific is a binding, written as {placeholder}
throughout this document and supplied by the consuming repository. The bindings
are:
| Binding | Meaning | Example |
|---|---|---|
{repoId} |
The repocontext repository id, as reported by repocontext_list_repos. Not your working directory, and not a worktree name. |
my-repo |
{owner}/{repo} |
The GitHub repository that mirrors items as issues. | my-org/my-repo |
{ghAccount} |
The GitHub account every gh call authenticates as. |
my-github-account |
{homeRegion} |
The region claims are taken in. Derive it from the cluster, never from a geography. lattice_list_regions gives the routable ids; repocontext_claim_status on any claimed item gives the region a claim actually records, and that is the value the tag must match. Whether it is load-bearing or merely informational is a property of the deployment, not of the tag - see the tag table below. |
local |
{conventionsDoc} |
The repository's contribution conventions: branch naming, commit rules, labels. | .github/copilot-instructions.md |
{implementationAgent} |
The agent a worker delegates feature implementation to, if the repository has one. | feature-dev |
A consuming repository supplies these in one small override file rather than by
editing this document. See bindings.example.md and the
adoption guide.
If the bindings are not available, stop and report. Do not guess a
repository, an account, or a region: a gh call under the wrong identity and a
claim taken in the wrong region both fail in ways that are expensive to unpick.
Responsibility split - two stores, one source of truth each. Neither store copies the other's content, because two writable copies with no transaction between them diverge, and the divergence surfaces weeks later.
| Concern | Owner |
|---|---|
| Item identity, specification, human-visible priority, oversight, audit trail, notifications | GitHub issues |
| Dependency graph, code anchors, claims, resume pointers, durable learnings | repocontext memory |
A human can reprioritise or respecify without an agent in the loop, because the thing they edit (the issue) is the thing that is authoritative.
Item schema
Topic backlog. One entry per item. The id is derived deterministically
from the item - issue-2057, never a generated GUID - so a retry merges in
place instead of creating a near-duplicate. The id is the mirrored issue number,
which makes identity and mirroring the same act (see
Entry gating).
A memory record has exactly four author-settable scalars besides the
kind that repocontext_remember fixes at creation (Decision, Note or
Memory), and this is the constraint the whole schema is built around. repocontext_update accepts
title, body, author and provenance on a memory record and rejects
every other name - update(fields: { "priority": "P0" }) fails with "The
field 'priority' is not a settable scalar on a Memory record". There is no
generic field bag. createdAt is set by the store at creation and is never
authored. So an item's structured attributes must be carried by the two
collection-valued members that do accept arbitrary content: tags and
links.
| Carrier | CRDT | Concurrency behaviour |
|---|---|---|
Scalars (title, body, author, provenance) |
LWW register | Two concurrent writers: one write is silently lost. |
tags |
add-wins OR-Set | Two concurrent writers: both survive, so the collision is visible. |
links |
OrMap<string, OrSet> per relation |
Two concurrent writers: both survive, converging per relation. |
The allocation follows directly from that table.
Attribute tags
Single-valued, low-cardinality attributes are carried as key:value tags.
Tags are returned by scan and recall and are matchable by search, so an
attribute expressed this way is filterable without reading bodies. Arbitrary
: and / characters round-trip intact, so a branch name is a legal tag value.
| Tag | Meaning |
|---|---|
backlog |
Plain marker tag. Every item carries it. |
priority:P0 .. priority:P3 |
Ordering priority. |
phase:research | phase:implementation | phase:integration |
Which phase of its grouping the item belongs to. Set at authoring, never changed by a worker, and it never carries execution state - a phase:complete or phase:review tag is a defect, not a status. |
homeRegion:<region> |
The region in which claims for this item are taken. Verify this is enforced before relying on it. The intent is that a write served from any other region, under a claim taken in this one, fails closed on the write path (ForeignRegion), because the underlying lock is cluster-scoped and so the claim is not observable there - but that is a property of the deployment, not of the tag. On a single-region deployment lattice_list_regions reports only current, claims report region local, and a geographic value such as uksouth is not enforced at all: a claim from anywhere succeeds. Treat a value the cluster does not route as a defect in the binding, and note the failure is worse than a no-op - an unenforced safety assumption that the protocol documents as enforced is more dangerous than an absent one, because it is relied upon. |
baseBranch:<branch> |
The branch this item's pull request targets. For an item in a grouping this is the epic branch, never main. |
state:complete | state:parked |
The item's terminal state, and the only execution state carried on the item. The vocabulary is closed: those two values are the whole of it, absent means the item is live, and any other state: value is a defect that quarantines the item out of the ready set rather than leaving it claimable. The state: tag vocabulary is the single definition of the enumerated set and of the verdict for an unrecognised value; Recording completion says when the terminal tag may be written. |
resource:<name> |
Optional, repeatable-by-name but one tag per distinct resource. Names a scarce non-file resource the item needs exclusively - a shared test box, a physical device, a deployment slot, a rate-limited external account. Two items naming the same resource may never be in flight together, however disjoint their code radii are. The resource: prefix is what makes the constraint load-bearing, because step 8 of the ready set matches on it. A bare descriptive tag asserting the same requirement - needs-box-exclusive, say - is read by no step and excludes nothing, however plainly it reads to a human. |
Five tags are mandatory on every item: backlog, priority:, phase:,
homeRegion: and baseBranch:. The rest are optional. state: is deliberately
not among them, because its absence is exactly what "live" means.
Exactly one tag per prefix. Two priority: tags on one item means two
authors wrote concurrently. Add-wins is what makes that visible rather than
silent, so it is reported as a defect and reconciled, never resolved by picking
one arbitrarily.
A missing mandatory tag is the more dangerous direction, because it is
silent. A duplicate is loud - two values sit where one belongs, and any reader
trips over it. An omission gives a reader nothing to see: the item is still
enumerated by the topic scan, still survives every narrowing step, and simply
never gets claimed. homeRegion: is the worst of them, because a worker filters
candidates to its own region, so an item naming no region matches no worker at
all. It looks healthy in every listing and starves indefinitely. Author the five
together, and detect the omission at step 1 of the ready set rather than
trusting authoring discipline to hold.
Keep attribute tags low-churn. OR-Set dots accumulate per add, so an
attribute rewritten every run would grow a long-lived item record without bound.
That is why per-run execution state - attempts, claims, leases, review rounds -
is deliberately not a tag (see below). The one exception is the terminal
state: tag, which is written at most once in an item's life and so costs a
single dot.
The state: tag vocabulary
This section is the single definition of the state: vocabulary. Both
agent prompts point here rather than restating the values, so there is one
place to change and the two cannot drift apart.
| Value | Recognised | Meaning | Ready-set verdict |
|---|---|---|---|
(no state: tag) |
yes | The item is live. | Admitted, subject to the remaining narrowing steps. |
state:complete |
yes | Terminal. The pull request is merged into its baseBranch, or the equivalent durable act has happened. See Recording completion. |
Dropped at step 2, silently. That is what the tag is for. |
state:parked |
yes | Terminal until a human re-admits it. See Parking an item. | Dropped at step 2, silently. |
any other state:<value> |
no | Undefined. The item's real condition is unknown. | Dropped at step 2 and reported as a defect. Never admitted. |
Match on the prefix state:, never on the two literals. A computation that
drops exactly the strings state:complete and state:parked does not treat an
unrecognised value as an unknown state. It treats it as no state at all -
byte-for-byte the same verdict it reaches for an item carrying no state: tag
whatsoever. The tag is then inert: the item reads live and is offered as
claimable, exactly as if nothing had been written. So the ready set partitions
every state:-prefixed tag into recognised and unrecognised, and both
partitions exclude. Only the reporting differs.
The unrecognised verdict fails safe, and the asymmetry is the whole reason the choice is not a judgement call. Excluding an item that is really live costs one missed candidate, it is named in the defect report, and it is repaired by a single write. Admitting an item that is really finished costs a worker an entire session redoing merged work, and it is silent: nothing in the store distinguishes a finished item read as live from one nobody has started. The two errors are not comparable, so the direction of failure is fixed here in advance rather than decided per reader.
It is reported, never silently absorbed. A quarantined item is named in the ready-set defect report together with its offending value, because an exclusion nobody sees is starvation - which is the failure the fail-safe verdict would otherwise trade into. The project manager reconciles the item to a recognised value - under a claim, like any other write - or removes the tag if the item is genuinely live.
The invented values already seen in practice. All four predate this section, and the earlier two-literal computation ignored every one of them:
| Invented value | What its author meant | How a two-literal computation reads it |
|---|---|---|
state:delivered |
the work is merged | live and claimable |
state:done |
the work is finished | live and claimable |
state:in-review |
a pull request is open and under review | live and claimable |
state:ready |
the item is live | live and claimable |
Three of the four assert that the work is finished or nearly so, and every one
of them read as work nobody had started. state:ready happened to fail safe -
it means live and was read as live - and that accident is precisely why an audit
that counts presence proves nothing. Presence looked healthy; validity did
not. Only the enumeration of distinct values answers the question that has
any bearing on behaviour.
A non-terminal condition is expressed by holding the claim, not by a tag.
This is the redirect for anyone reaching for a status value that is not in the
table above, and it is what lets the vocabulary stay closed at two. A live
fenced claim already means exactly "someone is working on this", and the ready
set already drops claimed items at step 2; attempts and review rounds are
carried by the claim markers. state:delivered and state:in-review are
attempts to record a real and meaningful condition that is already
representable - so hold the claim, or release honestly and write the resume
block. Do not invent a tag: the store accepts any string, so the invention
succeeds at the moment of writing and then does nothing.
Audit the vocabulary by enumerating distinct values, not by counting
presence. One scan of the backlog topic, project every state:-prefixed tag
to its distinct set, and diff that set against the two recognised values. It is
cheap, and it is the only form of the check that discriminates: a count of how
many items carry a state: tag is satisfied equally by a recognised value and
by an invented one.
Recording completion
An item is complete when, and only when, it carries state:complete. The
tag is the record; nothing else is. Without a defined encoding every worker
invents one, and the inventions do not agree - which is how an item ends up
tagged phase:complete (clobbering a reserved authoring attribute) or described
as finished only in prose inside body, where the ready-set computation cannot
see it.
Three rules make the tag trustworthy:
- It is written by the claim holder, under its fencing token, and only after
the thing it asserts is true. For an implementation item that means after
the pull request is merged into its
baseBranch, not after CI goes green and not after a review passes. Green-and-unmerged is not complete: the lifecycle transition isClaimed --> Complete, and an item tagged complete while its pull request is still open is a defect the next ready-set computation reports. An item that produces no pull request completes on the equivalent durable act, not on a weaker one. A research item's product is its findings, so it completes once those findings are recorded somewhere that outlives the item - the mirrored issue and durable memory - and never merely because the run ended. The item's ownbodydoes not count: it is a resume pointer rather than a deliverable, and a finding that exists only in the worker's context is lost the moment the session does. A design-integration item carries the further gate described under the grouping model: it may not complete while a grouping it emitted still lacks its dependency DAG. - It is terminal and it is not a status field. There is no
state:review, nostate:in-progress, nostate:blocked. Everything short of terminal is derived from state that already exists elsewhere and is authoritative there: blocked fromblockedBy, claimed from the claim surface, admitted from the mirrored issue's labels. Adding a mutable status tag would duplicate all three and churn the record besides. bodyexplains, the tag decides. The resume block says what landed and what is left, for a human and for the next claim holder. It is prose and is never parsed. No agent may infer completeness from it.
A worker that cannot merge - because it lost its claim, because review requested
changes, or because it ran out of run - does not tag the item complete. It
writes an honest resumeNote, posts an outcome comment carrying
result=released on the mirrored issue, and leaves the item live for the next
holder. result is a field of that outcome comment, never an argument to
repocontext_release_claim or any other tool. That is the normal path, not a
failure.
When the dispatch reserves the merge to the project manager
The rules above assume the claim holder merges its own pull request. A project manager may instead reserve the merge to itself - the right call when several items land on one shared integration branch and each merge changes what the next item is tested against, so the ordering is a project-manager decision rather than a worker one.
The reservation does not weaken the completion rule; it moves who satisfies it. Read it as a two-party sequence:
- The worker never tags the item complete, and that is not a failure path.
It is the same stand-down with an outcome comment carrying
result=releaseddescribed above, reached for a different reason: not that the worker could not merge, but that it was instructed not to. It writes an honest resume block, releases the claim under its fencing token, and reports plainly that the merge is outstanding and reserved. - The project manager tags the item complete once it has merged, under a
claim it takes itself. The tag still means "the pull request is merged into
its
baseBranch" and nothing weaker. A reserved merge is the one case where the party that satisfies that condition is not the party that held the claim while the work was done, which is exactly why it has to be written down: an unstated exception to a rule this load-bearing is indistinguishable from a worker getting it wrong. - A dispatch that reserves the merge must say so explicitly. A worker whose instructions and this protocol disagree should honour the dispatch and report the divergence rather than silently pick one, because a divergence is far more often a project-manager omission than a deliberate choice.
The failure this prevents is a genuinely finished item sitting live in the ready set - offering itself to a second worker who would redo the work - because the only party permitted to tag it was forbidden to perform the act that defines the tag.
A directly-deployed item still gets a ledger row, authored at completion
Not every item reaches a worker through the backlog. A project manager may deploy a worker directly against a GitHub issue, and should: it is the right move for a one-off defect that needs no dependency ordering and no ready-set arbitration. Such an item is never claimed, holds no lease and no fencing token, and never appears in a ready set. Nothing about that is wrong.
What is wrong is leaving no trace. The project manager authors a ledger row
for a directly-deployed item when it completes, carrying the five mandatory
tags, state:complete, pr:<n>, and direct-deployment.
Three rules make the row safe:
- Author it at completion, never while the work is in flight. A live row for work already underway is claimable: nothing in the store distinguishes it from work nobody has started, so the next ready-set computation offers it and a second worker can be dispatched onto an item that is already half done. Authoring at completion means the row is born terminal and can never be selected. This is the whole reason the row is retroactive rather than eager, and it is not a compromise.
- Record what happened, and nothing more. State plainly that the item was directly deployed, never claimed, held no lease, and never entered a ready set. Do not synthesise claim or lease history for a run that took none. The distinction is worth being precise about, because the scruple that stops an agent writing the row at all is a good instinct pointed at the wrong target: fabricating a claim history would be invention, whereas a row saying "completed via direct deployment, never claimed, closed by PR #N" asserts nothing untrue. Truthful and retroactive is not the same as invented.
- The project manager writes it, not the worker. Items are a project-manager
artefact - the lifecycle opens
[*] --> Drafted: authored by the project manager- and a worker authoring its own ledger row inverts that at the moment it is least able to be checked, as it stands down. A worker that notices it has no item to complete should report the absence rather than do something claim-shaped.
Why the gap is worth closing at all, given that the ready set already fails
safe here: a missing target is reported as a dangling blockedBy defect and is
never treated as satisfied, so no work is silently released. The cost is subtler
and worse. exists: false reads identically for a benign ledger gap and for
an item deleted while it was still gating work, and the second is precisely the
failure that rule exists to catch. Every uncreated row makes the defect signal
less able to mean anything. A complete ledger is what keeps the alarm credible,
and the throughput the ledger reports honest.
Evidence a worker may rely on
The rules above say when a worker may assert completion. This one says what makes the assertion admissible. A worker that reports a measurement its own instrument fabricated is neither lying nor careless: it has simply never been told that an instrument is a thing which can itself be wrong, and a number is the most persuasive object an agent can put in front of a reviewer.
Validate an instrument against a case with a known non-zero answer before you trust a zero from it. A zero from an unvalidated instrument is indistinguishable from a broken instrument, and it is the most expensive kind of wrong answer because it looks like a result rather than a failure. Find a case whose correct outcome requires the quantity to be non-zero, run it, and check the instrument agrees. That is one cheap run. In this protocol's first live use, a worker instrumented an exception path, measured zero, and reported a hypothesis refuted. The probe wrote to standard output, which the test runner captures and surfaces only for failing tests, so on every passing run it reported zero by construction. The true count on a single case was 117. By then the project manager had already recorded the refutation as durable fact.
A measurement taken over passing runs says nothing about the failing run. The failing run is by definition the one where behaviour differed. Reset per run, capture per run, keep the failure. This is independent of the rule above and does not substitute for it: the same worker moved its measurement onto the failing run and still got a fabricated zero, because the instrument was unchanged.
Cite the artifact, not the impression. Evidence is a machine-readable per-case result that survives whatever verbosity the run happened to use. A console tally is not evidence: the same worker captured "1 failed, 3 passed" on a four-case fixture and could not say which case had failed, and had to reproduce an event that had already happened.
Say which claim you are making. "Observed failing, changed, observed passing" and "correct by construction and not observed to fail since" are different claims of different strength. Both are legitimate and the second is often the best available. Blurring them is not legitimate, because a reader who is not told which one you mean will assume the stronger.
The project manager carries the mirror of this obligation. Evidence is not authoritative merely because it is numeric, and a durable memory written from an unvalidated measurement propagates one worker's broken instrument into every later session that reads it. Ask what the instrument was and whether it was checked, prefer a retraction to a defence, and correct the durable record the moment a measurement is withdrawn.
The item body
body holds a pointer to the mirrored GitHub issue - not a copy of its
specification - plus the resume block for the most recent attempt:
lastLocation: branch / pull request number / sha of the last attempt.resumeNote: a short "what is done, what is left".
body is an LWW register, and that is safe here only because these fields are
written exclusively by the current fenced claim holder, so there is never more
than one writer. LWW is not unsafe in general; it is unsafe when unserialised.
The fenced claim is what serialises it. Nothing else may write body while a
claim is live.
The resume block is advisory. A resuming worker re-decides from it and never continues blindly, because an abandoned run leaves the branch behind but not the reasoning that produced it.
What is deliberately not on the item record
attemptsis derived from the mirrored issue's claim-comment trail, not stored. GitHub already owns the audit trail, counting comments needs no reverse index, and a per-attempt counter on the item record would be exactly the unbounded-churn write the OR-Set dot cost warns against.- Claims, leases and fencing tokens live in the fenced claim/lease surface and on short-lived per-run worker records, never on the item.
Never set a TTL on a backlog item. Expiry is silent and unlogged, so a
lapsed item that other items declare blockedBy starves its dependents
invisibly, with no event anywhere to explain it. Retire an item deliberately
with forget. This is no longer an exception to the coordination rule: the
Coordination
section now forbids a TTL on coordination entries generally, on the same
reasoning. A backlog item - a ledger entry, not a handoff - is simply the case
where the harm is most concrete. Omitting ttlSeconds is not enough on its
own: repocontext_remember gives a newly created entry the repository's default
memory TTL whenever the host configures one (RepoContextTtlOptions.DefaultMemoryTtl,
unset by default), and ttlSeconds accepts only a positive value, so no call
can opt an item out. A repository that hosts a backlog leaves that default unset.
Recording baseBranch: on the item is what makes a retry land correctly.
Leaving it to worker convention means a resumed or reassigned attempt targets
whatever the worker assumes, which for a sub-item of an epic is usually main -
exactly the case the epic branch exists to avoid. An item that is partOf an
epic and carries baseBranch:main is a defect, reported rather than
silently accepted.
Cross-checking attempts against the fencing token
The claim-marker count is only as good as the workers who post markers, and a
missing marker produces no signal: an item drawn six times with no marker reads
attempts = 0, exactly like an item never drawn, so the poison threshold is
unreachable. Every sweep that derives attempts therefore also reads the item's
fencingToken from repocontext_claim_status. The claim grant itself advances
it, so it cannot be forgotten; it is absent until the first grant.
The token is an upper bound, not an attempt count. A same-owner re-claim
after a lease lapse advances it without being a new attempt (see
Detecting and picking up a dropped lease), and a renew that omits
leaseSeconds makes such lapses common. So it cross-checks the marker count and
never replaces it:
fencingToken |
Claim markers | Reading | Action |
|---|---|---|---|
| absent or 0 | 0 | never claimed | none - attempts is 0 |
| > 0 | 0 | claimed but unmarked | defect - report it; attempts is unknown, not zero |
| = markers | > 0 | consistent | use the marker count as attempts |
| > markers | > 0 | token ahead of the trail | use the marker count as attempts, and report the gap: markers were dropped, or claims lapsed |
| < markers | > 0 | trail ahead of the token | defect - report it; a marker with no grant behind it |
A sweep never reports attempts = 0 for an item whose token is above zero.
Relation vocabulary - the backlog extension
These extend the small, stable
knowledge-linking vocabulary
rather than competing with it. partOf and related are the documented
relations used unchanged, and the five additions follow the same discipline:
few, stable, one direction authored, named for what they assert.
They are documented here so tooling that audits memory - the daily Memory Accuracy automation in particular - recognises them and does not prune them as unknown relations.
| Relation | Authored on | Points at | Meaning |
|---|---|---|---|
blockedBy |
the dependent item | item keys | Every target must be complete before this item is claimable. |
anchoredTo |
the item | file / symbol keys | The code this item concerns. Gives digest-drift staleness for free. |
claims |
a per-run worker record | the item | This run asserts ownership of the item. |
partOf |
the sub-item | the epic item | Grouping membership. The documented relation, used unchanged. |
integrates |
the integration item | the epic it closes out | Marks exactly one item per grouping as its integration join. |
informs |
a research grouping | the implementation grouping it produced | Keeps the rationale behind a decomposition discoverable from the work it caused. |
related |
either | items, gotchas, decisions | Near-duplicate items, and the learnings a prior attempt produced. The documented relation, used unchanged. |
Two rules follow from the store's semantics rather than from taste:
claimslives on a short-lived per-run record, never on a long-lived one. OR-Set dots accumulate per add, so an edge asserted and released every run grows a long-lived record without bound.- Edges make a collision detectable, not preventable. There is no
compare-and-swap anywhere in this surface:
repocontext_updatepreconditions on record existence and, once the record is claimed, on the fencing token - never on a field's value. Aclaimsedge is therefore an audit record of who tried, not a lock. Mutual exclusion comes from the fenced claim/lease surface, whose monotonic fencing tokens and bounded, expiry-reclaimed leases give real exclusion and a real stale-claim reaper.
Parked blockers and the ruling route
This is the single definition for authoring, dependency reporting, and
parking. A parked item awaits a human ruling, not scheduled agent progress.
"Stalled" means never ready without human intervention, not "not ready
yet"; it is a derived report, never a new state: tag.
Before any repocontext_remember or repocontext_update that adds
blockedBy, recall every proposed target in its home region and apply the
table below. Validate the whole proposed write before sending it. If any
target is rejected, make no write (including unrelated fields or other edges
in that request) and report a hard validation error naming the dependent and
each rejected blocker key. A named ruling route does not make a parked target
an acceptable new dependency. Do not silently omit the edge and create an
apparently independent item.
These are agent preflight rules, not a claim that the generic memory API interprets backlog tags or provides a cross-record transaction. A target may be parked after the read or after a previously valid edge was written, so authoring validation alone is insufficient.
| Target observation | Authoring verdict | Dependency classification | Required report |
|---|---|---|---|
| missing | reject | invalid | dependent, blocker |
| invalid | reject | invalid | dependent, blocker |
| parked | reject | stalled | dependent, blocker, ruling route |
| live | allow | waiting | dependent, blocker |
| complete | allow | satisfied | none |
The predicates are disjoint: missing means exists: false; otherwise
invalid means multiple state: tags or any unrecognised value; parked
and complete mean exactly their one recognised tag; live means no
state: tag. A live blocker may be claimed or awaiting a claim; neither is a
terminal state tag. The completion-evidence defect checks still apply.
Every ready-set computation and PM sweep evaluates existing edges before
candidate narrowing, even if other work is ready. Page the entire backlog
scan, retain each dependent's outbound blockedBy keys, and recall any
targets not resolved by the scan; a truncated neighbors result is not a
complete dependency set. Classify every edge, including those whose
dependent is filtered out later. Report every stalled edge in the listing
with both keys, the blocker's ruling owner, and the question/issue pointer.
Report a missing ruling route as an additional defect, never as ordinary
waiting. For a dependent with mixed blockers, invalid takes precedence over
stalled, then waiting, then satisfied; retain all edge diagnostics.
Only all-satisfied (including no dependencies) survives dependency narrowing.
Do not delete an edge, tag a blocker complete, or unpark it to make work ready.
Parking requires a named route out before writing state:parked. Record
exactly one non-empty ruling-owner:<role-or-person> tag and exactly one
non-empty ruling-topic:<decision-key> tag alongside it in the same fenced
memory write. In the resume note and mirrored issue comment, name the exact
question that owner must rule on and the issue where they can answer it.
pm-ruling-required and needs-design-decision may remain descriptive tags;
they do not replace this explicit route. For example, ruling-owner:pm and
ruling-topic:design-decision must point to the actual undecided design
question, not merely say "needs a decision".
| Ruling owner named | Ruling question named | Parking verdict |
|---|---|---|
| no | no | reject |
| no | yes | reject |
| yes | no | reject |
| yes | yes | allow |
Here "named" requires the corresponding unique, non-empty tag and the concrete owner/question recorded in the resume note and issue comment. Reject an incomplete route as a hard validation error naming the item and missing owner/question. Existing parks without a route remain excluded and are reported to the PM for repair, never automatically unparked. After a human ruling and re-admission, remove the parking and route tags under the claim and recompute dependencies; unparking a blocker makes its dependents waiting, not satisfied, until that blocker actually completes.
Why anchoredTo matters
Linking an item to the files it concerns captures those targets' content digests
at link time, so repocontext_recall reports the item stale once the code
drifts - measured against the index, so an edit counts once it is re-ingested.
An item whose anchor moved auto-flags "re-validate the spec before
spending a run on it". This is the one capability GitHub issues cannot provide,
and it doubles as the poison-item mitigation. An anchor whose target is not in
the index - deleted, or not created yet - also reads stale, and so does an
edge written before its target was indexed: it captured no digest, and stays
stale until the edge is written again. recall also lists a target that has no
live record under danglingLinks as well as staleLinks, which separates an
anchor pointing at nothing - deleted code, or code that has not reached the
indexed branch yet - from one whose code has moved on.
Combined with repocontext_related, anchors also give each item a blast
radius, so two items touching the same code can be serialised at selection
time rather than colliding at merge time. Selecting for disjointness is the
primary throughput mechanism; a concurrency cap is only a backstop for an
unavoidably overlapping ready set.
The grouping model - three phases
A grouping is a set of items delivered together: an epic and its sub-items,
joined by partOf edges from sub-item to epic. A grouping runs in up to three
phases.
- Research and design (optional, for a large or uncertain epic). One item
per research area, fanned out to research agents. Research items produce
memory entries, docs and proposals rather than code, so their blast radius is
empty and they parallelise perfectly. The phase terminates in a
design-integration item that reconciles the findings and emits the
implementation grouping, linked to it with
informs. - Implementation. Seam-first fan-out: land the contract as one small fast
item, then fan out implementations against it. Prefer wide DAGs to deep
chains - a
blockedByedge that exists only because of how the work was described is not a real dependency. - Integration. The close-out item described below.
flowchart TB
subgraph P1["Phase 1 - research grouping (optional, leaf: never nested)"]
direction TB
RA["Research area A"]
RB["Research area B"]
RC["Research area C"]
RI["Design integration<br/>reconcile findings, emit grouping"]
RA --> RI
RB --> RI
RC --> RI
end
subgraph P2["Phase 2 - implementation grouping"]
direction TB
SEAM["Contract / seam item<br/><i>small, fast, unblocks everything</i>"]
F1["Fan-out A"]
F2["Fan-out B"]
F3["Fan-out C"]
SEAM --> F1
SEAM --> F2
SEAM --> F3
end
RI -->|informs| SEAM
F1 --> INT
F2 --> INT
F3 --> INT
INT["Phase 3 - integration item<br/><i>exclusive claim, others quiesced</i>"]
INT --> DONE(["Epic closed"])
classDef free fill:#dbeddb,stroke:#2d7a2d,color:#0b2e0b
classDef excl fill:#f6e3c5,stroke:#a8721a,color:#3a2606
class RA,RB,RC free
class INT,RI excl
Green items have empty or disjoint blast radii and run concurrently without restriction; amber items are exclusive joins.
Termination rule, and it is load-bearing: a research grouping does not itself
get a research grouping. It is a leaf phase. Without this rule an agent asked
to plan an epic can recurse indefinitely into planning the planning. An item
tagged phase:research may not author a further research grouping; whatever it
emits is an implementation grouping. Research is also not the default - where
the shape of the work is already understood, a research phase is pure
critical-path depth.
A phase:research item's deliverable is a durable memory entry plus an issue
comment - not a branch, and not a pull request. State this when dispatching
one, because the default assumption of a worker built to ship code is that it
must produce a diff. Three consequences follow:
- It does not need a branch at all, which sidesteps the branch-naming rules
above entirely. A session whose workspace was auto-provisioned with a
generated branch name - frequently one carrying a username and no
<type>/prefix, both of which many repositories forbid outright - simply never pushes it, and the non-conforming name never reaches the remote. - A well-evidenced negative is a completed item, not a failed one. Say so at dispatch. A research item exists to be capable of killing the work that would otherwise follow it, and a worker that believes a negative reflects on it will reach for an encouraging maybe. The cheapest possible outcome of a research phase is discovering early that the implementation phase must not be built.
- It must not touch the implementation surface. A research item that edits
src/has silently become an implementation item without being admitted as one, and its changes bypass the grouping its findings were meant to shape.
If a research item genuinely must produce a file, that is a signal it was mis-scoped as research - and the file needs a conforming branch arranged deliberately, rather than an auto-generated one pushed by default.
The integration item
Every grouping terminates in exactly one designated integration item, which is
blockedBy every fan-out item in the grouping and carries an integrates edge
to the epic it closes out.
It exists to absorb the risk that maximum parallelism creates: N pull requests, each green in isolation against a different base, none ever tested against the others. The failure it catches is not a merge conflict (those are visible) but the epic passing every sub-item's acceptance criteria while failing its own. Its remit is therefore conflict reconciliation, a full cross-package test run rather than the per-package targeted runs the sub-items ran, and verification against the epic's acceptance criteria.
Three rules attach to it:
- It is exclusive. It spans the grouping's whole blast radius by design, so it cannot be selected for disjointness like a normal item. It requires an exclusive claim with the grouping's other workers quiesced. This is the one deliberate exception to the disjointness rule.
- A grouping is not complete until its integration item is complete. An epic cannot be closed by its sub-items alone, however green they are.
- A design-integration item may not complete while the grouping it emitted lacks a mermaid dependency DAG, and the gate applies transitively to anything those groupings go on to emit. A generated grouping is held to exactly the standard a hand-authored one is. That is the case that matters most, because a human has least visibility into a decomposition an agent assembled, so the obligation must not be launderable through a layer of automation.
Branch inheritance
An epic gets one shared branch and its sub-item pull requests target that
branch; the epic reaches main as a single fully-gated pull request once its
integration item passes. Concretely:
- the epic record carries
baseBranch:<type>/epic/<epic-slug>; - every sub-item inherits that value as its own
baseBranch:tag; - sub-item branches are named
<type>/epic/<epic-slug>-<item-slug>; - an item that is
partOfan epic and carriesbaseBranch:mainis reported as a defect.
The final separator is a hyphen, not a slash, and this is forced by git rather
than chosen. An earlier revision of this document prescribed nesting sub-items
as <type>/epic/<epic-slug>/<item-slug>. That form is unimplementable
whenever the epic branch is parked on the bare slug - which the first rule above
also mandates - because git stores a branch as a file at refs/heads/<name>, so
refs/heads/X and refs/heads/X/anything cannot coexist. The remote refuses it:
cannot lock ref 'refs/heads/fix/epic/my-epic/my-item':
'refs/heads/fix/epic/my-epic' exists
This is a directory/file ref conflict, not a policy or permissions failure, and no naming choice on the sub-item's side avoids it. It was found by a worker attempting the push, having been reviewed twice in prose without either reader noticing - reading a ref name does not tell you git will refuse it.
Note the corollary, because it is the part that gives false assurance: a CI branch-name guard will happily accept the nested form, since it is a well-formed lower-case name under an allowed prefix. A guard that validates a string the underlying system then rejects is worse than no guard on that dimension, because it converts "unverified" into "verified" without adding verification.
The alternative - parking the epic on <type>/epic/<epic-slug>/integration and
leaving the namespace free for true nesting - does work, and is the better shape
for a grouping created from scratch. It is not adopted as the default because it
costs a rename of the epic branch and a rebase of every in-flight sub-item if
adopted mid-grouping. Choose it at epic-creation time or not at all.
Two invariants matter more than the name, and are what a reviewer should actually check. The name is a convenience; these are correctness:
- the sub-item branch is descended from the epic branch -
git merge-base --is-ancestor origin/<epic-branch> HEADexits0; - the sub-item's pull request targets the epic branch, never
main.
A sub-item that satisfies both under an off-convention name is fine and is reported as a naming nit. A sub-item that satisfies neither under a perfectly conventional name has silently bypassed the epic, and its work will not be collected by the integration item.
Do not use a workspace rename_branch affordance to satisfy this rule without
checking its output. In at least one environment it applies a configured
prefix that injects a username and omits the <type>/ prefix entirely,
producing a name this repository forbids outright and which fails the CI guard.
Rename the branch directly and verify the resulting name.
Computing the ready set
The ready set is the items claimable right now. It is always computed as a topic scan plus per-candidate depth-1 checks, and never as a single graph query.
Why it cannot be one call. repocontext_neighbors is navigation, not query:
it walks outbound edges only, with depth clamped to [1, 3] and
maxNodes to [1, 100]. There is no reverse index over memory links - the
reverse cross-reference index serves repocontext_related for symbols only -
so "who is blocked by me?" and "what did completing X unblock?" cannot be asked
directly. They require either an explicitly authored inverse edge or a topic
scan. Do not design a protocol around a reverse lookup this surface cannot
serve. Scan-plus-check is fine at hundreds of items; this is a coordination
graph, not a queue engine.
The computation:
repocontext_scanscopeMemoryTopic, topicbacklog, paging on the continuation token, to enumerate every live item. Verify each item carries the five mandatory tags as you page, and report any item that does not rather than silently passing over it. The check is free here, because every item is already in hand, and this is the only step that sees all of them - so an item malformed in a way that hides it from the later narrowing steps is caught here or not at all. Before narrowing candidates, validate parking routes and classify all existing dependency edges using Parked blockers and the ruling route; retain the stalled and invalid diagnostics even when the ready set is non-empty.Match every
state:-prefixed tag against the closed vocabulary in Thestate:tag vocabulary, and drop the item on both outcomes. A recognised terminal value -state:completeorstate:parked- drops it silently, because that is what the tag is for. Any otherstate:value drops it too, and is reported as a defect: it is an unknown state, not an absent one, and admitting it is how a finished item gets offered to a second worker. Do not match on the two literals alone - that reads an unrecognised value as no state at all and admits it. Also drop items held under a live fenced claim (repocontext_claim_status). Completeness is read from the tag and from nothing else - never from prose inbody, and never from a merged-looking pull request.Drop grouping records. A grouping (an epic, or any item that other items declare themselves
partOf) is a container, not a unit of work. It is completed by its integration item, never claimed directly. Omitting this step lets a worker claim the epic itself and duplicate the entire fan-out that the decomposition just created.Build the exclusion set during the step-1 scan, at no extra cost: collect the target of every
partOfedge you encounter as you page through the topic. Do not attempt this as a reverse lookup - "who ispartOfme?" is exactly the reverse-index query this surface cannot serve, which is why the check has to be a by-product of the enumeration rather than a per-candidate probe.The same conclusion can be reached from the data alone, and belt-and-braces is cheap here: a grouping should also carry
blockedByits own integration item, which drops it at step 4 anyway. Author both. The redundancy is one-way safe - it can only ever remove a container from the ready set, never admit one.For each remaining candidate, apply the dependency classifications from step 1 per Parked blockers and the ruling route. Only all-satisfied survives; ordinary waiting, stalled, and invalid dependencies stay out of the ready set.
Drop survivors whose mirrored issue is not admitted (see Entry gating). This is checked after the
blockedBynarrowing, so it costs one issue read per survivor rather than one per item in the topic.Sort by
(priority, createdAt, id), then pick from the top three to five. Ordering deterministically is fine and is not a defect:repocontext_claimis real mutual exclusion, so two workers converging on the same item resolve to exactly one proceeding and the other observing a clean refusal it can act on immediately. Jitter is a cheap way to spread the fan-out across candidates and avoid spending a round on a refusal, so it remains worth applying - but it is an optimisation, and no worker may rely on it for correctness.Prefer a candidate whose blast radius - its
anchoredToanchors plusrepocontext_relatedon them - is disjoint from the radii of in-flight items.Exclude, rather than merely deprioritise, a candidate whose
resource:tags collide with an in-flight item's. Step 7 is a preference computed overanchoredTofiles, so a scarce non-file resource is invisible to it: an item whose real constraint is "needs exclusive use of the shared test box" may have an empty code radius and will therefore look maximally disjoint and sort to the front. That is the exact inversion of the truth. Resource collision is a hard exclusion, not a tie-break, because the failure it prevents - two agents recreating the same container under one another - is not a merge conflict that surfaces loudly but a corrupted experiment that reports a plausible wrong answer.
A scan is a bulk read and therefore does not evaluate TTL or link
staleness: stale and staleLinks come back null there, meaning "not
evaluated" rather than "not stale". Staleness must be read with recall on the
specific candidate.
Defect conditions the ready-set computation must surface
These are reported, never silently absorbed:
Stalled by a parked blocker, or parked without a ruling route. Apply Parked blockers and the ruling route before narrowing, report both item keys and the human ruling needed, and keep stalled distinct from ordinary waiting even when other work is ready.
Dangling
blockedBy. A target that returnsexists: falseis a defect, not a satisfied dependency. Treating an absent blocker as complete is how a deleted item silently releases work that was deliberately gated on it.Stale item. An
anchoredTotarget drifted, sorecallreports the itemstale. Re-validate the spec before spending a run on it.Claimed but unmarked. A
fencingTokenabove zero with no claim marker, or a token that disagrees with the marker count in either direction. With no marker,attemptsis unknown, not zero, so the item is not fresh. See Cross-checkingattemptsagainst the fencing token.Duplicate attribute tag. Two tags sharing a
key:prefix means two concurrent authors. Reconcile; never pick one arbitrarily.Unrecognised
state:value. Astate:-prefixed tag whose value is neithercompletenorparked. Step 2 quarantines the item out of the ready set and names it here with its offending value, because the tag asserts a condition the protocol cannot interpret, and the safe reading of an uninterpretable assertion is not "live". Reconcile it to a recognised value under a claim, or remove the tag if the item really is live. See Thestate:tag vocabulary.Execution state on a
phase:tag.phase:carries the authored phase and nothing else, sophase:completeorphase:reviewmeans a worker wrote a status into a reserved attribute - and, because add-wins never replaces, the item's real phase is either lost or now duplicated. Reconcile to the authored phase plus astate:tag if one is warranted.Item tagged
state:completewith an unmerged pull request. Completion was claimed before the merge that defines it. The item is not complete; the merge is outstanding work.Green, mergeable pull request on an item with no live claim and no
state:complete. The attempt died between CI passing and the merge. This is the cheapest possible resume and should be picked up before any fresh item.Ready set empty while pending is not. There is no cycle detection in the store, so a dependency cycle is silent permanent starvation. Alarm rather than exit quietly.
Ready set empty and pending empty. Exit immediately. Every tick otherwise spends a whole session for nothing.
baseBranch:mainon an item that ispartOfan epic. See branch inheritance above.A grouping whose fan-out is complete but whose integration item is not. The grouping is not complete; do not close the epic.
Item lifecycle
stateDiagram-v2
[*] --> Drafted: authored by the project manager
Drafted --> Gated: mirrored to a GitHub issue
Gated --> Ready: admitted (human, or human-authored at source)
Ready --> Blocked: a blockedBy target is incomplete
Blocked --> Ready: every blocker completes
Ready --> Stalled: a blocker is parked
Blocked --> Stalled: a blocker is parked
Stalled --> Blocked: human re-admits blocker, work still pending
Stalled --> Ready: every blocker completes
Ready --> Quarantined: an unrecognised state tag value is present
Quarantined --> Ready: reconciled to a recognised value, under a claim
Ready --> Claimed: fenced claim acquired (homeRegion only)
Claimed --> Ready: lease expires, or the worker releases
Claimed --> Complete: PR merged into the base branch, or equivalent durable act
Claimed --> Parked: attempts exceed the poison threshold
Claimed --> Parked: the holder parks it deliberately
Parked --> Ready: a human respecifies and re-admits
Complete --> [*]
Claimed --> Ready on lease expiry is the normal path, not an exception. Stale
claims are the common case, so a claim is always lease-bounded and reclaimed on
expiry rather than held by a flag that a killed session leaves set forever.
Parking an item - both sides, and not only on exhaustion
Parking is encoded on both sides, and the tag is the half that has effect.
The ready set drops parked items at step 2 by reading the state:parked tag
on the item record; the stale label on the mirrored issue is what makes
the park visible to a human. A park that writes only the label is not a park:
the item survives every step of the ready set and is claimable again on the next
tick, so the guard silently does nothing. Write both, and write them under the
fencing token of the claim you hold, in the order tag then label - if the run
dies between them the item is already out of the ready set and the sweep can
finish the visible half.
First validate and record the ruling route. That section also governs dependents stranded by this transition and existing parks missing a route; neither may disappear as ordinary "not ready" work.
The poison threshold is three, and it is a floor on parking, not the only route to it. An item whose claim-marker count has reached three is parked by whichever worker takes it there. But a holder that establishes, at any attempt number, that the item cannot proceed as specified parks it deliberately and does not wait to burn the remaining attempts. The two cases that matter:
- The work is blocked on a decision, a defect, or a dependency that is not
itself an item, so no
blockedByedge can express it. - The specification is wrong, not merely hard - the item as written cannot be satisfied.
Releasing instead is the failure mode this exists to prevent. Releasing the claim
and posting an outcome comment carrying result=released on the mirrored issue
leaves the item live and immediately re-claimable, so the next worker draws it,
re-derives the same finding, and releases in turn; the fleet spends a session per
tick relearning one conclusion. A deliberate park costs one attempt and states
the conclusion once.
Say why, in both places a later reader will look. The resumeNote carries
the finding for the next holder, and the outcome comment carrying result=parked
on the mirrored issue carries it for the human who must decide. A park with no
stated reason is indistinguishable from a crash and will be unparked without the
finding being addressed.
Parking is safe to get wrong in one direction only. It can only ever remove
an item from the ready set, and Parked --> Ready requires a human, so an
over-eager park costs a human glance while an omitted one costs an unbounded
loop. Park when in doubt. Parking is a finding, not a failed attempt, and a
worker that parks correctly on its first attempt has done its job.
The lease is shorter than the work - renew before, never after
The cluster applies a short default lease when leaseSeconds is omitted -
commonly 30 seconds - and clamps every request to a host-configured ceiling.
Do not assume either figure: read the leaseSeconds and leaseExpiresAtUtc your
grant actually returns, because a request above the ceiling is clamped silently
and a deadline diaried from the length you asked for is already late. A
build-and-test cycle on a non-trivial repository exceeds a 30-second lease many
times over. The consequence is not hypothetical and was observed on the first
live run of this protocol: two independent workers each had a claim lapse
mid-build, while actively working the item.
The short default is deliberate, and claim and renew_claim apply it
identically. There is no divergence between the two surfaces - this was
measured on a live deployment, both arms with leaseSeconds genuinely omitted,
and both granted the same length. The rationale for keeping the default short is
that a caller which did not name a lease length is exactly the caller that should
not be granted a long one. A host that needs longer claims raises the ceiling
an explicit request may reach, not the default.
Both recovered correctly - repocontext_claim_status showed no other holder and
no queue, and the re-claim returned a fencing token incremented by exactly one -
so the mechanism behaved as designed. The gap is that the lease duration is
shorter than the shortest useful unit of work, which turns a safety property
into a routine occurrence.
Why that matters more than a retry: during the lapse the item is, to any other
worker computing the ready set, simply available. Step 2 drops items "held
under a live fenced claim", and nothing holds this one: the lease has lapsed, so
repocontext_claim_status reports isHeld: false and a repocontext_claim would
be granted. The record still reports claimed: true with the last fencing token,
which is what separates a lapse from an item nobody has started (claimed: false,
no fencingToken), but a ready set that goes by the lock cannot see the
difference. A sibling recomputing in that window would
have found the item available and begun duplicate work on an item another worker
was mid-build on. Nothing prevented that. Only the timing did.
Rules, in force for every worker:
- Always pass
leaseSecondsexplicitly - onrenew_claimas well as onclaim. The short default is shorter than almost any real operation and will lapse under a single test run. On a renew the omission is worse than on a claim: a renew that omitsleaseSecondsresolves to that same short default and therefore shortens a claim you are currently holding for longer. It still reportsgranted: true; the loss surfaces only on the next renew, asgranted: false, reason: "superseded", at a call site that did nothing wrong. A renew that shortens its lease is flagged byleaseShortenedin the result - and note that anullthere means the prior lease could not be read, so it is "unknown", never "nothing shrank". - Renew immediately BEFORE any long operation, never after it. Treat a build, a test run, or anything expected to exceed roughly two minutes as requiring a renewal first. Renewing afterwards is renewing during the window you needed to be covered for.
- On discovering a lapsed claim, re-claim and then CHECK THE FENCING TOKEN. If it incremented as expected and the holder is you, that is a clean re-claim; proceed, and report it. If the holder is not you, or the token did not move as expected, stop and report - somebody else has been working the item, and continuing would produce two divergent attempts at one unit of work.
- Never write anything under a token you know to be stale.
Fixes worth making to the surface itself, in preference order: raise the
host-configured ceiling above a realistic build time (many hosts already do -
check what your grants actually return before assuming otherwise), or make it
per-phase, since a research item and a build item have very different natural
durations; auto-renew on a timer for the lifetime of a long child process rather
than asking a worker to predict its duration; and distinguish "lease expired
while work was in progress" from "never claimed" in the ready set, so a lapse
degrades to a warning rather than to availability. That last needs no change to
the surface, only to the ready-set computation: repocontext_claim_status
already separates the two (claimed: true with isHeld: false, against no
fencingToken at all).
This was surfaced only because a worker volunteered an unflattering detail it had already recovered from. A protocol that discourages that reporting would have shipped this gap silently.
Detecting and picking up a dropped lease
Raising the lease only makes a lapse rarer. It does not say what a lapse means or who may act on it, and that is the part that has to be specified, because the store cannot answer the only question that matters.
The central difficulty: a lapse has two causes and the surface cannot tell them apart. An expired lease means either
- the holder is gone - crashed, killed, context-exhausted, session ended - and the item genuinely needs picking up; or
- the holder is alive and working, and merely failed to renew in time.
Both present identically: no live lease (isHeld: false) on a record that still
reports claimed: true. Nothing in the lock, the item record, or
the ready set distinguishes them, and there is no liveness signal independent of
the renewal itself. Treating every lapse as case 1 duplicates live work; treating
every lapse as case 2 leaks items permanently to dead agents. Neither default is
safe, so the protocol makes the distinction unnecessary rather than pretending
to resolve it.
Detection is pull, not push. Renewal IS the liveness probe. A worker is never notified that its lease expired; it finds out only by attempting a renew (or a fenced write) and being refused. There is no callback and no interrupt. A worker that never renews never learns it was evicted, and will keep working - which is precisely case 2 seen from the inside. This is why renewal is mandatory before long operations rather than merely advisable: it is the only mechanism by which a worker discovers it has lost the item.
What fencing does and does not protect, which is the load-bearing point. The monotonic fencing token makes store writes safe: a superseded worker's write is rejected, so two workers can never both mutate the item record. It protects nothing else. Git, GitHub, the filesystem and any deployed environment are outside the fence. A superseded worker can still push a branch, open a pull request, comment on an issue, or recreate a container, and none of those will be refused on account of a stale token.
Therefore:
- Renew immediately before every externally-visible side effect, and verify the token, not merely that the call succeeded. Push, pull-request creation, issue comments, and any environment mutation are all gated on a fresh, verified renewal. A renewal that returns is not enough; the token it reports must be the one you hold.
- On refusal, abort without side effects. Do not push "just this branch", do not open the pull request, do not comment. Report and stop.
- Never destroy your own work on discovering you were superseded. The branch and commits from an evicted attempt are the takeover's most useful input. Leave them, and say in your report exactly where they are. Deleting them converts a recoverable handover into a restart.
A lapsed item is quarantined before it becomes claimable. It does not re-enter the ready set the instant the lease expires. It becomes eligible only after a quarantine interval that comfortably exceeds the longest plausible renewal gap - one full lease is the working default. This is what buys the distinction the store cannot make: an alive-but-late holder reclaims its own item inside the quarantine and continues (its fence increments, nothing else changes), whereas a genuinely dead holder never does, and the item is released to others only after that window closes. The cost is bounded latency on genuine failures; the benefit is that the common case stops being a race.
One full lease can be the wrong quarantine, and elapsed time is the wrong
evidence. The working default above assumes the lease approximates the work.
It often does not, and the gap depends on a value you must measure rather than
assume. MaxLockLeaseDuration is a configured ceiling whose shipped default is
300 seconds, against turns that routinely run for hours; where it stands at that
default, a live worker's claim spends almost all of its life presenting as
lapsed. This was observed on the first real run - a productive worker sat at
fence 12, mid-implementation, while claim_status reported isHeld: false and
the item showed no unmet blockers. To any agent computing a ready set it was
indistinguishable from abandoned work, and the lock would have granted it on
request.
Do not read the 300-second figure as the value in force. It is a default, not a
constant, and a deployment may raise it: on the reference deployment a claim
requesting 1800 seconds was measured being granted 1800 seconds, so the
ceiling there is at least that. The bundled RepoContext container host, for one,
raises it to 1800 seconds by default through LATTICE_MAX_LOCK_LEASE_SECONDS
(accepted range 30 to 7200). Read the leaseSeconds your own grant returns
and reason from it. A quarantine measured in lease multiples is no protection
wherever the clamp is short, because the window it names has already elapsed in
the ordinary case.
Whatever the clamp, quarantine on evidence of work, not elapsed time. An item whose previous claimant shows a branch pushed, an issue comment, or a fencing token that has moved within the last hour is a live holder, whatever the lease says, and must not be taken over. Only the sustained absence of all three licenses a takeover. This inverts the default deliberately, because the two errors are not symmetric: waiting on genuinely dead work costs bounded latency, whereas taking over live work destroys an entire session's unpushed output at the moment it finally tries to write, and destroys it silently, since the evicted worker learns of the eviction only when its next fenced write is refused.
Taking over is an explicit, evidenced act. A worker claiming an item whose previous claim lapsed must:
- Read the resume block first (
lastLocation,resumeNote) and treat it as advisory. An abandoned run leaves its branch behind but not the reasoning that produced it, and the resume note was written before whatever ended the run - so it describes an intent, not a verified state. Re-derive. - Verify the recorded branch against the remote rather than trusting
lastLocation. It may not have been pushed at all - the most common shape, since eviction tends to happen mid-build, before any push. - Never force-push or rewrite the prior attempt's branch. Build on it or start beside it; do not destroy the only record of what the previous holder did.
- Post a takeover marker on the mirrored issue naming the prior owner, the
prior fencing token, and the new one. This is what makes
attemptscountable - it is derived from the claim-comment trail, not stored - and it is the only human-visible trace that an item changed hands. - Check for a contradicting marker before doing any work. If the prior owner posted activity after the takeover marker, it was case 2 and is still alive: stop, report the collision, and let a human adjudicate. Two agents silently working one item is the failure this whole section exists to prevent.
A takeover counts as an attempt. It is not a free retry. Repeated takeovers
on one item drive it toward the poison threshold and into Parked, which is
correct: an item that keeps evicting its holders is either mis-specified or too
large, and both need a human rather than another attempt.
A claim marker records a CLAIMANT, not a grant. A worker whose lease lapses and who re-claims its own item is continuing the same work under a new fencing token; custody never changed. It must not post a second claim marker - the marker's purpose is to record who holds the item, and that did not change, so a second one is noise in the exact trail the parking sweep counts. Disclose the fence movement in the item body and in the outcome marker instead.
The corollary is load-bearing: a lapse-and-re-claim by the same owner does not count as an attempt. Counting it would park an item purely for taking longer than one lease, which inverts what parking is for - it exists to catch items that keep evicting their holders, not items that are simply long. Only a genuine change of custody, evidenced by a takeover marker naming a different prior owner, is an attempt.
Reporting a clean re-claim is mandatory, not optional. A worker that lapses and successfully re-claims its own item inside the quarantine has had a near-miss, not a non-event. Report it. Both instances of this on the protocol's first run were reported voluntarily by workers that had already recovered, and that is the only reason the gap was found at all - had they stayed silent, the protocol would have shipped with a race nobody had observed.
Mirroring to GitHub
Mirroring exists so a human can see and steer the backlog without reading agent memory. It is deliberately narrow.
- Item to issue on creation. Every item is mirrored, and the issue number
becomes the item id (
issue-2057). Identity and mirroring are the same act, so an unmirrored item does not exist. - Epics mirror as GitHub epics with native sub-issues, matching the existing convention that an epic is a container closed by its sub-issues' pull requests, never by one pull request of its own.
- State transitions mirror as an issue comment or a label - claimed,
released, parked, complete. This trail is also what
attemptsis counted from. - Mirroring is one-way for content. A human editing the issue body is the
source of truth; the item's
bodypoints at the issue rather than copying it. An agent never writes the item's specification back onto the issue, and never reconciles a divergence by overwriting the human's text. - Never mirrored: claims, leases, fencing tokens, anchors and blast radii. They churn far faster than an issue timeline should, and they are execution state rather than specification.
Entry gating - mirror-first, admit-by-label
An agent-writable backlog otherwise grows without bound and lets the fleet pick its own homework. The gate is both halves of that choice, because each closes a different hole, and it is enforced at step 5 of the ready-set computation:
- Visibility is mandatory and structural. Every item is mirrored to a GitHub issue at creation and takes its id from that issue. There is no such thing as an unmirrored item, so nothing can be enqueued invisibly.
- Agent-authored items additionally require human admission. An item an
agent proposed is opened carrying the existing
needs-specificationlabel and is excluded from the ready set while that label is present. A human removes the label to admit it. An item a human filed, or one the product owner approved in conversation with the project manager, is admitted at creation.
The label's name understates what it does. needs-specification is an
admission gate, not a to-do that an agent discharges by writing a
specification. Two consequences follow, and both have been tripped in practice:
- Writing the specification does not admit the item. An agent may draft the spec, post it on the issue and record it in the mirrored item; only a human may then remove the label. An agent that files an item, specifies it, and clears the label has proposed the work and authorised it in the same breath, which is the hole this gate exists to close - and it is worth strictly more when the proposing agent is the one persuaded by its own argument, because there is then no independent check anywhere in the loop. The project manager is barred from this explicitly in its own boundaries; the prohibition applies to every agent.
- A fully-specified issue that still carries the label is not a labelling defect and must not be "corrected". The label reports that admission is outstanding, not that prose is missing. Any agent auditing or tidying labels must leave it alone.
This reuses the repository's existing needs-specification and stale label
ladder rather than inventing a parallel state machine, and it keeps admission on
the GitHub side where a human can exercise it without an agent in the loop -
consistent with GitHub owning oversight.
Parked items ride the same ladder: an item at the poison threshold of three
attempts, or one a holder parked deliberately, carries stale on the issue
alongside the state:parked tag that the ready set actually reads, rather than
burning a whole session per scheduled tick. See "Parking an item" for why both
halves are written and why exhaustion is not the only route. Unparking is a
human act, exactly as admission is.
Worked example
Two items, one blocked by the other, both anchored to real code and both
belonging to epic issue-2099.
# 1. The blocker. The issue is filed first, so its number is the item id.
remember(repoId: "{repoId}", topic: "backlog", id: "issue-2100",
kind: "Note", author: "backlog-pm",
title: "Add the WAL shard batching seam",
body: "Spec: https://github.com/{owner}/{repo}/issues/2100",
tags: ["backlog", "priority:P1", "phase:implementation",
"homeRegion:{homeRegion}", "baseBranch:feat/epic/wal-batching"],
addLinks: {
"partOf": ["repo/{repoId}/mem/backlog/issue-2099"],
"anchoredTo": ["repo/{repoId}/file/src/lattice/BPlusTree/Grains/IWalShardGrain.cs"]
})
# 2. The dependent. blockedBy is authored on the DEPENDENT, pointing back.
remember(repoId: "{repoId}", topic: "backlog", id: "issue-2101",
kind: "Note", author: "backlog-pm",
title: "Batch the shipper poll against the new seam",
body: "Spec: https://github.com/{owner}/{repo}/issues/2101",
tags: ["backlog", "priority:P1", "phase:implementation",
"homeRegion:{homeRegion}", "baseBranch:feat/epic/wal-batching"],
addLinks: {
"partOf": ["repo/{repoId}/mem/backlog/issue-2099"],
"blockedBy": ["repo/{repoId}/mem/backlog/issue-2100"],
"anchoredTo": ["repo/{repoId}/file/src/lattice/BPlusTree/Grains/IWalShardGrain.cs"]
})
Reading it back:
scanscopeMemoryTopictopicbacklogenumerates both, with their tags.neighbors(key: "repo/{repoId}/mem/backlog/issue-2101", relation: "blockedBy", depth: 1)returnsissue-2100, which is incomplete, soissue-2101is excluded from the ready set.issue-2100names no blocker and is ready.- Completing
issue-2100movesissue-2101into the ready set on the next computation. Nothing pushes that transition, because there is no reverse index; it is observed by the next scan-plus-check pass. - Deleting
issue-2100instead makesissue-2101'sblockedBytarget returnexists: false. That is reported as a defect, not treated as satisfied. - Once an edit to the anchored file is re-ingested into the index,
recallreports both itemsstale, because theiranchoredTotarget's digest no longer matches the one captured when the edge was written. Both are re-validated before a run is spent on them. - Epic
issue-2099stays open until the item carryingintegratesto it completes, even onceissue-2100andissue-2101are both merged.