Table of Contents

The agent-operated backlog protocol

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at backlog-protocol.md, and llms.txt lists every page.

This is the generic base for an agent-operated backlog: a durable work queue that lives in repocontext memory, is drained concurrently by agent sessions under fenced claims, and is mirrored to GitHub issues for human oversight.

It is written to be repository-neutral and is the single source of truth for the protocol. The repository that hosts it consumes it unmodified, so it cannot rot into a stale copy: if this document is wrong, that repository's own backlog agents are wrong with it.

The backlog topic is a specialisation of ordinary repocontext memory: ordinary memory entries, an extended relation vocabulary, and rules that make the graph safe for several agents to drain concurrently. Read it before authoring, claiming, or completing a backlog item. The agent definitions that implement it are backlog-pm.base.md and backlog-worker.base.md; this document defines the data they operate on. Background on the fencing mechanism is in The agent-operated backlog.

Bindings

Everything repository-specific is a binding, written as {placeholder} throughout this document and supplied by the consuming repository. The bindings are:

Binding Meaning Example
{repoId} The repocontext repository id, as reported by repocontext_list_repos. Not your working directory, and not a worktree name. my-repo
{owner}/{repo} The GitHub repository that mirrors items as issues. my-org/my-repo
{ghAccount} The GitHub account every gh call authenticates as. my-github-account
{homeRegion} The region claims are taken in. Derive it from the cluster, never from a geography. lattice_list_regions gives the routable ids; repocontext_claim_status on any claimed item gives the region a claim actually records, and that is the value the tag must match. Whether it is load-bearing or merely informational is a property of the deployment, not of the tag - see the tag table below. local
{conventionsDoc} The repository's contribution conventions: branch naming, commit rules, labels. .github/copilot-instructions.md
{implementationAgent} The agent a worker delegates feature implementation to, if the repository has one. feature-dev

A consuming repository supplies these in one small override file rather than by editing this document. See bindings.example.md and the adoption guide.

If the bindings are not available, stop and report. Do not guess a repository, an account, or a region: a gh call under the wrong identity and a claim taken in the wrong region both fail in ways that are expensive to unpick.

Responsibility split - two stores, one source of truth each. Neither store copies the other's content, because two writable copies with no transaction between them diverge, and the divergence surfaces weeks later.

Concern Owner
Item identity, specification, human-visible priority, oversight, audit trail, notifications GitHub issues
Dependency graph, code anchors, claims, resume pointers, durable learnings repocontext memory

A human can reprioritise or respecify without an agent in the loop, because the thing they edit (the issue) is the thing that is authoritative.

Item schema

Topic backlog. One entry per item. The id is derived deterministically from the item - issue-2057, never a generated GUID - so a retry merges in place instead of creating a near-duplicate. The id is the mirrored issue number, which makes identity and mirroring the same act (see Entry gating).

A memory record has exactly four author-settable scalars besides the kind that repocontext_remember fixes at creation (Decision, Note or Memory), and this is the constraint the whole schema is built around. repocontext_update accepts title, body, author and provenance on a memory record and rejects every other name - update(fields: { "priority": "P0" }) fails with "The field 'priority' is not a settable scalar on a Memory record". There is no generic field bag. createdAt is set by the store at creation and is never authored. So an item's structured attributes must be carried by the two collection-valued members that do accept arbitrary content: tags and links.

Carrier CRDT Concurrency behaviour
Scalars (title, body, author, provenance) LWW register Two concurrent writers: one write is silently lost.
tags add-wins OR-Set Two concurrent writers: both survive, so the collision is visible.
links OrMap<string, OrSet> per relation Two concurrent writers: both survive, converging per relation.

The allocation follows directly from that table.

Attribute tags

Single-valued, low-cardinality attributes are carried as key:value tags. Tags are returned by scan and recall and are matchable by search, so an attribute expressed this way is filterable without reading bodies. Arbitrary : and / characters round-trip intact, so a branch name is a legal tag value.

Tag Meaning
backlog Plain marker tag. Every item carries it.
priority:P0 .. priority:P3 Ordering priority.
phase:research | phase:implementation | phase:integration Which phase of its grouping the item belongs to. Set at authoring, never changed by a worker, and it never carries execution state - a phase:complete or phase:review tag is a defect, not a status.
homeRegion:<region> The region in which claims for this item are taken. Verify this is enforced before relying on it. The intent is that a write served from any other region, under a claim taken in this one, fails closed on the write path (ForeignRegion), because the underlying lock is cluster-scoped and so the claim is not observable there - but that is a property of the deployment, not of the tag. On a single-region deployment lattice_list_regions reports only current, claims report region local, and a geographic value such as uksouth is not enforced at all: a claim from anywhere succeeds. Treat a value the cluster does not route as a defect in the binding, and note the failure is worse than a no-op - an unenforced safety assumption that the protocol documents as enforced is more dangerous than an absent one, because it is relied upon.
baseBranch:<branch> The branch this item's pull request targets. For an item in a grouping this is the epic branch, never main.
state:complete | state:parked The item's terminal state, and the only execution state carried on the item. The vocabulary is closed: those two values are the whole of it, absent means the item is live, and any other state: value is a defect that quarantines the item out of the ready set rather than leaving it claimable. The state: tag vocabulary is the single definition of the enumerated set and of the verdict for an unrecognised value; Recording completion says when the terminal tag may be written.
resource:<name> Optional, repeatable-by-name but one tag per distinct resource. Names a scarce non-file resource the item needs exclusively - a shared test box, a physical device, a deployment slot, a rate-limited external account. Two items naming the same resource may never be in flight together, however disjoint their code radii are. The resource: prefix is what makes the constraint load-bearing, because step 8 of the ready set matches on it. A bare descriptive tag asserting the same requirement - needs-box-exclusive, say - is read by no step and excludes nothing, however plainly it reads to a human.

Five tags are mandatory on every item: backlog, priority:, phase:, homeRegion: and baseBranch:. The rest are optional. state: is deliberately not among them, because its absence is exactly what "live" means.

Exactly one tag per prefix. Two priority: tags on one item means two authors wrote concurrently. Add-wins is what makes that visible rather than silent, so it is reported as a defect and reconciled, never resolved by picking one arbitrarily.

A missing mandatory tag is the more dangerous direction, because it is silent. A duplicate is loud - two values sit where one belongs, and any reader trips over it. An omission gives a reader nothing to see: the item is still enumerated by the topic scan, still survives every narrowing step, and simply never gets claimed. homeRegion: is the worst of them, because a worker filters candidates to its own region, so an item naming no region matches no worker at all. It looks healthy in every listing and starves indefinitely. Author the five together, and detect the omission at step 1 of the ready set rather than trusting authoring discipline to hold.

Keep attribute tags low-churn. OR-Set dots accumulate per add, so an attribute rewritten every run would grow a long-lived item record without bound. That is why per-run execution state - attempts, claims, leases, review rounds - is deliberately not a tag (see below). The one exception is the terminal state: tag, which is written at most once in an item's life and so costs a single dot.

The state: tag vocabulary

This section is the single definition of the state: vocabulary. Both agent prompts point here rather than restating the values, so there is one place to change and the two cannot drift apart.

Value Recognised Meaning Ready-set verdict
(no state: tag) yes The item is live. Admitted, subject to the remaining narrowing steps.
state:complete yes Terminal. The pull request is merged into its baseBranch, or the equivalent durable act has happened. See Recording completion. Dropped at step 2, silently. That is what the tag is for.
state:parked yes Terminal until a human re-admits it. See Parking an item. Dropped at step 2, silently.
any other state:<value> no Undefined. The item's real condition is unknown. Dropped at step 2 and reported as a defect. Never admitted.

Match on the prefix state:, never on the two literals. A computation that drops exactly the strings state:complete and state:parked does not treat an unrecognised value as an unknown state. It treats it as no state at all - byte-for-byte the same verdict it reaches for an item carrying no state: tag whatsoever. The tag is then inert: the item reads live and is offered as claimable, exactly as if nothing had been written. So the ready set partitions every state:-prefixed tag into recognised and unrecognised, and both partitions exclude. Only the reporting differs.

The unrecognised verdict fails safe, and the asymmetry is the whole reason the choice is not a judgement call. Excluding an item that is really live costs one missed candidate, it is named in the defect report, and it is repaired by a single write. Admitting an item that is really finished costs a worker an entire session redoing merged work, and it is silent: nothing in the store distinguishes a finished item read as live from one nobody has started. The two errors are not comparable, so the direction of failure is fixed here in advance rather than decided per reader.

It is reported, never silently absorbed. A quarantined item is named in the ready-set defect report together with its offending value, because an exclusion nobody sees is starvation - which is the failure the fail-safe verdict would otherwise trade into. The project manager reconciles the item to a recognised value - under a claim, like any other write - or removes the tag if the item is genuinely live.

The invented values already seen in practice. All four predate this section, and the earlier two-literal computation ignored every one of them:

Invented value What its author meant How a two-literal computation reads it
state:delivered the work is merged live and claimable
state:done the work is finished live and claimable
state:in-review a pull request is open and under review live and claimable
state:ready the item is live live and claimable

Three of the four assert that the work is finished or nearly so, and every one of them read as work nobody had started. state:ready happened to fail safe - it means live and was read as live - and that accident is precisely why an audit that counts presence proves nothing. Presence looked healthy; validity did not. Only the enumeration of distinct values answers the question that has any bearing on behaviour.

A non-terminal condition is expressed by holding the claim, not by a tag. This is the redirect for anyone reaching for a status value that is not in the table above, and it is what lets the vocabulary stay closed at two. A live fenced claim already means exactly "someone is working on this", and the ready set already drops claimed items at step 2; attempts and review rounds are carried by the claim markers. state:delivered and state:in-review are attempts to record a real and meaningful condition that is already representable - so hold the claim, or release honestly and write the resume block. Do not invent a tag: the store accepts any string, so the invention succeeds at the moment of writing and then does nothing.

Audit the vocabulary by enumerating distinct values, not by counting presence. One scan of the backlog topic, project every state:-prefixed tag to its distinct set, and diff that set against the two recognised values. It is cheap, and it is the only form of the check that discriminates: a count of how many items carry a state: tag is satisfied equally by a recognised value and by an invented one.

Recording completion

An item is complete when, and only when, it carries state:complete. The tag is the record; nothing else is. Without a defined encoding every worker invents one, and the inventions do not agree - which is how an item ends up tagged phase:complete (clobbering a reserved authoring attribute) or described as finished only in prose inside body, where the ready-set computation cannot see it.

Three rules make the tag trustworthy:

  • It is written by the claim holder, under its fencing token, and only after the thing it asserts is true. For an implementation item that means after the pull request is merged into its baseBranch, not after CI goes green and not after a review passes. Green-and-unmerged is not complete: the lifecycle transition is Claimed --> Complete, and an item tagged complete while its pull request is still open is a defect the next ready-set computation reports. An item that produces no pull request completes on the equivalent durable act, not on a weaker one. A research item's product is its findings, so it completes once those findings are recorded somewhere that outlives the item - the mirrored issue and durable memory - and never merely because the run ended. The item's own body does not count: it is a resume pointer rather than a deliverable, and a finding that exists only in the worker's context is lost the moment the session does. A design-integration item carries the further gate described under the grouping model: it may not complete while a grouping it emitted still lacks its dependency DAG.
  • It is terminal and it is not a status field. There is no state:review, no state:in-progress, no state:blocked. Everything short of terminal is derived from state that already exists elsewhere and is authoritative there: blocked from blockedBy, claimed from the claim surface, admitted from the mirrored issue's labels. Adding a mutable status tag would duplicate all three and churn the record besides.
  • body explains, the tag decides. The resume block says what landed and what is left, for a human and for the next claim holder. It is prose and is never parsed. No agent may infer completeness from it.

A worker that cannot merge - because it lost its claim, because review requested changes, or because it ran out of run - does not tag the item complete. It writes an honest resumeNote, posts an outcome comment carrying result=released on the mirrored issue, and leaves the item live for the next holder. result is a field of that outcome comment, never an argument to repocontext_release_claim or any other tool. That is the normal path, not a failure.

When the dispatch reserves the merge to the project manager

The rules above assume the claim holder merges its own pull request. A project manager may instead reserve the merge to itself - the right call when several items land on one shared integration branch and each merge changes what the next item is tested against, so the ordering is a project-manager decision rather than a worker one.

The reservation does not weaken the completion rule; it moves who satisfies it. Read it as a two-party sequence:

  • The worker never tags the item complete, and that is not a failure path. It is the same stand-down with an outcome comment carrying result=released described above, reached for a different reason: not that the worker could not merge, but that it was instructed not to. It writes an honest resume block, releases the claim under its fencing token, and reports plainly that the merge is outstanding and reserved.
  • The project manager tags the item complete once it has merged, under a claim it takes itself. The tag still means "the pull request is merged into its baseBranch" and nothing weaker. A reserved merge is the one case where the party that satisfies that condition is not the party that held the claim while the work was done, which is exactly why it has to be written down: an unstated exception to a rule this load-bearing is indistinguishable from a worker getting it wrong.
  • A dispatch that reserves the merge must say so explicitly. A worker whose instructions and this protocol disagree should honour the dispatch and report the divergence rather than silently pick one, because a divergence is far more often a project-manager omission than a deliberate choice.

The failure this prevents is a genuinely finished item sitting live in the ready set - offering itself to a second worker who would redo the work - because the only party permitted to tag it was forbidden to perform the act that defines the tag.

A directly-deployed item still gets a ledger row, authored at completion

Not every item reaches a worker through the backlog. A project manager may deploy a worker directly against a GitHub issue, and should: it is the right move for a one-off defect that needs no dependency ordering and no ready-set arbitration. Such an item is never claimed, holds no lease and no fencing token, and never appears in a ready set. Nothing about that is wrong.

What is wrong is leaving no trace. The project manager authors a ledger row for a directly-deployed item when it completes, carrying the five mandatory tags, state:complete, pr:<n>, and direct-deployment.

Three rules make the row safe:

  • Author it at completion, never while the work is in flight. A live row for work already underway is claimable: nothing in the store distinguishes it from work nobody has started, so the next ready-set computation offers it and a second worker can be dispatched onto an item that is already half done. Authoring at completion means the row is born terminal and can never be selected. This is the whole reason the row is retroactive rather than eager, and it is not a compromise.
  • Record what happened, and nothing more. State plainly that the item was directly deployed, never claimed, held no lease, and never entered a ready set. Do not synthesise claim or lease history for a run that took none. The distinction is worth being precise about, because the scruple that stops an agent writing the row at all is a good instinct pointed at the wrong target: fabricating a claim history would be invention, whereas a row saying "completed via direct deployment, never claimed, closed by PR #N" asserts nothing untrue. Truthful and retroactive is not the same as invented.
  • The project manager writes it, not the worker. Items are a project-manager artefact - the lifecycle opens [*] --> Drafted: authored by the project manager - and a worker authoring its own ledger row inverts that at the moment it is least able to be checked, as it stands down. A worker that notices it has no item to complete should report the absence rather than do something claim-shaped.

Why the gap is worth closing at all, given that the ready set already fails safe here: a missing target is reported as a dangling blockedBy defect and is never treated as satisfied, so no work is silently released. The cost is subtler and worse. exists: false reads identically for a benign ledger gap and for an item deleted while it was still gating work, and the second is precisely the failure that rule exists to catch. Every uncreated row makes the defect signal less able to mean anything. A complete ledger is what keeps the alarm credible, and the throughput the ledger reports honest.

Evidence a worker may rely on

The rules above say when a worker may assert completion. This one says what makes the assertion admissible. A worker that reports a measurement its own instrument fabricated is neither lying nor careless: it has simply never been told that an instrument is a thing which can itself be wrong, and a number is the most persuasive object an agent can put in front of a reviewer.

  • Validate an instrument against a case with a known non-zero answer before you trust a zero from it. A zero from an unvalidated instrument is indistinguishable from a broken instrument, and it is the most expensive kind of wrong answer because it looks like a result rather than a failure. Find a case whose correct outcome requires the quantity to be non-zero, run it, and check the instrument agrees. That is one cheap run. In this protocol's first live use, a worker instrumented an exception path, measured zero, and reported a hypothesis refuted. The probe wrote to standard output, which the test runner captures and surfaces only for failing tests, so on every passing run it reported zero by construction. The true count on a single case was 117. By then the project manager had already recorded the refutation as durable fact.

  • A measurement taken over passing runs says nothing about the failing run. The failing run is by definition the one where behaviour differed. Reset per run, capture per run, keep the failure. This is independent of the rule above and does not substitute for it: the same worker moved its measurement onto the failing run and still got a fabricated zero, because the instrument was unchanged.

  • Cite the artifact, not the impression. Evidence is a machine-readable per-case result that survives whatever verbosity the run happened to use. A console tally is not evidence: the same worker captured "1 failed, 3 passed" on a four-case fixture and could not say which case had failed, and had to reproduce an event that had already happened.

  • Say which claim you are making. "Observed failing, changed, observed passing" and "correct by construction and not observed to fail since" are different claims of different strength. Both are legitimate and the second is often the best available. Blurring them is not legitimate, because a reader who is not told which one you mean will assume the stronger.

The project manager carries the mirror of this obligation. Evidence is not authoritative merely because it is numeric, and a durable memory written from an unvalidated measurement propagates one worker's broken instrument into every later session that reads it. Ask what the instrument was and whether it was checked, prefer a retraction to a defence, and correct the durable record the moment a measurement is withdrawn.

The item body

body holds a pointer to the mirrored GitHub issue - not a copy of its specification - plus the resume block for the most recent attempt:

  • lastLocation: branch / pull request number / sha of the last attempt.
  • resumeNote: a short "what is done, what is left".

body is an LWW register, and that is safe here only because these fields are written exclusively by the current fenced claim holder, so there is never more than one writer. LWW is not unsafe in general; it is unsafe when unserialised. The fenced claim is what serialises it. Nothing else may write body while a claim is live.

The resume block is advisory. A resuming worker re-decides from it and never continues blindly, because an abandoned run leaves the branch behind but not the reasoning that produced it.

What is deliberately not on the item record

  • attempts is derived from the mirrored issue's claim-comment trail, not stored. GitHub already owns the audit trail, counting comments needs no reverse index, and a per-attempt counter on the item record would be exactly the unbounded-churn write the OR-Set dot cost warns against.
  • Claims, leases and fencing tokens live in the fenced claim/lease surface and on short-lived per-run worker records, never on the item.

Never set a TTL on a backlog item. Expiry is silent and unlogged, so a lapsed item that other items declare blockedBy starves its dependents invisibly, with no event anywhere to explain it. Retire an item deliberately with forget. This is no longer an exception to the coordination rule: the Coordination section now forbids a TTL on coordination entries generally, on the same reasoning. A backlog item - a ledger entry, not a handoff - is simply the case where the harm is most concrete. Omitting ttlSeconds is not enough on its own: repocontext_remember gives a newly created entry the repository's default memory TTL whenever the host configures one (RepoContextTtlOptions.DefaultMemoryTtl, unset by default), and ttlSeconds accepts only a positive value, so no call can opt an item out. A repository that hosts a backlog leaves that default unset.

Recording baseBranch: on the item is what makes a retry land correctly. Leaving it to worker convention means a resumed or reassigned attempt targets whatever the worker assumes, which for a sub-item of an epic is usually main - exactly the case the epic branch exists to avoid. An item that is partOf an epic and carries baseBranch:main is a defect, reported rather than silently accepted.

Cross-checking attempts against the fencing token

The claim-marker count is only as good as the workers who post markers, and a missing marker produces no signal: an item drawn six times with no marker reads attempts = 0, exactly like an item never drawn, so the poison threshold is unreachable. Every sweep that derives attempts therefore also reads the item's fencingToken from repocontext_claim_status. The claim grant itself advances it, so it cannot be forgotten; it is absent until the first grant.

The token is an upper bound, not an attempt count. A same-owner re-claim after a lease lapse advances it without being a new attempt (see Detecting and picking up a dropped lease), and a renew that omits leaseSeconds makes such lapses common. So it cross-checks the marker count and never replaces it:

fencingToken Claim markers Reading Action
absent or 0 0 never claimed none - attempts is 0
> 0 0 claimed but unmarked defect - report it; attempts is unknown, not zero
= markers > 0 consistent use the marker count as attempts
> markers > 0 token ahead of the trail use the marker count as attempts, and report the gap: markers were dropped, or claims lapsed
< markers > 0 trail ahead of the token defect - report it; a marker with no grant behind it

A sweep never reports attempts = 0 for an item whose token is above zero.

Relation vocabulary - the backlog extension

These extend the small, stable knowledge-linking vocabulary rather than competing with it. partOf and related are the documented relations used unchanged, and the five additions follow the same discipline: few, stable, one direction authored, named for what they assert.

They are documented here so tooling that audits memory - the daily Memory Accuracy automation in particular - recognises them and does not prune them as unknown relations.

Relation Authored on Points at Meaning
blockedBy the dependent item item keys Every target must be complete before this item is claimable.
anchoredTo the item file / symbol keys The code this item concerns. Gives digest-drift staleness for free.
claims a per-run worker record the item This run asserts ownership of the item.
partOf the sub-item the epic item Grouping membership. The documented relation, used unchanged.
integrates the integration item the epic it closes out Marks exactly one item per grouping as its integration join.
informs a research grouping the implementation grouping it produced Keeps the rationale behind a decomposition discoverable from the work it caused.
related either items, gotchas, decisions Near-duplicate items, and the learnings a prior attempt produced. The documented relation, used unchanged.

Two rules follow from the store's semantics rather than from taste:

  • claims lives on a short-lived per-run record, never on a long-lived one. OR-Set dots accumulate per add, so an edge asserted and released every run grows a long-lived record without bound.
  • Edges make a collision detectable, not preventable. There is no compare-and-swap anywhere in this surface: repocontext_update preconditions on record existence and, once the record is claimed, on the fencing token - never on a field's value. A claims edge is therefore an audit record of who tried, not a lock. Mutual exclusion comes from the fenced claim/lease surface, whose monotonic fencing tokens and bounded, expiry-reclaimed leases give real exclusion and a real stale-claim reaper.

Parked blockers and the ruling route

This is the single definition for authoring, dependency reporting, and parking. A parked item awaits a human ruling, not scheduled agent progress. "Stalled" means never ready without human intervention, not "not ready yet"; it is a derived report, never a new state: tag.

Before any repocontext_remember or repocontext_update that adds blockedBy, recall every proposed target in its home region and apply the table below. Validate the whole proposed write before sending it. If any target is rejected, make no write (including unrelated fields or other edges in that request) and report a hard validation error naming the dependent and each rejected blocker key. A named ruling route does not make a parked target an acceptable new dependency. Do not silently omit the edge and create an apparently independent item.

These are agent preflight rules, not a claim that the generic memory API interprets backlog tags or provides a cross-record transaction. A target may be parked after the read or after a previously valid edge was written, so authoring validation alone is insufficient.

Target observation Authoring verdict Dependency classification Required report
missing reject invalid dependent, blocker
invalid reject invalid dependent, blocker
parked reject stalled dependent, blocker, ruling route
live allow waiting dependent, blocker
complete allow satisfied none

The predicates are disjoint: missing means exists: false; otherwise invalid means multiple state: tags or any unrecognised value; parked and complete mean exactly their one recognised tag; live means no state: tag. A live blocker may be claimed or awaiting a claim; neither is a terminal state tag. The completion-evidence defect checks still apply.

Every ready-set computation and PM sweep evaluates existing edges before candidate narrowing, even if other work is ready. Page the entire backlog scan, retain each dependent's outbound blockedBy keys, and recall any targets not resolved by the scan; a truncated neighbors result is not a complete dependency set. Classify every edge, including those whose dependent is filtered out later. Report every stalled edge in the listing with both keys, the blocker's ruling owner, and the question/issue pointer. Report a missing ruling route as an additional defect, never as ordinary waiting. For a dependent with mixed blockers, invalid takes precedence over stalled, then waiting, then satisfied; retain all edge diagnostics. Only all-satisfied (including no dependencies) survives dependency narrowing. Do not delete an edge, tag a blocker complete, or unpark it to make work ready.

Parking requires a named route out before writing state:parked. Record exactly one non-empty ruling-owner:<role-or-person> tag and exactly one non-empty ruling-topic:<decision-key> tag alongside it in the same fenced memory write. In the resume note and mirrored issue comment, name the exact question that owner must rule on and the issue where they can answer it. pm-ruling-required and needs-design-decision may remain descriptive tags; they do not replace this explicit route. For example, ruling-owner:pm and ruling-topic:design-decision must point to the actual undecided design question, not merely say "needs a decision".

Ruling owner named Ruling question named Parking verdict
no no reject
no yes reject
yes no reject
yes yes allow

Here "named" requires the corresponding unique, non-empty tag and the concrete owner/question recorded in the resume note and issue comment. Reject an incomplete route as a hard validation error naming the item and missing owner/question. Existing parks without a route remain excluded and are reported to the PM for repair, never automatically unparked. After a human ruling and re-admission, remove the parking and route tags under the claim and recompute dependencies; unparking a blocker makes its dependents waiting, not satisfied, until that blocker actually completes.

Why anchoredTo matters

Linking an item to the files it concerns captures those targets' content digests at link time, so repocontext_recall reports the item stale once the code drifts - measured against the index, so an edit counts once it is re-ingested. An item whose anchor moved auto-flags "re-validate the spec before spending a run on it". This is the one capability GitHub issues cannot provide, and it doubles as the poison-item mitigation. An anchor whose target is not in the index - deleted, or not created yet - also reads stale, and so does an edge written before its target was indexed: it captured no digest, and stays stale until the edge is written again. recall also lists a target that has no live record under danglingLinks as well as staleLinks, which separates an anchor pointing at nothing - deleted code, or code that has not reached the indexed branch yet - from one whose code has moved on.

Combined with repocontext_related, anchors also give each item a blast radius, so two items touching the same code can be serialised at selection time rather than colliding at merge time. Selecting for disjointness is the primary throughput mechanism; a concurrency cap is only a backstop for an unavoidably overlapping ready set.

The grouping model - three phases

A grouping is a set of items delivered together: an epic and its sub-items, joined by partOf edges from sub-item to epic. A grouping runs in up to three phases.

  1. Research and design (optional, for a large or uncertain epic). One item per research area, fanned out to research agents. Research items produce memory entries, docs and proposals rather than code, so their blast radius is empty and they parallelise perfectly. The phase terminates in a design-integration item that reconciles the findings and emits the implementation grouping, linked to it with informs.
  2. Implementation. Seam-first fan-out: land the contract as one small fast item, then fan out implementations against it. Prefer wide DAGs to deep chains - a blockedBy edge that exists only because of how the work was described is not a real dependency.
  3. Integration. The close-out item described below.
flowchart TB
  subgraph P1["Phase 1 - research grouping (optional, leaf: never nested)"]
    direction TB
    RA["Research area A"]
    RB["Research area B"]
    RC["Research area C"]
    RI["Design integration<br/>reconcile findings, emit grouping"]
    RA --> RI
    RB --> RI
    RC --> RI
  end

  subgraph P2["Phase 2 - implementation grouping"]
    direction TB
    SEAM["Contract / seam item<br/><i>small, fast, unblocks everything</i>"]
    F1["Fan-out A"]
    F2["Fan-out B"]
    F3["Fan-out C"]
    SEAM --> F1
    SEAM --> F2
    SEAM --> F3
  end

  RI -->|informs| SEAM
  F1 --> INT
  F2 --> INT
  F3 --> INT
  INT["Phase 3 - integration item<br/><i>exclusive claim, others quiesced</i>"]
  INT --> DONE(["Epic closed"])

  classDef free fill:#dbeddb,stroke:#2d7a2d,color:#0b2e0b
  classDef excl fill:#f6e3c5,stroke:#a8721a,color:#3a2606
  class RA,RB,RC free
  class INT,RI excl

Green items have empty or disjoint blast radii and run concurrently without restriction; amber items are exclusive joins.

Termination rule, and it is load-bearing: a research grouping does not itself get a research grouping. It is a leaf phase. Without this rule an agent asked to plan an epic can recurse indefinitely into planning the planning. An item tagged phase:research may not author a further research grouping; whatever it emits is an implementation grouping. Research is also not the default - where the shape of the work is already understood, a research phase is pure critical-path depth.

A phase:research item's deliverable is a durable memory entry plus an issue comment - not a branch, and not a pull request. State this when dispatching one, because the default assumption of a worker built to ship code is that it must produce a diff. Three consequences follow:

  • It does not need a branch at all, which sidesteps the branch-naming rules above entirely. A session whose workspace was auto-provisioned with a generated branch name - frequently one carrying a username and no <type>/ prefix, both of which many repositories forbid outright - simply never pushes it, and the non-conforming name never reaches the remote.
  • A well-evidenced negative is a completed item, not a failed one. Say so at dispatch. A research item exists to be capable of killing the work that would otherwise follow it, and a worker that believes a negative reflects on it will reach for an encouraging maybe. The cheapest possible outcome of a research phase is discovering early that the implementation phase must not be built.
  • It must not touch the implementation surface. A research item that edits src/ has silently become an implementation item without being admitted as one, and its changes bypass the grouping its findings were meant to shape.

If a research item genuinely must produce a file, that is a signal it was mis-scoped as research - and the file needs a conforming branch arranged deliberately, rather than an auto-generated one pushed by default.

The integration item

Every grouping terminates in exactly one designated integration item, which is blockedBy every fan-out item in the grouping and carries an integrates edge to the epic it closes out.

It exists to absorb the risk that maximum parallelism creates: N pull requests, each green in isolation against a different base, none ever tested against the others. The failure it catches is not a merge conflict (those are visible) but the epic passing every sub-item's acceptance criteria while failing its own. Its remit is therefore conflict reconciliation, a full cross-package test run rather than the per-package targeted runs the sub-items ran, and verification against the epic's acceptance criteria.

Three rules attach to it:

  • It is exclusive. It spans the grouping's whole blast radius by design, so it cannot be selected for disjointness like a normal item. It requires an exclusive claim with the grouping's other workers quiesced. This is the one deliberate exception to the disjointness rule.
  • A grouping is not complete until its integration item is complete. An epic cannot be closed by its sub-items alone, however green they are.
  • A design-integration item may not complete while the grouping it emitted lacks a mermaid dependency DAG, and the gate applies transitively to anything those groupings go on to emit. A generated grouping is held to exactly the standard a hand-authored one is. That is the case that matters most, because a human has least visibility into a decomposition an agent assembled, so the obligation must not be launderable through a layer of automation.

Branch inheritance

An epic gets one shared branch and its sub-item pull requests target that branch; the epic reaches main as a single fully-gated pull request once its integration item passes. Concretely:

  • the epic record carries baseBranch:<type>/epic/<epic-slug>;
  • every sub-item inherits that value as its own baseBranch: tag;
  • sub-item branches are named <type>/epic/<epic-slug>-<item-slug>;
  • an item that is partOf an epic and carries baseBranch:main is reported as a defect.

The final separator is a hyphen, not a slash, and this is forced by git rather than chosen. An earlier revision of this document prescribed nesting sub-items as <type>/epic/<epic-slug>/<item-slug>. That form is unimplementable whenever the epic branch is parked on the bare slug - which the first rule above also mandates - because git stores a branch as a file at refs/heads/<name>, so refs/heads/X and refs/heads/X/anything cannot coexist. The remote refuses it:

cannot lock ref 'refs/heads/fix/epic/my-epic/my-item':
'refs/heads/fix/epic/my-epic' exists

This is a directory/file ref conflict, not a policy or permissions failure, and no naming choice on the sub-item's side avoids it. It was found by a worker attempting the push, having been reviewed twice in prose without either reader noticing - reading a ref name does not tell you git will refuse it.

Note the corollary, because it is the part that gives false assurance: a CI branch-name guard will happily accept the nested form, since it is a well-formed lower-case name under an allowed prefix. A guard that validates a string the underlying system then rejects is worse than no guard on that dimension, because it converts "unverified" into "verified" without adding verification.

The alternative - parking the epic on <type>/epic/<epic-slug>/integration and leaving the namespace free for true nesting - does work, and is the better shape for a grouping created from scratch. It is not adopted as the default because it costs a rename of the epic branch and a rebase of every in-flight sub-item if adopted mid-grouping. Choose it at epic-creation time or not at all.

Two invariants matter more than the name, and are what a reviewer should actually check. The name is a convenience; these are correctness:

  1. the sub-item branch is descended from the epic branch - git merge-base --is-ancestor origin/<epic-branch> HEAD exits 0;
  2. the sub-item's pull request targets the epic branch, never main.

A sub-item that satisfies both under an off-convention name is fine and is reported as a naming nit. A sub-item that satisfies neither under a perfectly conventional name has silently bypassed the epic, and its work will not be collected by the integration item.

Do not use a workspace rename_branch affordance to satisfy this rule without checking its output. In at least one environment it applies a configured prefix that injects a username and omits the <type>/ prefix entirely, producing a name this repository forbids outright and which fails the CI guard. Rename the branch directly and verify the resulting name.

Computing the ready set

The ready set is the items claimable right now. It is always computed as a topic scan plus per-candidate depth-1 checks, and never as a single graph query.

Why it cannot be one call. repocontext_neighbors is navigation, not query: it walks outbound edges only, with depth clamped to [1, 3] and maxNodes to [1, 100]. There is no reverse index over memory links - the reverse cross-reference index serves repocontext_related for symbols only - so "who is blocked by me?" and "what did completing X unblock?" cannot be asked directly. They require either an explicitly authored inverse edge or a topic scan. Do not design a protocol around a reverse lookup this surface cannot serve. Scan-plus-check is fine at hundreds of items; this is a coordination graph, not a queue engine.

The computation:

  1. repocontext_scan scope MemoryTopic, topic backlog, paging on the continuation token, to enumerate every live item. Verify each item carries the five mandatory tags as you page, and report any item that does not rather than silently passing over it. The check is free here, because every item is already in hand, and this is the only step that sees all of them - so an item malformed in a way that hides it from the later narrowing steps is caught here or not at all. Before narrowing candidates, validate parking routes and classify all existing dependency edges using Parked blockers and the ruling route; retain the stalled and invalid diagnostics even when the ready set is non-empty.

  2. Match every state:-prefixed tag against the closed vocabulary in The state: tag vocabulary, and drop the item on both outcomes. A recognised terminal value - state:complete or state:parked - drops it silently, because that is what the tag is for. Any other state: value drops it too, and is reported as a defect: it is an unknown state, not an absent one, and admitting it is how a finished item gets offered to a second worker. Do not match on the two literals alone - that reads an unrecognised value as no state at all and admits it. Also drop items held under a live fenced claim (repocontext_claim_status). Completeness is read from the tag and from nothing else - never from prose in body, and never from a merged-looking pull request.

  3. Drop grouping records. A grouping (an epic, or any item that other items declare themselves partOf) is a container, not a unit of work. It is completed by its integration item, never claimed directly. Omitting this step lets a worker claim the epic itself and duplicate the entire fan-out that the decomposition just created.

    Build the exclusion set during the step-1 scan, at no extra cost: collect the target of every partOf edge you encounter as you page through the topic. Do not attempt this as a reverse lookup - "who is partOf me?" is exactly the reverse-index query this surface cannot serve, which is why the check has to be a by-product of the enumeration rather than a per-candidate probe.

    The same conclusion can be reached from the data alone, and belt-and-braces is cheap here: a grouping should also carry blockedBy its own integration item, which drops it at step 4 anyway. Author both. The redundancy is one-way safe - it can only ever remove a container from the ready set, never admit one.

  4. For each remaining candidate, apply the dependency classifications from step 1 per Parked blockers and the ruling route. Only all-satisfied survives; ordinary waiting, stalled, and invalid dependencies stay out of the ready set.

  5. Drop survivors whose mirrored issue is not admitted (see Entry gating). This is checked after the blockedBy narrowing, so it costs one issue read per survivor rather than one per item in the topic.

  6. Sort by (priority, createdAt, id), then pick from the top three to five. Ordering deterministically is fine and is not a defect: repocontext_claim is real mutual exclusion, so two workers converging on the same item resolve to exactly one proceeding and the other observing a clean refusal it can act on immediately. Jitter is a cheap way to spread the fan-out across candidates and avoid spending a round on a refusal, so it remains worth applying - but it is an optimisation, and no worker may rely on it for correctness.

  7. Prefer a candidate whose blast radius - its anchoredTo anchors plus repocontext_related on them - is disjoint from the radii of in-flight items.

  8. Exclude, rather than merely deprioritise, a candidate whose resource: tags collide with an in-flight item's. Step 7 is a preference computed over anchoredTo files, so a scarce non-file resource is invisible to it: an item whose real constraint is "needs exclusive use of the shared test box" may have an empty code radius and will therefore look maximally disjoint and sort to the front. That is the exact inversion of the truth. Resource collision is a hard exclusion, not a tie-break, because the failure it prevents - two agents recreating the same container under one another - is not a merge conflict that surfaces loudly but a corrupted experiment that reports a plausible wrong answer.

A scan is a bulk read and therefore does not evaluate TTL or link staleness: stale and staleLinks come back null there, meaning "not evaluated" rather than "not stale". Staleness must be read with recall on the specific candidate.

Defect conditions the ready-set computation must surface

These are reported, never silently absorbed:

  • Stalled by a parked blocker, or parked without a ruling route. Apply Parked blockers and the ruling route before narrowing, report both item keys and the human ruling needed, and keep stalled distinct from ordinary waiting even when other work is ready.

  • Dangling blockedBy. A target that returns exists: false is a defect, not a satisfied dependency. Treating an absent blocker as complete is how a deleted item silently releases work that was deliberately gated on it.

  • Stale item. An anchoredTo target drifted, so recall reports the item stale. Re-validate the spec before spending a run on it.

  • Claimed but unmarked. A fencingToken above zero with no claim marker, or a token that disagrees with the marker count in either direction. With no marker, attempts is unknown, not zero, so the item is not fresh. See Cross-checking attempts against the fencing token.

  • Duplicate attribute tag. Two tags sharing a key: prefix means two concurrent authors. Reconcile; never pick one arbitrarily.

  • Unrecognised state: value. A state:-prefixed tag whose value is neither complete nor parked. Step 2 quarantines the item out of the ready set and names it here with its offending value, because the tag asserts a condition the protocol cannot interpret, and the safe reading of an uninterpretable assertion is not "live". Reconcile it to a recognised value under a claim, or remove the tag if the item really is live. See The state: tag vocabulary.

  • Execution state on a phase: tag. phase: carries the authored phase and nothing else, so phase:complete or phase:review means a worker wrote a status into a reserved attribute - and, because add-wins never replaces, the item's real phase is either lost or now duplicated. Reconcile to the authored phase plus a state: tag if one is warranted.

  • Item tagged state:complete with an unmerged pull request. Completion was claimed before the merge that defines it. The item is not complete; the merge is outstanding work.

  • Green, mergeable pull request on an item with no live claim and no state:complete. The attempt died between CI passing and the merge. This is the cheapest possible resume and should be picked up before any fresh item.

  • Ready set empty while pending is not. There is no cycle detection in the store, so a dependency cycle is silent permanent starvation. Alarm rather than exit quietly.

  • Ready set empty and pending empty. Exit immediately. Every tick otherwise spends a whole session for nothing.

  • baseBranch:main on an item that is partOf an epic. See branch inheritance above.

  • A grouping whose fan-out is complete but whose integration item is not. The grouping is not complete; do not close the epic.

Item lifecycle

stateDiagram-v2
  [*] --> Drafted: authored by the project manager
  Drafted --> Gated: mirrored to a GitHub issue
  Gated --> Ready: admitted (human, or human-authored at source)
  Ready --> Blocked: a blockedBy target is incomplete
  Blocked --> Ready: every blocker completes
  Ready --> Stalled: a blocker is parked
  Blocked --> Stalled: a blocker is parked
  Stalled --> Blocked: human re-admits blocker, work still pending
  Stalled --> Ready: every blocker completes
  Ready --> Quarantined: an unrecognised state tag value is present
  Quarantined --> Ready: reconciled to a recognised value, under a claim
  Ready --> Claimed: fenced claim acquired (homeRegion only)
  Claimed --> Ready: lease expires, or the worker releases
  Claimed --> Complete: PR merged into the base branch, or equivalent durable act
  Claimed --> Parked: attempts exceed the poison threshold
  Claimed --> Parked: the holder parks it deliberately
  Parked --> Ready: a human respecifies and re-admits
  Complete --> [*]

Claimed --> Ready on lease expiry is the normal path, not an exception. Stale claims are the common case, so a claim is always lease-bounded and reclaimed on expiry rather than held by a flag that a killed session leaves set forever.

Parking an item - both sides, and not only on exhaustion

Parking is encoded on both sides, and the tag is the half that has effect. The ready set drops parked items at step 2 by reading the state:parked tag on the item record; the stale label on the mirrored issue is what makes the park visible to a human. A park that writes only the label is not a park: the item survives every step of the ready set and is claimable again on the next tick, so the guard silently does nothing. Write both, and write them under the fencing token of the claim you hold, in the order tag then label - if the run dies between them the item is already out of the ready set and the sweep can finish the visible half.

First validate and record the ruling route. That section also governs dependents stranded by this transition and existing parks missing a route; neither may disappear as ordinary "not ready" work.

The poison threshold is three, and it is a floor on parking, not the only route to it. An item whose claim-marker count has reached three is parked by whichever worker takes it there. But a holder that establishes, at any attempt number, that the item cannot proceed as specified parks it deliberately and does not wait to burn the remaining attempts. The two cases that matter:

  • The work is blocked on a decision, a defect, or a dependency that is not itself an item, so no blockedBy edge can express it.
  • The specification is wrong, not merely hard - the item as written cannot be satisfied.

Releasing instead is the failure mode this exists to prevent. Releasing the claim and posting an outcome comment carrying result=released on the mirrored issue leaves the item live and immediately re-claimable, so the next worker draws it, re-derives the same finding, and releases in turn; the fleet spends a session per tick relearning one conclusion. A deliberate park costs one attempt and states the conclusion once.

Say why, in both places a later reader will look. The resumeNote carries the finding for the next holder, and the outcome comment carrying result=parked on the mirrored issue carries it for the human who must decide. A park with no stated reason is indistinguishable from a crash and will be unparked without the finding being addressed.

Parking is safe to get wrong in one direction only. It can only ever remove an item from the ready set, and Parked --> Ready requires a human, so an over-eager park costs a human glance while an omitted one costs an unbounded loop. Park when in doubt. Parking is a finding, not a failed attempt, and a worker that parks correctly on its first attempt has done its job.

The lease is shorter than the work - renew before, never after

The cluster applies a short default lease when leaseSeconds is omitted - commonly 30 seconds - and clamps every request to a host-configured ceiling. Do not assume either figure: read the leaseSeconds and leaseExpiresAtUtc your grant actually returns, because a request above the ceiling is clamped silently and a deadline diaried from the length you asked for is already late. A build-and-test cycle on a non-trivial repository exceeds a 30-second lease many times over. The consequence is not hypothetical and was observed on the first live run of this protocol: two independent workers each had a claim lapse mid-build, while actively working the item.

The short default is deliberate, and claim and renew_claim apply it identically. There is no divergence between the two surfaces - this was measured on a live deployment, both arms with leaseSeconds genuinely omitted, and both granted the same length. The rationale for keeping the default short is that a caller which did not name a lease length is exactly the caller that should not be granted a long one. A host that needs longer claims raises the ceiling an explicit request may reach, not the default.

Both recovered correctly - repocontext_claim_status showed no other holder and no queue, and the re-claim returned a fencing token incremented by exactly one - so the mechanism behaved as designed. The gap is that the lease duration is shorter than the shortest useful unit of work, which turns a safety property into a routine occurrence.

Why that matters more than a retry: during the lapse the item is, to any other worker computing the ready set, simply available. Step 2 drops items "held under a live fenced claim", and nothing holds this one: the lease has lapsed, so repocontext_claim_status reports isHeld: false and a repocontext_claim would be granted. The record still reports claimed: true with the last fencing token, which is what separates a lapse from an item nobody has started (claimed: false, no fencingToken), but a ready set that goes by the lock cannot see the difference. A sibling recomputing in that window would have found the item available and begun duplicate work on an item another worker was mid-build on. Nothing prevented that. Only the timing did.

Rules, in force for every worker:

  • Always pass leaseSeconds explicitly - on renew_claim as well as on claim. The short default is shorter than almost any real operation and will lapse under a single test run. On a renew the omission is worse than on a claim: a renew that omits leaseSeconds resolves to that same short default and therefore shortens a claim you are currently holding for longer. It still reports granted: true; the loss surfaces only on the next renew, as granted: false, reason: "superseded", at a call site that did nothing wrong. A renew that shortens its lease is flagged by leaseShortened in the result - and note that a null there means the prior lease could not be read, so it is "unknown", never "nothing shrank".
  • Renew immediately BEFORE any long operation, never after it. Treat a build, a test run, or anything expected to exceed roughly two minutes as requiring a renewal first. Renewing afterwards is renewing during the window you needed to be covered for.
  • On discovering a lapsed claim, re-claim and then CHECK THE FENCING TOKEN. If it incremented as expected and the holder is you, that is a clean re-claim; proceed, and report it. If the holder is not you, or the token did not move as expected, stop and report - somebody else has been working the item, and continuing would produce two divergent attempts at one unit of work.
  • Never write anything under a token you know to be stale.

Fixes worth making to the surface itself, in preference order: raise the host-configured ceiling above a realistic build time (many hosts already do - check what your grants actually return before assuming otherwise), or make it per-phase, since a research item and a build item have very different natural durations; auto-renew on a timer for the lifetime of a long child process rather than asking a worker to predict its duration; and distinguish "lease expired while work was in progress" from "never claimed" in the ready set, so a lapse degrades to a warning rather than to availability. That last needs no change to the surface, only to the ready-set computation: repocontext_claim_status already separates the two (claimed: true with isHeld: false, against no fencingToken at all).

This was surfaced only because a worker volunteered an unflattering detail it had already recovered from. A protocol that discourages that reporting would have shipped this gap silently.

Detecting and picking up a dropped lease

Raising the lease only makes a lapse rarer. It does not say what a lapse means or who may act on it, and that is the part that has to be specified, because the store cannot answer the only question that matters.

The central difficulty: a lapse has two causes and the surface cannot tell them apart. An expired lease means either

  1. the holder is gone - crashed, killed, context-exhausted, session ended - and the item genuinely needs picking up; or
  2. the holder is alive and working, and merely failed to renew in time.

Both present identically: no live lease (isHeld: false) on a record that still reports claimed: true. Nothing in the lock, the item record, or the ready set distinguishes them, and there is no liveness signal independent of the renewal itself. Treating every lapse as case 1 duplicates live work; treating every lapse as case 2 leaks items permanently to dead agents. Neither default is safe, so the protocol makes the distinction unnecessary rather than pretending to resolve it.

Detection is pull, not push. Renewal IS the liveness probe. A worker is never notified that its lease expired; it finds out only by attempting a renew (or a fenced write) and being refused. There is no callback and no interrupt. A worker that never renews never learns it was evicted, and will keep working - which is precisely case 2 seen from the inside. This is why renewal is mandatory before long operations rather than merely advisable: it is the only mechanism by which a worker discovers it has lost the item.

What fencing does and does not protect, which is the load-bearing point. The monotonic fencing token makes store writes safe: a superseded worker's write is rejected, so two workers can never both mutate the item record. It protects nothing else. Git, GitHub, the filesystem and any deployed environment are outside the fence. A superseded worker can still push a branch, open a pull request, comment on an issue, or recreate a container, and none of those will be refused on account of a stale token.

Therefore:

  • Renew immediately before every externally-visible side effect, and verify the token, not merely that the call succeeded. Push, pull-request creation, issue comments, and any environment mutation are all gated on a fresh, verified renewal. A renewal that returns is not enough; the token it reports must be the one you hold.
  • On refusal, abort without side effects. Do not push "just this branch", do not open the pull request, do not comment. Report and stop.
  • Never destroy your own work on discovering you were superseded. The branch and commits from an evicted attempt are the takeover's most useful input. Leave them, and say in your report exactly where they are. Deleting them converts a recoverable handover into a restart.

A lapsed item is quarantined before it becomes claimable. It does not re-enter the ready set the instant the lease expires. It becomes eligible only after a quarantine interval that comfortably exceeds the longest plausible renewal gap - one full lease is the working default. This is what buys the distinction the store cannot make: an alive-but-late holder reclaims its own item inside the quarantine and continues (its fence increments, nothing else changes), whereas a genuinely dead holder never does, and the item is released to others only after that window closes. The cost is bounded latency on genuine failures; the benefit is that the common case stops being a race.

One full lease can be the wrong quarantine, and elapsed time is the wrong evidence. The working default above assumes the lease approximates the work. It often does not, and the gap depends on a value you must measure rather than assume. MaxLockLeaseDuration is a configured ceiling whose shipped default is 300 seconds, against turns that routinely run for hours; where it stands at that default, a live worker's claim spends almost all of its life presenting as lapsed. This was observed on the first real run - a productive worker sat at fence 12, mid-implementation, while claim_status reported isHeld: false and the item showed no unmet blockers. To any agent computing a ready set it was indistinguishable from abandoned work, and the lock would have granted it on request.

Do not read the 300-second figure as the value in force. It is a default, not a constant, and a deployment may raise it: on the reference deployment a claim requesting 1800 seconds was measured being granted 1800 seconds, so the ceiling there is at least that. The bundled RepoContext container host, for one, raises it to 1800 seconds by default through LATTICE_MAX_LOCK_LEASE_SECONDS (accepted range 30 to 7200). Read the leaseSeconds your own grant returns and reason from it. A quarantine measured in lease multiples is no protection wherever the clamp is short, because the window it names has already elapsed in the ordinary case.

Whatever the clamp, quarantine on evidence of work, not elapsed time. An item whose previous claimant shows a branch pushed, an issue comment, or a fencing token that has moved within the last hour is a live holder, whatever the lease says, and must not be taken over. Only the sustained absence of all three licenses a takeover. This inverts the default deliberately, because the two errors are not symmetric: waiting on genuinely dead work costs bounded latency, whereas taking over live work destroys an entire session's unpushed output at the moment it finally tries to write, and destroys it silently, since the evicted worker learns of the eviction only when its next fenced write is refused.

Taking over is an explicit, evidenced act. A worker claiming an item whose previous claim lapsed must:

  1. Read the resume block first (lastLocation, resumeNote) and treat it as advisory. An abandoned run leaves its branch behind but not the reasoning that produced it, and the resume note was written before whatever ended the run - so it describes an intent, not a verified state. Re-derive.
  2. Verify the recorded branch against the remote rather than trusting lastLocation. It may not have been pushed at all - the most common shape, since eviction tends to happen mid-build, before any push.
  3. Never force-push or rewrite the prior attempt's branch. Build on it or start beside it; do not destroy the only record of what the previous holder did.
  4. Post a takeover marker on the mirrored issue naming the prior owner, the prior fencing token, and the new one. This is what makes attempts countable - it is derived from the claim-comment trail, not stored - and it is the only human-visible trace that an item changed hands.
  5. Check for a contradicting marker before doing any work. If the prior owner posted activity after the takeover marker, it was case 2 and is still alive: stop, report the collision, and let a human adjudicate. Two agents silently working one item is the failure this whole section exists to prevent.

A takeover counts as an attempt. It is not a free retry. Repeated takeovers on one item drive it toward the poison threshold and into Parked, which is correct: an item that keeps evicting its holders is either mis-specified or too large, and both need a human rather than another attempt.

A claim marker records a CLAIMANT, not a grant. A worker whose lease lapses and who re-claims its own item is continuing the same work under a new fencing token; custody never changed. It must not post a second claim marker - the marker's purpose is to record who holds the item, and that did not change, so a second one is noise in the exact trail the parking sweep counts. Disclose the fence movement in the item body and in the outcome marker instead.

The corollary is load-bearing: a lapse-and-re-claim by the same owner does not count as an attempt. Counting it would park an item purely for taking longer than one lease, which inverts what parking is for - it exists to catch items that keep evicting their holders, not items that are simply long. Only a genuine change of custody, evidenced by a takeover marker naming a different prior owner, is an attempt.

Reporting a clean re-claim is mandatory, not optional. A worker that lapses and successfully re-claims its own item inside the quarantine has had a near-miss, not a non-event. Report it. Both instances of this on the protocol's first run were reported voluntarily by workers that had already recovered, and that is the only reason the gap was found at all - had they stayed silent, the protocol would have shipped with a race nobody had observed.

Mirroring to GitHub

Mirroring exists so a human can see and steer the backlog without reading agent memory. It is deliberately narrow.

  • Item to issue on creation. Every item is mirrored, and the issue number becomes the item id (issue-2057). Identity and mirroring are the same act, so an unmirrored item does not exist.
  • Epics mirror as GitHub epics with native sub-issues, matching the existing convention that an epic is a container closed by its sub-issues' pull requests, never by one pull request of its own.
  • State transitions mirror as an issue comment or a label - claimed, released, parked, complete. This trail is also what attempts is counted from.
  • Mirroring is one-way for content. A human editing the issue body is the source of truth; the item's body points at the issue rather than copying it. An agent never writes the item's specification back onto the issue, and never reconciles a divergence by overwriting the human's text.
  • Never mirrored: claims, leases, fencing tokens, anchors and blast radii. They churn far faster than an issue timeline should, and they are execution state rather than specification.

Entry gating - mirror-first, admit-by-label

An agent-writable backlog otherwise grows without bound and lets the fleet pick its own homework. The gate is both halves of that choice, because each closes a different hole, and it is enforced at step 5 of the ready-set computation:

  1. Visibility is mandatory and structural. Every item is mirrored to a GitHub issue at creation and takes its id from that issue. There is no such thing as an unmirrored item, so nothing can be enqueued invisibly.
  2. Agent-authored items additionally require human admission. An item an agent proposed is opened carrying the existing needs-specification label and is excluded from the ready set while that label is present. A human removes the label to admit it. An item a human filed, or one the product owner approved in conversation with the project manager, is admitted at creation.

The label's name understates what it does. needs-specification is an admission gate, not a to-do that an agent discharges by writing a specification. Two consequences follow, and both have been tripped in practice:

  • Writing the specification does not admit the item. An agent may draft the spec, post it on the issue and record it in the mirrored item; only a human may then remove the label. An agent that files an item, specifies it, and clears the label has proposed the work and authorised it in the same breath, which is the hole this gate exists to close - and it is worth strictly more when the proposing agent is the one persuaded by its own argument, because there is then no independent check anywhere in the loop. The project manager is barred from this explicitly in its own boundaries; the prohibition applies to every agent.
  • A fully-specified issue that still carries the label is not a labelling defect and must not be "corrected". The label reports that admission is outstanding, not that prose is missing. Any agent auditing or tidying labels must leave it alone.

This reuses the repository's existing needs-specification and stale label ladder rather than inventing a parallel state machine, and it keeps admission on the GitHub side where a human can exercise it without an agent in the loop - consistent with GitHub owning oversight.

Parked items ride the same ladder: an item at the poison threshold of three attempts, or one a holder parked deliberately, carries stale on the issue alongside the state:parked tag that the ready set actually reads, rather than burning a whole session per scheduled tick. See "Parking an item" for why both halves are written and why exhaustion is not the only route. Unparking is a human act, exactly as admission is.

Worked example

Two items, one blocked by the other, both anchored to real code and both belonging to epic issue-2099.

# 1. The blocker. The issue is filed first, so its number is the item id.
remember(repoId: "{repoId}", topic: "backlog", id: "issue-2100",
         kind: "Note", author: "backlog-pm",
         title: "Add the WAL shard batching seam",
         body: "Spec: https://github.com/{owner}/{repo}/issues/2100",
         tags: ["backlog", "priority:P1", "phase:implementation",
                "homeRegion:{homeRegion}", "baseBranch:feat/epic/wal-batching"],
         addLinks: {
           "partOf":     ["repo/{repoId}/mem/backlog/issue-2099"],
           "anchoredTo": ["repo/{repoId}/file/src/lattice/BPlusTree/Grains/IWalShardGrain.cs"]
         })

# 2. The dependent. blockedBy is authored on the DEPENDENT, pointing back.
remember(repoId: "{repoId}", topic: "backlog", id: "issue-2101",
         kind: "Note", author: "backlog-pm",
         title: "Batch the shipper poll against the new seam",
         body: "Spec: https://github.com/{owner}/{repo}/issues/2101",
         tags: ["backlog", "priority:P1", "phase:implementation",
                "homeRegion:{homeRegion}", "baseBranch:feat/epic/wal-batching"],
         addLinks: {
           "partOf":     ["repo/{repoId}/mem/backlog/issue-2099"],
           "blockedBy":  ["repo/{repoId}/mem/backlog/issue-2100"],
           "anchoredTo": ["repo/{repoId}/file/src/lattice/BPlusTree/Grains/IWalShardGrain.cs"]
         })

Reading it back:

  • scan scope MemoryTopic topic backlog enumerates both, with their tags.
  • neighbors(key: "repo/{repoId}/mem/backlog/issue-2101", relation: "blockedBy", depth: 1) returns issue-2100, which is incomplete, so issue-2101 is excluded from the ready set. issue-2100 names no blocker and is ready.
  • Completing issue-2100 moves issue-2101 into the ready set on the next computation. Nothing pushes that transition, because there is no reverse index; it is observed by the next scan-plus-check pass.
  • Deleting issue-2100 instead makes issue-2101's blockedBy target return exists: false. That is reported as a defect, not treated as satisfied.
  • Once an edit to the anchored file is re-ingested into the index, recall reports both items stale, because their anchoredTo target's digest no longer matches the one captured when the edge was written. Both are re-validated before a run is spent on them.
  • Epic issue-2099 stays open until the item carrying integrates to it completes, even once issue-2100 and issue-2101 are both merged.