Table of Contents

RepoContext MCP container - "codebase memory in a box"

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.

This sample runs the RepoContext MCP server as a single, restart-durable container alongside its embedding companion, and demonstrates the core durability guarantee end to end:

start -> add a repo under the mounted workspace -> recall -> restart -> context is still present.

The same box also serves the retrieval and token-economics tools for you to use and test: explainable search, the budgeted context bundle, its reuse economics, and usage accounting (see Retrieval and token economics).

Three containers, one private network:

  • repocontext - the MCP host image (apps/repocontext/Dockerfile). Its ONLY application listener is the MCP endpoint on port 8080 (plus the HTTP health probes and the Prometheus /metrics scrape endpoint). No gRPC facade and no Explorer UI are exposed. It runs the default local durability profile: Orleans ADO.NET grain storage and reminders over a single SQLite file, plus the file-backed Lattice WAL - all under /data, which is a named volume, so state survives docker compose restart, docker compose down, and image upgrades. Zero external services.
  • embedder - the ONNX Runtime model companion (apps/embedding-onnx/Dockerfile). It stays a SEPARATE container so the MCP host keeps its single-listener surface. The host's default embedding provider is pointed at it via LATTICE_EMBEDDING_ENDPOINT. Its model weights are baked into an image layer, so it needs no cache volume and no download on first run.
  • azurite-backup-sink - a blob-only Azurite instance (mcr.microsoft.com/azure-storage/azurite:latest) that receives the scheduled captures of the durable agent-memory tree, so the backup does not live in the store it protects. Its storage is a host bind mount at ${REPOCONTEXT_BACKUP_PATH:-./backup-sink} rather than a named volume, so docker compose down -v cannot reach it; set REPOCONTEXT_BACKUP_PATH to an absolute path to keep it outside the repository. Its blob endpoint is published on host port 11000 (override with REPOCONTEXT_BACKUP_SINK_PORT) so a restore can list it. See Agent-memory backup and recovery.

For a long-lived, tuned deployment rather than a first run - the CPU and memory grants, the reduced scan cadence, the image pin, the rollback ladder, and how to rebuild the whole thing from nothing - see the local deployment runbook and the tracked docker-compose.tuning.yml beside this file. The walkthrough below uses the untuned defaults and does not need either.

Optional CPU pinning

Two variables, REPOCONTEXT_CPUSET and EMBEDDER_CPUSET, pin each service to a fixed set of CPUs. Both are unset by default and should stay that way unless you have a reason; unset, Compose omits the key from the resolved document entirely, so ignoring them deploys exactly what this sample deployed before they existed.

They matter only once a CPU quota is in force, which the tuning overlay supplies. A container entitled to 4 CPUs but visible on 16 has its threads scattered across 16 run queues, and CFS charges a whole 5 ms slice to each queue a thread wakes on, so the quota is exhausted by reservation rather than by work: measured 5.4% throttled periods at ~17% quota utilisation, against 0.0% at ~26% when pinned, varying nothing else. Sizing the cpuset to the grant removes the throttling.

Derive the values from your own grants rather than copying them - one range per service, sized to the ceiling of that service's actual cpus grant in your .env (an EMBEDDER_CPUS raised above what scripts/New-TuningEnv.ps1 derives needs an EMBEDDER_CPUSET re-sized to match), the ranges must not overlap or you trade throttling for contention, and together they must fit the host's logical CPU count - grants whose ceilings sum past it cannot both be pinned, and pinning is not the fix for that. See .env.example for the form and the runbook section "CPU scatter under a fractional quota" for the evidence, the limits of what it claims, and why you must not enable it in the middle of a measurement.

Prerequisites

  • Docker with Compose v2.
  • An .env file in this directory: copy .env.example to .env and set its two paths for this machine before running any compose command below. REPOCONTEXT_MEMORY_ARCHIVE_PATH has no default, so compose refuses every command - up, build, ps, logs - until it is set (see What actually protects it).
  • Memory: derive the grant with scripts/New-TuningEnv.ps1 (issue #2779) rather than copying a figure from here. This is much more than a default Docker VM allocation, so give the VM headroom above whatever the script derives. Two figures matter and they are not the same number. Steady state is fitted at 3 GiB plus 1 MB per indexed file, so a ~8,200-file index settles near 11 GiB. The reclamation peak is what a box costs to reach that steady state, and it is higher: on an 8,224-file index whose WAL garbage collection had been blocked, releasing the backlog ran at 450-590% CPU and drove the working set to at least 13.81 GiB before it fell back to 10.55 GiB, while a 12 GiB limit crash-looped twice in 16 minutes (issue #3252). An earlier revision of this list recommended "at least 12 GiB" against a 10.2 GiB steady state; both figures are withdrawn. The peak is a migration cost paid once, the first time a backlog of stuck WAL is released, so an operator upgrading into a WAL GC fix needs more headroom than one already running healthy.
  • Memory: 13.81 GiB is a floor on the peak, not the peak. It is the largest of 16 samples about 103 seconds apart, so the true maximum is at least that and may be higher - nothing observed the gaps. Quote it as a lower bound. The fall to 10.55 GiB afterwards is the other half of the reading, and the useful half: a 3.26 GiB give-back is what distinguishes a bounded transient from a leak, which would not have receded.
  • Memory: do not assume the derived grant covers the migration burst. On the corpus above New-TuningEnv.ps1 derives 13.26 GiB and the observed peak exceeded it by about 568 MiB, so the script's 20% headroom band is all that stands between a reclamation burst and the ceiling, and this burst was larger than the band. Grant above the derived figure for the migration run specifically, then re-derive once the estate is healthy. That 568 MiB is pinned to one corpus and is not a constant shortfall - the derived grant is a function of file count and moves with it, so the same repository derives 13.26 GiB at 8,239 files and 13.97 GiB at 8,846. Do not read a later derivation that happens to exceed 13.81 as evidence the gap has closed: a larger corpus raises the peak too, and only the grant side of that comparison was re-measured. Compare a peak against a grant only at the same corpus.
  • Memory, and this decides whether the comparison just made means anything, because it is weaker than it looks in one direction and stronger in another. First, a working set measured under a generous cap is an upper bound on need, not a requirement - .NET collects less eagerly the further it is from its ceiling, so 13.81 GiB observed at an 18 GiB grant does not establish that 13.81 GiB is needed. A peak cannot be transported across caps: at 18 GiB the collector worked against a 13.5 GiB managed ceiling, whereas at 13.26 GiB it would work against 9.94 GiB and collect far harder, far earlier. That trajectory was never run. Second, pulling the other way, peak RSS is not the quantity a grant is tested against at all: the grant bounds RSS, but the GC hard limit is what throws, and it binds first at about 75% of the grant (12 GiB grants 9 GiB, 18 GiB grants 13.5 GiB). A box at 13.26 GiB would fault against 9.94 GiB of managed heap long before RSS could reach 13.26, so "the peak exceeded the grant" understates the exposure rather than overstating it.
  • Memory, what is actually established, stated on the plane that throws: a 12 GiB grant (9.00 GiB managed ceiling) crash-looped, and an 18 GiB grant (13.5 GiB managed ceiling) ran clean. Nothing has been measured at 13.26 GiB in either direction. The derived grant therefore offers a managed ceiling only about 10% above one that demonstrably crash-looped on this corpus. That thin margin over a measured failure is the reason to provision above it for a migration run - not the raw RSS comparison, which weighs a number produced under one cap against a different cap.
  • Memory: everything above is deploy-time fitting, which is a known limitation rather than the settled answer. New-TuningEnv.ps1 is host-specific by construction - its constants were fitted against one corpus on one host and are re-derived by nobody afterwards - and it goes stale in place, because its only corpus input is the indexed file count, so adding a repository to the workspace or removing one moves the requirement without moving the grant. Nothing re-derives the grant when that happens. Adapting sizing to the granted resources at runtime, instead of predicting it at deploy time, is tracked in issue #3255. What has landed from it so far is the measured consequence: the host publishes its peak commitment and its peak occupancy of the granted ceiling, records a managed-heap exhaustion on the data mount, and at the next start refuses a grant no larger than a recorded exhaustion ceiling and warns when the previous run's peak does not fit (see Measured requirement and startup admission). The runtime adaptation depends on issue #3133: the runtime's own high-load signal is published at 90% of the cgroup limit while the GC hard limit binds at 75% of it, so the threshold sits at 1.2x the limit at every grant and can never fire. It has been confirmed at both 12 GiB and 18 GiB with byte-exact matching percentages, so it is scale-invariant rather than a misconfiguration of one deployment.
  • Memory: under-provisioning does not present as memory pressure. A cgroup limit becomes the .NET GC heap hard limit, so the process is never OOM-killed and there is no restart, exit code or resource event. The visible symptom is a STORAGE error while reading grain state, because the allocation that fails is a leaf-snapshot deserialisation; the leaf then activates cold and replays its whole WAL window, raising pressure further (issue #2364). The orleans.lattice.leaf.snapshot.load_failures counter names the real cause directly (reason=resource_exhausted), but it only fires once an allocation has already failed, so it reports an arrival rather than warning of an approach. Alert instead on the signals that engage before exhaustion: lattice_repocontext_heap_committed_bytes / lattice_repocontext_heap_limit_bytes for heap-ceiling adherence (check lattice_repocontext_heap_high_load_threshold_reachable first - a 0 means the runtime's own pressure threshold can never fire at any grant, issue #3133), orleans.lattice.wal.replay.permit_adaptations with outcome=withheld, trigger=occupancy for the proactive replay-concurrency backpressure, and orleans.lattice.leaf.snapshot.hydration_admissions with outcome=queued. All publish before they fire - the heap gauges from process start, the replay counter once the replay gate is first sized, and the hydration counter per tree at that tree's first leaf hydration, each arm primed at zero - so a flat zero is a measured zero, and an absent series means either that nothing has replayed or hydrated yet in this process or that the running image predates the instrument.
  • Build context differs per image: the host image's is the REPOSITORY ROOT (it ProjectReferences the just-built src/ bits), so its service sets context: ../..; the embedder builds from its own apps/embedding-onnx directory, which has no ProjectReference into src/ and so keeps a small context. Its one shared source folder, the container cgroup readers in src/lattice/Internal/Cgroups, arrives as the BuildKit named context cgroups that the embedder service declares under additional_contexts. Run compose from this directory either way.

The mounted workspace

Set REPO_PATH to the absolute path of a directory the box may see. It is mounted READ-ONLY at /workspace inside the container, so the box can never mutate the code it indexes. This is a workspace root, not a single repository: mount a broad parent and register individual repositories under it at runtime with the repocontext_add_repo tool. It defaults to ../../.. from this directory, which in an ordinary clone is this repository's parent, so this repo is one registerable child. From a git worktree the same default resolves to the worktree collection directory instead, silently - see the worktree trap - so set REPO_PATH explicitly, in .env or the environment.

export REPO_PATH=/absolute/path/to/some/parent    # PowerShell: $env:REPO_PATH="C:\path\to\parent"

The published port

The MCP endpoint and health probes are published on host port 8080 by default. That port is a common one to already have in use, and a clash shows up only as an opaque bind failure when you run up, so it is overridable with REPOCONTEXT_PORT:

export REPOCONTEXT_PORT=18080                     # PowerShell: $env:REPOCONTEXT_PORT="18080"
docker compose up -d

Only the host side moves; inside the container the listener stays on 8080. Every localhost:8080 in the walkthrough below then becomes localhost:$REPOCONTEXT_PORT.

Every path passed to repocontext_add_repo is resolved to its real location - .. traversal and symlink escape are both defeated - and must resolve under /workspace (set by LATTICE_WORKSPACE_ROOT); a path outside it is refused.

Choosing an embedding companion

The sample brings up the ONNX Runtime companion (apps/embedding-onnx) by default. It bakes its weights into an image layer, so a cold start needs no model download, and a cuda-flavoured build selects its accelerator - CPU or NVIDIA - at runtime via EMBED_PROVIDER (the default cpu build is CPU-only; see Running the embedder on an NVIDIA GPU).

The original Onyx companion (apps/embedding) remains available as a fallback, selected with an override file. Build the embedder when you switch: both companions build under the same compose image name, so a plain up -d would reuse the cached ONNX image rather than build the Onyx one (switching back needs the same rebuild):

docker compose -f docker-compose.yml -f docker-compose.onyx.yml build embedder
docker compose -f docker-compose.yml -f docker-compose.onyx.yml up -d

Nothing else changes in either direction. Both serve the same contract on the same port, so the service name and LATTICE_EMBEDDING_ENDPOINT are identical, and they produce numerically identical vectors (same pinned model revision, fp32, same tokenizer and pooling), so switching does not invalidate an existing /data volume. The ONNX image is roughly an order of magnitude smaller (about 1.3 GB against 13 GB).

Those commands are for the untuned walkthrough stack. On the tuned deployment keep -f docker-compose.tuning.yml in both commands, before the Onyx file, or the switch silently drops the image pin, the grants and the tuned cadence; see Rolling back the embedder for the three-file form and the -ExpectedConfigFileCount 3 the provenance check then needs.

Running the embedder on an NVIDIA GPU

The default build is CPU-only, and it stays CPU-only on a GPU host: the cpu flavour has no GPU support compiled in. Enabling the GPU needs all three of these together, because they do different jobs - the build arg picks a different ONNX Runtime package, the environment variable binds the accelerator at runtime, and the reservation is what actually exposes the device to the container. In docker-compose.yml under the embedder service:

    build:
      args:
        ONNX_FLAVOR: cuda
    environment:
      EMBED_PROVIDER: cuda
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

Then rebuild, because a plain up -d would reuse the cached CPU image:

docker compose build embedder
docker compose up -d

Requires the NVIDIA Container Toolkit on the host. The cuda flavour is a much larger image but still includes the CPU provider, so it serves CPU hosts too, and an unrecognised EMBED_PROVIDER falls back to the CPU rather than failing to boot. That fallback is silent by design, so verify what actually bound:

docker compose logs embedder | grep "Embedding server listening"
# ... listening on port 9000 using Cuda with model model.onnx (768-dim) ...

# Or ask the health endpoint. Both images are chiseled and carry no curl, and the
# embedder port is only exposed on the compose network, so borrow its namespace:
docker run --rm --network "container:$(docker compose ps -q embedder)" \
  curlimages/curl -s http://localhost:9000/api/health
# {"status":"ok","provider":"Cuda","model":"model.onnx","dimension":768}

The provider field is the provider EMBED_PROVIDER resolved to, reported once a session has been created on it. So a provider of Cpu here means EMBED_PROVIDER did not resolve to cuda - it is unset, or a value other than cuda, gpu or nvidia - and the fix is the service's environment. A missing toolkit or device reservation cannot produce Cpu: with cuda selected the server has no CPU fallback of its own, and a session it cannot create stops the process at startup, so check docker compose logs embedder instead. The Onyx fallback companion takes a different route to the same place (clear CUDA_VISIBLE_DEVICES and add the reservation); see apps/embedding.

Walkthrough

From this directory:

# 0. Copy the example environment file, then edit both paths in it for this
#    machine. REPOCONTEXT_MEMORY_ARCHIVE_PATH is REQUIRED - compose refuses every
#    command without it - and REPO_PATH picks the workspace root (see above).
cp .env.example .env

# 1. Start the stack. The host waits for the embedder to become healthy and for
#    the backup sink to start.
docker compose up -d --build

# 2. Wait for the host to come up. /health/live returns 200 once the process and
#    the silo host are alive, which is what the remaining steps actually need.
curl -fsS http://localhost:8080/health/live

#    /health/ready is a stricter, orchestrator-facing probe, and this walkthrough
#    deliberately does NOT gate on it. It is the conjunction of the lifecycle phase
#    AND the vector plane having demonstrated a working semantic query, so it can
#    stay 503 long after the box is up and answering MCP calls. Observe it, but do
#    not wait on it or treat a 503 as a failed deployment - read "Interpreting a
#    persistent 503" below first. The -o/-w form reports the code without failing
#    the shell, which `curl -fsS` would do on any non-2xx.
curl -sS -o /dev/null -w 'ready: %{http_code}\n' http://localhost:8080/health/ready

# 3. Register a repository under the mounted workspace over MCP (repocontext_add_repo
#    with a path under /workspace). Use your MCP client of choice against
#    http://localhost:8080 (the MCP streamable-HTTP endpoint). For example, with
#    the reference `mcp` CLI:
#      mcp call http://localhost:8080 repocontext_add_repo '{"path":"/workspace/my-repo"}'
#    Omit repoId to derive it from the final path segment, or set it explicitly:
#      mcp call http://localhost:8080 repocontext_add_repo '{"path":"/workspace/my-repo","repoId":"demo"}'
#    List what is registered at any time:
#      mcp call http://localhost:8080 repocontext_list_repos '{}'

# 4. Recall: query the box (repocontext_search / repocontext_recall) and confirm it
#    returns the ingested context.

# 5. Restart the container - a full process restart (not a recreation: the same
#    container and its /data volume are kept) that evicts the in-memory projection
#    and forces a WAL replay / cold rebuild on next access.
docker compose restart repocontext
curl -fsS http://localhost:8080/health/live
#    As in step 2, /health/ready may stay 503 after the restart without meaning the
#    restart failed. Step 6, not the probe, is the proof that the data survived.

# 6. Recall again. The context is still present: it was replayed from the WAL and
#    SQLite state on the /data volume, proving durability across a restart.

Retrieval and token economics

Once a repository is registered (step 3 above), the same box exposes the epic's retrieval and token-economics tools over the same MCP endpoint - no extra service, no second listener. Every call below targets the demo repo id from step 3; substitute your own. Examples use the reference mcp CLI against http://localhost:8080.

# A. Explainable search. Every hit carries a machine-readable `reasons` array
#    saying WHY it ranked (semantic proximity and matched chunk/symbol, or the
#    specific keyword fields hit - path/name, symbol, tag, topic, content, key),
#    so an agent can justify a selection instead of trusting an opaque score.
mcp call http://localhost:8080 repocontext_search \
  '{"repoId":"demo","query":"where is the readiness health probe wired","k":5}'

# B. Budgeted context bundle. repocontext_context packs the ranked, explained
#    source for a task into ONE response under a HARD token ceiling: the reported
#    `responseTokens` (the response as delivered, envelope included) never exceeds
#    the ceiling it reports as `budgetTokens` - the `responseBudgetTokens` you asked
#    for, clamped to 1-200000 and 8192 when omitted - `totalTokens` is the narrower
#    sum of packed source, and `truncated` /
#    `retryBudgetTokens` say whether more would fit at a larger budget. `detail`
#    trades richness for budget - 'paths' (cheapest) -> 'outline' (declared-symbol
#    skeleton) -> 'slices' (bounded body text, richest), or 'auto' (default) which
#    picks the richest level that fits. Pass a `session` id so the box remembers
#    what it delivered.
mcp call http://localhost:8080 repocontext_context \
  '{"repoId":"demo","task":"explain the readiness health check","responseBudgetTokens":4000,"detail":"auto","session":"agent-1"}'

# C. Reuse economics. Repeat on the SAME `session`. Units the session already
#    holds are suppressed - acknowledged under `reused`, never re-charged and never
#    counted against `top` or the budget - so the second answer pays only for the
#    NEW context. (You can also feed the prior entries' unit receipts back via
#    `seen`, or a whole-file 'path@hash' claim via `known`; the server-side
#    `session` bookkeeping does it for you.)
mcp call http://localhost:8080 repocontext_context \
  '{"repoId":"demo","task":"explain the readiness health check and how drain flips it","responseBudgetTokens":4000,"detail":"auto","session":"agent-1"}'

# D. Usage accounting. repocontext_stats reports the aggregate token economics over
#    a bounded recent window: calls answered, response tokens spent, whole-file
#    reads replaced, and the NET tokens saved by budgeting plus reuse.
mcp call http://localhost:8080 repocontext_stats '{}'

With the embedder companion healthy, search and the bundle rank semantically; with it unreachable they degrade to a deterministic keyword rank - the bundle still answers either way. See docs/lattice.api.mcp.repocontext/retrieval-economics.md for the full model.

Tear down (state on the named volumes is preserved unless you pass -v):

docker compose down          # keeps the data volume (and the Onyx overlay's model cache, if used)
docker compose down -v       # also deletes durable state (start clean)

down -v deletes the index and the authored agent memory, because both live on the same /data volume. Read Agent memory versus the code index before using it: repocontext_reset_index rebuilds an index with no loss at all, and the memory archive that makes down -v survivable is bounded by its export interval rather than complete.

Agent memory versus the code index

Two kinds of state share the /data volume, and only one of them can be recreated.

What it is If it is destroyed The safe gesture
Code index Structural, content, symbol, xref, session and vector planes, derived from files on disk Re-run repocontext_add_repo; back in minutes repocontext_reset_index
Agent memory Every repocontext_remember note, decision, gotcha and glossary entry Gone; it is the store of record and derives from nothing Keep an archive (below)

docker compose down -v destroys both. That is the defect behind issue #2601: the gesture is documented as the ordinary way to start clean, and it silently takes the irreplaceable half with it. It has already happened once, to epic #2368's own memory.

Why the two are not simply on separate volumes

Because they cannot be, and because it would not have helped.

They cannot be: a Lattice tree's durable state spans a WAL root that is one directory for the whole storage provider, and a grain store that is a single SQLite file shared by every tree. The /data/wal/repo-context-* subdirectories look like separable locations but are a naming convention inside one root, and the memory tree's pages sit interleaved with every other tree's in /data/repocontext.db. There is no memory-only path to mount elsewhere.

It would not have helped: docker compose down -v removes every named volume the project declares, not just the one you had in mind. A second declared volume dies in the same command as the first.

What actually protects it

A bind mount, /memory-archive, which is not a project-declared volume and so is not removed by down -v. The host exports memory there periodically and, when it starts against an empty store, restores from it.

# The archive path is REQUIRED and has no default. A relative one would resolve
# against the directory you invoked compose from, which is how the only working
# backup of durable agent memory ended up inside an ephemeral git worktree
# (issue #2627). Set it to an absolute path outside every checkout and worktree.
REPOCONTEXT_MEMORY_ARCHIVE_PATH=/srv/repocontext-memory docker compose up -d

Copying .env.example to .env sets it for you; compose loads .env on every command, so up, down, ps, and logs all pick it up. Without it, every compose command in this directory fails by name rather than quietly choosing a directory nobody picked.

Verify where it actually landed, rather than where you meant it to land:

docker inspect "$(docker compose ps -q repocontext)" \
  --format '{{range .Mounts}}{{.Destination}} <- {{.Type}} {{.Source}}{{"\n"}}{{end}}'
pwsh -File scripts/Assert-ContainerProvenance.ps1   # check 5 of 7 refuses a doomed path
Variable Default Meaning
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR /memory-archive in this sample; unset (feature off) otherwise Container path the archive is written to. Unset disables the whole mechanism.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_INTERVAL_SECONDS 300 Export cadence. This is the size of the window an ungraceful stop loses. A positive value below 30 is raised to 30, and one above the longest delay a timer can wait (about 49.7 days) is lowered to it; a zero, negative or unparseable value, or one too large to represent as a duration (for example 1e20), falls back to 300.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_RESTORE auto auto restores into an empty store, or into one whose restore-state marker records that an earlier restore was left partial; always restores on every start; off never restores; none and false are accepted for off and on-empty for auto, and any other value falls back to auto.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_STOP_TIMEOUT_SECONDS 20 Budget for the final export during a graceful stop. A positive value is clamped to 1-60; a zero, negative or unparseable value falls back to 20.
REPOCONTEXT_MEMORY_ARCHIVE_PATH none - required HOST path bound at /memory-archive. Deliberately has no default: a relative one resolves against the compose invocation directory (issue #2627). Must be absolute and outside every checkout and worktree.

Restore this way by hand at any time - stop the box, put the archive files in place, start it against an empty store:

docker compose down -v
ls "$REPOCONTEXT_MEMORY_ARCHIVE_PATH"/   # repo-context-memory.snapshot (+ .previous.snapshot)
docker compose up -d
docker compose logs repocontext | grep -i -E 'memory durability|durable memory'

The memory durability lines are the startup statement of where memory lives and what protects it; they do not say whether anything was restored. The restore outcome is logged separately, a few seconds after start, on a line naming Durable memory - Durable memory was restored from the archive at ... when records were merged back - and that is the line that confirms the import.

What this does not do

It is not a backup, and it does not make down -v safe.

  • Everything written since the last export is lost. The exposure is the export interval, plus whatever a non-graceful stop discards. A graceful stop (down, down -v, stop) exports once more on the way out and closes most of that window; a kill -9 or a host crash does not.
  • It archives memory only. The index is not in the archive, by design - it rebuilds from source.
  • It does not remove the co-location. Memory still shares a volume with rebuildable state, and the host says so at startup, at warning level, every time. repocontext_reset_index remains the correct way to rebuild an index: it drops the derived planes and preserves memory outright, with no window at all.
  • It is not the scheduled backup that issue #2602 added (the azurite-backup-sink service). That one captures the same memory tree to an external sink, with manifests and retention, and restores only the backup id an operator names; this one owns automatic restore-on-empty for memory alone. Only this mechanism restores automatically at startup.

See docs/lattice.api.mcp.repocontext/memory-durability.md for the full model.

Health probing

The runtime image is distroless and shell-less, so probing is HTTP-only - there is no shell-exec healthcheck:

  • GET /health/live - process + silo host alive (liveness). This is the probe the walkthrough gates on, and the one an orchestrator uses to decide whether to restart the container.

  • GET /health/ready - readiness (routing). On this sample's local durability profile it is the conjunction of two independent components, and both must be healthy for a 200 (the Azure profile adds a third, the scaling-signal check):

    • lifecycle - the silo has joined, the activation-time WAL replay is done, the durable stores were proven reachable, and MCP is serving;
    • vector plane - semantic retrieval has been demonstrated to work. A deployment with no embedder bound (keyword-only), and a host with no repository registered yet, both count as ready here: there is no vector plane to wait for in the first case and nothing to serve in the second. This host always binds its embedding provider, though, so a missing or unreachable embedder does not make it keyword-only: searches report keyword.vector_plane_unavailable and, once a repository is registered, readiness stays down.

    So it is not-ready during startup replay and during drain, but those are not the only causes, and a sustained 503 is far more likely to be the vector-plane component than either of them.

  • GET /health/silo - grain liveness. It re-checks silo membership and makes a trivial grain call on every request, so it goes red for a silo that died after reaching readiness while the process kept listening (issue #2666) - which neither /health/live (always green) nor /health/ready (which never re-checks the silo once it has flipped) can do. It answers 200 only when healthy and 503 while the silo is still starting or once it is unhealthy, with the three-way verdict in the body. It is what the container's Docker healthcheck targets, through the exec-form --healthcheck self-probe the shell-less image needs, and it feeds neither of the two probes above. Docker only records that verdict: under this sample's restart: unless-stopped a container is restarted when its process exits, never because it is unhealthy.

  • GET /health/backup - whether the durable agent-memory tree is actually being captured to the azurite-backup-sink service. It is tagged as neither liveness nor readiness, so a failing backup never restarts the container or pulls it from rotation; the body names the tree in scope, what the sink holds and the last failure text, and it answers 503 only for an unhealthy verdict.

Interpreting a persistent 503

A /health/ready 503 that does not clear, on a container that is otherwise up, does not on its own mean the deployment is broken, and must not be used by itself as a rollback signal. The response body names each component and its verdict on its own line beneath the aggregate status, so the component holding readiness down is readable straight from the probe (issue #2962). That names the component, not the cause, so the steps below still apply: read the body first, then work through them to narrow why that component is unhappy. Before issue #2962 the endpoint returned a bare Unhealthy with no per-component breakdown, so a 503 was ambiguous until it had been narrowed by hand. Five steps, each one ruling out a cause the previous step left open:

  1. curl -fsS http://localhost:8080/health/live. A 200 says the process and the silo host are alive, so whatever is unhealthy is not the process. If this also fails, the container really is unhealthy - that is the case to act on.

  2. Make any MCP call against http://localhost:8080 (repocontext_list_repos is the cheapest). If it answers, the MCP surface is serving, which satisfies the lifecycle component and leaves the vector plane as the one holding readiness down.

  3. Run a repocontext_search and read the retrievalPath on the result. A value of keyword.vector_plane_unavailable confirms it: semantic retrieval is unavailable and the box has fallen back to deterministic keyword recall.

  4. Check the embedder with docker compose ps. Step 3 tells you the vector plane is at fault but not which side of it, and the two sides need opposite responses. An embedder container that is missing, exited, or (unhealthy) is itself the cause, and is directly actionable: restore it and readiness can recover on its own. An embedder reporting (healthy) while readiness stays 503 rules the embedder out and places the fault host-side, in the vector plane, where restarting the embedder achieves nothing. Use docker compose ps rather than probing the embedder directly - its port is not published to the host.

  5. Separate "never been ready" from "was ready and has since lost it" on /metrics. The two need different responses and steps 1 to 4 cannot tell them apart:

    curl -fsS http://localhost:8080/metrics \
      | grep -E 'repocontext_retrieval_(ready_seconds|unavailable)'
    

    repocontext_retrieval_ready_seconds_count is stamped once per process, on the first transition into a ready phase. Its absence therefore means the retrieval plane has never been ready in this container's current lifetime; its presence alongside a 503 means the plane was ready and has since lost it. Its phase label records which phase it first reached (serving, keyword_only, or nothing_registered). repocontext_retrieval_unavailable_total counts fault episodes under a closed cause label: the three capability-loss values of step 3's retrievalPath, so it separates a vector plane that cannot serve (keyword.vector_plane_unavailable) from an index that has drifted from its sources (keyword.index_degraded) and from a withheld exact fallback (keyword.exact_fallback_suppressed), plus probe (a readiness probe rather than a real query saw the plane unable to serve), saturated (an admission gate refused the plane's open past its bound) and unknown.

Issuing a query yourself does not clear it, and the host is already trying. A warmup service issues the same semantic query from application start, retrying with backoff (waits of 2, 4, 8, 16 and 32 seconds, then every 30 seconds) until the plane answers or shutdown begins, and once it has answered it re-checks readiness every 30 seconds and re-drives the query whenever readiness has been revoked. So a persistent 503 is never "nobody has queried it yet" - it is that warmup failing repeatedly. In particular, a box that has a repository registered but holds no vectors for it stays not-ready by design: the search reports keyword.vector_plane_unavailable, and running another search by hand returns the same thing and changes nothing. (A box with no repository registered is the opposite case and reports ready, because there is nothing it could be asked to serve.)

Readiness also lags a fault on purpose. Once the plane has served, a fault must persist for 30 seconds before readiness is revoked, and any successful retrieval inside that window clears the episode outright - so a 503 can appear up to half a minute after the fault that caused it, and a brief blip may never surface at all.

In that state the box is still usable and the whole walkthrough still completes: registration, keyword search, repocontext_context, and durability across a restart all work, and steps 3 to 6 demonstrate exactly that. What is degraded is semantic ranking, not the service. Treat it as a capability to restore, not as a deployment to roll back.

The same listener also serves GET /metrics, a Prometheus text exposition of every instrument on a meter whose name starts with orleans.lattice - the core meter and every per-package meter, Orleans.Lattice.Api.Mcp.RepoContext included - plus the Microsoft.Orleans and System.Runtime runtime meters. It needs no second port and no sidecar:

curl -fsS http://localhost:8080/metrics | head -n 20

Notes on durability and shutdown

  • All durable local state (the WAL directory and the SQLite database file) lives under /data, a named volume. The host fails fast at startup if that path cannot be created or is not writable by its non-root UID; a missing directory is created rather than refused.
  • On SIGTERM (a docker stop / restart) the host flips readiness to not-ready first, then drains: the silo deactivates and the WAL commit-log flushes buffered records before exit, so an in-flight write is durable after restart.
  • PID 1 is an init process, and that is what makes the SIGTERM land at all. init: true in docker-compose.yml has Docker bind-mount its own static docker-init binary and run it as PID 1, with the host as its child. It needs nothing in the distroless runtime image and changes no application code. Two kernel behaviours make it necessary, and both attach to PID 1 rather than to the application: no default action is taken for a signal delivered to PID 1 that PID 1 has installed no handler for, so a well-behaved process can be unkillable by SIGTERM purely by being PID 1; and PID 1 inherits every orphaned descendant and must wait() on it, which the .NET host does not do. In the epic #2368 gate runs a container reached a state where neither docker kill nor docker rm -f would reap PID 1 and it had to be SIGKILLed, costing that run its drain and leaving the next one unbanked state to replay (issue #2576). This is independent of the grace period below: init decides whether the drain starts, the grace period decides how long it may take, and setting one without the other leaves half the failure in place.
  • That drain's budget is 180 seconds in this sample, and it belongs to the host, not to Docker. The host sets HostOptions.ShutdownTimeout from the grace period the deployment declares - 180s from this sample's declared 240s, or the 90s default (RepoContextHostBuilder.ShutdownBudget) when nothing is declared; Docker's stop_grace_period defaults to 10 seconds. The two are enforced independently and the smaller wins, so without an explicit stop_grace_period the process is SIGKILLed at 10s with the drain still running and the budget is dead configuration (issue #2389). The stop_grace_period: 240s in docker-compose.yml is what makes it reachable. If you run this image under your own orchestration you must grant the same budget there - Kubernetes has the identical trap under a different name, since terminationGracePeriodSeconds defaults to 30s.
  • The drain reports its own duration, so the budget can be derived rather than bisected. docker logs carries `RepoContext drain complete in s, consuming

    % of the 180s host shutdown budget`. Read it together with its severity, because there are three distinct outcomes and the level is what separates them: - **No completion line at all.** The container was killed mid-drain, so `stop_grace_period` is smaller than the drain (issue #2389). The exit code will not tell you, because a killed container reports `137` and the next `docker start` overwrites it. - **`drain complete` at `Warning`.** The drain finished but consumed 70% or more of the budget. Nothing has failed; treat it as a lead indicator, because drain time grows with resident state. - **`drain ABANDONED after 180s` at `Error`, and the container exits `70`.** The *host* stopped waiting and deactivation was abandoned part-way. Raise the service's `stop_grace_period` and the `LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD` that declares it, together and to the same value; the budget is derived from the second and bounded by the first, so raising either alone achieves nothing.

  • The abandoned drain is also visible without reading the log (issue #2401). An orchestrator does not read logs, it reads the exit code, and before #2401 the abandoned case did not reliably produce a distinctive one. Measured against a real generic host, the outcome was not simply zero but undetermined: a hosted service that absorbed the shutdown cancellation left RunAsync returning normally, so the process exited 0 and the abandonment was recorded as a clean stop; one that rethrew it let the exception escape RunAsync unhandled, aborting the process in a way indistinguishable from a real crash. The host now assigns the code itself when the overrun latches: 0 if the drain completed inside the budget, 70 if it was abandoned. 70 is EX_SOFTWARE in the sysexits.h convention, chosen to avoid 0/1/2, Docker's reserved 125-127, and the 128 + signal band that holds 137 (SIGKILL) and 143 (SIGTERM) - the neighbouring conditions it exists to be told apart from.
    • Note what this does not do. docker-compose.yml uses restart: unless-stopped, under which Docker restarts on any exit code, so the code neither triggers nor suppresses a restart. What it changes is what is recorded: docker inspect --format '{{.State.ExitCode}}' reports 70, docker ps -a shows Exited (70), and under Kubernetes the container terminates with reason Error rather than Completed. That is what an alert can be written against.
    • There is deliberately no way to turn it off. A switch restoring 0 would remove the evidence rather than the problem.
  • That last line exists because of issue #2397, and the reason it is needed is not obvious. The host raises ApplicationStopped even when the shutdown budget expired and it gave up waiting - so a signal bound only to that event reported drain complete in 90.0s for a drain that did not complete. The failure looked like success. The overrun is now raised by an alarm armed when the drain begins, so it is reported at the moment the budget expires rather than depending on a callback that may never arrive.
  • Drain time scales with resident state: the same 400-file rig drained in 33.9s before its vector trees had landed and 67.2s once they had, which was already three quarters of the 90s budget the host then ran with. #2397 nevertheless did not raise it, because measurements on a live box show the resident set that a drain must flush has no observed ceiling (idle-deactivation sweeps ranging from 1 to 4,418 activations, still climbing between readings). A fixed ceiling on an unbounded quantity moves the threshold without changing the failure mode, so #2397 shipped the diagnostic instead of a new number, and #2402 - which proposed raising it - did not ship one either. Issue #3304 later did raise it, to 180s, against two consecutive drains that no longer fitted (89.7s and 91.9s against 90s); see the container quickstart.
  • The budget is no longer written down as an independent constant. Since issue #2402 the host derives it from the grace period the deployment declares through LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD, taking 75% of it or all but a two-second unwind reserve, whichever is smaller. The 120s the sample declared when that derivation landed yielded exactly the 90s the container had always run with, so nothing moved at the time (the sample now declares 240s, deriving 180s, since issue #3304); what changed is that there is one number to set instead of two, and a budget larger than the grace period can no longer be expressed. That matters because such a budget is not merely useless - the container kills the process at the real grace period regardless, so the ABANDONED line above is armed for an instant that never arrives and is never emitted, which is the silent teardown of issue #2389 all over again.
  • The variable declares the grant; it is not the grant. Nothing inside the container can read the real stop_grace_period, so a deployment that declares 120s while granting 20s runs a 90s budget under a 20s guillotine and cannot detect it. Keeping the two adjacent in the same compose service is the mitigation, and RepoContextComposeShutdownBudgetTests asserts they are equal here - but that adjacency is a convention, not an enforcement. Change them together, always.
  • None of this bounds the resident set. Raising both values past your own observed drain is a legitimate local remedy, but it buys time rather than fixing the shape.

Verifying what you actually deployed

Everything above describes what the tracked compose file declares. Nothing above establishes that a container now running received any of it.

docker compose up reads the compose files in its own working directory, whatever branch built the image it starts, and its output names no branch, no commit, and no directory. The image and the runtime configuration are therefore two independent inputs, and only the first is obviously version-controlled. An operator standing in one checkout can deploy a candidate image under a different checkout's configuration and see nothing at all to say so.

That is not hypothetical. Two gate runs of epic #2368 did exactly this: the candidate image ran under the baseline's runtime config, the container's own compose label resolved to the main checkout, and LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD was absent from the process environment while sitting merged on the candidate branch. Both runs observed the absence of the fix's effect and concluded the fix was absent. The observation was real and correctly made. The discriminator was in a channel nobody was reading.

Note what would not have helped. A check comparing tracked files to other tracked files would have been green throughout both runs, because the repository agreed with itself perfectly. Only a reading taken from the running process separates "the source does not carry the fix" from "the source carries it and this container never received it".

cd samples/RepoContextContainer
pwsh -File ./scripts/Assert-ContainerProvenance.ps1

It exits 0 only when every check agrees, and with a distinct non-zero code on each refusal path (2 refused, 3 the container could not be interrogated, 4 the expected configuration could not be read), so automation can gate on it. Until #2718 it never called exit at all and leaked 128 from a git probe that is SUPPOSED to fail; the full contract is in the local deployment runbook.

It performs seven checks against the running container and refuses unless all seven agree, printing every value it read either way:

  1. Compose provenance. The container's own com.docker.compose.project.working_dir label resolves to the checkout you are standing in, every file in com.docker.compose.project.config_files lies under it and exists on disk, and the number of those files is the number you expected (default 2, see below). An override file merged in from elsewhere is how a resolved document stops matching the tracked one.
  2. Git provenance. That directory is a git worktree and its HEAD is the commit you expect, reported as a value you read rather than an inference you make. The expectation is sourced from your checkout, never from the one the container resolved to; defaulting it to the latter would compare a value against itself and pass unconditionally, which is worse than omitting the check because it would report as checked.
  3. Image provenance. The image id the container is executing still matches what its image reference resolves to now, which catches a container left in place across a rebuild. Ids, never tags: a tag is a mutable pointer, so comparing tag to tag compares two names for whatever is current.
  4. Environment provenance. A candidate-only setting is present in the container's own environment with the expected value, expectation read from your checkout's docker-compose.yml and actual read from the running process. This is the non-redundant one. Checks 1 to 3 can all pass while an override file, an edit, or a stale container leaves the value unset, and it is the direct executable form of the warning above that the variable declares the grant rather than being it. Durations are compared parsed, never literally: Compose normalises 120s to 2m0s, so a text comparison would accuse a correctly configured stack of exactly this defect, and the obvious remedy for that accusation is to change a deployment that was already right.
  5. Archive durability. The bind mount holding durable agent memory resolves to an absolute host path outside every git worktree and checkout. It is the only check on the WRITE path, and it keys on the archive path alone, never on the compose directory (issue #2627).
  6. Build provenance. The image the container is executing was built from the expected commit, read from the image's own org.opencontainers.image.revision label or failing that a candidate-<sha> tag, and cross-checked against chronology. It fails closed. Checks 1 and 2 adjudicate a checkout; this one adjudicates the image (issue #2686).
  7. Workspace provenance. The /workspace bind the container indexes out of exists, is a bind rather than a volume, and resolves to an absolute host path - and, when you name them, it is the expected workspace root and a registered repository is indexed from the expected repository root (issue #2617).

The parameters not discussed below steer what those checks read. -ContainerName names the container to interrogate; by default it is the one carrying the compose service label for -ComposeServiceName (default repocontext). -ExpectedCheckout (default: the checkout holding the script's compose files) is the directory check 1 compares against and the checkout whose docker-compose.yml check 4 reads its expectation from, and -ExpectedCommit (default: that checkout's HEAD) is the commit checks 2 and 6 compare against; -ExpectedCommitDate (default: that commit's date), -CandidateTagPrefix (default candidate-) and -ClockSkewToleranceSeconds (default 120) feed check 6's tag fallback and chronology arm; and -ArchiveDestination (default /memory-archive) and -WorkspaceDestination (default /workspace) are the container paths checks 5 and 7 adjudicate.

The count assertion, and why it is not a walk of tracked files

The stack's real deployment is two compose files. The tracked docker-compose.yml carries a build: stanza and no image:, so on its own it cannot resolve an image for up -d --no-build at all. A second file supplies the image pin, the memory limit, the CPU caps, and the scan-cadence variables every prior measurement on a given box was taken against.

That second file used to be an untracked, gitignored docker-compose.override.yml, which put the entire tuned configuration on exactly one machine with no review, no history, and no diff (issue #2609). It is now the tracked docker-compose.tuning.yml, layered by name:

docker compose -f docker-compose.yml -f docker-compose.tuning.yml up -d --no-build

The rationale behind every value in it, the build-and-tag ladder that produces the image pin, and the rollback procedure are in the local deployment runbook.

The count is still two, so this check needs no new argument. Naming the files explicitly also suppresses an automatic docker-compose.override.yml, which is what makes the deployed configuration reproducible from the checkout alone; layering a personal override on top of the tuning file makes it three, and you say so with -ExpectedConfigFileCount 3.

The opt-in CPU pinning described above does not move the count: it is two variables on the existing services, not a third file.

A machine-local override remains legitimate, and it is why check 1 reads the container's label rather than walking the repository: it has to be able to fail on a file git has never heard of. The obvious remedy for a compose-provenance failure - relaunch from the checkout you meant - silently drops whichever second file the original launch had, leaving no image pin, no memory limit and a different scan cadence, while the tree looks perfectly correct and every path the container reports still resolves under the right directory. Only the count dissents.

The general form is worth stating, because it is not specific to compose: fixing a provenance defect by changing the launch directory is itself a provenance change, and it is not self-verifying.

Pass -ExpectedConfigFileCount 1 if you genuinely mean to run without an override. Making that an explicit act is the point - dropping the override should be something you said, not something that happened.

What a green run does not establish. That the checkout was clean when up ran, since HEAD is a commit and uncommitted compose edits are invisible here - and a GIT_COMMIT stamped from a dirty tree names a commit the image does not exactly contain. That any setting you did not name reached the process: check 4 proves only the settings passed as -ExpectedSetting. That the memory archive's content is good, or that the indexed workspace is current: checks 5 and 7 adjudicate where the archive landed and which tree is bound, not what state either is in. That the workspace root or the indexed repository is the one you meant, unless you named them with -ExpectedWorkspaceRoot and -ExpectedRepositoryRoot (the report prints NOT ESTABLISHED for each arm you did not ask for; -ExpectedRepositoryRoot also needs -IndexedRoot, the indexedRoot values repocontext_list_repos reports, and refuses without it rather than passing), or that every registered repository is correctly rooted - the indexed-root arm passes when one matches. Or anything whatever about a container you did not name. What a green run does establish, since check 6 (issue #2686), is that the image the container is executing was built from the expected commit, on the evidence of its own revision label or candidate-<sha> tag cross-checked against chronology. Check 3 as defaulted detects only a stale container; pass -ExpectedImageId if you need to pin an exact image.

This is an operator check and is deliberately not wired into CI. It needs a running container, and a fixture that skipped when Docker was absent would produce exactly the false green it exists to prevent.

It is also not the same instrument as the cold-start rig's Assert-RigComposeIsolation, and neither subsumes the other. That guard validates the declaration - what docker compose config resolved - before anything runs. This validates the deployment - what a container already running actually received. The rig has never had this failure mode, because Get-RigComposeFile pins the file it resolves, so the different-checkout drift cannot arise there. A green rig guard therefore says nothing about this class, and the two must not be collapsed. The boundary was drawn deliberately by issue #2576, whose handoff named both the remedy and its location: a post-up precondition on the container's own environment.

The adjudication is pure and separated from the acquisition, so the refusing direction is exercised against fabricated disagreements rather than assumed:

pwsh -File ./scripts/Test-ContainerProvenance.ps1

Every one of the seven checks has fixtures it accepts and fixtures it refuses, and several refusal fixtures are reconstructed from real gate-run and incident readings rather than invented. A check only ever observed passing is indistinguishable from one that cannot fail, which is the same reason the suite itself is worth measuring rather than trusting: commit first, then make one check return no violations unconditionally, re-run, and confirm the assertions that fail are the ones covering that check and no others.

Two sibling suites cover what that pure suite cannot see, since it neither runs git nor runs the script: Test-ArchiveGitReading.ps1 checks the archive check's pinned git messages against the git actually installed, and Test-ProvenanceExitCode.ps1 checks that the script's exit status says what its printed verdict says.

Measuring approximate retrieval on a running container

scripts/Invoke-AnnQueryProbe.ps1 issues real retrieval queries against a running container and reports what moved on repocontext_retrieval_ann_search_total. It exists because a scrape on its own cannot answer the question it appears to answer: that counter is written only on the per-query path, so a deploy that never issued a query leaves every arm at its primed zero, and that reading is byte-identical to a plane that was consulted and answered exactly nothing approximately.

pwsh -File ./scripts/Invoke-AnnQueryProbe.ps1
pwsh -File ./scripts/Invoke-AnnQueryProbe.ps1 -RepoId lattice -Repetitions 4

It scrapes the three state arms before and after, issues its queries over the MCP endpoint on 8080 (the container's only application listener), and reports the per-arm delta.

With no -RepoId it probes every repository repocontext_list_repos reports. Its other parameters are -BaseUri (default http://localhost:8080; pass the published port if you moved it with REPOCONTEXT_PORT), -Queries (a built-in spread of natural-language queries when omitted, each issued against every selected repository), -Repetitions (default 1), -K hits per query (default 5), -TimeoutSeconds per request (default 60), and -JsonOutputPath, which writes the full machine-readable result document.

It refuses to report rather than reporting a zero it cannot stand behind. The refusal is the feature, and there are two of them, kept deliberately distinct because they have different owners:

  • issued 0 (exit 2) - the probe never got a query out. Nothing can be concluded about the instrument; the fault is the probe's or the environment's.
  • issued N, succeeded 0 (exit 2) - queries went out and every one failed. The instrument reading is still inadmissible, but the fault is now the container's and is worth diagnosing.

Collapsing those two into one "no data" would discard exactly the bit that says whose problem it is. Neither is the environment failure that stops the probe before it can query at all - a container that does not answer /health/live, a failed MCP handshake, or a failed repocontext_list_repos - which exits 3.

retrievalPath on a search result is not evidence about the approximate arm. The approximate index's declared retrieval path is a property of the index, not of a query - one index serves every repository, so a state-tracking declaration would be wrong the moment two repositories were in different states. It therefore reads semantic.approximate unconditionally, including when the exact fallback answered with complete recall, and NormalizeSemantic fails closed the same way by resolving anything unrecognised to it. The declaration deliberately under-promises. Reading it as confirmation that approximate search ran is confidently wrong, and the probe prints that caveat rather than assuming you know it.

An absent arm is not a zero. The probe distinguishes "the series is present and reads 0" from "the series is not on the endpoint at all". The first is a measurement. The second means the series was refused at creation, and the probe sends you to lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total before you conclude anything from it.

It reads /health/ready, and prints the answer verbatim. This is a different endpoint from /health/live and answers a different question. Liveness asks whether anything is there; readiness asks whether this box can actually serve semantic retrieval, and on the repocontext host it is the conjunction of the lifecycle component and the vector plane. When the vector plane is down the readiness body already says so, in specific and self-limiting terms, and it even names the retrievalPath discrimination you would otherwise have to rediscover.

The endpoint is easy to miss, and has been missed: the container healthcheck runs a grain-liveness self-probe rather than an HTTP readiness call, so Docker can report healthy straight through a total retrieval outage, and the acceptance playbook's only outbound call is /metrics. The probe therefore reads it explicitly rather than assuming something upstream already did.

A 503 here is a successful probe result, not a probe failure. It is the system diagnosing itself, which is more authoritative than anything this harness can infer from a counter delta, so the probe prints the status and the full body and says as much. It is deliberately not a gate: a 503 is the expected reading on a rig whose vector plane is down, and refusing to continue would suppress the very measurement the harness exists to take.

The total across all three arms is the liveness witness. The arms partition the whole query population - the index records an outcome for every query including bootstrapping - so a moving total proves the instrument is capable of reporting, independently of which arm moved. A zero on approximate beside a non-zero total is a measured absence. A zero beside a zero total is not a reading at all. Because the arms carry no caller tag, a total delta larger than the probe's own succeeded count is reported as CONTAMINATED rather than claimed: concurrent internal retrieval is indistinguishable from the probe's own, and pretending otherwise would attribute traffic the probe did not generate.

The probe does not deploy, does not score any acceptance predicate, and does not tune retrieval parameters. Tuning until the approximate arm fires would encode the answer into the instrument.

Its own refusal paths are regression-tested rather than proven once:

pwsh -File ./scripts/Test-AnnQueryProbe.ps1

Twelve scenarios against a real HTTP listener that the suite runs as a background job (on localhost, port 18080 unless you pass -Port), invoking the probe as a separate pwsh process, covering both refusals, the absent-arm case, the contaminated delta, the suppressed-fallback state, and all three readiness shapes (ready, not-ready-with-a-diagnosis, and a readiness endpoint that cannot be read at all). The suite asserts its own scenario count is non-zero before reporting, for the same reason the probe asserts its issued count: a harness that ran nothing reports success in a way that is indistinguishable from a harness that ran everything and found nothing wrong.

Measuring an approximate-index build

scripts/Invoke-AnnBuildProbe.ps1 measures how long the approximate index takes to converge for one repository on a running container - the wall-clock figure an A/B of the build path is scored on. It registers the repository itself with repocontext_add_repo (a write, and for a repository that is already registered a fresh indexing pass), then polls repocontext_health for that repository every -PollSeconds (default 10) and reports the time to converge, the vectors indexed, and the sample series (written as JSON with -JsonOutputPath).

pwsh -File ./scripts/Invoke-AnnBuildProbe.ps1 -RepoPath /workspace/my-repo

-RepoPath is the in-container path under the mounted workspace, -RepoId defaults to its final segment, as repocontext_add_repo itself derives it, -BaseUri defaults to http://localhost:8080, and -TimeoutSeconds (default 240) bounds each MCP request it makes.

Convergence is not the approximate plane reporting Ready on its own. A build over a corpus that has not been embedded yet reaches Ready at once with nothing in it, so the probe waits for Ready with the indexed vector count caught up to a non-zero embedded coverage. It exits 0 when that happens and 2 when -MaxWaitMinutes (default 60) elapses first - an arm that does not converge is a result to record, not a harness fault - and it fails outright if the registration still fails after its retries.

It deliberately does not read /health/ready, which answers a different question (see above), and it scores on time to converge rather than on the repocontext.ann.build.stage.duration histogram, because only the former is reported by every build an A/B might compare. Read the stage split afterwards to explain a difference, not to score one.

Scripts

Every script in scripts/, and every parameter it accepts. All parameters are optional except -RepoPath on Invoke-AnnBuildProbe.ps1.

Script What it does Parameters (default) Described in
New-TuningEnv.ps1 Derives this host's resource knobs for docker-compose.tuning.yml and writes them to .env. -WorkspacePath (the repository root), -OutFile (.env beside the compose files), -DryRun, -CorpusOnly (measure and print the corpus, then exit), -IgnoreHostLoad, -ExpectedCorpusFiles, -CorpusTolerance (0.02), -Force, and two seams for driving its refusals deterministically, -HostMemoryBytes and -HostAvailableMemoryBytes Runbook: before you start the stack
Assert-TuningEnv.ps1 Refuses a tuning .env whose knobs are unset or carry a retired sentinel. -EnvFile (.env beside the compose files), -Reading (a hashtable checked instead of a file), -BuildCommit, -RepositoryPath (the repository holding the script), -Quiet Runbook: before you start the stack
Test-TuningIntegrity.ps1 Conformance suite for the tuning and attribution guards. -Quiet (the summary line and any failures only) Runbook: before you start the stack
Assert-DeployManifest.ps1 Records the deployment's attribution-relevant configuration and refuses a multi-variable step nobody acknowledged. -Label (a timestamp), -ComposeDirectory (this directory), -ManifestPath (.deploy/manifests/<label>.manifest here), -BaselinePath (the newest manifest already there), -AcceptMultipleDeltas with -Reason, -ContainerName (repocontextcontainer-repocontext-1), -EmbedContainerName (repocontextcontainer-embedder-1), -CgroupCpuQuota (read from the container), -Quiet, and the test seams -DeclaredReading, -EffectiveReading and -ComposeConfigFilesLabel Runbook: restart, drain, and verification
Test-DeployManifest.ps1 Shows the deploy manifest's declared-versus-effective check both firing and visibly not firing. -SkipAcquisition (skip the real compose resolution) Runbook: restart, drain, and verification
Assert-ContainerProvenance.ps1 The seven-check provenance guard over a running container. See its section. Verifying what you actually deployed
Test-ContainerProvenance.ps1, Test-ArchiveGitReading.ps1, Test-ProvenanceExitCode.ps1 The provenance guard's own suites. None. Verifying what you actually deployed
Invoke-AnnQueryProbe.ps1, Test-AnnQueryProbe.ps1 The retrieval-plane query probe and its suite. See its section. Measuring approximate retrieval on a running container
Invoke-AnnBuildProbe.ps1 The approximate-index build probe. See its section. Measuring an approximate-index build
Test-BackupSinkDurability.ps1 Shows the backup sink's host directory surviving docker compose down -v. -ProjectName (repocontext-backup-durability-probe), -SinkPath (a fresh directory under the system temp path), -Force (run even while another compose project's containers are up) Container quickstart
_deployManifest.ps1, _mcpClient.ps1, _provenance.ps1, _tuningKnobs.ps1 Shared functions the scripts above dot-source; not run on their own. None. -