RepoContext MCP container - "codebase memory in a box"
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.This sample runs the RepoContext MCP server as a single, restart-durable container alongside its embedding companion, and demonstrates the core durability guarantee end to end:
start -> add a repo under the mounted workspace -> recall -> restart -> context is still present.
The same box also serves the retrieval and token-economics tools for you to use and test: explainable search, the budgeted context bundle, its reuse economics, and usage accounting (see Retrieval and token economics).
Three containers, one private network:
repocontext- the MCP host image (apps/repocontext/Dockerfile). Its ONLY application listener is the MCP endpoint on port 8080 (plus the HTTP health probes and the Prometheus/metricsscrape endpoint). No gRPC facade and no Explorer UI are exposed. It runs the defaultlocaldurability profile: Orleans ADO.NET grain storage and reminders over a single SQLite file, plus the file-backed Lattice WAL - all under/data, which is a named volume, so state survivesdocker compose restart,docker compose down, and image upgrades. Zero external services.embedder- the ONNX Runtime model companion (apps/embedding-onnx/Dockerfile). It stays a SEPARATE container so the MCP host keeps its single-listener surface. The host's default embedding provider is pointed at it viaLATTICE_EMBEDDING_ENDPOINT. Its model weights are baked into an image layer, so it needs no cache volume and no download on first run.azurite-backup-sink- a blob-only Azurite instance (mcr.microsoft.com/azure-storage/azurite:latest) that receives the scheduled captures of the durable agent-memory tree, so the backup does not live in the store it protects. Its storage is a host bind mount at${REPOCONTEXT_BACKUP_PATH:-./backup-sink}rather than a named volume, sodocker compose down -vcannot reach it; setREPOCONTEXT_BACKUP_PATHto an absolute path to keep it outside the repository. Its blob endpoint is published on host port11000(override withREPOCONTEXT_BACKUP_SINK_PORT) so a restore can list it. See Agent-memory backup and recovery.
For a long-lived, tuned deployment rather than a first run - the CPU and
memory grants, the reduced scan cadence, the image pin, the rollback ladder, and
how to rebuild the whole thing from nothing - see the
local deployment runbook
and the tracked docker-compose.tuning.yml beside this file. The walkthrough below
uses the untuned defaults and does not need either.
Optional CPU pinning
Two variables, REPOCONTEXT_CPUSET and EMBEDDER_CPUSET, pin each service to a
fixed set of CPUs. Both are unset by default and should stay that way unless
you have a reason; unset, Compose omits the key from the resolved document
entirely, so ignoring them deploys exactly what this sample deployed before they
existed.
They matter only once a CPU quota is in force, which the tuning overlay supplies. A container entitled to 4 CPUs but visible on 16 has its threads scattered across 16 run queues, and CFS charges a whole 5 ms slice to each queue a thread wakes on, so the quota is exhausted by reservation rather than by work: measured 5.4% throttled periods at ~17% quota utilisation, against 0.0% at ~26% when pinned, varying nothing else. Sizing the cpuset to the grant removes the throttling.
Derive the values from your own grants rather than copying them - one range per
service, sized to the ceiling of that service's actual cpus grant in your
.env (an EMBEDDER_CPUS raised above what scripts/New-TuningEnv.ps1 derives
needs an EMBEDDER_CPUSET re-sized to match), the ranges must not overlap
or you trade throttling for contention, and together they must fit the host's
logical CPU count - grants whose ceilings sum past it cannot both be pinned, and
pinning is not the fix for that. See
.env.example for the form and the runbook section
"CPU scatter under a fractional quota" for the evidence, the limits of what it
claims, and why you must not enable it in the middle of a measurement.
Prerequisites
- Docker with Compose v2.
- An
.envfile in this directory: copy.env.exampleto.envand set its two paths for this machine before running any compose command below.REPOCONTEXT_MEMORY_ARCHIVE_PATHhas no default, so compose refuses every command -up,build,ps,logs- until it is set (see What actually protects it). - Memory: derive the grant with
scripts/New-TuningEnv.ps1(issue #2779) rather than copying a figure from here. This is much more than a default Docker VM allocation, so give the VM headroom above whatever the script derives. Two figures matter and they are not the same number. Steady state is fitted at 3 GiB plus 1 MB per indexed file, so a ~8,200-file index settles near 11 GiB. The reclamation peak is what a box costs to reach that steady state, and it is higher: on an 8,224-file index whose WAL garbage collection had been blocked, releasing the backlog ran at 450-590% CPU and drove the working set to at least 13.81 GiB before it fell back to 10.55 GiB, while a 12 GiB limit crash-looped twice in 16 minutes (issue #3252). An earlier revision of this list recommended "at least 12 GiB" against a 10.2 GiB steady state; both figures are withdrawn. The peak is a migration cost paid once, the first time a backlog of stuck WAL is released, so an operator upgrading into a WAL GC fix needs more headroom than one already running healthy. - Memory: 13.81 GiB is a floor on the peak, not the peak. It is the largest of 16 samples about 103 seconds apart, so the true maximum is at least that and may be higher - nothing observed the gaps. Quote it as a lower bound. The fall to 10.55 GiB afterwards is the other half of the reading, and the useful half: a 3.26 GiB give-back is what distinguishes a bounded transient from a leak, which would not have receded.
- Memory: do not assume the derived grant covers the migration burst. On the
corpus above
New-TuningEnv.ps1derives 13.26 GiB and the observed peak exceeded it by about 568 MiB, so the script's 20% headroom band is all that stands between a reclamation burst and the ceiling, and this burst was larger than the band. Grant above the derived figure for the migration run specifically, then re-derive once the estate is healthy. That 568 MiB is pinned to one corpus and is not a constant shortfall - the derived grant is a function of file count and moves with it, so the same repository derives 13.26 GiB at 8,239 files and 13.97 GiB at 8,846. Do not read a later derivation that happens to exceed 13.81 as evidence the gap has closed: a larger corpus raises the peak too, and only the grant side of that comparison was re-measured. Compare a peak against a grant only at the same corpus. - Memory, and this decides whether the comparison just made means anything, because it is weaker than it looks in one direction and stronger in another. First, a working set measured under a generous cap is an upper bound on need, not a requirement - .NET collects less eagerly the further it is from its ceiling, so 13.81 GiB observed at an 18 GiB grant does not establish that 13.81 GiB is needed. A peak cannot be transported across caps: at 18 GiB the collector worked against a 13.5 GiB managed ceiling, whereas at 13.26 GiB it would work against 9.94 GiB and collect far harder, far earlier. That trajectory was never run. Second, pulling the other way, peak RSS is not the quantity a grant is tested against at all: the grant bounds RSS, but the GC hard limit is what throws, and it binds first at about 75% of the grant (12 GiB grants 9 GiB, 18 GiB grants 13.5 GiB). A box at 13.26 GiB would fault against 9.94 GiB of managed heap long before RSS could reach 13.26, so "the peak exceeded the grant" understates the exposure rather than overstating it.
- Memory, what is actually established, stated on the plane that throws: a 12 GiB grant (9.00 GiB managed ceiling) crash-looped, and an 18 GiB grant (13.5 GiB managed ceiling) ran clean. Nothing has been measured at 13.26 GiB in either direction. The derived grant therefore offers a managed ceiling only about 10% above one that demonstrably crash-looped on this corpus. That thin margin over a measured failure is the reason to provision above it for a migration run - not the raw RSS comparison, which weighs a number produced under one cap against a different cap.
- Memory: everything above is deploy-time fitting, which is a known limitation
rather than the settled answer.
New-TuningEnv.ps1is host-specific by construction - its constants were fitted against one corpus on one host and are re-derived by nobody afterwards - and it goes stale in place, because its only corpus input is the indexed file count, so adding a repository to the workspace or removing one moves the requirement without moving the grant. Nothing re-derives the grant when that happens. Adapting sizing to the granted resources at runtime, instead of predicting it at deploy time, is tracked in issue #3255. What has landed from it so far is the measured consequence: the host publishes its peak commitment and its peak occupancy of the granted ceiling, records a managed-heap exhaustion on the data mount, and at the next start refuses a grant no larger than a recorded exhaustion ceiling and warns when the previous run's peak does not fit (see Measured requirement and startup admission). The runtime adaptation depends on issue #3133: the runtime's own high-load signal is published at 90% of the cgroup limit while the GC hard limit binds at 75% of it, so the threshold sits at 1.2x the limit at every grant and can never fire. It has been confirmed at both 12 GiB and 18 GiB with byte-exact matching percentages, so it is scale-invariant rather than a misconfiguration of one deployment. - Memory: under-provisioning does not present as memory pressure. A cgroup limit
becomes the .NET GC heap hard limit, so the process is never OOM-killed and
there is no restart, exit code or resource event. The visible symptom is a
STORAGE error while reading grain state, because the allocation that fails is a
leaf-snapshot deserialisation; the leaf then activates cold and replays its
whole WAL window, raising pressure further (issue #2364). The
orleans.lattice.leaf.snapshot.load_failurescounter names the real cause directly (reason=resource_exhausted), but it only fires once an allocation has already failed, so it reports an arrival rather than warning of an approach. Alert instead on the signals that engage before exhaustion:lattice_repocontext_heap_committed_bytes / lattice_repocontext_heap_limit_bytesfor heap-ceiling adherence (checklattice_repocontext_heap_high_load_threshold_reachablefirst - a0means the runtime's own pressure threshold can never fire at any grant, issue #3133),orleans.lattice.wal.replay.permit_adaptationswithoutcome=withheld, trigger=occupancyfor the proactive replay-concurrency backpressure, andorleans.lattice.leaf.snapshot.hydration_admissionswithoutcome=queued. All publish before they fire - the heap gauges from process start, the replay counter once the replay gate is first sized, and the hydration counter per tree at that tree's first leaf hydration, each arm primed at zero - so a flat zero is a measured zero, and an absent series means either that nothing has replayed or hydrated yet in this process or that the running image predates the instrument. - Build context differs per image: the host image's is the REPOSITORY ROOT (it
ProjectReferences the just-built
src/bits), so its service setscontext: ../..; the embedder builds from its ownapps/embedding-onnxdirectory, which has noProjectReferenceintosrc/and so keeps a small context. Its one shared source folder, the container cgroup readers insrc/lattice/Internal/Cgroups, arrives as the BuildKit named contextcgroupsthat the embedder service declares underadditional_contexts. Run compose from this directory either way.
The mounted workspace
Set REPO_PATH to the absolute path of a directory the box may see. It is mounted
READ-ONLY at /workspace inside the container, so the box can never mutate the
code it indexes. This is a workspace root, not a single repository: mount a broad
parent and register individual repositories under it at runtime with the
repocontext_add_repo tool. It defaults to ../../.. from this directory, which in
an ordinary clone is this repository's parent, so this repo is one registerable
child. From a git worktree the same default resolves to the worktree collection
directory instead, silently - see
the worktree trap -
so set REPO_PATH explicitly, in .env or the environment.
export REPO_PATH=/absolute/path/to/some/parent # PowerShell: $env:REPO_PATH="C:\path\to\parent"
The published port
The MCP endpoint and health probes are published on host port 8080 by default.
That port is a common one to already have in use, and a clash shows up only as an
opaque bind failure when you run up, so it is overridable with
REPOCONTEXT_PORT:
export REPOCONTEXT_PORT=18080 # PowerShell: $env:REPOCONTEXT_PORT="18080"
docker compose up -d
Only the host side moves; inside the container the listener stays on 8080. Every
localhost:8080 in the walkthrough below then becomes
localhost:$REPOCONTEXT_PORT.
Every path passed to repocontext_add_repo is resolved to its real location - ..
traversal and symlink escape are both defeated - and must resolve under
/workspace (set by LATTICE_WORKSPACE_ROOT); a path outside it is refused.
Choosing an embedding companion
The sample brings up the ONNX Runtime companion
(apps/embedding-onnx) by default. It
bakes its weights into an image layer, so a cold start needs no model download,
and a cuda-flavoured build selects its accelerator - CPU or NVIDIA - at
runtime via EMBED_PROVIDER (the default cpu build is CPU-only; see
Running the embedder on an NVIDIA GPU).
The original Onyx companion
(apps/embedding) remains available as a
fallback, selected with an override file. Build the embedder when you switch:
both companions build under the same compose image name, so a plain up -d
would reuse the cached ONNX image rather than build the Onyx one (switching back
needs the same rebuild):
docker compose -f docker-compose.yml -f docker-compose.onyx.yml build embedder
docker compose -f docker-compose.yml -f docker-compose.onyx.yml up -d
Nothing else changes in either direction. Both serve the same contract on the
same port, so the service name and LATTICE_EMBEDDING_ENDPOINT are identical,
and they produce numerically identical vectors (same pinned model revision,
fp32, same tokenizer and pooling), so switching does not invalidate an existing
/data volume. The ONNX image is roughly an order of magnitude smaller
(about 1.3 GB against 13 GB).
Those commands are for the untuned walkthrough stack. On the tuned deployment
keep -f docker-compose.tuning.yml in both commands, before the Onyx file, or
the switch silently drops the image pin, the grants and the tuned cadence; see
Rolling back the embedder
for the three-file form and the -ExpectedConfigFileCount 3 the provenance
check then needs.
Running the embedder on an NVIDIA GPU
The default build is CPU-only, and it stays CPU-only on a GPU host: the cpu
flavour has no GPU support compiled in. Enabling the GPU needs all three of these
together, because they do different jobs - the build arg picks a different ONNX
Runtime package, the environment variable binds the accelerator at runtime, and
the reservation is what actually exposes the device to the container. In
docker-compose.yml under the embedder service:
build:
args:
ONNX_FLAVOR: cuda
environment:
EMBED_PROVIDER: cuda
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Then rebuild, because a plain up -d would reuse the cached CPU image:
docker compose build embedder
docker compose up -d
Requires the NVIDIA Container Toolkit on the host. The cuda flavour is a much
larger image but still includes the CPU provider, so it serves CPU hosts too, and
an unrecognised EMBED_PROVIDER falls back to the CPU rather than failing to
boot. That fallback is silent by design, so verify what actually bound:
docker compose logs embedder | grep "Embedding server listening"
# ... listening on port 9000 using Cuda with model model.onnx (768-dim) ...
# Or ask the health endpoint. Both images are chiseled and carry no curl, and the
# embedder port is only exposed on the compose network, so borrow its namespace:
docker run --rm --network "container:$(docker compose ps -q embedder)" \
curlimages/curl -s http://localhost:9000/api/health
# {"status":"ok","provider":"Cuda","model":"model.onnx","dimension":768}
The provider field is the provider EMBED_PROVIDER resolved to, reported once
a session has been created on it. So a provider of Cpu here means
EMBED_PROVIDER did not resolve to cuda - it is unset, or a value other than
cuda, gpu or nvidia - and the fix is the service's environment. A missing
toolkit or device reservation cannot produce Cpu: with cuda selected the
server has no CPU fallback of its own, and a session it cannot create stops the
process at startup, so check docker compose logs embedder instead. The Onyx
fallback companion takes a different route to the same place (clear
CUDA_VISIBLE_DEVICES and add the reservation); see
apps/embedding.
Walkthrough
From this directory:
# 0. Copy the example environment file, then edit both paths in it for this
# machine. REPOCONTEXT_MEMORY_ARCHIVE_PATH is REQUIRED - compose refuses every
# command without it - and REPO_PATH picks the workspace root (see above).
cp .env.example .env
# 1. Start the stack. The host waits for the embedder to become healthy and for
# the backup sink to start.
docker compose up -d --build
# 2. Wait for the host to come up. /health/live returns 200 once the process and
# the silo host are alive, which is what the remaining steps actually need.
curl -fsS http://localhost:8080/health/live
# /health/ready is a stricter, orchestrator-facing probe, and this walkthrough
# deliberately does NOT gate on it. It is the conjunction of the lifecycle phase
# AND the vector plane having demonstrated a working semantic query, so it can
# stay 503 long after the box is up and answering MCP calls. Observe it, but do
# not wait on it or treat a 503 as a failed deployment - read "Interpreting a
# persistent 503" below first. The -o/-w form reports the code without failing
# the shell, which `curl -fsS` would do on any non-2xx.
curl -sS -o /dev/null -w 'ready: %{http_code}\n' http://localhost:8080/health/ready
# 3. Register a repository under the mounted workspace over MCP (repocontext_add_repo
# with a path under /workspace). Use your MCP client of choice against
# http://localhost:8080 (the MCP streamable-HTTP endpoint). For example, with
# the reference `mcp` CLI:
# mcp call http://localhost:8080 repocontext_add_repo '{"path":"/workspace/my-repo"}'
# Omit repoId to derive it from the final path segment, or set it explicitly:
# mcp call http://localhost:8080 repocontext_add_repo '{"path":"/workspace/my-repo","repoId":"demo"}'
# List what is registered at any time:
# mcp call http://localhost:8080 repocontext_list_repos '{}'
# 4. Recall: query the box (repocontext_search / repocontext_recall) and confirm it
# returns the ingested context.
# 5. Restart the container - a full process restart (not a recreation: the same
# container and its /data volume are kept) that evicts the in-memory projection
# and forces a WAL replay / cold rebuild on next access.
docker compose restart repocontext
curl -fsS http://localhost:8080/health/live
# As in step 2, /health/ready may stay 503 after the restart without meaning the
# restart failed. Step 6, not the probe, is the proof that the data survived.
# 6. Recall again. The context is still present: it was replayed from the WAL and
# SQLite state on the /data volume, proving durability across a restart.
Retrieval and token economics
Once a repository is registered (step 3 above), the same box exposes the epic's
retrieval and token-economics tools over the same MCP endpoint - no extra service,
no second listener. Every call below targets the demo repo id from step 3;
substitute your own. Examples use the reference mcp CLI against
http://localhost:8080.
# A. Explainable search. Every hit carries a machine-readable `reasons` array
# saying WHY it ranked (semantic proximity and matched chunk/symbol, or the
# specific keyword fields hit - path/name, symbol, tag, topic, content, key),
# so an agent can justify a selection instead of trusting an opaque score.
mcp call http://localhost:8080 repocontext_search \
'{"repoId":"demo","query":"where is the readiness health probe wired","k":5}'
# B. Budgeted context bundle. repocontext_context packs the ranked, explained
# source for a task into ONE response under a HARD token ceiling: the reported
# `responseTokens` (the response as delivered, envelope included) never exceeds
# the ceiling it reports as `budgetTokens` - the `responseBudgetTokens` you asked
# for, clamped to 1-200000 and 8192 when omitted - `totalTokens` is the narrower
# sum of packed source, and `truncated` /
# `retryBudgetTokens` say whether more would fit at a larger budget. `detail`
# trades richness for budget - 'paths' (cheapest) -> 'outline' (declared-symbol
# skeleton) -> 'slices' (bounded body text, richest), or 'auto' (default) which
# picks the richest level that fits. Pass a `session` id so the box remembers
# what it delivered.
mcp call http://localhost:8080 repocontext_context \
'{"repoId":"demo","task":"explain the readiness health check","responseBudgetTokens":4000,"detail":"auto","session":"agent-1"}'
# C. Reuse economics. Repeat on the SAME `session`. Units the session already
# holds are suppressed - acknowledged under `reused`, never re-charged and never
# counted against `top` or the budget - so the second answer pays only for the
# NEW context. (You can also feed the prior entries' unit receipts back via
# `seen`, or a whole-file 'path@hash' claim via `known`; the server-side
# `session` bookkeeping does it for you.)
mcp call http://localhost:8080 repocontext_context \
'{"repoId":"demo","task":"explain the readiness health check and how drain flips it","responseBudgetTokens":4000,"detail":"auto","session":"agent-1"}'
# D. Usage accounting. repocontext_stats reports the aggregate token economics over
# a bounded recent window: calls answered, response tokens spent, whole-file
# reads replaced, and the NET tokens saved by budgeting plus reuse.
mcp call http://localhost:8080 repocontext_stats '{}'
With the embedder companion healthy, search and the bundle rank semantically;
with it unreachable they degrade to a deterministic keyword rank - the bundle still
answers either way. See
docs/lattice.api.mcp.repocontext/retrieval-economics.md
for the full model.
Tear down (state on the named volumes is preserved unless you pass -v):
docker compose down # keeps the data volume (and the Onyx overlay's model cache, if used)
docker compose down -v # also deletes durable state (start clean)
down -v deletes the index and the authored agent memory, because both live
on the same /data volume. Read
Agent memory versus the code index before
using it: repocontext_reset_index rebuilds an index with no loss at all, and
the memory archive that makes down -v survivable is bounded by its export
interval rather than complete.
Agent memory versus the code index
Two kinds of state share the /data volume, and only one of them can be
recreated.
| What it is | If it is destroyed | The safe gesture | |
|---|---|---|---|
| Code index | Structural, content, symbol, xref, session and vector planes, derived from files on disk | Re-run repocontext_add_repo; back in minutes |
repocontext_reset_index |
| Agent memory | Every repocontext_remember note, decision, gotcha and glossary entry |
Gone; it is the store of record and derives from nothing | Keep an archive (below) |
docker compose down -v destroys both. That is the defect behind issue #2601:
the gesture is documented as the ordinary way to start clean, and it silently
takes the irreplaceable half with it. It has already happened once, to epic
#2368's own memory.
Why the two are not simply on separate volumes
Because they cannot be, and because it would not have helped.
They cannot be: a Lattice tree's durable state spans a WAL root that is one
directory for the whole storage provider, and a grain store that is a single
SQLite file shared by every tree. The /data/wal/repo-context-* subdirectories
look like separable locations but are a naming convention inside one root, and
the memory tree's pages sit interleaved with every other tree's in
/data/repocontext.db. There is no memory-only path to mount elsewhere.
It would not have helped: docker compose down -v removes every named volume
the project declares, not just the one you had in mind. A second declared volume
dies in the same command as the first.
What actually protects it
A bind mount, /memory-archive, which is not a project-declared volume and
so is not removed by down -v. The host exports memory there periodically and,
when it starts against an empty store, restores from it.
# The archive path is REQUIRED and has no default. A relative one would resolve
# against the directory you invoked compose from, which is how the only working
# backup of durable agent memory ended up inside an ephemeral git worktree
# (issue #2627). Set it to an absolute path outside every checkout and worktree.
REPOCONTEXT_MEMORY_ARCHIVE_PATH=/srv/repocontext-memory docker compose up -d
Copying .env.example to .env sets it for you; compose loads .env on every
command, so up, down, ps, and logs all pick it up. Without it, every
compose command in this directory fails by name rather than quietly choosing a
directory nobody picked.
Verify where it actually landed, rather than where you meant it to land:
docker inspect "$(docker compose ps -q repocontext)" \
--format '{{range .Mounts}}{{.Destination}} <- {{.Type}} {{.Source}}{{"\n"}}{{end}}'
pwsh -File scripts/Assert-ContainerProvenance.ps1 # check 5 of 7 refuses a doomed path
| Variable | Default | Meaning |
|---|---|---|
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR |
/memory-archive in this sample; unset (feature off) otherwise |
Container path the archive is written to. Unset disables the whole mechanism. |
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_INTERVAL_SECONDS |
300 |
Export cadence. This is the size of the window an ungraceful stop loses. A positive value below 30 is raised to 30, and one above the longest delay a timer can wait (about 49.7 days) is lowered to it; a zero, negative or unparseable value, or one too large to represent as a duration (for example 1e20), falls back to 300. |
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_RESTORE |
auto |
auto restores into an empty store, or into one whose restore-state marker records that an earlier restore was left partial; always restores on every start; off never restores; none and false are accepted for off and on-empty for auto, and any other value falls back to auto. |
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_STOP_TIMEOUT_SECONDS |
20 |
Budget for the final export during a graceful stop. A positive value is clamped to 1-60; a zero, negative or unparseable value falls back to 20. |
REPOCONTEXT_MEMORY_ARCHIVE_PATH |
none - required | HOST path bound at /memory-archive. Deliberately has no default: a relative one resolves against the compose invocation directory (issue #2627). Must be absolute and outside every checkout and worktree. |
Restore this way by hand at any time - stop the box, put the archive files in place, start it against an empty store:
docker compose down -v
ls "$REPOCONTEXT_MEMORY_ARCHIVE_PATH"/ # repo-context-memory.snapshot (+ .previous.snapshot)
docker compose up -d
docker compose logs repocontext | grep -i -E 'memory durability|durable memory'
The memory durability lines are the startup statement of where memory lives and
what protects it; they do not say whether anything was restored. The restore
outcome is logged separately, a few seconds after start, on a line naming
Durable memory - Durable memory was restored from the archive at ... when
records were merged back - and that is the line that confirms the import.
What this does not do
It is not a backup, and it does not make down -v safe.
- Everything written since the last export is lost. The exposure is the export
interval, plus whatever a non-graceful stop discards. A graceful stop
(
down,down -v,stop) exports once more on the way out and closes most of that window; akill -9or a host crash does not. - It archives memory only. The index is not in the archive, by design - it rebuilds from source.
- It does not remove the co-location. Memory still shares a volume with
rebuildable state, and the host says so at startup, at warning level, every
time.
repocontext_reset_indexremains the correct way to rebuild an index: it drops the derived planes and preserves memory outright, with no window at all. - It is not the scheduled backup that issue #2602 added (the
azurite-backup-sinkservice). That one captures the same memory tree to an external sink, with manifests and retention, and restores only the backup id an operator names; this one owns automatic restore-on-empty for memory alone. Only this mechanism restores automatically at startup.
See docs/lattice.api.mcp.repocontext/memory-durability.md for the full model.
Health probing
The runtime image is distroless and shell-less, so probing is HTTP-only - there is no shell-exec healthcheck:
GET /health/live- process + silo host alive (liveness). This is the probe the walkthrough gates on, and the one an orchestrator uses to decide whether to restart the container.GET /health/ready- readiness (routing). On this sample's local durability profile it is the conjunction of two independent components, and both must be healthy for a 200 (the Azure profile adds a third, the scaling-signal check):- lifecycle - the silo has joined, the activation-time WAL replay is done, the durable stores were proven reachable, and MCP is serving;
- vector plane - semantic retrieval has been demonstrated to work. A
deployment with no embedder bound (keyword-only), and a host with no repository
registered yet, both count as ready here: there is no vector plane to wait for
in the first case and nothing to serve in the second. This host always binds
its embedding provider, though, so a missing or unreachable
embedderdoes not make it keyword-only: searches reportkeyword.vector_plane_unavailableand, once a repository is registered, readiness stays down.
So it is not-ready during startup replay and during drain, but those are not the only causes, and a sustained 503 is far more likely to be the vector-plane component than either of them.
GET /health/silo- grain liveness. It re-checks silo membership and makes a trivial grain call on every request, so it goes red for a silo that died after reaching readiness while the process kept listening (issue #2666) - which neither/health/live(always green) nor/health/ready(which never re-checks the silo once it has flipped) can do. It answers200only when healthy and503while the silo is still starting or once it is unhealthy, with the three-way verdict in the body. It is what the container's Docker healthcheck targets, through the exec-form--healthcheckself-probe the shell-less image needs, and it feeds neither of the two probes above. Docker only records that verdict: under this sample'srestart: unless-stoppeda container is restarted when its process exits, never because it is unhealthy.GET /health/backup- whether the durable agent-memory tree is actually being captured to theazurite-backup-sinkservice. It is tagged as neither liveness nor readiness, so a failing backup never restarts the container or pulls it from rotation; the body names the tree in scope, what the sink holds and the last failure text, and it answers503only for an unhealthy verdict.
Interpreting a persistent 503
A /health/ready 503 that does not clear, on a container that is otherwise up,
does not on its own mean the deployment is broken, and must not be used by
itself as a rollback signal. The response body names each component and its verdict on its own line beneath the aggregate status, so the component holding readiness down is readable straight from the probe (issue #2962). That names the component, not the cause, so the steps below still apply: read the body first, then work through them to narrow why that component is unhappy. Before issue #2962 the endpoint returned a bare Unhealthy with no
per-component breakdown, so a 503 was ambiguous until it had been narrowed by hand. Five steps,
each one ruling out a cause the previous step left open:
curl -fsS http://localhost:8080/health/live. A 200 says the process and the silo host are alive, so whatever is unhealthy is not the process. If this also fails, the container really is unhealthy - that is the case to act on.Make any MCP call against
http://localhost:8080(repocontext_list_reposis the cheapest). If it answers, the MCP surface is serving, which satisfies the lifecycle component and leaves the vector plane as the one holding readiness down.Run a
repocontext_searchand read theretrievalPathon the result. A value ofkeyword.vector_plane_unavailableconfirms it: semantic retrieval is unavailable and the box has fallen back to deterministic keyword recall.Check the embedder with
docker compose ps. Step 3 tells you the vector plane is at fault but not which side of it, and the two sides need opposite responses. Anembeddercontainer that is missing, exited, or(unhealthy)is itself the cause, and is directly actionable: restore it and readiness can recover on its own. Anembedderreporting(healthy)while readiness stays 503 rules the embedder out and places the fault host-side, in the vector plane, where restarting the embedder achieves nothing. Usedocker compose psrather than probing the embedder directly - its port is not published to the host.Separate "never been ready" from "was ready and has since lost it" on
/metrics. The two need different responses and steps 1 to 4 cannot tell them apart:curl -fsS http://localhost:8080/metrics \ | grep -E 'repocontext_retrieval_(ready_seconds|unavailable)'repocontext_retrieval_ready_seconds_countis stamped once per process, on the first transition into a ready phase. Its absence therefore means the retrieval plane has never been ready in this container's current lifetime; its presence alongside a 503 means the plane was ready and has since lost it. Itsphaselabel records which phase it first reached (serving,keyword_only, ornothing_registered).repocontext_retrieval_unavailable_totalcounts fault episodes under a closedcauselabel: the three capability-loss values of step 3'sretrievalPath, so it separates a vector plane that cannot serve (keyword.vector_plane_unavailable) from an index that has drifted from its sources (keyword.index_degraded) and from a withheld exact fallback (keyword.exact_fallback_suppressed), plusprobe(a readiness probe rather than a real query saw the plane unable to serve),saturated(an admission gate refused the plane's open past its bound) andunknown.
Issuing a query yourself does not clear it, and the host is already trying. A
warmup service issues the same semantic query from application start, retrying with
backoff (waits of 2, 4, 8, 16 and 32 seconds, then every 30 seconds) until the plane
answers or shutdown begins, and once it has answered it re-checks readiness every
30 seconds and re-drives the query whenever readiness has been revoked. So a
persistent 503 is never "nobody has queried it yet" - it is that warmup failing
repeatedly. In particular, a box that has a repository registered but holds no
vectors for it stays not-ready by design: the search reports
keyword.vector_plane_unavailable, and running another search by hand returns the
same thing and changes nothing. (A box with no repository registered is the
opposite case and reports ready, because there is nothing it could be asked to
serve.)
Readiness also lags a fault on purpose. Once the plane has served, a fault must persist for 30 seconds before readiness is revoked, and any successful retrieval inside that window clears the episode outright - so a 503 can appear up to half a minute after the fault that caused it, and a brief blip may never surface at all.
In that state the box is still usable and the whole walkthrough still completes:
registration, keyword search, repocontext_context, and durability across a restart
all work, and steps 3 to 6 demonstrate exactly that. What is degraded is semantic
ranking, not the service. Treat it as a capability to restore, not as a deployment
to roll back.
The same listener also serves GET /metrics, a Prometheus text exposition of every
instrument on a meter whose name starts with orleans.lattice - the core meter and
every per-package meter, Orleans.Lattice.Api.Mcp.RepoContext included - plus the
Microsoft.Orleans and System.Runtime runtime meters. It needs no second port and
no sidecar:
curl -fsS http://localhost:8080/metrics | head -n 20
Notes on durability and shutdown
- All durable local state (the WAL directory and the SQLite database file) lives
under
/data, a named volume. The host fails fast at startup if that path cannot be created or is not writable by its non-root UID; a missing directory is created rather than refused. - On
SIGTERM(adocker stop/restart) the host flips readiness to not-ready first, then drains: the silo deactivates and the WAL commit-log flushes buffered records before exit, so an in-flight write is durable after restart. - PID 1 is an init process, and that is what makes the
SIGTERMland at all.init: trueindocker-compose.ymlhas Docker bind-mount its own staticdocker-initbinary and run it as PID 1, with the host as its child. It needs nothing in the distroless runtime image and changes no application code. Two kernel behaviours make it necessary, and both attach to PID 1 rather than to the application: no default action is taken for a signal delivered to PID 1 that PID 1 has installed no handler for, so a well-behaved process can be unkillable bySIGTERMpurely by being PID 1; and PID 1 inherits every orphaned descendant and mustwait()on it, which the .NET host does not do. In the epic #2368 gate runs a container reached a state where neitherdocker killnordocker rm -fwould reap PID 1 and it had to beSIGKILLed, costing that run its drain and leaving the next one unbanked state to replay (issue #2576). This is independent of the grace period below:initdecides whether the drain starts, the grace period decides how long it may take, and setting one without the other leaves half the failure in place. - That drain's budget is 180 seconds in this sample, and it belongs to the
host, not to Docker. The host sets
HostOptions.ShutdownTimeoutfrom the grace period the deployment declares - 180s from this sample's declared240s, or the 90s default (RepoContextHostBuilder.ShutdownBudget) when nothing is declared; Docker'sstop_grace_perioddefaults to 10 seconds. The two are enforced independently and the smaller wins, so without an explicitstop_grace_periodthe process isSIGKILLed at 10s with the drain still running and the budget is dead configuration (issue #2389). Thestop_grace_period: 240sindocker-compose.ymlis what makes it reachable. If you run this image under your own orchestration you must grant the same budget there - Kubernetes has the identical trap under a different name, sinceterminationGracePeriodSecondsdefaults to 30s. - The drain reports its own duration, so the budget can be derived rather than
bisected.
docker logscarries `RepoContext drain complete ins, consuming % of the 180s host shutdown budget`. Read it together with its severity, because there are three distinct outcomes and the level is what separates them: - **No completion line at all.** The container was killed mid-drain, so `stop_grace_period` is smaller than the drain (issue #2389). The exit code will not tell you, because a killed container reports `137` and the next `docker start` overwrites it. - **`drain complete` at `Warning`.** The drain finished but consumed 70% or more of the budget. Nothing has failed; treat it as a lead indicator, because drain time grows with resident state. - **`drain ABANDONED after 180s` at `Error`, and the container exits `70`.** The *host* stopped waiting and deactivation was abandoned part-way. Raise the service's `stop_grace_period` and the `LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD` that declares it, together and to the same value; the budget is derived from the second and bounded by the first, so raising either alone achieves nothing.
- The abandoned drain is also visible without reading the log (issue #2401).
An orchestrator does not read logs, it reads the exit code, and before #2401
the abandoned case did not reliably produce a distinctive one. Measured against
a real generic host, the outcome was not simply zero but undetermined: a
hosted service that absorbed the shutdown cancellation left
RunAsyncreturning normally, so the process exited0and the abandonment was recorded as a clean stop; one that rethrew it let the exception escapeRunAsyncunhandled, aborting the process in a way indistinguishable from a real crash. The host now assigns the code itself when the overrun latches:0if the drain completed inside the budget,70if it was abandoned.70isEX_SOFTWAREin thesysexits.hconvention, chosen to avoid0/1/2, Docker's reserved125-127, and the128 + signalband that holds137(SIGKILL) and143(SIGTERM) - the neighbouring conditions it exists to be told apart from.- Note what this does not do.
docker-compose.ymlusesrestart: unless-stopped, under which Docker restarts on any exit code, so the code neither triggers nor suppresses a restart. What it changes is what is recorded:docker inspect --format '{{.State.ExitCode}}'reports70,docker ps -ashowsExited (70), and under Kubernetes the container terminates with reasonErrorrather thanCompleted. That is what an alert can be written against. - There is deliberately no way to turn it off. A switch restoring
0would remove the evidence rather than the problem.
- Note what this does not do.
- That last line exists because of issue #2397, and the reason it is needed is
not obvious. The host raises
ApplicationStoppedeven when the shutdown budget expired and it gave up waiting - so a signal bound only to that event reporteddrain complete in 90.0sfor a drain that did not complete. The failure looked like success. The overrun is now raised by an alarm armed when the drain begins, so it is reported at the moment the budget expires rather than depending on a callback that may never arrive. - Drain time scales with resident state: the same 400-file rig drained in 33.9s before its vector trees had landed and 67.2s once they had, which was already three quarters of the 90s budget the host then ran with. #2397 nevertheless did not raise it, because measurements on a live box show the resident set that a drain must flush has no observed ceiling (idle-deactivation sweeps ranging from 1 to 4,418 activations, still climbing between readings). A fixed ceiling on an unbounded quantity moves the threshold without changing the failure mode, so #2397 shipped the diagnostic instead of a new number, and #2402 - which proposed raising it - did not ship one either. Issue #3304 later did raise it, to 180s, against two consecutive drains that no longer fitted (89.7s and 91.9s against 90s); see the container quickstart.
- The budget is no longer written down as an independent constant. Since issue #2402
the host derives it from the grace period the deployment declares through
LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD, taking 75% of it or all but a two-second unwind reserve, whichever is smaller. The120sthe sample declared when that derivation landed yielded exactly the 90s the container had always run with, so nothing moved at the time (the sample now declares240s, deriving 180s, since issue #3304); what changed is that there is one number to set instead of two, and a budget larger than the grace period can no longer be expressed. That matters because such a budget is not merely useless - the container kills the process at the real grace period regardless, so theABANDONEDline above is armed for an instant that never arrives and is never emitted, which is the silent teardown of issue #2389 all over again. - The variable declares the grant; it is not the grant. Nothing inside the
container can read the real
stop_grace_period, so a deployment that declares120swhile granting20sruns a 90s budget under a 20s guillotine and cannot detect it. Keeping the two adjacent in the same compose service is the mitigation, andRepoContextComposeShutdownBudgetTestsasserts they are equal here - but that adjacency is a convention, not an enforcement. Change them together, always. - None of this bounds the resident set. Raising both values past your own observed drain is a legitimate local remedy, but it buys time rather than fixing the shape.
Verifying what you actually deployed
Everything above describes what the tracked compose file declares. Nothing above establishes that a container now running received any of it.
docker compose up reads the compose files in its own working directory,
whatever branch built the image it starts, and its output names no branch, no
commit, and no directory. The image and the runtime configuration are therefore
two independent inputs, and only the first is obviously version-controlled. An
operator standing in one checkout can deploy a candidate image under a different
checkout's configuration and see nothing at all to say so.
That is not hypothetical. Two gate runs of epic #2368 did exactly this: the
candidate image ran under the baseline's runtime config, the container's own
compose label resolved to the main checkout, and
LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD was absent from the process environment
while sitting merged on the candidate branch. Both runs observed the absence of
the fix's effect and concluded the fix was absent. The observation was real and
correctly made. The discriminator was in a channel nobody was reading.
Note what would not have helped. A check comparing tracked files to other tracked files would have been green throughout both runs, because the repository agreed with itself perfectly. Only a reading taken from the running process separates "the source does not carry the fix" from "the source carries it and this container never received it".
cd samples/RepoContextContainer
pwsh -File ./scripts/Assert-ContainerProvenance.ps1
It exits 0 only when every check agrees, and with a distinct non-zero code on
each refusal path (2 refused, 3 the container could not be interrogated, 4
the expected configuration could not be read), so automation can gate on it. Until
#2718 it never called exit at all and leaked 128 from a git probe that is
SUPPOSED to fail; the full contract is in the
local deployment runbook.
It performs seven checks against the running container and refuses unless all seven agree, printing every value it read either way:
- Compose provenance. The container's own
com.docker.compose.project.working_dirlabel resolves to the checkout you are standing in, every file incom.docker.compose.project.config_fileslies under it and exists on disk, and the number of those files is the number you expected (default 2, see below). An override file merged in from elsewhere is how a resolved document stops matching the tracked one. - Git provenance. That directory is a git worktree and its HEAD is the commit you expect, reported as a value you read rather than an inference you make. The expectation is sourced from your checkout, never from the one the container resolved to; defaulting it to the latter would compare a value against itself and pass unconditionally, which is worse than omitting the check because it would report as checked.
- Image provenance. The image id the container is executing still matches what its image reference resolves to now, which catches a container left in place across a rebuild. Ids, never tags: a tag is a mutable pointer, so comparing tag to tag compares two names for whatever is current.
- Environment provenance. A candidate-only setting is present in the
container's own environment with the expected value, expectation read from
your checkout's
docker-compose.ymland actual read from the running process. This is the non-redundant one. Checks 1 to 3 can all pass while an override file, an edit, or a stale container leaves the value unset, and it is the direct executable form of the warning above that the variable declares the grant rather than being it. Durations are compared parsed, never literally: Compose normalises120sto2m0s, so a text comparison would accuse a correctly configured stack of exactly this defect, and the obvious remedy for that accusation is to change a deployment that was already right. - Archive durability. The bind mount holding durable agent memory resolves to an absolute host path outside every git worktree and checkout. It is the only check on the WRITE path, and it keys on the archive path alone, never on the compose directory (issue #2627).
- Build provenance. The image the container is executing was built from the
expected commit, read from the image's own
org.opencontainers.image.revisionlabel or failing that acandidate-<sha>tag, and cross-checked against chronology. It fails closed. Checks 1 and 2 adjudicate a checkout; this one adjudicates the image (issue #2686). - Workspace provenance. The
/workspacebind the container indexes out of exists, is a bind rather than a volume, and resolves to an absolute host path - and, when you name them, it is the expected workspace root and a registered repository is indexed from the expected repository root (issue #2617).
The parameters not discussed below steer what those checks read.
-ContainerName names the container to interrogate; by default it is the one
carrying the compose service label for -ComposeServiceName (default
repocontext). -ExpectedCheckout (default: the checkout holding the script's
compose files) is the directory check 1 compares against and the checkout whose
docker-compose.yml check 4 reads its expectation from, and -ExpectedCommit
(default: that checkout's HEAD) is the commit checks 2 and 6 compare against;
-ExpectedCommitDate (default: that commit's date), -CandidateTagPrefix
(default candidate-) and -ClockSkewToleranceSeconds (default 120) feed
check 6's tag fallback and chronology arm; and -ArchiveDestination (default
/memory-archive) and -WorkspaceDestination (default /workspace) are the
container paths checks 5 and 7 adjudicate.
The count assertion, and why it is not a walk of tracked files
The stack's real deployment is two compose files. The tracked
docker-compose.yml carries a build: stanza and no image:, so on its own it
cannot resolve an image for up -d --no-build at all. A second file supplies the
image pin, the memory limit, the CPU caps, and the scan-cadence variables every
prior measurement on a given box was taken against.
That second file used to be an untracked, gitignored
docker-compose.override.yml, which put the entire tuned configuration on
exactly one machine with no review, no history, and no diff (issue #2609). It is
now the tracked docker-compose.tuning.yml, layered by name:
docker compose -f docker-compose.yml -f docker-compose.tuning.yml up -d --no-build
The rationale behind every value in it, the build-and-tag ladder that produces the image pin, and the rollback procedure are in the local deployment runbook.
The count is still two, so this check needs no new argument. Naming the files
explicitly also suppresses an automatic docker-compose.override.yml, which
is what makes the deployed configuration reproducible from the checkout alone;
layering a personal override on top of the tuning file makes it three, and you
say so with -ExpectedConfigFileCount 3.
The opt-in CPU pinning described above does not move the count: it is two variables on the existing services, not a third file.
A machine-local override remains legitimate, and it is why check 1 reads the container's label rather than walking the repository: it has to be able to fail on a file git has never heard of. The obvious remedy for a compose-provenance failure - relaunch from the checkout you meant - silently drops whichever second file the original launch had, leaving no image pin, no memory limit and a different scan cadence, while the tree looks perfectly correct and every path the container reports still resolves under the right directory. Only the count dissents.
The general form is worth stating, because it is not specific to compose: fixing a provenance defect by changing the launch directory is itself a provenance change, and it is not self-verifying.
Pass -ExpectedConfigFileCount 1 if you genuinely mean to run without an
override. Making that an explicit act is the point - dropping the override
should be something you said, not something that happened.
What a green run does not establish. That the checkout was clean when up
ran, since HEAD is a commit and uncommitted compose edits are invisible here - and
a GIT_COMMIT stamped from a dirty tree names a commit the image does not
exactly contain. That any setting you did not name reached the process: check 4
proves only the settings passed as -ExpectedSetting. That the memory archive's
content is good, or that the indexed workspace is current: checks 5 and 7
adjudicate where the archive landed and which tree is bound, not what state
either is in. That the workspace root or the indexed repository is the one you
meant, unless you named them with -ExpectedWorkspaceRoot and
-ExpectedRepositoryRoot (the report prints NOT ESTABLISHED for each arm you
did not ask for; -ExpectedRepositoryRoot also needs -IndexedRoot, the
indexedRoot values repocontext_list_repos reports, and refuses without it
rather than passing), or that every registered repository is correctly rooted -
the indexed-root arm passes when one matches. Or anything whatever about a container
you did not name. What a green run does establish, since check 6 (issue
#2686), is that the image the container is executing was built from the expected
commit, on the evidence of its own revision label or candidate-<sha> tag
cross-checked against chronology. Check 3 as defaulted detects only a stale
container; pass -ExpectedImageId if you need to pin an exact image.
This is an operator check and is deliberately not wired into CI. It needs a running container, and a fixture that skipped when Docker was absent would produce exactly the false green it exists to prevent.
It is also not the same instrument as the cold-start rig's
Assert-RigComposeIsolation, and neither subsumes the other. That guard
validates the declaration - what docker compose config resolved - before
anything runs. This validates the deployment - what a container already running
actually received. The rig has never had this failure mode, because
Get-RigComposeFile pins the file it resolves, so the different-checkout drift
cannot arise there. A green rig guard therefore says nothing about this class,
and the two must not be collapsed. The boundary was drawn deliberately by issue
#2576, whose handoff named both the remedy and its location: a post-up
precondition on the container's own environment.
The adjudication is pure and separated from the acquisition, so the refusing direction is exercised against fabricated disagreements rather than assumed:
pwsh -File ./scripts/Test-ContainerProvenance.ps1
Every one of the seven checks has fixtures it accepts and fixtures it refuses, and several refusal fixtures are reconstructed from real gate-run and incident readings rather than invented. A check only ever observed passing is indistinguishable from one that cannot fail, which is the same reason the suite itself is worth measuring rather than trusting: commit first, then make one check return no violations unconditionally, re-run, and confirm the assertions that fail are the ones covering that check and no others.
Two sibling suites cover what that pure suite cannot see, since it neither runs
git nor runs the script: Test-ArchiveGitReading.ps1 checks the archive check's
pinned git messages against the git actually installed, and
Test-ProvenanceExitCode.ps1 checks that the script's exit status says what its
printed verdict says.
Measuring approximate retrieval on a running container
scripts/Invoke-AnnQueryProbe.ps1 issues real retrieval queries against a
running container and reports what moved on
repocontext_retrieval_ann_search_total. It exists because a scrape on its own
cannot answer the question it appears to answer: that counter is written only on
the per-query path, so a deploy that never issued a query leaves every arm at
its primed zero, and that reading is byte-identical to a plane that was
consulted and answered exactly nothing approximately.
pwsh -File ./scripts/Invoke-AnnQueryProbe.ps1
pwsh -File ./scripts/Invoke-AnnQueryProbe.ps1 -RepoId lattice -Repetitions 4
It scrapes the three state arms before and after, issues its queries over the
MCP endpoint on 8080 (the container's only application listener), and reports
the per-arm delta.
With no -RepoId it probes every repository repocontext_list_repos reports.
Its other parameters are -BaseUri (default http://localhost:8080; pass the
published port if you moved it with REPOCONTEXT_PORT), -Queries (a built-in
spread of natural-language queries when omitted, each issued against every
selected repository), -Repetitions (default 1), -K hits per query
(default 5), -TimeoutSeconds per request (default 60), and
-JsonOutputPath, which writes the full machine-readable result document.
It refuses to report rather than reporting a zero it cannot stand behind. The refusal is the feature, and there are two of them, kept deliberately distinct because they have different owners:
- issued 0 (exit 2) - the probe never got a query out. Nothing can be concluded about the instrument; the fault is the probe's or the environment's.
- issued N, succeeded 0 (exit 2) - queries went out and every one failed. The instrument reading is still inadmissible, but the fault is now the container's and is worth diagnosing.
Collapsing those two into one "no data" would discard exactly the bit that says
whose problem it is. Neither is the environment failure that stops the probe
before it can query at all - a container that does not answer /health/live, a
failed MCP handshake, or a failed repocontext_list_repos - which exits 3.
retrievalPath on a search result is not evidence about the approximate arm.
The approximate index's declared retrieval path is a property of the index, not of a
query - one index serves every repository, so a state-tracking declaration would
be wrong the moment two repositories were in different states. It therefore
reads semantic.approximate unconditionally, including when the exact fallback
answered with complete recall, and NormalizeSemantic fails closed the same
way by resolving anything unrecognised to it. The declaration deliberately
under-promises. Reading it as confirmation that approximate search ran is
confidently wrong, and the probe prints that caveat rather than assuming you
know it.
An absent arm is not a zero. The probe distinguishes "the series is present
and reads 0" from "the series is not on the endpoint at all". The first is a
measurement. The second means the series was refused at creation, and the probe
sends you to lattice_metrics_series and
lattice_metrics_dropped_measurements_by_family_total before you conclude
anything from it.
It reads /health/ready, and prints the answer verbatim. This is a
different endpoint from /health/live and answers a different question.
Liveness asks whether anything is there; readiness asks whether this box can
actually serve semantic retrieval, and on the repocontext host it is the
conjunction of the lifecycle component and the vector plane. When the vector
plane is down the readiness body already says so, in specific and self-limiting
terms, and it even names the retrievalPath discrimination you would otherwise
have to rediscover.
The endpoint is easy to miss, and has been missed: the container healthcheck
runs a grain-liveness self-probe rather than an HTTP readiness call, so Docker
can report healthy straight through a total retrieval outage, and the
acceptance playbook's only outbound call is /metrics. The probe therefore
reads it explicitly rather than assuming something upstream already did.
A 503 here is a successful probe result, not a probe failure. It is the system diagnosing itself, which is more authoritative than anything this harness can infer from a counter delta, so the probe prints the status and the full body and says as much. It is deliberately not a gate: a 503 is the expected reading on a rig whose vector plane is down, and refusing to continue would suppress the very measurement the harness exists to take.
The total across all three arms is the liveness witness. The arms partition
the whole query population - the index records an outcome for every
query including bootstrapping - so a moving total proves the instrument is
capable of reporting, independently of which arm moved. A zero on
approximate beside a non-zero total is a measured absence. A zero beside a
zero total is not a reading at all. Because the arms carry no caller tag, a
total delta larger than the probe's own succeeded count is reported as
CONTAMINATED rather than claimed: concurrent internal retrieval is
indistinguishable from the probe's own, and pretending otherwise would attribute
traffic the probe did not generate.
The probe does not deploy, does not score any acceptance predicate, and does not tune retrieval parameters. Tuning until the approximate arm fires would encode the answer into the instrument.
Its own refusal paths are regression-tested rather than proven once:
pwsh -File ./scripts/Test-AnnQueryProbe.ps1
Twelve scenarios against a real HTTP listener that the suite runs as a background
job (on localhost, port 18080 unless you pass -Port), invoking the probe as
a separate pwsh process, covering both refusals,
the absent-arm case, the contaminated delta, the suppressed-fallback state, and
all three readiness shapes (ready, not-ready-with-a-diagnosis, and a readiness
endpoint that cannot be read at all).
The suite asserts its own scenario count is non-zero before reporting, for the
same reason the probe asserts its issued count: a harness that ran nothing
reports success in a way that is indistinguishable from a harness that ran
everything and found nothing wrong.
Measuring an approximate-index build
scripts/Invoke-AnnBuildProbe.ps1 measures how long the approximate index takes
to converge for one repository on a running container - the wall-clock figure an
A/B of the build path is scored on. It registers the repository itself with
repocontext_add_repo (a write, and for a repository that is already registered a
fresh indexing pass), then polls repocontext_health for that repository every
-PollSeconds (default 10) and reports the time to converge, the vectors indexed,
and the sample series (written as JSON with -JsonOutputPath).
pwsh -File ./scripts/Invoke-AnnBuildProbe.ps1 -RepoPath /workspace/my-repo
-RepoPath is the in-container path under the mounted workspace, -RepoId
defaults to its final segment, as repocontext_add_repo itself derives it,
-BaseUri defaults to http://localhost:8080, and -TimeoutSeconds (default
240) bounds each MCP request it makes.
Convergence is not the approximate plane reporting Ready on its own. A build
over a corpus that has not been embedded yet reaches Ready at once with nothing
in it, so the probe waits for Ready with the indexed vector count caught up to a
non-zero embedded coverage. It exits 0 when that happens and 2 when
-MaxWaitMinutes (default 60) elapses first - an arm that does not converge is a
result to record, not a harness fault - and it fails outright if the registration
still fails after its retries.
It deliberately does not read /health/ready, which answers a different question
(see above), and it scores on time to converge rather than on the
repocontext.ann.build.stage.duration histogram, because only the former is
reported by every build an A/B might compare. Read the stage split afterwards to
explain a difference, not to score one.
Scripts
Every script in scripts/, and every parameter it accepts. All
parameters are optional except -RepoPath on Invoke-AnnBuildProbe.ps1.
| Script | What it does | Parameters (default) | Described in |
|---|---|---|---|
New-TuningEnv.ps1 |
Derives this host's resource knobs for docker-compose.tuning.yml and writes them to .env. |
-WorkspacePath (the repository root), -OutFile (.env beside the compose files), -DryRun, -CorpusOnly (measure and print the corpus, then exit), -IgnoreHostLoad, -ExpectedCorpusFiles, -CorpusTolerance (0.02), -Force, and two seams for driving its refusals deterministically, -HostMemoryBytes and -HostAvailableMemoryBytes |
Runbook: before you start the stack |
Assert-TuningEnv.ps1 |
Refuses a tuning .env whose knobs are unset or carry a retired sentinel. |
-EnvFile (.env beside the compose files), -Reading (a hashtable checked instead of a file), -BuildCommit, -RepositoryPath (the repository holding the script), -Quiet |
Runbook: before you start the stack |
Test-TuningIntegrity.ps1 |
Conformance suite for the tuning and attribution guards. | -Quiet (the summary line and any failures only) |
Runbook: before you start the stack |
Assert-DeployManifest.ps1 |
Records the deployment's attribution-relevant configuration and refuses a multi-variable step nobody acknowledged. | -Label (a timestamp), -ComposeDirectory (this directory), -ManifestPath (.deploy/manifests/<label>.manifest here), -BaselinePath (the newest manifest already there), -AcceptMultipleDeltas with -Reason, -ContainerName (repocontextcontainer-repocontext-1), -EmbedContainerName (repocontextcontainer-embedder-1), -CgroupCpuQuota (read from the container), -Quiet, and the test seams -DeclaredReading, -EffectiveReading and -ComposeConfigFilesLabel |
Runbook: restart, drain, and verification |
Test-DeployManifest.ps1 |
Shows the deploy manifest's declared-versus-effective check both firing and visibly not firing. | -SkipAcquisition (skip the real compose resolution) |
Runbook: restart, drain, and verification |
Assert-ContainerProvenance.ps1 |
The seven-check provenance guard over a running container. | See its section. | Verifying what you actually deployed |
Test-ContainerProvenance.ps1, Test-ArchiveGitReading.ps1, Test-ProvenanceExitCode.ps1 |
The provenance guard's own suites. | None. | Verifying what you actually deployed |
Invoke-AnnQueryProbe.ps1, Test-AnnQueryProbe.ps1 |
The retrieval-plane query probe and its suite. | See its section. | Measuring approximate retrieval on a running container |
Invoke-AnnBuildProbe.ps1 |
The approximate-index build probe. | See its section. | Measuring an approximate-index build |
Test-BackupSinkDurability.ps1 |
Shows the backup sink's host directory surviving docker compose down -v. |
-ProjectName (repocontext-backup-durability-probe), -SinkPath (a fresh directory under the system temp path), -Force (run even while another compose project's containers are up) |
Container quickstart |
_deployManifest.ps1, _mcpClient.ps1, _provenance.ps1, _tuningKnobs.ps1 |
Shared functions the scripts above dot-source; not run on their own. | None. | - |