---
title: "RepoContext MCP container - \"codebase memory in a box\""
url: "https://nsta1.github.io/Orleans.Lattice/samples/RepoContextContainer/README.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/samples/RepoContextContainer/README.md"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
---
# RepoContext MCP container - "codebase memory in a box"

Part of [Samples](../index.md).

This sample runs the RepoContext MCP server as a single,
restart-durable container alongside its embedding companion, and demonstrates the
core durability guarantee end to end:

**start -> add a repo under the mounted workspace -> recall -> restart -> context
is still present.**

The same box also serves the retrieval and token-economics tools for you to use and
test: explainable search, the budgeted context bundle, its reuse economics, and
usage accounting (see [Retrieval and token economics](#retrieval-and-token-economics)).

Three containers, one private network:

- **`repocontext`** - the MCP host image (`apps/repocontext/Dockerfile`). Its
  ONLY application listener is the MCP endpoint on port 8080 (plus the HTTP health
  probes and the Prometheus `/metrics` scrape endpoint). No gRPC facade and no
  Explorer UI are exposed. It runs the default
  `local` durability profile: Orleans ADO.NET grain storage and reminders over a
  single SQLite file, plus the file-backed Lattice WAL - all under `/data`, which
  is a named volume, so state survives `docker compose restart`, `docker compose
  down`, and image upgrades. Zero external services.
- **`embedder`** - the ONNX Runtime model companion
  (`apps/embedding-onnx/Dockerfile`). It stays a SEPARATE container so the MCP
  host keeps its single-listener surface. The host's default embedding provider
  is pointed at it via `LATTICE_EMBEDDING_ENDPOINT`. Its model weights are baked
  into an image layer, so it needs no cache volume and no download on first run.
- **`azurite-backup-sink`** - a blob-only Azurite instance
  (`mcr.microsoft.com/azure-storage/azurite:latest`) that receives the scheduled
  captures of the durable agent-memory tree, so the backup does not live in the
  store it protects. Its storage is a host bind mount at
  `${REPOCONTEXT_BACKUP_PATH:-./backup-sink}` rather than a named volume, so
  `docker compose down -v` cannot reach it; set `REPOCONTEXT_BACKUP_PATH` to an
  absolute path to keep it outside the repository. Its blob endpoint is published
  on host port `11000` (override with `REPOCONTEXT_BACKUP_SINK_PORT`) so a restore
  can list it. See
  [Agent-memory backup and recovery](../../docs/lattice.api.mcp.repocontext/container/agent-memory-backup-and-recovery.md).

For a **long-lived, tuned** deployment rather than a first run - the CPU and
memory grants, the reduced scan cadence, the image pin, the rollback ladder, and
how to rebuild the whole thing from nothing - see the
[local deployment runbook](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md)
and the tracked `docker-compose.tuning.yml` beside this file. The walkthrough below
uses the untuned defaults and does not need either.

### Optional CPU pinning

Two variables, `REPOCONTEXT_CPUSET` and `EMBEDDER_CPUSET`, pin each service to a
fixed set of CPUs. Both are **unset by default and should stay that way unless
you have a reason**; unset, Compose omits the key from the resolved document
entirely, so ignoring them deploys exactly what this sample deployed before they
existed.

They matter only once a **CPU quota** is in force, which the tuning overlay
supplies. A container entitled to 4 CPUs but visible on 16 has its threads
scattered across 16 run queues, and CFS charges a whole 5 ms slice to each queue
a thread wakes on, so the quota is exhausted by *reservation* rather than by
work: measured 5.4% throttled periods at ~17% quota utilisation, against **0.0%
at ~26%** when pinned, varying nothing else. Sizing the cpuset to the grant
removes the throttling.

Derive the values from your own grants rather than copying them - one range per
service, sized to the ceiling of that service's *actual* `cpus` grant in your
`.env` (an `EMBEDDER_CPUS` raised above what `scripts/New-TuningEnv.ps1` derives
needs an `EMBEDDER_CPUSET` re-sized to match), the ranges **must not overlap**
or you trade throttling for contention, and together they must fit the host's
logical CPU count - grants whose ceilings sum past it cannot both be pinned, and
pinning is not the fix for that. See
[.env.example](https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/samples/RepoContextContainer/.env.example) for the form and the runbook section
"CPU scatter under a fractional quota" for the evidence, the limits of what it
claims, and why you must not enable it in the middle of a measurement.

## Prerequisites

- Docker with Compose v2.
- An `.env` file in this directory: copy `.env.example` to `.env` and set its
  two paths for this machine before running any compose command below.
  `REPOCONTEXT_MEMORY_ARCHIVE_PATH` has no default, so compose refuses every
  command - `up`, `build`, `ps`, `logs` - until it is set (see
  [What actually protects it](#what-actually-protects-it)).
- Memory: derive the grant with
  [`scripts/New-TuningEnv.ps1`](https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/samples/RepoContextContainer/scripts/New-TuningEnv.ps1) (issue #2779) rather
  than copying a figure from here. This is much more than a default Docker VM
  allocation, so give the VM headroom above whatever the script derives. **Two
  figures matter and they are not the same number.** *Steady state* is fitted at
  3 GiB plus 1 MB per indexed file, so a ~8,200-file index settles near 11 GiB.
  *The reclamation peak* is what a box costs to **reach** that steady state, and
  it is higher: on an 8,224-file index whose WAL garbage collection had been
  blocked, releasing the backlog ran at 450-590% CPU and drove the working set to
  **at least 13.81 GiB** before it fell back to 10.55 GiB, while a 12 GiB limit
  crash-looped twice in 16 minutes (issue #3252). An earlier revision of this
  list recommended "at least 12 GiB" against a 10.2 GiB steady state; **both
  figures are withdrawn.** The peak is a *migration* cost paid once, the first
  time a backlog of stuck WAL is released, so an operator upgrading into a WAL GC
  fix needs more headroom than one already running healthy.
- Memory: **13.81 GiB is a floor on the peak, not the peak.** It is the largest
  of 16 samples about 103 seconds apart, so the true maximum is at least that and
  may be higher - nothing observed the gaps. Quote it as a lower bound. The fall
  to 10.55 GiB afterwards is the other half of the reading, and the useful half:
  a 3.26 GiB give-back is what distinguishes a bounded transient from a leak,
  which would not have receded.
- Memory: **do not assume the derived grant covers the migration burst.** On the
  corpus above `New-TuningEnv.ps1` derives 13.26 GiB and the observed peak
  exceeded it by about 568 MiB, so the script's 20% headroom band is all that
  stands between a reclamation burst and the ceiling, and this burst was larger
  than the band. Grant **above** the derived figure for the migration run
  specifically, then re-derive once the estate is healthy. That 568 MiB is
  **pinned to one corpus and is not a constant shortfall** - the derived grant is
  a function of file count and moves with it, so the same repository derives
  13.26 GiB at 8,239 files and 13.97 GiB at 8,846. Do not read a later derivation
  that happens to exceed 13.81 as evidence the gap has closed: a larger corpus
  raises the peak too, and only the grant side of that comparison was
  re-measured. Compare a peak against a grant only at the *same* corpus.
- Memory, and this decides whether the comparison just made means anything,
  because it is weaker than it looks in one direction and stronger in another.
  First, a working set measured under a generous cap is an **upper bound on need,
  not a requirement** - .NET collects less eagerly the further it is from its
  ceiling, so 13.81 GiB observed at an 18 GiB grant does not establish that
  13.81 GiB is *needed*. A peak cannot be transported across caps: at 18 GiB the
  collector worked against a 13.5 GiB managed ceiling, whereas at 13.26 GiB it
  would work against 9.94 GiB and collect far harder, far earlier. That
  trajectory was never run. Second, pulling the **other** way, peak RSS is not
  the quantity a grant is tested against at all: **the grant bounds RSS, but the
  GC hard limit is what throws**, and it binds first at about 75% of the grant
  (12 GiB grants 9 GiB, 18 GiB grants 13.5 GiB). A box at 13.26 GiB would fault
  against 9.94 GiB of managed heap long before RSS could reach 13.26, so "the
  peak exceeded the grant" *understates* the exposure rather than overstating it.
- Memory, what is actually established, stated on the plane that throws: a 12 GiB
  grant (9.00 GiB managed ceiling) **crash-looped**, and an 18 GiB grant
  (13.5 GiB managed ceiling) ran **clean**. Nothing has been measured at 13.26 GiB
  in either direction. The derived grant therefore offers a managed ceiling only
  about **10% above one that demonstrably crash-looped** on this corpus. That
  thin margin over a measured failure is the reason to provision above it for a
  migration run - not the raw RSS comparison, which weighs a number produced
  under one cap against a different cap.
- Memory: **everything above is deploy-time fitting, which is a known limitation
  rather than the settled answer.** `New-TuningEnv.ps1` is host-specific by
  construction - its constants were fitted against one corpus on one host and are
  re-derived by nobody afterwards - and it goes stale in place, because its only
  corpus input is the indexed file count, so adding a repository to the workspace
  or removing one moves the requirement without moving the grant. Nothing
  re-derives the grant when that happens. Adapting sizing to the granted
  resources **at runtime**, instead of predicting it at deploy time, is tracked in
  issue #3255. What has landed from it so far is the measured consequence: the host
  publishes its peak commitment and its peak occupancy of the granted ceiling,
  records a managed-heap exhaustion on the data mount, and at the next start
  refuses a grant no larger than a recorded exhaustion ceiling and warns when the
  previous run's peak does not fit (see
  [Measured requirement and startup admission](../../docs/lattice.api.mcp.repocontext/container/graceful-shutdown.md#measured-requirement-and-startup-admission)).
  The runtime adaptation depends on issue #3133: the runtime's own high-load signal
  is published at 90% of the cgroup limit while the GC hard limit binds at 75% of it, so the threshold sits
  at 1.2x the limit at *every* grant and can never fire. It has been confirmed at
  both 12 GiB and 18 GiB with byte-exact matching percentages, so it is
  scale-invariant rather than a misconfiguration of one deployment.
- Memory: under-provisioning does not present as memory pressure. A cgroup limit
  becomes the .NET GC heap hard limit, so the process is never OOM-killed and
  there is no restart, exit code or resource event. The visible symptom is a
  STORAGE error while reading grain state, because the allocation that fails is a
  leaf-snapshot deserialisation; the leaf then activates cold and replays its
  whole WAL window, raising pressure further (issue #2364). The
  `orleans.lattice.leaf.snapshot.load_failures` counter names the real cause
  directly (`reason=resource_exhausted`), but it only fires once an allocation
  has **already** failed, so it reports an arrival rather than warning of an
  approach. Alert instead on the signals that engage *before* exhaustion:
  `lattice_repocontext_heap_committed_bytes /
  lattice_repocontext_heap_limit_bytes` for heap-ceiling adherence (check
  `lattice_repocontext_heap_high_load_threshold_reachable` first - a `0` means
  the runtime's own pressure threshold can never fire at any grant, issue
  #3133),
  `orleans.lattice.wal.replay.permit_adaptations` with
  `outcome=withheld, trigger=occupancy` for the proactive replay-concurrency
  backpressure, and `orleans.lattice.leaf.snapshot.hydration_admissions` with
  `outcome=queued`. All publish before they fire - the heap gauges from process
  start, the replay counter once the replay gate is first sized, and the
  hydration counter per tree at that tree's first leaf hydration, each arm primed
  at zero - so a flat zero is a measured zero, and an absent series means either
  that nothing has replayed or hydrated yet in this process or that the running
  image predates the instrument.
- Build context differs per image: the host image's is the REPOSITORY ROOT (it
  ProjectReferences the just-built `src/` bits), so its service sets
  `context: ../..`; the embedder builds from its own `apps/embedding-onnx`
  directory, which has no `ProjectReference` into `src/` and so keeps a small
  context. Its one shared source folder, the container cgroup readers in
  `src/lattice/Internal/Cgroups`, arrives as the BuildKit named context `cgroups`
  that the embedder service declares under `additional_contexts`. Run compose
  from this directory either way.

## The mounted workspace

Set `REPO_PATH` to the absolute path of a directory the box may see. It is mounted
READ-ONLY at `/workspace` inside the container, so the box can never mutate the
code it indexes. This is a *workspace root*, not a single repository: mount a broad
parent and register individual repositories under it at runtime with the
`repocontext_add_repo` tool. It defaults to `../../..` from this directory, which in
an ordinary clone is this repository's parent, so this repo is one registerable
child. From a git worktree the same default resolves to the worktree collection
directory instead, silently - see
[the worktree trap](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#the-worktree-trap) -
so set `REPO_PATH` explicitly, in `.env` or the environment.

```bash
export REPO_PATH=/absolute/path/to/some/parent    # PowerShell: $env:REPO_PATH="C:\path\to\parent"
```

## The published port

The MCP endpoint and health probes are published on host port **8080** by default.
That port is a common one to already have in use, and a clash shows up only as an
opaque bind failure when you run `up`, so it is overridable with
`REPOCONTEXT_PORT`:

```bash
export REPOCONTEXT_PORT=18080                     # PowerShell: $env:REPOCONTEXT_PORT="18080"
docker compose up -d
```

Only the host side moves; inside the container the listener stays on 8080. Every
`localhost:8080` in the walkthrough below then becomes
`localhost:$REPOCONTEXT_PORT`.

Every path passed to `repocontext_add_repo` is resolved to its real location - `..`
traversal and symlink escape are both defeated - and must resolve under
`/workspace` (set by `LATTICE_WORKSPACE_ROOT`); a path outside it is refused.

## Choosing an embedding companion

The sample brings up the ONNX Runtime companion
([`apps/embedding-onnx`](https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/apps/embedding-onnx/README.md)) by default. It
bakes its weights into an image layer, so a cold start needs no model download,
and a `cuda`-flavoured build selects its accelerator - CPU or NVIDIA - at
runtime via `EMBED_PROVIDER` (the default `cpu` build is CPU-only; see
[Running the embedder on an NVIDIA GPU](#running-the-embedder-on-an-nvidia-gpu)).

The original Onyx companion
([`apps/embedding`](https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/apps/embedding/README.md)) remains available as a
fallback, selected with an override file. Build the embedder when you switch:
both companions build under the same compose image name, so a plain `up -d`
would reuse the cached ONNX image rather than build the Onyx one (switching back
needs the same rebuild):

```bash
docker compose -f docker-compose.yml -f docker-compose.onyx.yml build embedder
docker compose -f docker-compose.yml -f docker-compose.onyx.yml up -d
```

Nothing else changes in either direction. Both serve the same contract on the
same port, so the service name and `LATTICE_EMBEDDING_ENDPOINT` are identical,
and they produce **numerically identical vectors** (same pinned model revision,
fp32, same tokenizer and pooling), so switching does not invalidate an existing
`/data` volume. The ONNX image is roughly an order of magnitude smaller
(about 1.3 GB against 13 GB).

Those commands are for the untuned walkthrough stack. On the tuned deployment
keep `-f docker-compose.tuning.yml` in both commands, before the Onyx file, or
the switch silently drops the image pin, the grants and the tuned cadence; see
[Rolling back the embedder](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#rolling-back-the-embedder)
for the three-file form and the `-ExpectedConfigFileCount 3` the provenance
check then needs.

### Running the embedder on an NVIDIA GPU

The default build is CPU-only, and it stays CPU-only on a GPU host: the `cpu`
flavour has no GPU support compiled in. Enabling the GPU needs all three of these
together, because they do different jobs - the build arg picks a different ONNX
Runtime package, the environment variable binds the accelerator at runtime, and
the reservation is what actually exposes the device to the container. In
`docker-compose.yml` under the `embedder` service:

```yaml
    build:
      args:
        ONNX_FLAVOR: cuda
    environment:
      EMBED_PROVIDER: cuda
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
```

Then rebuild, because a plain `up -d` would reuse the cached CPU image:

```bash
docker compose build embedder
docker compose up -d
```

Requires the NVIDIA Container Toolkit on the host. The `cuda` flavour is a much
larger image but still includes the CPU provider, so it serves CPU hosts too, and
an unrecognised `EMBED_PROVIDER` falls back to the CPU rather than failing to
boot. That fallback is silent by design, so verify what actually bound:

```bash
docker compose logs embedder | grep "Embedding server listening"
# ... listening on port 9000 using Cuda with model model.onnx (768-dim) ...

# Or ask the health endpoint. Both images are chiseled and carry no curl, and the
# embedder port is only exposed on the compose network, so borrow its namespace:
docker run --rm --network "container:$(docker compose ps -q embedder)" \
  curlimages/curl -s http://localhost:9000/api/health
# {"status":"ok","provider":"Cuda","model":"model.onnx","dimension":768}
```

The `provider` field is the provider `EMBED_PROVIDER` resolved to, reported once
a session has been created on it. So a `provider` of `Cpu` here means
`EMBED_PROVIDER` did not resolve to `cuda` - it is unset, or a value other than
`cuda`, `gpu` or `nvidia` - and the fix is the service's `environment`. A missing
toolkit or device reservation cannot produce `Cpu`: with `cuda` selected the
server has no CPU fallback of its own, and a session it cannot create stops the
process at startup, so check `docker compose logs embedder` instead. The Onyx
fallback companion takes a different route to the same place (clear
`CUDA_VISIBLE_DEVICES` and add the reservation); see
[`apps/embedding`](https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/apps/embedding/README.md).

## Walkthrough

From this directory:

```bash
# 0. Copy the example environment file, then edit both paths in it for this
#    machine. REPOCONTEXT_MEMORY_ARCHIVE_PATH is REQUIRED - compose refuses every
#    command without it - and REPO_PATH picks the workspace root (see above).
cp .env.example .env

# 1. Start the stack. The host waits for the embedder to become healthy and for
#    the backup sink to start.
docker compose up -d --build

# 2. Wait for the host to come up. /health/live returns 200 once the process and
#    the silo host are alive, which is what the remaining steps actually need.
curl -fsS http://localhost:8080/health/live

#    /health/ready is a stricter, orchestrator-facing probe, and this walkthrough
#    deliberately does NOT gate on it. It is the conjunction of the lifecycle phase
#    AND the vector plane having demonstrated a working semantic query, so it can
#    stay 503 long after the box is up and answering MCP calls. Observe it, but do
#    not wait on it or treat a 503 as a failed deployment - read "Interpreting a
#    persistent 503" below first. The -o/-w form reports the code without failing
#    the shell, which `curl -fsS` would do on any non-2xx.
curl -sS -o /dev/null -w 'ready: %{http_code}\n' http://localhost:8080/health/ready

# 3. Register a repository under the mounted workspace over MCP (repocontext_add_repo
#    with a path under /workspace). Use your MCP client of choice against
#    http://localhost:8080 (the MCP streamable-HTTP endpoint). For example, with
#    the reference `mcp` CLI:
#      mcp call http://localhost:8080 repocontext_add_repo '{"path":"/workspace/my-repo"}'
#    Omit repoId to derive it from the final path segment, or set it explicitly:
#      mcp call http://localhost:8080 repocontext_add_repo '{"path":"/workspace/my-repo","repoId":"demo"}'
#    List what is registered at any time:
#      mcp call http://localhost:8080 repocontext_list_repos '{}'

# 4. Recall: query the box (repocontext_search / repocontext_recall) and confirm it
#    returns the ingested context.

# 5. Restart the container - a full process restart (not a recreation: the same
#    container and its /data volume are kept) that evicts the in-memory projection
#    and forces a WAL replay / cold rebuild on next access.
docker compose restart repocontext
curl -fsS http://localhost:8080/health/live
#    As in step 2, /health/ready may stay 503 after the restart without meaning the
#    restart failed. Step 6, not the probe, is the proof that the data survived.

# 6. Recall again. The context is still present: it was replayed from the WAL and
#    SQLite state on the /data volume, proving durability across a restart.
```

## Retrieval and token economics

Once a repository is registered (step 3 above), the same box exposes the epic's
retrieval and token-economics tools over the same MCP endpoint - no extra service,
no second listener. Every call below targets the `demo` repo id from step 3;
substitute your own. Examples use the reference `mcp` CLI against
`http://localhost:8080`.

```bash
# A. Explainable search. Every hit carries a machine-readable `reasons` array
#    saying WHY it ranked (semantic proximity and matched chunk/symbol, or the
#    specific keyword fields hit - path/name, symbol, tag, topic, content, key),
#    so an agent can justify a selection instead of trusting an opaque score.
mcp call http://localhost:8080 repocontext_search \
  '{"repoId":"demo","query":"where is the readiness health probe wired","k":5}'

# B. Budgeted context bundle. repocontext_context packs the ranked, explained
#    source for a task into ONE response under a HARD token ceiling: the reported
#    `responseTokens` (the response as delivered, envelope included) never exceeds
#    the ceiling it reports as `budgetTokens` - the `responseBudgetTokens` you asked
#    for, clamped to 1-200000 and 8192 when omitted - `totalTokens` is the narrower
#    sum of packed source, and `truncated` /
#    `retryBudgetTokens` say whether more would fit at a larger budget. `detail`
#    trades richness for budget - 'paths' (cheapest) -> 'outline' (declared-symbol
#    skeleton) -> 'slices' (bounded body text, richest), or 'auto' (default) which
#    picks the richest level that fits. Pass a `session` id so the box remembers
#    what it delivered.
mcp call http://localhost:8080 repocontext_context \
  '{"repoId":"demo","task":"explain the readiness health check","responseBudgetTokens":4000,"detail":"auto","session":"agent-1"}'

# C. Reuse economics. Repeat on the SAME `session`. Units the session already
#    holds are suppressed - acknowledged under `reused`, never re-charged and never
#    counted against `top` or the budget - so the second answer pays only for the
#    NEW context. (You can also feed the prior entries' unit receipts back via
#    `seen`, or a whole-file 'path@hash' claim via `known`; the server-side
#    `session` bookkeeping does it for you.)
mcp call http://localhost:8080 repocontext_context \
  '{"repoId":"demo","task":"explain the readiness health check and how drain flips it","responseBudgetTokens":4000,"detail":"auto","session":"agent-1"}'

# D. Usage accounting. repocontext_stats reports the aggregate token economics over
#    a bounded recent window: calls answered, response tokens spent, whole-file
#    reads replaced, and the NET tokens saved by budgeting plus reuse.
mcp call http://localhost:8080 repocontext_stats '{}'
```

With the `embedder` companion healthy, search and the bundle rank semantically;
with it unreachable they degrade to a deterministic keyword rank - the bundle still
answers either way. See
[docs/lattice.api.mcp.repocontext/retrieval-economics.md](../../docs/lattice.api.mcp.repocontext/retrieval-economics.md)
for the full model.

Tear down (state on the named volumes is preserved unless you pass `-v`):

```bash
docker compose down          # keeps the data volume (and the Onyx overlay's model cache, if used)
docker compose down -v       # also deletes durable state (start clean)
```

`down -v` deletes the index **and the authored agent memory**, because both live
on the same `/data` volume. Read
[Agent memory versus the code index](#agent-memory-versus-the-code-index) before
using it: `repocontext_reset_index` rebuilds an index with no loss at all, and
the memory archive that makes `down -v` survivable is bounded by its export
interval rather than complete.

## Agent memory versus the code index

Two kinds of state share the `/data` volume, and only one of them can be
recreated.

| | What it is | If it is destroyed | The safe gesture |
|---|---|---|---|
| **Code index** | Structural, content, symbol, xref, session and vector planes, derived from files on disk | Re-run `repocontext_add_repo`; back in minutes | `repocontext_reset_index` |
| **Agent memory** | Every `repocontext_remember` note, decision, gotcha and glossary entry | Gone; it is the store of record and derives from nothing | Keep an archive (below) |

`docker compose down -v` destroys both. That is the defect behind issue #2601:
the gesture is documented as the ordinary way to start clean, and it silently
takes the irreplaceable half with it. It has already happened once, to epic
#2368's own memory.

### Why the two are not simply on separate volumes

Because they cannot be, and because it would not have helped.

They cannot be: a Lattice tree's durable state spans a WAL root that is one
directory for the whole storage **provider**, and a grain store that is a single
SQLite file shared by every tree. The `/data/wal/repo-context-*` subdirectories
look like separable locations but are a naming convention inside one root, and
the memory tree's pages sit interleaved with every other tree's in
`/data/repocontext.db`. There is no memory-only path to mount elsewhere.

It would not have helped: `docker compose down -v` removes **every** named volume
the project declares, not just the one you had in mind. A second declared volume
dies in the same command as the first.

### What actually protects it

A **bind mount**, `/memory-archive`, which is not a project-declared volume and
so is not removed by `down -v`. The host exports memory there periodically and,
when it starts against an empty store, restores from it.

```bash
# The archive path is REQUIRED and has no default. A relative one would resolve
# against the directory you invoked compose from, which is how the only working
# backup of durable agent memory ended up inside an ephemeral git worktree
# (issue #2627). Set it to an absolute path outside every checkout and worktree.
REPOCONTEXT_MEMORY_ARCHIVE_PATH=/srv/repocontext-memory docker compose up -d
```

Copying `.env.example` to `.env` sets it for you; compose loads `.env` on every
command, so `up`, `down`, `ps`, and `logs` all pick it up. Without it, every
compose command in this directory fails by name rather than quietly choosing a
directory nobody picked.

Verify where it actually landed, rather than where you meant it to land:

```bash
docker inspect "$(docker compose ps -q repocontext)" \
  --format '{{range .Mounts}}{{.Destination}} <- {{.Type}} {{.Source}}{{"\n"}}{{end}}'
pwsh -File scripts/Assert-ContainerProvenance.ps1   # check 5 of 7 refuses a doomed path
```

| Variable | Default | Meaning |
|---|---|---|
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR` | `/memory-archive` in this sample; unset (feature off) otherwise | Container path the archive is written to. Unset disables the whole mechanism. |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_INTERVAL_SECONDS` | `300` | Export cadence. This is the size of the window an ungraceful stop loses. A positive value below 30 is raised to 30, and one above the longest delay a timer can wait (about 49.7 days) is lowered to it; a zero, negative or unparseable value, or one too large to represent as a duration (for example `1e20`), falls back to 300. |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_RESTORE` | `auto` | `auto` restores into an empty store, or into one whose restore-state marker records that an earlier restore was left partial; `always` restores on every start; `off` never restores; `none` and `false` are accepted for `off` and `on-empty` for `auto`, and any other value falls back to `auto`. |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_STOP_TIMEOUT_SECONDS` | `20` | Budget for the final export during a graceful stop. A positive value is clamped to 1-60; a zero, negative or unparseable value falls back to 20. |
| `REPOCONTEXT_MEMORY_ARCHIVE_PATH` | none - **required** | HOST path bound at `/memory-archive`. Deliberately has no default: a relative one resolves against the compose invocation directory (issue #2627). Must be absolute and outside every checkout and worktree. |

Restore this way by hand at any time - stop the box, put the archive files in
place, start it against an empty store:

```bash
docker compose down -v
ls "$REPOCONTEXT_MEMORY_ARCHIVE_PATH"/   # repo-context-memory.snapshot (+ .previous.snapshot)
docker compose up -d
docker compose logs repocontext | grep -i -E 'memory durability|durable memory'
```

The `memory durability` lines are the startup statement of where memory lives and
what protects it; they do not say whether anything was restored. The restore
outcome is logged separately, a few seconds after start, on a line naming
`Durable memory` - `Durable memory was restored from the archive at ...` when
records were merged back - and that is the line that confirms the import.

### What this does not do

It is **not a backup**, and it does not make `down -v` safe.

* Everything written since the last export is lost. The exposure is the export
  interval, plus whatever a non-graceful stop discards. A graceful stop
  (`down`, `down -v`, `stop`) exports once more on the way out and closes most
  of that window; a `kill -9` or a host crash does not.
* It archives **memory only**. The index is not in the archive, by design - it
  rebuilds from source.
* It does not remove the co-location. Memory still shares a volume with
  rebuildable state, and the host says so at startup, at warning level, every
  time. `repocontext_reset_index` remains the correct way to rebuild an index:
  it drops the derived planes and preserves memory outright, with no window at
  all.
* It is not the scheduled backup that issue #2602 added (the
  `azurite-backup-sink` service). That one captures the same memory tree to an
  external sink, with manifests and retention, and restores only the backup id
  an operator names; this one owns automatic restore-on-empty for memory alone.
  Only this mechanism restores automatically at startup.

See
[docs/lattice.api.mcp.repocontext/memory-durability.md](../../docs/lattice.api.mcp.repocontext/memory-durability.md)
for the full model.

## Health probing

The runtime image is distroless and shell-less, so probing is HTTP-only - there is
no shell-exec healthcheck:

- `GET /health/live` - process + silo host alive (liveness). This is the probe the
  walkthrough gates on, and the one an orchestrator uses to decide whether to
  restart the container.
- `GET /health/ready` - readiness (routing). On this sample's local durability
  profile it is the **conjunction of two independent components**, and both must be
  healthy for a 200 (the Azure profile adds a third, the scaling-signal check):
  - **lifecycle** - the silo has joined, the activation-time WAL replay is done, the
    durable stores were proven reachable, and MCP is serving;
  - **vector plane** - semantic retrieval has been demonstrated to work. A
    deployment with no embedder bound (keyword-only), and a host with no repository
    registered yet, both count as ready here: there is no vector plane to wait for
    in the first case and nothing to serve in the second. This host always binds
    its embedding provider, though, so a missing or unreachable `embedder` does
    not make it keyword-only: searches report `keyword.vector_plane_unavailable`
    and, once a repository is registered, readiness stays down.

  So it is not-ready during startup replay and during drain, but those are not the
  only causes, and a sustained 503 is far more likely to be the vector-plane
  component than either of them.
- `GET /health/silo` - grain liveness. It re-checks silo membership and makes a
  trivial grain call on every request, so it goes red for a silo that died after
  reaching readiness while the process kept listening (issue #2666) - which
  neither `/health/live` (always green) nor `/health/ready` (which never
  re-checks the silo once it has flipped) can do. It answers `200` only when
  healthy and `503` while the silo is still starting or once it is unhealthy,
  with the three-way verdict in the body. It is what the container's Docker
  healthcheck targets, through the exec-form `--healthcheck` self-probe the
  shell-less image needs, and it feeds neither of the two probes above. Docker
  only records that verdict: under this sample's `restart: unless-stopped` a
  container is restarted when its process exits, never because it is unhealthy.
- `GET /health/backup` - whether the durable agent-memory tree is actually being
  captured to the `azurite-backup-sink` service. It is tagged as neither liveness
  nor readiness, so a failing backup never restarts the container or pulls it
  from rotation; the body names the tree in scope, what the sink holds and the
  last failure text, and it answers `503` only for an unhealthy verdict.

### Interpreting a persistent 503

A `/health/ready` 503 that does not clear, on a container that is otherwise up,
does **not** on its own mean the deployment is broken, and must not be used by
itself as a rollback signal. The response body names each component and its verdict on its own line beneath the aggregate status, so the component holding readiness down is readable straight from the probe (issue #2962). That names the component, not the cause, so the steps below still apply: read the body first, then work through them to narrow why that component is unhappy. Before issue #2962 the endpoint returned a bare `Unhealthy` with no
per-component breakdown, so a 503 was ambiguous until it had been narrowed by hand. Five steps,
each one ruling out a cause the previous step left open:

1. `curl -fsS http://localhost:8080/health/live`. A 200 says the process and the
   silo host are alive, so whatever is unhealthy is not the process. If this also
   fails, the container really is unhealthy - that is the case to act on.
2. Make any MCP call against `http://localhost:8080` (`repocontext_list_repos` is
   the cheapest). If it answers, the MCP surface is serving, which satisfies the
   lifecycle component and leaves the vector plane as the one holding readiness
   down.
3. Run a `repocontext_search` and read the `retrievalPath` on the result. A value
   of `keyword.vector_plane_unavailable` confirms it: semantic retrieval is
   unavailable and the box has fallen back to deterministic keyword recall.
4. Check the embedder with `docker compose ps`. Step 3 tells you the vector plane
   is at fault but not which side of it, and the two sides need opposite responses.
   An `embedder` container that is missing, exited, or `(unhealthy)` is itself the
   cause, and is directly actionable: restore it and readiness can recover on its
   own. An `embedder` reporting `(healthy)` while readiness stays 503 rules the
   embedder out and places the fault host-side, in the vector plane, where
   restarting the embedder achieves nothing. Use `docker compose ps` rather than
   probing the embedder directly - its port is not published to the host.
5. Separate **"never been ready"** from **"was ready and has since lost it"** on
   `/metrics`. The two need different responses and steps 1 to 4 cannot tell them
   apart:

   ```bash
   curl -fsS http://localhost:8080/metrics \
     | grep -E 'repocontext_retrieval_(ready_seconds|unavailable)'
   ```

   `repocontext_retrieval_ready_seconds_count` is stamped **once per process**, on
   the first transition into a ready phase. Its absence therefore means the
   retrieval plane has never been ready in this container's current lifetime; its
   presence alongside a 503 means the plane was ready and has since lost it. Its
   `phase` label records which phase it first reached (`serving`, `keyword_only`,
   or `nothing_registered`). `repocontext_retrieval_unavailable_total` counts fault
   episodes under a closed `cause` label: the three capability-loss values of step
   3's `retrievalPath`, so it separates a vector plane that cannot serve
   (`keyword.vector_plane_unavailable`) from an index that has drifted from its
   sources (`keyword.index_degraded`) and from a withheld exact fallback
   (`keyword.exact_fallback_suppressed`), plus `probe` (a readiness probe rather
   than a real query saw the plane unable to serve), `saturated` (an admission gate
   refused the plane's open past its bound) and `unknown`.

**Issuing a query yourself does not clear it, and the host is already trying.** A
warmup service issues the same semantic query from application start, retrying with
backoff (waits of 2, 4, 8, 16 and 32 seconds, then every 30 seconds) until the plane
answers or shutdown begins, and once it has answered it re-checks readiness every
30 seconds and re-drives the query whenever readiness has been revoked. So a
persistent 503 is never "nobody has queried it yet" - it is that warmup failing
repeatedly. In particular, a box that has a repository **registered** but holds no
vectors for it stays not-ready by design: the search reports
`keyword.vector_plane_unavailable`, and running another search by hand returns the
same thing and changes nothing. (A box with **no** repository registered is the
opposite case and reports ready, because there is nothing it could be asked to
serve.)

Readiness also lags a fault on purpose. Once the plane has served, a fault must
persist for **30 seconds** before readiness is revoked, and any successful retrieval
inside that window clears the episode outright - so a 503 can appear up to half a
minute after the fault that caused it, and a brief blip may never surface at all.

In that state **the box is still usable and the whole walkthrough still completes**:
registration, keyword search, `repocontext_context`, and durability across a restart
all work, and steps 3 to 6 demonstrate exactly that. What is degraded is semantic
ranking, not the service. Treat it as a capability to restore, not as a deployment
to roll back.

The same listener also serves `GET /metrics`, a Prometheus text exposition of every
instrument on a meter whose name starts with `orleans.lattice` - the core meter and
every per-package meter, `Orleans.Lattice.Api.Mcp.RepoContext` included - plus the
`Microsoft.Orleans` and `System.Runtime` runtime meters. It needs no second port and
no sidecar:

```bash
curl -fsS http://localhost:8080/metrics | head -n 20
```

## Notes on durability and shutdown

- All durable local state (the WAL directory and the SQLite database file) lives
  under `/data`, a named volume. The host fails fast at startup if that path
  cannot be created or is not writable by its non-root UID; a missing directory is
  created rather than refused.
- On `SIGTERM` (a `docker stop` / `restart`) the host flips readiness to not-ready
  first, then drains: the silo deactivates and the WAL commit-log flushes buffered
  records before exit, so an in-flight write is durable after restart.
- **PID 1 is an init process, and that is what makes the `SIGTERM` land at all.**
  `init: true` in `docker-compose.yml` has Docker bind-mount its own static
  `docker-init` binary and run it as PID 1, with the host as its child. It needs
  nothing in the distroless runtime image and changes no application code. Two
  kernel behaviours make it necessary, and both attach to PID 1 rather than to
  the application: no default action is taken for a signal delivered to PID 1
  that PID 1 has installed no handler for, so a well-behaved process can be
  unkillable by `SIGTERM` purely by being PID 1; and PID 1 inherits every
  orphaned descendant and must `wait()` on it, which the .NET host does not do.
  In the epic #2368 gate runs a container reached a state where neither
  `docker kill` nor `docker rm -f` would reap PID 1 and it had to be
  `SIGKILL`ed, costing that run its drain and leaving the next one unbanked
  state to replay (issue #2576). This is **independent of the grace period
  below**: `init` decides whether the drain starts, the grace period decides how
  long it may take, and setting one without the other leaves half the failure in
  place.
- **That drain's budget is 180 seconds in this sample, and it belongs to the
  host, not to Docker.** The host sets `HostOptions.ShutdownTimeout` from the
  grace period the deployment declares - 180s from this sample's declared `240s`,
  or the 90s default (`RepoContextHostBuilder.ShutdownBudget`) when nothing is
  declared; Docker's `stop_grace_period` defaults to **10 seconds**. The two are
  enforced independently and the smaller wins, so without an explicit
  `stop_grace_period` the process is `SIGKILL`ed at 10s with the drain still
  running and the budget is dead configuration (issue #2389). The
  `stop_grace_period: 240s` in `docker-compose.yml` is what makes it reachable.
  If you run this image under your own orchestration you must grant the same
  budget there - Kubernetes has the identical trap under a different name, since
  `terminationGracePeriodSeconds` defaults to 30s.
- The drain reports its own duration, so the budget can be derived rather than
  bisected. `docker logs` carries `RepoContext drain complete in <n>s, consuming

% of the 180s host shutdown budget`. Read it together with its severity, because there are three distinct outcomes and the level is what separates them: - **No completion line at all.** The container was killed mid-drain, so `stop_grace_period` is smaller than the drain (issue #2389). The exit code will not tell you, because a killed container reports `137` and the next `docker start` overwrites it. - **`drain complete` at `Warning`.** The drain finished but consumed 70% or more of the budget. Nothing has failed; treat it as a lead indicator, because drain time grows with resident state. - **`drain ABANDONED after 180s` at `Error`, and the container exits `70`.** The *host* stopped waiting and deactivation was abandoned part-way. Raise the service's `stop_grace_period` and the `LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD` that declares it, together and to the same value; the budget is derived from the second and bounded by the first, so raising either alone achieves nothing. - **The abandoned drain is also visible without reading the log** (issue #2401). An orchestrator does not read logs, it reads the exit code, and before #2401 the abandoned case did not reliably produce a distinctive one. Measured against a real generic host, the outcome was not simply zero but *undetermined*: a hosted service that absorbed the shutdown cancellation left `RunAsync` returning normally, so the process exited `0` and the abandonment was recorded as a clean stop; one that rethrew it let the exception escape `RunAsync` unhandled, aborting the process in a way indistinguishable from a real crash. The host now assigns the code itself when the overrun latches: `0` if the drain completed inside the budget, `70` if it was abandoned. `70` is `EX_SOFTWARE` in the `sysexits.h` convention, chosen to avoid `0`/`1`/`2`, Docker's reserved `125`-`127`, and the `128 + signal` band that holds `137` (`SIGKILL`) and `143` (`SIGTERM`) - the neighbouring conditions it exists to be told apart from. - Note what this does **not** do. `docker-compose.yml` uses `restart: unless-stopped`, under which Docker restarts on any exit code, so the code neither triggers nor suppresses a restart. What it changes is what is recorded: `docker inspect --format '{{.State.ExitCode}}'` reports `70`, `docker ps -a` shows `Exited (70)`, and under Kubernetes the container terminates with reason `Error` rather than `Completed`. That is what an alert can be written against. - There is deliberately no way to turn it off. A switch restoring `0` would remove the evidence rather than the problem. - That last line exists because of issue #2397, and the reason it is needed is not obvious. The host raises `ApplicationStopped` **even when the shutdown budget expired and it gave up waiting** - so a signal bound only to that event reported `drain complete in 90.0s` for a drain that did not complete. The failure looked like success. The overrun is now raised by an alarm armed when the drain begins, so it is reported at the moment the budget expires rather than depending on a callback that may never arrive. - Drain time scales with resident state: the same 400-file rig drained in 33.9s before its vector trees had landed and 67.2s once they had, which was already three quarters of the 90s budget the host then ran with. #2397 nevertheless did **not** raise it, because measurements on a live box show the resident set that a drain must flush has no observed ceiling (idle-deactivation sweeps ranging from 1 to 4,418 activations, still climbing between readings). A fixed ceiling on an unbounded quantity moves the threshold without changing the failure mode, so #2397 shipped the diagnostic instead of a new number, and #2402 - which proposed raising it - did not ship one either. Issue #3304 later did raise it, to 180s, against two consecutive drains that no longer fitted (89.7s and 91.9s against 90s); see the [container quickstart](../../docs/lattice.api.mcp.repocontext/container/graceful-shutdown.md). - The budget is no longer written down as an independent constant. Since issue #2402 the host derives it from the grace period the deployment declares through `LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD`, taking 75% of it or all but a two-second unwind reserve, whichever is smaller. The `120s` the sample declared when that derivation landed yielded exactly the 90s the container had always run with, so nothing moved at the time (the sample now declares `240s`, deriving 180s, since issue #3304); what changed is that there is one number to set instead of two, and a budget larger than the grace period can no longer be expressed. That matters because such a budget is not merely useless - the container kills the process at the real grace period regardless, so the `ABANDONED` line above is armed for an instant that never arrives and is never emitted, which is the silent teardown of issue #2389 all over again. - **The variable declares the grant; it is not the grant.** Nothing inside the container can read the real `stop_grace_period`, so a deployment that declares `120s` while granting `20s` runs a 90s budget under a 20s guillotine and cannot detect it. Keeping the two adjacent in the same compose service is the mitigation, and `RepoContextComposeShutdownBudgetTests` asserts they are equal here - but that adjacency is **a convention, not an enforcement**. Change them together, always. - None of this bounds the resident set. Raising both values past your own observed drain is a legitimate local remedy, but it buys time rather than fixing the shape.

## Verifying what you actually deployed

Everything above describes what the tracked compose file declares. Nothing above
establishes that a container now running received any of it.

`docker compose up` reads the compose files in its **own working directory**,
whatever branch built the image it starts, and its output names no branch, no
commit, and no directory. The image and the runtime configuration are therefore
two independent inputs, and only the first is obviously version-controlled. An
operator standing in one checkout can deploy a candidate image under a different
checkout's configuration and see nothing at all to say so.

That is not hypothetical. Two gate runs of epic #2368 did exactly this: the
candidate image ran under the baseline's runtime config, the container's own
compose label resolved to the main checkout, and
`LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD` was absent from the process environment
while sitting merged on the candidate branch. Both runs observed the absence of
the fix's effect and concluded the fix was absent. The observation was real and
correctly made. The discriminator was in a channel nobody was reading.

Note what would **not** have helped. A check comparing tracked files to other
tracked files would have been green throughout both runs, because the repository
agreed with itself perfectly. Only a reading taken from the running process
separates "the source does not carry the fix" from "the source carries it and
this container never received it".

```bash
cd samples/RepoContextContainer
pwsh -File ./scripts/Assert-ContainerProvenance.ps1
```

It exits `0` only when every check agrees, and with a distinct non-zero code on
each refusal path (`2` refused, `3` the container could not be interrogated, `4`
the expected configuration could not be read), so automation can gate on it. Until
#2718 it never called `exit` at all and leaked `128` from a git probe that is
SUPPOSED to fail; the full contract is in the
[local deployment runbook](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#what-its-exit-code-means).

It performs seven checks against the running container and refuses unless all
seven agree, printing every value it read either way:

1. **Compose provenance.** The container's own
   `com.docker.compose.project.working_dir` label resolves to the checkout you
   are standing in, every file in `com.docker.compose.project.config_files` lies
   under it and exists on disk, and **the number of those files is the number you
   expected** (default 2, see below). An override file merged in from elsewhere
   is how a resolved document stops matching the tracked one.
2. **Git provenance.** That directory is a git worktree and its HEAD is the
   commit you expect, reported as a value you read rather than an inference you
   make. The expectation is sourced from *your* checkout, never from the one the
   container resolved to; defaulting it to the latter would compare a value
   against itself and pass unconditionally, which is worse than omitting the
   check because it would report as checked.
3. **Image provenance.** The image id the container is executing still matches
   what its image reference resolves to now, which catches a container left in
   place across a rebuild. Ids, never tags: a tag is a mutable pointer, so
   comparing tag to tag compares two names for whatever is current.
4. **Environment provenance.** A candidate-only setting is **present in the
   container's own environment** with the expected value, expectation read from
   your checkout's `docker-compose.yml` and actual read from the running
   process. This is the non-redundant one. Checks 1 to 3 can all pass while an
   override file, an edit, or a stale container leaves the value unset, and it is
   the direct executable form of the warning above that the variable declares the
   grant rather than being it. Durations are compared **parsed, never
   literally**: Compose normalises `120s` to `2m0s`, so a text comparison would
   accuse a correctly configured stack of exactly this defect, and the obvious
   remedy for that accusation is to change a deployment that was already right.
5. **Archive durability.** The bind mount holding durable agent memory resolves
   to an absolute host path outside every git worktree and checkout. It is the
   only check on the WRITE path, and it keys on the archive path alone, never on
   the compose directory (issue #2627).
6. **Build provenance.** The image the container is executing was built from the
   expected commit, read from the image's own `org.opencontainers.image.revision`
   label or failing that a `candidate-<sha>` tag, and cross-checked against
   chronology. It fails closed. Checks 1 and 2 adjudicate a checkout; this one
   adjudicates the image (issue #2686).
7. **Workspace provenance.** The `/workspace` bind the container indexes out of
   exists, is a bind rather than a volume, and resolves to an absolute host path -
   and, when you name them, it is the expected workspace root and a registered
   repository is indexed from the expected repository root (issue #2617).

The parameters not discussed below steer what those checks read.
`-ContainerName` names the container to interrogate; by default it is the one
carrying the compose service label for `-ComposeServiceName` (default
`repocontext`). `-ExpectedCheckout` (default: the checkout holding the script's
compose files) is the directory check 1 compares against and the checkout whose
`docker-compose.yml` check 4 reads its expectation from, and `-ExpectedCommit`
(default: that checkout's HEAD) is the commit checks 2 and 6 compare against;
`-ExpectedCommitDate` (default: that commit's date), `-CandidateTagPrefix`
(default `candidate-`) and `-ClockSkewToleranceSeconds` (default `120`) feed
check 6's tag fallback and chronology arm; and `-ArchiveDestination` (default
`/memory-archive`) and `-WorkspaceDestination` (default `/workspace`) are the
container paths checks 5 and 7 adjudicate.

### The count assertion, and why it is not a walk of tracked files

The stack's real deployment is **two** compose files. The tracked
`docker-compose.yml` carries a `build:` stanza and no `image:`, so on its own it
cannot resolve an image for `up -d --no-build` at all. A second file supplies the
image pin, the memory limit, the CPU caps, and the scan-cadence variables every
prior measurement on a given box was taken against.

That second file used to be an untracked, gitignored
`docker-compose.override.yml`, which put the entire tuned configuration on
exactly one machine with no review, no history, and no diff (issue #2609). It is
now the **tracked** `docker-compose.tuning.yml`, layered by name:

```bash
docker compose -f docker-compose.yml -f docker-compose.tuning.yml up -d --no-build
```

The rationale behind every value in it, the build-and-tag ladder that produces
the image pin, and the rollback procedure are in the
[local deployment runbook](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md).

The count is still two, so this check needs no new argument. Naming the files
explicitly also **suppresses** an automatic `docker-compose.override.yml`, which
is what makes the deployed configuration reproducible from the checkout alone;
layering a personal override on top of the tuning file makes it three, and you
say so with `-ExpectedConfigFileCount 3`.

The opt-in CPU pinning described above does **not** move the count: it is two
variables on the existing services, not a third file.

A machine-local override remains legitimate, and it is why check 1 reads the
container's label rather than walking the repository: **it has to be able to fail
on a file git has never heard of.** The obvious remedy for a compose-provenance
failure - relaunch from the checkout you meant - **silently drops** whichever
second file the original launch had, leaving no image pin, no memory limit and a
different scan cadence, while the tree looks perfectly correct and every path the
container reports still resolves under the right directory. Only the count
dissents.

The general form is worth stating, because it is not specific to compose:
*fixing a provenance defect by changing the launch directory is itself a
provenance change, and it is not self-verifying.*

Pass `-ExpectedConfigFileCount 1` if you genuinely mean to run without an
override. Making that an explicit act is the point - dropping the override
should be something you said, not something that happened.

**What a green run does not establish.** That the checkout was clean when `up`
ran, since HEAD is a commit and uncommitted compose edits are invisible here - and
a `GIT_COMMIT` stamped from a dirty tree names a commit the image does not
exactly contain. That any setting you did not name reached the process: check 4
proves only the settings passed as `-ExpectedSetting`. That the memory archive's
*content* is good, or that the indexed workspace is *current*: checks 5 and 7
adjudicate where the archive landed and which tree is bound, not what state
either is in. That the workspace root or the indexed repository is the one you
meant, unless you named them with `-ExpectedWorkspaceRoot` and
`-ExpectedRepositoryRoot` (the report prints `NOT ESTABLISHED` for each arm you
did not ask for; `-ExpectedRepositoryRoot` also needs `-IndexedRoot`, the
`indexedRoot` values `repocontext_list_repos` reports, and refuses without it
rather than passing), or that every registered repository is correctly rooted -
the indexed-root arm passes when one matches. Or anything whatever about a container
you did not name. What a green run **does** establish, since check 6 (issue
#2686), is that the image the container is executing was built from the expected
commit, on the evidence of its own revision label or `candidate-<sha>` tag
cross-checked against chronology. Check 3 as defaulted detects only a stale
container; pass `-ExpectedImageId` if you need to pin an exact image.

This is an operator check and is deliberately **not** wired into CI. It needs a
running container, and a fixture that skipped when Docker was absent would
produce exactly the false green it exists to prevent.

It is also **not** the same instrument as the cold-start rig's
`Assert-RigComposeIsolation`, and neither subsumes the other. That guard
validates the *declaration* - what `docker compose config` resolved - before
anything runs. This validates the *deployment* - what a container already running
actually received. The rig has never had this failure mode, because
`Get-RigComposeFile` pins the file it resolves, so the different-checkout drift
cannot arise there. A green rig guard therefore says nothing about this class,
and the two must not be collapsed. The boundary was drawn deliberately by issue
#2576, whose handoff named both the remedy and its location: a post-up
precondition on the container's own environment.

The adjudication is pure and separated from the acquisition, so the refusing
direction is exercised against fabricated disagreements rather than assumed:

```bash
pwsh -File ./scripts/Test-ContainerProvenance.ps1
```

Every one of the seven checks has fixtures it accepts and fixtures it refuses,
and several refusal fixtures are reconstructed from real gate-run and incident
readings rather than invented. A check only ever
observed passing is indistinguishable from one that cannot fail, which is the
same reason the suite itself is worth measuring rather than trusting: commit
first, then make one check return no violations unconditionally, re-run, and
confirm the assertions that fail are the ones covering that check and no others.

Two sibling suites cover what that pure suite cannot see, since it neither runs
git nor runs the script: `Test-ArchiveGitReading.ps1` checks the archive check's
pinned git messages against the git actually installed, and
`Test-ProvenanceExitCode.ps1` checks that the script's exit status says what its
printed verdict says.

## Measuring approximate retrieval on a running container

`scripts/Invoke-AnnQueryProbe.ps1` issues **real retrieval queries** against a
running container and reports what moved on
`repocontext_retrieval_ann_search_total`. It exists because a scrape on its own
cannot answer the question it appears to answer: that counter is written only on
the per-query path, so a deploy that never issued a query leaves every arm at
its primed zero, and **that reading is byte-identical to a plane that was
consulted and answered exactly nothing approximately.**

```bash
pwsh -File ./scripts/Invoke-AnnQueryProbe.ps1
pwsh -File ./scripts/Invoke-AnnQueryProbe.ps1 -RepoId lattice -Repetitions 4
```

It scrapes the three `state` arms before and after, issues its queries over the
MCP endpoint on 8080 (the container's only application listener), and reports
the per-arm delta.

With no `-RepoId` it probes every repository `repocontext_list_repos` reports.
Its other parameters are `-BaseUri` (default `http://localhost:8080`; pass the
published port if you moved it with `REPOCONTEXT_PORT`), `-Queries` (a built-in
spread of natural-language queries when omitted, each issued against every
selected repository), `-Repetitions` (default `1`), `-K` hits per query
(default `5`), `-TimeoutSeconds` per request (default `60`), and
`-JsonOutputPath`, which writes the full machine-readable result document.

**It refuses to report rather than reporting a zero it cannot stand behind.** The
refusal is the feature, and there are two of them, kept deliberately distinct
because they have different owners:

- **issued 0** (exit 2) - the probe never got a query out. Nothing can be
  concluded about the instrument; the fault is the probe's or the environment's.
- **issued N, succeeded 0** (exit 2) - queries went out and every one failed. The
  instrument reading is still inadmissible, but the fault is now the container's
  and is worth diagnosing.

Collapsing those two into one "no data" would discard exactly the bit that says
whose problem it is. Neither is the environment failure that stops the probe
before it can query at all - a container that does not answer `/health/live`, a
failed MCP handshake, or a failed `repocontext_list_repos` - which exits `3`.

**`retrievalPath` on a search result is not evidence about the approximate arm.**
The approximate index's declared retrieval path is a property of the index, not of a
query - one index serves every repository, so a state-tracking declaration would
be wrong the moment two repositories were in different states. It therefore
reads `semantic.approximate` unconditionally, **including when the exact fallback
answered with complete recall**, and `NormalizeSemantic` fails closed the same
way by resolving anything unrecognised to it. The declaration deliberately
under-promises. Reading it as confirmation that approximate search ran is
confidently wrong, and the probe prints that caveat rather than assuming you
know it.

**An absent arm is not a zero.** The probe distinguishes "the series is present
and reads 0" from "the series is not on the endpoint at all". The first is a
measurement. The second means the series was refused at creation, and the probe
sends you to `lattice_metrics_series` and
`lattice_metrics_dropped_measurements_by_family_total` before you conclude
anything from it.

**It reads `/health/ready`, and prints the answer verbatim.** This is a
different endpoint from `/health/live` and answers a different question.
Liveness asks whether anything is there; readiness asks whether this box can
actually serve semantic retrieval, and on the repocontext host it is the
conjunction of the lifecycle component and the vector plane. When the vector
plane is down the readiness body already says so, in specific and self-limiting
terms, and it even names the `retrievalPath` discrimination you would otherwise
have to rediscover.

The endpoint is easy to miss, and has been missed: the container healthcheck
runs a grain-liveness self-probe rather than an HTTP readiness call, so Docker
can report `healthy` straight through a total retrieval outage, and the
acceptance playbook's only outbound call is `/metrics`. The probe therefore
reads it explicitly rather than assuming something upstream already did.

A 503 here is a **successful** probe result, not a probe failure. It is the
system diagnosing itself, which is more authoritative than anything this harness
can infer from a counter delta, so the probe prints the status and the full body
and says as much. It is deliberately **not** a gate: a 503 is the expected
reading on a rig whose vector plane is down, and refusing to continue would
suppress the very measurement the harness exists to take.

**The total across all three arms is the liveness witness.** The arms partition
the whole query population - the index records an outcome for every
query including bootstrapping - so a moving total proves the instrument is
capable of reporting, independently of which arm moved. A zero on
`approximate` beside a non-zero total is a measured absence. A zero beside a
zero total is not a reading at all. Because the arms carry no caller tag, a
total delta larger than the probe's own succeeded count is reported as
`CONTAMINATED` rather than claimed: concurrent internal retrieval is
indistinguishable from the probe's own, and pretending otherwise would attribute
traffic the probe did not generate.

The probe does not deploy, does not score any acceptance predicate, and does not
tune retrieval parameters. Tuning until the approximate arm fires would encode
the answer into the instrument.

Its own refusal paths are regression-tested rather than proven once:

```bash
pwsh -File ./scripts/Test-AnnQueryProbe.ps1
```

Twelve scenarios against a real HTTP listener that the suite runs as a background
job (on `localhost`, port `18080` unless you pass `-Port`), invoking the probe as
a separate `pwsh` process, covering both refusals,
the absent-arm case, the contaminated delta, the suppressed-fallback state, and
all three readiness shapes (ready, not-ready-with-a-diagnosis, and a readiness
endpoint that cannot be read at all).
The suite asserts its own scenario count is non-zero before reporting, for the
same reason the probe asserts its issued count: a harness that ran nothing
reports success in a way that is indistinguishable from a harness that ran
everything and found nothing wrong.

## Measuring an approximate-index build

`scripts/Invoke-AnnBuildProbe.ps1` measures how long the approximate index takes
to converge for one repository on a running container - the wall-clock figure an
A/B of the build path is scored on. It registers the repository itself with
`repocontext_add_repo` (a write, and for a repository that is already registered a
fresh indexing pass), then polls `repocontext_health` for that repository every
`-PollSeconds` (default 10) and reports the time to converge, the vectors indexed,
and the sample series (written as JSON with `-JsonOutputPath`).

```bash
pwsh -File ./scripts/Invoke-AnnBuildProbe.ps1 -RepoPath /workspace/my-repo
```

`-RepoPath` is the in-container path under the mounted workspace, `-RepoId`
defaults to its final segment, as `repocontext_add_repo` itself derives it,
`-BaseUri` defaults to `http://localhost:8080`, and `-TimeoutSeconds` (default
`240`) bounds each MCP request it makes.

**Convergence is not the approximate plane reporting `Ready` on its own.** A build
over a corpus that has not been embedded yet reaches `Ready` at once with nothing
in it, so the probe waits for `Ready` with the indexed vector count caught up to a
non-zero embedded coverage. It exits `0` when that happens and `2` when
`-MaxWaitMinutes` (default 60) elapses first - an arm that does not converge is a
result to record, not a harness fault - and it fails outright if the registration
still fails after its retries.

It deliberately does not read `/health/ready`, which answers a different question
(see above), and it scores on time to converge rather than on the
`repocontext.ann.build.stage.duration` histogram, because only the former is
reported by every build an A/B might compare. Read the stage split afterwards to
explain a difference, not to score one.

## Scripts

Every script in [`scripts/`](https://github.com/NSTA1/Orleans.Lattice/tree/release/9.9/samples/RepoContextContainer/scripts), and every parameter it accepts. All
parameters are optional except `-RepoPath` on `Invoke-AnnBuildProbe.ps1`.

| Script | What it does | Parameters (default) | Described in |
|---|---|---|---|
| `New-TuningEnv.ps1` | Derives this host's resource knobs for `docker-compose.tuning.yml` and writes them to `.env`. | `-WorkspacePath` (the repository root), `-OutFile` (`.env` beside the compose files), `-DryRun`, `-CorpusOnly` (measure and print the corpus, then exit), `-IgnoreHostLoad`, `-ExpectedCorpusFiles`, `-CorpusTolerance` (`0.02`), `-Force`, and two seams for driving its refusals deterministically, `-HostMemoryBytes` and `-HostAvailableMemoryBytes` | [Runbook: before you start the stack](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#before-you-start-the-stack) |
| `Assert-TuningEnv.ps1` | Refuses a tuning `.env` whose knobs are unset or carry a retired sentinel. | `-EnvFile` (`.env` beside the compose files), `-Reading` (a hashtable checked instead of a file), `-BuildCommit`, `-RepositoryPath` (the repository holding the script), `-Quiet` | [Runbook: before you start the stack](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#before-you-start-the-stack) |
| `Test-TuningIntegrity.ps1` | Conformance suite for the tuning and attribution guards. | `-Quiet` (the summary line and any failures only) | [Runbook: before you start the stack](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#before-you-start-the-stack) |
| `Assert-DeployManifest.ps1` | Records the deployment's attribution-relevant configuration and refuses a multi-variable step nobody acknowledged. | `-Label` (a timestamp), `-ComposeDirectory` (this directory), `-ManifestPath` (`.deploy/manifests/<label>.manifest` here), `-BaselinePath` (the newest manifest already there), `-AcceptMultipleDeltas` with `-Reason`, `-ContainerName` (`repocontextcontainer-repocontext-1`), `-EmbedContainerName` (`repocontextcontainer-embedder-1`), `-CgroupCpuQuota` (read from the container), `-Quiet`, and the test seams `-DeclaredReading`, `-EffectiveReading` and `-ComposeConfigFilesLabel` | [Runbook: restart, drain, and verification](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#restart-drain-and-verification) |
| `Test-DeployManifest.ps1` | Shows the deploy manifest's declared-versus-effective check both firing and visibly not firing. | `-SkipAcquisition` (skip the real compose resolution) | [Runbook: restart, drain, and verification](../../docs/lattice.api.mcp.repocontext/local-deployment-runbook.md#restart-drain-and-verification) |
| `Assert-ContainerProvenance.ps1` | The seven-check provenance guard over a running container. | See its section. | [Verifying what you actually deployed](#verifying-what-you-actually-deployed) |
| `Test-ContainerProvenance.ps1`, `Test-ArchiveGitReading.ps1`, `Test-ProvenanceExitCode.ps1` | The provenance guard's own suites. | None. | [Verifying what you actually deployed](#verifying-what-you-actually-deployed) |
| `Invoke-AnnQueryProbe.ps1`, `Test-AnnQueryProbe.ps1` | The retrieval-plane query probe and its suite. | See its section. | [Measuring approximate retrieval on a running container](#measuring-approximate-retrieval-on-a-running-container) |
| `Invoke-AnnBuildProbe.ps1` | The approximate-index build probe. | See its section. | [Measuring an approximate-index build](#measuring-an-approximate-index-build) |
| `Test-BackupSinkDurability.ps1` | Shows the backup sink's host directory surviving `docker compose down -v`. | `-ProjectName` (`repocontext-backup-durability-probe`), `-SinkPath` (a fresh directory under the system temp path), `-Force` (run even while another compose project's containers are up) | [Container quickstart](../../docs/lattice.api.mcp.repocontext/container/agent-memory-backup-and-recovery.md#what-survives-and-what-does-not) |
| `_deployManifest.ps1`, `_mcpClient.ps1`, `_provenance.ps1`, `_tuningKnobs.ps1` | Shared functions the scripts above dot-source; not run on their own. | None. | - |
