---
title: "Retrieval and token economics"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.mcp.repocontext/retrieval-economics.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.api.mcp.repocontext/retrieval-economics.md"
package: "Orleans.Lattice.Api.Mcp.RepoContext"
status: "unreleased"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.mcp.repocontext/llms-full.txt"
---
# Retrieval and token economics

Part of the [Api.Mcp.RepoContext documentation](README.md).

The repository-context surface is only useful to an agent if the context it
returns is worth the tokens it costs. This topic covers the retrieval and
token-economics capabilities layered on top of the [record model](record-model.md):
explainable search, the graph-navigation tools, the budgeted context bundle, its
reuse economics, usage accounting, and the shared token counter they all budget
in. Every tool here is read-only and clears the same fail-closed authorization
gate as the rest of the surface.

## The shared token counter

A single byte-pair-encoding (BPE) token counter underpins every figure on this
page, so a token cost reported by one tool means the same thing everywhere. It is
constructed once from a tokenizer **profile** and reused by the reconcile path
(the per-file `TokenCount` on each file node) and the retrieval surface (bundle
budgets and outline costs).

The profile is selected by the `LATTICE_REPOCONTEXT_TOKENIZER` environment
variable: `o200k` (the default, matching current-generation models) or `cl100k`.
Because the same counter is used to compute the stored per-file counts and to pack
a bundle, a bundle's `totalTokens` is an exact sum over the same encoding the
consuming model uses, not an estimate. The bundle's `responseTokens` builds on that
exact content sum but is itself a deliberately conservative **estimate**, because it
also accounts for JSON envelope and the SDK's dual emission (see
[The budgeted context bundle](#the-budgeted-context-bundle)).

## Explainable search

`repocontext_search` ranks records against a natural-language query and returns
each hit hydrated from the store. Beyond the `mode` field (`semantic`, `keyword`,
or `empty`) that reports which path answered, every hit carries a machine-readable
**`reasons`** list: server-derived, deterministic, ordinal-ordered, bounded, and
never null.

- A **semantic** hit lists `semantic`, the matched chunk kind (`chunk:symbol`,
  `chunk:file`, or `chunk:memory`), and `symbol:<fqName>` when the match was a
  symbol vector or `topic:<topic>` when it was a memory-entry vector.
- A **keyword** hit lists whichever projected fields the query terms actually hit,
  in a fixed high-signal-first order: `path-name-match`, `symbol:<fqName>`,
  `tag:<tag>`, `topic-match`, `content-match`, and `key-match`.

The reasons let an agent (or a human reviewing a trace) understand *why* a result
ranked where it did, rather than treating the ranking as opaque - and they let a
caller decide which hits are worth a full read before spending the tokens.

## Graph navigation

Three read-only tools let an agent navigate the code graph without reading whole
files, each a bounded read over stored records that never touches the workspace on
disk (except `repocontext_changed`, which walks the workspace only through the
fail-closed boundary):

- **`repocontext_outline`** returns a file's declared-symbol skeleton - each
  symbol's kind, signature, and 1-based line span, ordered by position - plus the
  token cost of reading the whole file. It is the cheapest way to grasp a file's
  shape and decide whether a full read is worth the tokens.
- **`repocontext_related`** resolves a file's structural neighbourhood: the
  type-names it references (outbound imports), the indexed symbols that reference
  its declarations (inbound dependents, resolved to their declaring files), and
  the test types that cover it. Dependents and tests come from the reverse
  cross-reference projection, so the lookup is bounded rather than a
  whole-repository scan.
- **`repocontext_changed`** reports how the current workspace has drifted from the
  index - files added, updated, and removed - by comparing content digests without
  invoking git, and lists the indexed files that depend on the changed ones (the
  reverse-reference impact set), so an agent sees the blast radius of a set of
  edits before re-indexing. The walk is rooted at the repository's *indexed* root
  and reuses the filters it was ingested with, so the report always compares the
  same path space the index was built in; the supplied path is a scope, so a
  directory inside the repository restricts the report to that subtree, and a path
  outside the indexed root is refused rather than compared. Unchanged files are
  settled by a stat against the stored size and ingest anchor instead of being
  re-read, the same fast path the periodic reconcile uses, so a whole-repository
  drift report stays cheap on a large tree.

## The budgeted context bundle

`repocontext_context` is the headline capability: it collapses the
search -> recall -> read loop into a single round trip that can never overrun the
context budget. Given a natural-language task and a token budget, it searches the
store, resolves the top hits to unique files, and packs each file at a **detail
level** under a hard token ceiling:

- `paths` - the path only.
- `outline` - the declared-symbol skeleton, reusing the outline projection.
- `slices` - bounded body text: the file's content projection, which holds at
  most its first 65,536 characters.
- `auto` (the default) - the richest level that still yields a non-empty bundle,
  with the concrete level reported back in `detail`.

Every entry carries its match `reasons`, its exact BPE `tokenCount`, and the
whole-file `fullReadTokenCount`. The bundle reports **two** figures, and it is the
first that the ceiling bounds:

- `responseTokens` - the estimated cost of the response **as the caller receives
  it**: the delivered content plus each entry's JSON envelope (path, reasons,
  content hash, per-unit receipts), multiplied by the MCP SDK's dual-emission
  factor, because every tool result is serialized twice - once as structured
  content and once as text. The factor is a conservative 3.5 rather than 2,
  because the text copy is escaped JSON, which tokenizes worse than the structured
  copy. This **never exceeds** `budgetTokens`; an empty bundle reports `0`.
- `totalTokens` - the narrower exact BPE sum of the packed source text alone.
  Useful as "how much source did I get", but it is not what the budget bounds:
  charging content alone once let a bundle reporting a few thousand tokens land
  as a response many times that size (issue #1811).

The estimate is deliberately conservative, so a bundle may come in slightly under
the ceiling but never over it. When even the cheapest entry does not fit,
the tool **fails closed**: `entries` is empty and `retryBudgetTokens` reports a
budget guaranteed to admit at least one entry on a retry (null when the search
matched nothing, so no larger budget would help). A `truncated` flag marks a
bundle that had to drop lower-ranked candidates. The `top`, `responseBudgetTokens`,
and `detail` arguments are validated and clamped, never trusted to drive unbounded
work: `top` to at most 50 files (default 10) and `responseBudgetTokens` to at most
200000 tokens (default 8192), with 0 or less meaning the default, and an
unrecognised `detail` to `auto`.

## Reuse economics

The bundle never makes an agent pay twice for context it already holds. Each
delivered **unit** - a path pointer, a body span, or an outline symbol - carries a
stable opaque `receipt`, and each entry carries a per-version `contentHash`. A
unit is a **descriptor, not a copy of the text**: the delivered text lives once,
on its entry's `content`, and the units correspond one-to-one, in order, to that
content's newline-separated segments. (Carrying the text on the units too would
put every byte of source on the wire twice within a single payload, and four
times across the emitted pair - see issue #1811.) A caller feeds prior knowledge
back in three ways:

- Hand receipts back in **`seen`** to suppress exactly those units; the rest of
  the file still arrives.
- Assert whole-file possession in **`known`** as `path@hash`.
- Pass a **`session`** id to persist this bookkeeping across calls: the session
  auto-suppresses units it already delivered and validates `known` claims, so a
  multi-call conversation converges on delivering each unit once.

The load-bearing guard is that a whole-file claim is honoured **only** for a
version that was actually delivered as a complete body. The session store records
possession only for `slices` (whole-body) deliveries; a `known` claim is validated
only against recorded possession. So partial evidence (an outline or a path) can
never be promoted to whole-file possession, and without a session a `known` claim
can never validate (fail closed). Suppressed content is acknowledged in `reused`
and is **never** charged against `top` or the token budget - a fully-reused file
does not consume a result slot, so the freed budget backfills lower-ranked
candidates.

The per-session bookkeeping lives on the `repo-context-session` tree as a
grow-only CRDT with a finite time-to-live; see
[record-model.md](record-model.md#session-reuse-bookkeeping) for the storage
model.

## Usage accounting

`repocontext_stats` reports whether the surface actually reduces context cost. Over
a bounded recent window (the last hour, held in one-minute buckets, so `windowSeconds`
reads 3600) it returns only summed token figures:

- `calls` - how many context calls were answered.
- `responseTokens` - the response tokens they spent, charged at each bundle's own
  `responseTokens`: the conservative wire-cost estimate described above, not a BPE count.
- `readsReplacedTokens` - the whole-file read tokens they conservatively replaced,
  credited only for delivered whole-file-equivalent content (`slices` detail),
  never for discovery, partial detail, or content the caller already held.
- `netSavedTokens` - the net tokens saved, `readsReplacedTokens - responseTokens` (a
  signed figure; see below).
- `windowSeconds` - the length of the reporting window.

Crediting is deliberately conservative: reused or suppressed content is structurally
excluded, and only `slices` deliveries earn read-replacement credit. It is not a strict
floor, though: each `slices` entry is credited at the file's stored whole-file token
count, so a file longer than the 65,536-character content projection a `slices` body
carries is credited in full for the prefix it delivered. Because crediting is this
conservative, `netSavedTokens` is
**signed** and routinely negative for discovery-heavy or reuse-light usage - that is
correct, not a defect. It turns positive as a task delivers real bodies (`slices`) and
reuses a `session` so repeated context is suppressed and never re-charged, and it is
deliberately not clamped at zero so the surface can honestly report when it is not yet
paying for itself. The charge side has one exception: an empty bundle - one that failed
closed, or whose search matched nothing - reports `responseTokens` 0 although its
scaffolding still ships, so it is counted as a call that spent nothing. The figures are recorded per answered context call on a
bounded in-memory window and are also emitted as
`System.Diagnostics.Metrics` counters carrying a low-cardinality `command`
tag, so a host already scraping OpenTelemetry sees them flow through the existing
[telemetry](../lattice.api.mcp.telemetry/README.md) surface with no bespoke plumbing.
The tool carries no body, query, path, or repository identity - aggregate figures
only.

### Emitted instruments

Every instrument this package publishes is listed below, all on the
`Orleans.Lattice.Api.Mcp.RepoContext` meter, so one scraper subscription covers the
whole surface. Each carries only low-cardinality tags - never a path, query, or any
body text. Every instrument also carries the repository-wide derived `tenant` tag
(`LatticeTenantLabel.TagTenant`), which the Tags column leaves implicit: it is fixed to the
platform sentinel `_platform_` on every instrument except `repocontext.vectorplane.rederive`,
which derives it from its `tree` tag and so reports the `default` tenant for these bare
tree names. Twelve instruments carry a repository id, deliberately. Four are on the
approximate-index build plane, `repocontext.ann.build.slice` and
`repocontext.ann.build.step.in_flight_seconds`, `repocontext.ann.vectors`, and
`repocontext.ann.partitions` - the exception argued in the
`repocontext.ann.build.slice` row below: the approximate-index build plane is one
coordinator per onboarded repository and embedding space, an operator-chosen set of
single or low double digits, and issue #2855 established that without it fifteen
failing repositories and one succeeding repository were the same series. The other
eight are the `repocontext.ingest.*` family (issue #3151), on the same precedent: the
tag is bounded by the repositories onboarded on the host, and it is what makes a
stalled ingest attributable, because a silo-wide reading would let a healthy
repository hide a wedged one. One
instrument, `lattice.repocontext.memory.restore`, is named outside the
`repocontext.` prefix the others share, so a selector written against that prefix
does not match it.

| Instrument | Kind | Unit | Tags | What it records |
|---|---|---|---|---|
| `repocontext.calls` | `Counter<long>` | `{call}` | `command` | Answered repocontext calls, tagged by tool name. Only the context bundle records usage, so `command` reads `repocontext_context` on every series these three counters carry. |
| `repocontext.response_tokens` | `Counter<long>` | `{token}` | `command` | The response tokens those calls spent, charged at each bundle's conservative `responseTokens` wire-cost estimate rather than a BPE count. |
| `repocontext.reads_replaced_tokens` | `Counter<long>` | `{token}` | `command` | The whole-file read tokens they conservatively replaced. Credited only for `slices`-detail deliveries, at each file's stored whole-file token count, so it is a floor except for a file longer than the 65,536-character content projection a `slices` body carries, which is credited in full for the prefix it delivered. |
| `repocontext.retrieval.ready_seconds` | `Histogram<double>` | `s` | `phase` | Seconds from host start to the retrieval plane first reporting ready, tagged by the phase it reached. Recorded once per process, so it is the cold-start time-to-retrieval-ready figure. |
| `repocontext.retrieval.unavailable` | `Counter<long>` | `{event}` | `cause` | Observed vector-plane fault episodes that made semantic retrieval unavailable, tagged by cause: `keyword.vector_plane_unavailable`, `keyword.index_degraded`, `keyword.exact_fallback_suppressed`, `probe` (a readiness probe rather than a query observed it), `saturated` (the vector plane's open was refused by an admission gate past its declared bound, so the plane is not expected to arm at the present capacity - issue #3286), or `unknown`. A non-zero rate is what distinguishes a keyword answer caused by a real capability loss from an intended keyword-only deployment. All six arms are pre-minted at zero when the readiness state is constructed, so each exists from process start and 'retrieval has never degraded on this process' is a measured absence rather than an absent measurement; before this the counter published no series at all until the first fault, so the healthy reading and an unwired instrument were identical. If an arm is *absent* rather than zero, read `lattice_metrics_series` against the collector ceiling and `lattice_metrics_dropped_measurements_by_family_total` before concluding anything, because a series whose first occurrence falls after a ceiling is reached is refused at creation. |
| `repocontext.retrieval.ann.search` | `Counter<long>` | `{query}` | `state` | Semantic searches partitioned by the approximate-plane state that answered them: `bootstrapping` (the plane could not answer, so the fallback ladder ran), `exhaustive` (answered by scanning the vectors it holds), or `approximate` (answered from its trained partitioning). Because **every** answered query is counted, the total is a denominator: `approximate` pinned at zero beside a rising total is a measured absence of trained serving, not an absent measurement. All three arms are pre-minted at zero when the reporter is constructed, so each exists from process start; if an arm is *absent* rather than zero, read `lattice_metrics_series` against the collector ceiling and `lattice_metrics_dropped_measurements_by_family_total` before concluding anything, because a series whose first occurrence falls after a ceiling is reached is refused at creation. |
| `repocontext.ann.sweep` | `Counter<long>` | `{sweep}` | `outcome`, `cause` | Approximate-index build sweeps partitioned by outcome: `armed` (armed at least one build coordinator), `empty` (completed without arming anything, either because it observed no repository in the store listing or because every repository it observed declined arming - this arm reports what the sweep observed, never that the store is empty), or `faulted` (threw, so nothing is scheduled until a sweep gets through). Denominate by the **total across all three arms**, never by the configured sweep interval: a completing sweep advances at that interval, but `faulted` advances on the far faster retry backoff (250 ms doubling to a 30-second ceiling). All three arms are pre-minted at zero when the reporter is constructed, so all three reading zero means no sweep has completed yet. That does not localise the fault and in particular does not establish that the sweep loop is not running, which the service's startup log line reports directly; an arm that is *absent* rather than zero points at collector saturation, read from `lattice_metrics_series` and `lattice_metrics_dropped_measurements_by_family_total`. The `faulted` arm alone carries a second tag, `cause`, resolved where the fault is raised: `authority-unavailable` (resolving the run credential threw, so nothing was attempted), `listing-unavailable` (the repository listing threw), `plane-rejected` (a coordinator was reached and refused the arming call, which will not clear on retry), `dependency-unavailable` (a coordinator could not be reached; a grain-call timeout is a busy coordinator and lands on `repocontext.ann.sweep.arming` as `deferred` instead), or `unexpected` (unclassified, and the only cause that should page). The causes are deliberately not pre-minted, so read them only once `faulted` is non-zero. |
| `repocontext.ann.sweep.arming` | `Counter<long>` | `{repository}` | `result` | Arming calls the sweep made, partitioned by what the coordinator answered: `armed` (the coordinator accepted, so a build is scheduled), `deferred` (the call timed out because the coordinator is non-reentrant and already inside a long build turn, which is the expected answer from a healthy coordinator mid-build and deliberately **not** a fault), or `faulted` (the call threw something other than a timeout). Its denominator is **repository visits**, where `repocontext.ann.sweep` counts **sweeps**, and one sweep visits every repository in the listing - so the two are different populations, neither decomposes the other, and a ratio between them means nothing. This series exists because the sweep outcome could not carry the deferral at all: a sweep arming one repository and deferring nine reported `armed` exactly as one that armed ten, and a sweep on which **every** coordinator deferred fell through to `empty`, indistinguishable from a sweep over a store holding no repositories - two states at opposite extremes, one observation. A rising `deferred` beside `armed` at zero is a wholly wedged build plane; a `result` partition totalling zero beside a completed sweep is a genuinely empty listing. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from `lattice_metrics_series` and `lattice_metrics_dropped_measurements_by_family_total`. |
| `repocontext.vectorplane.rederive` | `Counter<long>` | `{event}` | `tree`, `outcome` | Rebuildable vector-plane tree fall-off observations and re-derivations: `observed` (an allowlisted fall-off was seen and a reset triggered), `completed`, `failed` (a transient fault stands and a later pass retries it), `denied` (the access gate refused the reset, which is deterministic and will not clear on retry), `suppressed` (a fall-off was seen inside the post-failure backoff window, so no reset was attempted), or `refused` (the tree is not a rebuildable derived tree, so re-derivation is declined fail-closed and the fault propagates). The `tree` tag is one of the fixed vector-tree names, never a repository id; only the vector-metadata and vector-membership paths are guarded, so in practice it is one of those two, and `refused` is a defence-in-depth arm no current call site reaches. |
| `repocontext.ann.build.corpus` | `Counter<long>` | `{build}` | `coverage` | Approximate-index builds that reached `Ready`, partitioned by how much of the repository's vector prefix the read-path access gate admitted: `nonempty` (the build holds vectors, so the corpus read plainly succeeded and no probe was taken), `unrestricted` (it holds nothing and the whole prefix is admitted, so the repository genuinely has no vectors), `filtered` (it holds nothing and the gate narrowed the prefix, so an unknown subset was withheld), `denied` (it holds nothing because the gate refused the prefix outright, so the read never happened and the index state is **unknown** rather than empty), or `unknown` (it holds nothing and the coverage probe could not answer). A denied *range* read returns a clean, successful, empty result rather than throwing, so without this partition an authorization failure and an empty repository are the same observation. Because **every** completed build is counted, the total is a denominator: `denied` pinned at zero beside a rising total is a measured absence of denial, not an absent measurement. All five series are pre-minted, so `denied` is present and reads `0` on a healthy host rather than being missing. |
| `repocontext.ann.build.denial_terminal` | `Counter<long>` | `{coordinator}` | none | Build coordinators that observed enough consecutive uninterpretable corpus reads to conclude the host is refusing them, and keep retrying on a lengthening backoff that settles at one attempt roughly every five minutes, instead of on the two-second phase cadence. Counted once per denial episode, so it separates "a denial happened and is being retried" from "this deployment is permanently refused and the approximate plane will never build" - which a monotonically rising `repocontext.ann.build.corpus{coverage="denied"}` cannot do on its own. Any non-zero value warrants an operator: an index that should exist does not, and it will not appear by itself. |
| `repocontext.ann.build.slice` | `Counter<long>` | `{step}` | `repository`, `space`, `phase`, `progress`, `cause` | Build steps taken by an approximate-index coordinator, partitioned by what the step achieved, compared against the same coordinator's previous reading: `advanced` (it banked vectors or resolved partitions, so the build got closer to serving), `churned` (it moved the build phase while banking no vector and resolving no partition, so the step did something and got the build nowhere), `starved` (its slice hit the ingest deadline having consumed nothing, so the source could not be read fast enough to bank a single vector), `idle` (the step completed and changed nothing measurable, which is what a coordinator stepping over a corpus it cannot consume looks like), or `faulted` (the step threw rather than completing, so the store-of-record read could not be served at all - a different condition from a read that was served and returned nothing, and one whose remedy is the projection or the store behind it rather than the access gate or the embedding throughput). This is the only series on the approximate-build plane that fires **before** a build reaches `Ready`. Every other one is terminal - `repocontext.ann.build.corpus` at `Ready`, `repocontext.ann.build.denial_terminal` at a terminal denial, `repocontext.ann.partitioning` on a plane that has finished building, `repocontext.ann.sweep` at arming and never again - so without it the whole interval between `sweep{outcome="armed"}` and `Ready` emits nothing, and a coordinator grinding through slices that bank nothing is byte-identical in telemetry to one that never took a step. Both present as every arm of every counter at its primed zero. Read it beside `repocontext.ann.sweep{outcome="armed"}`: a total pinned at zero beside a non-zero `armed` means the coordinator is **not stepping**, which is a scheduling fault; a rising `starved` means a source read that cannot complete inside the slice; a `churned` arm rising without bound beside a flat `advanced` means it is stepping and moving between phases while banking nothing, which is the #2791 `Training -> Persisting -> Training` livelock and which read as steadily rising `advanced` until #2818 split the arm (a healthy build churns at most once per phase transition, so its churn arm is bounded by the phase count); a rising `idle` means it is stepping over a corpus it cannot consume; and a rising `faulted` means it is stepping and throwing, which is the reading that would otherwise be indistinguishable from not stepping at all, because a step that throws never reaches the record taken after it completes. The `faulted` arm alone carries a second tag, `cause`, naming the class of fault the tick raised: `scan-page-stalled`, `projection-stale`, `dependency-unavailable`, `response-timeout`, `plane-rejected`, `saturated`, or `unexpected`. `response-timeout` means only that a response deadline expired: a keepalive call can queue behind a busy build, so a timeout proves neither busyness nor unreachability. Check progress and logs before declaring a stall. `dependency-unavailable` remains the transport or silo/message-rejection arm. It exists because the run-12 build faulted on every tick and localising it needed an 8 MB container log read by hand, which established that 39 of 39 faults were a single condition - a grain-call timeout reading the vector-index tree - that the counter alone could not name. Causes are classified across the **whole tick**, not only the build step: five further call sites on a tick can throw, three of them before the step is counted at all, and a fault at any of those previously left the tick silent in this series entirely. The twenty-two `(phase, progress)` arms of a plane are pre-minted at zero the first time that plane is armed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from `lattice_metrics_series` and `lattice_metrics_dropped_measurements_by_family_total`. The `cause` values are deliberately **not** pre-minted: priming them would mint seven faulted-arm series on a host that has never faulted, so anything counting series rather than values would read a healthy host as a faulting one. The consequence is that a cause reading zero is **uninterpretable rather than innocent** - it says only that no fault of that class has been recorded since process start, which is equally what a healthy host and a mis-wired classifier look like. Read a cause only once `faulted` itself is non-zero, and denominate the causes against it: they sum to it exactly. Three further dimensions - `repository`, `space`, and `phase` - name **where** a step was taken, and were added by issue #2855 because without them the series could not answer the question the epic's acceptance run turns on. `phase` is the load-bearing one: it names the stage of the build the step was executing when it was counted, and on the `faulted` arm it is read at the fault site from the index's own progress rather than snapshotted on entry, so it is exact rather than approximate. That precision is not incidental. A build step entered in `training` trains the index and then persists the trained partitioning **within the same step**, so an entry snapshot would file a persist fault under `training` and destroy the one distinction the tag exists to draw: a fault reading `ingesting` is a corpus that could not be read, which is an independent defect, while a fault reading `persisting` is a trained index that could not be **written into the vector-index tree** - and if that tree is itself the subject of an open fault, the persist fault is a downstream symptom of it rather than a second defect, so scoring the two separately would count one defect twice. The values are `coordinating`, `opening`, `ingesting`, `training`, `persisting`, and `reconciling`; the first two precede any build step and are therefore reachable on the `faulted` arm alone, which is why a plane primes twenty-two arms rather than thirty - the eight combinations a completing step can never reach are deliberately not minted. `repository` and `space` bound the plane the step belongs to. Their cardinality is the product of onboarded repositories and embedding spaces, which is exactly the number of durable indexes and build coordinators the host already runs - one apiece - so it is operator-chosen and small, and is **not** the per-grain cardinality class of issue #2518. `space` reads as `{model-id}/{dimension}`, or `unspecified` on a plane whose space is not yet resolved. A plane re-derived onto a new embedding model is a different build over a different corpus, so merging the two would hide a migration mid-flight. Priming is per plane and happens when a plane is first armed rather than when the process starts, because a process cannot know which planes exist and a primed series for a plane nobody armed claims a build nobody asked for; the consequence is that a host with no armed plane emits no series for this instrument at all, which is the correct reading of a host that has never built an index. |
| `repocontext.ann.build.slice.items` | `Counter<long>` | `{item}` | none | Source items consumed by the approximate-index build slices that completed, summed across slices; a slice that throws records nothing. It exists to make **vectors per slice** obtainable from metrics alone, and that is a narrower gap than it sounds: the obvious way to compute it - dividing `ann.vectorsIndexed` from a health payload by `repocontext.ann.build.slice{progress="advanced"}` - is **invalid**, because the first is cumulative and is inherited across a restart whenever the index was restored from durable state, while the second is process-scoped and resets at the deploy boundary. The ratio of two counters with different epochs is not a rate of anything. With this arm the quantity is a delta of two series that share an epoch. Read it against `repocontext.ann.build.stage.duration` below: a slice consuming far fewer items than `IngestBatchSize` allows, while the stage split shows the time in `key_assign` or `source_wait`, is a per-item round trip rather than a slice that ran out of budget. A slice that consumed nothing records nothing here, so this counter does not double as a slice count - `repocontext.ann.build.slice` is that. |
| `repocontext.ann.build.stage.duration` | `Histogram<double>` | `s` | `stage` | Seconds one stage of a build slice took, accumulated across the items of that slice and tagged by stage: `source_wait` (awaiting the source enumerator, which on this host streams over grain calls), `key_assign` (mapping identifiers to index keys, including a durable block reservation when one falls due), `index_upsert` (the in-memory insert), and `key_flush` (the single batched write that makes the slice's key-map records durable). Separating them is the whole point, and the retrieval path has had the same split for the same reason: an end-to-end build rate cannot distinguish a slow source from slow key assignment from a slow index, and those are three different owners with three different fixes. Before this existed the distinction had to be made by **reading the source**, which is not a thing an operator can do against a running container. The four do not cover the whole slice: the source count taken while the expected count is still unknown, the ingest checkpoint that then persists the slice's vector chunks and build state, and the loop's own bookkeeping belong to no stage, so the four sum to less than the slice's elapsed time. A slice records all four only once it has checkpointed; a slice that throws records nothing here. Deliberately **not** zero-primed: priming a histogram fabricates a zero-valued sample that reads as a real measurement of an instantaneous stage and destroys the distribution the instrument exists to report, so an absent stage here means no slice has completed rather than a stage that took no time. The timings arrive through the vector package's public `IVectorIndexBuildObserver` seam, which the host binds as each index's `DurableVectorIndexOptions.BuildObserver` (see [the vector index configuration](../lattice.vector/configuration.md)); sampling is skipped entirely when no observer is bound, so a build with no reporter wired makes no extra clock reads - which matters because the slice budget is measured against that same clock. |
| `repocontext.ann.build.step.in_flight_seconds` | `ObservableGauge<double>` | `s` | `repository`, `space`, `phase` | How long the approximate-index build step currently executing has been inside its present phase, or `0` when no step is executing. Every other instrument on this plane - `repocontext.ann.build.slice` included - is recorded at a **terminal** moment: after the step returns, after the open finishes or throws, or once the build reaches `Ready`. A phase that never terminates therefore emits nothing on any of them, and the resulting all-zero reading is byte-identical to a coordinator that is not stepping at all. Run 14 of epic #2368 hit exactly that: a build sat inside a single non-reentrant coordinator turn for over four minutes, and the only evidence available anywhere was Orleans' own generic "request has been active for 00:04:00" warning in a container log, while six dedicated arms on this plane all read zero. An observable gauge is the **only** instrument shape that can close that gap, because the collector drives it rather than the step reaching an end it may never reach - a counter or histogram is written by the code path being measured, so a path that does not complete suppresses its own evidence (issue #3130). The clock restarts on every phase entry rather than running for the whole step, because the question is not "has this step been slow" but **which half of it is not returning**: an `ingesting` arm climbing without bound is a corpus-read defect, a `persisting` one is an index-write defect, and the two have different owners and different remedies. Every primed plane emits every one of the six phases, so a healthy idle build reads as six zeros rather than as an absence - a gauge that emitted only while a step was in flight would make health and a mis-wired instrument the same observation, which is the conflation issue #2952 removed from the slice counter. Steps are tracked by token rather than by plane, so an outstanding step that never returns cannot be erased by the next step that does. Priming is per plane and happens when a plane is first seen, so a host with no armed plane emits no series for this instrument at all. |
| `repocontext.ann.partitioning` | `Counter<long>` | `{observation}` | `state` | Approximate-index planes observed at each maintenance turn, partitioned by whether the plane holds a partitioning and, when it does not, by why: `partitioned` (it answers from a trained partitioning), `unpartitioned-small` (it holds none and its corpus is below `MinimumTrainingCount`, which is the correct state at that size), or `unpartitioned-large` (it holds none although its corpus is at or above the minimum). The third arm is the entire reason the partition exists: before issue #2706 those two cases were the same observation, so a plane holding 7.5x the threshold across zero partitions was indistinguishable from one that was simply too small to train, and every semantic query was answered by brute-force scan with nothing reporting it. A sustained non-zero `unpartitioned-large` means exactly that, and warrants an operator. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from `lattice_metrics_series` and `lattice_metrics_dropped_measurements_by_family_total`, not a statement about the plane.
| `repocontext.ann.repartition` | `Counter<long>` | `{training}` | `outcome` | Training passes taken because a plane's corpus crossed the training minimum (1,024 vectors) after an earlier training had declined to partition it, by outcome: `partitioned` (the pass produced a partitioning, so the plane now serves approximate) or `declined` (the corpus met the minimum yet still resolved to fewer than two partitions, so the plane stays exhaustive and exact and the next attempt waits until the corpus has doubled). This series reports the *repair*, where `repocontext.ann.partitioning` reports the *state*, so it is expected to read zero forever on a deployment that partitioned on its first build; a single `partitioned` is one latched plane healing itself. A rising `declined` means the corpus keeps crossing the training minimum while the partition count still will not resolve, which further corpus growth inside one activation will not change quickly. Both arms are pre-minted at zero. |
| `repocontext.ann.index.load` | `Counter<long>` | `{attempt}` | `outcome`, `reason` | One arm per load attempt: `fresh`, `resumed`, `faulted`, `deferred`, `refused`, or `discarded`. `discarded` means an open successfully removed unverifiable derived state and will rebuild from source; it is not also `fresh`. Reasons are bounded: discarded has `count_mismatch`, `embedding_space_change`, or `unloadable_record`; faulted has `timeout`, `embedding_space_change`, `unloadable_record`, or `other`; refused has `admission_refused`. Fresh/resumed/deferred carry `none`. No exception text or keys enter labels. Every legal pair is pre-minted at zero. Deferred means the open budget yielded with progress retained, or the restore found a record missing from one store read but returned by another and kept the durable index to retry (issue #3905); refused means the replay admission gate rejected work, not a load fault. Resumed includes progress retained after either yield or a fault. A valid manifest with an older build-state generation is adopted by existing recovery, not discarded. Warnings identify discard/fault reasons, and a discard of committed state names the generation, vector count and partition count it destroyed; refusal logging retains the admission distinction. **Label migration:** consumers pinned to the old exact label set must aggregate away `reason`, for example `sum by (outcome, tenant) (repocontext_ann_index_load_total)`. |
| `repocontext.ann.vectors` | `ObservableGauge<long>` | `{vector}` | `repository`, `space`, `count` = `held`, `expected` | Last published `VectorIndexBuildProgress.VectorsIndexed` and `VectorsExpected` for each registry handle. Held is the resident indexed count, not the source size or a durability guarantee; expected is the last source count observed by the build, not a live census. Labels match the build-slice plane. Open/restored and failed-step progress are published too; unopened handles read zero, no handles emit no series, and disposing the registry retires its series. |
| `repocontext.ann.partitions` | `ObservableGauge<long>` | `{partition}` | `repository`, `space` | `VectorIndexBuildProgress.PartitionsTotal` from the same handle snapshot; zero means no trained partitioning, including a small completed exhaustive index. It is not the number of persisted partitions. The same lifetime and label rules as `repocontext.ann.vectors` apply. |
| `repocontext.retrieval.duration` | `Histogram<double>` | `s` | `tool`, `path` | End-to-end seconds for one retrieval tool call, tagged by the tool (`search`, `context`, `outline`, `related`) and by the retrieval path that answered it (`semantic.exact`, `semantic.approximate`, one of the `keyword.*` causes, `not_applicable` for a graph read that consults no vector plane, or `unresolved` for a call that ended - cancelled or faulted - before a path was settled). The path tag is what makes the figure interpretable rather than merely present: a fast keyword answer and a fast approximate answer mean opposite things about the health of the box. Recorded from a `finally`, **once per call, on every termination including cancellation and failure**, so the count is a true call total and is the denominator for the stage series below; the only measurement it can lose is one whose process died mid-call. A `context` call records exactly one row under `tool="context"` and none under `tool="search"` even though it runs a search internally, so the two tools' latencies never contaminate each other. |
| `repocontext.retrieval.stage.duration` | `Histogram<double>` | `s` | `stage`, `path` | Seconds spent inside one stage of a retrieval, tagged by stage - `embed` (the network hop to the embedding service, covering both its availability probe and the query embed), `vector_search` (the index scan), `hydrate` (reading each matched identity back from the store of record, which the index never returns a second copy of), `keyword_scan` (the BM25 fallback) - and by the same `path` value the enclosing call resolved to. Separating them is the whole point: an end-to-end figure alone cannot distinguish a slow embedder from a slow index from a slow store, and those are three different owners with three different fixes. **A stage is recorded if and only if it ran**, including when it ran and then failed, so this series is deliberately sparse and its zero does not describe itself. Denominate it with `repocontext.retrieval.duration` above: no `embed` beside a rising call total is a *measured* absence of the semantic path - an intended keyword-only host - whereas both reading zero means no retrieval ran at all. |
| `repocontext.retrieval.exact_gather.faults` | `Counter<long>` | `{fault}` | `fault` | Exact k-nearest-neighbour gathers that faulted, partitioned by the class of fault, which is the distinction the whole fallback ladder turns on: `stalled` (the tree abandoned its own page fill), `timed_out` (a call the gather issued never answered), `exhausted` (the gather could not allocate), `abandoned` (a deadline this process owns cancelled it - **not** a caller walking away, which is deliberately excluded so one client cannot arm a backoff shared by every other), or `propagated` (the fault said something about the index's contents rather than about capacity). A sixth arm, `deterministic`, is not a cause but a **verdict on the episode**: a capacity-shaped fault that has recurred consecutively with no intervening success has stopped behaving like capacity, and is reported rather than absorbed. The first four are absorbed into the exact-scan breaker's backoff and answered with keyword recall classified `keyword.exact_fallback_suppressed`; `propagated` and `deterministic` are **not** absorbed and surface as `keyword.index_degraded`. **Do not read a flat `propagated` count beside a climbing absorbed one as load.** That is the reading that ran issue #2948's six-hour total retrieval outage as capacity pressure: all 25 of its gather faults were individually textbook timeouts, so the per-event classification was right about each one and the aggregate was still wrong. A deterministic defect and sustained load produce the same per-event classification; what separates them is the fault **rate**, which load cannot hold at one hundred percent. That is what `deterministic` measures, and the first faults of an episode stay on their own cause arm, so a real episode reads as a short run of (say) `timed_out` followed by a long run of `deterministic` - which names both what faulted and that it stopped being transient. This series exists because issue #2749 had to be diagnosed by counting exception type names in a container's log - the ladder absorbed `ScanPageStalledException` only, which is a *subclass* of the `TimeoutException` the deployment actually raised, so the absorbed set matched nothing that was happening and the breaker's backoff never escalated past its initial delay. All six arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from `lattice_metrics_series` and `lattice_metrics_dropped_measurements_by_family_total`. |
| `repocontext.retrieval.exact_scan.budget` | `Counter<long>` | `{evaluation}` | `outcome` | Fallback budget evaluations: `unbounded`, `corpus_unknown`, `within_budget`, `exceeded`. The last is predicted exhaustion and prevents a gather; it is not an elapsed-time timeout. ANN-served queries and breaker skips never reach this decision; explicit exact mode bypasses it. All four arms are zero-primed. |
| `repocontext.retrieval.exact_scan.duration` | `Counter<double>` | `s` | (none) | Cumulative gather wall seconds, recorded on exit even on failure or caller cancellation. Includes metadata/payload waits; excludes cache hits, cache storage and ranking. Not CPU time or CPU share: overlapping gathers can exceed one second per second, and an in-flight gather has not recorded its duration yet. Zero-primed counter, not a synthetic histogram sample. |
| `repocontext.retrieval.exact_scan.gathers` | `Counter<long>` | `{gather}` | `outcome` | Terminal gathers: `completed`, `faulted`, `cancelled` (caller cancellation only). Includes empty successful scans, excludes cache hits and suppressed scans. All three arms are zero-primed. Detailed fault classification and breaker behavior remain on `repocontext.retrieval.exact_gather.faults`. |
| `repocontext.retrieval.exact_scan.pages` | `Counter<long>` | `{page}` | (none) | Returned logical metadata pages, including empty terminal pages. Not internal shard page fills or RPCs. Recorded immediately after enumeration returns, before decode/payload hydration, so later faults retain completed-page work. A partial page that faults before returning is not counted. Zero-primed. |
| `repocontext.retrieval.exact_scan.vectors` | `Counter<long>` | `{vector}` | (none) | Metadata records in returned gather pages, including records filtered for embedding-space mismatch or missing payloads. Not ranked candidates or matches. Excludes pagination lookahead records and partial pages that never return. Recorded once per page, not once per vector; cache hits add nothing. Zero-primed. |
| `repocontext.bootstrap.pass_arm_faults` | `Counter<long>` | `{fault}` | `arm`, `kind` | Indexing-pass arms that faulted, by the arm that faulted (`retire`, `ingest-files`, `ingest-symbols`, `ingest-memory`) and the fault kind (`scan-page-stalled`, otherwise the exception type name). Read it as a diagnosis of *which stage* aborted a pass, which previously required correlating log lines by timestamp. It is deliberately **not** denominated by passes started, because a single pass can fault on more than one arm. A rising `arm="retire"` is the one value that does not mean lost work: a retirement fault defers every removal to the next pass and the additions and updates already computed are still committed, so it reports deferral rather than an aborted pass. Carries no repository id - the accompanying warning does, at a cardinality logs can afford. |
| `repocontext.bootstrap.phase_cancelled` | `Counter<long>` | `{cancellation}` | `phase` | Indexing runs cancelled part-way through a phase, tagged by the phase that was executing (`Walking`, `Reconciling`, `Applying`, `Vectorising`). Durable structural writes already committed survive a cancellation, but everything the cancelled phase had accumulated and not yet banked is lost and a re-run pays for it again, so this counts **discarded work**, not merely a stopped run. All four phase series are **zero-primed when the service is constructed**, so a flat zero is a reading - this deployment has discarded nothing - rather than the absence a counter reports before its first `Add`. A zero does **not** mean indexing is converging: a run that completes having made no progress is not a cancellation and is invisible here. Carries no repository id - the accompanying log line does, at a cardinality logs can afford. |
| `repocontext.bootstrap.phase_cancelled.discarded_time` | `Counter<long>` | `ms` | `phase` | Running total of run time thrown away by the cancellations above, tagged by the same phase. A counter rather than a histogram deliberately: the question this answers is "how much work has this deployment discarded", which is a total, and a counter is the only one of the two shapes that can be zero-primed without fabricating a sample that never happened. Zero-primed for the same four phases. Read it beside `repocontext.bootstrap.phase_cancelled` to get mean discarded time per cancellation; a rising total against a flat count is one long-running phase being abandoned repeatedly, which is the shape that starves a repository of an index indefinitely. |
| `repocontext.bootstrap.memory_marker_scan` | `Counter<long>` | `{walk}` | `outcome` | Walks of the embedded-memory-key marker range, partitioned by how the walk ended: `complete` (it exhausted the range in a single pass, having consumed no banked progress), `resumed` (it exhausted the range after consuming progress banked by an earlier pass that faulted), or `banked` (a page faulted, so it banked the pages already read and resumes from them on the next pass). The marker scan walks its range in small resumable pages so a walk that cannot finish inside one page-fill ceiling still converges within a bounded number of passes, instead of restarting from the beginning and never finishing at all (issue #2071). Whether it actually converges was previously observable only as the presence or absence of a warning in the host log, and the call site itself records why that is not enough: "the warning stopped" is a much weaker signal than "the range was exhausted", because the warning also stops when the scan is never reached at all. These arms make that three-state question readable from the scrape - all three at their primed zero means the scan was never reached, which no log grep can distinguish from a scan that ran and completed. `resumed` is split out of `complete` deliberately: folding them together would conflate "never needed to bank" with "banked and recovered", which are opposite answers about whether the resumable cursor is live, so the mechanism would be unobservable exactly when it is working. A rising `banked` with both completion arms flat is the thrash the cursor exists to prevent. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations. Named `banked` rather than `faulted` (which is what the structurally similar arm on `repocontext.ann.index.load` is called) because the fault here is swallowed and the pass continues successfully with a usable skip signal, so `faulted` would make the scrape assert something false. Deliberately a counter and not a histogram: the quantity of interest is how many walks ended each way, and priming a histogram would fabricate a zero-valued sample that reads as a real measurement. A `complete` reading is **not** evidence that bootstrap ingestion as a whole is healthy - the coverage-probe stand-down paths beside it are silent by design and are tracked separately (issue #2964). |
| `repocontext.bootstrap.coverage_probe` | `Counter<long>` | `{probe}` | `arm`, `outcome` | Bootstrap embedding-coverage resolutions, partitioned by the ingestion arm that resolved coverage (`file`, `symbol`, `sweep`) and by how the resolution ended (`conclusive`, `gate_pruned`, `probe_failed`). Each arm decides, before it does its work, whether it can trust an absence of embedding-membership keys as evidence that embeddings are genuinely missing. When the store's read-path access gate prunes that membership probe the arm cannot trust the absence, so it stands down: it skips the back-fill sweep and proceeds with reduced work. That stand-down was previously invisible on the scrape, and invisible in a way worse than an ordinary missing signal, because it is produced by a correctly functioning safety gate - the probe answered, the gate did its job, nothing on the path looks wrong at the point the signal is lost, so there is no error, no fault, and no warning an operator has any reason to expect (issue #2964). On the scrape, "did less work because there was less to do" and "did less work because it was not allowed to look" render identically. These arms separate them. Read the instrument four ways: **all nine zero** means no bootstrap coverage resolution was reached at all; **`conclusive` only** is the healthy steady state; **`gate_pruned` > 0** is a standing misconfiguration that is actionable and never clears by waiting, because the ingestor cannot read its own membership keys; and **`probe_failed` > 0 with `gate_pruned` == 0 in the same arm** means that zero proves nothing, since the gate check sits structurally below the probe-failure branch at every site, so a failing probe masks whatever the gate would have done. Four limits on the reading, each load-bearing. First, the `arm` dimension localises which consumer stood down, not which grant is missing: all three probes funnel through one membership-probe seam, so there is exactly one grant to repair regardless of how many arms report. Second, five further gate-pruning decision sites exist that are deliberately not instrumented here because they do not stand an arm down, and their absence from this instrument is not evidence the gate is not pruning there. Third, `conclusive` means "this arm resolved coverage it can trust", not "a network probe succeeded" - the file arm may be served from the coverage digest without probing at all. Fourth, the file arm has two guard returns above this seam (no embedding provider bound, and nothing to embed), so all-nine-zero is also the normal steady state of a keyword-only deployment with no embedder bound; that benign cause and a genuinely unreached seam are **not separable by this instrument alone**, and a reader must check whether an embedding provider is registered to tell them apart. All nine pairs are pre-minted at zero when the reporter is constructed, on the same path the arms are charged from, so an arm reading zero and an arm being absent are different observations. Deliberately a counter and not a histogram: the quantity of interest is how many resolutions ended each way, and priming a histogram would fabricate a zero-valued sample that reads as a real measurement. |
| `repocontext.bootstrap.coverage_verdict` | `Counter<long>` | `{verdict}` | `reason` | What an indexing pass concluded about embedding coverage, and therefore which cadence the next pass runs at, partitioned by the reason the verdict was reached: `converged` (coverage was measured and no gaps remain, so the gap scan stands down to its periodic interval), `arm_failure` (an arm threw, so the pass's measurement is not admissible; a retire- or file-arm fault also clears convergence and keeps the scan armed, whereas a symbol- or memory-arm fault, which says nothing about file coverage, is counted here but leaves the file arm's banked verdict - and so the cadence - in place), `gap_found` (coverage was measured and gaps remain, so the scan stays armed on every pass), and `probe_unmeasurable` (the pass could not establish coverage at all, because the membership probe was refused, the embed call deferred under saturation, or the gap scan was skipped). The last arm is why the instrument exists. A refused probe was previously folded into the same boolean as a measured gap (issue #3340), so an unmeasurable pass escalated the gap scan from its periodic cadence to every pass, and the escalated full-corpus sweep was itself the load that refused the next probe - a closed loop whose only external symptom was unbounded write-ahead-log growth. The fix stops an unmeasurable pass from clearing convergence, which means the escalation now stops silently, and a silent stand-down is precisely the failure mode that would make the fix indistinguishable from the defect it replaces. Read it as a partition, not a rate: sustained `probe_unmeasurable` means coverage is not being measured at all and any prior `converged` reading is stale rather than true, which is actionable and does not clear by waiting; `gap_found` falling away while `converged` rises is convergence; `arm_failure` is the loud arm and pairs with `repocontext.bootstrap.pass_arm_faults`. It does not duplicate `repocontext.bootstrap.coverage_probe` above - that instrument reports how each arm's probe ended, this one reports what the pass then did about it. All four arms are pre-minted at zero when the reporter is constructed, on the same path they are charged from, so `probe_unmeasurable` reading zero is a measured absence rather than a series that was never registered. |
| `repocontext.bootstrap.symbol_walk` | `Counter<long>` | `{walk}` | `outcome` | Passes of the symbol arm's whole-symbol-space range walk, partitioned by how the pass ended: `complete` (it closed a circuit in a single pass, having consumed no banked progress), `resumed` (it closed a circuit after consuming progress banked by an earlier pass that faulted), or `banked` (a page read faulted, so it banked the continuation token it had reached and resumes from there on the next pass). Before issue #2953 a faulted page discarded every page the pass had already read and the next pass restarted at the head, re-issuing the identical leaf reads. That re-drive is not a bystander to the stall that caused it: scan-page issued leaf reads measure at roughly 98% of cold WAL replay permit demand on a deployed container, against 1.6% for tombstone compaction and 0.16% for the WAL-GC blocked-leaf sweep, so a restarting walk regenerates exactly the load that made it fault. The cycle has no exit - the walk never closes a circuit, so the arm never banks a snapshot, so the tree's WAL cursor floor never advances and nothing is reclaimed. `resumed` is split out of `complete` for the same reason as on `memory_marker_scan`: folding them together would conflate "never needed to bank" with "banked and recovered", which are opposite answers about whether the cursor is live, and at the tree the two passes are byte-identical - so without this split the fix would be unfalsifiable in the deployment it exists to fix. A rising `banked` with both completion arms flat is the thrash the cursor exists to prevent, and is the reading that falsifies the fix. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations. Named `banked` rather than `faulted` because the fault is rethrown to the arm's own fault accounting after the pass has landed its partial work, so the bank is a durable side effect of the pass rather than its outcome. The memory arm has no counterpart arm here and is deliberately still re-driving: its walk builds the live-key set that drives the orphan sweep, so a partial walk would present as a partial live-key set and delete live vectors. |
| `repocontext.ingest.files_scanned` | `Counter<long>` | `{file}` | `repository` | Files the indexing walk discovered after filtering, cumulative across reconcile passes. Advances while a pass is still walking, and on every pass of a converged repository, because every pass re-scans the tree: scanned rising with `repocontext.ingest.files_embedded` flat is a completed no-change reconcile, not a stall (issue #3151). |
| `repocontext.ingest.files` | `Counter<long>` | `{file}` | `repository`, `outcome` | Files the reconcile planned, partitioned by `outcome` = `added`, `updated`, `removed`, `unchanged`. Every arm is primed at zero when a repository's first pass begins on the silo. |
| `repocontext.ingest.files_embedded` | `Counter<long>` | `{file}` | `repository` | Files whose vectors the file embedding arm stored - the headline "is embedding happening" series. Advances while a pass is still vectorising. |
| `repocontext.ingest.symbols_embedded` | `Counter<long>` | `{symbol}` | `repository` | Symbol passages whose vectors the symbol embedding arm stored. That arm can run long after file coverage is complete, so this rising while `files_embedded` is flat is a healthy back-fill, not a stall. |
| `repocontext.ingest.files_content_projected` | `Counter<long>` | `{file}` | `repository` | Files whose searchable content projection was written (added, updated and back-filled files). |
| `repocontext.ingest.passes` | `Counter<long>` | `{pass}` | `repository`, `outcome` | Reconcile passes, partitioned by `outcome` = `completed` (reconciled the whole tree and the job grain recorded it), `failed` (stopped on an error), `cancelled` (host shutdown or repository removal). Every pass that begins settles into exactly one arm; all are primed at zero on a repository's first pass. |
| `repocontext.ingest.pass.duration` | `Histogram<double>` | `s` | `repository`, `outcome` | Wall-clock seconds one reconcile pass took, from the runner beginning it to it settling, with the same `outcome` as `repocontext.ingest.passes`. Deliberately not primed, which would fabricate a zero-second pass. |
| `repocontext.ingest.last_completed_pass_age` | `ObservableGauge<double>` | `s` | `repository` | Seconds since the newest completed pass for each repository the silo has begun a pass for (until one completes, since the first pass began). The alertable "ingest has stalled" signal: it stays below the reconcile interval plus one pass duration on a healthy repository and climbs without bound while passes fail, hang or never finish. On a multi-silo host read the minimum across silos. |
| `lattice.repocontext.memory.restore` | `Counter<long>` | `{attempt}` | `outcome` | Durable-memory restore attempts made by the [memory archive](memory-durability.md#the-memory-archive) at startup, partitioned by what the attempt did to the memory tree: `restored` (a snapshot was imported, and the tree was then counted and found to hold at least what the snapshot carried), `partial` (an import wrote records and did not finish, so the tree holds strictly more than nothing and strictly less than the archive), `nothingtorestore` (the tree already holds memory records that no incomplete restore put there, so `auto` mode declines), `notattempted` (restore is `off`, or no snapshot exists yet), or `failed` (every candidate snapshot was refused and nothing was written). A non-zero `partial` always warrants an operator: the tree is short of the archive yet presents as a populated store, and nothing else reports it - the next `auto` restore heals it rather than declining. One attempt is made per process start, and only on a host that configures an archive directory: with `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR` unset the reporter is never registered and the instrument does not exist. All five arms are pre-minted at zero when the reporter is constructed, so a zero on `partial` beside a non-zero total is a measured absence of damage rather than an absent measurement. An attempt that throws before it reaches an outcome is logged at `Error` and is not counted. |

Subtracting the growth of `repocontext.response_tokens` from the growth of
`repocontext.reads_replaced_tokens` over the same one-hour window gives the signed net
saving `repocontext_stats` reports, so the dashboard and the tool agree when read over
that window. The counters themselves are cumulative since the process started, so their
raw difference is the lifetime net saving instead.

### Reading exact-gather cost

The Overview dashboard's **Exact KNN** panels expose work that process CPU and
memory counters cannot attribute. Compare `rate(repocontext_retrieval_exact_scan_vectors_total[5m])`
and `rate(repocontext_retrieval_exact_scan_pages_total[5m])` with ANN build
progress over the same interval. They count returned logical pages and visited
metadata, not internal shard work, bytes read, or CPU consumption. A failed
partial page is deliberately not estimated.

`rate(repocontext_retrieval_exact_scan_duration_seconds_total[5m])` measures
aggregate gather wall seconds per second, including awaits and faults. It is not
a CPU percentage or proof that exact work caused ANN contention. Gathers publish
their duration on exit; overlapping gathers can exceed one second per second.
For mean gather time, divide that rate by the summed rate of
`repocontext_retrieval_exact_scan_gathers_total` across all outcomes. These are the
series names an OpenTelemetry Prometheus exporter renders, which the dashboard
queries; the repocontext container's own `/metrics` exposition appends no unit word,
so there the duration counter is `repocontext_retrieval_exact_scan_duration_total`
while the other four keep the names above.

The `budget` counter partitions every evaluation, so `outcome="exceeded"` over
the summed budget rate is the fraction prevented by the configured prediction.
Tune the metadata tree's nominal/stall page budgets against measured pages and
vectors, not match count. An ANN answer or open-breaker skip does not evaluate
this budget, and explicit exact mode bypasses it. Cache hits do not gather.
All five counters and every outcome arm are zero-primed when the retrieval
reporter is constructed, with only the platform tenant and closed outcome tags,
never repository, space, key or query. An absent series means missing reporting
or collector drops, not zero work. Check collector health before interpreting it.

## Enabling the surface

All of these tools are read-only and are contributed to any caller whose data read-or-write permission unlocks the repository-context group; none requires `enableWrites`. The one exception to unconditional contribution is `repocontext_changed`: it walks a caller-supplied path, so it is contributed only when `AddRepoContextTools` is given a `workspaceRoot`, which the minimal registration below does not pass. Register the module as
a companion to `AddLatticeMcp`, exactly as for the rest of the surface:

```csharp verify
using Orleans.Lattice.Api.Mcp.RepoContext;
using Microsoft.Extensions.DependencyInjection;

var services = new ServiceCollection();
services.AddLatticeMcp(o => o.RequireAuthorization = true);
services.AddRepoContextTools();
```

Bind an `IEmbeddingProvider` for semantic search and bundles; without one, search
and the bundle still answer by keyword. For a ready-to-run local deployment see the
[container quickstart](container.md); the
[RepoContext MCP container sample](../../samples/RepoContextContainer/README.md)
runs the box end to end and walks the explainable-search, budgeted-bundle, reuse,
and stats tools against it.
