Retrieval and token economics
This page documents Orleans.Lattice.Api.Mcp.RepoContext, which is unreleased, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at retrieval-economics.md, and llms.txt lists every page.The repository-context surface is only useful to an agent if the context it returns is worth the tokens it costs. This topic covers the retrieval and token-economics capabilities layered on top of the record model: explainable search, the graph-navigation tools, the budgeted context bundle, its reuse economics, usage accounting, and the shared token counter they all budget in. Every tool here is read-only and clears the same fail-closed authorization gate as the rest of the surface.
The shared token counter
A single byte-pair-encoding (BPE) token counter underpins every figure on this
page, so a token cost reported by one tool means the same thing everywhere. It is
constructed once from a tokenizer profile and reused by the reconcile path
(the per-file TokenCount on each file node) and the retrieval surface (bundle
budgets and outline costs).
The profile is selected by the LATTICE_REPOCONTEXT_TOKENIZER environment
variable: o200k (the default, matching current-generation models) or cl100k.
Because the same counter is used to compute the stored per-file counts and to pack
a bundle, a bundle's totalTokens is an exact sum over the same encoding the
consuming model uses, not an estimate. The bundle's responseTokens builds on that
exact content sum but is itself a deliberately conservative estimate, because it
also accounts for JSON envelope and the SDK's dual emission (see
The budgeted context bundle).
Explainable search
repocontext_search ranks records against a natural-language query and returns
each hit hydrated from the store. Beyond the mode field (semantic, keyword,
or empty) that reports which path answered, every hit carries a machine-readable
reasons list: server-derived, deterministic, ordinal-ordered, bounded, and
never null.
- A semantic hit lists
semantic, the matched chunk kind (chunk:symbol,chunk:file, orchunk:memory), andsymbol:<fqName>when the match was a symbol vector ortopic:<topic>when it was a memory-entry vector. - A keyword hit lists whichever projected fields the query terms actually hit,
in a fixed high-signal-first order:
path-name-match,symbol:<fqName>,tag:<tag>,topic-match,content-match, andkey-match.
The reasons let an agent (or a human reviewing a trace) understand why a result ranked where it did, rather than treating the ranking as opaque - and they let a caller decide which hits are worth a full read before spending the tokens.
Graph navigation
Three read-only tools let an agent navigate the code graph without reading whole
files, each a bounded read over stored records that never touches the workspace on
disk (except repocontext_changed, which walks the workspace only through the
fail-closed boundary):
repocontext_outlinereturns a file's declared-symbol skeleton - each symbol's kind, signature, and 1-based line span, ordered by position - plus the token cost of reading the whole file. It is the cheapest way to grasp a file's shape and decide whether a full read is worth the tokens.repocontext_relatedresolves a file's structural neighbourhood: the type-names it references (outbound imports), the indexed symbols that reference its declarations (inbound dependents, resolved to their declaring files), and the test types that cover it. Dependents and tests come from the reverse cross-reference projection, so the lookup is bounded rather than a whole-repository scan.repocontext_changedreports how the current workspace has drifted from the index - files added, updated, and removed - by comparing content digests without invoking git, and lists the indexed files that depend on the changed ones (the reverse-reference impact set), so an agent sees the blast radius of a set of edits before re-indexing. The walk is rooted at the repository's indexed root and reuses the filters it was ingested with, so the report always compares the same path space the index was built in; the supplied path is a scope, so a directory inside the repository restricts the report to that subtree, and a path outside the indexed root is refused rather than compared. Unchanged files are settled by a stat against the stored size and ingest anchor instead of being re-read, the same fast path the periodic reconcile uses, so a whole-repository drift report stays cheap on a large tree.
The budgeted context bundle
repocontext_context is the headline capability: it collapses the
search -> recall -> read loop into a single round trip that can never overrun the
context budget. Given a natural-language task and a token budget, it searches the
store, resolves the top hits to unique files, and packs each file at a detail
level under a hard token ceiling:
paths- the path only.outline- the declared-symbol skeleton, reusing the outline projection.slices- bounded body text: the file's content projection, which holds at most its first 65,536 characters.auto(the default) - the richest level that still yields a non-empty bundle, with the concrete level reported back indetail.
Every entry carries its match reasons, its exact BPE tokenCount, and the
whole-file fullReadTokenCount. The bundle reports two figures, and it is the
first that the ceiling bounds:
responseTokens- the estimated cost of the response as the caller receives it: the delivered content plus each entry's JSON envelope (path, reasons, content hash, per-unit receipts), multiplied by the MCP SDK's dual-emission factor, because every tool result is serialized twice - once as structured content and once as text. The factor is a conservative 3.5 rather than 2, because the text copy is escaped JSON, which tokenizes worse than the structured copy. This never exceedsbudgetTokens; an empty bundle reports0.totalTokens- the narrower exact BPE sum of the packed source text alone. Useful as "how much source did I get", but it is not what the budget bounds: charging content alone once let a bundle reporting a few thousand tokens land as a response many times that size (issue #1811).
The estimate is deliberately conservative, so a bundle may come in slightly under
the ceiling but never over it. When even the cheapest entry does not fit,
the tool fails closed: entries is empty and retryBudgetTokens reports a
budget guaranteed to admit at least one entry on a retry (null when the search
matched nothing, so no larger budget would help). A truncated flag marks a
bundle that had to drop lower-ranked candidates. The top, responseBudgetTokens,
and detail arguments are validated and clamped, never trusted to drive unbounded
work: top to at most 50 files (default 10) and responseBudgetTokens to at most
200000 tokens (default 8192), with 0 or less meaning the default, and an
unrecognised detail to auto.
Reuse economics
The bundle never makes an agent pay twice for context it already holds. Each
delivered unit - a path pointer, a body span, or an outline symbol - carries a
stable opaque receipt, and each entry carries a per-version contentHash. A
unit is a descriptor, not a copy of the text: the delivered text lives once,
on its entry's content, and the units correspond one-to-one, in order, to that
content's newline-separated segments. (Carrying the text on the units too would
put every byte of source on the wire twice within a single payload, and four
times across the emitted pair - see issue #1811.) A caller feeds prior knowledge
back in three ways:
- Hand receipts back in
seento suppress exactly those units; the rest of the file still arrives. - Assert whole-file possession in
knownaspath@hash. - Pass a
sessionid to persist this bookkeeping across calls: the session auto-suppresses units it already delivered and validatesknownclaims, so a multi-call conversation converges on delivering each unit once.
The load-bearing guard is that a whole-file claim is honoured only for a
version that was actually delivered as a complete body. The session store records
possession only for slices (whole-body) deliveries; a known claim is validated
only against recorded possession. So partial evidence (an outline or a path) can
never be promoted to whole-file possession, and without a session a known claim
can never validate (fail closed). Suppressed content is acknowledged in reused
and is never charged against top or the token budget - a fully-reused file
does not consume a result slot, so the freed budget backfills lower-ranked
candidates.
The per-session bookkeeping lives on the repo-context-session tree as a
grow-only CRDT with a finite time-to-live; see
record-model.md for the storage
model.
Usage accounting
repocontext_stats reports whether the surface actually reduces context cost. Over
a bounded recent window (the last hour, held in one-minute buckets, so windowSeconds
reads 3600) it returns only summed token figures:
calls- how many context calls were answered.responseTokens- the response tokens they spent, charged at each bundle's ownresponseTokens: the conservative wire-cost estimate described above, not a BPE count.readsReplacedTokens- the whole-file read tokens they conservatively replaced, credited only for delivered whole-file-equivalent content (slicesdetail), never for discovery, partial detail, or content the caller already held.netSavedTokens- the net tokens saved,readsReplacedTokens - responseTokens(a signed figure; see below).windowSeconds- the length of the reporting window.
Crediting is deliberately conservative: reused or suppressed content is structurally
excluded, and only slices deliveries earn read-replacement credit. It is not a strict
floor, though: each slices entry is credited at the file's stored whole-file token
count, so a file longer than the 65,536-character content projection a slices body
carries is credited in full for the prefix it delivered. Because crediting is this
conservative, netSavedTokens is
signed and routinely negative for discovery-heavy or reuse-light usage - that is
correct, not a defect. It turns positive as a task delivers real bodies (slices) and
reuses a session so repeated context is suppressed and never re-charged, and it is
deliberately not clamped at zero so the surface can honestly report when it is not yet
paying for itself. The charge side has one exception: an empty bundle - one that failed
closed, or whose search matched nothing - reports responseTokens 0 although its
scaffolding still ships, so it is counted as a call that spent nothing. The figures are recorded per answered context call on a
bounded in-memory window and are also emitted as
System.Diagnostics.Metrics counters carrying a low-cardinality command
tag, so a host already scraping OpenTelemetry sees them flow through the existing
telemetry surface with no bespoke plumbing.
The tool carries no body, query, path, or repository identity - aggregate figures
only.
Emitted instruments
Every instrument this package publishes is listed below, all on the
Orleans.Lattice.Api.Mcp.RepoContext meter, so one scraper subscription covers the
whole surface. Each carries only low-cardinality tags - never a path, query, or any
body text. Every instrument also carries the repository-wide derived tenant tag
(LatticeTenantLabel.TagTenant), which the Tags column leaves implicit: it is fixed to the
platform sentinel _platform_ on every instrument except repocontext.vectorplane.rederive,
which derives it from its tree tag and so reports the default tenant for these bare
tree names. Twelve instruments carry a repository id, deliberately. Four are on the
approximate-index build plane, repocontext.ann.build.slice and
repocontext.ann.build.step.in_flight_seconds, repocontext.ann.vectors, and
repocontext.ann.partitions - the exception argued in the
repocontext.ann.build.slice row below: the approximate-index build plane is one
coordinator per onboarded repository and embedding space, an operator-chosen set of
single or low double digits, and issue #2855 established that without it fifteen
failing repositories and one succeeding repository were the same series. The other
eight are the repocontext.ingest.* family (issue #3151), on the same precedent: the
tag is bounded by the repositories onboarded on the host, and it is what makes a
stalled ingest attributable, because a silo-wide reading would let a healthy
repository hide a wedged one. One
instrument, lattice.repocontext.memory.restore, is named outside the
repocontext. prefix the others share, so a selector written against that prefix
does not match it.
| Instrument | Kind | Unit | Tags | What it records |
|---|---|---|---|---|
repocontext.calls |
Counter<long> |
{call} |
command |
Answered repocontext calls, tagged by tool name. Only the context bundle records usage, so command reads repocontext_context on every series these three counters carry. |
repocontext.response_tokens |
Counter<long> |
{token} |
command |
The response tokens those calls spent, charged at each bundle's conservative responseTokens wire-cost estimate rather than a BPE count. |
repocontext.reads_replaced_tokens |
Counter<long> |
{token} |
command |
The whole-file read tokens they conservatively replaced. Credited only for slices-detail deliveries, at each file's stored whole-file token count, so it is a floor except for a file longer than the 65,536-character content projection a slices body carries, which is credited in full for the prefix it delivered. |
repocontext.retrieval.ready_seconds |
Histogram<double> |
s |
phase |
Seconds from host start to the retrieval plane first reporting ready, tagged by the phase it reached. Recorded once per process, so it is the cold-start time-to-retrieval-ready figure. |
repocontext.retrieval.unavailable |
Counter<long> |
{event} |
cause |
Observed vector-plane fault episodes that made semantic retrieval unavailable, tagged by cause: keyword.vector_plane_unavailable, keyword.index_degraded, keyword.exact_fallback_suppressed, probe (a readiness probe rather than a query observed it), saturated (the vector plane's open was refused by an admission gate past its declared bound, so the plane is not expected to arm at the present capacity - issue #3286), or unknown. A non-zero rate is what distinguishes a keyword answer caused by a real capability loss from an intended keyword-only deployment. All six arms are pre-minted at zero when the readiness state is constructed, so each exists from process start and 'retrieval has never degraded on this process' is a measured absence rather than an absent measurement; before this the counter published no series at all until the first fault, so the healthy reading and an unwired instrument were identical. If an arm is absent rather than zero, read lattice_metrics_series against the collector ceiling and lattice_metrics_dropped_measurements_by_family_total before concluding anything, because a series whose first occurrence falls after a ceiling is reached is refused at creation. |
repocontext.retrieval.ann.search |
Counter<long> |
{query} |
state |
Semantic searches partitioned by the approximate-plane state that answered them: bootstrapping (the plane could not answer, so the fallback ladder ran), exhaustive (answered by scanning the vectors it holds), or approximate (answered from its trained partitioning). Because every answered query is counted, the total is a denominator: approximate pinned at zero beside a rising total is a measured absence of trained serving, not an absent measurement. All three arms are pre-minted at zero when the reporter is constructed, so each exists from process start; if an arm is absent rather than zero, read lattice_metrics_series against the collector ceiling and lattice_metrics_dropped_measurements_by_family_total before concluding anything, because a series whose first occurrence falls after a ceiling is reached is refused at creation. |
repocontext.ann.sweep |
Counter<long> |
{sweep} |
outcome, cause |
Approximate-index build sweeps partitioned by outcome: armed (armed at least one build coordinator), empty (completed without arming anything, either because it observed no repository in the store listing or because every repository it observed declined arming - this arm reports what the sweep observed, never that the store is empty), or faulted (threw, so nothing is scheduled until a sweep gets through). Denominate by the total across all three arms, never by the configured sweep interval: a completing sweep advances at that interval, but faulted advances on the far faster retry backoff (250 ms doubling to a 30-second ceiling). All three arms are pre-minted at zero when the reporter is constructed, so all three reading zero means no sweep has completed yet. That does not localise the fault and in particular does not establish that the sweep loop is not running, which the service's startup log line reports directly; an arm that is absent rather than zero points at collector saturation, read from lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total. The faulted arm alone carries a second tag, cause, resolved where the fault is raised: authority-unavailable (resolving the run credential threw, so nothing was attempted), listing-unavailable (the repository listing threw), plane-rejected (a coordinator was reached and refused the arming call, which will not clear on retry), dependency-unavailable (a coordinator could not be reached; a grain-call timeout is a busy coordinator and lands on repocontext.ann.sweep.arming as deferred instead), or unexpected (unclassified, and the only cause that should page). The causes are deliberately not pre-minted, so read them only once faulted is non-zero. |
repocontext.ann.sweep.arming |
Counter<long> |
{repository} |
result |
Arming calls the sweep made, partitioned by what the coordinator answered: armed (the coordinator accepted, so a build is scheduled), deferred (the call timed out because the coordinator is non-reentrant and already inside a long build turn, which is the expected answer from a healthy coordinator mid-build and deliberately not a fault), or faulted (the call threw something other than a timeout). Its denominator is repository visits, where repocontext.ann.sweep counts sweeps, and one sweep visits every repository in the listing - so the two are different populations, neither decomposes the other, and a ratio between them means nothing. This series exists because the sweep outcome could not carry the deferral at all: a sweep arming one repository and deferring nine reported armed exactly as one that armed ten, and a sweep on which every coordinator deferred fell through to empty, indistinguishable from a sweep over a store holding no repositories - two states at opposite extremes, one observation. A rising deferred beside armed at zero is a wholly wedged build plane; a result partition totalling zero beside a completed sweep is a genuinely empty listing. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total. |
repocontext.vectorplane.rederive |
Counter<long> |
{event} |
tree, outcome |
Rebuildable vector-plane tree fall-off observations and re-derivations: observed (an allowlisted fall-off was seen and a reset triggered), completed, failed (a transient fault stands and a later pass retries it), denied (the access gate refused the reset, which is deterministic and will not clear on retry), suppressed (a fall-off was seen inside the post-failure backoff window, so no reset was attempted), or refused (the tree is not a rebuildable derived tree, so re-derivation is declined fail-closed and the fault propagates). The tree tag is one of the fixed vector-tree names, never a repository id; only the vector-metadata and vector-membership paths are guarded, so in practice it is one of those two, and refused is a defence-in-depth arm no current call site reaches. |
repocontext.ann.build.corpus |
Counter<long> |
{build} |
coverage |
Approximate-index builds that reached Ready, partitioned by how much of the repository's vector prefix the read-path access gate admitted: nonempty (the build holds vectors, so the corpus read plainly succeeded and no probe was taken), unrestricted (it holds nothing and the whole prefix is admitted, so the repository genuinely has no vectors), filtered (it holds nothing and the gate narrowed the prefix, so an unknown subset was withheld), denied (it holds nothing because the gate refused the prefix outright, so the read never happened and the index state is unknown rather than empty), or unknown (it holds nothing and the coverage probe could not answer). A denied range read returns a clean, successful, empty result rather than throwing, so without this partition an authorization failure and an empty repository are the same observation. Because every completed build is counted, the total is a denominator: denied pinned at zero beside a rising total is a measured absence of denial, not an absent measurement. All five series are pre-minted, so denied is present and reads 0 on a healthy host rather than being missing. |
repocontext.ann.build.denial_terminal |
Counter<long> |
{coordinator} |
none | Build coordinators that observed enough consecutive uninterpretable corpus reads to conclude the host is refusing them, and keep retrying on a lengthening backoff that settles at one attempt roughly every five minutes, instead of on the two-second phase cadence. Counted once per denial episode, so it separates "a denial happened and is being retried" from "this deployment is permanently refused and the approximate plane will never build" - which a monotonically rising repocontext.ann.build.corpus{coverage="denied"} cannot do on its own. Any non-zero value warrants an operator: an index that should exist does not, and it will not appear by itself. |
repocontext.ann.build.slice |
Counter<long> |
{step} |
repository, space, phase, progress, cause |
Build steps taken by an approximate-index coordinator, partitioned by what the step achieved, compared against the same coordinator's previous reading: advanced (it banked vectors or resolved partitions, so the build got closer to serving), churned (it moved the build phase while banking no vector and resolving no partition, so the step did something and got the build nowhere), starved (its slice hit the ingest deadline having consumed nothing, so the source could not be read fast enough to bank a single vector), idle (the step completed and changed nothing measurable, which is what a coordinator stepping over a corpus it cannot consume looks like), or faulted (the step threw rather than completing, so the store-of-record read could not be served at all - a different condition from a read that was served and returned nothing, and one whose remedy is the projection or the store behind it rather than the access gate or the embedding throughput). This is the only series on the approximate-build plane that fires before a build reaches Ready. Every other one is terminal - repocontext.ann.build.corpus at Ready, repocontext.ann.build.denial_terminal at a terminal denial, repocontext.ann.partitioning on a plane that has finished building, repocontext.ann.sweep at arming and never again - so without it the whole interval between sweep{outcome="armed"} and Ready emits nothing, and a coordinator grinding through slices that bank nothing is byte-identical in telemetry to one that never took a step. Both present as every arm of every counter at its primed zero. Read it beside repocontext.ann.sweep{outcome="armed"}: a total pinned at zero beside a non-zero armed means the coordinator is not stepping, which is a scheduling fault; a rising starved means a source read that cannot complete inside the slice; a churned arm rising without bound beside a flat advanced means it is stepping and moving between phases while banking nothing, which is the #2791 Training -> Persisting -> Training livelock and which read as steadily rising advanced until #2818 split the arm (a healthy build churns at most once per phase transition, so its churn arm is bounded by the phase count); a rising idle means it is stepping over a corpus it cannot consume; and a rising faulted means it is stepping and throwing, which is the reading that would otherwise be indistinguishable from not stepping at all, because a step that throws never reaches the record taken after it completes. The faulted arm alone carries a second tag, cause, naming the class of fault the tick raised: scan-page-stalled, projection-stale, dependency-unavailable, response-timeout, plane-rejected, saturated, or unexpected. response-timeout means only that a response deadline expired: a keepalive call can queue behind a busy build, so a timeout proves neither busyness nor unreachability. Check progress and logs before declaring a stall. dependency-unavailable remains the transport or silo/message-rejection arm. It exists because the run-12 build faulted on every tick and localising it needed an 8 MB container log read by hand, which established that 39 of 39 faults were a single condition - a grain-call timeout reading the vector-index tree - that the counter alone could not name. Causes are classified across the whole tick, not only the build step: five further call sites on a tick can throw, three of them before the step is counted at all, and a fault at any of those previously left the tick silent in this series entirely. The twenty-two (phase, progress) arms of a plane are pre-minted at zero the first time that plane is armed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total. The cause values are deliberately not pre-minted: priming them would mint seven faulted-arm series on a host that has never faulted, so anything counting series rather than values would read a healthy host as a faulting one. The consequence is that a cause reading zero is uninterpretable rather than innocent - it says only that no fault of that class has been recorded since process start, which is equally what a healthy host and a mis-wired classifier look like. Read a cause only once faulted itself is non-zero, and denominate the causes against it: they sum to it exactly. Three further dimensions - repository, space, and phase - name where a step was taken, and were added by issue #2855 because without them the series could not answer the question the epic's acceptance run turns on. phase is the load-bearing one: it names the stage of the build the step was executing when it was counted, and on the faulted arm it is read at the fault site from the index's own progress rather than snapshotted on entry, so it is exact rather than approximate. That precision is not incidental. A build step entered in training trains the index and then persists the trained partitioning within the same step, so an entry snapshot would file a persist fault under training and destroy the one distinction the tag exists to draw: a fault reading ingesting is a corpus that could not be read, which is an independent defect, while a fault reading persisting is a trained index that could not be written into the vector-index tree - and if that tree is itself the subject of an open fault, the persist fault is a downstream symptom of it rather than a second defect, so scoring the two separately would count one defect twice. The values are coordinating, opening, ingesting, training, persisting, and reconciling; the first two precede any build step and are therefore reachable on the faulted arm alone, which is why a plane primes twenty-two arms rather than thirty - the eight combinations a completing step can never reach are deliberately not minted. repository and space bound the plane the step belongs to. Their cardinality is the product of onboarded repositories and embedding spaces, which is exactly the number of durable indexes and build coordinators the host already runs - one apiece - so it is operator-chosen and small, and is not the per-grain cardinality class of issue #2518. space reads as {model-id}/{dimension}, or unspecified on a plane whose space is not yet resolved. A plane re-derived onto a new embedding model is a different build over a different corpus, so merging the two would hide a migration mid-flight. Priming is per plane and happens when a plane is first armed rather than when the process starts, because a process cannot know which planes exist and a primed series for a plane nobody armed claims a build nobody asked for; the consequence is that a host with no armed plane emits no series for this instrument at all, which is the correct reading of a host that has never built an index. |
repocontext.ann.build.slice.items |
Counter<long> |
{item} |
none | Source items consumed by the approximate-index build slices that completed, summed across slices; a slice that throws records nothing. It exists to make vectors per slice obtainable from metrics alone, and that is a narrower gap than it sounds: the obvious way to compute it - dividing ann.vectorsIndexed from a health payload by repocontext.ann.build.slice{progress="advanced"} - is invalid, because the first is cumulative and is inherited across a restart whenever the index was restored from durable state, while the second is process-scoped and resets at the deploy boundary. The ratio of two counters with different epochs is not a rate of anything. With this arm the quantity is a delta of two series that share an epoch. Read it against repocontext.ann.build.stage.duration below: a slice consuming far fewer items than IngestBatchSize allows, while the stage split shows the time in key_assign or source_wait, is a per-item round trip rather than a slice that ran out of budget. A slice that consumed nothing records nothing here, so this counter does not double as a slice count - repocontext.ann.build.slice is that. |
repocontext.ann.build.stage.duration |
Histogram<double> |
s |
stage |
Seconds one stage of a build slice took, accumulated across the items of that slice and tagged by stage: source_wait (awaiting the source enumerator, which on this host streams over grain calls), key_assign (mapping identifiers to index keys, including a durable block reservation when one falls due), index_upsert (the in-memory insert), and key_flush (the single batched write that makes the slice's key-map records durable). Separating them is the whole point, and the retrieval path has had the same split for the same reason: an end-to-end build rate cannot distinguish a slow source from slow key assignment from a slow index, and those are three different owners with three different fixes. Before this existed the distinction had to be made by reading the source, which is not a thing an operator can do against a running container. The four do not cover the whole slice: the source count taken while the expected count is still unknown, the ingest checkpoint that then persists the slice's vector chunks and build state, and the loop's own bookkeeping belong to no stage, so the four sum to less than the slice's elapsed time. A slice records all four only once it has checkpointed; a slice that throws records nothing here. Deliberately not zero-primed: priming a histogram fabricates a zero-valued sample that reads as a real measurement of an instantaneous stage and destroys the distribution the instrument exists to report, so an absent stage here means no slice has completed rather than a stage that took no time. The timings arrive through the vector package's public IVectorIndexBuildObserver seam, which the host binds as each index's DurableVectorIndexOptions.BuildObserver (see the vector index configuration); sampling is skipped entirely when no observer is bound, so a build with no reporter wired makes no extra clock reads - which matters because the slice budget is measured against that same clock. |
repocontext.ann.build.step.in_flight_seconds |
ObservableGauge<double> |
s |
repository, space, phase |
How long the approximate-index build step currently executing has been inside its present phase, or 0 when no step is executing. Every other instrument on this plane - repocontext.ann.build.slice included - is recorded at a terminal moment: after the step returns, after the open finishes or throws, or once the build reaches Ready. A phase that never terminates therefore emits nothing on any of them, and the resulting all-zero reading is byte-identical to a coordinator that is not stepping at all. Run 14 of epic #2368 hit exactly that: a build sat inside a single non-reentrant coordinator turn for over four minutes, and the only evidence available anywhere was Orleans' own generic "request has been active for 00:04:00" warning in a container log, while six dedicated arms on this plane all read zero. An observable gauge is the only instrument shape that can close that gap, because the collector drives it rather than the step reaching an end it may never reach - a counter or histogram is written by the code path being measured, so a path that does not complete suppresses its own evidence (issue #3130). The clock restarts on every phase entry rather than running for the whole step, because the question is not "has this step been slow" but which half of it is not returning: an ingesting arm climbing without bound is a corpus-read defect, a persisting one is an index-write defect, and the two have different owners and different remedies. Every primed plane emits every one of the six phases, so a healthy idle build reads as six zeros rather than as an absence - a gauge that emitted only while a step was in flight would make health and a mis-wired instrument the same observation, which is the conflation issue #2952 removed from the slice counter. Steps are tracked by token rather than by plane, so an outstanding step that never returns cannot be erased by the next step that does. Priming is per plane and happens when a plane is first seen, so a host with no armed plane emits no series for this instrument at all. |
repocontext.ann.partitioning |
Counter<long> |
{observation} |
state |
Approximate-index planes observed at each maintenance turn, partitioned by whether the plane holds a partitioning and, when it does not, by why: partitioned (it answers from a trained partitioning), unpartitioned-small (it holds none and its corpus is below MinimumTrainingCount, which is the correct state at that size), or unpartitioned-large (it holds none although its corpus is at or above the minimum). The third arm is the entire reason the partition exists: before issue #2706 those two cases were the same observation, so a plane holding 7.5x the threshold across zero partitions was indistinguishable from one that was simply too small to train, and every semantic query was answered by brute-force scan with nothing reporting it. A sustained non-zero unpartitioned-large means exactly that, and warrants an operator. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total, not a statement about the plane. |
repocontext.ann.repartition |
Counter<long> |
{training} |
outcome |
Training passes taken because a plane's corpus crossed the training minimum (1,024 vectors) after an earlier training had declined to partition it, by outcome: partitioned (the pass produced a partitioning, so the plane now serves approximate) or declined (the corpus met the minimum yet still resolved to fewer than two partitions, so the plane stays exhaustive and exact and the next attempt waits until the corpus has doubled). This series reports the repair, where repocontext.ann.partitioning reports the state, so it is expected to read zero forever on a deployment that partitioned on its first build; a single partitioned is one latched plane healing itself. A rising declined means the corpus keeps crossing the training minimum while the partition count still will not resolve, which further corpus growth inside one activation will not change quickly. Both arms are pre-minted at zero. |
repocontext.ann.index.load |
Counter<long> |
{attempt} |
outcome, reason |
One arm per load attempt: fresh, resumed, faulted, deferred, refused, or discarded. discarded means an open successfully removed unverifiable derived state and will rebuild from source; it is not also fresh. Reasons are bounded: discarded has count_mismatch, embedding_space_change, or unloadable_record; faulted has timeout, embedding_space_change, unloadable_record, or other; refused has admission_refused. Fresh/resumed/deferred carry none. No exception text or keys enter labels. Every legal pair is pre-minted at zero. Deferred means the open budget yielded with progress retained, or the restore found a record missing from one store read but returned by another and kept the durable index to retry (issue #3905); refused means the replay admission gate rejected work, not a load fault. Resumed includes progress retained after either yield or a fault. A valid manifest with an older build-state generation is adopted by existing recovery, not discarded. Warnings identify discard/fault reasons, and a discard of committed state names the generation, vector count and partition count it destroyed; refusal logging retains the admission distinction. Label migration: consumers pinned to the old exact label set must aggregate away reason, for example sum by (outcome, tenant) (repocontext_ann_index_load_total). |
repocontext.ann.vectors |
ObservableGauge<long> |
{vector} |
repository, space, count = held, expected |
Last published VectorIndexBuildProgress.VectorsIndexed and VectorsExpected for each registry handle. Held is the resident indexed count, not the source size or a durability guarantee; expected is the last source count observed by the build, not a live census. Labels match the build-slice plane. Open/restored and failed-step progress are published too; unopened handles read zero, no handles emit no series, and disposing the registry retires its series. |
repocontext.ann.partitions |
ObservableGauge<long> |
{partition} |
repository, space |
VectorIndexBuildProgress.PartitionsTotal from the same handle snapshot; zero means no trained partitioning, including a small completed exhaustive index. It is not the number of persisted partitions. The same lifetime and label rules as repocontext.ann.vectors apply. |
repocontext.retrieval.duration |
Histogram<double> |
s |
tool, path |
End-to-end seconds for one retrieval tool call, tagged by the tool (search, context, outline, related) and by the retrieval path that answered it (semantic.exact, semantic.approximate, one of the keyword.* causes, not_applicable for a graph read that consults no vector plane, or unresolved for a call that ended - cancelled or faulted - before a path was settled). The path tag is what makes the figure interpretable rather than merely present: a fast keyword answer and a fast approximate answer mean opposite things about the health of the box. Recorded from a finally, once per call, on every termination including cancellation and failure, so the count is a true call total and is the denominator for the stage series below; the only measurement it can lose is one whose process died mid-call. A context call records exactly one row under tool="context" and none under tool="search" even though it runs a search internally, so the two tools' latencies never contaminate each other. |
repocontext.retrieval.stage.duration |
Histogram<double> |
s |
stage, path |
Seconds spent inside one stage of a retrieval, tagged by stage - embed (the network hop to the embedding service, covering both its availability probe and the query embed), vector_search (the index scan), hydrate (reading each matched identity back from the store of record, which the index never returns a second copy of), keyword_scan (the BM25 fallback) - and by the same path value the enclosing call resolved to. Separating them is the whole point: an end-to-end figure alone cannot distinguish a slow embedder from a slow index from a slow store, and those are three different owners with three different fixes. A stage is recorded if and only if it ran, including when it ran and then failed, so this series is deliberately sparse and its zero does not describe itself. Denominate it with repocontext.retrieval.duration above: no embed beside a rising call total is a measured absence of the semantic path - an intended keyword-only host - whereas both reading zero means no retrieval ran at all. |
repocontext.retrieval.exact_gather.faults |
Counter<long> |
{fault} |
fault |
Exact k-nearest-neighbour gathers that faulted, partitioned by the class of fault, which is the distinction the whole fallback ladder turns on: stalled (the tree abandoned its own page fill), timed_out (a call the gather issued never answered), exhausted (the gather could not allocate), abandoned (a deadline this process owns cancelled it - not a caller walking away, which is deliberately excluded so one client cannot arm a backoff shared by every other), or propagated (the fault said something about the index's contents rather than about capacity). A sixth arm, deterministic, is not a cause but a verdict on the episode: a capacity-shaped fault that has recurred consecutively with no intervening success has stopped behaving like capacity, and is reported rather than absorbed. The first four are absorbed into the exact-scan breaker's backoff and answered with keyword recall classified keyword.exact_fallback_suppressed; propagated and deterministic are not absorbed and surface as keyword.index_degraded. Do not read a flat propagated count beside a climbing absorbed one as load. That is the reading that ran issue #2948's six-hour total retrieval outage as capacity pressure: all 25 of its gather faults were individually textbook timeouts, so the per-event classification was right about each one and the aggregate was still wrong. A deterministic defect and sustained load produce the same per-event classification; what separates them is the fault rate, which load cannot hold at one hundred percent. That is what deterministic measures, and the first faults of an episode stay on their own cause arm, so a real episode reads as a short run of (say) timed_out followed by a long run of deterministic - which names both what faulted and that it stopped being transient. This series exists because issue #2749 had to be diagnosed by counting exception type names in a container's log - the ladder absorbed ScanPageStalledException only, which is a subclass of the TimeoutException the deployment actually raised, so the absorbed set matched nothing that was happening and the breaker's backoff never escalated past its initial delay. All six arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations; an absent arm is a collector fault, read from lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total. |
repocontext.retrieval.exact_scan.budget |
Counter<long> |
{evaluation} |
outcome |
Fallback budget evaluations: unbounded, corpus_unknown, within_budget, exceeded. The last is predicted exhaustion and prevents a gather; it is not an elapsed-time timeout. ANN-served queries and breaker skips never reach this decision; explicit exact mode bypasses it. All four arms are zero-primed. |
repocontext.retrieval.exact_scan.duration |
Counter<double> |
s |
(none) | Cumulative gather wall seconds, recorded on exit even on failure or caller cancellation. Includes metadata/payload waits; excludes cache hits, cache storage and ranking. Not CPU time or CPU share: overlapping gathers can exceed one second per second, and an in-flight gather has not recorded its duration yet. Zero-primed counter, not a synthetic histogram sample. |
repocontext.retrieval.exact_scan.gathers |
Counter<long> |
{gather} |
outcome |
Terminal gathers: completed, faulted, cancelled (caller cancellation only). Includes empty successful scans, excludes cache hits and suppressed scans. All three arms are zero-primed. Detailed fault classification and breaker behavior remain on repocontext.retrieval.exact_gather.faults. |
repocontext.retrieval.exact_scan.pages |
Counter<long> |
{page} |
(none) | Returned logical metadata pages, including empty terminal pages. Not internal shard page fills or RPCs. Recorded immediately after enumeration returns, before decode/payload hydration, so later faults retain completed-page work. A partial page that faults before returning is not counted. Zero-primed. |
repocontext.retrieval.exact_scan.vectors |
Counter<long> |
{vector} |
(none) | Metadata records in returned gather pages, including records filtered for embedding-space mismatch or missing payloads. Not ranked candidates or matches. Excludes pagination lookahead records and partial pages that never return. Recorded once per page, not once per vector; cache hits add nothing. Zero-primed. |
repocontext.bootstrap.pass_arm_faults |
Counter<long> |
{fault} |
arm, kind |
Indexing-pass arms that faulted, by the arm that faulted (retire, ingest-files, ingest-symbols, ingest-memory) and the fault kind (scan-page-stalled, otherwise the exception type name). Read it as a diagnosis of which stage aborted a pass, which previously required correlating log lines by timestamp. It is deliberately not denominated by passes started, because a single pass can fault on more than one arm. A rising arm="retire" is the one value that does not mean lost work: a retirement fault defers every removal to the next pass and the additions and updates already computed are still committed, so it reports deferral rather than an aborted pass. Carries no repository id - the accompanying warning does, at a cardinality logs can afford. |
repocontext.bootstrap.phase_cancelled |
Counter<long> |
{cancellation} |
phase |
Indexing runs cancelled part-way through a phase, tagged by the phase that was executing (Walking, Reconciling, Applying, Vectorising). Durable structural writes already committed survive a cancellation, but everything the cancelled phase had accumulated and not yet banked is lost and a re-run pays for it again, so this counts discarded work, not merely a stopped run. All four phase series are zero-primed when the service is constructed, so a flat zero is a reading - this deployment has discarded nothing - rather than the absence a counter reports before its first Add. A zero does not mean indexing is converging: a run that completes having made no progress is not a cancellation and is invisible here. Carries no repository id - the accompanying log line does, at a cardinality logs can afford. |
repocontext.bootstrap.phase_cancelled.discarded_time |
Counter<long> |
ms |
phase |
Running total of run time thrown away by the cancellations above, tagged by the same phase. A counter rather than a histogram deliberately: the question this answers is "how much work has this deployment discarded", which is a total, and a counter is the only one of the two shapes that can be zero-primed without fabricating a sample that never happened. Zero-primed for the same four phases. Read it beside repocontext.bootstrap.phase_cancelled to get mean discarded time per cancellation; a rising total against a flat count is one long-running phase being abandoned repeatedly, which is the shape that starves a repository of an index indefinitely. |
repocontext.bootstrap.memory_marker_scan |
Counter<long> |
{walk} |
outcome |
Walks of the embedded-memory-key marker range, partitioned by how the walk ended: complete (it exhausted the range in a single pass, having consumed no banked progress), resumed (it exhausted the range after consuming progress banked by an earlier pass that faulted), or banked (a page faulted, so it banked the pages already read and resumes from them on the next pass). The marker scan walks its range in small resumable pages so a walk that cannot finish inside one page-fill ceiling still converges within a bounded number of passes, instead of restarting from the beginning and never finishing at all (issue #2071). Whether it actually converges was previously observable only as the presence or absence of a warning in the host log, and the call site itself records why that is not enough: "the warning stopped" is a much weaker signal than "the range was exhausted", because the warning also stops when the scan is never reached at all. These arms make that three-state question readable from the scrape - all three at their primed zero means the scan was never reached, which no log grep can distinguish from a scan that ran and completed. resumed is split out of complete deliberately: folding them together would conflate "never needed to bank" with "banked and recovered", which are opposite answers about whether the resumable cursor is live, so the mechanism would be unobservable exactly when it is working. A rising banked with both completion arms flat is the thrash the cursor exists to prevent. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations. Named banked rather than faulted (which is what the structurally similar arm on repocontext.ann.index.load is called) because the fault here is swallowed and the pass continues successfully with a usable skip signal, so faulted would make the scrape assert something false. Deliberately a counter and not a histogram: the quantity of interest is how many walks ended each way, and priming a histogram would fabricate a zero-valued sample that reads as a real measurement. A complete reading is not evidence that bootstrap ingestion as a whole is healthy - the coverage-probe stand-down paths beside it are silent by design and are tracked separately (issue #2964). |
repocontext.bootstrap.coverage_probe |
Counter<long> |
{probe} |
arm, outcome |
Bootstrap embedding-coverage resolutions, partitioned by the ingestion arm that resolved coverage (file, symbol, sweep) and by how the resolution ended (conclusive, gate_pruned, probe_failed). Each arm decides, before it does its work, whether it can trust an absence of embedding-membership keys as evidence that embeddings are genuinely missing. When the store's read-path access gate prunes that membership probe the arm cannot trust the absence, so it stands down: it skips the back-fill sweep and proceeds with reduced work. That stand-down was previously invisible on the scrape, and invisible in a way worse than an ordinary missing signal, because it is produced by a correctly functioning safety gate - the probe answered, the gate did its job, nothing on the path looks wrong at the point the signal is lost, so there is no error, no fault, and no warning an operator has any reason to expect (issue #2964). On the scrape, "did less work because there was less to do" and "did less work because it was not allowed to look" render identically. These arms separate them. Read the instrument four ways: all nine zero means no bootstrap coverage resolution was reached at all; conclusive only is the healthy steady state; gate_pruned > 0 is a standing misconfiguration that is actionable and never clears by waiting, because the ingestor cannot read its own membership keys; and probe_failed > 0 with gate_pruned == 0 in the same arm means that zero proves nothing, since the gate check sits structurally below the probe-failure branch at every site, so a failing probe masks whatever the gate would have done. Four limits on the reading, each load-bearing. First, the arm dimension localises which consumer stood down, not which grant is missing: all three probes funnel through one membership-probe seam, so there is exactly one grant to repair regardless of how many arms report. Second, five further gate-pruning decision sites exist that are deliberately not instrumented here because they do not stand an arm down, and their absence from this instrument is not evidence the gate is not pruning there. Third, conclusive means "this arm resolved coverage it can trust", not "a network probe succeeded" - the file arm may be served from the coverage digest without probing at all. Fourth, the file arm has two guard returns above this seam (no embedding provider bound, and nothing to embed), so all-nine-zero is also the normal steady state of a keyword-only deployment with no embedder bound; that benign cause and a genuinely unreached seam are not separable by this instrument alone, and a reader must check whether an embedding provider is registered to tell them apart. All nine pairs are pre-minted at zero when the reporter is constructed, on the same path the arms are charged from, so an arm reading zero and an arm being absent are different observations. Deliberately a counter and not a histogram: the quantity of interest is how many resolutions ended each way, and priming a histogram would fabricate a zero-valued sample that reads as a real measurement. |
repocontext.bootstrap.coverage_verdict |
Counter<long> |
{verdict} |
reason |
What an indexing pass concluded about embedding coverage, and therefore which cadence the next pass runs at, partitioned by the reason the verdict was reached: converged (coverage was measured and no gaps remain, so the gap scan stands down to its periodic interval), arm_failure (an arm threw, so the pass's measurement is not admissible; a retire- or file-arm fault also clears convergence and keeps the scan armed, whereas a symbol- or memory-arm fault, which says nothing about file coverage, is counted here but leaves the file arm's banked verdict - and so the cadence - in place), gap_found (coverage was measured and gaps remain, so the scan stays armed on every pass), and probe_unmeasurable (the pass could not establish coverage at all, because the membership probe was refused, the embed call deferred under saturation, or the gap scan was skipped). The last arm is why the instrument exists. A refused probe was previously folded into the same boolean as a measured gap (issue #3340), so an unmeasurable pass escalated the gap scan from its periodic cadence to every pass, and the escalated full-corpus sweep was itself the load that refused the next probe - a closed loop whose only external symptom was unbounded write-ahead-log growth. The fix stops an unmeasurable pass from clearing convergence, which means the escalation now stops silently, and a silent stand-down is precisely the failure mode that would make the fix indistinguishable from the defect it replaces. Read it as a partition, not a rate: sustained probe_unmeasurable means coverage is not being measured at all and any prior converged reading is stale rather than true, which is actionable and does not clear by waiting; gap_found falling away while converged rises is convergence; arm_failure is the loud arm and pairs with repocontext.bootstrap.pass_arm_faults. It does not duplicate repocontext.bootstrap.coverage_probe above - that instrument reports how each arm's probe ended, this one reports what the pass then did about it. All four arms are pre-minted at zero when the reporter is constructed, on the same path they are charged from, so probe_unmeasurable reading zero is a measured absence rather than a series that was never registered. |
repocontext.bootstrap.symbol_walk |
Counter<long> |
{walk} |
outcome |
Passes of the symbol arm's whole-symbol-space range walk, partitioned by how the pass ended: complete (it closed a circuit in a single pass, having consumed no banked progress), resumed (it closed a circuit after consuming progress banked by an earlier pass that faulted), or banked (a page read faulted, so it banked the continuation token it had reached and resumes from there on the next pass). Before issue #2953 a faulted page discarded every page the pass had already read and the next pass restarted at the head, re-issuing the identical leaf reads. That re-drive is not a bystander to the stall that caused it: scan-page issued leaf reads measure at roughly 98% of cold WAL replay permit demand on a deployed container, against 1.6% for tombstone compaction and 0.16% for the WAL-GC blocked-leaf sweep, so a restarting walk regenerates exactly the load that made it fault. The cycle has no exit - the walk never closes a circuit, so the arm never banks a snapshot, so the tree's WAL cursor floor never advances and nothing is reclaimed. resumed is split out of complete for the same reason as on memory_marker_scan: folding them together would conflate "never needed to bank" with "banked and recovered", which are opposite answers about whether the cursor is live, and at the tree the two passes are byte-identical - so without this split the fix would be unfalsifiable in the deployment it exists to fix. A rising banked with both completion arms flat is the thrash the cursor exists to prevent, and is the reading that falsifies the fix. All three arms are pre-minted at zero when the reporter is constructed, so an arm reading zero and an arm being absent are different observations. Named banked rather than faulted because the fault is rethrown to the arm's own fault accounting after the pass has landed its partial work, so the bank is a durable side effect of the pass rather than its outcome. The memory arm has no counterpart arm here and is deliberately still re-driving: its walk builds the live-key set that drives the orphan sweep, so a partial walk would present as a partial live-key set and delete live vectors. |
repocontext.ingest.files_scanned |
Counter<long> |
{file} |
repository |
Files the indexing walk discovered after filtering, cumulative across reconcile passes. Advances while a pass is still walking, and on every pass of a converged repository, because every pass re-scans the tree: scanned rising with repocontext.ingest.files_embedded flat is a completed no-change reconcile, not a stall (issue #3151). |
repocontext.ingest.files |
Counter<long> |
{file} |
repository, outcome |
Files the reconcile planned, partitioned by outcome = added, updated, removed, unchanged. Every arm is primed at zero when a repository's first pass begins on the silo. |
repocontext.ingest.files_embedded |
Counter<long> |
{file} |
repository |
Files whose vectors the file embedding arm stored - the headline "is embedding happening" series. Advances while a pass is still vectorising. |
repocontext.ingest.symbols_embedded |
Counter<long> |
{symbol} |
repository |
Symbol passages whose vectors the symbol embedding arm stored. That arm can run long after file coverage is complete, so this rising while files_embedded is flat is a healthy back-fill, not a stall. |
repocontext.ingest.files_content_projected |
Counter<long> |
{file} |
repository |
Files whose searchable content projection was written (added, updated and back-filled files). |
repocontext.ingest.passes |
Counter<long> |
{pass} |
repository, outcome |
Reconcile passes, partitioned by outcome = completed (reconciled the whole tree and the job grain recorded it), failed (stopped on an error), cancelled (host shutdown or repository removal). Every pass that begins settles into exactly one arm; all are primed at zero on a repository's first pass. |
repocontext.ingest.pass.duration |
Histogram<double> |
s |
repository, outcome |
Wall-clock seconds one reconcile pass took, from the runner beginning it to it settling, with the same outcome as repocontext.ingest.passes. Deliberately not primed, which would fabricate a zero-second pass. |
repocontext.ingest.last_completed_pass_age |
ObservableGauge<double> |
s |
repository |
Seconds since the newest completed pass for each repository the silo has begun a pass for (until one completes, since the first pass began). The alertable "ingest has stalled" signal: it stays below the reconcile interval plus one pass duration on a healthy repository and climbs without bound while passes fail, hang or never finish. On a multi-silo host read the minimum across silos. |
lattice.repocontext.memory.restore |
Counter<long> |
{attempt} |
outcome |
Durable-memory restore attempts made by the memory archive at startup, partitioned by what the attempt did to the memory tree: restored (a snapshot was imported, and the tree was then counted and found to hold at least what the snapshot carried), partial (an import wrote records and did not finish, so the tree holds strictly more than nothing and strictly less than the archive), nothingtorestore (the tree already holds memory records that no incomplete restore put there, so auto mode declines), notattempted (restore is off, or no snapshot exists yet), or failed (every candidate snapshot was refused and nothing was written). A non-zero partial always warrants an operator: the tree is short of the archive yet presents as a populated store, and nothing else reports it - the next auto restore heals it rather than declining. One attempt is made per process start, and only on a host that configures an archive directory: with LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR unset the reporter is never registered and the instrument does not exist. All five arms are pre-minted at zero when the reporter is constructed, so a zero on partial beside a non-zero total is a measured absence of damage rather than an absent measurement. An attempt that throws before it reaches an outcome is logged at Error and is not counted. |
Subtracting the growth of repocontext.response_tokens from the growth of
repocontext.reads_replaced_tokens over the same one-hour window gives the signed net
saving repocontext_stats reports, so the dashboard and the tool agree when read over
that window. The counters themselves are cumulative since the process started, so their
raw difference is the lifetime net saving instead.
Reading exact-gather cost
The Overview dashboard's Exact KNN panels expose work that process CPU and
memory counters cannot attribute. Compare rate(repocontext_retrieval_exact_scan_vectors_total[5m])
and rate(repocontext_retrieval_exact_scan_pages_total[5m]) with ANN build
progress over the same interval. They count returned logical pages and visited
metadata, not internal shard work, bytes read, or CPU consumption. A failed
partial page is deliberately not estimated.
rate(repocontext_retrieval_exact_scan_duration_seconds_total[5m]) measures
aggregate gather wall seconds per second, including awaits and faults. It is not
a CPU percentage or proof that exact work caused ANN contention. Gathers publish
their duration on exit; overlapping gathers can exceed one second per second.
For mean gather time, divide that rate by the summed rate of
repocontext_retrieval_exact_scan_gathers_total across all outcomes. These are the
series names an OpenTelemetry Prometheus exporter renders, which the dashboard
queries; the repocontext container's own /metrics exposition appends no unit word,
so there the duration counter is repocontext_retrieval_exact_scan_duration_total
while the other four keep the names above.
The budget counter partitions every evaluation, so outcome="exceeded" over
the summed budget rate is the fraction prevented by the configured prediction.
Tune the metadata tree's nominal/stall page budgets against measured pages and
vectors, not match count. An ANN answer or open-breaker skip does not evaluate
this budget, and explicit exact mode bypasses it. Cache hits do not gather.
All five counters and every outcome arm are zero-primed when the retrieval
reporter is constructed, with only the platform tenant and closed outcome tags,
never repository, space, key or query. An absent series means missing reporting
or collector drops, not zero work. Check collector health before interpreting it.
Enabling the surface
All of these tools are read-only and are contributed to any caller whose data read-or-write permission unlocks the repository-context group; none requires enableWrites. The one exception to unconditional contribution is repocontext_changed: it walks a caller-supplied path, so it is contributed only when AddRepoContextTools is given a workspaceRoot, which the minimal registration below does not pass. Register the module as
a companion to AddLatticeMcp, exactly as for the rest of the surface:
using Orleans.Lattice.Api.Mcp.RepoContext;
using Microsoft.Extensions.DependencyInjection;
var services = new ServiceCollection();
services.AddLatticeMcp(o => o.RequireAuthorization = true);
services.AddRepoContextTools();
Bind an IEmbeddingProvider for semantic search and bundles; without one, search
and the bundle still answer by keyword. For a ready-to-run local deployment see the
container quickstart; the
RepoContext MCP container sample
runs the box end to end and walks the explainable-search, budgeted-bundle, reuse,
and stats tools against it.