Orleans.Lattice.Api.Mcp.RepoContext meter
This page documents Orleans.Lattice.Dashboards 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at orleans-lattice-api-mcp-repocontext-meter.md, and llms.txt lists every page.
Part of Metric-to-panel coverage map.
The repository-context MCP surface's telemetry, published by the opt-in lattice.api.mcp.repocontext add-on. This meter is named unlike every other meter on this page: it carries the assembly-style name Orleans.Lattice.Api.Mcp.RepoContext rather than a dotted-lowercase orleans.lattice.* name, and all but one of its instruments carry a bare repocontext. prefix (the exception, lattice.repocontext.memory.restore, carries a lattice.repocontext. prefix). The two halves of that mismatch behave differently, and the difference is load-bearing. The container's Prometheus exposition subscribes by meter name and compares case-insensitively, so Orleans.Lattice.Api.Mcp.RepoContext does match an orleans.lattice prefix and these series are collected; the collector pins that exact meter name in a test. A selector or a doc guard written against the instrument names does not match, because those begin repocontext. or lattice.repocontext. rather than orleans.lattice.. That is why this section exists as its own table rather than as rows under an existing meter, and why a doc-coverage fixture for this package needs its own instrument-name prefix.
The Overview dashboard charts exact-KNN gather cost and budget outcomes; the remaining instruments are explicitly marked "not charted". RepoContextMetricsToPanelMapTests checks source-to-row and row-to-source coverage, reconciles each row's charted claim with actual panel expressions, and rejects unknown Prometheus tokens. Its positive control over orleans_lattice_ tokens guards against broken extraction. These checks cover the bare repocontext. prefix, not lattice.repocontext.memory.restore. The exact-KNN panels are written in the .AddPrometheusExporter() spelling described above, so the gather wall-time panel's repocontext_retrieval_exact_scan_duration_seconds_total names a series the container's own exposition never emits - it exposes that instrument as repocontext_retrieval_exact_scan_duration_total - and that panel renders no data against the container; the other exact-KNN instruments carry annotation units, so their series names are the same on both expositions. The container serves /metrics on the same listener as MCP and the health probes. Subscribe to both this meter and the core orleans.lattice meter when diagnosing contention: core leaf/WAL signals are not published on the repository-context meter.
Separately, a Histogram<T> renders on this endpoint as a Prometheus summary carrying _sum and _count and no _bucket series, so a histogram_quantile panel over it returns nothing, and the common dashboard idiom of appending or vector(0) then substitutes a literal zero that is indistinguishable from a genuine sustained-zero fault. That applies to every histogram in the table below - repocontext.retrieval.ready_seconds, repocontext.retrieval.duration, repocontext.retrieval.stage.duration, repocontext.ann.build.stage.duration and repocontext.ingest.pass.duration: chart them as _sum and _count, and treat any quantile panel against this endpoint as unavailable rather than as measured. The full account is in the container guide. Author the panels and update this column when they land.
The container host publishes a second meter of its own from apps/repocontext - its health, drain-forecast, heap, garbage-collection, grain-storage-lock and backup instruments. No bundled dashboard charts it either, and it sits outside every doc-coverage gate, because those enumerate src/ only, so the meter and its instruments are catalogued in the container guide rather than in a table here.
Every instrument here carries the derived tenant label described in The derived tenant label. All but one carry the reserved _platform_ value, because the repository-context surface is platform-owned and their series have no owning tree. The exception is repocontext.vectorplane.rederive, which derives the label from its tree tag: the repository-context trees have bare, unsegmented ids, so its series read default, the legacy-adoption tenant, rather than _platform_.
| Instrument | Type | Tags | Dashboard | Panel(s) |
|---|---|---|---|---|
repocontext.calls |
counter ({call}, per-operation) |
command, tenant |
(none) | not charted |
repocontext.response_tokens |
counter ({token}) |
command, tenant |
(none) | not charted |
repocontext.reads_replaced_tokens |
counter ({token}) |
command, tenant |
(none) | not charted |
repocontext.ann.sweep |
counter ({sweep}) |
outcome = armed, empty, faulted; cause = authority-unavailable, listing-unavailable, plane-rejected, dependency-unavailable, unexpected (on the faulted arm only); tenant |
(none) | not charted - the three outcome arms are zero-primed; the five cause values partition faulted only and are deliberately not, so read a cause only once faulted is non-zero |
repocontext.ann.sweep.arming |
counter ({repository}) |
result = armed, deferred, faulted; tenant |
(none) | not charted - denominated by repository visits rather than by sweeps, so it must not be ratioed against repocontext.ann.sweep |
repocontext.ann.build.corpus |
counter ({build}) |
coverage = nonempty, unrestricted, filtered, denied, unknown; tenant |
(none) | not charted |
repocontext.ann.build.denial_terminal |
counter ({coordinator}) |
tenant |
(none) | not charted |
repocontext.ann.build.slice |
counter ({step}) |
repository = the onboarded repository id; space = {model-id}/{dimension} or unspecified; phase = coordinating, opening, ingesting, training, persisting, reconciling; progress = advanced, churned, starved, idle, faulted; cause = scan-page-stalled, projection-stale, dependency-unavailable, response-timeout, plane-rejected, saturated, unexpected (on the faulted arm only); tenant |
(none) | not charted - twenty-two (phase, progress) arms are zero-primed per (repository, space) plane when that plane is first armed, the seven cause values deliberately are not, so a cause reading zero is uninterpretable rather than innocent: read it only once faulted itself is non-zero, against which the causes sum exactly. coordinating and opening precede any build step, so they are primed on the faulted arm alone and the remaining eight combinations are deliberately never minted. phase on a fault is read at the fault site rather than snapshotted on entry, because a step entered in training also runs the persist within the same step; without that, a failed persist and a failed corpus read are the same series (issue #2855). Cardinality is repositories x embedding spaces, one durable index and one coordinator apiece, so it is operator-chosen and small |
repocontext.ann.build.slice.items |
counter ({item}) |
tenant |
(none) | not charted - the denominator that makes vectors-per-slice obtainable from metrics alone. Do not compute that ratio by dividing ann.vectorsIndexed from a health payload by repocontext.ann.build.slice: the first is cumulative and is inherited across a restart whenever the index was restored from durable state, the second is process-scoped and resets at the deploy boundary, so the two carry different epochs and their ratio is not a rate of anything. A slice that consumed nothing records nothing here, so this is not a slice count, and a slice that throws records nothing either, even when the items it consumed before the fault were checkpointed |
repocontext.ann.build.stage.duration |
histogram (s) |
stage = source_wait, key_assign, index_upsert, key_flush; tenant |
(none) | not charted - subject to the summary rendering described above, so read as rate(_sum)/rate(_count) and treat histogram_quantile against this endpoint as unavailable. source_wait, key_assign and index_upsert accumulate across the slice's items, and key_flush times its one batched key-map write. The four do not cover the whole slice: the source count taken while the expected count is still unknown, releasing the source enumerator, the ingest checkpoint that then persists the slice's vector chunks and build state, and the loop's own bookkeeping belong to no stage, so the four sum to less than the slice's elapsed time. A slice records all four only once it has checkpointed and written its build state, including a slice that consumed nothing; a slice that throws records nothing here. Deliberately not zero-primed: priming fabricates a zero-valued sample that reads as a real measurement of an instantaneous stage. An absent stage therefore means no slice has completed, not a stage that took no time |
repocontext.ann.build.step.in_flight_seconds |
observable gauge (s) |
repository = the onboarded repository id; space = {model-id}/{dimension} or unspecified; phase = coordinating, opening, ingesting, training, persisting, reconciling; tenant |
(none) | not charted - read as a ceiling, not a rate: any arm sitting above the coordinator's phase-tick period means a step is inside one non-reentrant turn and is not returning, which no terminal instrument on this plane can report because a step that does not complete never writes one (#3130). All six phases are zero-primed per (repository, space) plane, so an absent arm means the build did not ship, while six zeros mean no step is in flight rather than a broken instrument. The clock restarts on each phase entry, so whichever arm is non-zero names the half of the step that is stuck: ingesting is a corpus read, persisting is an index write. Cardinality is repositories x embedding spaces x six phases, one durable index and one coordinator apiece, so it is operator-chosen and small |
repocontext.ann.partitioning |
counter ({observation}) |
state = partitioned, unpartitioned-small, unpartitioned-large; tenant |
(none) | not charted |
repocontext.ann.repartition |
counter ({training}) |
outcome = partitioned, declined; tenant |
(none) | not charted |
repocontext.ann.index.load |
counter ({attempt}) |
outcome = fresh, resumed, faulted, deferred, refused, discarded; reason; tenant |
(none) | not charted. Discard reasons: count_mismatch, embedding_space_change, unloadable_record; fault reasons: timeout, embedding_space_change, unloadable_record, other; refused: admission_refused; remaining outcomes: none. unloadable_record is recorded only for a record absent on every read path or unreadable once returned; one read omitted and another returned is deferred. Every legal pair is pre-minted. Aggregate away reason for consumers of the former label set. |
repocontext.ann.vectors |
observable gauge ({vector}) |
repository; space; count = held, expected; tenant |
(none) | not charted. Resident and last-counted expected vectors from build progress, not a fresh source census. |
repocontext.ann.partitions |
observable gauge ({partition}) |
repository, space, tenant |
(none) | not charted. Trained partition count from build progress; zero is untrained, not necessarily empty. |
repocontext.retrieval.ann.search |
counter ({query}) |
state = bootstrapping, exhaustive, approximate; tenant |
(none) | not charted |
repocontext.retrieval.exact_gather.faults |
counter ({fault}) |
fault = stalled, timed_out, exhausted, abandoned, propagated, deterministic; tenant |
(none) | not charted - all six fault arms are zero-primed. The first four are absorbed as capacity and backed off; propagated is an index-integrity fault. Do not read a flat propagated count beside a climbing absorbed one as load - that reading ran issue #2948's six-hour total retrieval outage as capacity pressure. What separates a deterministic defect from load is the fault rate, which load cannot hold at one hundred percent, and deterministic is the arm that reports it. Either non-absorbed arm warrants an operator |
repocontext.retrieval.exact_scan.budget |
counter ({evaluation}) |
outcome = unbounded, corpus_unknown, within_budget, exceeded; tenant |
Overview | Exact KNN budget and gather outcomes - predicted exhaustion suppresses work before it starts; explicit exact mode bypasses the budget |
repocontext.retrieval.exact_scan.duration |
counter (s) |
tenant |
Overview | Exact KNN gather wall time - seconds/s including awaits and failed/cancelled gathers, not CPU utilization; overlapping gathers can exceed one |
repocontext.retrieval.exact_scan.gathers |
counter ({gather}) |
outcome = completed, faulted, cancelled; tenant |
Overview | Exact KNN budget and gather outcomes - terminal gathers only, not cache hits or skipped scans |
repocontext.retrieval.exact_scan.pages |
counter ({page}) |
tenant |
Overview | Exact KNN gather work - returned logical metadata pages, including empty terminal pages, not internal shard fills; completed pages remain counted after a later fault |
repocontext.retrieval.exact_scan.vectors |
counter ({vector}) |
tenant |
Overview | Exact KNN gather work - visited metadata records, including filtered spaces/missing payloads, not matches; excludes lookahead and partial pages that never return |
repocontext.retrieval.ready_seconds |
histogram (s) |
phase = serving, keyword_only, nothing_registered; tenant |
(none) | not charted |
repocontext.retrieval.unavailable |
counter ({event}) |
cause, tenant |
(none) | not charted |
repocontext.vectorplane.rederive |
counter ({event}) |
tree; outcome = observed, refused, completed, failed, denied, suppressed; tenant = default |
(none) | not charted - the one instrument on this meter whose tenant label is derived from its tree tag rather than fixed at _platform_; the repository-context trees have bare ids, so it reads default |
repocontext.retrieval.duration |
histogram (s) |
tool = search, context, outline, related; path = the resolved retrieval path, not_applicable (graph read), or unresolved (call ended before a path was settled); tenant |
(none) | not charted |
repocontext.retrieval.stage.duration |
histogram (s) |
stage = embed, vector_search, hydrate, keyword_scan; path as above; tenant |
(none) | not charted |
repocontext.bootstrap.pass_arm_faults |
counter ({fault}) |
arm = retire, ingest-files, ingest-symbols, ingest-memory; kind = scan-page-stalled or an exception type name; tenant |
(none) | not charted |
repocontext.bootstrap.phase_cancelled |
counter ({cancellation}) |
phase = Walking, Reconciling, Applying, Vectorising; tenant |
(none) | not charted - zero-primed for all four phases, so any non-zero value is an indexing run whose in-flight work was discarded |
repocontext.bootstrap.phase_cancelled.discarded_time |
counter (ms) |
phase = Walking, Reconciling, Applying, Vectorising; tenant |
(none) | not charted - running total of run time thrown away by the cancellations above; zero-primed for the same four phases |
repocontext.bootstrap.memory_marker_scan |
counter ({walk}) |
outcome = complete, resumed, banked; tenant |
(none) | not charted - read as a three-state pair, not a rate: banked rising with both completion arms flat is a marker scan that banks forever without converging (#2071), a non-zero resumed is the resumable cursor proven working end to end, and all three at their primed zero means the scan was never reached at all. A complete reading is not evidence the bootstrap as a whole is healthy (#2964). |
repocontext.bootstrap.coverage_probe |
counter ({probe}) |
arm = file, symbol, sweep; outcome = conclusive, gate_pruned, probe_failed; tenant |
(none) | not charted - read as a four-state grid, not a rate. gate_pruned > 0 on any arm is a standing misconfiguration that never clears by waiting: the ingestor cannot read its own membership keys, so it stands the back-fill sweep down and reports a converged bootstrap forever (#2964). probe_failed > 0 with gate_pruned == 0 in the same arm means that zero proves nothing, because the gate check sits below the probe-failure branch at every site. conclusive only is healthy. All nine at their primed zero means no coverage resolution was reached - which is also the normal steady state of a keyword-only deployment with no embedder bound, and this instrument alone cannot separate those two; check whether an embedding provider is registered. The arm tag localises which consumer stood down, not which grant is missing: there is one membership grant behind all three. |
repocontext.bootstrap.coverage_verdict |
counter ({verdict}) |
reason = converged, arm_failure, gap_found, probe_unmeasurable; tenant |
(none) | not charted - read as a four-way partition, not a rate. probe_unmeasurable > 0 sustained means the indexing pass cannot measure embedding coverage at all, so any prior converged reading is stale rather than true and the gap scan is standing down on an unmeasured verdict (#3340); it is actionable and does not clear by waiting. gap_found is the healthy armed state while a back-fill is outstanding and converged the healthy settled one, so gap_found falling away while converged rises is convergence. arm_failure is the loud arm and pairs with repocontext.bootstrap.pass_arm_faults. Distinct from repocontext.bootstrap.coverage_probe, which reports how each arm's probe ended rather than what the pass concluded from it. All four at their primed zero means no pass reached the verdict seam. |
repocontext.bootstrap.symbol_walk |
counter ({walk}) |
outcome = complete, resumed, banked; tenant |
(none) | not charted - read as a three-state pair, not a rate, exactly as memory_marker_scan above. banked rising with both completion arms flat is a symbol walk that banks forever without ever closing a circuit, which is the stall condition (#2953): the arm never banks a snapshot, so the symbol tree's WAL cursor floor never advances and its WAL is never reclaimed. A non-zero resumed is the resumable cursor proven working end to end - and it is the only reading that distinguishes a resumed walk from one that silently restarted at the head, because at the tree the two are identical. All three at their primed zero means the symbol arm was never reached. |
repocontext.ingest.files_scanned |
counter ({file}) |
repository = the onboarded repository id; tenant |
(none) | not charted - read against repocontext.ingest.files_embedded: scanned rising with embedded flat is a completed no-change reconcile, not a stall (#3151) |
repocontext.ingest.files |
counter ({file}) |
repository; outcome = added, updated, removed, unchanged; tenant |
(none) | not charted |
repocontext.ingest.files_embedded |
counter ({file}) |
repository, tenant |
(none) | not charted |
repocontext.ingest.symbols_embedded |
counter ({symbol}) |
repository, tenant |
(none) | not charted |
repocontext.ingest.files_content_projected |
counter ({file}) |
repository, tenant |
(none) | not charted |
repocontext.ingest.passes |
counter ({pass}) |
repository; outcome = completed, failed, cancelled; tenant |
(none) | not charted |
repocontext.ingest.pass.duration |
histogram (s) |
repository; outcome = completed, failed, cancelled; tenant |
(none) | not charted |
repocontext.ingest.last_completed_pass_age |
observable gauge (s) |
repository, tenant |
(none) | not charted - the alertable stall signal; alert on it climbing past the reconcile interval plus one pass duration, taking the minimum across silos |
lattice.repocontext.memory.restore |
counter ({attempt}) |
outcome = restored, partial, nothingtorestore, notattempted, failed; tenant |
(none) | not charted - durable-memory restore attempts, partitioned by what each did to the memory tree. The instrument is created only on a host that enables the durable-memory archive (LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR set), which attempts one restore when the archive service starts, so a host that has not opted in carries no series. Where it is enabled, every outcome is pre-minted at zero when the instrument is created and each attempt that returns a result is counted (one that throws is logged instead), so the total is a denominator and a zero on partial beside a non-zero total is a measured absence of damage. A non-zero partial always warrants an operator: the import wrote records and did not finish, so the tree is short of the archive while presenting as a populated store, and nothing else reports it. The only instrument on this meter not named repocontext.*, so the guards described above do not reach it |
The two retrieval-latency histograms are a matched pair and neither is readable alone. repocontext.retrieval.duration is recorded from a finally on every call, including a cancelled or faulted one (tagged path="unresolved" when it ended before a path was settled), so its _count is a true call total. repocontext.retrieval.stage.duration is recorded only for stages that actually ran, so it is deliberately sparse and its zero does not describe itself: it is the call total that turns the absence into a measurement. No embed beside a rising call total on repocontext.retrieval.duration is an intended keyword-only host; both at zero means no retrieval ran at all. Both are subject to the summary rendering described above, so read them as rate(..._sum[5m]) / rate(..._count[5m]) - a mean - and treat histogram_quantile against this endpoint as unavailable rather than as returning a wrong answer.
Several of these are partitions of a total rather than free-standing counts, and reading them as free-standing counts inverts their meaning. repocontext.ann.sweep counts every sweep including the faulting one, and a faulting sweep is re-run on a retry backoff rather than at the sweep interval, so outcome="faulted" must not be denominated by the sweep interval. repocontext.retrieval.ann.search counts every query including the ones the approximate plane could not answer. repocontext.ann.build.corpus counts every approximate-index build that reached Ready, including the ordinary coverage="nonempty" ones, which is what makes coverage="denied" reading zero beside a rising total a measured absence of authorization denial rather than an absent measurement - a denied range read returns a clean empty result rather than throwing, so without that denominator an authorization failure and an empty repository are the same observation. In each case a zero on one tag value beside a non-zero total is a measured absence. Every arm of these partitions is pre-minted at zero when its reporter is constructed, so an arm reading zero and an arm being absent are different observations rather than the same one: a series whose first occurrence falls after the collector reaches a ceiling is refused at creation and never appears, so an absent arm is a collector fault to be read from lattice_metrics_series and lattice_metrics_dropped_measurements_by_family_total, not a statement about the subsystem. All arms reading zero means no iteration has completed yet, which does not on its own establish that the loop is not running; the service's startup line reports that directly.
The derived tenant label
Every instrument on every meter carries a tenant tag. It is derived from the tree id rather than measured, and it is emitted on tenancy-on and tenancy-off clusters alike, so a panel or a named query is byte-identical in both deployment modes - there are no tenancy-on and tenancy-off variants of a query.
The value is one of three kinds:
| Value | Means |
|---|---|
| a tenant id | The series belongs to that tenant. Tenancy composes tree ids as t/{tenantId}/{name} and ownership is re-derived from that prefix. |
default |
The reserved legacy-adoption tenant, which owns every bare unsegmented tree id - and therefore every series on a cluster with tenancy off. It is a real, queryable tenant. |
_platform_ |
A reserved sentinel for series that belong to the platform and to no tenant: the _lattice_ and sys- tree namespaces, a malformed id in the reserved t/ namespace, and every instrument carrying no tree dimension at all. |
_platform_ is a sentinel rather than an absent label, deliberately. If platform-owned series were simply untagged, a tenant-scoped matcher would exclude them only by accident of absence, and any later change that started tagging them would silently widen every existing query. Naming the platform explicitly means {tenant="acme"} excludes it by stating so. The value opens with an underscore, which the tenant-id grammar forbids, so it can never collide with a real tenant.
Do not derive a tenant by regex over the tree label. Tree ownership is a genuine three-way classification and a single regex cannot reproduce it: tenant acme maps cleanly to tree=~"^t/acme/.*", but the default tenant's adopted legacy ids are bare, so its matcher becomes tree!~"^t/.*" - which also matches the _lattice_ and sys- platform namespaces and leaks platform-internal series into a tenant's view. An instrument with no tree tag cannot be scoped that way at all.
The label is cardinality-neutral. tree -> tenant is a function, so it attaches to series that already exist rather than multiplying them: two measurements that shared a series before still share one after, because equal tree ids always derive equal tenant labels.
A small number of instruments are documented as unscopable - cluster-level and per-peer telemetry that has no owning tree by construction. Those carry _platform_ rather than being left untagged, for the reason above.
orleans.lattice.tenancy remains the only meter whose instruments are about tenancy (quota, usage, enforcement). The tenant label described here is a dimension on everything else, which is a different thing: it says whose the series is, not what it measures.