---
title: "Configuration - Container quickstart"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.mcp.repocontext/container/configuration.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.api.mcp.repocontext/container.md?plain=1#L59-L221"
package: "Orleans.Lattice.Api.Mcp.RepoContext"
status: "unreleased"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.mcp.repocontext/llms-full.txt"
---
# Configuration

Part of [Container quickstart](../container.md).

The host is configured entirely by environment variables. The common ones:

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_DURABILITY` | `local` | The durability profile (`local`, `postgres`, `azure`; `postgresql` is accepted as an alias of `postgres`). |
| `LATTICE_DATA_ROOT` | `/data` | Root for all durable local state; must be a writable host mount. |
| `LATTICE_MCP_PORT` | `8080` | The MCP listener port. The health probes and the `/metrics` scrape endpoint are served on it too; it is the container's only application listener. A value that is not an integer in 1-65535 fails startup. |
| `LATTICE_WORKSPACE_ROOT` | `/workspace` | The read-only root that runtime-registered repositories must resolve under; a path escaping it is refused. |
| `LATTICE_EMBEDDING_ENDPOINT` | `http://localhost:9000` | The separate embedding companion's base address. The embedding provider is always bound, so this repoints it at the companion rather than switching semantic search on; semantic search degrades to keyword ranking whenever that address cannot be reached. Must be an absolute URI or startup fails. |
| `LATTICE_WAL_DIR` / `LATTICE_SQLITE_PATH` | under the data root | Override the WAL directory or SQLite file path individually. |
| `LATTICE_WAL_COMPACTION_MAX_DEAD_BYTES` | `0` (disabled) when unset; the sample compose sets `134217728` (128 MiB) | An absolute ceiling, in bytes, on the dead (trimmed but not yet reclaimed) bytes a single file-WAL **shard** may hold before a compaction is forced. Compaction is otherwise triggered only by a **ratio** - dead bytes reaching half the file - which bounds waste relative to live data rather than absolutely, so a growing log roughly doubles in size before any space returns to the filesystem. The library ships this at `0` (disabled), leaving that ratio as the only trigger, and so does this host when the variable is unset; the sample compose file names a value because issue #3223 measured what the ratio alone does here, with the ratio declining at 200/1268 = 0.158 on every sweep, `orleans_lattice_wal_compactions_total` at zero on all three trigger tags, and physical WAL reaching 10174 MiB against the 8192 MiB ceiling the row below then carried. The gate is per shard, not per tree, so the tree-wide bound is this value times the shard count: 128 MiB over the 8 shards of `repo-context-vector-index` bounds dead at 1024 MiB - one eighth of the 8 GiB ceiling it was sized against, and one twentieth of the 20 GiB the sample declares today - where the ratio permits one half. Size it to FIRE against measured dead per shard - a ceiling above it changes nothing at all - and against the cost, which is write amplification of live/ceiling per byte reclaimed (here 1068/128 = 8.3x, against 1.0x for the ratio) paid as a synchronous whole-shard rewrite whose duration is set by live bytes and so does not shrink as the ceiling shrinks. Only `0` disables it; a non-zero value below the 64 KiB `CompactionMinimumDeadBytes` floor is **not** inert but maximally aggressive, and is refused at startup (issue #3210). It bounds only dead bytes, so it cannot by itself bring a tree under `LATTICE_WAL_MAX_RETAINED_BYTES` when retained alone already exceeds it. `orleans_lattice_wal_compactions_total{trigger="ceiling"}` reports when this trigger fires. |
| `LATTICE_WAL_PIN_BUCKETS` | `8` | How many persisted slots the WAL materialiser retention-floor pin state is split across, so an advancing floor rewrites a fraction of the pin blob rather than all of it. Accepts 1-256; `1` is the library's legacy single-slot write path. Widening self-migrates on activation and leaves the legacy slot intact, so reverting to `1` is a safe rollback that over-retains WAL rather than over-trimming it. |
| `LATTICE_WAL_PIN_SHED_CEILING_SECONDS` | `120` | Ceiling on how long one durable pin shard may shed steady-state pin reports *continuously* before one report is forced through regardless (issue #3310). The shed path is the only write path carrying an advancing `CheckpointOffset`, so an unbounded shed freezes the WAL GC durable offset floor and retained WAL grows without bound. Accepts 0-3600; `0` disarms the ceiling and is the exact rollback to the pre-fix behaviour, which over-retains WAL rather than over-trimming it. Forcing cannot overstate durability - the leaf clamps the reported offset to `min(checkpoint, covered)` before the report is built - so this bounds *when* a true value is published, never *what*. Setting it too low is the real hazard: it forces writes faster than the pin store can absorb and re-saturates the queue the shedding protects. |
| `LATTICE_SETMANY_FANOUT_BUDGET_SECONDS` | `30` | Budget for a single batch-write fan-out across shards, in seconds (issue #3384). A batch write whose keys span more than one shard scatters to every shard owning a key in the batch and awaits them all; without a budget one unresponsive shard holds the whole fan-out, and its caller, for as long as it stays unresponsive. With the budget, a fan-out still unsettled when it expires is refused with `LatticeSaturatedException` (source `SetManyFanOut`) and rolls nothing back: shards that already committed stay committed and branches still in flight keep running. A batch whose keys all land on one shard is a single call this budget does not bound. Accepts 0-3600. **Read the zero the opposite way round to `LATTICE_WAL_PIN_SHED_CEILING_SECONDS` above**: there a zero disarms a bound, whereas here `0` resolves to an *infinite* budget - the pre-fix behaviour, under which a fan-out waits indefinitely - and is the exact rollback. It deliberately does not mean `TimeSpan.Zero`, which the library refuses because a zero budget would refuse every fan-out and turn a rollback into an outage. Setting it too low is the real hazard: a budget below the time a healthy fan-out legitimately takes converts ordinary slowness into failed writes, so size it against measured fan-out duration and not against a latency target. |
| `LATTICE_WAL_MAX_RETAINED_BYTES` | `1073741824` (1 GiB) when unset; the sample compose sets `21474836480` (20 GiB) | An advisory per-tree ceiling on the WAL bytes held on disk. On reaching it, each garbage-collection pass lowers its trim frontier toward 80% of the ceiling - but only within the frontier that was already safe to trim, so it never blocks a write, never trims past a live consumer cursor, and never loses a record (what it trims has already been materialised into leaves). The library ships this disabled because a correct value is a fraction of the volume the WAL lives on, which a library cannot know; this container declares its own volume, so it names one. Left disabled, every tree reports `orleans_lattice_wal_gc_backlog_bytes_unavailable_total{reason="policy_disabled"}` and reclaims nothing, so on-disk occupancy only ever grows. Costs one physical-size probe per WAL partition per pass. Accepts an integer of at least `67108864` (64 MiB), or `0` to restore the library default; anything else refuses startup. Size it at no less than twice the largest tree's logical working set: the ceiling is compared against physical bytes summed over a tree's partitions, and WAL compaction only reclaims once a shard is 50% dead, so steady-state physical settles at about 2x logical and a tighter ceiling can never be satisfied or disarmed (issue #3234). Raise it on a deployment indexing a substantially larger corpus. This container was resized from 4 GiB to 8 GiB by issue #3242 as headroom, not in response to a measured overrun: its `vector-index` tree grew from the 1077 MB logical working set the 4 GiB was calibrated against to about 1672 MB, moving the 2x floor from 2154 MB to about 3344 MB and consuming most of the margin that calibration had left. An earlier revision of this row stated the 4096 MB ceiling had already been overrun; that figure came from reading a Prometheus `summary` series as an instantaneous value when `sum/count` on a summary is a lifetime mean, and it is withdrawn. Measure a tree's logical working set with the `orleans_lattice_storage_wal_bytes` gauge, which is sampled at scrape time; never derive it by differencing a summary. A ceiling calibrated against a measured working set acquires an expiry date the moment that workload can grow, so re-check it against the tree rather than treating a past calibration as durable. Raised from 8 GiB to 20 GiB by issue #3094 after `wal_gc_ceiling_unsatisfiable` fired continuously on this container. The 8 GiB ceiling firing at all implies a logical working set above 4 GiB; physical was measured at 5.864 GiB and logical is bounded by physical, so the 2x floor lay somewhere in 8.0 to 11.7 GiB - at or above the ceiling, which is why the policy could never be satisfied or disarmed. 20 GiB clears the top of that range with about 1.7x headroom. After the raise the counter stayed at 0 and on-disk occupancy stopped climbing monotonically, oscillating within about 25 MB instead. |
| `LATTICE_POSTGRES_CONNECTION_STRING` / `LATTICE_AZURE_STORAGE_CONNECTION_STRING` | (unset) | Required by the `postgres` / `azure` profiles. |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR` | (unset, feature off) | Directory the agent-memory archive is exported to and restored from. Point it at a path **outside** the data volume - a bind mount rather than a named volume - or it dies with the state it exists to outlive. See [Memory durability](../memory-durability.md). |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_INTERVAL_SECONDS` | `300` | Export cadence; the size of the window an ungraceful stop loses. Positive values below 30 are raised to 30, and values above the longest delay a timer can wait (about 49.7 days) are lowered to it; a zero, negative or unparseable value, or one too large to represent as a duration (for example `1e20`), falls back to `300`. |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_RESTORE` | `auto` | `auto` restores into a store holding no memory, or into one whose restore-state marker records that an earlier restore was left partial (so a half-imported tree is healed rather than mistaken for a populated one); `always` restores every start; `off` never restores. `none` and `false` are accepted for `off` and `on-empty` for `auto`; any other value falls back to `auto`. |
| `LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_STOP_TIMEOUT_SECONDS` | `20` | Budget for the final export during a graceful stop. A positive value is clamped to 1-60; a zero, negative or unparseable value falls back to `20`. It shares the stop grace period with the drain, so it is deliberately a fraction of it. |
| `LATTICE_REPOCONTEXT_HEAP_ADMISSION_OVERRIDE` | (unset) | Suppresses the startup refusal described under [Measured requirement and startup admission](graceful-shutdown.md#measured-requirement-and-startup-admission). Its value must be the **exact** recorded byte count being disbelieved, which the refusal message prints and `lattice_repocontext_heap_insufficient_limit_bytes` publishes; any other value, including `true` or `1`, is ignored and the refusal stands. It is deliberately not a boolean, so it cannot be set once and left to suppress correct refusals for ever, and it stops matching by itself once a larger exhaustion is recorded. |

A profile is a preset, not a straitjacket: each store it selects can be overridden on its own, and the remaining variables name the cluster and the embedding space. An unrecognised value for any of the four provider variables fails startup rather than falling back silently:

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_WAL_PROVIDER` | `azure` under the `azure` profile, otherwise `file` | Selects the WAL provider on its own. Accepts `file` or `azure` (`azuretable`). |
| `LATTICE_GRAIN_STORAGE` | the profile's store (`sqlite` / `postgres` / `azure`) | Selects the grain-storage provider on its own. Accepts `sqlite`, `postgres` (`postgresql`), or `azure` (`azuretable`). |
| `LATTICE_REMINDERS` | the profile's store | Selects the reminders provider on its own; same accepted values as the grain store. |
| `LATTICE_CLUSTERING` | `azure` under the `azure` profile, otherwise `localhost` | Selects the clustering provider. Accepts `localhost` (`local`) or `azure`. |
| `LATTICE_AZURE_WAL_TABLE` | `RepoContextWal` | The Azure Table the WAL writes to when the Azure WAL provider is selected. |
| `LATTICE_EMBEDDING_MODEL` | `nomic-ai/nomic-embed-text-v1` | The embedding model id requested from the companion. |
| `LATTICE_EMBEDDING_DIMENSION` | `768` | The embedding vector dimension; must match the model the companion serves. A non-positive value fails startup. |
| `LATTICE_CLUSTER_ID` | `repo-context` | The Orleans cluster id. |
| `LATTICE_SERVICE_ID` | `repo-context` | The Orleans service id. |

Selecting any Azure-backed store without `LATTICE_AZURE_STORAGE_CONNECTION_STRING` refuses to start rather than silently degrading durability. Changing `LATTICE_EMBEDDING_MODEL` or `LATTICE_EMBEDDING_DIMENSION` is a **new embedding space**, so it builds a wholly separate approximate index under its own prefix - which is what `LATTICE_REPOCONTEXT_ANN_INDEX_RECLAMATION` below then retires the superseded one for.

The background reconcile cadence (see [Background reconcile and change detection](../container.md#background-reconcile-and-change-detection)) is tuned by six further variables. The two periodic reconcile deadlines - the full walk and the embedding gap scan - are declared in wall clock but **counted in reconcile passes**: each is divided by the widest scheduled reconcile spacing (`LATTICE_RECONCILE_INTERVAL_SECONDS` plus `LATTICE_RECONCILE_JITTER_SECONDS`), rounded up, and clamped to at least one pass. That is what makes them hold on a large repository, where a pass routinely takes longer than its own scheduled spacing and a wall-clock deadline would be past on arrival every single time. The coverage-digest audit is not pass-counted; see its row:

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_SELFINDEX_TICK_SECONDS` | `15` | How often each repository's self-index grain ticks; the reconcile cannot fire more often than this. A value longer than a grain timer can wait (about 49.7 days) ticks at that ceiling. |
| `LATTICE_RECONCILE_INTERVAL_SECONDS` | `900` | Base interval between periodic content reconciles. A small value (with zero jitter) makes the reconcile effectively continuous, bounded by the tick. |
| `LATTICE_RECONCILE_JITTER_SECONDS` | `300` | Maximum extra random interval added on top of the reconcile interval to desync repositories. |
| `LATTICE_FULL_WALK_INTERVAL_SECONDS` | `3600` | How often a reconcile is forced to ignore the directory-modification-time prune cache and stat every file, bounding how stale an in-place content edit can be. Counted in passes: at the shipped defaults it is 3 reconciles, so 2 in every 3 prune. Set it at or below one reconcile spacing and it degenerates to 1 pass, meaning every reconcile walks in full and pruning never engages. |
| `LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS` | `1200` | How often a reconcile re-checks every content-unchanged file for an embedding gap - a file whose structural record is committed but whose vector never landed. Detection reads the per-page **vector-coverage digest**: a fixed 257 rows on a tree of its own, whatever the corpus size. Counted in passes, and at the shipped defaults this is deliberately 1 - the shortest window the scheduler can express, so a gap is found on the next reconcile rather than up to four hours later. Because it is 1, the periodic cadence has no effect at the defaults: the scan runs on every pass whatever the coverage verdict says, and the host logs a startup warning saying so. Set it above one reconcile spacing (for example `2400` at the defaults) for the cadence to engage. The former `14400` default was a ration against a probe that cost two membership reads per indexed source; that cost is gone, and so is the reason for the ration. |
| `LATTICE_COVERAGE_DIGEST_AUDIT_INTERVAL_SECONDS` | `86400` | How often the self-index grain's gap sweep re-derives the vector-coverage digest exhaustively from the membership tree, rather than trusting the incremental maintenance the write path performs. Unlike the two rows above it is applied in wall clock, not reconcile passes: the audit falls due this long, plus up to a minute of jitter, after the previous one, and runs when the sweep next checks. The effective-configuration report currently prints it rounded up to whole reconcile spacings, which is not the cadence the sweep applies. This is the O(sources) read the gap scan used to be, and it is what bounds digest drift - from an interrupted write, or from a self-heal that reset the membership tree under a surviving digest. Raising it widens the window in which a drifted digest can mask a gap; lowering it re-imports the cost this design removed. |

> **These interval variables are a matched set.** `LATTICE_FULL_WALK_INTERVAL_SECONDS` and `LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS` are wall-clock values that are converted once into **pass counts** by dividing by the reconcile spacing (`LATTICE_RECONCILE_INTERVAL_SECONDS` plus `LATTICE_RECONCILE_JITTER_SECONDS`). Changing the reconcile interval therefore silently re-denominates both. Raising it far enough that the full-walk interval floors to a single pass switches directory-modification-time pruning off entirely - no error, and the prune cache is written on every run but never read. If you raise the reconcile interval, restate both. The host logs the derived pass counts next to the configured seconds at startup (`full walk 120 s = 24 pass(es) ...; pruning can engage: True`), and warns when the arithmetic has disabled pruning or has collapsed the embedding gap scan to every pass, so the conversion never has to be worked out by hand.

Two further variables tune the indexing role and per-file token counting, and two select the semantic-retrieval path and size the vector cache:

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_REPOCONTEXT_INDEXING_ROLE` | `hub` | The cluster's indexing role: `hub` (the authoritative indexer that walks, reconciles, prunes, and re-embeds) or `spoke` (a read-only replica whose index pass is inert). An absent or unrecognised value falls back to `hub`. |
| `LATTICE_REPOCONTEXT_TOKENIZER` | `o200k` | The BPE tokenizer profile the per-file token counter uses: `o200k` (OpenAI o200k_base) or `cl100k` (OpenAI cl100k_base). An absent or unrecognised value falls back to `o200k`. |
| `LATTICE_REPOCONTEXT_SEMANTIC_RETRIEVAL` | `approximate` | Which semantic retrieval path is bound: `approximate` routes semantic search through the persisted approximate nearest-neighbour index (bounded recall, sub-linear query cost, survives a restart), and `exact` routes it through the complete-recall brute-force scan instead, whose cost is proportional to the corpus. An absent or unrecognised value falls back to `approximate`. A host set to `exact` maintains no approximate index at all, so the build coordinator below is inert for it. Documented in full under [Semantic search](../semantic-search.md#the-two-paths). |
| `LATTICE_VECTOR_CACHE_TTL_SECONDS` | `30` | How long (in seconds) a warm decoded-vector candidate set is trusted before it is re-gathered from the store; `0` disables the cache. |

Four variables govern the **adaptive indexing pacer** (issue #3447), the silo-wide controller that smooths the embedding drain so a large pass no longer lands as one sustained spike. It sits in front of every embedding batch and never changes what a pass embeds, only when each batch runs: a failed or slow batch (over 250 ms and over 2.5x the learned per-passage latency baseline), a throttled vector tree, or GC memory load at its high-load threshold doubles an inter-batch delay from 250 ms up to the ceiling, and each clean batch walks it back down. Only a batch at least half the size of the largest seen is judged slow or moves the baseline, so a small remainder batch is neither misread as congestion nor drags the baseline down, and the baseline is never lowered by a batch that ran under a backoff delay. A vector tree that stays throttled for 60 seconds while the delay is held at its ceiling is not moved by indexing (a materialiser drain-lag minimum is the usual cause), so it stops counting as congestion until it clears or worsens; the WAL still paces throttled appends itself (issue #3456). A saturated vector tree holds the next batch for up to 30 seconds, an in-flight `repocontext_search` or `repocontext_context` call holds it for up to 2 seconds, and after each work slice the drain rests. While a foreground request is open or the drain is backing off, the approximate-index build defers its slices (at most 15 phase ticks in a row) and the self-index grain postpones a due coverage-digest audit (at most 12 sweeps in a row), so both are slowed and never starved. What the pacer is doing is reported as `pacing` on `repocontext_index_status`; see [Adaptive pacing](../tools.md#adaptive-pacing).

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_REPOCONTEXT_PACING` | `true` | Whether the pacer is on. `false` restores back-to-back batches with no delay, yield, rest, or deferral. An absent or unrecognised value falls back to `true`. |
| `LATTICE_REPOCONTEXT_PACING_SLICE_SECONDS` | `60` | How long the drain works before it rests. `0` switches the duty cycle off (batches are still delayed under congestion). |
| `LATTICE_REPOCONTEXT_PACING_REST_SECONDS` | `5` | How long the drain rests after each work slice. `0` switches the duty cycle off. |
| `LATTICE_REPOCONTEXT_PACING_MAX_DELAY_SECONDS` | `5` | The ceiling on the congestion-driven inter-batch delay. |

The next variables are the kill switches for the approximate index's own housekeeping, which default on, and the cadence of its build sweep. They are documented in full under [Scheduling the approximate index build](../semantic-search.md#scheduling-the-approximate-index-build):

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_REPOCONTEXT_ANN_INDEX_SCHEDULING` | `true` | Whether the approximate index build is scheduled by its durable, reminder-anchored coordinator - which is what lets a restored volume converge to a serving index with no client traffic at all, and what resumes a build interrupted by a process death. Set `false` and no index is built at all: every semantic query is answered by the exact scan with complete recall, within the bounds described under [When the exact fallback is declined](../semantic-search.md#when-the-exact-fallback-is-declined). An absent or unrecognised value falls back to `true`. |
| `LATTICE_REPOCONTEXT_ANN_INDEX_RECLAMATION` | `true` | Whether an index that has just reached `Ready` retires the sibling prefixes of its own repository whose embedding-space fingerprint is no longer live. A model or dimension change otherwise leaves the previous index resident forever. Set `false` to keep a superseded space for a deliberate roll-back. An absent or unrecognised value falls back to `true`. |
| `LATTICE_REPOCONTEXT_ANN_SWEEP_INTERVAL_SECONDS` | `900` | How often the build sweep re-arms every registered repository's coordinator. Floored at 60 seconds: a shorter value is raised to the floor, and the startup line says so rather than leaving the setting to look ignored. |

> **The sweep cadence is deliberately not part of the matched set above.** It used to be: the sweep took its interval from `LATTICE_RECONCILE_INTERVAL_SECONDS`, so raising that variable to quiesce walk load - a reasonable action, with nothing in its name to suggest otherwise - throttled index arming by the same factor. That is worse than a slow sweep. Two things arm a coordinator, this sweep and the self-index grain finishing a vectorising pass; a converged repository whose index was never built has no vectorising pass to finish, so the sweep is its **only** arming path, and the vectorising pass was paced by the reconcile interval too. Raising it did not slow one path of two, it slowed the only two there are. The index then serves nothing while the retrieval counter records `state="bootstrapping"`, which at the metric is indistinguishable from a genuine index defect. `LATTICE_REPOCONTEXT_ANN_SWEEP_INTERVAL_SECONDS` defaults to 900 seconds, which is the reconcile interval's own default, so a host that configures neither variable sweeps at exactly the cadence it always did.

Five more variables bound how long the approximate index may hold its build coordinator's turn while it opens and ingests. An absent or malformed value falls back to the default:

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_REPOCONTEXT_ANN_OPEN_SLICE_BUDGET_SECONDS` | `5` | Wall-clock ceiling on one attempt to open (restore) the durable index. A stopped attempt banks what it walked and the next continues past it, so this slices one long open into short ones. `0` removes the bound. A value above about 49.7 days (the longest wait a timer accepts) is held at that ceiling. |
| `LATTICE_REPOCONTEXT_ANN_OPEN_SLICE_MAX_EXTENSIONS` | `6` | How many further open-slice budget periods an open slice that has banked nothing may take before the budget fires anyway. `0` reproduces the elapsed-only bound. |
| `LATTICE_REPOCONTEXT_ANN_INGEST_SLICE_BUDGET_SECONDS` | `5` | Wall-clock ceiling on one ingest slice of the build, so the coordinator's keep-alive reminder and arming calls are answered while a build runs. `0` removes the bound, leaving the 4,096-vector slice batch as the only one. A value above about 49.7 days (the longest wait a timer accepts) is held at that ceiling. |
| `LATTICE_REPOCONTEXT_ANN_OPEN_MAX_CONSECUTIVE_REFUSALS` | `12` | Consecutive admission refusals after which the open declares itself terminally saturated. Declaring does not stop retrying. `0` removes the count bound. |
| `LATTICE_REPOCONTEXT_ANN_OPEN_REFUSAL_TERMINAL_SECONDS` | `600` | How long an unbroken run of admission refusals may last before the same terminal state is declared, whichever bound is reached first. `0` removes the elapsed bound. |

Three further variables bound resources whose defaults come from a runtime fact or a fixed assumption rather than from the deployment's real limit, so a constrained container can state the limit it actually has:

| Variable | Default | Purpose |
|---|---|---|
| `LATTICE_WAL_MAX_CONCURRENT_REPLAYS` | `0` (defer to the library) | The per-silo ceiling on concurrent activation-time leaf WAL replays. Each permit admits one whole-readable-window replay holding multi-MiB buffers, so although the default ceiling is derived from CPU figures the gate behaves as a memory admission gate: the resource a reactivation storm exhausts first is managed memory, not CPU (issue #2784). `0` defers to the library, which sizes the gate from the **lesser** of `Environment.ProcessorCount` and the container's enforced cgroup CPU grant. Accepts `auto` or an integer 0-256; anything else fails startup rather than being silently ignored. `auto` defers to exactly the same derivation as `0`, and exists so that a deployment can state that it *meant* to derive - `0` reads as a pinned zero to anyone scanning the file, which is not what it does. |
| `LATTICE_MAX_LOCK_LEASE_SECONDS` | `1800` | The ceiling this host clamps every named-lock lease to, including the claim leases agents take through `repocontext_claim`. It bounds how long a crashed holder can pin an item while still covering a full build-and-test cycle. Accepts 30-7200; anything else fails startup. |
| `LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD` | `120s` when unset; this sample declares `240s` | The container grace period this deployment grants between `SIGTERM` and `SIGKILL`, declared to the process that has to fit inside it. The host derives its shutdown budget from it (75% of the grant, or all but a two-second unwind reserve, whichever is smaller), so the sample's declared `240s` yields a 180s budget and the `120s` default yields 90s. Accepts a positive number of seconds up to `3600`, with an optional `s` suffix; compound forms such as `1m30s` are refused rather than misread, and so is a grant of two seconds or less, which leaves no budget once the reserve is taken. It must equal the `stop_grace_period` on the same service - see [Where the budget comes from](graceful-shutdown.md#where-the-budget-comes-from-and-why-it-is-not-a-free-parameter). |

> **You no longer need to set the replay ceiling beside your CPU limit.** `Environment.ProcessorCount` honours a container CPU quota only while `DOTNET_PROCESSOR_COUNT` does not override it, and that variable takes precedence over the quota-derived value. A container granted 6 CPUs whose environment also carries `DOTNET_PROCESSOR_COUNT=16` therefore used to size this gate at 16, not 6, with nothing inside the process able to tell the difference - the condition this sample's own compose file predicts, measured live at a 2.67x oversubscription. The library now reads the enforced cgroup quota itself and takes the **lesser** of the two figures, so the default follows the grant without an operator declaring anything, and it does so adaptively rather than being pinned to whatever the grant was on the day it was written (issue #2816). Pinning this variable is still honoured and still supersedes both figures, but it is now an override rather than a requirement, and a pinned value has the usual hazard of silently diverging from a `cpus` limit that later changes. The host logs the resolved ceiling once at startup alongside the configured option, the `Environment.ProcessorCount` the runtime reported, and the CPU grant it read, so the effective figure and any disagreement between the two inputs can be read off the log instead of inferred from the host's vCPU count.

An opt-in family of `LATTICE_REPOCONTEXT_GIT_*` variables switches a repository from the mounted workspace to a git remote; see [Index source strategies](index-source-strategies.md).

## Garbage collection on a multi-GiB heap

This host's steady-state working set is measured in GiB, and Workstation GC is the wrong collector at that size. Workstation GC collects a single heap and its blocking gen2 phases are effectively single-threaded, so one collection walks the whole heap on one thread with every other thread in the process suspended. A deployment of this host at about 11 GiB resident, held on Workstation GC by an explicit `DOTNET_gcServer: "0"` in its compose overlay, had the runtime attribute a **252 second** pause to the collector, against a 30 second Orleans request timeout.

| Variable | Default | Purpose |
|---|---|---|
| `DOTNET_gcServer` | unset, which runs **Server** GC: the host is built with the ASP.NET Core Web SDK, which writes `System.GC.Server: true` into its `runtimeconfig.json` | `1` selects Server GC, which collects several heaps in parallel; `0` forces Workstation GC. |
| `DOTNET_GCHeapCount` | one heap per processor | Bounds how many heaps Server GC creates. **Read as hexadecimal** - see below. Inert under Workstation GC (`DOTNET_gcServer=0`); because this host defaults to Server GC, it applies whenever `DOTNET_gcServer` is unset or `1`. |

Neither is set in the base compose file, because the right heap count is a property of your CPU grant rather than of any file in this repository. Left unset, the base sample therefore runs Server GC with a heap count the collector chooses for itself; the effective-configuration report below states which (`GC.Mode`, `GC.HeapCount`). The sample's tuning overlay (`docker-compose.tuning.yml`) does set both on the `repocontext` service - `DOTNET_gcServer: "1"`, and a `DOTNET_GCHeapCount` it refuses to start without - taking the heap count from the CPU grant that `scripts/New-TuningEnv.ps1` derives for that service on this host.

> **`DOTNET_PROCESSOR_COUNT` cannot double as the heap count.** Server GC sizes its heap count from the processor count, which is exactly what `DOTNET_PROCESSOR_COUNT` overrides - the same variable the replay ceiling above discusses. That variable may legitimately be pinned **above** the container's CPU grant for the thread pool's sake, and reusing it as a heap count would then create one heap per phantom processor on a heap already near its ceiling. One variable, two jobs, opposite requirements. `DOTNET_GCHeapCount` separates them: derive it from the container's actual CPU grant, independently of `DOTNET_PROCESSOR_COUNT`, and leave the processor count to the pools that genuinely want it.

> **Write the heap count in hexadecimal.** The collector reads its numeric **environment variables** as hex, while the same settings in `runtimeconfig.json` are decimal. `DOTNET_GCHeapCount=10` therefore asks for **16** heaps and `=16` asks for **22**, silently and with no error. Values below `10` read identically either way, which is what makes this easy to miss on a small box and then get wrong on a large one. Prefer an explicit `0x` prefix.

Verify rather than assume. The effective-configuration report states `GC.Mode`, `GC.HeapCount` (the figure the collector **resolved**, not the one declared), the process's memory ceiling, and the accumulated `GC.GetTotalPauseDuration`, and it raises a `GC HAZARD` **warning** when the process runs Workstation GC against a large ceiling, when a heap count is declared under Workstation GC and is therefore inert, or when the resolved heap count disagrees with the number that was written. Grep the log for `GC HAZARD`.

**The claim is narrow on purpose.** Server GC with a bounded heap count removes the class of pause that is multi-minute, process-wide, and attributed to the collector by the runtime itself. It is not a general remedy for stalls: measurement of the same container found collector pauses accounted for under a third of long-silence time and did not explain its largest timeout burst at all. A stall the runtime does not attribute to the collector needs its own diagnosis, and `GC.GetTotalPauseDuration` is the quantity to reach for rather than gaps between log timestamps.

## Thread pools on a CPU-limited container

The collector is not the only pool sized from a number that a CPU limit does not constrain. The `embedder` service in the same sample sizes its ONNX Runtime intra-op pool - the threads that parallelise a single inference - and its own default reads the **host core count** while ignoring the cgroup quota entirely.

| Variable | Default | Purpose |
|---|---|---|
| `EMBED_INTRA_THREADS` | derived from the enforced cgroup CPU quota | Sizes the ONNX Runtime intra-op thread pool. Accepts `auto` or a non-negative integer; `0` hands the decision back to ONNX Runtime, and `auto` derives from the grant deliberately. Leaving it **unset** also derives. A value that is present but unusable - a negative number, a non-number, a typo of the token - fails startup rather than deriving, because absence and a typo are different facts about a deployment and deriving from both makes them indistinguishable (issue #2887). |

It is not set in the base compose file, for the same reason the heap count is not: the right value is a property of your CPU grant rather than of any file in this repository. Left unset, the embedder reads the quota itself, which is the recommended configuration. The tuning overlay does declare it, and refuses to start without it, using the embedder CPU grant that `scripts/New-TuningEnv.ps1` derives for this host.

**The cost of getting it wrong is worse than proportional.** Measured on a 4.0-CPU grant (`cpu.max = "400000 100000"`) on a 16-core host: an intra-op pool of **16**, a 4x oversubscription, with the kernel throttling **296 of 298** consecutive scheduling periods and the pool accumulating **346.3 CPU-seconds stalled against 118.8 CPU-seconds run**. The arithmetic behind that ratio is elementary once written down: sixteen threads drain a 400ms quota in 400/16 = **25ms** of wall time and are then frozen for the remaining **75ms** of the period, predicting 75:25 = **3.0** stalled per unit run against **2.91** measured, within 3%. Throughput does not merely fall by the oversubscription ratio, because ONNX Runtime synchronises its intra-op threads at **every operator boundary** and a transformer inference crosses hundreds of them; a freeze landing mid-barrier stalls the whole operator rather than one thread. The observed embedding rate was 1.8 files per minute, projecting roughly 77 hours for a single 8,315-file checkout.

> **`DOTNET_PROCESSOR_COUNT` cannot double as the thread count either.** This is the same variable, doing a third job with a third set of requirements. Whenever it is set on the `repocontext` service - the tuning overlay declares it there by name only, so it is absent unless your environment or `.env` supplies a value - it **overrides** `Environment.ProcessorCount` and wins over the quota, and copying that service's environment block onto the embedder - an entirely ordinary thing to do - would silently restore the oversubscription. The embedder reads `/sys/fs/cgroup/cpu.max` directly and is immune to it, and the WAL replay gate now does the same (issue #2816); this variable remains the hazard for every pool that does **not**. When the two disagree the embedder logs a `CPU GRANT MISMATCH` warning naming both figures and the resulting factor, because a process that believes it has sixteen processors under a four-CPU grant will oversubscribe **every** pool sized from that belief.

**Deriving it, if you choose to declare it.** Use the container's actual CPU grant, rounded **up**: `cpus: "4.5"` becomes `5`. That matches what .NET itself reports for the same limit, so the declared pool never disagrees with the runtime in the unsafe direction. Do not use the host core count, and do not use `DOTNET_PROCESSOR_COUNT`.

> **A declared value does not follow the grant.** If you later change `cpus` and leave `EMBED_INTRA_THREADS` pinned, the pair silently diverges and nothing in the container will object - the number is no longer wrong in a way any single file reveals. That is the entire hazard of pinning one, and it is why the derived default is recommended. The same reasoning is why `LATTICE_WAL_MAX_CONCURRENT_REPLAYS` is no longer something the base sample asks you to declare; the tuning overlay still pins it, to the CPU grant `scripts/New-TuningEnv.ps1` derives, and refuses to start without one. If you do pin either, keep it beside the `cpus` limit so the two are checkable against each other, and prefer `auto` over the derived default's numeric equivalent when what you mean is "derive": `auto` reaches the same number while recording that the derivation was chosen rather than inherited.

Verify rather than assume. The embedder states its resolved intra-op count once at startup, marked `DECLARED` when an operator supplied it and `DERIVED` when it did not, so the effective figure can be read off the log rather than inferred from this file. Reading a compose file tells you what was written; only the log tells you what the process resolved.

## Reading the effective configuration off the log

The container's real settings arrive partly from files this repository does not track - the per-host `.env` that supplies the tuning overlay's derived values, and any machine-local override - so reading this repository does not tell you what a running process resolved. The host therefore states its own resolved configuration once at startup, on the `Repository-context effective configuration:` prefix, and that report supersedes any file when the two disagree:

- a summary line counting the settings, how many differ from the host default, how many were not declared at all, and how many supplied `LATTICE_` variables nothing reads;
- one line per setting, carrying the value this process resolved, marked `[OVERRIDDEN...]` when it differs from the host default;
- a `SCOPE:` line, described below;
- one line per supplied variable recognised only by a family prefix (a `LATTICE_REPOCONTEXT_GIT_*` member), with its value withheld;
- a **warning** per supplied `LATTICE_` variable that nothing in this host binds;
- a **warning** per hazardous garbage-collector configuration, prefixed `GC HAZARD`.

The same reporter then states, at warning level and on its own `Repository-context memory durability:` prefix, where agent memory lives and what does and does not protect it; see [Memory durability](../memory-durability.md).

Grep the log for `SUPPLIED BUT NOT READ` to find a variable an operator set that never reaches anything - the silent failure that motivated the report. Values are printed through an allowlist, so a key that is not classified as safe to print renders as `<redacted: unclassified>` rather than leaking; a variable matched only by a prefix renders as `<withheld: matched by prefix only>`, because the host recognises the family without having verified that member individually.

**The report covers one input channel, and says so.** The `SCOPE:` line states that it covers settings resolved from the process environment - the `LATTICE_` variables plus the `DOTNET_` garbage-collector variables - together with the runtime facts stated as such (`Environment.ProcessorCount` and the collector's resolved mode, heap count, memory ceiling and pause total), and that it does **not** cover `LatticeOptions` configured in code through `ConfigureLattice` - `WalRetention` among them - nor any value supplied through some other channel. So a setting absent from the report is a setting outside its scope, not a setting proven unset. Read a silence that way and nothing else in the report has to be qualified by hand.

**Every value states where it came from, and a runtime fact is not a setting.** A value an operator supplied is marked `(DECLARED)`; a value nothing supplied is marked `(DEFAULTED, not declared)` and must not be read as configured; an observation such as `GC.Mode` or `GC.HeapCount` is marked `(RUNTIME FACT, not a declared setting)`, so nobody goes looking for a variable of that name. The distinction is load-bearing for the collector lines in particular: a resolved heap count of `6` says nothing about whether `DOTNET_GCHeapCount` was set, and the two lines together are what let you tell a declaration that was applied from one that was misread or ignored.

The set of keys the report treats as read is derived, not restated: the package publishes them as `RepoContextEnvironmentVariables`, whose `All` and `Prefixes` are built from the option classes' own constants, and the host folds that set into its own. A key added to an option class and published there is covered by the report without a second edit, which is what stops the two drifting apart.

Next: [Index source strategies](index-source-strategies.md). Contents: [Container quickstart](../container.md).
