Table of Contents

Configuration

This page documents Orleans.Lattice.Api.Mcp.RepoContext, which is unreleased, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at configuration.md, and llms.txt lists every page.

Part of Container quickstart.

The host is configured entirely by environment variables. The common ones:

Variable Default Purpose
LATTICE_DURABILITY local The durability profile (local, postgres, azure; postgresql is accepted as an alias of postgres).
LATTICE_DATA_ROOT /data Root for all durable local state; must be a writable host mount.
LATTICE_MCP_PORT 8080 The MCP listener port. The health probes and the /metrics scrape endpoint are served on it too; it is the container's only application listener. A value that is not an integer in 1-65535 fails startup.
LATTICE_WORKSPACE_ROOT /workspace The read-only root that runtime-registered repositories must resolve under; a path escaping it is refused.
LATTICE_EMBEDDING_ENDPOINT http://localhost:9000 The separate embedding companion's base address. The embedding provider is always bound, so this repoints it at the companion rather than switching semantic search on; semantic search degrades to keyword ranking whenever that address cannot be reached. Must be an absolute URI or startup fails.
LATTICE_WAL_DIR / LATTICE_SQLITE_PATH under the data root Override the WAL directory or SQLite file path individually.
LATTICE_WAL_COMPACTION_MAX_DEAD_BYTES 0 (disabled) when unset; the sample compose sets 134217728 (128 MiB) An absolute ceiling, in bytes, on the dead (trimmed but not yet reclaimed) bytes a single file-WAL shard may hold before a compaction is forced. Compaction is otherwise triggered only by a ratio - dead bytes reaching half the file - which bounds waste relative to live data rather than absolutely, so a growing log roughly doubles in size before any space returns to the filesystem. The library ships this at 0 (disabled), leaving that ratio as the only trigger, and so does this host when the variable is unset; the sample compose file names a value because issue #3223 measured what the ratio alone does here, with the ratio declining at 200/1268 = 0.158 on every sweep, orleans_lattice_wal_compactions_total at zero on all three trigger tags, and physical WAL reaching 10174 MiB against the 8192 MiB ceiling the row below then carried. The gate is per shard, not per tree, so the tree-wide bound is this value times the shard count: 128 MiB over the 8 shards of repo-context-vector-index bounds dead at 1024 MiB - one eighth of the 8 GiB ceiling it was sized against, and one twentieth of the 20 GiB the sample declares today - where the ratio permits one half. Size it to FIRE against measured dead per shard - a ceiling above it changes nothing at all - and against the cost, which is write amplification of live/ceiling per byte reclaimed (here 1068/128 = 8.3x, against 1.0x for the ratio) paid as a synchronous whole-shard rewrite whose duration is set by live bytes and so does not shrink as the ceiling shrinks. Only 0 disables it; a non-zero value below the 64 KiB CompactionMinimumDeadBytes floor is not inert but maximally aggressive, and is refused at startup (issue #3210). It bounds only dead bytes, so it cannot by itself bring a tree under LATTICE_WAL_MAX_RETAINED_BYTES when retained alone already exceeds it. orleans_lattice_wal_compactions_total{trigger="ceiling"} reports when this trigger fires.
LATTICE_WAL_PIN_BUCKETS 8 How many persisted slots the WAL materialiser retention-floor pin state is split across, so an advancing floor rewrites a fraction of the pin blob rather than all of it. Accepts 1-256; 1 is the library's legacy single-slot write path. Widening self-migrates on activation and leaves the legacy slot intact, so reverting to 1 is a safe rollback that over-retains WAL rather than over-trimming it.
LATTICE_WAL_PIN_SHED_CEILING_SECONDS 120 Ceiling on how long one durable pin shard may shed steady-state pin reports continuously before one report is forced through regardless (issue #3310). The shed path is the only write path carrying an advancing CheckpointOffset, so an unbounded shed freezes the WAL GC durable offset floor and retained WAL grows without bound. Accepts 0-3600; 0 disarms the ceiling and is the exact rollback to the pre-fix behaviour, which over-retains WAL rather than over-trimming it. Forcing cannot overstate durability - the leaf clamps the reported offset to min(checkpoint, covered) before the report is built - so this bounds when a true value is published, never what. Setting it too low is the real hazard: it forces writes faster than the pin store can absorb and re-saturates the queue the shedding protects.
LATTICE_SETMANY_FANOUT_BUDGET_SECONDS 30 Budget for a single batch-write fan-out across shards, in seconds (issue #3384). A batch write whose keys span more than one shard scatters to every shard owning a key in the batch and awaits them all; without a budget one unresponsive shard holds the whole fan-out, and its caller, for as long as it stays unresponsive. With the budget, a fan-out still unsettled when it expires is refused with LatticeSaturatedException (source SetManyFanOut) and rolls nothing back: shards that already committed stay committed and branches still in flight keep running. A batch whose keys all land on one shard is a single call this budget does not bound. Accepts 0-3600. Read the zero the opposite way round to LATTICE_WAL_PIN_SHED_CEILING_SECONDS above: there a zero disarms a bound, whereas here 0 resolves to an infinite budget - the pre-fix behaviour, under which a fan-out waits indefinitely - and is the exact rollback. It deliberately does not mean TimeSpan.Zero, which the library refuses because a zero budget would refuse every fan-out and turn a rollback into an outage. Setting it too low is the real hazard: a budget below the time a healthy fan-out legitimately takes converts ordinary slowness into failed writes, so size it against measured fan-out duration and not against a latency target.
LATTICE_WAL_MAX_RETAINED_BYTES 1073741824 (1 GiB) when unset; the sample compose sets 21474836480 (20 GiB) An advisory per-tree ceiling on the WAL bytes held on disk. On reaching it, each garbage-collection pass lowers its trim frontier toward 80% of the ceiling - but only within the frontier that was already safe to trim, so it never blocks a write, never trims past a live consumer cursor, and never loses a record (what it trims has already been materialised into leaves). The library ships this disabled because a correct value is a fraction of the volume the WAL lives on, which a library cannot know; this container declares its own volume, so it names one. Left disabled, every tree reports orleans_lattice_wal_gc_backlog_bytes_unavailable_total{reason="policy_disabled"} and reclaims nothing, so on-disk occupancy only ever grows. Costs one physical-size probe per WAL partition per pass. Accepts an integer of at least 67108864 (64 MiB), or 0 to restore the library default; anything else refuses startup. Size it at no less than twice the largest tree's logical working set: the ceiling is compared against physical bytes summed over a tree's partitions, and WAL compaction only reclaims once a shard is 50% dead, so steady-state physical settles at about 2x logical and a tighter ceiling can never be satisfied or disarmed (issue #3234). Raise it on a deployment indexing a substantially larger corpus. This container was resized from 4 GiB to 8 GiB by issue #3242 as headroom, not in response to a measured overrun: its vector-index tree grew from the 1077 MB logical working set the 4 GiB was calibrated against to about 1672 MB, moving the 2x floor from 2154 MB to about 3344 MB and consuming most of the margin that calibration had left. An earlier revision of this row stated the 4096 MB ceiling had already been overrun; that figure came from reading a Prometheus summary series as an instantaneous value when sum/count on a summary is a lifetime mean, and it is withdrawn. Measure a tree's logical working set with the orleans_lattice_storage_wal_bytes gauge, which is sampled at scrape time; never derive it by differencing a summary. A ceiling calibrated against a measured working set acquires an expiry date the moment that workload can grow, so re-check it against the tree rather than treating a past calibration as durable. Raised from 8 GiB to 20 GiB by issue #3094 after wal_gc_ceiling_unsatisfiable fired continuously on this container. The 8 GiB ceiling firing at all implies a logical working set above 4 GiB; physical was measured at 5.864 GiB and logical is bounded by physical, so the 2x floor lay somewhere in 8.0 to 11.7 GiB - at or above the ceiling, which is why the policy could never be satisfied or disarmed. 20 GiB clears the top of that range with about 1.7x headroom. After the raise the counter stayed at 0 and on-disk occupancy stopped climbing monotonically, oscillating within about 25 MB instead.
LATTICE_POSTGRES_CONNECTION_STRING / LATTICE_AZURE_STORAGE_CONNECTION_STRING (unset) Required by the postgres / azure profiles.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR (unset, feature off) Directory the agent-memory archive is exported to and restored from. Point it at a path outside the data volume - a bind mount rather than a named volume - or it dies with the state it exists to outlive. See Memory durability.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_INTERVAL_SECONDS 300 Export cadence; the size of the window an ungraceful stop loses. Positive values below 30 are raised to 30, and values above the longest delay a timer can wait (about 49.7 days) are lowered to it; a zero, negative or unparseable value, or one too large to represent as a duration (for example 1e20), falls back to 300.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_RESTORE auto auto restores into a store holding no memory, or into one whose restore-state marker records that an earlier restore was left partial (so a half-imported tree is healed rather than mistaken for a populated one); always restores every start; off never restores. none and false are accepted for off and on-empty for auto; any other value falls back to auto.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_STOP_TIMEOUT_SECONDS 20 Budget for the final export during a graceful stop. A positive value is clamped to 1-60; a zero, negative or unparseable value falls back to 20. It shares the stop grace period with the drain, so it is deliberately a fraction of it.
LATTICE_REPOCONTEXT_HEAP_ADMISSION_OVERRIDE (unset) Suppresses the startup refusal described under Measured requirement and startup admission. Its value must be the exact recorded byte count being disbelieved, which the refusal message prints and lattice_repocontext_heap_insufficient_limit_bytes publishes; any other value, including true or 1, is ignored and the refusal stands. It is deliberately not a boolean, so it cannot be set once and left to suppress correct refusals for ever, and it stops matching by itself once a larger exhaustion is recorded.

A profile is a preset, not a straitjacket: each store it selects can be overridden on its own, and the remaining variables name the cluster and the embedding space. An unrecognised value for any of the four provider variables fails startup rather than falling back silently:

Variable Default Purpose
LATTICE_WAL_PROVIDER azure under the azure profile, otherwise file Selects the WAL provider on its own. Accepts file or azure (azuretable).
LATTICE_GRAIN_STORAGE the profile's store (sqlite / postgres / azure) Selects the grain-storage provider on its own. Accepts sqlite, postgres (postgresql), or azure (azuretable).
LATTICE_REMINDERS the profile's store Selects the reminders provider on its own; same accepted values as the grain store.
LATTICE_CLUSTERING azure under the azure profile, otherwise localhost Selects the clustering provider. Accepts localhost (local) or azure.
LATTICE_AZURE_WAL_TABLE RepoContextWal The Azure Table the WAL writes to when the Azure WAL provider is selected.
LATTICE_EMBEDDING_MODEL nomic-ai/nomic-embed-text-v1 The embedding model id requested from the companion.
LATTICE_EMBEDDING_DIMENSION 768 The embedding vector dimension; must match the model the companion serves. A non-positive value fails startup.
LATTICE_CLUSTER_ID repo-context The Orleans cluster id.
LATTICE_SERVICE_ID repo-context The Orleans service id.

Selecting any Azure-backed store without LATTICE_AZURE_STORAGE_CONNECTION_STRING refuses to start rather than silently degrading durability. Changing LATTICE_EMBEDDING_MODEL or LATTICE_EMBEDDING_DIMENSION is a new embedding space, so it builds a wholly separate approximate index under its own prefix - which is what LATTICE_REPOCONTEXT_ANN_INDEX_RECLAMATION below then retires the superseded one for.

The background reconcile cadence (see Background reconcile and change detection) is tuned by six further variables. The two periodic reconcile deadlines - the full walk and the embedding gap scan - are declared in wall clock but counted in reconcile passes: each is divided by the widest scheduled reconcile spacing (LATTICE_RECONCILE_INTERVAL_SECONDS plus LATTICE_RECONCILE_JITTER_SECONDS), rounded up, and clamped to at least one pass. That is what makes them hold on a large repository, where a pass routinely takes longer than its own scheduled spacing and a wall-clock deadline would be past on arrival every single time. The coverage-digest audit is not pass-counted; see its row:

Variable Default Purpose
LATTICE_SELFINDEX_TICK_SECONDS 15 How often each repository's self-index grain ticks; the reconcile cannot fire more often than this. A value longer than a grain timer can wait (about 49.7 days) ticks at that ceiling.
LATTICE_RECONCILE_INTERVAL_SECONDS 900 Base interval between periodic content reconciles. A small value (with zero jitter) makes the reconcile effectively continuous, bounded by the tick.
LATTICE_RECONCILE_JITTER_SECONDS 300 Maximum extra random interval added on top of the reconcile interval to desync repositories.
LATTICE_FULL_WALK_INTERVAL_SECONDS 3600 How often a reconcile is forced to ignore the directory-modification-time prune cache and stat every file, bounding how stale an in-place content edit can be. Counted in passes: at the shipped defaults it is 3 reconciles, so 2 in every 3 prune. Set it at or below one reconcile spacing and it degenerates to 1 pass, meaning every reconcile walks in full and pruning never engages.
LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS 1200 How often a reconcile re-checks every content-unchanged file for an embedding gap - a file whose structural record is committed but whose vector never landed. Detection reads the per-page vector-coverage digest: a fixed 257 rows on a tree of its own, whatever the corpus size. Counted in passes, and at the shipped defaults this is deliberately 1 - the shortest window the scheduler can express, so a gap is found on the next reconcile rather than up to four hours later. Because it is 1, the periodic cadence has no effect at the defaults: the scan runs on every pass whatever the coverage verdict says, and the host logs a startup warning saying so. Set it above one reconcile spacing (for example 2400 at the defaults) for the cadence to engage. The former 14400 default was a ration against a probe that cost two membership reads per indexed source; that cost is gone, and so is the reason for the ration.
LATTICE_COVERAGE_DIGEST_AUDIT_INTERVAL_SECONDS 86400 How often the self-index grain's gap sweep re-derives the vector-coverage digest exhaustively from the membership tree, rather than trusting the incremental maintenance the write path performs. Unlike the two rows above it is applied in wall clock, not reconcile passes: the audit falls due this long, plus up to a minute of jitter, after the previous one, and runs when the sweep next checks. The effective-configuration report currently prints it rounded up to whole reconcile spacings, which is not the cadence the sweep applies. This is the O(sources) read the gap scan used to be, and it is what bounds digest drift - from an interrupted write, or from a self-heal that reset the membership tree under a surviving digest. Raising it widens the window in which a drifted digest can mask a gap; lowering it re-imports the cost this design removed.

These interval variables are a matched set. LATTICE_FULL_WALK_INTERVAL_SECONDS and LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS are wall-clock values that are converted once into pass counts by dividing by the reconcile spacing (LATTICE_RECONCILE_INTERVAL_SECONDS plus LATTICE_RECONCILE_JITTER_SECONDS). Changing the reconcile interval therefore silently re-denominates both. Raising it far enough that the full-walk interval floors to a single pass switches directory-modification-time pruning off entirely - no error, and the prune cache is written on every run but never read. If you raise the reconcile interval, restate both. The host logs the derived pass counts next to the configured seconds at startup (full walk 120 s = 24 pass(es) ...; pruning can engage: True), and warns when the arithmetic has disabled pruning or has collapsed the embedding gap scan to every pass, so the conversion never has to be worked out by hand.

Two further variables tune the indexing role and per-file token counting, and two select the semantic-retrieval path and size the vector cache:

Variable Default Purpose
LATTICE_REPOCONTEXT_INDEXING_ROLE hub The cluster's indexing role: hub (the authoritative indexer that walks, reconciles, prunes, and re-embeds) or spoke (a read-only replica whose index pass is inert). An absent or unrecognised value falls back to hub.
LATTICE_REPOCONTEXT_TOKENIZER o200k The BPE tokenizer profile the per-file token counter uses: o200k (OpenAI o200k_base) or cl100k (OpenAI cl100k_base). An absent or unrecognised value falls back to o200k.
LATTICE_REPOCONTEXT_SEMANTIC_RETRIEVAL approximate Which semantic retrieval path is bound: approximate routes semantic search through the persisted approximate nearest-neighbour index (bounded recall, sub-linear query cost, survives a restart), and exact routes it through the complete-recall brute-force scan instead, whose cost is proportional to the corpus. An absent or unrecognised value falls back to approximate. A host set to exact maintains no approximate index at all, so the build coordinator below is inert for it. Documented in full under Semantic search.
LATTICE_VECTOR_CACHE_TTL_SECONDS 30 How long (in seconds) a warm decoded-vector candidate set is trusted before it is re-gathered from the store; 0 disables the cache.

Four variables govern the adaptive indexing pacer (issue #3447), the silo-wide controller that smooths the embedding drain so a large pass no longer lands as one sustained spike. It sits in front of every embedding batch and never changes what a pass embeds, only when each batch runs: a failed or slow batch (over 250 ms and over 2.5x the learned per-passage latency baseline), a throttled vector tree, or GC memory load at its high-load threshold doubles an inter-batch delay from 250 ms up to the ceiling, and each clean batch walks it back down. Only a batch at least half the size of the largest seen is judged slow or moves the baseline, so a small remainder batch is neither misread as congestion nor drags the baseline down, and the baseline is never lowered by a batch that ran under a backoff delay. A vector tree that stays throttled for 60 seconds while the delay is held at its ceiling is not moved by indexing (a materialiser drain-lag minimum is the usual cause), so it stops counting as congestion until it clears or worsens; the WAL still paces throttled appends itself (issue #3456). A saturated vector tree holds the next batch for up to 30 seconds, an in-flight repocontext_search or repocontext_context call holds it for up to 2 seconds, and after each work slice the drain rests. While a foreground request is open or the drain is backing off, the approximate-index build defers its slices (at most 15 phase ticks in a row) and the self-index grain postpones a due coverage-digest audit (at most 12 sweeps in a row), so both are slowed and never starved. What the pacer is doing is reported as pacing on repocontext_index_status; see Adaptive pacing.

Variable Default Purpose
LATTICE_REPOCONTEXT_PACING true Whether the pacer is on. false restores back-to-back batches with no delay, yield, rest, or deferral. An absent or unrecognised value falls back to true.
LATTICE_REPOCONTEXT_PACING_SLICE_SECONDS 60 How long the drain works before it rests. 0 switches the duty cycle off (batches are still delayed under congestion).
LATTICE_REPOCONTEXT_PACING_REST_SECONDS 5 How long the drain rests after each work slice. 0 switches the duty cycle off.
LATTICE_REPOCONTEXT_PACING_MAX_DELAY_SECONDS 5 The ceiling on the congestion-driven inter-batch delay.

The next variables are the kill switches for the approximate index's own housekeeping, which default on, and the cadence of its build sweep. They are documented in full under Scheduling the approximate index build:

Variable Default Purpose
LATTICE_REPOCONTEXT_ANN_INDEX_SCHEDULING true Whether the approximate index build is scheduled by its durable, reminder-anchored coordinator - which is what lets a restored volume converge to a serving index with no client traffic at all, and what resumes a build interrupted by a process death. Set false and no index is built at all: every semantic query is answered by the exact scan with complete recall, within the bounds described under When the exact fallback is declined. An absent or unrecognised value falls back to true.
LATTICE_REPOCONTEXT_ANN_INDEX_RECLAMATION true Whether an index that has just reached Ready retires the sibling prefixes of its own repository whose embedding-space fingerprint is no longer live. A model or dimension change otherwise leaves the previous index resident forever. Set false to keep a superseded space for a deliberate roll-back. An absent or unrecognised value falls back to true.
LATTICE_REPOCONTEXT_ANN_SWEEP_INTERVAL_SECONDS 900 How often the build sweep re-arms every registered repository's coordinator. Floored at 60 seconds: a shorter value is raised to the floor, and the startup line says so rather than leaving the setting to look ignored.

The sweep cadence is deliberately not part of the matched set above. It used to be: the sweep took its interval from LATTICE_RECONCILE_INTERVAL_SECONDS, so raising that variable to quiesce walk load - a reasonable action, with nothing in its name to suggest otherwise - throttled index arming by the same factor. That is worse than a slow sweep. Two things arm a coordinator, this sweep and the self-index grain finishing a vectorising pass; a converged repository whose index was never built has no vectorising pass to finish, so the sweep is its only arming path, and the vectorising pass was paced by the reconcile interval too. Raising it did not slow one path of two, it slowed the only two there are. The index then serves nothing while the retrieval counter records state="bootstrapping", which at the metric is indistinguishable from a genuine index defect. LATTICE_REPOCONTEXT_ANN_SWEEP_INTERVAL_SECONDS defaults to 900 seconds, which is the reconcile interval's own default, so a host that configures neither variable sweeps at exactly the cadence it always did.

Five more variables bound how long the approximate index may hold its build coordinator's turn while it opens and ingests. An absent or malformed value falls back to the default:

Variable Default Purpose
LATTICE_REPOCONTEXT_ANN_OPEN_SLICE_BUDGET_SECONDS 5 Wall-clock ceiling on one attempt to open (restore) the durable index. A stopped attempt banks what it walked and the next continues past it, so this slices one long open into short ones. 0 removes the bound. A value above about 49.7 days (the longest wait a timer accepts) is held at that ceiling.
LATTICE_REPOCONTEXT_ANN_OPEN_SLICE_MAX_EXTENSIONS 6 How many further open-slice budget periods an open slice that has banked nothing may take before the budget fires anyway. 0 reproduces the elapsed-only bound.
LATTICE_REPOCONTEXT_ANN_INGEST_SLICE_BUDGET_SECONDS 5 Wall-clock ceiling on one ingest slice of the build, so the coordinator's keep-alive reminder and arming calls are answered while a build runs. 0 removes the bound, leaving the 4,096-vector slice batch as the only one. A value above about 49.7 days (the longest wait a timer accepts) is held at that ceiling.
LATTICE_REPOCONTEXT_ANN_OPEN_MAX_CONSECUTIVE_REFUSALS 12 Consecutive admission refusals after which the open declares itself terminally saturated. Declaring does not stop retrying. 0 removes the count bound.
LATTICE_REPOCONTEXT_ANN_OPEN_REFUSAL_TERMINAL_SECONDS 600 How long an unbroken run of admission refusals may last before the same terminal state is declared, whichever bound is reached first. 0 removes the elapsed bound.

Three further variables bound resources whose defaults come from a runtime fact or a fixed assumption rather than from the deployment's real limit, so a constrained container can state the limit it actually has:

Variable Default Purpose
LATTICE_WAL_MAX_CONCURRENT_REPLAYS 0 (defer to the library) The per-silo ceiling on concurrent activation-time leaf WAL replays. Each permit admits one whole-readable-window replay holding multi-MiB buffers, so although the default ceiling is derived from CPU figures the gate behaves as a memory admission gate: the resource a reactivation storm exhausts first is managed memory, not CPU (issue #2784). 0 defers to the library, which sizes the gate from the lesser of Environment.ProcessorCount and the container's enforced cgroup CPU grant. Accepts auto or an integer 0-256; anything else fails startup rather than being silently ignored. auto defers to exactly the same derivation as 0, and exists so that a deployment can state that it meant to derive - 0 reads as a pinned zero to anyone scanning the file, which is not what it does.
LATTICE_MAX_LOCK_LEASE_SECONDS 1800 The ceiling this host clamps every named-lock lease to, including the claim leases agents take through repocontext_claim. It bounds how long a crashed holder can pin an item while still covering a full build-and-test cycle. Accepts 30-7200; anything else fails startup.
LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD 120s when unset; this sample declares 240s The container grace period this deployment grants between SIGTERM and SIGKILL, declared to the process that has to fit inside it. The host derives its shutdown budget from it (75% of the grant, or all but a two-second unwind reserve, whichever is smaller), so the sample's declared 240s yields a 180s budget and the 120s default yields 90s. Accepts a positive number of seconds up to 3600, with an optional s suffix; compound forms such as 1m30s are refused rather than misread, and so is a grant of two seconds or less, which leaves no budget once the reserve is taken. It must equal the stop_grace_period on the same service - see Where the budget comes from.

You no longer need to set the replay ceiling beside your CPU limit. Environment.ProcessorCount honours a container CPU quota only while DOTNET_PROCESSOR_COUNT does not override it, and that variable takes precedence over the quota-derived value. A container granted 6 CPUs whose environment also carries DOTNET_PROCESSOR_COUNT=16 therefore used to size this gate at 16, not 6, with nothing inside the process able to tell the difference - the condition this sample's own compose file predicts, measured live at a 2.67x oversubscription. The library now reads the enforced cgroup quota itself and takes the lesser of the two figures, so the default follows the grant without an operator declaring anything, and it does so adaptively rather than being pinned to whatever the grant was on the day it was written (issue #2816). Pinning this variable is still honoured and still supersedes both figures, but it is now an override rather than a requirement, and a pinned value has the usual hazard of silently diverging from a cpus limit that later changes. The host logs the resolved ceiling once at startup alongside the configured option, the Environment.ProcessorCount the runtime reported, and the CPU grant it read, so the effective figure and any disagreement between the two inputs can be read off the log instead of inferred from the host's vCPU count.

An opt-in family of LATTICE_REPOCONTEXT_GIT_* variables switches a repository from the mounted workspace to a git remote; see Index source strategies.

Garbage collection on a multi-GiB heap

This host's steady-state working set is measured in GiB, and Workstation GC is the wrong collector at that size. Workstation GC collects a single heap and its blocking gen2 phases are effectively single-threaded, so one collection walks the whole heap on one thread with every other thread in the process suspended. A deployment of this host at about 11 GiB resident, held on Workstation GC by an explicit DOTNET_gcServer: "0" in its compose overlay, had the runtime attribute a 252 second pause to the collector, against a 30 second Orleans request timeout.

Variable Default Purpose
DOTNET_gcServer unset, which runs Server GC: the host is built with the ASP.NET Core Web SDK, which writes System.GC.Server: true into its runtimeconfig.json 1 selects Server GC, which collects several heaps in parallel; 0 forces Workstation GC.
DOTNET_GCHeapCount one heap per processor Bounds how many heaps Server GC creates. Read as hexadecimal - see below. Inert under Workstation GC (DOTNET_gcServer=0); because this host defaults to Server GC, it applies whenever DOTNET_gcServer is unset or 1.

Neither is set in the base compose file, because the right heap count is a property of your CPU grant rather than of any file in this repository. Left unset, the base sample therefore runs Server GC with a heap count the collector chooses for itself; the effective-configuration report below states which (GC.Mode, GC.HeapCount). The sample's tuning overlay (docker-compose.tuning.yml) does set both on the repocontext service - DOTNET_gcServer: "1", and a DOTNET_GCHeapCount it refuses to start without - taking the heap count from the CPU grant that scripts/New-TuningEnv.ps1 derives for that service on this host.

DOTNET_PROCESSOR_COUNT cannot double as the heap count. Server GC sizes its heap count from the processor count, which is exactly what DOTNET_PROCESSOR_COUNT overrides - the same variable the replay ceiling above discusses. That variable may legitimately be pinned above the container's CPU grant for the thread pool's sake, and reusing it as a heap count would then create one heap per phantom processor on a heap already near its ceiling. One variable, two jobs, opposite requirements. DOTNET_GCHeapCount separates them: derive it from the container's actual CPU grant, independently of DOTNET_PROCESSOR_COUNT, and leave the processor count to the pools that genuinely want it.

Write the heap count in hexadecimal. The collector reads its numeric environment variables as hex, while the same settings in runtimeconfig.json are decimal. DOTNET_GCHeapCount=10 therefore asks for 16 heaps and =16 asks for 22, silently and with no error. Values below 10 read identically either way, which is what makes this easy to miss on a small box and then get wrong on a large one. Prefer an explicit 0x prefix.

Verify rather than assume. The effective-configuration report states GC.Mode, GC.HeapCount (the figure the collector resolved, not the one declared), the process's memory ceiling, and the accumulated GC.GetTotalPauseDuration, and it raises a GC HAZARD warning when the process runs Workstation GC against a large ceiling, when a heap count is declared under Workstation GC and is therefore inert, or when the resolved heap count disagrees with the number that was written. Grep the log for GC HAZARD.

The claim is narrow on purpose. Server GC with a bounded heap count removes the class of pause that is multi-minute, process-wide, and attributed to the collector by the runtime itself. It is not a general remedy for stalls: measurement of the same container found collector pauses accounted for under a third of long-silence time and did not explain its largest timeout burst at all. A stall the runtime does not attribute to the collector needs its own diagnosis, and GC.GetTotalPauseDuration is the quantity to reach for rather than gaps between log timestamps.

Thread pools on a CPU-limited container

The collector is not the only pool sized from a number that a CPU limit does not constrain. The embedder service in the same sample sizes its ONNX Runtime intra-op pool - the threads that parallelise a single inference - and its own default reads the host core count while ignoring the cgroup quota entirely.

Variable Default Purpose
EMBED_INTRA_THREADS derived from the enforced cgroup CPU quota Sizes the ONNX Runtime intra-op thread pool. Accepts auto or a non-negative integer; 0 hands the decision back to ONNX Runtime, and auto derives from the grant deliberately. Leaving it unset also derives. A value that is present but unusable - a negative number, a non-number, a typo of the token - fails startup rather than deriving, because absence and a typo are different facts about a deployment and deriving from both makes them indistinguishable (issue #2887).

It is not set in the base compose file, for the same reason the heap count is not: the right value is a property of your CPU grant rather than of any file in this repository. Left unset, the embedder reads the quota itself, which is the recommended configuration. The tuning overlay does declare it, and refuses to start without it, using the embedder CPU grant that scripts/New-TuningEnv.ps1 derives for this host.

The cost of getting it wrong is worse than proportional. Measured on a 4.0-CPU grant (cpu.max = "400000 100000") on a 16-core host: an intra-op pool of 16, a 4x oversubscription, with the kernel throttling 296 of 298 consecutive scheduling periods and the pool accumulating 346.3 CPU-seconds stalled against 118.8 CPU-seconds run. The arithmetic behind that ratio is elementary once written down: sixteen threads drain a 400ms quota in 400/16 = 25ms of wall time and are then frozen for the remaining 75ms of the period, predicting 75:25 = 3.0 stalled per unit run against 2.91 measured, within 3%. Throughput does not merely fall by the oversubscription ratio, because ONNX Runtime synchronises its intra-op threads at every operator boundary and a transformer inference crosses hundreds of them; a freeze landing mid-barrier stalls the whole operator rather than one thread. The observed embedding rate was 1.8 files per minute, projecting roughly 77 hours for a single 8,315-file checkout.

DOTNET_PROCESSOR_COUNT cannot double as the thread count either. This is the same variable, doing a third job with a third set of requirements. Whenever it is set on the repocontext service - the tuning overlay declares it there by name only, so it is absent unless your environment or .env supplies a value - it overrides Environment.ProcessorCount and wins over the quota, and copying that service's environment block onto the embedder - an entirely ordinary thing to do - would silently restore the oversubscription. The embedder reads /sys/fs/cgroup/cpu.max directly and is immune to it, and the WAL replay gate now does the same (issue #2816); this variable remains the hazard for every pool that does not. When the two disagree the embedder logs a CPU GRANT MISMATCH warning naming both figures and the resulting factor, because a process that believes it has sixteen processors under a four-CPU grant will oversubscribe every pool sized from that belief.

Deriving it, if you choose to declare it. Use the container's actual CPU grant, rounded up: cpus: "4.5" becomes 5. That matches what .NET itself reports for the same limit, so the declared pool never disagrees with the runtime in the unsafe direction. Do not use the host core count, and do not use DOTNET_PROCESSOR_COUNT.

A declared value does not follow the grant. If you later change cpus and leave EMBED_INTRA_THREADS pinned, the pair silently diverges and nothing in the container will object - the number is no longer wrong in a way any single file reveals. That is the entire hazard of pinning one, and it is why the derived default is recommended. The same reasoning is why LATTICE_WAL_MAX_CONCURRENT_REPLAYS is no longer something the base sample asks you to declare; the tuning overlay still pins it, to the CPU grant scripts/New-TuningEnv.ps1 derives, and refuses to start without one. If you do pin either, keep it beside the cpus limit so the two are checkable against each other, and prefer auto over the derived default's numeric equivalent when what you mean is "derive": auto reaches the same number while recording that the derivation was chosen rather than inherited.

Verify rather than assume. The embedder states its resolved intra-op count once at startup, marked DECLARED when an operator supplied it and DERIVED when it did not, so the effective figure can be read off the log rather than inferred from this file. Reading a compose file tells you what was written; only the log tells you what the process resolved.

Reading the effective configuration off the log

The container's real settings arrive partly from files this repository does not track - the per-host .env that supplies the tuning overlay's derived values, and any machine-local override - so reading this repository does not tell you what a running process resolved. The host therefore states its own resolved configuration once at startup, on the Repository-context effective configuration: prefix, and that report supersedes any file when the two disagree:

  • a summary line counting the settings, how many differ from the host default, how many were not declared at all, and how many supplied LATTICE_ variables nothing reads;
  • one line per setting, carrying the value this process resolved, marked [OVERRIDDEN...] when it differs from the host default;
  • a SCOPE: line, described below;
  • one line per supplied variable recognised only by a family prefix (a LATTICE_REPOCONTEXT_GIT_* member), with its value withheld;
  • a warning per supplied LATTICE_ variable that nothing in this host binds;
  • a warning per hazardous garbage-collector configuration, prefixed GC HAZARD.

The same reporter then states, at warning level and on its own Repository-context memory durability: prefix, where agent memory lives and what does and does not protect it; see Memory durability.

Grep the log for SUPPLIED BUT NOT READ to find a variable an operator set that never reaches anything - the silent failure that motivated the report. Values are printed through an allowlist, so a key that is not classified as safe to print renders as <redacted: unclassified> rather than leaking; a variable matched only by a prefix renders as <withheld: matched by prefix only>, because the host recognises the family without having verified that member individually.

The report covers one input channel, and says so. The SCOPE: line states that it covers settings resolved from the process environment - the LATTICE_ variables plus the DOTNET_ garbage-collector variables - together with the runtime facts stated as such (Environment.ProcessorCount and the collector's resolved mode, heap count, memory ceiling and pause total), and that it does not cover LatticeOptions configured in code through ConfigureLattice - WalRetention among them - nor any value supplied through some other channel. So a setting absent from the report is a setting outside its scope, not a setting proven unset. Read a silence that way and nothing else in the report has to be qualified by hand.

Every value states where it came from, and a runtime fact is not a setting. A value an operator supplied is marked (DECLARED); a value nothing supplied is marked (DEFAULTED, not declared) and must not be read as configured; an observation such as GC.Mode or GC.HeapCount is marked (RUNTIME FACT, not a declared setting), so nobody goes looking for a variable of that name. The distinction is load-bearing for the collector lines in particular: a resolved heap count of 6 says nothing about whether DOTNET_GCHeapCount was set, and the two lines together are what let you tell a declaration that was applied from one that was misread or ignored.

The set of keys the report treats as read is derived, not restated: the package publishes them as RepoContextEnvironmentVariables, whose All and Prefixes are built from the option classes' own constants, and the host folds that set into its own. A key added to an option class and published there is covered by the report without a second edit, which is what stops the two drifting apart.