docker-compose.tuning.yml
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at docker-compose-tuning-yml.md, and llms.txt lists every page.Part of RepoContextContainer source.
# Tuned local deployment overlay - TRACKED, and OPT-IN.
#
# Layer this over the base compose file explicitly:
#
# docker compose -f docker-compose.yml -f docker-compose.tuning.yml up -d --no-build
#
# It is deliberately NOT named docker-compose.override.yml. An override file is
# loaded AUTOMATICALLY and silently, so it changes what every reader of this
# directory gets from a plain `docker compose up -d`. This file changes nothing
# unless you name it, which is what makes it safe to track.
#
# WHY THIS FILE EXISTS (issue #2609). Every value below previously lived only in
# an untracked, gitignored docker-compose.override.yml on one developer's
# machine, and the instructions for operating the stack lived only in a wipeable
# agent-memory entry. Both were load-bearing for epic #2368's gate, and neither
# had review, history, or a diff. The rationale, the measurement behind each
# value, and the recovery procedure are in the runbook:
#
# docs/lattice.api.mcp.repocontext/local-deployment-runbook.md
#
# The runbook and this file are kept in step by
# LocalDeploymentRunbookHygieneTests, which resolves the two compose files
# with `docker compose config` and asserts the runbook's settings table
# enumerates exactly what the resolved document declares.
#
# WHAT THIS FILE IS NOT EVIDENCE OF. Agreement between this file and the runbook
# says nothing whatsoever about any running container: `docker compose up` reads
# the compose files in its OWN working directory, whichever checkout that is, and
# nothing in its output names a branch or a commit. Epic #2368's gate runs 1 and
# 2 both failed exactly there. To establish what a running container actually
# received, read it from the container:
#
# pwsh -File ./scripts/Assert-ContainerProvenance.ps1
#
# ---------------------------------------------------------------------------
# EVERY RESOURCE KNOB BELOW IS DERIVED, NOT TRANSCRIBED (issue #2779).
#
# This block used to read, in full:
#
# "Host these values were measured on: 16 logical CPUs, 55.7 GiB RAM.
# Do not copy them onto a different host without re-measuring. In particular
# DOTNET_PROCESSOR_COUNT below is a literal host core count, and a literal
# copied from someone else's deployment is the failure mode the base compose
# file's LATTICE_WAL_MAX_CONCURRENT_REPLAYS comment warns about at length."
#
# That warning was exactly right, which is why it is quoted rather than
# deleted. What changed is that it is now ENFORCED instead of advised. Every
# resource knob below is a `${VAR:?...}` reference with NO default, so this
# file cannot carry a literal from anyone's machine - including the one it was
# written on. The warning had been sitting directly above the literals it
# warned about for the whole of epic #2368, and was not enough on its own.
#
# WHAT THAT ENFORCEMENT DOES NOT REACH, stated because the paragraph above read
# for a while as though it did (issue #2863). `${VAR:?...}` errors when VAR is
# unset or empty. It does not, and cannot, inspect the value: compose
# interpolation has no value predicate, only presence ones. So it establishes
# that SOMETHING was supplied, never that what was supplied means what the
# supplier intended.
#
# That gap is not academic here, because every knob below has a falsy value in
# its valid syntax that means "ignore me" to its final consumer. `0` is
# non-empty, so it satisfies every guard on this page while selecting: no CPU
# limit and no memory limit at all from Docker; the host core count from ONNX
# Runtime and from the CLR's GC; and the runtime-derived WAL replay ceiling this
# overlay exists to pin. An operator who forgot to export a variable and one who
# pinned it to `0` on purpose produce byte-identical deployments, and the guard
# reports both as satisfied.
#
# No check can recover a distinction the encoding destroyed, so `0` is retired
# as a spelling. It is refused by the preflight below, and the deliberate
# "derive it" case is spelled `auto` on the two knobs whose parser we own
# (REPOCONTEXT_MAX_CONCURRENT_REPLAYS, EMBEDDER_INTRA_THREADS). The other five
# are read finally by Docker or the CLR, whose vocabulary is not ours to extend,
# so they take a positive value or nothing:
#
# pwsh -File ./scripts/Assert-TuningEnv.ps1
#
# New-TuningEnv.ps1 runs it over the .env it has just written, so the ordinary
# path is already covered and this is for a file edited by hand afterwards.
#
# WHAT THE LITERALS COST, so the next person does not restore one. The two
# mem_limit values summed to 17 GiB, so on a 16 GiB host this stack could not
# start at all, and nothing said why. Worse, DOTNET_PROCESSOR_COUNT: "16"
# overrode a cgroup-aware default and held the WAL replay concurrency gate at
# 16 permits against a 6.0-CPU quota - a 2.67x oversubscription, measured on
# the live deployment, and the deployment half of the root cause in #2692.
#
# Derive the values for THIS host and corpus:
#
# pwsh -File ./scripts/New-TuningEnv.ps1
#
# It writes .env in this directory. Read .env.example for what each variable
# means, what it is derived FROM, and which of them are policy rather than
# measurement - the distinction is stated per-variable and deliberately.
# ---------------------------------------------------------------------------
services:
repocontext:
# The base file declares `build:` and no `image:`, so without this pin the
# documented `up -d --no-build` step cannot resolve an image at all. The
# build-and-tag ladder that produces it (candidate-<sha> -> local, with the
# displaced tag kept as rollback-<timestamp>) is in the runbook.
image: repocontext-mcp:local
environment:
# ---- Scan cadence -------------------------------------------------
# The base file's cadence makes the periodic reconcile effectively
# continuous, which is the right default for a sample demonstrating that
# an edit shows up within seconds. It is far more than a long-lived local
# deployment needs, and it is the source of the continuous embedder load.
#
# The periodic deadlines below are counted in PASSES, not wall clock
# (interval / (reconcile + jitter), rounded up), so the reconcile interval
# and the jitter had to move together with them or the pass counts would
# not change as intended.
#
# TRADE-OFF, stated because it is real: an in-place content edit that does
# not bump its directory's modification time is now noticed within about
# an hour instead of about two minutes. An edit that DOES touch a
# directory mtime is still picked up on the next reconcile.
LATTICE_SELFINDEX_TICK_SECONDS: "30"
LATTICE_RECONCILE_INTERVAL_SECONDS: "60"
# Non-zero jitter stops passes across repositories phase-locking into
# simultaneous walks, which is what produced the observed load spikes.
LATTICE_RECONCILE_JITTER_SECONDS: "15"
# The full re-stat of every file, and the single heaviest operation.
LATTICE_FULL_WALK_INTERVAL_SECONDS: "3600"
# Two membership reads per indexed source, so on a converged repository
# this dominates a pass. Costs no healing latency: finding an actual gap
# forces an immediate in-pass scan regardless of this interval.
LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS: "3600"
# ---- Runtime sizing -----------------------------------------------
# Server GC with an EXPLICIT heap count (issue #2596).
#
# This was DOTNET_gcServer: "0", commented "one heap instead of sixteen".
# That traded away parallel collection to protect the memory limit,
# because Server GC would otherwise size its heap count from
# DOTNET_PROCESSOR_COUNT below, which is pinned to 16 for a completely
# unrelated reason (the WAL replay concurrency gate). One variable, two
# jobs, opposite requirements.
#
# The cost of that trade went unmeasured until gate run 2, whose log
# measures it directly: 283 whole-process silence gaps of 5s or more,
# longest 29.3s, totalling 20.9% of wall-clock, against a 30s Orleans
# request timeout. Every timed-out call (127 of 127) began executing
# within 0.5s of enqueue and then froze, so this is a stop-the-world
# pause rather than queueing contention. Orleans logged "Silo not running
# with ServerGC turned on" at startup throughout.
#
# GCHeapCount decouples the collector from PROCESSOR_COUNT, so the
# footprint concern that motivated "0" is addressed directly instead of
# by giving up parallel collection. It matched the cpus cap below, NOT
# the 16 that used to be pinned beneath it.
#
# Under #2779 that correspondence is no longer a coincidence maintained
# by hand: this and `cpus` below read the SAME derived value, so they
# cannot drift apart the way two independently-edited literals could.
# The original pairing was correct and is preserved by construction.
DOTNET_gcServer: "1"
DOTNET_GCHeapCount: "${REPOCONTEXT_GC_HEAP_COUNT:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
# ---- WAL replay concurrency ---------------------------------------
# DOTNET_PROCESSOR_COUNT: "16" USED TO BE HERE AS A PIN. THE PIN IS GONE
# (issue #2779) and is not coming back. What IS here, at the bottom of
# this block, is the same name declared as an UNSET pass-through - which
# is a different object and is explained there (issue #2931).
#
# Do not restore the pin. Its comment read, in full:
#
# "DO NOT REMOVE. This decouples the WAL replay concurrency gate from
# the CPU cap below. BPlusLeafGrain sizes that process-wide gate from
# Environment.ProcessorCount when the option is left non-positive,
# once, as a structural constant. Environment.ProcessorCount is
# cgroup-aware, so the cpus limit below would otherwise shrink the
# gate silently and unattributably. Pinning the reported count holds
# the gate at the value every previous field measurement on this box
# was taken against, while the cgroup quota still bounds actual CPU
# consumption.
#
# This is also why the embedder derives its own thread count from the
# cgroup quota rather than from the processor count (issue #2610):
# this variable overrides Environment.ProcessorCount and wins over the
# quota, so a service that copied this environment block would be
# oversubscribed."
#
# Every mechanical sentence in that is accurate, and the last paragraph
# correctly identifies the variable as an oversubscription hazard for any
# service that copies it. The justification for keeping it is what fails:
# it holds the gate "at the value every previous field measurement on
# this box was taken against" - preserving COMPARABILITY WITH PRIOR RUNS,
# while the thing held constant was the defect those runs were measuring.
# A comment can be meticulous about one axis and silent about another;
# this one reasons carefully about continuity and never about cost.
#
# THE COST, measured on the live deployment rather than argued:
# - the gate sized to 16 permits against a `cpus: 6.0` quota (2.67x);
# - `docker stats` bursting to 638/751/681% against a 600% quota;
# - `Configured WalMaterialiserMaxConcurrentReplays=0` in the gate's own
# log line, proving the supported knob was never set;
# - and 16 concurrent whole-window replays holding multi-MiB buffers,
# which is the deployment half of the heap exhaustion in #2692.
# BPlusLeafGrain.Activation.cs:154-172 (#2278) already names this exact
# shape as self-reinforcing, in the code this variable was overriding.
#
# AND THE COMPARABILITY IT BOUGHT IS WORTHLESS, which is the part that
# closes the argument rather than merely rebutting it. Every prior field
# measurement this pin was protecting was taken IN the 2.67x-oversubscribed
# condition. They are comparable to each other and to nothing that should
# ever be run again, so there is no longer a continuity argument to weigh
# against the cost - the baseline the pin preserved was itself the defect.
# If you are tempted to restore this variable for that good reason, that
# is the reason it does not hold.
#
# The gate's decoupling from the CPU cap is a REAL requirement and is
# still met - by the option that exists for it (#2279) instead of by a
# runtime-wide variable with two jobs. This one has a single job, is
# read by the library in preference to Environment.ProcessorCount, is
# refused at startup if unparseable, and - unlike the pin - does not also
# resize the GC, the thread pool, and every other ProcessorCount consumer
# in the process as a side effect.
#
# ---- The name, declared UNSET, so the value is attributable (#2931) ----
#
# Everything above is about the PIN and remains true. This is about the
# NAME, and the two were conflated by deleting the line outright.
#
# Deleting it did remove the pin. It also removed the variable, and with
# it any way to express the value without editing this tracked file. The
# cost landed on attribution rather than on tuning: run 9's predicate was
# amended to bind that run to run 8's grants INCLUDING
# DOTNET_PROCESSOR_COUNT=16, so that run 9 would be a code-only
# comparison. The variable was by then not merely different, it was
# ABSENT, and Environment.ProcessorCount was resolving from the cgroup.
# A replay-gate ceiling then moved 16 -> 6 in a step that also removed a
# pin, giving one movement two sufficient causes and no way to separate
# them. That is a measurement lost, not a value mis-set.
#
# DECLARED WITH NO VALUE, which is a compose idiom and not an omission.
# A null-valued mapping entry is resolved from the shell environment or
# .env, and is OMITTED FROM THE CONTAINER ENTIRELY when neither supplies
# it. Verified on this rig rather than assumed:
#
# unset -> `docker compose config` renders `null`, and in the
# container `${DOTNET_PROCESSOR_COUNT+x}` is empty,
# i.e. the variable is ABSENT, not empty-string;
# set in .env -> present with that value.
#
# So today's absent state is preserved BYTE FOR BYTE - no effective value
# moves by adding this line - while a controlled run can pin it without
# editing a tracked file, which is what #2931 asked for.
#
# IT IS STILL AN OVERSUBSCRIPTION HAZARD. The final paragraph of the old
# comment above is correct and applies unchanged to anyone who sets this:
# it overrides Environment.ProcessorCount process-wide and wins over the
# cgroup quota. Setting it is a deliberate act for a controlled
# comparison, and Assert-DeployManifest.ps1 reports the resolved count
# and WHERE IT CAME FROM on every deploy, so a set value cannot be a
# silent one and neither can an unset one.
DOTNET_PROCESSOR_COUNT:
LATTICE_WAL_MAX_CONCURRENT_REPLAYS: "${REPOCONTEXT_MAX_CONCURRENT_REPLAYS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it from this container's CPU grant, or set it by hand from .env.example. No default by design (issue #2779).}"
# Diagnostic for issue #2252, NOT tuning, and so deliberately left
# commented out rather than tracked as a default. Uncomment only while
# investigating vector-writer disposition counts, and record it in the
# runbook's local-only deltas section while it is on.
# Logging__LogLevel__Orleans.Lattice.Api.Mcp.RepoContext.RepoContextVectorWriter: "Debug"
# Bounds a runaway without starving normal operation. This service measured
# about 92% of one core in steady state before any of these limits, so the
# grant is headroom rather than a working limit.
#
# It is now also the SOURCE for DOTNET_GCHeapCount and for the WAL replay
# gate above: the derive script sizes all three from this one number, which
# is what the base compose file's LATTICE_WAL_MAX_CONCURRENT_REPLAYS
# comment asks for when it says to set the gate together with the CPU limit
# it derives from, in the same file.
cpus: "${REPOCONTEXT_CPUS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
# Sized above the measured steady-state plateau, not below it. See the
# runbook: a cap below the working set does not present as a resource event,
# because .NET sizes its heap hard limit from the cgroup limit and collects
# harder as it approaches it rather than being OOM-killed at it.
#
# That is precisely why this one must not be a literal (#2779): the failure
# mode of an undersized grant is a STORAGE fault, not a container kill, so
# a value copied from a larger host degrades the stack while `docker ps`
# still reports it healthy. The derive script sizes it from the corpus and
# clamps it to what the host can actually offer. It REFUSES when the host
# cannot offer even the floor, and when the clamp binds between the floor
# and the requirement it grants the ceiling and names the shortfall - it
# never grants less silently (#2832).
mem_limit: "${REPOCONTEXT_MEM_LIMIT:?unset - run ./scripts/New-TuningEnv.ps1 to derive it from this host and corpus, or set it by hand from .env.example. No default by design: a memory grant nobody measured is the defect this overlay exists to stop (issue #2779).}"
embedder:
environment:
# Workstation GC. Unlike the sibling service above, this one has no large
# managed heap worth collecting in parallel: its footprint is dominated by
# the resident ONNX model, which is native and outside the GC's control.
DOTNET_gcServer: "0"
# Pins the ONNX intra-op thread pool to the cgroup grant below.
#
# ONNX Runtime sizes that pool from HOST cores (16 here) and does not
# consult the cgroup quota, so under the 4.0-CPU grant the pool ran 4x
# oversubscribed and the kernel throttled it in 296 of 298 consecutive
# scheduling periods during vectorising. This is the third instance in
# this one deployment of a runtime sizing a pool from host cores under a
# fractional grant; see DOTNET_GCHeapCount and DOTNET_PROCESSOR_COUNT on
# the sibling service above.
#
# Issue #2610 derives this value from the cgroup automatically, so on an
# image built from that commit or later the pin only restates what the
# server would have chosen for itself. It is tracked anyway, because the
# deployment this file describes still runs a pre-#2610 image and because
# an explicit value is what the deployed configuration actually carries.
# Removing it is precisely how the derived path gets exercised, which has
# not happened yet on this box.
#
# Under #2779 it stays explicit for those reasons, but is no longer a
# literal: it is derived from the SAME number as the cgroup grant below,
# so the restatement is guaranteed to agree with what #2610 would pick
# rather than merely happening to. A literal `4` beside a `cpus` value
# that someone later edits is how the two silently stop matching.
EMBED_INTRA_THREADS: "${EMBEDDER_INTRA_THREADS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
# Leaves the balance of the host's cores for interactive work. The embedder
# sizes its ONNX intra-op thread pool from this grant, read from the cgroup
# filesystem (issue #2610), so changing this number changes the pool.
cpus: "${EMBEDDER_CPUS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
# The resident ONNX model needs room to work above its own footprint.
#
# Unlike the sibling service's grant, this one is WORKLOAD-derived, not
# corpus-derived: it is dominated by the model resident in the image, so it
# does not grow with the repository. The derive script therefore treats it
# as a near-constant and only clamps it against the host. It is still a
# variable because it is half of the sum that has to fit: the two
# mem_limits together used to demand 17 GiB on any host, which is why a
# 16 GiB machine could not start this stack at all (issue #2779).
mem_limit: "${EMBEDDER_MEM_LIMIT:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"