Table of Contents

docker-compose.tuning.yml

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at docker-compose-tuning-yml.md, and llms.txt lists every page.

Part of RepoContextContainer source.

# Tuned local deployment overlay - TRACKED, and OPT-IN.
#
# Layer this over the base compose file explicitly:
#
#   docker compose -f docker-compose.yml -f docker-compose.tuning.yml up -d --no-build
#
# It is deliberately NOT named docker-compose.override.yml. An override file is
# loaded AUTOMATICALLY and silently, so it changes what every reader of this
# directory gets from a plain `docker compose up -d`. This file changes nothing
# unless you name it, which is what makes it safe to track.
#
# WHY THIS FILE EXISTS (issue #2609). Every value below previously lived only in
# an untracked, gitignored docker-compose.override.yml on one developer's
# machine, and the instructions for operating the stack lived only in a wipeable
# agent-memory entry. Both were load-bearing for epic #2368's gate, and neither
# had review, history, or a diff. The rationale, the measurement behind each
# value, and the recovery procedure are in the runbook:
#
#   docs/lattice.api.mcp.repocontext/local-deployment-runbook.md
#
# The runbook and this file are kept in step by
# LocalDeploymentRunbookHygieneTests, which resolves the two compose files
# with `docker compose config` and asserts the runbook's settings table
# enumerates exactly what the resolved document declares.
#
# WHAT THIS FILE IS NOT EVIDENCE OF. Agreement between this file and the runbook
# says nothing whatsoever about any running container: `docker compose up` reads
# the compose files in its OWN working directory, whichever checkout that is, and
# nothing in its output names a branch or a commit. Epic #2368's gate runs 1 and
# 2 both failed exactly there. To establish what a running container actually
# received, read it from the container:
#
#   pwsh -File ./scripts/Assert-ContainerProvenance.ps1
#
# ---------------------------------------------------------------------------
# EVERY RESOURCE KNOB BELOW IS DERIVED, NOT TRANSCRIBED (issue #2779).
#
# This block used to read, in full:
#
#   "Host these values were measured on: 16 logical CPUs, 55.7 GiB RAM.
#    Do not copy them onto a different host without re-measuring. In particular
#    DOTNET_PROCESSOR_COUNT below is a literal host core count, and a literal
#    copied from someone else's deployment is the failure mode the base compose
#    file's LATTICE_WAL_MAX_CONCURRENT_REPLAYS comment warns about at length."
#
# That warning was exactly right, which is why it is quoted rather than
# deleted. What changed is that it is now ENFORCED instead of advised. Every
# resource knob below is a `${VAR:?...}` reference with NO default, so this
# file cannot carry a literal from anyone's machine - including the one it was
# written on. The warning had been sitting directly above the literals it
# warned about for the whole of epic #2368, and was not enough on its own.
#
# WHAT THAT ENFORCEMENT DOES NOT REACH, stated because the paragraph above read
# for a while as though it did (issue #2863). `${VAR:?...}` errors when VAR is
# unset or empty. It does not, and cannot, inspect the value: compose
# interpolation has no value predicate, only presence ones. So it establishes
# that SOMETHING was supplied, never that what was supplied means what the
# supplier intended.
#
# That gap is not academic here, because every knob below has a falsy value in
# its valid syntax that means "ignore me" to its final consumer. `0` is
# non-empty, so it satisfies every guard on this page while selecting: no CPU
# limit and no memory limit at all from Docker; the host core count from ONNX
# Runtime and from the CLR's GC; and the runtime-derived WAL replay ceiling this
# overlay exists to pin. An operator who forgot to export a variable and one who
# pinned it to `0` on purpose produce byte-identical deployments, and the guard
# reports both as satisfied.
#
# No check can recover a distinction the encoding destroyed, so `0` is retired
# as a spelling. It is refused by the preflight below, and the deliberate
# "derive it" case is spelled `auto` on the two knobs whose parser we own
# (REPOCONTEXT_MAX_CONCURRENT_REPLAYS, EMBEDDER_INTRA_THREADS). The other five
# are read finally by Docker or the CLR, whose vocabulary is not ours to extend,
# so they take a positive value or nothing:
#
#   pwsh -File ./scripts/Assert-TuningEnv.ps1
#
# New-TuningEnv.ps1 runs it over the .env it has just written, so the ordinary
# path is already covered and this is for a file edited by hand afterwards.
#
# WHAT THE LITERALS COST, so the next person does not restore one. The two
# mem_limit values summed to 17 GiB, so on a 16 GiB host this stack could not
# start at all, and nothing said why. Worse, DOTNET_PROCESSOR_COUNT: "16"
# overrode a cgroup-aware default and held the WAL replay concurrency gate at
# 16 permits against a 6.0-CPU quota - a 2.67x oversubscription, measured on
# the live deployment, and the deployment half of the root cause in #2692.
#
# Derive the values for THIS host and corpus:
#
#   pwsh -File ./scripts/New-TuningEnv.ps1
#
# It writes .env in this directory. Read .env.example for what each variable
# means, what it is derived FROM, and which of them are policy rather than
# measurement - the distinction is stated per-variable and deliberately.
# ---------------------------------------------------------------------------

services:
  repocontext:
    # The base file declares `build:` and no `image:`, so without this pin the
    # documented `up -d --no-build` step cannot resolve an image at all. The
    # build-and-tag ladder that produces it (candidate-<sha> -> local, with the
    # displaced tag kept as rollback-<timestamp>) is in the runbook.
    image: repocontext-mcp:local
    environment:
      # ---- Scan cadence -------------------------------------------------
      # The base file's cadence makes the periodic reconcile effectively
      # continuous, which is the right default for a sample demonstrating that
      # an edit shows up within seconds. It is far more than a long-lived local
      # deployment needs, and it is the source of the continuous embedder load.
      #
      # The periodic deadlines below are counted in PASSES, not wall clock
      # (interval / (reconcile + jitter), rounded up), so the reconcile interval
      # and the jitter had to move together with them or the pass counts would
      # not change as intended.
      #
      # TRADE-OFF, stated because it is real: an in-place content edit that does
      # not bump its directory's modification time is now noticed within about
      # an hour instead of about two minutes. An edit that DOES touch a
      # directory mtime is still picked up on the next reconcile.
      LATTICE_SELFINDEX_TICK_SECONDS: "30"
      LATTICE_RECONCILE_INTERVAL_SECONDS: "60"
      # Non-zero jitter stops passes across repositories phase-locking into
      # simultaneous walks, which is what produced the observed load spikes.
      LATTICE_RECONCILE_JITTER_SECONDS: "15"
      # The full re-stat of every file, and the single heaviest operation.
      LATTICE_FULL_WALK_INTERVAL_SECONDS: "3600"
      # Two membership reads per indexed source, so on a converged repository
      # this dominates a pass. Costs no healing latency: finding an actual gap
      # forces an immediate in-pass scan regardless of this interval.
      LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS: "3600"

      # ---- Runtime sizing -----------------------------------------------
      # Server GC with an EXPLICIT heap count (issue #2596).
      #
      # This was DOTNET_gcServer: "0", commented "one heap instead of sixteen".
      # That traded away parallel collection to protect the memory limit,
      # because Server GC would otherwise size its heap count from
      # DOTNET_PROCESSOR_COUNT below, which is pinned to 16 for a completely
      # unrelated reason (the WAL replay concurrency gate). One variable, two
      # jobs, opposite requirements.
      #
      # The cost of that trade went unmeasured until gate run 2, whose log
      # measures it directly: 283 whole-process silence gaps of 5s or more,
      # longest 29.3s, totalling 20.9% of wall-clock, against a 30s Orleans
      # request timeout. Every timed-out call (127 of 127) began executing
      # within 0.5s of enqueue and then froze, so this is a stop-the-world
      # pause rather than queueing contention. Orleans logged "Silo not running
      # with ServerGC turned on" at startup throughout.
      #
      # GCHeapCount decouples the collector from PROCESSOR_COUNT, so the
      # footprint concern that motivated "0" is addressed directly instead of
      # by giving up parallel collection. It matched the cpus cap below, NOT
      # the 16 that used to be pinned beneath it.
      #
      # Under #2779 that correspondence is no longer a coincidence maintained
      # by hand: this and `cpus` below read the SAME derived value, so they
      # cannot drift apart the way two independently-edited literals could.
      # The original pairing was correct and is preserved by construction.
      DOTNET_gcServer: "1"
      DOTNET_GCHeapCount: "${REPOCONTEXT_GC_HEAP_COUNT:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
      # ---- WAL replay concurrency ---------------------------------------
      # DOTNET_PROCESSOR_COUNT: "16" USED TO BE HERE AS A PIN. THE PIN IS GONE
      # (issue #2779) and is not coming back. What IS here, at the bottom of
      # this block, is the same name declared as an UNSET pass-through - which
      # is a different object and is explained there (issue #2931).
      #
      # Do not restore the pin. Its comment read, in full:
      #
      #   "DO NOT REMOVE. This decouples the WAL replay concurrency gate from
      #    the CPU cap below. BPlusLeafGrain sizes that process-wide gate from
      #    Environment.ProcessorCount when the option is left non-positive,
      #    once, as a structural constant. Environment.ProcessorCount is
      #    cgroup-aware, so the cpus limit below would otherwise shrink the
      #    gate silently and unattributably. Pinning the reported count holds
      #    the gate at the value every previous field measurement on this box
      #    was taken against, while the cgroup quota still bounds actual CPU
      #    consumption.
      #
      #    This is also why the embedder derives its own thread count from the
      #    cgroup quota rather than from the processor count (issue #2610):
      #    this variable overrides Environment.ProcessorCount and wins over the
      #    quota, so a service that copied this environment block would be
      #    oversubscribed."
      #
      # Every mechanical sentence in that is accurate, and the last paragraph
      # correctly identifies the variable as an oversubscription hazard for any
      # service that copies it. The justification for keeping it is what fails:
      # it holds the gate "at the value every previous field measurement on
      # this box was taken against" - preserving COMPARABILITY WITH PRIOR RUNS,
      # while the thing held constant was the defect those runs were measuring.
      # A comment can be meticulous about one axis and silent about another;
      # this one reasons carefully about continuity and never about cost.
      #
      # THE COST, measured on the live deployment rather than argued:
      #   - the gate sized to 16 permits against a `cpus: 6.0` quota (2.67x);
      #   - `docker stats` bursting to 638/751/681% against a 600% quota;
      #   - `Configured WalMaterialiserMaxConcurrentReplays=0` in the gate's own
      #     log line, proving the supported knob was never set;
      #   - and 16 concurrent whole-window replays holding multi-MiB buffers,
      #     which is the deployment half of the heap exhaustion in #2692.
      # BPlusLeafGrain.Activation.cs:154-172 (#2278) already names this exact
      # shape as self-reinforcing, in the code this variable was overriding.
      #
      # AND THE COMPARABILITY IT BOUGHT IS WORTHLESS, which is the part that
      # closes the argument rather than merely rebutting it. Every prior field
      # measurement this pin was protecting was taken IN the 2.67x-oversubscribed
      # condition. They are comparable to each other and to nothing that should
      # ever be run again, so there is no longer a continuity argument to weigh
      # against the cost - the baseline the pin preserved was itself the defect.
      # If you are tempted to restore this variable for that good reason, that
      # is the reason it does not hold.
      #
      # The gate's decoupling from the CPU cap is a REAL requirement and is
      # still met - by the option that exists for it (#2279) instead of by a
      # runtime-wide variable with two jobs. This one has a single job, is
      # read by the library in preference to Environment.ProcessorCount, is
      # refused at startup if unparseable, and - unlike the pin - does not also
      # resize the GC, the thread pool, and every other ProcessorCount consumer
      # in the process as a side effect.
      #
      # ---- The name, declared UNSET, so the value is attributable (#2931) ----
      #
      # Everything above is about the PIN and remains true. This is about the
      # NAME, and the two were conflated by deleting the line outright.
      #
      # Deleting it did remove the pin. It also removed the variable, and with
      # it any way to express the value without editing this tracked file. The
      # cost landed on attribution rather than on tuning: run 9's predicate was
      # amended to bind that run to run 8's grants INCLUDING
      # DOTNET_PROCESSOR_COUNT=16, so that run 9 would be a code-only
      # comparison. The variable was by then not merely different, it was
      # ABSENT, and Environment.ProcessorCount was resolving from the cgroup.
      # A replay-gate ceiling then moved 16 -> 6 in a step that also removed a
      # pin, giving one movement two sufficient causes and no way to separate
      # them. That is a measurement lost, not a value mis-set.
      #
      # DECLARED WITH NO VALUE, which is a compose idiom and not an omission.
      # A null-valued mapping entry is resolved from the shell environment or
      # .env, and is OMITTED FROM THE CONTAINER ENTIRELY when neither supplies
      # it. Verified on this rig rather than assumed:
      #
      #   unset        -> `docker compose config` renders `null`, and in the
      #                   container `${DOTNET_PROCESSOR_COUNT+x}` is empty,
      #                   i.e. the variable is ABSENT, not empty-string;
      #   set in .env  -> present with that value.
      #
      # So today's absent state is preserved BYTE FOR BYTE - no effective value
      # moves by adding this line - while a controlled run can pin it without
      # editing a tracked file, which is what #2931 asked for.
      #
      # IT IS STILL AN OVERSUBSCRIPTION HAZARD. The final paragraph of the old
      # comment above is correct and applies unchanged to anyone who sets this:
      # it overrides Environment.ProcessorCount process-wide and wins over the
      # cgroup quota. Setting it is a deliberate act for a controlled
      # comparison, and Assert-DeployManifest.ps1 reports the resolved count
      # and WHERE IT CAME FROM on every deploy, so a set value cannot be a
      # silent one and neither can an unset one.
      DOTNET_PROCESSOR_COUNT:
      LATTICE_WAL_MAX_CONCURRENT_REPLAYS: "${REPOCONTEXT_MAX_CONCURRENT_REPLAYS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it from this container's CPU grant, or set it by hand from .env.example. No default by design (issue #2779).}"

      # Diagnostic for issue #2252, NOT tuning, and so deliberately left
      # commented out rather than tracked as a default. Uncomment only while
      # investigating vector-writer disposition counts, and record it in the
      # runbook's local-only deltas section while it is on.
      # Logging__LogLevel__Orleans.Lattice.Api.Mcp.RepoContext.RepoContextVectorWriter: "Debug"

    # Bounds a runaway without starving normal operation. This service measured
    # about 92% of one core in steady state before any of these limits, so the
    # grant is headroom rather than a working limit.
    #
    # It is now also the SOURCE for DOTNET_GCHeapCount and for the WAL replay
    # gate above: the derive script sizes all three from this one number, which
    # is what the base compose file's LATTICE_WAL_MAX_CONCURRENT_REPLAYS
    # comment asks for when it says to set the gate together with the CPU limit
    # it derives from, in the same file.
    cpus: "${REPOCONTEXT_CPUS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
    # Sized above the measured steady-state plateau, not below it. See the
    # runbook: a cap below the working set does not present as a resource event,
    # because .NET sizes its heap hard limit from the cgroup limit and collects
    # harder as it approaches it rather than being OOM-killed at it.
    #
    # That is precisely why this one must not be a literal (#2779): the failure
    # mode of an undersized grant is a STORAGE fault, not a container kill, so
    # a value copied from a larger host degrades the stack while `docker ps`
    # still reports it healthy. The derive script sizes it from the corpus and
    # clamps it to what the host can actually offer. It REFUSES when the host
    # cannot offer even the floor, and when the clamp binds between the floor
    # and the requirement it grants the ceiling and names the shortfall - it
    # never grants less silently (#2832).
    mem_limit: "${REPOCONTEXT_MEM_LIMIT:?unset - run ./scripts/New-TuningEnv.ps1 to derive it from this host and corpus, or set it by hand from .env.example. No default by design: a memory grant nobody measured is the defect this overlay exists to stop (issue #2779).}"

  embedder:
    environment:
      # Workstation GC. Unlike the sibling service above, this one has no large
      # managed heap worth collecting in parallel: its footprint is dominated by
      # the resident ONNX model, which is native and outside the GC's control.
      DOTNET_gcServer: "0"
      # Pins the ONNX intra-op thread pool to the cgroup grant below.
      #
      # ONNX Runtime sizes that pool from HOST cores (16 here) and does not
      # consult the cgroup quota, so under the 4.0-CPU grant the pool ran 4x
      # oversubscribed and the kernel throttled it in 296 of 298 consecutive
      # scheduling periods during vectorising. This is the third instance in
      # this one deployment of a runtime sizing a pool from host cores under a
      # fractional grant; see DOTNET_GCHeapCount and DOTNET_PROCESSOR_COUNT on
      # the sibling service above.
      #
      # Issue #2610 derives this value from the cgroup automatically, so on an
      # image built from that commit or later the pin only restates what the
      # server would have chosen for itself. It is tracked anyway, because the
      # deployment this file describes still runs a pre-#2610 image and because
      # an explicit value is what the deployed configuration actually carries.
      # Removing it is precisely how the derived path gets exercised, which has
      # not happened yet on this box.
      #
      # Under #2779 it stays explicit for those reasons, but is no longer a
      # literal: it is derived from the SAME number as the cgroup grant below,
      # so the restatement is guaranteed to agree with what #2610 would pick
      # rather than merely happening to. A literal `4` beside a `cpus` value
      # that someone later edits is how the two silently stop matching.
      EMBED_INTRA_THREADS: "${EMBEDDER_INTRA_THREADS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
    # Leaves the balance of the host's cores for interactive work. The embedder
    # sizes its ONNX intra-op thread pool from this grant, read from the cgroup
    # filesystem (issue #2610), so changing this number changes the pool.
    cpus: "${EMBEDDER_CPUS:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"
    # The resident ONNX model needs room to work above its own footprint.
    #
    # Unlike the sibling service's grant, this one is WORKLOAD-derived, not
    # corpus-derived: it is dominated by the model resident in the image, so it
    # does not grow with the repository. The derive script therefore treats it
    # as a near-constant and only clamps it against the host. It is still a
    # variable because it is half of the sum that has to fit: the two
    # mem_limits together used to demand 17 GiB on any host, which is why a
    # 16 GiB machine could not start this stack at all (issue #2779).
    mem_limit: "${EMBEDDER_MEM_LIMIT:?unset - run ./scripts/New-TuningEnv.ps1 to derive it, or set it by hand from .env.example. No default by design (issue #2779).}"