docker-compose.yml
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at docker-compose-yml.md, and llms.txt lists every page.Part of RepoContextContainer source.
# RepoContext MCP container - "codebase memory in a box" (issue #1435).
#
# Brings up the restart-durable MCP host alongside its ONNX Runtime embedding
# companion. The host runs the default `local` durability profile (SQLite grain
# storage + reminders and the file WAL, all under the /data named volume), so
# ingested context survives `docker compose restart` and `docker compose down`
# (without `-v`). See README.md for the start -> bootstrap -> recall -> restart
# -> verify walkthrough.
#
# `docker compose down -v` destroys /data, and that volume holds the authored
# agent memory as well as the rebuildable index. The /memory-archive bind mount
# on the repocontext service exists for exactly that gesture: it is not a
# declared volume, so `-v` does not remove it, and the host restores from it
# into an empty store at startup. It is bounded by the export interval and is
# not a backup - `repocontext_reset_index` remains the lossless way to rebuild
# an index. See the comments on that service (issue #2601).
#
# Build context is the repository root for the host image (it ProjectReferences
# the just-built src/ bits); the embedder builds from its own app directory. Run
# from this directory.
services:
# The embedding companion stays a SEPARATE container so the MCP host keeps its
# single-listener surface. Its baked HEALTHCHECK gates the host's startup.
#
# This is the ONNX Runtime companion (apps/embedding-onnx). It serves the same
# HTTP contract on the same port as the Onyx companion it replaced, and emits
# numerically identical vectors, so it is a drop-in: no host configuration
# changes and no re-indexing of an existing volume. To go back to the Onyx
# companion, layer docker-compose.onyx.yml over this file.
#
# NVIDIA: the default `cpu` flavour has NO GPU support compiled in, so running
# it on a GPU host changes nothing. Enabling the GPU takes all three of the
# edits below - the build arg selects a different ONNX Runtime package, the env
# var selects the accelerator at runtime, and the reservation is what actually
# exposes the device to the container:
#
# build:
# args:
# ONNX_FLAVOR: cuda # (1) build the CUDA-capable image
# environment:
# EMBED_PROVIDER: cuda # (2) bind the CUDA execution provider
# deploy: # (3) hand the GPU to the container
# resources:
# reservations:
# devices:
# - driver: nvidia
# count: 1
# capabilities: [gpu]
#
# Rebuild after changing the build arg (`docker compose build embedder`); a
# plain `up -d` reuses the cached cpu image. The cuda flavour still includes the
# CPU provider, so that one image also serves CPU hosts, and an unrecognised
# EMBED_PROVIDER falls back to the CPU instead of failing to boot. Confirm what
# actually bound with `curl localhost:9000/api/health` inside the network - it
# reports the live provider, so a silent CPU fallback is visible.
embedder:
build:
context: ../../apps/embedding-onnx
dockerfile: Dockerfile
# The container cgroup readers are compiled from their one canonical copy
# in the core library rather than mirrored into the app (issue #2817). The
# Dockerfile copies them from this BuildKit named context.
additional_contexts:
cgroups: ../../src/lattice/Internal/Cgroups
args:
# cpu (default) or cuda. The cuda flavour also serves CPU hosts, so one
# image covers both and EMBED_PROVIDER alone selects the accelerator.
ONNX_FLAVOR: cpu
environment:
# cpu, or cuda on an NVIDIA host started with a device reservation.
# An unknown value falls back to the CPU rather than failing to boot.
EMBED_PROVIDER: cpu
# EMBED_INTRA_THREADS is deliberately NOT set here, and deliberately not
# given a placeholder value either, for the same reason
# LATTICE_WAL_MAX_CONCURRENT_REPLAYS is left unset below: the correct value
# is a property of YOUR CPU grant, and copying a literal out of somebody
# else's deployment is precisely the failure the setting exists to prevent.
#
# WHAT IT IS. It sizes the ONNX Runtime intra-op thread pool - the threads
# that parallelise a single inference. Left unset, this server derives it
# from the container's ENFORCED cgroup CPU quota, which is the right answer
# in every case we have measured. You only need to declare it if you want
# to depart from that.
#
# WHY IT IS DERIVED FROM THE QUOTA AND NOT FROM THE CORE COUNT. ONNX
# Runtime's own default sizes the pool from the HOST core count and does
# not consult the cgroup quota. Under any CPU limit that oversubscribes by
# the ratio between the two, and the cost is not proportional - it is far
# worse. Measured on a 4.0-CPU grant on a 16-core host (issue #2606): a
# pool of 16, the kernel throttling 296 of 298 consecutive scheduling
# periods, and 346.3 CPU-seconds stalled against 118.8 running. The
# arithmetic is elementary once stated: 16 threads drain a 400ms quota in
# 25ms and are then frozen for the remaining 75ms, so 75:25 = 3.0 stalled
# per unit run, against 2.91 measured. Because ONNX Runtime synchronises
# intra-op threads at EVERY operator boundary and a transformer crosses
# hundreds per inference, a freeze landing mid-barrier stalls the whole
# operator rather than one thread, so throughput degrades super-linearly.
#
# WHY DOTNET_PROCESSOR_COUNT CANNOT BE REUSED FOR THIS. It used to be set
# on the repocontext service, where it sized an unrelated WAL replay gate.
# It OVERRIDES Environment.ProcessorCount and wins over the quota, so
# copying that service's environment block onto this one - an entirely
# ordinary thing to do - would silently restore the oversubscription. This
# server reads the quota directly and is immune to it, and logs a warning
# naming both figures when the two disagree.
#
# That pin is GONE as of #2779, which found it holding the replay gate at
# 16 permits against a 6.0-CPU quota on the reference host: a 2.67x
# oversubscription, and the deployment half of #2692. The warning above is
# kept rather than deleted because the HAZARD outlived the instance - the
# variable is still global to the process and still outranks the quota, so
# anyone reintroducing it anywhere reintroduces exactly this.
#
# HOW TO DERIVE IT IF YOU DO SET IT. Use the container's actual CPU grant,
# rounded UP: `cpus: "4.5"` -> 5. Not the host core count, and not
# DOTNET_PROCESSOR_COUNT. IF YOU LATER CHANGE THE GRANT BELOW, A DECLARED
# VALUE DOES NOT FOLLOW IT - that is the whole hazard of pinning one, and
# is why leaving it unset is the recommended configuration. The startup log
# reports the resolved count and whether it was DECLARED or DERIVED, so
# check there rather than inferring it from this file.
expose:
- "9000"
# Run PID 1 as an init process (issue #2576). Docker bind-mounts its own
# static `docker-init` binary and runs it as PID 1, so this works unchanged
# on the distroless, shell-less runtime image - nothing has to be baked in.
#
# This is a SEPARATE defect from the stop_grace_period below, and neither
# substitutes for the other: the grace period governs how long the drain is
# given, while init governs whether PID 1 behaves like a process supervisor
# at all. Without it the server itself is PID 1, which reaps no orphaned
# children and is subject to the kernel's rule that PID 1 takes no default
# action for a signal it has installed no handler for.
init: true
# The embedder holds NO durable state: its weights are baked into an image
# layer and the base file mounts it no volume, so there is nothing to flush
# and Docker's 10s default would in fact be adequate here. It is set
# explicitly anyway, and deliberately short, so that the value is a decision
# on the record rather than an omission. The repocontext service below was
# left to the same default for years while its documentation promised a
# "generous shutdown budget" (issue #2389); the cost of that was not the
# default itself but that nothing distinguished "10s is enough here" from
# "nobody considered it". A reader of this file can now tell the two apart.
stop_grace_period: 30s
# Survive a daemon/host restart like the repocontext service does. Without
# this the embedder's default `no` policy leaves it stopped after a Docker
# daemon restart while repocontext (unless-stopped) comes back on its own,
# degrading search to keyword mode until the embedder is started by hand.
restart: unless-stopped
# OPT-IN CPU PINNING. Unset by default, and unset it renders to NOTHING -
# Compose omits the key entirely rather than emitting an empty one, so a
# stack that does not set this variable resolves a byte-identical document
# to one built before the knob existed. See the repocontext service below
# for what it does, when to use it, and when NOT to (issue #2623).
cpuset: "${EMBEDDER_CPUSET:-}"
# Cap container log growth: the default json-file driver is otherwise
# unbounded. Rotate at 20 MiB x 5 files (100 MiB ceiling per container).
logging:
driver: json-file
options:
max-size: "20m"
max-file: "5"
# The RepoContext MCP host: its ONLY application listener is the MCP endpoint on
# 8080 (plus the HTTP health probes and the /metrics scrape endpoint). No gRPC
# facade, no Explorer UI.
repocontext:
build:
context: ../..
dockerfile: apps/repocontext/Dockerfile
args:
# The commit being built, stamped into the image as the OCI label
# org.opencontainers.image.revision. Export it before building:
#
# $env:GIT_COMMIT = (git rev-parse HEAD) # pwsh
# export GIT_COMMIT=$(git rev-parse HEAD) # sh
#
# WHY THIS IS NOT OPTIONAL DECORATION. Without it the only record of
# what an image contains is the HEAD of whatever checkout someone
# happened to be standing in, and that is a fact about the checkout, not
# about the image. Bucket-4 gate run 4 made exactly that inference and
# measured eleven hours against an image built 45 commits behind what it
# believed (issue #2686).
#
# Unset resolves to the empty string, which
# scripts/Assert-ContainerProvenance.ps1 reads as unresolved: it then
# falls back to a `candidate-<sha>` image tag, and REFUSES if neither
# channel can name the built commit.
GIT_COMMIT: ${GIT_COMMIT:-}
depends_on:
embedder:
condition: service_healthy
# Wait for the backup sink so the first capture at startup has somewhere
# to write. service_started rather than service_healthy on purpose: a
# sink that is slow to come up must never gate the MCP host's own
# readiness, and the capture path retries with backoff anyway.
azurite-backup-sink:
condition: service_started
environment:
# Default local, durable, zero-external-dependency profile.
LATTICE_DURABILITY: local
# All durable local state lives here (a named volume): the file WAL
# directory and the SQLite database file.
LATTICE_DATA_ROOT: /data
# Where the durable AGENT MEMORY archive is written (issue #2601).
#
# Read what this does and does not protect before relying on it.
#
# WHAT IS AT RISK. Two kinds of state share the /data volume above. The
# code index (structural, content, symbol, xref, session, and the vector
# planes) is DERIVED: destroy it and re-running repocontext_add_repo
# rebuilds it in minutes. Agent memory - every repocontext_remember note,
# decision, gotcha, and glossary entry - is AUTHORED, is the store of
# record, and is rebuildable from nothing. `docker compose down -v`
# destroys both, and the asymmetry is the whole defect: the gesture is
# documented as the ordinary way to reset an index, and it silently takes
# the irreplaceable half with it. It has already happened once, to epic
# #2368's own memory.
#
# WHY NOT SIMPLY A SECOND VOLUME. Because it cannot be done and would not
# help. It cannot be done: a Lattice tree's durable state spans a WAL root
# that is one directory for the whole PROVIDER, and a grain store that is
# one SQLite file for every tree - neither is per-tree addressable, so
# there is no memory-only path to mount elsewhere. And it would not help:
# `down -v` removes every named volume the project declares, so a second
# declared volume dies in the same gesture as the first.
#
# WHAT THIS DOES INSTEAD. A BIND MOUNT (below) is not a project volume, so
# `down -v` does not touch it. The host periodically exports memory there
# and restores it into an empty store at startup, so a `down -v` followed
# by `up` comes back with memory intact.
#
# WHAT IT DOES NOT DO. It is not a backup, and it does not make `down -v`
# safe. Everything written since the last export is lost, so the exposure
# is bounded by the interval below plus whatever a non-graceful stop
# discards. `repocontext_reset_index` remains the correct, lossless way to
# rebuild an index: it drops the derived planes and preserves memory
# outright, with no window at all. Reach for `down -v` only when you mean
# to destroy the container's state, and expect to lose memory written
# since the last export.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR: /memory-archive
# How often memory is exported. This is the size of the window you lose to
# an ungraceful stop, so lower it if a session's notes are expensive to
# recreate. A graceful stop exports once more on the way down, which closes
# most of the window for `down` / `down -v` / `stop` specifically; a SIGKILL
# or a host crash gets no such chance.
LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_INTERVAL_SECONDS: "300"
LATTICE_MCP_PORT: "8080"
# WAL materialiser pin buckets: how many persisted slots the retention-floor
# pin state is split across, so an advancing floor rewrites a fraction of the
# blob instead of all of it. The library default is 1 (the legacy write path,
# kept so upgrading an existing deployment changes nothing); this box opts in
# because its measured pin blobs reached ~1.4 MB rewritten tens of thousands
# of times. Widening self-migrates on activation and the legacy slot is
# retained, so setting this back to "1" and restarting is a safe rollback
# that over-retains WAL rather than over-trimming it.
LATTICE_WAL_PIN_BUCKETS: "8"
# Ceiling on how long ONE durable pin shard may shed steady-state pin
# reports CONTINUOUSLY before a single report is forced through anyway.
#
# WHY THIS EXISTS (issue #3310, measured on this box). The steady-state
# pin write is shed under pin-store pressure (issue #2014) and is the ONLY
# write path carrying an advancing CheckpointOffset. The paths that are
# exempt from the gate - the birth seed and the deactivation flush - still
# FEED it, so under continuous churn the window re-arms faster than it
# lapses and never closes. The durable pin offset then freezes, the WAL GC
# offset floor has nothing fresher to stand on, and retained WAL grows
# without bound. On repo-context-vector-index the floor sat frozen at one
# offset for forty minutes while the WAL grew 175 MB, and the shed counter
# read 12987 for that tree against 12, 2 and 0 for its three siblings.
#
# THIS IS NOT THE ISSUE #3300 FAILURE IN DISGUISE, and the distinction is
# the whole safety argument. It does not lower, weaken or bypass the
# retention floor - the leaf clamps every reported offset to
# min(checkpoint, covered) before the report is built, so a forced report
# is arithmetically incapable of publishing an offset it has not proven
# durable. It bounds WHEN a true value is published, never WHAT.
#
# The hazard of setting this TOO SMALL is the opposite one: forcing writes
# faster than the pin store can absorb them re-saturates the queue the
# shedding exists to protect. Read it as a ceiling on stall, not a target.
# At 120 s each pin shard contributes at most one forced write every two
# minutes, so eight shards cost a handful of writes a minute against the
# thousands of reports a busy tree sheds in the same window.
#
# Setting this to "0" disarms the ceiling and restores the pre-#3310
# behaviour exactly. That is a rollback to unbounded retention growth, not
# to data loss, which is the correct direction for this seam to fail.
# Watch orleans_lattice_materialiser_pin_shed_stall_seconds (a line that
# only climbs is the latch) and ..._pin_shed_forced_total (every increment
# is one deliberate override, and an alarm rather than routine).
LATTICE_WAL_PIN_SHED_CEILING_SECONDS: "120"
# Bound how long a batch write's per-shard fan-out may run before it is
# refused with a LatticeSaturatedException. The library ships this
# UNBOUNDED so that upgrading cannot regress any existing caller (#3386),
# which is the right default for a library and the wrong setting here:
# this deployment issues wide batch writes against trees that are
# demonstrably saturated, and an unbounded fan-out converts that
# saturation into an indefinitely-held call instead of a refusal the
# caller can see, retry, or shed. The held call is the wedge; a refusal is
# a signal. 30s is the figure #3348's measurements support and matches
# LATTICE WAL append dispatch's own 30s outer bound, so a fan-out branch
# cannot outlive the single WAL dispatch it is waiting for.
# The refusal rolls NOTHING back: a batch write is not atomic across
# shards, so committed branches stay committed and in-flight branches are
# left running rather than cancelled. The durable outcome is identical to
# an unbounded wait; only WHEN the caller learns changes.
# Rollback is "0", which restores the unbounded library default exactly -
# a rollback to the wedge, not to data loss.
LATTICE_SETMANY_FANOUT_BUDGET_SECONDS: "30"
# Advisory per-tree ceiling on the WAL bytes held on disk. The library ships
# this DISABLED and is right to: the correct value depends both on the
# volume the WAL lives on and on the largest tree's working set, neither of
# which a library can know. This deployment declares its own volume and has
# measured its trees, so it names one.
#
# Left unset, every tree here reported
# wal_gc_backlog_bytes_unavailable_total{reason="policy_disabled"}
# and reclaimed_bytes_total sat at zero on all but one tree, so a 3 GB WAL
# only ever grew.
#
# SIZE IT ABOVE 2x THE LARGEST TREE'S LOGICAL WORKING SET. This is not a
# style preference, it is what makes the ceiling reachable. Compaction is
# the mechanism that actually bounds the WAL, it is decided per shard in
# FileWalShard.CompactIfNeeded, and with CompactionMaximumDeadBytes
# disabled by default the sole trigger is the 0.5 dead-byte ratio. Designed
# steady-state physical occupancy is therefore about TWICE the live set.
# Since #3107 this ceiling is compared against physical bytes, so a value
# below 2x logical is unsatisfiable by construction: the tree breaches while
# perfectly healthy. Worse, WalBytePressureReclaimTarget (0.8) means it then
# never disarms either, leaving a permanently armed advisory alarm that
# cannot distinguish pathological growth from normal size.
#
# 20 GiB against repo-context-vector-index measured 2026-09-21. Logical was
# not read directly - it is BOUNDED BY INFERENCE from the alarm itself,
# which is the cheaper measurement and is exact enough to size against.
# CeilingUnsatisfiable fires iff ceiling < 2x logical (LatticeWalGc.cs
# L1997-2001), and it was firing against 8 GiB, so logical EXCEEDS 4 GiB.
# Physical was 5.864 GiB and rising at ~0.82 GiB/h, and logical <= physical,
# so 2x logical lies in (8.0, 11.7) GiB. 20 GiB clears the top of that band
# with ~1.7x headroom. Every other tree is under 40 MB, so the ceiling is
# inert on them, which is correct for a disaster backstop; they still pay
# one physical-size probe per partition per GC pass. The volume has ~681 GB
# free, so this is not disk-constrained.
#
# THIS VALUE HAS BEEN WRONG THREE TIMES, FOR THREE DIFFERENT REASONS. Read
# all three before changing it, because the rule above is easy to satisfy
# on the day and hard to keep satisfied.
#
# 1 GiB was sized from occupancy measured WHILE the #3229 starvation
# defect was active (it cited vector-index 1710 MB and vector-payload
# 1097 MB). Post-fix vector-payload is 17 MiB, so that calibration
# described the faulted state, not the healthy one, and 1 GiB had fallen
# below index's live payload alone. See #3234.
#
# 4 GiB was correct when written and was outgrown within hours. It was
# calibrated against 1,077 MiB logical; the same tree measured 2,226 MB
# logical the same day - 2.07x - which put 2x logical at 4,451 MB, i.e.
# 355 MB ABOVE the ceiling meant to contain it. The tree was at 71% of a
# designed steady state that sat above its own ceiling, so it was headed
# for a breach-while-healthy and a permanently-armed alarm. Nothing
# reported this: a tree in that state emits wal_gc_passes_total
# {outcome="stranded"}, which says "could not reclaim" and does not
# distinguish lagging consumers from an arithmetically unreachable
# ceiling. See #3242, which adds the missing in-band signal.
#
# 8 GiB repeated the 4 GiB mistake at a larger number, and the repetition
# is the point: it was calibrated against the same 2,226 MB logical figure
# that 4 GiB was, so it bought a 2x multiple of a reading that was already
# moving. It survived roughly two days. By 2026-09-21 the same tree was
# unsatisfiable again, with wal_gc_ceiling_unsatisfiable climbing steadily
# (40 -> 54 over one 36-minute window) - the in-band signal #3242 added,
# doing exactly its job and reporting a ceiling arithmetically unreachable
# rather than a lagging consumer. Note what did NOT happen: byte pressure
# never armed (over_ceiling stayed 0), so nothing was ever trimmed harder
# on account of this value being wrong, and the only cost was a misleading
# alarm - which is precisely the failure mode the paragraph above predicts.
#
# THE LESSON, WHICH IS THE ONLY DURABLE PART: this ceiling is expressed as
# a multiple of a MEASURED quantity, so it acquires an expiry date the
# moment the workload can grow. Re-measure logical retained per tree
# before trusting the current value - do not assume the last calibration
# still holds because it was right when it was made.
#
# It cannot lose data and cannot block a write: it only lowers the trim
# frontier WITHIN the frontier that was already safe, never past a live
# consumer cursor. Setting it to "0" restores the library default and is a
# safe rollback that over-retains rather than over-trims.
LATTICE_WAL_MAX_RETAINED_BYTES: "21474836480"
# Absolute per-SHARD ceiling on dead (trimmed but not yet physically
# reclaimed) WAL bytes, above which a compaction evaluation rewrites the
# shard whatever the ratio above says.
#
# WHY THIS EXISTS (issue #3223, measured on this box). FileWalShard
# .CompactIfNeeded is the only thing that returns bytes to the
# filesystem, and with this UNSET its only arm that can fire on a running
# tree is the 0.5 dead-byte RATIO. A ratio is not a bound. It fires at
# dead >= live, so designed steady state is TWICE the live set and the
# absolute waste grows with the tree. Measured on
# repo-context-vector-index the ratio was 200/1268 = 0.158, so the arm
# declined correctly and forever: wal_compactions_total read 0 on ALL
# THREE trigger tags, wal_compaction_reclaimed_bytes_total read 0, and
# physical size never decreased once in 37 samples (8971 -> 10174 MiB).
#
# THE MEASUREMENT THIS VALUE IS SIZED FROM. Per shard, over 8 shards:
# retained 1068 MiB (x8 = 8544; gauge storage_wal_bytes read 8564)
# dead 200 MiB (x8 = 1600)
# total 1268 MiB (x8 = 10144; du -sm on the volume read 10174)
# Both identities close to within 0.3%, so the du-versus-gauge gap IS the
# dead bucket. storage_wal_bytes is RETAINED-ONLY and the sum is not
# exported at all (#3205) - do not read that gauge as physical size.
#
# SIZING. 134217728 = 128 MiB per shard. Three things fix it, in order.
#
# 1. THE GATE IS SHARD-LOCAL. _deadBytes is one shard's, never the
# tree's, so the tree-wide bound is this value TIMES the shard count.
# The ceiling above already declares 8192 MiB per tree and
# vector-index has 8 shards, so each shard's share of physical budget
# is 1024 MiB. 128 MiB caps dead at ONE EIGHTH of that share, i.e.
# 1024 MiB tree-wide, where the ratio arm permits one HALF. The
# derivation is anchored on a number this file already owns and on the
# shard count - not on a corpus snapshot - which is deliberate: see
# the expiry-date lesson in the block above. It is not snapshot-free
# though: RESHARDING vector-index moves the tree-wide bound
# proportionally without touching this line.
#
# 2. IT MUST FIRE AGAINST MEASURED DEAD, AND ABOVE ~200 MiB IT DOES NOT.
# The untracked predecessor was 256 MiB, sized in #3109 against a
# corpus holding ~344 MB dead per shard. Present dead is 200 MiB per
# shard, so 256 MiB is INERT at the state actually measured. This
# deployment's recurring failure is settings that look applied and do
# nothing, so prefer the value whose effect is observable within one
# GC pass: 128 < 200, so the first evaluation after deploy compacts
# and wal_compactions_total{trigger="ceiling"} leaves zero. The
# asymmetry is the whole argument - too high forfeits the entire
# benefit, modestly too low costs a bounded, quantified amount of I/O.
# 192 MiB would also fire but sits at 96% of a single snapshot of a
# quantity that is still climbing, which is no margin at all.
#
# 3. THE COST IS WRITE AMPLIFICATION, AND 128 MiB BUYS IT CHEAPLY. A
# compaction rewrites every LIVE byte to reclaim the dead ones, so
# work per byte reclaimed is live/ceiling = 1068/128 = 8.3x, against
# 1.0x for the ratio arm. Dead reached 200 MiB per shard within the
# 2h42m since the container was recreated, so accrual is AT MOST
# ~74 MiB/h per shard while ingesting: a shard rewrites about every
# 1.7 h, i.e. ~620 MiB/h per shard and ~4.8 GiB/h over 8 shards, which
# is 1.4 MiB/s sustained on a volume with ~598 GB free. Bandwidth is
# therefore NOT the constraint. LATENCY is: Compact() runs
# synchronously inside TrimAsync while holding that shard's gate, so
# every rewrite stalls the shard for a ~1 GiB file copy. The stall
# SIZE is set by live bytes and does not shrink as this value shrinks;
# only the FREQUENCY rises. That is what rules out the small end far
# more than the amplification arithmetic does - 32 MiB would quadruple
# the stalls and reclaim not one byte more.
#
# DO NOT PROBE THIS WITH A SMALL VALUE (issue #3210). Low is the
# AGGRESSIVE end and only "0" disables. CompactionMinimumDeadBytes
# (64 KiB by default, and this host never sets it) is evaluated FIRST, so
# every shard reaching the ceiling comparison has already cleared the
# floor; a sub-floor ceiling is then unconditionally true and rewrites
# the whole shard on every trim. 128 MiB is 2048x that floor. Since #3211
# the registration-time validator refuses 0 < ceiling < floor outright,
# so getting this wrong downward now costs a refused start rather than
# silent amplification - but the validator only sees values this host
# passes through, so it is a backstop and not a licence to guess.
#
# WHAT THIS DOES NOT FIX, AND IT IS NOT A DETAIL. Reclaiming EVERY dead
# byte leaves retained at 8544 MiB, still above the 8192 MiB ceiling
# above - per shard, 1068 MiB of live data already exceeds its own
# 1024 MiB share. Compaction only ever moves dead -> free; only an
# advancing trim frontier moves retained -> dead, and the durable offset
# floor is currently refusing that trim (#3310). Expect physical to fall
# about 10174 -> 8574 MiB on the first pass and then settle below
# ~9568 MiB: a ~7.3 GiB cut to the designed steady state, and still a
# breach of the ceiling. NECESSARY, NOT SUFFICIENT.
#
# Setting this to "0" restores the library default and disarms the
# ceiling, which is a safe rollback to unbounded dead-byte growth rather
# than to data loss. Watch wal_compactions_total{trigger="ceiling"} and
# wal_compaction_reclaimed_bytes_total: both flat while entries_trimmed
# climbs for the same shard is this defect returning.
LATTICE_WAL_COMPACTION_MAX_DEAD_BYTES: "134217728"
# Per-silo ceiling on concurrent activation-time leaf WAL replays. Left
# UNSET here deliberately, so this sample behaves identically on every
# host: unset defers to the library, which sizes the gate from
# Environment.ProcessorCount.
#
# Set it when this container runs under a CPU limit (issue #2279).
# Environment.ProcessorCount is cgroup-aware only while
# DOTNET_PROCESSOR_COUNT does not override it, and an override wins
# silently - so a box granted 6 CPUs whose environment says 16 sizes this
# gate to 16 and oversubscribes the memory-admitting replay path 2.67x. Each permit admits
# one whole-readable-window WAL replay, which is exactly the expensive path
# a reactivation storm multiplies.
#
# LATTICE_WAL_MAX_CONCURRENT_REPLAYS: "<this container's CPU limit>"
#
# The placeholder is deliberate: copying a literal from someone else's
# deployment is the failure this setting exists to prevent, so there is no
# number here to copy. An unparseable value is refused at startup with a
# message naming the variable, which is the safe way to get it wrong.
#
# Set it TOGETHER WITH the CPU limit it is derived from, in the same file,
# and keep the two adjacent. A ceiling pinned somewhere the quota is not
# declared is a second claim about that quota which nothing can check - the
# hazard that produced #2279 to begin with, relocated rather than removed.
# Verify what actually bound: the silo logs the resolved ceiling once at
# startup, alongside the configured option and Environment.ProcessorCount,
# so a value that was ignored is visible rather than merely assumed.
# MEMORY. This box needs FAR more than a default Docker VM allocation, and
# under-provisioning it fails in a way that does not look like memory
# (issue #2364). TWO figures matter here and they are NOT the same number:
# what the box costs once it is healthy, and what it costs to GET there.
#
# STEADY STATE. The fitted requirement is 3 GiB fixed plus 1 MB per indexed
# file, so a ~8,200-file index settles near 11 GiB. An earlier revision of
# this note gave 10.2 GiB and recommended "at least 12 GiB"; both figures
# are WITHDRAWN. They predate scripts/New-TuningEnv.ps1 (issue #2779),
# which derives the grant from YOUR corpus and clamps it to your host. That
# script is the only number to use; nothing in this comment is copyable.
#
# RECLAMATION PEAK. Steady state excludes the cost of the operation a box
# must perform in order to REACH steady state, and that omission is what
# made the withdrawn figure unsafe. On an 8,224-file index whose WAL
# garbage collection had been blocked, the pass that finally released the
# backlog ran at 450-590% CPU and drove the container working set to
# 13.81 GiB - above the 12 GiB that was being recommended, and above the
# 13.26 GiB that scripts/New-TuningEnv.ps1 DERIVES for this same corpus -
# before falling back to 10.55 GiB (issue #3252). At a 12 GiB limit the
# same corpus crash-looped twice in 16 minutes (129 then 304
# OutOfMemoryException, ExitCode=0, OOMKilled=false, /health/ready 503);
# at 18 GiB it completed clean with zero of either.
#
# 13.81 IS A FLOOR ON THE PEAK, NOT THE PEAK. It is the largest of 16
# samples taken about 103 seconds apart, so the true maximum is at least
# this and may be higher; nothing was sampling between those points. Quote
# it as a lower bound. The fall to 10.55 GiB afterwards is the useful other
# half of the reading: a 3.26 GiB give-back is what distinguishes a bounded
# transient from a leak, which would not have receded.
#
# So an operator upgrading INTO a WAL GC fix needs more headroom than one
# already running healthy, and needs it ONCE. The peak is a MIGRATION cost
# paid the first time a backlog of stuck WAL is released, not a permanent
# requirement: provision for it for that run, then re-derive.
#
# DO NOT ASSUME THE DERIVED GRANT COVERS THIS. It may not. On the corpus
# above the derivation yields 13.26 GiB and the observed peak exceeded it
# by about 568 MiB, so the script's 20% headroom band is the only thing
# standing between a migration burst and the ceiling, and this burst was
# larger than the band. Grant ABOVE the derived figure for the migration
# run specifically.
#
# THAT 568 MiB IS PINNED TO ONE CORPUS AND IS NOT A CONSTANT SHORTFALL.
# The derived grant is a function of file count, so it MOVES: the same
# repository measured at 8,239 files derives 13.26 GiB and at 8,846 files
# derives 13.97 GiB. Do not read a later derivation that happens to exceed
# 13.81 as evidence that the gap has closed - a larger corpus raises the
# reclamation peak too, and only the grant side of that comparison has
# been re-measured. Compare a peak and a grant only at the SAME corpus.
#
# READ THE TWO KINDS OF FIGURE CORRECTLY, BECAUSE THE COMPARISON JUST MADE
# IS WEAKER THAN IT LOOKS AND THE DIRECTION OF THE ERROR MATTERS.
#
# First, a working set measured under a generous cap is an UPPER BOUND on
# need, not a requirement - .NET collects less eagerly the further it sits
# from its ceiling, so 13.81 GiB observed at an 18 GiB grant does not
# establish that 13.81 GiB is NEEDED. A peak cannot be transported across
# caps: at 18 GiB the collector was working against a 13.5 GiB managed
# ceiling, and at 13.26 GiB it would be working against 9.94 GiB and
# collecting far harder, far earlier. The 13.26 GiB trajectory was never
# run and is not the one observed.
#
# Second, and pulling the OTHER way, peak RSS is not even the quantity a
# grant is tested against. The grant bounds RSS; the GC hard limit is what
# THROWS, and it binds FIRST at about 75% of the grant. A box at 13.26 GiB
# would fault against 9.94 GiB of managed heap long before RSS could reach
# 13.26, so "the peak exceeded the grant" UNDERSTATES the exposure rather
# than overstating it.
#
# WHAT IS ACTUALLY ESTABLISHED, stated on the plane that throws:
#
# measured 12 GiB grant -> 9.00 GiB managed ceiling -> CRASH LOOP.
# measured 18 GiB grant -> 13.5 GiB managed ceiling -> clean.
# NOT measured, in either direction: anything at all at 13.26 GiB.
#
# The derived grant therefore offers a managed ceiling only about 10%
# above one that demonstrably crash-looped on this corpus. That is the
# reason to provision above it for a migration run - a thin margin over a
# measured failure - and NOT the raw RSS comparison, which compares a
# number produced under one cap against a different cap.
#
# THE GRANT IS NOT THE CEILING THAT THROWS, AND THAT IS WHY 12 LOOKED LIKE
# ENOUGH. Under a container memory limit .NET applies its default
# GCHeapHardLimitPercent to the grant, so the managed heap ceiling is about
# 75% of what you grant: a 12 GiB grant yields a 9 GiB ceiling and an
# 18 GiB grant yields 13.5 GiB. Every LIMIT named above is a grant and
# every WORKING SET named above is container RSS; neither is comparable
# with the other until you apply that 75%.
#
# No memory limit is declared here on purpose, for the same reason as the
# CPU ceiling above: the right number is a property of your corpus, not of
# this file. Derive yours with scripts/New-TuningEnv.ps1 rather than
# copying any figure from this comment.
#
# THAT DERIVATION IS DEPLOY-TIME AND HOST-SPECIFIC BY CONSTRUCTION, SO IT
# IS THE BEST ADVICE AVAILABLE TODAY RATHER THAN A SETTLED ANSWER. Its
# constants were fitted against one corpus on one host and are re-derived
# by nobody afterwards. It also goes stale in place: its only corpus input
# is the indexed file count, so adding a repository to the workspace or
# removing one moves the requirement without moving the grant, and nothing
# signals the drift. Adapting sizing to the granted resources AT RUNTIME,
# instead of predicting it at deploy time, is tracked in issue #3255. That
# work depends on issue #3133 below, because a runtime that cannot observe
# its own approach to the ceiling cannot adapt to it.
#
# The failure mode is why this is worth stating. A cgroup limit becomes the
# .NET GC heap HARD limit, so the process is not OOM-killed: there is no
# restart, no exit code and no resource event. Instead the allocation for a
# leaf-snapshot deserialisation fails, the snapshot read fails, and the leaf
# activates COLD and replays its whole WAL window - which raises pressure
# further. What an operator sees is a STORAGE error, so the box looks slow
# and flaky rather than short of memory.
#
# WATCH A LEADING INDICATOR, NOT ONLY THE LAGGING ONE.
# orleans.lattice.leaf.snapshot.load_failures with reason=resource_exhausted
# names the cause directly, but its precondition is that an allocation has
# ALREADY failed, so it cannot warn of an approach - only report an
# arrival. These three engage BEFORE exhaustion and are what to alert on:
#
# lattice_repocontext_heap_committed_bytes
# / lattice_repocontext_heap_limit_bytes
# Heap-ceiling adherence: the direct approach signal, already
# published by this image. Read
# lattice_repocontext_heap_high_load_threshold_reachable FIRST - a 0
# there means the runtime's own pressure threshold sits ABOVE the
# limit and can never fire (issue #3133), which makes this ratio the
# only pressure reading you have. That 0 is not a misconfiguration
# you can tune away: the threshold is published at 90% of the cgroup
# limit while the GC hard limit binds at 75% of it, so the ratio is
# 1.2 at EVERY grant. It has been confirmed at 12 GiB and at 18 GiB
# with byte-exact matching percentages, so it is scale-invariant.
#
# orleans.lattice.wal.replay.permit_adaptations
# with outcome=withheld, trigger=occupancy
# The PROACTIVE replay-concurrency backpressure, which withholds a
# permit at 75% managed-heap occupancy and restores it below 60%. A
# rising withheld rate is the silo reducing concurrency to stay under
# the ceiling, which is the mechanism working rather than a fault.
#
# orleans.lattice.leaf.snapshot.hydration_admissions with outcome=queued
# Cold-start hydrations being serialised rather than allowed to
# exhaust the heap together.
#
# All three are zero-primed, so a flat zero is a MEASURED zero and an
# ABSENT series means the running image predates the instrument.
# GARBAGE COLLECTION. Follows directly from the memory note above, because
# the hazard is the PAIR (issue #2596): this host uses Server GC (the Web SDK
# emits System.GC.Server=true), and changing that would put a multi-GiB heap under
# workstation collection, which walks a single heap with effectively single-threaded blocking
# gen2 phases. That is the right default for a small heap and the wrong one
# here. Once the heap is multi-GiB one collection walks all of it on one
# thread with every other thread in the process suspended. A deployment of
# this host at about 11 GiB resident had the runtime attribute a 252 second
# pause to the collector, against a 30 second Orleans request timeout.
#
# DOTNET_gcServer: "1"
# DOTNET_GCHeapCount: "<hex; see below>"
#
# Neither is set here, for the same reason as the two settings above: the
# right heap count is a property of your CPU grant, not of this file.
#
# DOTNET_GCHeapCount is what makes the first line safe to set at all. Server
# GC otherwise creates one heap per processor, and Environment.ProcessorCount
# is what DOTNET_PROCESSOR_COUNT overrides - the same variable the replay
# ceiling above discusses. That variable may legitimately be pinned ABOVE
# this container's CPU grant to hold the replay gate's permits, and reusing
# it as a heap count would then create one heap per phantom processor on a
# heap already at its ceiling. One variable, two jobs, opposite
# requirements. DOTNET_GCHeapCount separates them: derive it from the
# container's actual CPU grant, independently of DOTNET_PROCESSOR_COUNT.
#
# WRITE THE HEAP COUNT IN HEXADECIMAL. The collector reads its numeric
# ENVIRONMENT variables as hex, while the same settings in
# runtimeconfig.json are decimal. So "10" asks for 16 heaps and "16" asks
# for 22, silently and with no error. Values below 10 read the same either
# way, which is what makes this so easy to miss on a small box and then get
# wrong on a large one. Prefer an explicit 0x prefix.
#
# Verify rather than assume: this host reports GC.Mode, GC.HeapCount (the
# figure the collector RESOLVED, not the one declared) and its memory
# ceiling in the effective-configuration report at startup, and warns there
# when the combination is the hazardous one. Check the resolved heap count
# against what you wrote; a declaration that was misread or ignored is
# visible in that line rather than only in a latency graph.
#
# The claim here is narrow on purpose. This removes the class of pause that
# is multi-minute, process-wide, and attributed to the collector by the
# runtime itself. It is NOT a general remedy for stalls: measurement of the
# same container found collector pauses accounted for under a third of long
# silences and did not explain the largest timeout burst at all. Pauses the
# runtime does not attribute to the collector need their own diagnosis.
# The read-only workspace root. Repositories are registered at runtime by
# the client through the repocontext_add_repo tool, and every added path
# must resolve under this root - a path escaping it (via `..` or a symlink)
# is refused. Mount a broad parent at /workspace (below) and add individual
# repositories beneath it on demand.
LATTICE_WORKSPACE_ROOT: /workspace
# Point the default embedding provider at the separate companion container.
LATTICE_EMBEDDING_ENDPOINT: http://embedder:9000
# Background-indexing cadence. A short self-index tick and a short reconcile
# interval (with no jitter) make the periodic content reconcile effectively
# continuous: the walk's directory-modification-time pruning keeps each
# reconcile cheap, so edits and deletions are picked up within seconds. The
# full-walk interval bounds how long an in-place content edit (which does not
# bump a directory's modification time and is invisible to pruning) can go
# unnoticed before a forced full sweep re-stats every file. Both periodic
# deadlines below are declared in wall clock but counted in reconcile passes
# (interval / (reconcile interval + jitter), rounded up, minimum one), so they
# still hold when a pass takes longer than its own scheduled spacing - which
# is the normal case on a large repository, and which silently defeated the
# previous wall-clock deadlines. Here that is 24 passes to a full walk and
# 60 to an embedding gap scan.
LATTICE_SELFINDEX_TICK_SECONDS: "5"
LATTICE_RECONCILE_INTERVAL_SECONDS: "5"
LATTICE_RECONCILE_JITTER_SECONDS: "0"
LATTICE_FULL_WALK_INTERVAL_SECONDS: "120"
# How often a pass re-probes every unchanged file for a missing vector. The
# probe is two membership reads per indexed source, so it dominates a pass on
# a converged repository. Spacing it costs no healing latency: the self-index
# grain's out-of-band paged sweep forces an immediate in-pass scan whenever it
# actually finds a gap.
LATTICE_EMBEDDING_GAP_SCAN_INTERVAL_SECONDS: "300"
# The container's stop_grace_period, DECLARED to the process that has to fit
# inside it. The host derives its own shutdown budget from this (issue #2402)
# instead of holding a second, independent constant: 240s here yields a 180s
# budget, and the remainder is reserve for the host to unwind and report a
# cut-short drain before SIGKILL.
#
# 240s replaces the 120s this file shipped with, which derived a 90s budget
# (issue #3304). That budget was below the measured requirement, not near it:
# two consecutive real drains of this deployment took 89.7s and 91.9s, which
# is 99.7% and 102% of 90s, so a stop was expected to be abandoned and to exit
# 70. The host said so itself before the first of them, projecting a 112.2s
# drain from 3334 resident activations at the 33.7ms each the previous drain
# had measured. 240s is roughly 2x the 91.9s observed, because the host's own
# "at least 150s" advice is explicitly a no-headroom figure and the resident
# set grows.
#
# It must EQUAL the stop_grace_period set on this same service below, and
# RepoContextComposeShutdownBudgetTests asserts exactly that. The adjacency is
# a convention, not an enforcement: the process cannot read the real grace
# period from inside the container, so a value edited here but not below (or
# vice versa) leaves the budget derived from a stale figure, and the dangerous
# direction - declaring MORE than the container actually grants - is
# undetectable at run time and reintroduces the silent kill of issue #2389.
# Change the two together, always.
#
# This is a SAMPLE DEPLOYMENT'S DECLARATION and not the library default. The
# default behind RepoContextShutdownBudget.DefaultStopGracePeriod stays at
# 120s deliberately: it applies to a deployment that declares nothing, which
# may well be running under Docker's own 10s grace period, and a budget above
# the real grant is strictly worse than a small one - it arms the overrun
# alarm for an instant the process never reaches, so the drain-ABANDONED line
# is never emitted. The two numbers are free to differ and now do.
LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD: "240s"
# --------------------------------------------------------------------
# Durable agent-memory backup (issue #2602).
#
# WHAT THIS PROTECTS. The repo-context-memory tree holds the decisions,
# gotchas, conventions and glossary entries agents write across
# sessions. Unlike the code-index trees it is NOT derivable from the
# workspace: destroyed, it is gone. On 2026-09-10 a routine
# "docker compose down -v", intended only to clear the code index,
# destroyed several hundred such entries with no copy anywhere.
#
# WHAT IS BACKED UP. Only repo-context-memory, by name. The code-index
# trees are deliberately excluded: they are orders of magnitude larger
# and can be rebuilt by re-adding the repository, so capturing them
# would fill the sink with the one thing that does not need protecting
# and age out the one thing that does.
#
# WHERE IT GOES. The azurite-backup-sink service below, whose storage
# is a HOST BIND MOUNT and therefore not a compose-managed volume, so
# "docker compose down -v" cannot reach it. See that service for the
# precise list of gestures this does and does not survive.
#
# Presence of the connection string is what enables backup. Unset it
# and /health/backup reports DISABLED while the backup state gauge
# publishes state 0, rather than the deployment being silently
# indistinguishable from a backed-up one. This is the fixed, public
# Azurite development account, which is
# published in Microsoft's own documentation and is not a secret.
LATTICE_BACKUP_BLOB_CONNECTION_STRING: "DefaultEndpointsProtocol=http;AccountName=devstoreaccount1;AccountKey=Eby8vdM02xNOcqFlqUwJPLlmEtlCDXJ1OUzFT50uSRZ6IFsuFq2UVErCz4I6tq/K1SZFPTOtr/KBHBeksoGMGw==;BlobEndpoint=http://azurite-backup-sink:10000/devstoreaccount1;"
# Cadence. An initial FULL capture runs at startup - manifest
# validation rejects an incremental with no base, so a full must exist
# first - and incrementals follow on this interval. Both values are the
# library defaults, restated here so the configured cadence is visible
# in the deployment and not only in code.
LATTICE_BACKUP_INCREMENTAL_MINUTES: "60"
LATTICE_BACKUP_FULL_HOURS: "24"
# Retention keeps a backup satisfying EITHER bound, and always
# preserves the base chain a retained increment depends on.
LATTICE_BACKUP_RETENTION_KEEP_LAST: "60"
LATTICE_BACKUP_RETENTION_MAX_AGE_DAYS: "14"
# OPERATOR-DRIVEN RESTORE. Left unset. Set it to a specific backup id
# and restart to replay that backup into the memory tree before the
# capture cadence starts, then UNSET it again. There is no latest
# value and no automatic restore-on-empty: a restore is an explicit
# human act. The replay merges by hybrid-logical-clock order, so it
# cannot clobber entries written since the backup was captured.
#
# LATTICE_BACKUP_RESTORE_BACKUP_ID: "<backup-id>"
# --------------------------------------------------------------------
ports:
# The MCP streamable-HTTP endpoint (and the health probes) on the host.
#
# The HOST side is overridable because 8080 is one of the most commonly
# occupied ports on a developer machine, and a collision surfaces only as
# an opaque bind failure at `up` time. Set REPOCONTEXT_PORT (in .env or the
# environment) to publish elsewhere:
#
# REPOCONTEXT_PORT=18080 docker compose up -d
#
# The CONTAINER side stays 8080 and must keep matching LATTICE_MCP_PORT
# above - only the published port moves, so the health probes and the MCP
# endpoint are then at http://localhost:$REPOCONTEXT_PORT. The default is
# deliberately unchanged, so existing clients keep working untouched.
- "${REPOCONTEXT_PORT:-8080}:8080"
volumes:
# Durable state survives restart / recreation / image upgrade.
#
# It does NOT survive `docker compose down -v`, and that includes the agent
# memory that lives in here alongside the rebuildable index - see the
# LATTICE_REPOCONTEXT_MEMORY_ARCHIVE_DIR comment above for why the two
# cannot be split across volumes, and what is done instead.
- repocontext-data:/data
# The agent-memory archive (issue #2601). A BIND MOUNT, deliberately, and
# the distinction is the entire protection: `docker compose down -v`
# removes the named volumes a project DECLARES, and a bind mount is not
# one of them, so this directory outlives the gesture that destroys /data
# above. Declaring a second named volume would not have worked - `-v`
# takes every declared volume, not the first one.
#
# It auto-creates on the host at `up`, so there is no `docker volume
# create` step to forget - and that auto-creation is precisely why this
# variable HAS NO DEFAULT (issue #2627).
#
# A relative default resolves against the directory `docker compose` was
# INVOKED FROM, not against this file, and compose then creates whatever
# that produced rather than refusing. The documented gate procedure
# composes from the candidate git WORKTREE on purpose, so the previous
# default `./memory-archive` put the only surviving copy of durable agent
# memory inside a directory that `git worktree remove` deletes. That
# deployment is indistinguishable from a correct one by every observable
# signal: the archive is written, it rotates, it is current, and it
# verifies by decoded content. There is no channel to read better, so
# reporting the resolved path would not have been enough on its own.
#
# Refusing to interpolate is therefore the fix rather than a stricter
# default: any default at all is a path nobody chose. `${VAR:?...}` fails
# the compose command with the message below - before a container starts
# and before a directory is created - so the operator must state where
# their last copy of durable memory lives. Copy `.env.example` to `.env`
# in this directory and compose loads it on every command, including
# `down`, `ps` and `logs`.
#
# Set it to an ABSOLUTE path outside every git checkout and worktree. The
# provenance check verifies that (check 5 of 7 in
# scripts/Assert-ContainerProvenance.ps1), keyed on the ARCHIVE path
# alone - never on the compose directory, which is legitimately a worktree
# and which check 2 of 7 requires to be one.
#
# Do NOT point it inside the /data volume or at a path that resolves under
# it. The host detects that and says so at startup, because an archive
# there dies with the thing it is meant to survive.
- "${REPOCONTEXT_MEMORY_ARCHIVE_PATH:?set REPOCONTEXT_MEMORY_ARCHIVE_PATH to an ABSOLUTE host path outside every git checkout and worktree - copy .env.example to .env in samples/RepoContextContainer. This variable has no default because a relative one resolves against the compose invocation directory, which silently placed the only working memory backup inside an ephemeral git worktree (issue #2627).}:/memory-archive"
# The workspace the box can see, mounted READ-ONLY so the box can never
# mutate the code it ingests. Mount a broad PARENT directory here and let
# the client register individual repositories under it with
# repocontext_add_repo. Override with REPO_PATH; defaults to this repo's
# parent directory (so this repo is one registerable child, added at
# /workspace/<repo>). Note ../../.. resolves three levels above this compose
# file's directory (the repo root's parent), NOT the repo root itself.
- ${REPO_PATH:-../../..}:/workspace:ro
# Grain-liveness healthcheck (issue #2666). Before this, the repocontext
# service PUBLISHED no healthcheck at all while CONSUMING service_healthy
# from its dependencies above: the two services that can report health were
# green through an outage in which the silo itself vanished, and it had to be
# found by a human reading `docker ps` by eye.
#
# The runtime image is chiseled and shell-less (no curl/wget/nc), so the test
# is the host binary re-invoked with --healthcheck, exactly as the embedder
# service does. That self-probe GETs the host's own /health/silo endpoint,
# which - unlike /health/live (always green) and the latched /health/ready -
# re-exercises silo membership and a trivial grain call on EVERY probe. So it
# goes red for a silo that died after reaching readiness, which is the whole
# point: a probe that only proves Kestrel is listening cannot.
#
# start_period is sized for REAL silo startup (cluster join + WAL replay
# warmup), which is far longer than a stateless service's. Docker's health
# model is two-valued (exit 0/1); "starting" is realised by this window
# holding early non-zero exits out of the retry tally.
#
# What getting it too short actually costs is ACCURATE REPORTING, not
# stability (issue #2906). A Docker restart policy acts on process EXIT and
# never reads health, so the restart: unless-stopped set below does NOT
# restart an unhealthy-but-running container, and a silo still joining
# cannot crash-loop because of its health verdict. Health-triggered restart
# is a Swarm feature (an unhealthy task is rescheduled) and a Kubernetes one
# (a failing livenessProbe restarts the container); in neither case is it
# restart:. Do not re-derive the crash-loop model from this window's size.
#
# The real cost of too short a window is that a NORMAL BOOT gets reported
# unhealthy: docker ps shows it, the health log records it, and since issue
# #2905 the verdict is published onto the /metrics scrape, so it would fire
# a real alert on every start. An alert that cries wolf on every start is
# how a true wedge comes to be ignored, which is precisely the failure that
# issue #2868 documents. The window buys a truthful signal, not stability.
# The endpoint still REPORTS the three-way verdict (healthy/starting/unhealthy)
# in its body for observability; the self-probe echoes it to the health log.
healthcheck:
test: ["CMD", "dotnet", "Orleans.Lattice.Api.Mcp.RepoContext.Host.dll", "--healthcheck"]
interval: 15s
timeout: 10s
retries: 5
start_period: 180s
restart: unless-stopped
# Run PID 1 as an init process (issue #2576). Docker bind-mounts its own
# static `docker-init` binary in and runs it as PID 1, so this needs nothing
# in the distroless, shell-less runtime image and changes no application
# code: the host becomes a CHILD of PID 1 rather than PID 1 itself.
#
# Two things follow, and both were observed to matter during the epic #2368
# gate runs, where a container reached a state in which neither
# `docker kill` nor `docker rm -f` would reap PID 1 and it had to be
# SIGKILLed - taking the run's time-to-ready measurement with it:
#
# * ORPHAN REAPING. PID 1 inherits every orphaned descendant and must
# wait() on it. The .NET host installs no such reaper, so anything it
# orphans accumulates as a zombie for the life of the container.
# * SIGNAL SEMANTICS. The kernel applies NO default action to a signal
# delivered to PID 1 that PID 1 has installed no handler for. A process
# that is merely a well-behaved application when run as a child can
# therefore become unkillable-by-SIGTERM purely by being PID 1. tini
# always installs handlers and forwards, so signalling is never a
# property of the application's own handler registration.
#
# This is INDEPENDENT of stop_grace_period below. The grace period decides
# how long the drain is allowed; init decides whether the SIGTERM that
# starts the drain is honoured at all, and whether the container's process
# table stays clean while it runs. Setting one and not the other leaves the
# other half of the failure in place - which is exactly the state this file
# was in when #2576 was raised, with a fully considered 120s grace period
# sitting above a PID 1 that was not an init process.
init: true
# How long Docker waits between SIGTERM and SIGKILL. This MUST exceed the
# host's own shutdown budget or that budget is dead configuration: the host
# asks for the budget it derives from the declaration above to deactivate the
# silo and flush the WAL commit-log, and Docker's default grants 10s, so
# every `docker compose stop` / `restart` / `down` is a crash teardown and
# the graceful deactivation path never completes. That was issue #2389, and
# it was invisible because a SIGKILLed container reports only an exit code
# that the next start overwrites.
#
# 240s derives a 180s budget. It replaces the 120s / 90s pair this file
# shipped with, which was below the measured requirement rather than near it:
# two consecutive real drains took 89.7s and 91.9s against that 90s budget
# (issue #3304). The earlier reasoning for 120s was that it cleared the host
# budget with 30s to spare, and that reasoning is intact - what was wrong was
# the budget it cleared, not the margin above it.
#
# The bound to clear is still the HOST budget rather than any one corpus's
# drain time, and the measurements are why. A rig holding a 400-file
# synthetic index spent 67.2s in Orleans' GrainDeactivation stage alone, and
# the same rig earlier drained the same corpus in 33.9s before its vector
# trees had landed. Drain time scales with resident state.
#
# Note what that pair implies, because it is the trap this value has to
# survive: a `stop_grace_period` tuned to an observed drain would have been
# set to something like 45s after the first measurement and would then have
# started killing teardowns as the index grew, reintroducing #2389 silently
# and with a number that looked carefully chosen. Clearing the host budget is
# what makes the value stable as the corpus grows. 240s is set at roughly 2x
# the 91.9s observed for the same reason: the smallest grant that admits a
# measured drain grants no headroom, and the host says so when it prints one.
#
# THIS IS NOW THE ONLY NUMBER TO CHANGE. The host derives its own budget from
# the grant declared in LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD above (issue
# #2402), so raising the grace period raises the budget with it and a budget
# above the grant is unrepresentable. Keep the two in step: they are two
# halves of one statement, and the process cannot read this line.
# RepoContextComposeShutdownBudgetTests asserts they are equal.
#
# The drain is observable rather than inferred. The host logs
# "RepoContext drain complete in <n>s" once every hosted service has stopped,
# so `docker logs` after a stop tells you the measured requirement directly;
# if that line is ABSENT, the container was killed mid-drain and this value is
# too small.
stop_grace_period: 240s
# OPT-IN CPU PINNING - UNSET BY DEFAULT, AND DELIBERATELY SO (issue #2623).
#
# DO NOT SWITCH THIS ON IN THE MIDDLE OF A MEASUREMENT. Epic #2368 adopted a
# precondition that no service configuration may change between a gate run's
# T0 and its final scrape, after a mid-run service recreation voided gate
# run 3 for every criterion that spanned it. A change of this kind lands
# BEFORE a run's T0 and ALONE, or not at all. This knob exists unset so that
# enabling it is an ACT, on the record, at a moment somebody chose.
#
# WHAT IT FIXES. A container given a fractional CPU quota (a `cpus` grant
# such as 4.0, which docker-compose.tuning.yml reads from `.env`) and no
# cpuset is ENTITLED to 4 CPUs but VISIBLE on all of them - 16 on the host
# this was measured on. CFS bandwidth
# control vends that quota to per-CPU run queues in 5 ms slices, and a
# thread waking on a run queue draws a whole slice whether it then runs for
# 5 ms or 5 us; the unused remainder is returned only lazily. Threads
# scattered across many run queues therefore exhaust the quota by
# RESERVATION rather than by execution, and the kernel throttles the cgroup
# while its actual utilisation is a small fraction of its entitlement. The
# effect scales with the number of run queues threads can land on, which is
# the VISIBLE CPU count, not the quota. Sizing the cpuset to the grant makes
# those two numbers equal.
#
# MEASURED, not inferred. Throwaway alpine containers, identical synthetic
# load, identical `--cpus=4`, varying ONLY `--cpuset-cpus`:
#
# periods throttled ratio quota used
# cpuset 0-15 907 49 5.4% ~17%
# cpuset 0-3 953 0 0.0% ~26%
#
# The load confound is INVERTED there, which is what makes it decisive: the
# arm offering MORE work throttled ZERO, and the arm offering LESS throttled
# at 17% of its quota. Utilisation cannot produce that ordering; scatter
# can. A heavier first experiment reached 47.7% against 0.0% on the same
# single variable.
#
# WHAT IT BUYS, WITHOUT OVERCLAIM. Latency jitter and scheduling
# determinism. It licenses NO throughput claim: the measurement is synthetic
# load on alpine, and the transfer to this silo and to the ONNX embedder is
# an inference. And do NOT read a throttle ratio as a measure of thread-pool
# oversubscription - it answers a different question and has a high floor
# (an idle container measured ~31% on this host). See the runbook section
# "CPU scatter under a fractional quota".
#
# HOW TO DERIVE A VALUE. One range per service, sized to the CEILING of that
# service's ACTUAL `cpus` grant in `.env` (REPOCONTEXT_CPUS / EMBEDDER_CPUS,
# which docker-compose.tuning.yml reads), and the ranges must NOT overlap
# each other or you trade throttling for contention. Their sum must also fit
# the host's logical CPU count: grants whose ceilings sum past it cannot both
# be pinned (issue #4188), and raising EMBEDDER_CPUS above what
# New-TuningEnv.ps1 derives means re-sizing EMBEDDER_CPUSET to match. On the
# 16-CPU host the tuning overlay was measured on, with grants of 6.0 and 4.0,
# that was REPOCONTEXT_CPUSET=0-5 and EMBEDDER_CPUSET=6-9, leaving 10-15 for
# the host. These are NOT portable: Docker refuses to start a container whose
# cpuset names a CPU the host does not have, so a copied value fails loudly
# on a smaller machine. Set them in `.env` beside REPO_PATH; see .env.example.
#
# NOT CHANGED BY THIS KNOB, AND NO LONGER PRESENT: DOTNET_PROCESSOR_COUNT in
# docker-compose.tuning.yml was a number above the effective grant and the
# same SHAPE, but it sized the WAL replay concurrency gate rather than a
# thread pool. This comment used to add that moving it was "a separate
# change with a separate blast radius that must be measured on its own".
# That was right, and #2779 was that change: the pin is gone and the gate is
# sized by LATTICE_WAL_MAX_CONCURRENT_REPLAYS (#2279), derived from the same
# CPU grant as this cpuset. See "The pool-sizing class" in the runbook.
cpuset: "${REPOCONTEXT_CPUSET:-}"
# Cap container log growth: the periodic self-index reconcile (every
# LATTICE_RECONCILE_INTERVAL_SECONDS) logs continuously, so the default
# unbounded json-file driver would grow without limit. Rotate at
# 20 MiB x 5 files (100 MiB ceiling); older lines are dropped, not lost to disk.
logging:
driver: json-file
options:
max-size: "20m"
max-file: "5"
# The DEDICATED sink the agent-memory backups are written to (issue #2602).
#
# WHY A SEPARATE SERVICE AND A SEPARATE STORE. A backup that lives in the
# same store it protects is not a backup: the gesture that destroys the
# primary destroys the copy at the same instant, while the deployment still
# reports itself as backed up. That is strictly worse than having no backup,
# because absence is visible and false protection is not. So this is its own
# Azurite instance on its own storage, and the host is configured to write
# here rather than to the default in-cluster sink.
#
# WHY A BIND MOUNT AND NOT A NAMED VOLUME. A down -v removes every named
# volume declared in this project. If this sink used one, the exact command
# that caused the loss this wiring exists to prevent would also destroy the
# backups. A bind mount is not a compose-managed volume and is not declared
# in the volumes: block below, so down -v has nothing to remove.
#
# An external: true volume also survives down -v and was considered. It was
# rejected because compose refuses to start until an operator has run
# docker volume create by hand, and a backup that needs a manual pre-step is
# exactly the backup that will not exist on the machine that needs it.
#
# WHAT THIS DOES AND DOES NOT SURVIVE. Stated precisely, because claiming
# more coverage than was implemented is the failure mode this issue is about.
#
# SURVIVES: docker compose down -v, down, stop, restart, rm, image
# rebuild, docker volume prune, docker system prune, and
# deleting the repocontext-data volume by hand.
#
# DOES NOT SURVIVE: deleting the host directory itself (by default
# ./backup-sink, next to this file), git clean -xdf in this
# directory, deleting or reformatting the host disk, or losing
# the machine. This is a same-host copy, NOT an off-site backup.
#
# To place it outside the repository, or on another disk, set
# REPOCONTEXT_BACKUP_PATH to an absolute path in .env or the environment. To
# make it genuinely off-host, point LATTICE_BACKUP_BLOB_CONNECTION_STRING at
# a real Azure Storage account instead and remove this service.
azurite-backup-sink:
image: mcr.microsoft.com/azure-storage/azurite:latest
# Blob only: backups are blobs, so the queue and table endpoints are not
# started. --skipApiVersionCheck keeps a newer Azure SDK from being
# rejected by an older emulator build.
command: azurite-blob --blobHost 0.0.0.0 --location /data --skipApiVersionCheck
volumes:
- ${REPOCONTEXT_BACKUP_PATH:-./backup-sink}:/data
ports:
# Published so an operator can point Azure Storage Explorer or azcopy at
# the sink during a restore without exec-ing into the container.
# Overridable because 11000 may already be taken on the host.
- "${REPOCONTEXT_BACKUP_SINK_PORT:-11000}:10000"
healthcheck:
test: ["CMD", "nc", "-z", "127.0.0.1", "10000"]
interval: 5s
timeout: 3s
retries: 20
start_period: 5s
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "20m"
max-file: "5"
volumes:
# Durable Lattice state (WAL + SQLite): the heart of "codebase memory in a box".
# This volume is separate from the host bind mount used by the service above,
# which is not declared here and is therefore not removed with named volumes.
#
# NOTE: the agent-memory backup sink is likewise NOT declared here. It is a
# host bind mount on the azurite-backup-sink service, for the same reason and
# by the same mechanism. Do not tidy either of them into a named volume. The
# two are separate and neither replaces the other: #2601's /memory-archive is
# an automatic restore-on-empty of the memory tree, while this sink holds
# scheduled, retained, manifest-tracked backups that only an operator
# restores.
repocontext-data: