Table of Contents

Graceful shutdown

This page documents Orleans.Lattice.Api.Mcp.RepoContext, which is unreleased, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at graceful-shutdown.md, and llms.txt lists every page.

Part of Container quickstart.

On SIGTERM (a docker stop or restart) the host flips readiness to not-ready first, then drains: the silo deactivates and the WAL commit-log flushes buffered records before exit, so an in-flight write is durable after restart.

That SIGTERM only arrives usefully because PID 1 is an init process. The compose service sets init: true, so Docker bind-mounts its own static docker-init binary and runs it as PID 1 with the host as its child. Two things follow, and both are properties of PID 1 specifically rather than of the application. The kernel applies no default action to a signal delivered to PID 1 for which PID 1 has installed no handler, so a process that is perfectly well behaved as a child can be unkillable by SIGTERM purely by being PID 1; and PID 1 inherits every orphaned descendant and must wait() on it, which the .NET host does not do. During the epic #2368 gate runs a container reached a state in which neither docker kill nor docker rm -f would reap PID 1 and it had to be SIGKILLed, which cost that run its drain and left the next run unbanked state to replay. That was issue #2576.

This is independent of the grace period below, and neither substitutes for the other: init decides whether the SIGTERM that starts the drain is honoured at all, and stop_grace_period decides how long the drain that follows is allowed to take. The file was for a time in exactly the half-fixed state that makes the point - a carefully derived 120s grace period sitting above a PID 1 that was not an init process. RepoContextComposeInitProcessTests asserts both settings on every compose file in the repository that runs this image, and RepoContextComposeShutdownBehaviourTests demonstrates the resulting behaviour end to end: it brings a stack up under its own project name, issues a plain docker compose stop with no -t and no -f, and asserts the container exits 0 rather than 137.

The budget for that drain belongs to the host, not to Docker, and this sample sets it to 180 seconds. The host sets HostOptions.ShutdownTimeout from the grace period the deployment declares, and Docker's own stop_grace_period defaults to 10 seconds. A budget the container will not grant is dead configuration: the two are enforced independently, the smaller one wins, and the process is SIGKILLed at 10s with the drain still in flight. That was issue #2389, and this compose file's stop_grace_period: 240s is what makes the 180s reachable. It has to exceed the host budget rather than merely exceed some measured drain time, because a larger index moves the drain but not the bound; RepoContextComposeShutdownBudgetTests asserts that relationship so the two values cannot drift apart unnoticed. Since issue #2402 the budget is not written down independently at all: it is derived from the grace period the deployment declares through LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD, so a budget larger than the grant cannot be configured. See Where the budget comes from.

The pair was 120s/90s until issue #3304 and that was not enough. Two consecutive real drains of the deployed container took 89.7s and 91.9s, which is 99.7% and 102% of the 90s budget a 120s grant derives - the second was abandoned and exited 70. The sample now declares 240s, deriving 180s, which the same drains consume half of. Note what the near-miss shows: at 99.7% the deployment was already failing and the only thing separating the two runs was 2.2 seconds of growth. A budget a drain just fits is not a budget that fits.

RepoContextHostBuilder.ShutdownBudget stayed at 90s, deliberately. It is the default applied to a deployment that declares no grant at all, and such a deployment may be running under Docker's 10s. Raising it would raise that deployment's budget without raising its grant by one second, which is strictly worse than leaving it alone for the reason below. The way a deployment raises its budget is by declaring its grant, which is what this sample does.

If you run this image under your own orchestration, you must grant the same budget there. Kubernetes has the identical trap under a different name: terminationGracePeriodSeconds defaults to 30s, which is a sixth of the 180s this sample derives - and less than the 90s an undeclared deployment assumes.

The drain is observable rather than inferred, so the budget can be derived instead of bisected. The host logs one line when a drain starts and one when it completes:

RepoContext drain started with ... resident activations: ... The host shutdown budget is 180s; ...
RepoContext drain complete in 33.9s, consuming 18.8% of the 180s host shutdown budget. ...

There are three outcomes and the log distinguishes all three, which it did not before issue #2397.

What you see What happened What to do
Start line, no completion line The container was killed mid-drain. The grace period is smaller than the drain. Raise stop_grace_period above the host budget. This was issue #2389.
drain complete ... consuming NN% at Information The drain finished with headroom. NN% is what your corpus needs. Nothing.
drain complete ... consuming NN% at Warning The drain finished, but consumed 70% or more of the budget. Treat as a lead indicator: the next growth in the index may push it over.
drain ABANDONED after NNs at Error, and the container exits 70 The host stopped waiting. Deactivation was abandoned part-way. Raise stop_grace_period and the LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD that declares it, together and to the same value. This is what issue #3304 recorded against the shipped 120s/90s pair.

The last row is the one that needed issue #2397. A widespread belief - stated in an earlier revision of this very document - is that ApplicationStopped fires only after every hosted service has stopped, which would make the completion line self-evidently trustworthy. It is not true. HostShutdownTimeoutBehaviourTests demonstrates the actual behaviour against a real generic host: when HostOptions.ShutdownTimeout expires, the host stops waiting for the services and raises ApplicationStopped anyway. Before #2397 the signal was bound to that event and to nothing else, so an abandoned drain emitted drain complete in 90.0s - a confident false positive, which is worse than the silence it was assumed to be. The overrun is now reported at Error, from an alarm armed when the drain starts, so it is emitted at the instant the budget expires rather than depending on a completion callback that may never arrive.

The abandoned drain also reports itself in the exit code

An Error line only helps somebody who is already reading the log. The layer that acts on a stopped container automatically - your orchestrator - does not read logs, it reads the exit code, and before issue #2401 an abandoned drain did not reliably produce a distinctive one.

Measured against a real generic host rather than assumed, the pre-#2401 outcome was not merely zero, it was undetermined, and which of two outcomes you got depended on an internal choice of the silo's hosted service:

  • if the service absorbed the cancellation and returned (a force-stop), RunAsync returned normally and nothing assigned an exit code, so the process exited 0 - an abandoned drain recorded as a clean stop;
  • if the service rethrew it, the exception escaped RunAsync unhandled and the process aborted, which is indistinguishable from a genuine crash.

So the host now assigns the code itself, at the moment the overrun latches:

Exit code Meaning
0 The drain completed inside the host shutdown budget.
70 The host shutdown budget expired and the drain was abandoned part-way, so leaf activations were torn down without banking their projection checkpoints.

70 is EX_SOFTWARE in the BSD sysexits.h convention. The convention is not something any orchestrator interprets, so the value's job is to be distinct and documented: it avoids 0, 1 and 2 (success, generic failure, shell misuse), Docker's reserved 125-127, and the whole 128 + signal band - which is where 137 (SIGKILL, the killed-mid-drain case of issue #2389) and 143 (SIGTERM) live, and those are precisely the neighbouring conditions this code exists to be told apart from.

Be clear about what the code does and does not change. It is an observability signal, not a restart control. This compose file runs the container under restart: unless-stopped, and Docker restarts on that policy regardless of exit code, so nothing here suppresses or triggers a restart. What changes is what is recorded, which is what an alert can be written against:

$ docker inspect --format '{{.State.ExitCode}}' "$(docker compose ps -a -q repocontext)"
70
$ docker ps -a --filter name=repocontext
... Exited (70) 12 seconds ago

Under Kubernetes the same container terminates with reason Error rather than Completed, so an abandoned drain becomes visible in kubectl get pod and in lastState.terminated.exitCode instead of looking like an ordinary graceful stop.

There is deliberately no configuration knob to turn this off. A switch restoring 0 would remove the evidence rather than the problem, and an operator who does not want the signal wants the drain to fit inside its budget instead.

Measured drains for scale, and they are worth reading carefully. The same 400-file rig drained in 33.9s before its vector trees had landed and in 67.2s once they had - so drain time scales with resident state, and the second figure was already three quarters of the 90s the host then allowed. This is why the value to clear is the host budget rather than an observed drain: a stop_grace_period tuned to the first measurement would have looked carefully chosen and would have begun killing teardowns as the index grew, reintroducing the defect silently. Issue #3304 is that prediction coming true at full scale: the same deployment, grown to 3,334 resident activations, drained in 89.7s and then 91.9s.

It also means the host budget itself is a finite resource, not merely a formality. If a drain exceeds the budget the host abandons it, and no stop_grace_period can rescue that on its own - the budget has to rise with it, which is why the two values move together.

Why the budget is not raised to some larger fixed number on principle

The obvious response to a drain at 74.7% of budget is to raise the budget. Issue #2397 investigated that and deliberately did not, because the measurements do not support any particular replacement value, and a value that is not supported is worse than none: it looks chosen.

What the instrumentation on a live, actively-indexing box shows is that the quantity driving drain time has no observed ceiling. Over a three-hour window that box logged 135 idle-deactivation sweeps whose sizes ranged from 1 to 4,418 activations, with the high-water mark still rising between successive readings taken minutes apart. Per-leaf persistence cost over the same period had a marginal mean of roughly 520 ms (orleans_lattice_leaf_write_duration), sustained at about 2.6x concurrency. A drain must flush the resident dirty set, so its duration tracks that set - and a fixed ceiling on an unbounded quantity is the wrong shape of fix regardless of which fixed value is chosen. Raising 90s to 150s or 300s moves the threshold without changing the failure mode.

That is why issue #2397 shipped the diagnostic and not the number, and why issue #2402 - which proposed raising the number - did not ship one either.

Issue #3304 did raise it, and that is not a reversal of this argument. The distinction is evidence, not size. #2397 and #2402 were asked to raise the budget against a drain that still fitted, with no measurement indicating any particular replacement; #3304 raised it against two consecutive drains that did not fit, one of which was abandoned. Raising a value that is demonstrably too small is arithmetic. What this section rejects is raising a value that is currently adequate in the hope of staying ahead of an unbounded quantity, and that is still rejected: 180s is not claimed to be a ceiling, and the forecast below exists precisely because it will not stay sufficient on its own.

Where the budget comes from, and why it is not a free parameter

Issue #2402 asked for the budget to be raised, or made adaptive from observed residency. Neither is honest here, and the reason is worth stating because it is the opposite of the intuition.

stop_grace_period is a hard ceiling imposed from outside the process. Docker sends SIGTERM and then SIGKILL at the grace period whatever the host is doing, and the host can neither read that value nor change it. So a budget set above the grant buys no drain time whatsoever. What it does instead is strictly worse than leaving it alone: it arms the overrun alarm for an instant the process never lives to reach, so the drain ABANDONED line - the only evidence a drain was cut short - is never emitted. Raising the budget past the grace period therefore reintroduces the silent teardown of issue #2389 by way of the change meant to prevent it. Deriving the budget from residency has the same defect with extra steps: it would climb straight past a grant nothing can see.

So the budget is derived from the quantity that genuinely bounds it. The deployment declares its grace period to the process through LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD, and the host takes 75% of it, or all but a two-second unwind reserve, whichever is smaller. The sample's declared 240s yields 180s; the 120s assumed when nothing is declared yields 90s, which is what the container ran with when that derivation landed, so nothing moved at the time. What changed is that there is now one number to set instead of two independent ones, and a budget exceeding its grace period can no longer be expressed. The 75% was calibrated to reproduce that original pair rather than measured, and it is unchanged by issue #3304 because what it was calibrated against is the relationship between a grant and its budget rather than either number. The constant reserve exists because the cost it covers - emitting one log line and flushing it - is roughly fixed, so a pure percentage would leave only a second at a four-second grace period.

The residual risk, stated plainly: the environment variable declares the grant, it is not the grant. A deployment that declares 240s while granting 20s derives a 180s budget under a 20s guillotine, and by the same premise that motivates all of this - the real grace period is unobservable from inside the container - the process cannot detect it. Writing the two values adjacently in the same compose service is the mitigation, and RepoContextComposeShutdownBudgetTests asserts they are equal in the sample. That adjacency is a convention, not an enforcement. Change the two together, always.

None of this bounds the resident activation set, and drain time still scales with it. If your own box reports the Error line, raising both values past your observed drain buys time; it does not fix the cause.

Knowing before the stop: the drain forecast

Everything above is discovered at shutdown, which is the worst moment to learn it. The drain ABANDONED line and the exit 70 are honest, but by the time either is emitted the state they were warning about has already been torn down unbanked. Issue #2598 is the case in point: gate run 2 of epic #2368 drained past 102s against a 90s budget, exited 70 exactly as designed, and the first anyone knew of it was the corpse.

The budget cannot be moved on its own, for the reason the section above gives: it is bounded by a grant the process cannot see, so raising it alone converts a loud failure into a silent one. Moving the pair is what issue #3304 eventually did once the evidence demanded it. What this section adds is that the mismatch is visible while the container is running, hours before anybody types docker stop, so the pair can be moved before a stop proves it necessary rather than after.

Two mechanisms supply that, and they are complementary:

1. The last drain is remembered across the restart. The host writes a drain-history.txt under its data root: a marker when a drain starts, replaced by the measured outcome when it finishes. A container that starts and finds a start marker with no outcome knows its predecessor was killed mid-drain, which is direct evidence the real grace period is smaller than the drain needed. That is the one fact the running process genuinely cannot observe about itself, and it is observable across a restart precisely because the file outlives the process.

2. Drain cost is projected from live residency. The host samples Orleans' own activation working set every 10 seconds, divides the last measured drain by the residency it was measured against to get a per-activation cost, and multiplies by residency. It keeps the samples of the trailing 10 minutes and projects both the instant reading and the peak of that window. When the peak projection exceeds the budget, the host says so at Error rather than waiting for a stop to prove it; a peak projection that fits is reported at Information.

Gate a stop on the windowed peak, never on the instant reading (issue #3628). The resident set swings by an order of magnitude within minutes: the hourly compaction and reclaim walks (issue #3607) drive it from about 550 to over 8,000 and back. A stop taken on an instant "fits" reading of 786 activations began its drain fifteen seconds later with 1,791 resident, and was abandoned at the budget. The peak over the window is the reading that says the current wave has passed, so lattice_repocontext_projected_drain_peak_seconds is the gauge to check before a planned stop, and the logged verdict follows it.

A projection from an abandoned drain is a floor. An abandoned drain's duration is the budget it was cut at, not the drain, so dividing it by the starting residency understates the cost. When the record also carries the stranded count (below), the duration is divided by the activations the drain actually got through, which measures that part and errs high, because the partial work on the stranded activations is charged to it. When it does not, the projection is reported as at least its value, at Warning if it fits and Error if it does not, and lattice_repocontext_projected_drain_lower_bound reads 1. A floor that exceeds the budget is a certain overrun; a floor that fits proves nothing.

The loss is recorded as well as the duration. An abandoned drain writes strandedActivations into drain-history.txt: the resident count at the instant the host stopped waiting, which is the set torn down without banking its projection checkpoints. The abandonment log line reports it too, but that line goes with the container when docker compose up recreates it; the file is what survives. The next start names it in the Exceeded or Unproven line. A record written before this field existed reads with the count absent, never as zero.

The record describes the last drain, not necessarily the last stop. A stop that never raised the host's stopping signal (the host sleeping, or Docker Desktop restarting underneath the container) writes nothing, so the next start reports the earlier drain. Check observedAtUtc in the file against the stop you think it describes.

The forecast verdict, which compares the last recorded drain with this process's budget, is reported once in the startup log. The live projection is reported on its first poll and after that only when the peak projection flips between fitting and exceeding the budget, so a residency dip between two waves does not clear an over-budget verdict until the peak has aged out of the window. The verdicts are:

Verdict Meaning Severity
NoHistory No drain has been measured yet on this volume. Information
Fits The last drain fitted the budget with headroom. Information
Thin The last drain consumed 70% or more of the budget. Warning
Exceeded The last drain did not fit. The next stop will abandon. Error
KilledMidDrain A previous process was killed with a drain in flight. Error
Unproven The last drain was abandoned under a smaller budget than this process has. Its duration is a floor, not a measurement, so nothing is predicted either way. Warning

Unproven exists because of issue #3305, and the alternative was worse than it looks. A recorded abandonment is a fact about the budget in force when it was recorded, and replaying it against a larger budget produced a verdict that contradicted its own evidence - a live container printed last drain took 91.9s and does not fit this process's 180s shutdown budget (51% of it), and advised raising a grant to 123s to a deployment already declaring 240s. Every figure was individually correct; the fit boolean was evaluated against the previous 90s budget while the percentage rendered from the current 180s one. But the truncated drain is not evidence of fitting either: it was cut short at 91.9s, so 91.9s is a lower bound on what a complete drain costs. Reporting Fits would have replaced a false alarm with a false all-clear. Unproven says what is actually known, and self-clears on the first clean stop.

The Exceeded and Unproven lines carry the grace period the measurement actually requires, computed by inverting the derivation, so the remedy is a value to copy rather than a number to guess. A KilledMidDrain line has no measured duration to invert, because the killed drain recorded none, so it names the grant and budget in force and asks for the service's stop_grace_period and LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD to be raised together. A drain measured at 102.1s reports a required grant of 137s, which is what LATTICE_REPOCONTEXT_STOP_GRACE_PERIOD and the service's stop_grace_period must both be raised to. On an Unproven line the figure is a floor on a floor and the line says so, because the grant already declared necessarily covers it - which is what makes it read as confirmation rather than as an instruction to reduce.

Two properties hold structurally since issue #3305, and both are pinned by RepoContextDrainForecastTests. A verdict of Exceeded is now equivalent to the percentage printed beside it being at least 100%, so a line cannot contradict itself. And the grant an Exceeded line advises is never below the grant already declared, so following the remedy can never be a reduction.

Those lines also distinguish a declared grant from an assumed one, because the remedy differs. If the grace period was declared and the evidence contradicts it, the declaration is wrong and must be raised. If it was merely assumed because the variable is unset, the deployment may already grant enough and simply never said so. The same distinction is carried in the effective-configuration dump since issue #2593; before that fix a defaulted 120s and a declared 120s printed identically, which is how epic #2368's gate run 2 came to record a grace period nobody had actually set.

Ten gauges expose the same state on /metrics, so this is alertable without log scraping:

Gauge Meaning
lattice_repocontext_shutdown_budget_seconds The host drain budget in force.
lattice_repocontext_stop_grace_period_declared 1 when the grant was declared, 0 when assumed.
lattice_repocontext_last_drain_seconds The last measured drain duration.
lattice_repocontext_drain_forecast The verdict above, as its numeric value.
lattice_repocontext_resident_activations Activations resident now.
lattice_repocontext_projected_drain_seconds Projected drain at current residency.
lattice_repocontext_required_stop_grace_period_seconds The grant that projection would need.
lattice_repocontext_resident_activations_peak The highest residency sampled over the trailing 10 minutes.
lattice_repocontext_projected_drain_peak_seconds Projected drain at that peak. Gate a planned stop on this.
lattice_repocontext_projected_drain_lower_bound 1 when the projections are floors because the last drain was abandoned, 0 when scaled from a measured cost.

lattice_repocontext_drain_forecast == 3 or lattice_repocontext_drain_forecast == 4 is the alert worth having: it fires on Exceeded and KilledMidDrain, the two verdicts that predict a failure, and on nothing else. Write it as that pair rather than as >= 3: Unproven is 5 and would be swept up by the inequality, which is exactly the false alarm issue #3305 was about. The enum's ordinals are wire format for this gauge, so Unproven was appended rather than inserted and 0-4 are unchanged.

What this does not do, stated plainly so it is not over-read. It does not make the drain fit. A projection that exceeds the budget is a warning that the next stop will abandon, not a repair of it, and the remedy is still to raise the grace period and the variable that declares it together. Every reading is best-effort: a residency count the runtime will not supply is reported as unavailable and never as zero, and a projection needs a prior measured drain, so a first-ever start forecasts NoHistory and offers no projection at all. The point is only that the failure now announces itself while there is still time to act on it.

Garbage-collector pause time

The host runs a multi-GiB heap by design, so a collector pause is a first-class explanation for a request timeout. The runtime already emits the accumulated pause total by name, and the effective-configuration report names it at startup, where it is necessarily near zero. Two counters make the quantity queryable over a window rather than readable only from a log scrape:

Counter Meaning
lattice_repocontext_gc_pause_seconds_total Accumulated seconds this process has spent suspended for garbage collection since it started.
lattice_repocontext_gc_collections_total Garbage collections completed since start, summed across every generation.

Read the two together; the first cannot be read alone. A pause total of zero is otherwise indistinguishable between a collector that has run without suspending the process measurably and a collector that has not run at all. With the count beside it, a zero on lattice_repocontext_gc_pause_seconds_total against a rising lattice_repocontext_gc_collections_total is a measured absence of pause, and both at zero means no collection has happened yet.

Both are observable counters, sampled at scrape time from cumulative runtime figures, so a scrape gap loses resolution rather than corrupting the series and both exist from process start rather than appearing on a first occurrence. A series that is absent rather than zero therefore means the host did not construct the meter, or the collector refused the series at one of its ceilings; it never means the process has not paused. Neither carries a tenant dimension: a collector pause is a property of the host process and belongs to no tenant's traffic.

rate(lattice_repocontext_gc_pause_seconds_total[5m]) is the reading worth alerting on, because it is the fraction of wall-clock the process spent suspended and is directly comparable with a request-latency series.

Heap ceiling and commitment

Subscribing the runtime meter supplies what the process is using. Nothing in it supplies the limit that usage is measured against, and that limit is what this epic's heap work is judged by: the fixes for cold activation of oversized leaves (#2765) and for resident hydrated leaf working set (#2767) both claim adherence to the heap hard limit as their primary effect, and before these gauges neither claim was checkable from inside the container at all. It had to be inferred from docker stats, which carries one aggregate number with no attribution and no history.

Gauge Meaning
lattice_repocontext_heap_limit_bytes Memory the garbage collector believes it may use, in bytes.
lattice_repocontext_heap_committed_bytes Memory committed by the collector as of its last collection, in bytes.
lattice_repocontext_heap_high_load_threshold_bytes Commitment at which the collector begins treating memory as under pressure, in bytes. Computed against total physical/cgroup memory, not against the limit, so it may sit above it.
lattice_repocontext_heap_high_load_threshold_reachable 1 when the threshold sits at or below the limit and can fire; 0 when it sits above and is dead.

lattice_repocontext_heap_committed_bytes / lattice_repocontext_heap_limit_bytes is heap-ceiling adherence directly. The two are published separately rather than pre-divided so a query can read either alone and neither can drift from the other. A hard limit is enforced against committed memory rather than against live heap size, which is why the numerator is commitment and not dotnet_gc_last_collection_heap_size.

The ceiling is read from the runtime, never configured here. It resolves to whichever limit actually binds - a container memory limit, a configured GC heap hard limit, or physical memory - at the moment of the read, so the same code reports the truth on a 12 GiB container and on a 56 GiB developer machine without being told which it is on. That matters more than it sounds: the resource knobs on this image have already been found to be transcriptions of one developer machine (issue #2779), and a ceiling hard-coded here would be another one, with the added defect that it would look like a measurement. It also means a changed memory grant is followed without a restart and without a configuration change.

The high-load threshold does not sit below the limit, and assuming it does hides a real exhaustion (#3133). The runtime computes the two against different denominators: the threshold is a fraction (90% by default) of total physical or cgroup memory, while the limit is the GC hard limit, which in a container defaults to 75% of that same figure. On the 12 GiB deployment that puts the threshold at 10.80 GiB and the limit at 9.00 GiB - the threshold is 1.80 GiB above the ceiling, so the hard limit binds first and the threshold can never be crossed. The process OOMs with the pressure signal still reading as uncrossed, which reads as "memory is not under pressure" and is exactly backwards; an alert written on the crossing is silently dead, returning no data rather than erroring. Because the ordering is a property of the deployment's hard-limit percentage rather than of the code, it is measured rather than asserted: read lattice_repocontext_heap_high_load_threshold_reachable instead of recomputing two percentages against the cgroup limit. A 0 there means the threshold series is real but unreachable, and heap-ceiling adherence is the only pressure reading you have. Zero semantics differ between the four, so they are stated separately. All four are observable gauges sampled at scrape time and published from process start, so an absent series means the host did not construct the meter or the collector refused it at one of its ceilings; it never means memory is unbounded. The limit and the threshold are populated by the runtime before any collection has run, so a zero on either is a fault to investigate rather than a reading. Committed bytes is carried by the last collection's figures and so is genuinely zero until the first collection: read it against lattice_repocontext_gc_collections_total exactly as the pause total is read, and a zero beside a rising collection count is the only form that means the process has committed nothing. The reachability gauge is the exception and inverts the rule: a zero there is its most important reading, not a fault - it is the measured statement that the threshold is unreachable. None of the four carries a tenant dimension.

Measured requirement and startup admission

The gauges above answer "how much is committed right now against the ceiling". They do not answer "how close did this deployment ever come", which is the question a sizing decision actually turns on, and an instantaneous gauge cannot answer it: the peak occurs during ingest, so a scrape landing before or after it reports a comfortable figure that says nothing about the margin really consumed. Four further instruments close that (#3255), and the same measurement feeds a startup admission check.

Instrument Meaning
lattice_repocontext_heap_committed_peak_bytes Highest commitment observed this run: the high-water mark of lattice_repocontext_heap_committed_bytes.
lattice_repocontext_heap_peak_occupancy_ratio Share of the granted ceiling consumed at that peak, as a fraction of 1.
lattice_repocontext_heap_exhaustion_events_total Managed OutOfMemoryException observed this run.
lattice_repocontext_heap_insufficient_limit_bytes Largest ceiling at which this deployment has ever recorded running out of memory. Absent when none is recorded.

Size the next grant from the ratio, not from the byte count. lattice_repocontext_heap_peak_occupancy_ratio is normalised against whatever ceiling this process was given, so it transfers between hosts where a byte count does not. It is derived at scrape time from a single reading rather than divided across two series, so its numerator and denominator are provably the same observation. A value approaching 1 means the grant is being fully consumed and exhaustion is near; it does not mean the process is efficient.

The peak is sampled, so it is a floor on the true peak. A spike falling entirely between two samples is not seen. It therefore errs low, which is the safe direction for every consumer: the admission check only escalates on it, so under-reporting costs a missed warning rather than a wrong refusal.

Startup refuses a grant this deployment has already measured to be insufficient. Before #3255, a grant that was too small was accepted and then crash-looped. Measured on one corpus at a 12 GiB grant: two restarts in 16 minutes, 129 then 304 OutOfMemoryException, ExitCode=0, OOMKilled=false, and /health/ready answering 503. It reached the operator as a STORAGE error reading grain state, that is, as flakiness rather than as insufficiency. The check compares the ceiling the runtime reports now against lattice_repocontext_heap_insufficient_limit_bytes carried in heap-history.txt on the data mount:

  • Refuse when a ceiling is recorded and the grant is no larger than it. The process exits with a message naming both numbers and the path the evidence is at.
  • Warn when the previous run's peak does not fit inside the new ceiling, or when the previous run never recorded stopping cleanly.
  • Admit otherwise, and on every failure to read.

A refusal rests only on a measurement, never on a model. There is no byte constant, no corpus model and no fraction in the check. The deploy-time sizing model in New-TuningEnv.ps1 is a one-point fit by its own account - two parameters identified by a single observation - and a refusal is an outage, so it may not rest on an unidentified parameter. The check also fails open at every other turn: no record, an unreadable or corrupt record, an unrecognised format version, or a runtime reporting no usable ceiling all admit, because a wrong refusal is an outage in a distroless container with no shell while a wrong admission costs exactly the crash-loop that already exists.

Overriding a refusal. Set LATTICE_REPOCONTEXT_HEAP_ADMISSION_OVERRIDE to the exact recorded byte count being disbelieved - the value of lattice_repocontext_heap_insufficient_limit_bytes, which the refusal message also prints. A boolean flag would be set once, forgotten, and would then suppress correct refusals for ever; echoing the number cannot be set by accident and is self-invalidating, because a later exhaustion at a larger ceiling changes the recorded figure and the stale override stops matching. An honoured override warns on every start it suppresses and is recorded into heap-history.txt, so a later reader can see the evidence was deliberately disbelieved. The likeliest reason to need it is that the indexed corpus has shrunk since the exhaustion was recorded, which the check cannot see.

What the exhaustion counter cannot see, and why a zero on it is not a clean bill of health. It counts managed OutOfMemoryException observed first-chance, which is the form the documented failure takes: the exceptions are caught and surface as STORAGE errors, so nothing terminates and no unhandled-exception path ever sees them. A cgroup out-of-memory kill is a different path entirely - SIGKILL, no exception, no opportunity to record anything - and is counted here as zero. The absence of a recorded exhaustion is therefore not evidence that the grant was sufficient. Read this counter beside the container's restart count and OOMKilled status, which is where that path is visible. The only residue a kill leaves inside the container is a previous run's record still reading Admitted, which is ambiguous - a host reboot and an orchestrator rescheduling leave the same - so it warns and never refuses.

Zero semantics. All four are observable, published from process start with a real value, so an absent series means the host did not construct the meter and never that nothing happened; a zero on the exhaustion counter is a measured absence of managed exhaustion so far this run. lattice_repocontext_heap_insufficient_limit_bytes is the deliberate exception and publishes no measurement at all rather than a zero when nothing is recorded, because a zero there would read as an exhaustion at a ceiling of zero bytes, which no grant is smaller than. None of the four carries a tenant dimension.

Agent-memory backup protection

The container captures the durable agent-memory tree to an external sink on a cadence. Until issue #2640 that subsystem published no series at all, so a deployment on which every capture threw could not be alerted on by any rule: the evidence existed only as a log line and as a status object nothing exported. Five instruments close that, all of them observable and therefore published from process start with a real value rather than appearing on a first occurrence:

Instrument Meaning
lattice_repocontext_backup_state Protection as a numeric state (see below). Only 5 means the tree is captured as configured.
lattice_repocontext_backup_captures_total Captures this container has completed since it started.
lattice_repocontext_backup_last_full_entries Key descriptors carried by the most recent full capture's manifest.
lattice_repocontext_backup_sink_backups Backups of the tree found in the external sink, or -1 when the sink has not been enumerated.
lattice_repocontext_backup_incremental_fallbacks_total Incremental captures the capture service silently promoted to full ones.

The state values are 0 disabled, 1 failing with nothing ever captured, 2 configured but nothing captured yet, 3 capturing but the last full capture described zero entries, 4 failing after an earlier success, and 5 protected.

They are identifiers, not a severity scale, so alert on lattice_repocontext_backup_state != 5 rather than on a threshold. Only the last value means the tree is protected as configured; the rest are distinct conditions with no useful ordering between them. The distinction that matters most is between 1 and 4 - both are failing, but 1 means nothing this container produced is recoverable at all, while 4 means an earlier capture survives and only the cadence is broken.

The counts are what make a state readable, so read them together. A 0 on lattice_repocontext_backup_last_full_entries after a successful capture means an empty or wrongly-scoped selection was captured, which succeeds and reports success through every other signal while protecting nothing - that is what state 3 reports. On lattice_repocontext_backup_sink_backups, 0 and -1 are different facts: 0 means the sink was readable and empty, -1 means it was never successfully read. That series counts what is recoverable after the store is destroyed, which lattice_repocontext_backup_captures_total cannot, because that counts only what this process captured.

A series that is absent rather than valued means the host did not construct the meter, or the collector refused the series at one of its ceilings. It never means backup is healthy. This is why every instrument here is observable and constructed unconditionally, including on a deployment with no sink configured, which reports state 0 as a value: an instrument created on a first failure would be missing during exactly the window an alert is meant to cover, and a series first created late can be refused outright by the series cap.

SQLite grain-storage lock attribution

On the local profile every tree shares one SQLite file, and SQLite admits one writer at a time. A lock storm recorded on a deployed container (issue #2431) could only be attributed afterwards by proximity - which grain types appeared in the log lines near each database is locked - and log lines from concurrent activations interleave, so that was never attribution. It also could not say whether the busy window had actually been exhausted, or how many writers were queued when a write failed.

The SQLite grain-storage arm therefore wraps the provider in an attributing decorator. It changes no timeout, rethrows the provider's own exception unchanged, and ignores every failure that is not a lock failure. It also bounds how many writes contend at once, and re-issues a lock failure on the writes whose loss generates more writes (see Write convoy: bounded concurrency and re-issue below). For each failure whose cause chain holds SQLITE_BUSY or SQLITE_LOCKED it writes one GrainStorageLockContention warning (category Orleans.Lattice.Api.Mcp.RepoContext.Host.RepoContextLockAttributingGrainStorage) carrying:

Field Meaning
Operation read, write or clear.
GrainType, GrainId, StateName The grain whose own call failed - attribution, not proximity.
SqliteErrorCode, SqliteExtendedErrorCode The primary and extended result codes, so SQLITE_BUSY_SNAPSHOT (517) and SQLITE_BUSY_RECOVERY (261) are told apart from a plain busy.
ElapsedMs, BusyWindowMs, Wait How long the call ran against the busy window read from the provider's own connection string. Wait is exhausted when it failed at or after the window and early when SQLite refused the lock before it.
WritesAtEntry, WritesAtFailure, PeakWrites, ReadsAtFailure The write convoy - writes and clears in flight, counting the failed call when it is one - when the call started and when it failed, its high-water mark since start, and the reads in flight beside it.
Attempt, Retrying Which attempt of the call failed (1 for the first), and whether the decorator will re-issue it. Every failed attempt writes its own line.

Writes and clears form the convoy because both need the write lock; reads are reported beside it, because in WAL journal mode a reader does not queue for that lock. Seven instruments on the host meter carry the same evidence in aggregate:

Instrument Meaning
lattice_repocontext_grain_storage_lock_failures_total Lock failures by operation and wait, one per failed attempt.
lattice_repocontext_grain_storage_lock_retries_total Re-issued writes and clears by operation and outcome (recovered or gave_up), one per operation rather than per attempt.
lattice_repocontext_grain_storage_lock_convoy_width Writes in flight when each lock failure surfaced; _sum / _count is the mean width at failure.
lattice_repocontext_grain_storage_writes_in_flight Writes and clears in flight now.
lattice_repocontext_grain_storage_writes_in_flight_peak The most writes and clears in flight at once since start.
lattice_repocontext_grain_storage_write_gate_admissions_total Writes and clears reaching the write gate by operation and admission (immediate, queued, timed_out or unbounded), one per attempt.
lattice_repocontext_grain_storage_write_gate_queued Writes and clears waiting for admission now. These hold no connection and no lock, so unlike the in-flight gauge this depth costs the store nothing.

Read wait first. A convoy outlasting the busy window fails exhausted, with a wide WritesAtFailure; a lock SQLite refuses without waiting fails early, whatever the convoy. The two call for opposite remedies - widening the write path against throttling the convoy - which is why they are separate arms rather than one total. The peak gauge exists because a scrape interval is far wider than a convoy, so the instantaneous gauge can sample either side of one.

Zero semantics. Every arm of the failure counter, and the write and clear arms of the retry and gate-admission counters, are published at zero from process start, so an absent series means the host did not construct the meter or the collector refused it, never that no lock failure happened. The convoy-width summary is the exception: a zero sample would bias its mean, so it is absent until the first failure; read its absence against the failure counter. None carries a tenant dimension, because the grain store is shared by every tree on the host. Reminders share the same SQLite file but are not wrapped, so a reminder write can hold the lock without appearing in the convoy.

To provoke the condition rather than wait for it, hold the write lock from a second connection with BEGIN IMMEDIATE: every writer queued behind it waits out the window and fails exhausted. RepoContextLockAttributingGrainStorageTests does exactly that against Orleans' real ADO.NET provider and this host's schema.

Write convoy: bounded concurrency and re-issue

SQLite has one writer. The attribution above measured what ignoring that costs (issue #2419): a checkpoint write failing after 15016 ms against a 15000 ms busy window, with 106 peak concurrent writes and reads in flight 0. Past a small width the extra concurrency buys no throughput at all, because only one of those writers can hold the lock; the surplus is converted directly into exhausted windows. It is not confined to bulk maintenance either - a routine incremental reconcile of seventeen changed files reproduced it.

Two mechanisms address it, and they work in opposite directions: one reduces the contention, the other survives what remains.

The write gate bounds concurrency. A write or clear waits for admission before it reaches the provider, up to 8 at once, and a writer that is not admitted within 5 s proceeds anyway. Queueing at the gate is strictly cheaper than queueing inside SQLite: a writer waiting here holds no connection, takes no lock and burns none of its busy window, so when it is admitted its whole window is still ahead of it. A writer waiting inside SQLite is spending the budget that decides whether its write survives. Read ..._write_gate_queued as the gate doing its job and a rising admission="timed_out" arm as the bound being too tight - or the store too slow - for the offered load, because those writers are contending exactly as they did before the gate existed. Reads are never gated. The gate fails open by design: it sits in front of every grain-storage write the host makes, so its worst case has to be the behaviour that shipped without it.

Re-issue covers the writes whose loss generates more writes. A dropped write of that kind is not one lost write, it is the next burst: the leaf's write commits the durable projection checkpoint advance, and the advance is already applied in memory when it is made, so dropping it unretried leaves the leaf to re-enter replay over the span it had already applied (re-entered replay WITHOUT its persisted checkpoint having advanced ... partition gap 678 entries) and that replay issues more writes into the convoy that dropped the first one. So the decorator re-issues a lock failure on a write or clear to the leaf and its snapshots (leaf, leaf-snapshot, leaf-snapshot-segment), the shard root (shardroot), or the WAL materialiser pin store (wal-materialiser-pins, bucketed as wal-materialiser-pins~b{n}, whose stale published pin holds the WAL trim floor down - issue #3761 item 6). Up to two re-issues, after a backoff of 250 ms doubling per re-issue and jittered between half and all of it, so writers that failed together do not re-queue together. Neither the admission nor the convoy is held across that backoff.

The set is an allow-list, not everything: a state whose loss does not generate more writes is left to fail, because re-issuing it would add load to a contended writer for no reduction in future load.

A re-issue is safe because the grain-state write is one autocommit statement: a lock failure rolls it back whole, and the provider assigns the new ETag to the grain state only after the statement succeeds, so the re-issue presents the same ETag against an unchanged row. Nothing in that argument is specific to a state name. Reads are never re-issued - a WAL-mode reader does not queue for the writer lock - and nor is any failure that is not a lock failure, since a constraint violation is a verdict on the data rather than on contention. Orleans' ADO.NET provider still logs its own Error writing grain state line for each failed attempt, so read a recovered write through lattice_repocontext_grain_storage_lock_retries_total{outcome="recovered"} rather than by the absence of that line; outcome="gave_up" is a write whose lock failure reached the grain even after every allowed re-issue.

Every failed attempt is still attributed and counted as a lock failure, so a recovered write never reads as an uncontended one - the convoy stays visible in the exposition exactly when the retry starts absorbing it.