Performance: multi-silo scaling guide
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at performance-multi-silo.md, and llms.txt lists every page.Important
Scope: one tree, one storage account. Every figure here comes from a cluster serving a single tree (two for the cross-tree saga rows), with its write-ahead log and grain state on one Azure Storage account.
- Trees. Some of a tree's resources have a fixed count that does not grow with the cluster: its WAL partitions (
WalPartitions, default 8), its shards, and its transaction registry (TxRegistryShardCount, default 1). An estate serving many trees spreads more of them across its silos, so its aggregate write and atomic-saga throughput can scale further than one tree's. Reads, whose bound here lies outside the tree, would gain little. Multi-tree scaling has not been measured. - Storage accounts. One account is a test constraint, not a library limit. Lattice can spread a tree's WAL partitions across several storage accounts; see Multi-account fan-out. The
SetManyAsyncceiling below is that one account's ceiling.
This document is an approximate guide to how Orleans.Lattice throughput responds when you add silos. It is the horizontal-scaling companion to Performance: single-silo guide, which remains the authority on what one silo delivers and on the per-operation call shapes; read that first. Everything here answers one question that the single-silo document deliberately does not: when you put N hosts behind the same tree, what do you actually get?
The numbers come from a third benchmark tier ("Layer 3") that stands up a
fixed-size Orleans cluster on Azure Container Apps, drives it from an
out-of-cluster Orleans client, and records one cell per (workload, silo
count). It covers the same nine workloads as the single-silo Layer 2 table
and is regenerated by benchmark/performance-report.ps1 -Layer3, the same
harness that produces the Layer 1 and Layer 2 tables, either against a rig
it provisions and deletes itself or against one deployed beforehand and
named with -ReuseAca (see Re-running it yourself).
Read the curve, not the absolute numbers. The point of this tier is the shape - where throughput stops rewarding extra hosts - and the shape is robust in a way the absolute figures are not. Several deliberate choices below (one shared storage account per rig, a fixed shard count, ACA's memory ceiling, a client with bounded concurrency) hold the topology constant so the silo count is the only variable; each of them also caps the absolute ceiling. A production deployment tuned for throughput rather than for comparability will beat these figures.
The short version
| Workload | Shape from 1 to 8 silos | What the evidence says bounds it |
|---|---|---|
GetAsync (point read) |
~54 k keys/s at one silo, ~58 k at two, then rising to ~131 k at eight (2.43x) | From two silos on, time spent outside the grain call: the silo-side call time holds at ~4 ms while the client's calls take longer |
GetManyAsync (batched read) |
~312 k keys/s at one silo, ~235 k at two, then rising to ~384 k at six and ~382 k at eight (1.22x) | The client's bounded concurrency against a per-call time that grows with the silo count; not isolated further |
SetAsync, with or without a view |
~4 k keys/s at one silo, ~11-12 k at six, ~10-11 k at eight | Not isolated; the silo-side median write holds at 35-50 ms up to six silos, then rises at eight as Azure Tables timeouts appear |
SetManyAsync (batched write) |
~17.6 k keys/s at one silo, 14-26 k from two to eight, no sustained gain | The one shared storage account: server-side timeouts, and at eight silos a cohort that lost whole batches to them |
| 64-key atomic sagas (single-tree and cross-tree) | ~45-53 sagas/s at one silo, ~120-140 at eight (2.6-2.7x) | Per-saga latency under the client's bounded concurrency; the saga registry's admission budget was never reached |
| 2-key atomic sagas (single-tree and cross-tree) | ~250 and ~150 sagas/s at one silo, ~640 and ~330 at eight (2.2-2.6x) | Per-saga durable-write latency on the one grain-state storage account (#3591) |
No workload is flat any more, and none scales linearly. The rest of this document gives the grid, how it was measured, and the evidence behind each row of that table.
The scaling curve
xychart-beta
title "Speedup over 1 silo vs silo count"
x-axis "Silos" ["1", "2", "4", "6", "8"]
y-axis "Speedup (x)" 0 --> 9
line [1, 2, 4, 6, 8]
line [1, 1.07, 1.7, 2.18, 2.43]
line [1, 0.75, 0.99, 1.23, 1.22]
line [1, 1.23, 2.15, 2.64, 2.23]
line [1, 1.41, 2.48, 2.86, 2.94]
line [1, 0.79, 1.21, 1.47, 1.11]
line [1, 1.09, 1.7, 2.27, 2.62]
line [1, 1.14, 1.71, 2.1, 2.55]
line [1, 1.07, 1.45, 1.84, 2.15]
line [1, 1.27, 1.84, 2.65, 2.66]
Series order (xychart-beta renders no legend): ideal linear scaling (y = silos), then GetAsync (point read), then GetManyAsync (4,096 keys/call), then SetAsync (point write), then SetAsync (point write + async materialised view), then SetManyAsync (4,096 keys/call), then SetManyAtomicAsync (64 keys/saga), then SetManyAtomicAsync (2 keys/saga, single-tree), then BeginAtomicWrite cross-tree (2 keys/saga, 2 trees), then BeginAtomicWrite cross-tree (64 keys/saga, 2 trees).
xychart-beta
title "Reads: sustained throughput vs silo count"
x-axis "Silos" ["1", "2", "4", "6", "8"]
y-axis "Thousand keys/s" 0 --> 423
line [53.96, 57.77, 91.67, 117.55, 131.18]
line [311.88, 234.84, 308.17, 384.01, 381.53]
Series order (xychart-beta renders no legend): GetAsync (point read), then GetManyAsync (4,096 keys/call).
xychart-beta
title "Point and batched writes: sustained throughput vs silo count"
x-axis "Silos" ["1", "2", "4", "6", "8"]
y-axis "Thousand keys/s" 0 --> 29
line [4.37, 5.37, 9.38, 11.55, 9.75]
line [3.87, 5.44, 9.6, 11.09, 11.37]
line [17.59, 13.91, 21.22, 25.94, 19.5]
Series order (xychart-beta renders no legend): SetAsync (point write), then SetAsync (point write + async materialised view), then SetManyAsync (4,096 keys/call).
xychart-beta
title "Atomic and cross-tree sagas: sustained throughput vs silo count"
x-axis "Silos" ["1", "2", "4", "6", "8"]
y-axis "keys/s" 0 --> 9818
line [3410, 3701, 5790, 7726, 8925]
line [502, 573, 860, 1052, 1282]
line [304, 326, 440, 559, 655]
line [2902, 3692, 5340, 7690, 7712]
Series order (xychart-beta renders no legend): SetManyAtomicAsync (64 keys/saga), then SetManyAtomicAsync (2 keys/saga, single-tree), then BeginAtomicWrite cross-tree (2 keys/saga, 2 trees), then BeginAtomicWrite cross-tree (64 keys/saga, 2 trees).
The grid
How it was run. For each cohort the harness scales the silo Container App to N replicas, waits for every replica to report running, retires the superseded revision, settles for cluster membership, then runs the producer as a Container Apps job that joins the cluster as an Orleans client over the same Azure Table clustering store. The client opens several connections per silo (four by default) so grain calls spread across every gateway rather than funnelling through one, seeds the keyspace for the read workloads, warms the tree, and then drives the measurement window. Offered load is per-silo rung x N, so the demand presented to the cluster grows with the cluster: a flat line on the chart means extra hosts absorbed nothing, not that the load ran out.
Every cohort starts on empty storage. Before each cohort, with the silos parked at zero replicas, the harness deletes every table in the storage account except the clustering table and points the silos at a freshly named WAL table and grain-state table. No cohort inherits trees, tree-registry rows, or WAL backlog from an earlier one.
Every figure is true throughput, not offered load. A cell whose first
cohort completes at least 90% of what was offered has not found the
cluster's limit, so the harness discards that cohort as a probe, doubles
the per-silo rung and runs it again, up to three times. The rung a workload
settles on carries forward to the larger silo counts, so demand never falls
as the cluster grows. A cell still keeping up with its offered load after
the last escalation is published as a lower bound, marked >=, and so is a
cell whose producer could not generate the load it was asked for: that
measures the client, not the cluster. Each cell is the median of two
cohorts at the settled rung; GetAsync takes three, because its cohorts do
not always land at the same level. Between cells the silo app is parked at
zero replicas.
The client's concurrency is sized per workload. The client runs a fixed
number of flush slots per silo, recorded as flushConcurrency in the
table's metadata. Batched workloads need few, because one call carries 4,096
keys, and run 8 per silo. Point and saga workloads carry one key or one saga
per call, so at that default the client, not the cluster, would be the
limit; GetAsync runs 16 slots per silo and the point-write and atomic
workloads 64, and each atomic saga is its own flush unit. The point
workloads fan each slot out into that many individual calls, so their
in-flight call count is the slot count squared times N: it grows linearly
with the cluster, keeping per-silo demand constant.
Two identical rigs, split by workload family. To halve the wall-clock,
the sweep ran on two identically deployed rigs in the same region, each with
its own storage account. One ran the two reads and SetManyAsync; the other
ran the four atomic workloads and the two point-write workloads. Every curve
comes from a single rig, so no line mixes the two. As a cross-check the
second rig also ran SetManyAsync at every silo count: it reproduced the
same shape, a dip at two silos and no sustained gain after four, at 4-24%
below the first rig cell by cell, which is the spread to allow when
comparing absolute figures across rigs.
The code measured. Every cell ran the tip of the epic that removed the
multi-silo ceilings (#3496),
recorded as gitSha below. The point-write rows and one-silo re-runs of
GetManyAsync and the 64-key SetManyAtomicAsync saga ran on the same
library code with later fixes to the benchmark client's accounting and
grading applied, because those fixes changed what those cells measured;
those fixes grade a cohort producer-bound only when the producer really
fell behind (#3589),
book point-mode failures per key rather than per 4,096-key flush unit
(#3590), and re-run
a read cohort whose keyspace was never seeded instead of letting it cap the
cell. The six-silo GetManyAsync cell was re-run for that last reason: one
of its two original cohorts was unseeded and had been excluded.
| Operation | Silos | Offered | Sustained throughput | Speedup vs 1 silo | Per-silo efficiency | Per-call p50 | Per-call p99 |
|---|---|---|---|---|---|---|---|
GetAsync (point read) |
1 | ~80 k keys/s | ~54 k keys/s | 1x | 100% | ~30 us | ~290 us |
GetAsync (point read) |
2 | ~160 k keys/s | ~57.8 k keys/s | 1.07x | 54% | ~3.79 ms | ~14.46 ms |
GetAsync (point read) |
4 | ~320 k keys/s | ~91.7 k keys/s | 1.7x | 42% | ~4.07 ms | ~15.92 ms |
GetAsync (point read) |
6 | ~480 k keys/s | ~117.6 k keys/s | 2.18x | 36% | ~3.62 ms | ~17.59 ms |
GetAsync (point read) |
8 | ~640 k keys/s | ~131.2 k keys/s | 2.43x | 30% | ~3.94 ms | ~20.93 ms |
GetManyAsync (4,096 keys/call) |
1 | ~640 k keys/s | ~311.9 k keys/s | 1x | 100% | ~51.05 ms | ~168.93 ms |
GetManyAsync (4,096 keys/call) |
2 | ~640 k keys/s | ~234.8 k keys/s | 0.75x | 38% | ~237.9 ms | ~285.38 ms |
GetManyAsync (4,096 keys/call) |
4 | ~1.28 M keys/s | ~308.2 k keys/s | 0.99x | 25% | ~349 ms | ~451.36 ms |
GetManyAsync (4,096 keys/call) |
6 | ~1.92 M keys/s | ~384 k keys/s | 1.23x | 21% | ~446.96 ms | ~539.3 ms |
GetManyAsync (4,096 keys/call) |
8 | ~2.56 M keys/s | ~381.5 k keys/s | 1.22x | 15% | ~621.26 ms | ~713.69 ms |
SetAsync (point write) |
1 | ~6 k keys/s | ~4.4 k keys/s | 1x | 100% | ~46.04 ms | ~152.52 ms |
SetAsync (point write) |
2 | ~12 k keys/s | ~5.4 k keys/s | 1.23x | 61% | ~34.51 ms | ~89.21 ms |
SetAsync (point write) |
4 | ~24 k keys/s | ~9.4 k keys/s | 2.15x | 54% | ~43 ms | ~148.88 ms |
SetAsync (point write) |
6 | ~36 k keys/s | ~11.6 k keys/s | 2.64x | 44% | ~48.7 ms | ~144.36 ms |
SetAsync (point write) |
8 | ~48 k keys/s | ~9.7 k keys/s | 2.23x | 28% | ~67.86 ms | ~152.17 ms |
SetAsync (point write + async materialised view) |
1 | ~6 k keys/s | ~3.9 k keys/s | 1x | 100% | ~50.2 ms | ~141.06 ms |
SetAsync (point write + async materialised view) |
2 | ~12 k keys/s | ~5.4 k keys/s | 1.41x | 70% | ~41.31 ms | ~137.78 ms |
SetAsync (point write + async materialised view) |
4 | ~24 k keys/s | ~9.6 k keys/s | 2.48x | 62% | ~53.32 ms | ~170.14 ms |
SetAsync (point write + async materialised view) |
6 | ~36 k keys/s | ~11.1 k keys/s | 2.86x | 48% | ~47.1 ms | ~607.48 ms |
SetAsync (point write + async materialised view) |
8 | ~48 k keys/s | ~11.4 k keys/s | 2.94x | 37% | ~51.9 ms | ~168.48 ms |
SetManyAsync (4,096 keys/call) |
1 | ~24 k keys/s | ~17.6 k keys/s | 1x | 100% | ~1622.34 ms | ~1888.69 ms |
SetManyAsync (4,096 keys/call) |
2 | ~48 k keys/s | ~13.9 k keys/s | 0.79x | 40% | ~3717.43 ms | ~4142.52 ms |
SetManyAsync (4,096 keys/call) |
4 | ~96 k keys/s | ~21.2 k keys/s | 1.21x | 30% | ~5853.15 ms | ~6372.12 ms |
SetManyAsync (4,096 keys/call) |
6 | ~144 k keys/s | ~25.9 k keys/s | 1.47x | 25% | ~6895.64 ms | ~7388.44 ms |
SetManyAsync (4,096 keys/call) |
8 | ~192 k keys/s | ~19.5 k keys/s | 1.11x | 14% | ~16466.48 ms | ~20061.78 ms |
SetManyAtomicAsync (64 keys/saga) |
1 | ~8 k keys/s | ~3.4 k keys/s | 1x | 100% | ~476.5 ms | ~761.59 ms |
SetManyAtomicAsync (64 keys/saga) |
2 | ~8 k keys/s | ~3.7 k keys/s | 1.09x | 54% | ~1015.98 ms | ~1395.41 ms |
SetManyAtomicAsync (64 keys/saga) |
4 | ~16 k keys/s | ~5.8 k keys/s | 1.7x | 42% | ~1446.25 ms | ~1665.31 ms |
SetManyAtomicAsync (64 keys/saga) |
6 | ~24 k keys/s | ~7.7 k keys/s | 2.27x | 38% | ~1684.66 ms | ~2559.9 ms |
SetManyAtomicAsync (64 keys/saga) |
8 | ~32 k keys/s | ~8.9 k keys/s | 2.62x | 33% | ~2160.95 ms | ~2870.2 ms |
SetManyAtomicAsync (2 keys/saga, single-tree) |
1 | ~1 k keys/s | 502 keys/s | 1x | 100% | ~20.66 ms | ~73.04 ms |
SetManyAtomicAsync (2 keys/saga, single-tree) |
2 | ~2 k keys/s | 573 keys/s | 1.14x | 57% | ~38.56 ms | ~112.05 ms |
SetManyAtomicAsync (2 keys/saga, single-tree) |
4 | ~4 k keys/s | 860 keys/s | 1.71x | 43% | ~58.86 ms | ~142.6 ms |
SetManyAtomicAsync (2 keys/saga, single-tree) |
6 | ~6 k keys/s | ~1.1 k keys/s | 2.1x | 35% | ~80.03 ms | ~273.61 ms |
SetManyAtomicAsync (2 keys/saga, single-tree) |
8 | ~8 k keys/s | ~1.3 k keys/s | 2.55x | 32% | ~72.96 ms | ~152.04 ms |
BeginAtomicWrite cross-tree (2 keys/saga, 2 trees) |
1 | 500 keys/s | 304 keys/s | 1x | 100% | ~9.97 ms | ~68.4 ms |
BeginAtomicWrite cross-tree (2 keys/saga, 2 trees) |
2 | ~1 k keys/s | 326 keys/s | 1.07x | 54% | ~14.07 ms | ~83.96 ms |
BeginAtomicWrite cross-tree (2 keys/saga, 2 trees) |
4 | ~2 k keys/s | 440 keys/s | 1.45x | 36% | ~19.88 ms | ~96.76 ms |
BeginAtomicWrite cross-tree (2 keys/saga, 2 trees) |
6 | ~3 k keys/s | 559 keys/s | 1.84x | 31% | ~22.36 ms | ~100.46 ms |
BeginAtomicWrite cross-tree (2 keys/saga, 2 trees) |
8 | ~4 k keys/s | 655 keys/s | 2.15x | 27% | ~25.52 ms | ~102.4 ms |
BeginAtomicWrite cross-tree (64 keys/saga, 2 trees) |
1 | ~4 k keys/s | ~2.9 k keys/s | 1x | 100% | ~301.34 ms | ~481.09 ms |
BeginAtomicWrite cross-tree (64 keys/saga, 2 trees) |
2 | ~8 k keys/s | ~3.7 k keys/s | 1.27x | 64% | ~171.47 ms | ~316.63 ms |
BeginAtomicWrite cross-tree (64 keys/saga, 2 trees) |
4 | ~16 k keys/s | ~5.3 k keys/s | 1.84x | 46% | ~3677.4 ms | ~4006.31 ms |
BeginAtomicWrite cross-tree (64 keys/saga, 2 trees) |
6 | ~24 k keys/s | ~7.7 k keys/s | 2.65x | 44% | ~166.9 ms | ~499.32 ms |
BeginAtomicWrite cross-tree (64 keys/saga, 2 trees) |
8 | ~32 k keys/s | ~7.7 k keys/s | 2.66x | 33% | ~145.33 ms | ~305.57 ms |
Measured 2026-09-26 on Azure Container Apps (Consumption) in westus3 (.NET 10.0.x) at git sha 0981d6d27, n=2/3 cohorts per cell, silo counts 1,2,4,6,8. Offered load scales with the silo count and is raised until each cell plateaus below it, so every figure is true throughput rather than offered load; every cohort starts on freshly emptied storage. All silo counts share one Azure Storage account for the WAL - see the caveats below for what its metrics showed.
How to read this
Speedup vs 1 silo is the headline and Per-silo efficiency is the
warning light. Speedup is the cell's throughput over the same workload's
N=1 throughput; efficiency is that speedup divided by N. Perfect scaling is
speedup = N and efficiency = 100%. The knee is the silo count after
which efficiency falls away sharply - past it you are paying for hosts that
contend rather than contribute.
The one-silo cell is a different machine from the rest. At one silo every grain is local, so a call from the router to a shard root and from the shard root to a leaf never leaves the process. From two silos on, most of those hops cross the network. The step from one to two silos therefore carries a one-off latency cost that no later step repeats, and for the workloads dominated by per-call latency it shows as a speedup well below 2x at two silos (or, for the two batched workloads, below 1x). Reading speedup from two silos onwards, as well as from one, separates that one-off cost from how the cluster scales.
Offered is the load the cell was driven at, and it is above the
result by design. The harness raises the offered load until the cluster
stops keeping up, so a cell well below its Offered figure is the
cluster's limit, not load that was lost: the client is closed-loop and
bounds how much work it has in flight, so when the cluster cannot keep up
the producer is held back rather than queueing without limit. A >= cell
is the exception - either the cluster was still keeping up at the highest
load the harness offered, or the producer could not generate that load - so
the true ceiling is higher.
Throughput is cluster-wide; the latency quantiles are not. The
throughput column counts every key the whole cluster retired. The p50/p99
columns come from the duration histogram of a single silo's last productive
reporter window, because the histogram is per-process and this tier does not
aggregate quantiles across hosts (merging quantiles is not sound, and
summing them is meaningless). Treat the latency columns as representative
of one host under the cluster's share of the load, not as a cluster-wide
distribution, and as the time spent inside the silo, not the round trip the
client sees. For the atomic rows the instrument is the saga's terminal
broadcast phase (saga.broadcast.duration), not its end-to-end commit time.
Failures are data here, and they are storage timeouts. The engine
retries a WAL saturation refusal on its own back-off ladder, so
back-pressure is absorbed rather than counted, and the FINAL line carries
satRetries / satRecovered / satExhausted to make it readable. Across
the whole sweep satExhausted is zero in every cohort. What failed counts
is Azure Tables operations that did not report success. In every cohort
checked they are almost all server-side timeouts ("Operation could not be
completed within the specified time"), with a handful of
The specified entity already exists conflicts - a transaction retried
after it had in fact landed, so those keys are durable. A point workload books a failure per key; a batched
workload books the whole 4,096-key batch the timeout hit, which is why the
batched failure counts are much larger. Most cells report zero; the non-zero
ones are called out below. The harness grades a cohort a wedge, and leaves
it out of the cell, only when its producer never finished, no measurement
window was productive, or completions stopped for a minute with work still
in flight; otherwise it carries the failure count through rather than
discarding the cell.
Workload by workload
GetAsync: scales from two silos, on client round-trip time
Point reads deliver ~54 k keys/s at one silo, ~58 k at two, and then ~92 k, ~118 k and ~131 k at four, six and eight: 2.43x at eight silos, or 2.27x for the fourfold step from two to eight. All three cohorts of every cell are within about 20% of each other, with zero failures.
The client is closed-loop at 256 calls in flight per silo, and its producer is blocked on the cluster 88-98% of the time in every cohort, so the cluster is setting the pace. Dividing the calls in flight by the throughput gives the mean time a read takes as the client sees it: about 4.7 ms at one silo, 9 ms at two and 16 ms at eight. The silo-side time is far smaller: a p50 of ~30 us at one silo, where the whole read is in-process, and a flat ~3.6-4 ms at two to eight, where the shard root usually lives on another host. So from two silos on, the growth is not inside the grain call; it is in the client-to-gateway hop, the messaging layer and the queueing in front of them. Shard-root point reads are interleavable (#3474), so the fixed pool of 64 shard roots is not a serialisation point for them.
GetManyAsync: rises from two silos, on per-call time
Batched reads deliver ~312 k keys/s at one silo, drop to ~235 k at two, and rise to ~308 k at four and ~384 k and ~382 k at six and eight: 1.22x at eight silos, or 1.62x for the fourfold step from two to eight. Every cohort reports zero failures, and in every one the producer is blocked on the cluster 99% of the time, so the cluster sets the pace.
Each call reads a 4,096-key batch whose keys land on shards on every host, and the client holds eight calls in flight per silo. The silo-side median call time grows with the silo count: ~50 ms at one silo, ~240 ms at two, ~350 ms at four, ~450 ms at six and ~620 ms at eight. From two silos to eight the calls in flight grow fourfold while the call time grows about 2.6-fold, which is the ~1.6x gain measured. Why the call time grows is not isolated further; a batch completes only when its slowest shard answers, and each added host puts more of its shards across the network. The one-silo cell ran 16 producer clients per silo rather than 4, so that the producer could offer the load a single silo needed.
The read keyspace grows with the cell. The harness seeds as many keys as the producer has vehicles, and the vehicle count also sets the offered load, so a larger cell reads a larger keyspace: 128 k keys at one and two silos, 256 k at four, 384 k at six and 512 k at eight. Keyspace size affects the result. At six silos a 96 k-key keyspace delivered ~452 k keys/s in a single cohort, while 192 k and 384 k keys both delivered ~385 k; an earlier six-silo sweep at 384 k keys delivered ~342 k from a single cohort, which puts the run-to-run spread at about 12%. The one-silo cell read twice as many keys per silo as the others, so if anything its figure is understated.
SetAsync and SetAsync with a materialised view
Point writes deliver ~4.4 k keys/s at one silo, ~5.4 k at two, ~9.4 k at
four and ~11.6 k at six, then fall back to ~9.7 k at eight: 2.64x at six
silos. The median time a write spends inside the silo is 35-50 ms from
one silo to six, so the gain comes from more writes in flight across more
hosts rather than from faster writes. At eight the median rises to ~68 ms
and failures appear: the two eight-silo cohorts lost 827 and 1,405 keys, Azure Tables server timeouts with a few already exists conflicts,
against zero at every other silo count. The pace-setter is not isolated
further.
The materialised-view variant tracks plain SetAsync within about 12% from
one silo to six (~3.9 k, ~5.4 k, ~9.6 k and ~11.1 k keys/s) and reaches
~11.4 k at eight, where one cohort lost 148 keys to timeouts; the plain
writes' larger losses at eight account for most of that gap. It is lowest
relative to plain writes at one silo, where the view maintainer shares the
only host with the write path. Maintaining the view does not measurably
lower the write ceiling at two or more silos.
SetManyAsync: bound by the storage account from one silo
Batched writes deliver ~17.6 k keys/s at one silo and do not sustain a gain beyond it: ~13.9 k at two, ~21.2 k at four, ~25.9 k at six and ~19.5 k at eight. The median call already takes ~1.6 s at one silo, and ~7 s at six, so each call is waiting on the WAL rather than on compute, and extra silos add calls in flight without adding commit capacity.
The pace-setter is the one storage account the whole rig writes its WAL
through. It is the only resource in the path that every silo shares, the
second rig reproduced the same shape through its own account, and the
failures it produces grow with the load: zero in most cohorts, one
two-silo cohort with 53,248 keys failed, and one eight-silo cohort with
538,624 keys failed, which retired 13.4 k keys/s beside its sibling's
25.6 k. Those failures are storage timeouts, not WAL back-pressure
refusals. The lever is the harness's
multi-account WAL fan-out (BENCH_WAL_ACCOUNTS), which spreads a tree's
WAL partitions across several accounts; it is held out of this sweep so the
silo count stays the only variable.
Atomic and cross-tree sagas: per-saga latency
The four atomic workloads read most clearly in sagas per second:
| Workload | Keys per saga | Sagas/s at 1 / 2 / 4 / 6 / 8 silos |
|---|---|---|
SetManyAtomicAsync |
64 | 53 / 58 / 90 / 121 / 139 |
SetManyAtomicAsync |
2 | 251 / 287 / 430 / 526 / 641 |
BeginAtomicWrite cross-tree |
64 | 45 / 58 / 83 / 120 / 120 |
BeginAtomicWrite cross-tree |
2 | 152 / 163 / 220 / 280 / 328 |
All four rise with the silo count, by 2.2-2.7x at eight silos. The cost is mostly per saga rather than per key: a 2-key saga runs about two to five times as many sagas per second as a 64-key one, not thirty-two times. Spanning two trees costs a 2-key saga about half its rate; a 64-key saga loses at most 15%.
No cohort was refused. satRetries is zero in every atomic cohort, so
neither WAL back-pressure nor the saga decision registry turned a saga away.
The registry does have a real ceiling: it persists a tombstone per completed
saga for TxDecisionRetention, and LatticeOptions.TxRegistryAdmissionBudgetBytes
refuses new sagas with LatticeSaturatedException (source
TxRegistryCapacity) once a registry shard's row approaches the storage
provider's entity limit. The sweep runs the registry with eight shards
(LatticeOptions.TxRegistryShardCount, default 1), and at these rates it
never got there. A tree that needs more sagas per second than its registry
admits can raise the shard count; a refusal shows up as non-zero
satRetries on the FINAL line.
What sets the rate is how long each saga takes. The client keeps 64 sagas in
flight per silo, and its producer is blocked on the cluster for a growing
share of the time as silos are added (a third of the time or less at one
silo, 78-93% at eight), so the cluster is the limit and each saga's
end-to-end time grows with the silo count. For the 2-key saga the growth is
in its chain of serial durable writes - the leaf prepare and decision
commits and three saga checkpoints - whose p50s rise three- to fourfold
from one silo to eight (a checkpoint from 12 ms to 53 ms, a prepare from
43 ms to 135 ms), while each per-shard broadcast call
(saga.broadcast.shard.duration) stays at about 3 ms. Every one of those
writes lands in the rig's single grain-state storage account. That is the
saga analogue of the SetManyAsync storage bound, and it is tracked as
#3591. The 64-key
sagas carry a 64-key WAL write on top of the same chain; their few failed
keys (up to ~32 k in one cross-tree cohort at eight silos) are Azure Tables
server timeouts.
Caveats that bound these numbers
The N=1 cell is not a like-for-like control for Layer 2. Layer 2
publishes most of its rows at a fixed offered load (for example GetAsync
at ~19.9 k keys/s against 20 k offered), because its question is what one
host costs per call, not where it stops keeping up. Every Layer 3 cell is
driven until the cluster stops keeping up, so most N=1 cells here sit well
above the Layer 2 figure for the same workload, and the two are not
directly comparable. The check that stands in for a control is the one
described under the grid: the second rig ran SetManyAsync independently
and reproduced its shape. If a re-run shows one rig's N=1 cells moving far
from the other's, treat the whole curve as suspect before believing
anything it says about scaling.
The throughput basis differs from Layer 2's, deliberately. Layer 2 publishes the mean of the silo's per-second rate samples. That works when the producer is co-located and work retires smoothly. It does not work here: the Layer 3 client submits large batches, so a whole batch retires inside one sample and the samples between retirements are exactly zero. Layer 2's filter discards those zeros, which averages only the spikes - on a real warmed single-silo cohort that reported 9,637 keys/s against an offered load of 5,935 keys/s, a rate higher than the load, which cannot be a sustained throughput. Worse, the overstatement scales with burstiness, and burstiness varies with silo count, so it would bend the very curve this tier exists to measure. Layer 3 therefore publishes completed operations divided by the active measurement window. A Layer 3 cell and a Layer 2 cell are consequently not computed the same way; the Layer 3 figure is the conservative one.
One storage account per rig backs the whole cluster. Every silo's WAL
writes and grain state land in the same Azure Tables account, so a
flattening write curve can be measuring that account rather than the
cluster. For SetManyAsync and the 2-key sagas the evidence above says it
is. Multi-account WAL fan-out exists in the harness and is held out of this
sweep so that silo count stays the only variable; there is no equivalent
fan-out for grain state yet.
The shard count is fixed across every cell. All cells use the same 64-shard tree topology, so a change in the curve is a change in host count and not in shard granularity. At eight silos that leaves eight shard roots per silo.
The client's concurrency is bounded and scales with the silo count. The client runs a fixed number of flush slots per silo, sized per workload. The rung escalation raises the offered rate, not the slot count, so a cell where the client was pinned at its slot cap measures the cluster's response time under that concurrency, not an open-loop ceiling. Each section above says where that applies.
Harness fixes after this sweep started are listed in the changelog. Three harness bugs were found and fixed while the sweep ran, and the cells they affected were re-run: point-write fan-out booked a whole flush unit as failed on one faulted call (#3590), the report graded some cluster-bound or on-schedule cohorts as producer-bound (#3589), and an unseeded read cohort was accepted in place of a measurement, capping its cell at the starting load. A read preseed that failed part-way also restarted from the first key, and at six silos the 384 k-key seed needed three whole cohort attempts before one seeded (#3588); the seed now resumes from the last slice that landed, a fix made after these cells were measured. One harness defect remains open under the same issue: a new cohort's silo revision can overlap the previous cohort's for up to a minute. The retiring silos run under the previous cohort's cluster id, so they cannot join the new cluster, but they share its storage account until they stop.
Silo sizing matches Layer 2 on CPU, not on memory. The Layer 2 host is a
Standard_D4as_v5 (4 vCPU / 16 GiB). ACA Consumption caps a replica at 4
vCPU / 8 GiB with a fixed 1:2 CPU-to-memory ratio, so CPU parity is exact
and memory is half. This mostly affects cache-resident read workloads,
whose absolute numbers are correspondingly conservative; it does not affect
the scaling shape.
There is a residual per-call latency cost against Layer 2. The Layer 2
producer streams its events to the silo over a local TCP socket and the silo
issues the ILattice calls from inside its own process, so those calls
never leave the host.
The Layer 3 producer is a separate container reaching silo gateways over the
container environment's network, so every call pays a real network hop that
Layer 2 does not. Compare Layer 3 latencies to each other across silo
counts; compare them to Layer 2's only with that hop in mind.
The replay admission queue is widened for this tier. A cohort opens
against a cold tree and the warm-up needs every shard root live at once. The
replay admission gate sizes its queue from the host's CPU grant, which on a
4-vCPU replica refuses a fan-out that wide - it exists to stop a foreground
read queueing behind a background corpus walk, and the benchmark has no
foreground reader to protect. The harness raises
LatticeOptions.WalReplayPermitQueueDepthPerPermit on the Layer 3 silos
accordingly, to 64 against a library default of 4. The single-silo tiers do
not set it and are unaffected.
Two saturation budgets are finite on this tier. The library ships
LatticeOptions.SetManyFanOutBudget and
LatticeOptions.WalAdmissionSaturationCallBudget disabled (infinite). The
Layer 3 harness opts its silos into finite values - 30 s and 15 s by default,
set with -SetManyFanOutBudgetSec and -WalAdmissionCallBudgetSec - so a
sweep measures the bounded configuration. Pass 0 for both to measure the
out-of-the-box configuration instead.
Re-running it yourself
The sweep is a single switch, -Layer3, on the same harness that produces
the other two tiers. Run with no rig named,
pwsh benchmark/performance-report.ps1 -Layer3 provisions a rig, sweeps it
and deletes it; -NamePrefix <prefix> names that rig, and without it the
sweep mints a fresh prefix. To sweep a rig you have deployed yourself
instead, point the sweep at it with -ReuseAca.
Give the rig a prefix of three to nine lowercase letters and digits.
deploy-aca.ps1 needs at least three letters or digits, and the sweep
lowercases a -ReuseAca prefix, strips its hyphens and refuses one longer
than nine characters, so only that form is guaranteed to name the same rig in
both scripts. Run both from the same checkout: the rig's context file is
written locally, under benchmark/.run/aca/, and the sweep cannot find the
rig without it.
pwsh benchmark/azure-throughput/scripts/deploy-aca.ps1 -NamePrefix l3ab12cd3
pwsh benchmark/performance-report.ps1 -Layer3 -ReuseAca l3ab12cd3
deploy-aca.ps1 creates the resource group rg-<prefix> holding the storage
account, container registry, log workspace, container environment, silo app
and producer job. The sweep then runs every workload at 1, 2, 4, 6 and 8
silos and parks the cluster at zero replicas between cells and when it ends,
but it never deletes a rig it did not provision: remove the resource group
with az group delete --name rg-<prefix> --yes when you are done.
The switches that matter for a partial or resumed run:
-ReuseAca <prefix>runs against an already deployed rig instead of provisioning one, and leaves it in place (parked) afterwards.-KeepAcakeeps a rig the sweep provisioned itself: the resource group is preserved with the silos parked at zero, ready for a later-ReuseAcarun. It applies only to the self-provisioning form; a rig passed with-ReuseAcais never deleted.-Workloadsand-SiloCountsrestrict the sweep.pwsh -Filepasses an array argument as a single string, so to pass a list usepwsh -Command "& ./benchmark/performance-report.ps1 -Layer3 -ReuseAca <prefix> -SiloCounts 6,8".-Resumeskips cells that the rig's own run state (benchmark/.run/performance-report/<prefix>/state.json) already holds, so it applies to a re-run against the same rig with-ReuseAca <prefix>; combined with a larger-Nit tops an existing cell up with extra cohorts rather than re-running it.-SaturationRatio(default 0.9) is the completed-to-offered ratio at or above which a cell counts as still keeping up and is re-run at double the load;-MaxRungEscalations(default 3) bounds how many times that happens before the cell is published as a>=lower bound.-Layer3ClientsPerSilo(default 4) sets how many Orleans client connections the producer opens per silo, capped at 64 in total. Raise it when a cell is graded producer-bound, which means the producer could not generate the load it was asked for.-DryRunre-renders this document without touching Azure, from the most recently writtenstate.jsonunderbenchmark/.run/performance-report/, which must hold a Layer 3 sweep (the replay throws otherwise). Use the switch form,-Layer3 -DryRun: the dry-run path checks only the switch, so-Layer 3 -DryRunre-renders the single-silo document instead and leaves this one unchanged.
See Benchmarks for the runbook and prerequisites, and
benchmark/azure-throughput/throughput.md for the engine's own
documentation.