Architecture
This page documents Orleans.Lattice.Storage.AzureTable 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at architecture.md, and llms.txt lists every page.Orleans.Lattice.Storage.AzureTable is a durable implementation of the core IWalStorageProvider seam. The core WAL grain decides what to append, when to trim, and how consumers read; the Azure Table provider decides how those append, read, trim, and reconcile calls are represented in Azure Table Storage.
For the core WAL contract and placement model, see WAL Storage Providers. For replication consumption of the WAL, see Replication WAL.
High-level pipeline
flowchart LR
Writer[Core WAL writer] -->|IWalStorageProvider append| Provider[AzureTableWalStorageProvider]
Provider -->|entry transaction| Batch[(Per-append batch partition)]
Provider -->|ordered completion| Manifest[(Per-shard commit metadata)]
Manifest --> Tail[(Committed shard tail)]
Reader[Core WAL reader] -->|IWalStorageProvider read| Provider
Provider -->|committed batch order| Batch
Recovery[Activation reconciliation] --> Provider
Recovery -->|complete or remove interrupted batches| Manifest
The provider keeps each (tree, shard) stream ordered while allowing append payloads to land in per-batch storage partitions. Each commit writes the metadata of the batches pending at that moment in ascending start-offset order, and the stored shard tail is the high-water mark of the batches committed so far, so it never moves backward.
Storage layout
The table is split by behaviour, not by public .NET type:
| Stored data | Behaviour |
|---|---|
| Entry rows | Store one retained WAL entry payload and its offset. Rows for one append batch are written together in one Azure Table transaction. |
| Batch partition | Groups the entry rows for one append batch. Separate batches use separate partitions so concurrent shard work is not forced through one storage partition for the entry payloads. |
| Commit metadata | Records that a batch is committed and the inclusive offset range it covers. Reads enumerate committed metadata in offset order before fetching entry rows. |
| Tail metadata | Stores the highest committed offset for a shard. Highest-offset recovery is a point read rather than a scan over all entries. |
| Recovery markers or discoverable batch state | Lets reconciliation find interrupted appends after a crash and either complete the contiguous prefix or remove non-contiguous leftovers. |
Tree ids are encoded for Azure Table key safety. The table is created on first use. Use a non-default TableName when multiple deployments share one storage account.
Transactional batch contract
Every append preserves the public WAL invariants:
- The supplied offsets must be dense within each batch; across batches the producer allocates them densely, in allocation order.
- Entry payload rows for one batch are committed atomically by Azure Table Storage.
- Each commit writes the pending batches' metadata in ascending start-offset order.
- The stored tail never moves backward. Batches whose entry writes finish out of order can commit out of order, so the tail can pass a lower batch that is still in flight; if that batch then fails, its offsets stay an honest gap beneath the tail, as the core WAL contract permits.
- A failed append does not expose a partial visible batch.
- Trim can remove old retained rows, but it never moves the committed tail backward.
Azure Table transactions are limited to 100 entities and a 4 MiB total payload. The provider exposes AzureTableWalStorageProvider.MaxEntriesPerBatch = 100, and the core WAL batch defaults use the same figures: LatticeOptions.WalMaxBatchEntries defaults to 100 entries and WalMaxBatchBytes to 4 MiB of encoded payload, measured before compression. For throughput and pending-depth tuning, see WAL tuning.
A single entry has a much smaller limit. Each entry is stored as one entity whose whole encoded record - after compression, when it applies - is a single binary property, and the Azure Table service data model limits a binary property to 64 KiB (and an entity to 1 MiB). The provider neither splits a payload across properties or rows nor checks its size, so an entry whose stored payload exceeds 64 KiB is rejected by the service. Because a batch commits in one transaction, the whole append fails with it, including every other entry in the same batch, and the rejection is not a transient fault, so it is not retried in place. Compression can bring a compressible payload under the limit; an incompressible payload, or any payload with Compression set to LatticeCompression.None, is stored verbatim.
Commit pipeline
A normal append has three behavioural stages. The first two run concurrently, and the third starts only once both have landed:
- Prepare recovery state. The provider records enough information to distinguish an interrupted append from a committed append during the next reconciliation pass. With
EliminateCandidateRowOnHotPath = true, the normal path skips an extra recovery-marker write and relies on discoverable batch state plus the committed tail. - Write entries. Entry payload rows are written in a single Azure Table transaction for that batch.
- Complete in offset order. Each commit writes the metadata of the completions pending at that moment in ascending offset order and raises the shard tail to the highest end offset committed so far; the tail never moves backward. Under load, multiple completions can be coalesced into one transaction, bounded by Azure Table transaction limits.
PipelinePhaseTwoCommits = true lets a caller return after durable entry write and observation of the previous pending completion for the shard. It does not change ordering, recovery, or all-or-nothing durability; it changes which append observes a completion fault, and it introduces a bounded read visibility lag. PipelinedPhaseTwoFaultHandler exists so an idle shard can still report a completion fault for observability.
Read visibility lag under pipelining
The read path is derived from the commit metadata written in stage 3, so a batch becomes readable when its completion lands rather than when its append returns. Under PipelinePhaseTwoCommits = true the trailing batch on a shard is therefore durable but not yet readable for a short interval, and a read can omit it.
The reported highest offset does not share that lag. GetHighestOffsetAsync folds the contiguous run of already-durable batches the shard's live completion worker has accepted over the stored tail, so an offset returned by a completed append is never reported back as though it did not exist. The fold walks upward from the stored tail and stops at the first gap, which is exactly the run reconciliation would roll forward, so it can never claim a batch whose lower neighbour is still in flight. It degrades to the stored tail alone when this provider instance has no live worker for the shard - a fresh activation or another silo. With PipelinePhaseTwoCommits = false the worker is still live, but every append waits for its own completion, so the fold can add only batches whose appends have not yet returned. The stored tail keeps its former meaning everywhere else, including on the reconciliation path.
Nothing is lost. The entries are durable before the append returns, and reconciliation rolls the batch forward if the process restarts first. WAL consumers poll, and the core WAL grain tracks its next offset in memory rather than re-reading the tail, so the read lag is invisible on the canonical replication path.
A caller that genuinely needs read-after-write - a controlled hand-off, an operator consistency probe, or a test - awaits the provider's phase-two flush barrier, which drains the completions outstanding at the moment of the call and then rethrows any that failed. Prefer that barrier over sleeping or polling for an expected count.
PhaseTwoCoalescingWindow controls how long completion waits for more pending work before sending the coalesced transaction. PhaseTwoCommitTimeout bounds a wedged completion transaction so later work is not blocked indefinitely. The deadline abandons only the worker's wait, not the transaction: an abandoned submit stays fenced on its partition until it actually completes - or until the provider cancels it, 60 seconds after abandoning it, so the fence cannot be held forever - and the post-failure reconcile and tail read wait for that fence, so a late-landing transaction can never be overwritten by a resync that read the partition before it landed.
Overlap rejection
Every batch is stored in its own partition, keyed by its first offset, so storage on its own would accept a batch that starts inside one already written, and a read would then return the shared offsets twice. The provider rejects that batch with InvalidOperationException before writing anything, as the IWalStorageProvider contract requires.
The check costs nothing on the steady-state path. The provider keeps, per shard, an upper bound on the offsets that may be written, and an append starting above it is admitted with no storage call; the core WAL grain always appends above it. Any other append, such as an out-of-order arrival or a retry, is checked against the batches in motion on this instance and then with one query over the few partitions that could hold an overlapping entry. The first append on a shard reads the bound from the stored tail and any uncommitted batches above it. Reconciliation re-establishes it.
A re-append that starts at the same offset as a written batch is not rejected here: it collides on the batch's own rows and is resolved by the idempotent-replay check, which accepts an identical retry and fails anything else. The bound is only trusted while this instance is the shard's single writer, which the WAL grain's single activation provides, so a batch another process writes after this instance read the bound is not detected until the next reconciliation. The storage check looks for stored entry rows, so an offset whose rows a trim has already deleted is not detected either: an append that reuses only trimmed offsets is accepted.
Recovery and downgrade safety
On activation, and again after a failed flush, the core WAL grain calls the provider reconciliation hook. The provider compares the committed tail with interrupted append evidence and applies these rules:
- If an interrupted batch contiguously extends the committed tail, reconciliation completes it and advances the tail.
- If an interrupted batch is below the tail or above a gap, reconciliation removes it so the next append can use the correct next offset.
- A batch whose manifest row already exists is committed, even when the stored tail has not caught up with it. Reconciliation keeps it and re-anchors the tail on it; it never rolls a committed batch back.
- Reconciliation never lowers the stored tail. The tail write is conditional on the tail it read, so a concurrent commit makes the pass retry rather than overwrite.
- Reconciliation is idempotent: a clean shard has no work to do.
The post-failure call is not quiescent: pipelined completions accepted before the failure may still be committing. Reconciliation is serialised per shard and first waits, on this provider instance, for that shard's in-motion appends, its queued completions, and any completion transaction abandoned on its commit deadline to settle, so it never mistakes a live batch for an orphan. A writer on another process is covered only by the conditional writes above.
When EliminateCandidateRowOnHotPath is enabled, reconciliation recognizes both the legacy recovery-marker shape and the newer discoverable-batch shape. Upgrading from the legacy setting to the default setting is safe. Before downgrading back to the legacy setting, drain pending appends and allow reconciliation to complete on a deployment that still has the default setting enabled.
Read, trim, and capacity behaviour
Reads enumerate committed batch metadata in offset order, then stream entry rows lazily from each overlapping batch. GetHighestOffsetAsync reads the stored tail and folds over it the contiguous run of already-durable batches the shard's live completion worker has accepted, so it never lags behind a completed append. GetLowestOffsetAsync finds the first retained batch after trim.
The filtered replay read (ReadFilteredAsync, issue #3565) walks the same metadata, and bounds each batch query above by the window's last row key as well. It classifies each row from the routing prefix of its payload before decoding it: a compressed row is inflated into a pooled buffer rather than a new array, and a row the reader's filter excludes is neither decoded nor retained - except the window's last examined row, which, when excluded, is projected routing-only at the end of the scan so a resuming reader moves past it. The table service still returns every row in the window - the key lives inside the payload, so the service cannot select on it - which makes the saving the per-row decode and its allocations, not the transfer. The classification needs the routing reader AddAzureTableWalStorage supplies; a provider built through a public constructor decodes every row it examines, and returns the same rows.
Trim deletes old retained entry rows in bounded Azure Table transactions and removes matching commit metadata in order. A crash during trim can leave a stale retained prefix, but not a gap in the live tail; a later trim can resume cleanup.
Capacity planning is shared with core WAL tuning:
- Increase shard count to spread work across storage partitions.
- Keep
WalMaxBatchEntriesat or below the provider batch limit. - Keep each entry's encoded record under 64 KiB after compression; an entry is never split across properties or rows.
- Use
WalMaxPendingBatchescarefully; more pending batches increase pipeline depth and storage pressure. - Watch retry-attempt and retry-exhausted telemetry to distinguish transient retry storms from saturated storage.
- Use the core WAL saturation signal to coordinate admission and retry behaviour.
Compression and retry policies
Stored payload compression is per row. A row records enough metadata to decode itself, so changing Compression, CompressionMinPayloadBytes, or the registered compressor affects new rows only. Older rows remain readable as long as a compressor for their recorded algorithm is still registered; a row whose algorithm has none fails its read with NotSupportedException. The Zstandard fallback that AddAzureTableWalStorage registers keeps Zstd rows readable.
When the provider constructs its own Azure SDK client, it attaches RetryAttemptTrackingPolicy for retry telemetry. When HonorSaturationSignal is enabled and IWalSaturationSignal is available, it also attaches SaturationAwareRetryPolicy so retry attempts can short-circuit during saturated WAL pressure. A pre-built TableServiceClient bypasses provider-owned pipeline construction; the host owns any equivalent policies in that mode.
Relationship to replication
Replication writes and reads the WAL through the same core provider seam as single-cluster WAL users. This package does not define replication semantics; it supplies durable storage for the retained log that replication shippers, bootstrap, and fall-off-log handling depend on. See Replication package and Replication WAL.