Design

Placement, tiering and eviction

The core's tiering model rests on two numbers. A target score ranks each device. A blob score says how hot the data is. Placement matches the two. Reorganization changes a blob's score and then moves its bytes. Flushing moves bytes toward durable tiers, and eviction removes only data that is marked as a cache.

The Data Placement Engine

When a blob grows, the core asks the Data Placement Engine (DPE) for an ordered list of targets. Three policies are available, selected once per pool with dpe.dpe_type:

PolicyOrdering
max_bw (default)Preferred targets sorted by configured score, highest first. Ties are broken by measured write bandwidth for writes of 32 KB or more, and by latency for smaller writes. Fallback targets are sorted by score, lowest first.
round_robinRotates through each bucket.
randomShuffles each bucket.

All three policies use the same two-bucket structure. First, any target that cannot hold the whole extension is dropped. The rest are split into:

  • Preferred: targets whose score is at or below the blob's score, meaning the matching tier or a colder one.
  • Fallback: hotter tiers, used only when nothing preferred has room. The least over-provisioned is tried first.

After the DPE, the core applies filters of its own:

  • targets on dead nodes are removed;
  • a request's minimum persistence level excludes volatile tiers;
  • cache copies are capped away from durable tiers;
  • a device-health rule drops targets whose predicted remaining lifetime is under a day, and, for long-term data, under a week.

The core then walks the ordered list and allocates from each target in turn until the request is satisfied. A single extension can therefore span several devices and tiers.

Design decision: configured scores beat bandwidth predictions. An earlier version let measured or predicted bandwidth dominate placement. A mispredicted file tier then outranked RAM, which put the primary copy of a "DRAM cache" on disk and left an HBM tier permanently empty. Today the operator's score decides the tier, and performance figures only break ties within it. The cost is that a wrong score in the configuration is followed faithfully.

What this means in practice

  • New data lands on the hottest tier. A new blob defaults to score 1.0, so it goes to the highest-scored tier until that tier is full, then spills downward through the fallback order.
  • Scores mean something only relative to each other. A blob at 0.5 never lands on a 1.0 tier unless nothing at or below 0.5 has room. A fast tier scored above every blob's score receives data only through explicit promotion. The GPU paging work found exactly this: an HBM tier scored 1.0 never received 0.5-scored pages.
  • Placement is decided by the blob's owner using only the targets that owner has registered: its own node's devices, plus its neighbors' persistent devices when the neighborhood is larger than 1.

Capacity tracking

The data path debits and credits each block's physical footprint atomically. A periodic statistics task (every second by default) refreshes each target's free space, capacity, performance figures and predicted lifetime from its block device. It also recomputes automatic scores and evacuates targets predicted to fail soon. File-backed devices cap the free space they report at what the host filesystem actually has free.

Node-local cache copies are reported as free space in capacity queries, the way Linux reports the page cache, because they can be reclaimed on demand.

Reorganization: change the score, move the bytes

Reorganizing a blob means giving it a new score. If the score changes by more than a hysteresis threshold (0.05 by default), the core moves the bytes to match. A move does the following:

  1. Takes the blob's write token and drains readers.
  2. Reads the whole blob into a buffer.
  3. Places and writes a complete new copy on the target tier.
  4. Swaps the layout, bumps the placement generation (which invalidates zero-IPC readers), and logs the new layout.
  5. Frees the old extents, and only then.
1. Token drain readers 2. Read whole blob 3. Place + write new copy (2× space) 4. Swap + log new generation 5. Free old only after log A failure before step 4 leaves the original blob untouched.
Place-before-free. The cost is a transient doubling of the blob's footprint. In exchange, no failed move can lose data.

Four refinements keep moves cheap and safe:

  • Drop instead of copy. If a primary is being demoted below the score of a persistent replica that already holds the bytes, the primary's blocks are simply released.
  • Durability floor. A move never puts a blob on a less durable tier than the one it currently occupies.
  • No-room back-off. Scores are grouped into ten bands. A move into a band that recently had no room fails fast for one second, instead of repeatedly reading whole blobs only to be refused.
  • Graceful shutdown. Moves in flight are counted, so shutdown waits for them instead of tearing a blob in half.

The data organizer

The organizer is an optional background policy that rescores blobs automatically. It is off by default. When enabled, one or more periodic tasks each take a hash partition of all blobs and rescore them in turn, reorganizing those whose score changes. Five policies are built in:

PolicyRuleFits
frecencyHalf recency (exponential decay with a 10-minute half-life), half frequency (saturating access count).General-purpose caches
cyclicPromote a stable prefix of page numbers that fits the fast tier. It only promotes, never demotes.Iterative sweeps over the same data
hotsetRank by access count.Skewed reuse
scatterPromote once reads begin, or when a "gather phase" hint arrives.Write-then-read pipelines
grayscottDemote checkpoint tags by name prefix.Simulation checkpoint output

Applications can also broadcast an opaque phase hint that policies interpret, such as "the gather phase starts now".

Measured tradeoff. Filling a fast tier costs on the order of 270 µs per page moved. On a short GPU workload (a 45 ms run with a 48 MB fast tier), every promoting policy made the run 13–23% slower. With the slow tier also in RAM, every policy was within noise. The rule of thumb in the code: migration pays only when the time saved by later reads is large compared with the cost of moving the data, so short or one-pass workloads should leave the organizer off.

Each organizer round scans every blob in its partition, so the cost per round grows with the number of blobs. Access counts are not persisted across a restart. Score changes reach disk only through the periodic metadata snapshot, so up to about a minute of rescoring can be lost in a crash.

Flushing to persistent tiers

Data written to a volatile tier becomes durable in three ways:

  • Periodic flush (every 10 seconds by default). Blobs with blocks below the configured persistence level are moved to persistent tiers, using the same place-swap-log-free sequence as reorganization, under the write token. The flush works against a capacity budget computed up front and stops at the first "no room".
  • Sync of a tag (the core of fsync). The tag's blobs are moved up to persistent tiers, every device holding them is synced in parallel, and then the metadata log is synced. If no persistent tier exists, the sync succeeds with tmpfs semantics. A full tier fails with "no space", and device errors fail with an I/O error. fsync_mode: deferred turns this into a no-op that relies on the periodic flushes.
  • Graceful stop. A stop hook runs a flush within the stop grace period, so RAM-tier data reaches disk on a clean shutdown.

History. An earlier flush freed a blob's RAM blocks and then re-put them to disk. Under concurrent writes this truncated or zeroed data, losing hundreds of megabytes in large-scale runs. The flush now moves data atomically under the write token, the same way reorganization does.

Eviction

Eviction is deliberately narrow.

  • Triggered only by a failed placement. A put that cannot find room evicts on its own node, then retries once. There is no background high-watermark eviction.
  • Only droppable data is evicted to admit a write. Authoritative data is never evicted to make room for another write. Two kinds of data qualify: node-local cache copies, reclaimed first with their slot metadata kept, and blobs created as droppable, deleted whole, lowest score first.
  • Blobs being written are skipped, because the holder of the write token may be the very put that is waiting for room.

An explicit eviction request (broadcast to every node, with a tier threshold and a byte budget) uses the same two phases.

Each eviction pass scans and sorts all blobs on the node, because there is no index by score. A blob deleted to free space on one tier also loses whatever parts it had on other tiers, since eviction removes whole blobs.

Why eviction is narrow. A previous version put the droppability check on the incoming put rather than on the victims, which made authoritative data evictable. The current rule makes "out of space" an honest error for authoritative writes. The layers that need cache behavior (GPU page caches, HDF5 dataset images, node-local cache copies) mark their data droppable explicitly.

Configuration that matters

SettingDefaultEffect
storage[]RAM tier (score 1.0, 80% of DRAM) and a 10 GB file tier (score 0.2)Defines the tiers: device type, capacity, score and persistence level.
dpe.dpe_typemax_bwPlacement policy.
targets.neighborhood1 in the shipped configHow many nodes' persistent devices each owner can place on.
performance.score_difference_threshold0.05Reorganization hysteresis.
performance.flush_data_period_ms10000How often volatile data is moved to persistent tiers.
performance.flush_data_min_persistence1The persistence level flushes and syncs aim for.
performance.fsync_modedurabledeferred makes fsync return immediately.
performance.stat_targets_period_ms1000How fresh capacity and health figures are.
organizer, organizer_tasks, organizer_period_msnone, 1, 5000Background rescoring. These keys sit at the top level of the CTE pool entry.

Tradeoffs at a glance

DecisionBuysCosts
Operator scores rank tiersPredictable placement. A misprediction cannot demote the cache tier.A mis-scored tier is followed exactly as configured.
Copy-based moves under the write tokenMoves are atomic and failure-safe.Readers of that blob wait during the move, and the blob briefly needs twice the space.
Organizer off by defaultNo surprise migrations stealing bandwidth from short jobs.Cold data stays on the fast tier unless something demotes it.
Eviction only of droppable data, only on demandAuthoritative data is never silently discarded.A full tier of authoritative data returns "no space" instead of spilling, and each pass is a full scan.
Periodic flush plus explicit syncWrites run at RAM speed, with durability on request.Between flushes, data on volatile tiers exists only in memory.