Design

How CLIO Core is built.

How the storage side of CLIO Core is built: the Context Transfer Engine (CTE), the ChiMods stacked on it, the clio-fs filesystem, and the runtime underneath. These pages describe mechanisms and the reasons behind them, with the costs stated alongside the benefits. There is no code here; for APIs, see the Docs.

Pages

Foundation

Runtime and block devices

Tasks and coroutines, transports, workers, routing, and the block-device module every tier sits on.

CTE core

Data model and data path

Tags, blobs and targets; hash ownership; the put and get paths; concurrency control.

CTE core

Placement, tiering and eviction

Scores, the placement engine, reorganization, the organizer, flushing and eviction.

Reliability

Durability and recovery

What is persisted, the ordering rules, what survives what, and how a restart recovers.

ChiMods

The interposition chain

Replication, cache, indexer, compressor, stream and checkpoint, and why they stack in that order.

Filesystem

clio-fs and FUSE

A distributed POSIX namespace made of blobs: directory blocks, inode homes, leases, rename, and the FUSE daemon.

Adapters

I/O adapters

POSIX, STDIO and MPI-IO interception; the HDF5 VOL connector and VFD; ADIOS2; the GPU paths.

The stack

FRONT ENDS FUSE mount InterceptorsPOSIX, STDIO, MPI-IO HDF5 / ADIOS2VOL, VFD, engine C++ / Python / GPUcore client API NAMESPACE AND STREAMS Filesystem ChiMod (clio-fs)directory blocks, inodes, handles Streamsizes, ordered appends INTERPOSITION CHAIN Cache Indexer Compressor Replication STORAGE ENGINE CTE corehash-owned blob metadata · placement engine · write-ahead log · reorganize / flush / evict DEVICES (bdev pools) RAM / HBM / pinned NVMe / SSD file HDD / PFS file S3 / GCS All of it runs as tasks inside the CLIO runtime: one process per node, coroutines on a few worker threads.
Front ends reach the CTE through the namespace layer (FUSE, interceptors) or directly (HDF5, ADIOS2, the API). Every layer below the front ends is a ChiMod pool, and every arrow between them is a task.

Design principles

  1. No central metadata service. Every blob, tag, directory block and inode has exactly one home container chosen by hashing. There is no metadata server to scale or fail over. The cost is that whole-namespace questions (sizes, listings, searches) are broadcasts.
  2. One writer per object. Each object's home serializes its mutations. That makes versioning, coherence and crash ordering tractable. Parallelism comes from many objects, not from concurrent writers to one object.
  3. Everything is a blob. File data, directory blocks, inode records, extended attributes and staged appends are all ordinary CTE blobs. One tiering, replication and durability mechanism covers data and metadata alike.
  4. Mechanism in the core, policy in layers. The core provides replica slots, transform flags, invalidation, fault handlers and conditional puts. Replication, caching, compression, indexing and filesystem semantics are separate, optional ChiMods composed into a chain.
  5. Operators rank tiers; data carries a score. Placement matches a blob's score to configured tier scores, and measurements only break ties. Movement between tiers is explicit, by rescoring, flushing or eviction, and always copies before it frees.
  6. Durable on fsync, like ext4. Acknowledged writes run at memory speed. Power-loss and node-loss durability are paid for on request, through fsync and replication.
  7. Order operations so a crash leaks rather than corrupts. Logs trail data, new copies precede frees, records precede names, link counts rise before links appear. The failure mode is wasted space, never a reference to the wrong bytes.
  8. Fail closed. Lost volatile data reads as an I/O error, not zeros. Expired caches that cannot be revalidated fail, rather than serving stale data. Coherence stamps that cannot vouch for a file drop the cache.
  9. Keep hot paths off the network. Zero-IPC reads from shared memory, local path resolution from pushed caches, write sieves and write-behind, and node-local cache copies all exist to turn common operations into memory accesses.

One write, end to end

An application on node A appends 4 KiB to /mnt/clio/run/log.txt through the FUSE mount, then calls fsync.

  1. Kernel to daemon. FUSE delivers the write to the daemon on node A. On a single node, the kernel's cached size supplies the append offset. On a multi-node mount, the append is staged locally and merged by the file's stream home.
  2. Sieve. The bytes are copied into the client's write sieve: a 64 KiB shared-memory page for the file's 1 MiB page blob. The write returns immediately.
  3. Ship. Within about 500 µs the sieve page is sent as a put to the filesystem's chain. The cache layer routes it to node A. The file's page blob is owned by node B, the hash of (file tag, page number), so the authoritative put travels to B. B invalidates other nodes' cache copies and registers A's copy.
  4. Owner. On B, replication writes the primary (and any remote copies) and passes the put to the core. The core takes the blob's write token, drains readers, places the bytes (the RAM tier, since the blob's score defaults to 1.0), writes them, logs the new layout, and acknowledges.
  5. fsync. The daemon drains the sieve, publishes the file's size to its stream home, and waits for replication. It then syncs the file's tag: B moves the page blob to a persistent tier (place, swap, log, free) and fsyncs that device and its metadata log. The daemon then fsyncs the stream's size log and syncs the parent directory's tag, so the name is durable too. Only then does fsync return.

Any later reader on node A finds the bytes in its node-local cache copy and reads them through shared memory without a request. A reader on node C fetches them from B and gets a cache copy of its own, registered against B's current version.

Tradeoff summary

AreaOptimized forAccepted cost
OwnershipScale-out metadata with no single point of failureBroadcasts for sizes, listings, search; hot single blobs are serialized
ConcurrencyLock-free reads, crash-safe mutationsWriters to one blob spin-yield with no fairness
TieringPredictable placement, safe movesMoves copy whole blobs and need 2× space; the organizer often does not pay for short jobs
EvictionAuthoritative data is never silently droppedFull tiers of real data return ENOSPC; each eviction pass is a full scan
DurabilityMemory-speed writes, durability on requestSeconds of power-loss exposure for un-synced data; one copy by default
Metadata logUncontended per-worker logging, cheap snapshotsNo per-record checksums; recovery continues on partial state
ReplicationAvailability while degradedNot quorum-based: degraded writes keep fewer copies and are not re-replicated
CachingStrong coherence, local readsExtra local writes and owner fan-out on invalidation
FilesystemNever serves stale namespace stateMutations can stall ~12 s behind a partitioned holder; rename is multi-step
FUSEMicrosecond stat and write-behind data pathFast paths are single-node; errors surface at close or fsync
AdaptersDrop-in acceleration without code changesInterceptors relax fsync and MPI-IO semantics; HDF5 caching never speeds up writes
MembershipNo false "dead" verdicts at scaleSlow failure detection; no automatic container recovery

Component maturity

ComponentIn the default build and configStatus
CTE core, bdev, filesystem, streamYesProduction path, extensively hardened at 16–64 nodes and under file-system stress suites
Replication, cacheYesProduction path. Replication factor 1 by default.
IndexerYesStable; lexical search only
CompressorBuilt only with the compression option; not composed by defaultPinned codecs are mature. Dynamic codec selection is experimental.
Checkpoint (lazy copies)Built; created on demandNew
HDF5 VOL connector and VFDSource build optionCovered by differential compatibility suites
POSIX, STDIO and MPI-IO interceptorsSource build, with ELF interceptionUsable, with the semantic limits described on the adapters page
GPU vector, UVM, LLM hooksGPU / LLM build optionsResearch and prototype stage

These pages describe the dev branch as of October 2026. When the design changes, update the matching page. Each page is self-contained and links to its neighbors.