Design
How CLIO Core is built.
How the storage side of CLIO Core is built: the Context Transfer Engine (CTE), the ChiMods stacked on it, the clio-fs filesystem, and the runtime underneath. These pages describe mechanisms and the reasons behind them, with the costs stated alongside the benefits. There is no code here; for APIs, see the Docs.
Pages
Runtime and block devices
Tasks and coroutines, transports, workers, routing, and the block-device module every tier sits on.
Data model and data path
Tags, blobs and targets; hash ownership; the put and get paths; concurrency control.
Placement, tiering and eviction
Scores, the placement engine, reorganization, the organizer, flushing and eviction.
Durability and recovery
What is persisted, the ordering rules, what survives what, and how a restart recovers.
The interposition chain
Replication, cache, indexer, compressor, stream and checkpoint, and why they stack in that order.
clio-fs and FUSE
A distributed POSIX namespace made of blobs: directory blocks, inode homes, leases, rename, and the FUSE daemon.
I/O adapters
POSIX, STDIO and MPI-IO interception; the HDF5 VOL connector and VFD; ADIOS2; the GPU paths.
The stack
Design principles
- No central metadata service. Every blob, tag, directory block and inode has exactly one home container chosen by hashing. There is no metadata server to scale or fail over. The cost is that whole-namespace questions (sizes, listings, searches) are broadcasts.
- One writer per object. Each object's home serializes its mutations. That makes versioning, coherence and crash ordering tractable. Parallelism comes from many objects, not from concurrent writers to one object.
- Everything is a blob. File data, directory blocks, inode records, extended attributes and staged appends are all ordinary CTE blobs. One tiering, replication and durability mechanism covers data and metadata alike.
- Mechanism in the core, policy in layers. The core provides replica slots, transform flags, invalidation, fault handlers and conditional puts. Replication, caching, compression, indexing and filesystem semantics are separate, optional ChiMods composed into a chain.
- Operators rank tiers; data carries a score. Placement matches a blob's score to configured tier scores, and measurements only break ties. Movement between tiers is explicit, by rescoring, flushing or eviction, and always copies before it frees.
- Durable on fsync, like ext4. Acknowledged writes
run at memory speed. Power-loss and node-loss durability are paid for
on request, through
fsyncand replication. - Order operations so a crash leaks rather than corrupts. Logs trail data, new copies precede frees, records precede names, link counts rise before links appear. The failure mode is wasted space, never a reference to the wrong bytes.
- Fail closed. Lost volatile data reads as an I/O error, not zeros. Expired caches that cannot be revalidated fail, rather than serving stale data. Coherence stamps that cannot vouch for a file drop the cache.
- Keep hot paths off the network. Zero-IPC reads from shared memory, local path resolution from pushed caches, write sieves and write-behind, and node-local cache copies all exist to turn common operations into memory accesses.
One write, end to end
An application on node A appends 4 KiB to
/mnt/clio/run/log.txt through the FUSE mount, then calls
fsync.
- Kernel to daemon. FUSE delivers the write to the daemon on node A. On a single node, the kernel's cached size supplies the append offset. On a multi-node mount, the append is staged locally and merged by the file's stream home.
- Sieve. The bytes are copied into the client's write sieve: a 64 KiB shared-memory page for the file's 1 MiB page blob. The write returns immediately.
- Ship. Within about 500 µs the sieve page is sent as a put to the filesystem's chain. The cache layer routes it to node A. The file's page blob is owned by node B, the hash of (file tag, page number), so the authoritative put travels to B. B invalidates other nodes' cache copies and registers A's copy.
- Owner. On B, replication writes the primary (and any remote copies) and passes the put to the core. The core takes the blob's write token, drains readers, places the bytes (the RAM tier, since the blob's score defaults to 1.0), writes them, logs the new layout, and acknowledges.
fsync. The daemon drains the sieve, publishes the file's size to its stream home, and waits for replication. It then syncs the file's tag: B moves the page blob to a persistent tier (place, swap, log, free) and fsyncs that device and its metadata log. The daemon then fsyncs the stream's size log and syncs the parent directory's tag, so the name is durable too. Only then doesfsyncreturn.
Any later reader on node A finds the bytes in its node-local cache copy and reads them through shared memory without a request. A reader on node C fetches them from B and gets a cache copy of its own, registered against B's current version.
Tradeoff summary
| Area | Optimized for | Accepted cost |
|---|---|---|
| Ownership | Scale-out metadata with no single point of failure | Broadcasts for sizes, listings, search; hot single blobs are serialized |
| Concurrency | Lock-free reads, crash-safe mutations | Writers to one blob spin-yield with no fairness |
| Tiering | Predictable placement, safe moves | Moves copy whole blobs and need 2× space; the organizer often does not pay for short jobs |
| Eviction | Authoritative data is never silently dropped | Full tiers of real data return ENOSPC; each eviction pass is a full scan |
| Durability | Memory-speed writes, durability on request | Seconds of power-loss exposure for un-synced data; one copy by default |
| Metadata log | Uncontended per-worker logging, cheap snapshots | No per-record checksums; recovery continues on partial state |
| Replication | Availability while degraded | Not quorum-based: degraded writes keep fewer copies and are not re-replicated |
| Caching | Strong coherence, local reads | Extra local writes and owner fan-out on invalidation |
| Filesystem | Never serves stale namespace state | Mutations can stall ~12 s behind a partitioned holder; rename is multi-step |
| FUSE | Microsecond stat and write-behind data path | Fast paths are single-node; errors surface at close or fsync |
| Adapters | Drop-in acceleration without code changes | Interceptors relax fsync and MPI-IO semantics; HDF5 caching never speeds up writes |
| Membership | No false "dead" verdicts at scale | Slow failure detection; no automatic container recovery |
Component maturity
| Component | In the default build and config | Status |
|---|---|---|
| CTE core, bdev, filesystem, stream | Yes | Production path, extensively hardened at 16–64 nodes and under file-system stress suites |
| Replication, cache | Yes | Production path. Replication factor 1 by default. |
| Indexer | Yes | Stable; lexical search only |
| Compressor | Built only with the compression option; not composed by default | Pinned codecs are mature. Dynamic codec selection is experimental. |
| Checkpoint (lazy copies) | Built; created on demand | New |
| HDF5 VOL connector and VFD | Source build option | Covered by differential compatibility suites |
| POSIX, STDIO and MPI-IO interceptors | Source build, with ELF interception | Usable, with the semantic limits described on the adapters page |
| GPU vector, UVM, LLM hooks | GPU / LLM build options | Research and prototype stage |
These pages describe the
dev branch as of October 2026. When the design changes,
update the matching page. Each page is self-contained and links to its
neighbors.