Design
The filesystem: clio-fs and FUSE
clio-fs is a POSIX-shaped, distributed namespace built entirely out of CTE tags and blobs. It has no metadata server and no separate namespace log: directories, inodes and file data are all ordinary blobs, so they inherit the CTE's tiering, replication and durability. The FUSE daemon clio_cte_fuse mounts it as a normal directory on Linux, macOS and Windows.
Layers
Namespace model
Every object is a tag
Every file, directory and symbolic link is a CTE tag. The tag's
identifier doubles as the inode number, so stat and
readdir always agree. Identifiers are minted per container
and also encode the container that serves the inode, its
home. Ids are handed out only below a durably stored
high-water mark, so a restart never reissues a live inode number.
Files: data pages plus an inode record
- Data lives in the file's tag as 1 MiB page blobs named by page index. Sparse files need no special handling: a missing page is a hole and reads as zeros. Any other read failure is an I/O error, never zeros.
- Attributes live in a small inode-record blob in the same tag: type, size snapshot, link count, mode, owner, the three times, flags, and the symbolic link target if there is one.
- Extended attributes live in one global tag, one blob per file, kept out of the file's own tag so they do not distort its size.
- The logical size lives in the stream for that file, not in the tag. Writes raise it, truncates set it, appends reserve it. The inode record's size snapshot is used only to reconcile after a restart or failover.
Writes preallocate space for growth, doubling up to 64 KiB. A flat 64 KiB per page once made the RAM tier hold about three times the data written.
Directories: extendible-hash blocks
A directory is a tag whose entries live in directory block blobs. Block 0 also carries the directory's header: parent, own name, mode and owner, times.
- Names are placed by a hash. When a block exceeds a threshold (1024 entries by default), half its names move to a new block determined by the next hash bit. Blocks never merge.
- A lookup walks down from block 0 using the name's hash bits. No separate block directory has to be kept consistent.
- Each entry has a state: live, pending (reserved by an operation in progress, and invisible) or leaving (being removed or moved, still visible, frozen).
- The durable image of a block stores pending entries as absent and leaving entries as live. After a crash, any half-done multi-step operation is therefore simply aborted.
- Block images are padded to powers of two so that rewrites do not fragment into many small extents. This alone took file creation from 63 to 893 per second.
Because blocks are keyed by the directory's identifier and not its path, renaming a directory moves one entry. Its subtree does not move, so directory rename is constant-time at any depth.
Path resolution is local
Every path-based request is sent to the local filesystem container. That container resolves the path component by component from its own cache of directory blocks. Blocks this node does not own are fetched from their home once, and this node registers as a holder that the home will push changes to. A warm path walk therefore costs no I/O and no network round trip.
Every name is also published, asynchronously and in batches, as a CTE tag name. CTE search, query and the indexer therefore see real paths. This mirror is eventually consistent. The directory blocks remain the authority.
Where things live in a cluster
| State | Home |
|---|---|
| Directory block | The CTE owner of that block's blob, a hash of (directory, block number) |
| Inode (attributes, links, open handles, stream) | Chosen by hashing the file's path at creation and encoded in its identifier. A later rename does not move it. |
| Open file handle | The inode's home, encoded in the handle, so reads, writes and close go straight there |
| File data pages | The CTE owner of each page blob, hashed per page |
Homing inodes by the file's own path, rather than its directory, was a deliberate change. Previously every file in a directory landed on one container: 264 of 264 files on one node in a stress test, which created a hotspot. Data pages hash independently of inodes, so one large file's bandwidth spreads across the cluster.
The namespace went through three designs before this one: a single metadata home on container 0, then hash sharding with per-node namespace logs, and finally the current directory blocks stored as CTE blobs.
When a home's node is dead, its first live successor serves in its place. The successor loads the blocks and inode records from the CTE, which replication keeps reachable.
Operations
Create and open
Creating a file does three things, in order:
- Mints an identifier homed by the file's path.
- Stores the inode record, before the name exists anywhere.
- Publishes the name to the parent's directory block, as an exclusive insert.
If the name was taken in the meantime, the record is discarded. Of
N racing creators anywhere in the cluster, exactly one wins,
which is what O_EXCL requires. A second step on the inode's
home allocates the handle and applies O_TRUNC. If a racing
rename removes the inode between the two steps, the open re-resolves and
retries.
Directories
- mkdir reserves the name as pending, creates the new directory's first block on its own home, then commits or rolls back the entry.
- rmdir marks the entry as leaving and seals every block of the target directory. Sealing succeeds only if all of them are empty; pending entries count as non-empty. The blocks are then dropped and the entry removed. A racing create into a sealed block fails with ENOENT.
- readdir loads every block and takes one consistent snapshot, retrying if a block changed during collection. Pending entries, and leaving entries for names that are live elsewhere, are hidden. The listing is returned whole, not paged.
Rename
Rename is a multi-step protocol, not one atomic transaction. The design makes every intermediate state either invisible or harmless.
- Files.
- Freeze the source entry.
- Raise the inode's link count.
- Insert at the destination, replacing a non-directory.
- Remove the source.
- Drop the extra link, unlink the replaced file, and publish one tag-name change.
- Directories. Freeze the source, then walk parent pointers from the destination up to the root. Meeting the directory being moved means EINVAL, because it would move into its own subtree. Meeting an ancestor that is itself moving means EBUSY. Then seal an empty destination directory if one is being replaced, move the single entry, and update the moved directory's parent pointer.
- Crossing renames back off. Two-party operations that collide (for example renames in opposite directions) get EBUSY, roll back completely, and retry after a randomized back-off of 20 µs to 2 ms. This prevents both deadlock and directory cycles.
Atomicity caveat. During a file rename, the inode is
briefly reachable under both names. readdir hides the
duplicate. After a crash in that window, the file can come back under
both names as two consistent hard links. The link count was raised
first, so the result is a leak, never a dangling name.
RENAME_EXCHANGE is not supported. RENAME_NOREPLACE
is implemented as probe-then-rename, which leaves a small race
window.
Unlink and orphans
Unlink removes the name and lowers the link count on the inode's home. At zero links with no open handles, the inode is queued for purge. With handles still open, it becomes an orphan and is destroyed at the last close. The orphan list is persisted, so a restart destroys any orphans left behind (handles do not survive a restart).
Purges are batched: every 20 ms each container receives one message
listing up to 512 dead inodes, and deletes their pages locally. Before
this change, every unlink was a cluster-wide broadcast, and
tar extraction on four nodes took 45 s instead of 3 s.
Truncate
A truncate does the following, in order:
- Sets the stream size.
- Truncates the boundary page and physically zeroes the part that survives beyond the new size, so a later extension cannot expose old bytes.
- Deletes the whole pages beyond the new size. It probes past the recorded end, because clients push sizes lazily and stored pages can outrun the recorded size.
Permissions
The ChiMod stores mode, owner and group but does not enforce them.
Enforcement is delegated to the kernel: on Linux the daemon mounts with
default_permissions unless that is turned off. Access times
follow relatime by default.
Not provided
Advisory locks (flock, fcntl), hole punching
and range collapse or insert, RENAME_EXCHANGE, and paged
readdir.
Consistency model
Server side: one writer per object, push before acknowledge
Each directory block and each inode has exactly one home that applies every change to it. A change works like this:
- The home bumps the object's version.
- It writes the blob if the durable image changed.
- It pushes the change to every holder that caches the object.
- Only then does it acknowledge.
Commits for one object are grouped: one is in flight at a time, and changes that arrive in the meantime ride the next one. Holders apply a pushed change only if it matches the version they have, and on a gap they drop their copy and refetch it.
As a result, when a namespace operation returns, every node that caches the affected state already sees the change. For single-object operations that is close to linearizable.
Leases bound staleness
A cached copy is trusted for a 10-second lease. After that it is revalidated against the home. If the home cannot be reached, the operation fails rather than serving possibly stale data. In the other direction, a home never stops pushing to a holder whose lease might still be running. It waits out an unreachable holder, plus 2 seconds of slack for clock skew.
Cost of the lease design. While a node holding cached state is partitioned or dead but not yet declared dead, mutations of that state can block for up to about 12 seconds. The design chooses "slow but correct" over "fast but stale".
Restarts cannot resurrect stale caches
When a home reloads an object from its blob after a restart, it jumps the object's version far ahead (by 2³²) and persists that before serving. A cached copy that was ahead of the blob therefore cannot pass as current. The home's first commit after a restart pushes a full snapshot to every container, because it no longer knows who held the object.
Client side: where POSIX is relaxed
| Mechanism | Behavior | Scope |
|---|---|---|
| Shared-memory attribute mirror | The runtime publishes per-path records that clients read lock-free. A generation counter invalidates a subtree after a directory rename. Records for complete directories make "not found" authoritative without asking the server. | Single-node by default |
| Write-behind | write() can return before the runtime has the bytes. Failures latch per file and surface at fsync or close, like kernel writeback errors. Reads see in-flight bytes. | On by default |
| Close-to-open | Closing a file drains its pages, merges deferred appends and publishes its size, so the next opener anywhere sees the result. | Multi-node, or when the handle wrote |
Overlapping writes from different clients that have not been synced are not ordered last-writer-wins. Applications that write the same range from several nodes must synchronize themselves, as they would on most parallel filesystems.
Durability and crash consistency
The namespace persists exactly as far as the CTE persists its blobs, and
file sizes persist through the stream log. Every namespace operation
stores the inode records it changed before it acknowledges, and reports
ENOSPC or EIO if that store fails. Metadata is placed on a non-volatile
tier when one exists. A device-level flush, however, happens only on
fsync.
The crash-consistency techniques:
- The inode record is written before its name is published, so names never point at missing inodes.
- Durable directory images abort in-flight multi-step operations.
- A block split writes the new child block before the parent records the deeper split. A crash leaves an unreachable block, never a lost entry.
- Link counts are raised before names are added, so a crash can leak a count but never leave a dangling name.
- A directory whose entry exists but whose blocks are missing (a crash
between the two halves of rmdir) is repaired on the next
stat. - A corrupted split tree is detected and reported as an I/O error.
fsync through FUSE drains pages, publishes the size, waits
for replication, syncs the file's tag to a persistent device, fsyncs the
size log and syncs the parent directory, in that order. If a node that may
have held unsynced writes for the file died since the last sync,
fsync returns EIO.
The FUSE daemon
Platforms
| Platform | Backend | Notes |
|---|---|---|
| Linux | libfuse3 | Multi-threaded loop, default_permissions. Can also mount from an inherited file descriptor, which is how Apptainer's --fusemount works. An experimental low-level variant passes attributes in replies to cut kernel round trips. |
| macOS | macFUSE | Darwin attribute and extended-attribute callbacks. Source build only; the wheel does not include it. |
| Windows | WinFsp (FUSE3 compatibility layer) | A small compatibility header maps POSIX types and identities. No fallocate. |
Kernel cache settings
- Inode numbers from tag identifiers, so tools that compare inode numbers work.
- The page cache is on, so
mmapworks. Writeback caching is opt-in: trusting a lagging size corrupted git packs. - Direct I/O is chosen per file for write-only opens
(avoiding one 4 KiB request per dirtied page), for multi-node
O_APPEND(so the server picks the offset), and when two server files would otherwise share one page cache. - Attribute and negative-lookup caches last 1 s on a
single node and are disabled on multi-node mounts. The entry cache stays
at 1 s even on multi-node, because every
getattrandopenre-resolves the full path at the owner anyway. Keeping it took an 8-nodegitworkload from over 600 s to 58 s. - Unlinked-but-open files stay local to the mount: libfuse's hidden-file rename becomes a server unlink plus a local mapping, and the hidden names are filtered from listings.
Data path: bypassing the ChiMod
The filesystem ChiMod has read and write operations, but the FUSE daemon does not use them by default. Instead, the daemon goes straight to the file's page blobs through the CTE client.
- Writes are copied into the client's write sieve:
64 KiB shared-memory pages, shipped when full or every 500 µs. The
file's logical size is tracked in the client and published at
flush,fsyncorrelease. Pages go to the top of the filesystem's chain, so they are cached and replicated. - Reads are served from in-flight staging first, then the zero-IPC shared-memory path, then a request. Reads whose owner just died are retried for about 10 s, then fail with EIO.
- Appends trust the kernel's size on a single node and use deferred, node-local staging (merged at the stream's home) on multi-node mounts. A synchronous reservation mode is also available.
- Out of space is reported honestly. A refused page fails the next write and the close with ENOSPC, and truncates the file back to the last page that was stored.
- Close is fire-and-forget by default. Durability
lives in
flushandfsync.
Namespace callbacks are thin. Each issues one filesystem request, with
fast paths in front. stat reads the shared-memory mirror
first (about 10 µs, against about 290 µs through the ChiMod), and a
complete-directory record answers "not found" without asking the server.
Before unlink, rename and truncate, the daemon drains the file's pending
writes, so a late size update cannot resurrect a truncated file or land
data under an old name.
Performance notes
- Kernel transitions dominate. In the low-level variant's measurements, about 91% of per-operation wall time was kernel transit and only about 9.6 µs was daemon CPU time. That is why the design invests so much in caches that answer without a request.
- Representative costs:
statabout 31 µs through the ChiMod,readdir255 µs for 10 entries and 592 µs for 100, rename about 570 µs. A Linux kernelgit clonetook 171 s (ext4: 39 s), down from over 15 minutes before the FUSE overhaul. - Large directories spread across many blocks and
nodes.
readdirand a directory'sstatstill visit every block, so very large directories make those calls proportional to their block count. - Each mutation costs one blob write plus one push per holder of the object. Group commit amortizes bursts.
statof a file not homed locally always costs one size request to the stream home, even when its attributes are cached.
Tradeoffs at a glance
| Decision | Buys | Costs |
|---|---|---|
| Namespace stored as CTE blobs | One durability and replication story for data and metadata. No separate metadata store. | Metadata durability is bounded by the CTE configuration, for example RAM-only tiers. |
| Directory blocks keyed by directory identifier | Constant-time directory rename. Directories scale across nodes. | Listings and directory stats visit every block. Blocks never merge. |
| Local resolution with pushed caches | Warm path walks with no network traffic. | Each change pushes to every holder, and holders must be tracked. |
| Leases, failing closed | Never serves stale namespace state. | Mutations can stall about 12 s behind an unreachable holder, and cut-off nodes get EIO. |
| Multi-step protocols instead of transactions | No distributed locks or two-phase commit. Crashes leak instead of corrupting. | Rename is not one atomic step, and leaked link counts are possible after crashes. |
| FUSE data path straight to blobs | Avoids a server-side hop and per-page serialization. | Write-behind means errors surface at close or fsync, not at write. |
| Fast paths only on single node | Microsecond stat locally. | Multi-node mounts pay a request per stat or lookup miss. |