Design
I/O adapters
Besides the FUSE mount, applications can reach the CTE through adapters that plug into the I/O libraries they already use. There are two distinct designs. The interceptors make the CTE the file's only home. The HDF5 connectors keep the native file authoritative and use the CTE as a disposable cache.
| Adapter | Mechanism | Where the data lives |
|---|---|---|
| POSIX, STDIO, MPI-IO | Symbol interception (LD_PRELOAD) | Only in the CTE (clio-fs). No backing file. |
| HDF5 VOL connector | HDF5 plugin, above the library | Native file first; dataset images cached in the CTE |
| HDF5 VFD | HDF5 plugin, below the library | Native file first; byte ranges optionally cached in the CTE |
| ADIOS2 engine | ADIOS2 plugin engine | Only in the CTE |
| GPU paths (vector, HDF5-like datasets, GDS) | GPU producer API, or cuFile interception | Only in the CTE |
Interceptors: POSIX, STDIO, MPI-IO
How interception works
Each interceptor library defines the C library (or MPI) entry points
itself and is loaded with LD_PRELOAD. At startup it walks the
process's loaded libraries to find the real implementation of
each function, skipping itself. Every call that is not for a tracked file
goes straight to the real one. A load guard keeps open from
being intercepted while the interceptor is only half initialized.
The interceptors are Linux/ELF only, and require the ELF interception build option.
Which files are intercepted
Selection is opt-in by path: only paths containing a
clio:: component are taken over. The marker is stripped and
the rest becomes a clio-fs path. Everything else, including the
application's ordinary files, passes through untouched. This keeps the
blast radius small: an interceptor can never accidentally capture a
system library or a config file.
A shared descriptor layer
All interceptors (and the HDF5 VFD) use one descriptor layer in the filesystem client. An adapter owns only its handle table and seek offsets. Everything else lives in that layer:
- write-behind staging;
- size tracking;
- zero-IPC reads;
- error latching.
CTE descriptors start at 8192 so they never collide with kernel
descriptors. stat results are synthesized, with a fixed device
number, a path-derived inode number and 1 MiB blocks. Creating a file
whose parent directories exist on the host but not yet in clio-fs creates
them on the fly.
Write-behind
A write() copies the buffer into pooled shared-memory
staging, submits it, and returns, unless the file was opened
O_SYNC or O_DSYNC.
- In-flight window. Bounded per process at 256 writes or 64 MiB. That count was the best measured setting. A byte-only limit let 4 KiB writes queue 16k deep and cost 1.6× throughput.
- Reads see pending writes. A read that is fully covered is served from staging; a partially covered read waits for the overlapping writes.
- Errors surface late. A failed write is latched and
reported once, at
fsyncorclose, the way the kernel reports writeback errors. - Draining first.
stat,SEEK_END, truncate, unlink and rename all drain the file's pending writes before acting.
Semantics and limits
| Adapter | Provides | Does not provide |
|---|---|---|
| POSIX | open/creat, read/write, pread/pwrite, lseek, the stat family, ftruncate, fsync, close, unlink and remove. O_APPEND and O_TRUNC are honored. | mmap, vectored I/O, rename, dup, fcntl, directory iteration and fallocate are not intercepted. flock succeeds without locking. |
| STDIO | fopen/fdopen/freopen, character, line and block I/O, seeking, flush, close | No user-space buffering, so each character is a request. Formatted I/O (fprintf, fscanf) and setvbuf are not intercepted and must not be used on intercepted streams. |
| MPI-IO | Open, close, seek; independent, collective, shared and ordered read and write; nonblocking variants | Collectives run as independent I/O, and shared file pointers are per process. File views and derived datatypes are ignored (the data is treated as contiguous). Status objects are not filled in, and nonblocking calls complete at submission. |
Durability through interceptors. Through these paths there is no backing file on a parallel filesystem: the CTE's tier configuration alone decides durability.
- POSIX and STDIO.
fsyncdrains pending writes and reports errors, but does not force a sync to a persistent device the way the FUSE mount does. Data reaches persistent tiers through the periodic flush. - MPI-IO.
MPI_File_syncdoes nothing, and the result of close is not checked. A failed deferred write through MPI-IO can therefore go unreported. Applications that need strict error reporting should use the FUSE mount or the POSIX interceptor withO_SYNC.
HDF5: native file first, CTE as a cache
Both HDF5 connectors take the opposite stance from the interceptors. Every write goes to the native HDF5 file first, synchronously. The CTE holds only a cache that can be thrown away at any time. They add no data-loss window beyond native HDF5's own.
Coherence stamps
Because the native file can be changed by tools that know nothing about the cache (h5repack, rsync, an editor), each cache is guarded by a coherence stamp. The stamp records the file's device, inode, size and modification time.
- Written at close. When a writing session closes, the stamp is stored next to the cache.
- Checked at open. A missing, unreadable or mismatched stamp means the whole cache is dropped. The check fails closed.
- Ambiguous timestamps. Filesystem timestamps are coarse, so a modification time younger than one granule (10 ms by default, 1 s on Windows) is ambiguous. The stamp is then withheld, and the next open rebuilds the cache. An opt-in mode waits for the timestamp to settle instead, which is safe only for single-writer pipelines.
Modification during a session is undefined, as with native HDF5.
VOL connector versus VFD
| VOL connector | VFD | |
|---|---|---|
| Layer | Object level, above the HDF5 library. Sees datasets, types and selections. | Byte level, below the library. Sees only addresses, sizes and memory types. |
| What is cached | Whole dataset images, chunked into blobs (1 MiB by default) | Byte ranges of the file, in a clio-fs mirror (1 MiB pages) |
| Admission | Policy (on write, on read miss, or on second access), plus a self-tuning cost model that stages only if the native path measured slower than the tier, plus a per-device verdict | One switch: the read tier on or off |
| Read serving | On by default. Whole reads are served from the cache. Partial reads are served only if already cached and never trigger staging. | Off by default. When on, only ranges written or read in this session are served. |
| Metadata | Never cached (groups, attributes and links pass through) | Treated like any other byte range |
| Overhead when it cannot help | Near native: nothing touches the CTE until the first cacheable transfer | Near native with the read tier off |
| Weak spots | A partial write invalidates the whole dataset image. It does not stack over other VOL connectors. | Cannot distinguish hot data from metadata, and tiny HDF5 writes become many requests |
Much of the tuning came from metadata-heavy netCDF-4 workloads. Two examples:
- Lazy binding. The VOL connector does not touch the CTE until the first cacheable transfer, so 32,768 open/close cycles that touch only metadata cost nothing.
- No populate without read serving. With the VFD's read tier off, writes do not populate the cache at all: an attribute-heavy test went from 705 s to 10.8 s.
The VOL connector writes cache chunks as droppable blobs. When the tier is full, it backs off for a few seconds instead of slowing the authoritative write.
ADIOS2 engine
A plugin engine that maps each variable, step and rank to one blob in a tag named after the engine. Synchronous puts block. Deferred puts are staged into shared memory and complete at the end of the step. It stores no variable metadata (shape, offset, type), so readers must use the writer's rank layout and step numbering. Start-up is staggered across ranks, so that thousands of ranks do not all connect at once.
GPU paths
GPU kernels are pure producers. All tasks and buffers are allocated by the host, in memory registered with the runtime, before the kernel launches. Inside the kernel, the only operation is "send this prepared task" and wait for its completion flag. This keeps device code free of allocators and locks.
- GPU vector. A paged array whose pages are blobs, cached in a set-associative page cache in GPU memory shared by all thread blocks. Kernels hold pages through coroutine guards. Fetches and flushes are batched, up to 64 per request.
- The vector never writes back on its own. The caller owns flushing, and a page that is being fetched or flushed is never evicted. A refused fetch or flush stops the kernel with a reported cause, instead of continuing with wrong data.
- Prefetching means rescoring pages so they move up the tier stack. Measured, it did not pay off on the tested splits: one migration cost four to five page faults, and promoting into a full tier first needs a free.
- HDF5-like GPU datasets. One blob per chunk, with a small metadata blob written once per dataset.
- The cuFile (GDS) interceptor bounces through host memory into the POSIX interceptor. It is a compatibility shim, not GPU-direct I/O, and is off by default.
Tradeoffs at a glance
| Decision | Buys | Costs |
|---|---|---|
Opt-in clio:: paths | Interception never captures unrelated files. | Producers and consumers must both use the marker, an interceptor, FUSE or the API. |
| One shared descriptor layer | Every interceptor gets the same write-behind, error and size semantics. | A fixed 1 MiB page size for clio-fs files. |
| Write-behind in interceptors | Small-write throughput close to the page cache's. | Errors arrive at fsync or close. Interceptor fsync is not a device sync. |
| HDF5: native file is the authority | No new data-loss window. The cache can always be discarded. | Writes are never faster than native. The gain is on re-reads. |
| Cost-gated VOL admission | Caching is skipped where it would lose (fast local disk, tiny datasets). | The model needs measurements, so early accesses may be mis-admitted. |
| Simplified MPI-IO | A drop-in for common independent-I/O codes. | Not suitable for codes relying on views, shared pointers or strict error reporting. |