Open source BSD 3-Clause

A context filesystem
for scientific data.

CLIO Core turns the RAM, NVMe and disks in your machines into one tiered store, then mounts it as a normal directory. Applications read and write files without changes. You search what they wrote from Python.

Try it today

pip install iowarp-core

01 / Mount

Mount it like
any other directory.

Start the runtime, then mount the filesystem with clio_cte_fuse. Copy files in with cp, open them from a notebook, or point a simulation at the mount. Writes go straight through, so a file's size is exact as soon as write() returns. Linux, macOS and Windows each use their native FUSE driver.

FUSE tutorials for each platform
Linux
$ clio_run start &
$ mkdir -p ~/clio-mnt
$ CLIO_WITH_RUNTIME=0 clio_cte_fuse ~/clio-mnt -f &
$ cp -r ./results ~/clio-mnt/
$ ls ~/clio-mnt/results
Linux
libfuse3. Included in the pip wheel.
Windows
WinFsp, mounted as a drive letter such as Z:. Included in the pip wheel.
macOS
macFUSE, or FSKit on macOS 15.4 and newer. Build from source.

02 / Compose

Compose the storage
you already have.

A short YAML file lists your storage tiers. Each tier gets a capacity and a score from 0 to 1. New data lands on the fastest tier with room. Data you stop touching moves down to slower, larger tiers. A replica on a persistent tier keeps your files across restarts.

Compose RAM, NVMe and SSD tiers
  1. ram::hotDRAM, for the working setscore 1.0
  2. /mnt/nvmeLocal NVMe, for staged datascore 0.9
  3. /mnt/ssdSATA SSD, as a persistent copyscore 0.5
  4. /mnt/hddDisk, for capacityscore 0.2
  5. s3:// · gs://Object stores: import from them, or use a bucket as a tieroptional

03 / Query

Ask for files
by what they say.

The default setup indexes what you write. From Python you can search by keyword relevance (BM25) or by name pattern. You get back the files that match, ranked, and read them without going through the mount.

Query files from Python
python
import os
os.environ["CLIO_WITH_RUNTIME"] = "0"   # attach to clio_run start
os.environ["CLIO_CTE_POOL"] = "564.0"   # search lives in the indexer
import clio_cte_core_ext as cte

cte.clio_init(cte.RuntimeMode.kClient, False)
cte.initialize_cte("", cte.PoolQuery.Dynamic())
client = cte.get_cte_client()

paths = {}
for p in client.TagQuery(".*notes/.*", 0):
    t = cte.Tag(p).GetTagId()
    paths[(t.major_, t.minor_)] = p

for h in client.SemanticSearch(tag_regex=".*notes/.*", blob_regex="[0-9]+",
                               query_text="sea ice extent anomaly", k=5):
    if h.score > 0:
        print(round(h.score, 2), paths[(h.tag_id.major_, h.tag_id.minor_)])

04 / Observe

Watch it work
in the dashboard.

Every runtime serves a web dashboard on 127.0.0.1:8080. It shows the nodes in your cluster, the load on each worker thread, and every storage tier's free space and traffic. You can add or remove pools there too, with no separate process to install.

Navigate the dashboard
Cluster
One card per node, with CPU and memory meters and a task summary.
Pools
Every module on this node. Open the CTE pool to see each tier's score, free space and bytes moved.
Config
Key settings the daemon actually came up with, and every REST endpoint it serves.

05 / Agents

Bring your AI
to your data.

CLIO Core is the data layer of the Clio family. Clio Agent and Clio Coder are IOWarp's agents for science: one manages and analyzes data, the other writes scientific software. Anything they write through the mount is stored, tiered and indexed like the rest of your files.

Query that data from Python

06 / Install

Install the way
your site works.

The pip wheel runs on Linux, Windows and Apple Silicon Macs with Python 3.10 to 3.13. HPC sites can use Spack. Containers can use the Docker image. Build from source for GPUs, MPI, HDF5 or ADIOS2.

Full installation guide

pip

shell
$ pip install iowarp-core
$ python -c "import iowarp_core; print(iowarp_core.get_version())"

conda

shell
$ conda create -n iowarp -c iowarp -c conda-forge iowarp-core
$ conda activate iowarp

Docker

shell
$ docker run -d -p 9413:9413 --memory=8g --name iowarp \
    iowarp/deploy-cpu:latest clio_run start

Spack

shell
$ git clone --recurse-submodules https://github.com/iowarp/clio-core.git
$ spack repo add clio-core/installers/spack
$ spack install iowarp +fuse

Source

shell
$ git clone --recurse-submodules https://github.com/iowarp/clio-core.git
$ cd clio-core
$ cmake --preset release-fuse
$ cmake --build build -j"$(nproc)" && sudo cmake --install build

07 / Under the hood

Five engines,
one runtime.

The filesystem sits on a stack you can also use directly. Each layer has its own C++ and Python interfaces. The Design section explains how they fit together and which tradeoffs they make for speed and reliability.

Read the design

08 / Questions

Common questions.

Do my applications need to change?

No. Through the FUSE mount, CLIO Core looks like a local directory. Applications that use HDF5, ADIOS2 or MPI-IO can also use a plugin instead of the mount for lower overhead.

Does my data survive a restart?

Yes, with the default configuration. Every write is copied to a persistent disk tier and recorded in a metadata log under ~/.clio. clio_run start recovers from that log. RAM-only setups are volatile by design.

Can I use Amazon S3 or Google Cloud Storage?

Yes, in two ways. The assimilation engine imports s3:// and gs:// objects into the filesystem. A build with the cloud options enabled can also use a bucket as a storage tier. See the cloud storage tutorial for what each one supports.

Does it scale past one machine?

Yes. Runtimes on several nodes form a cluster from a hostfile. Each tier can place data on neighboring nodes, and the filesystem namespace has no central metadata server.

How mature is the project?

CLIO Core is under active development. The filesystem, the runtime and the Python bindings run in CI on Linux, macOS and Windows. Report an issue if something does not work.

The Clio ecosystem

Storage, agents and code,
from one team.

Mount it on a machine
you already have.