Docs · Reference
Tags, blobs, tiers and pools.
Six ideas explain almost everything CLIO Core does: tags and blobs, files as tags, tiers and scores, persistence, pools, and the module chain.
- Written for
- Anyone configuring or scripting CLIO Core
- Read time
- 8 minutes
Tags and blobs
The storage engine (the CTE) stores blobs: named byte ranges, up to any size. Blobs are grouped under tags, which are named containers. A tag is something like a directory. A blob is something like a file or a piece of one.
Both have names you choose, and both have numeric ids that the engine assigns. The Python and C++ APIs let you put and get blobs, list a tag's blobs, and query tags and blobs by regular expression.
Files are tags
The filesystem layer builds a POSIX namespace out of tags:
- Every file, directory and symbolic link is a tag named by its path. The tag id doubles as the inode number.
- A file's bytes are stored as page blobs of 1 MiB, named
"0","1","2"and so on by offset. - Each of those tags also holds one small blob named
~iwith the entry's attributes (mode, size, times). Directories have only that blob.
That is why the Python search APIs return a tag id and a blob name: together they mean "this page of this file". The filesystem design page explains the namespace, consistency and crash safety.
Tiers and scores
A storage tier (also called a target) is one place the engine can put bytes: a slice of RAM, a file on an NVMe drive, a file on a disk. Each tier has a capacity and a score from 0.0 (slowest) to 1.0 (fastest).
Blobs carry scores too. The data placement engine puts a new blob on the best tier that has room, using a policy you choose (max_bw, round_robin or random). Later, the organizer moves blobs whose score no longer matches their tier: hot data up, cold data down. With the frecency organizer, scores follow how recently and how often data is used.
| Tier type | bdev_type | path |
|---|---|---|
| DRAM | ram | ram::<name> |
| NVMe, SSD, HDD, parallel filesystem | file | A file path on that device |
| GPU memory | hbm, pinned | A name; needs a GPU build |
| Object store | s3, gcs | s3://bucket/prefix; needs a build with the cloud options |
| Latency testing | noop | Discards data |
There is no separate NVMe or SSD type. The device is whatever the file path lives on.
Persistence levels
Each tier declares how long its contents last: volatile (lost at shutdown; the default), temporary (survives a restart) or long_term. The replication layer keeps its persistent copies only on non-volatile tiers. A configuration whose tiers are all volatile cannot keep data across a restart.
The metadata log (performance.metadata_log_path) records which blobs exist and where. Without it, a disk tier's bytes survive but nothing remembers what they belong to.
Pools and modules
Everything inside the runtime is a module (also called a ChiMod) running in a pool. A pool has a name, a module type and an id such as 512.0. The compose section of clio.yaml lists the pools to create at startup. clio_run compose start file.yaml adds more later.
| Pool id | Module | Role |
|---|---|---|
301.0 | clio_bdev | Default RAM block device |
400.0 | clio_cae_core | Assimilation (imports) |
512.0 | clio_cte_core | Storage engine with your tiers |
561.0 | clio_cte_replication | Persistent copies |
562.0 | clio_cte_compressor | Optional compression |
563.0 | clio_cte_cache | Node-local copies; top of the chain |
564.0 | clio_cte_indexer | Keyword search index |
565.0 | clio_cte_stream | Byte streams: page layout, file sizes, appends |
560.0 | clio_cte_filesystem | The filesystem namespace |
The module chain
Replication, caching, indexing and compression are separate modules that all speak the storage engine's interface. Each one forwards to the next through next_pool_id. In the default configuration, a write from the filesystem flows like this:
filesystem 560.0 → cache 563.0 → indexer 564.0 → [compressor 562.0] → replication 561.0 → core 512.0Any client can bind to any layer. Set CLIO_CTE_POOL=564.0, for example, and your Python client talks to the indexer, which serves search and forwards everything else. The interposition chain design page describes each layer's guarantees.