Docs · Reference

Tags, blobs, tiers and pools.

Six ideas explain almost everything CLIO Core does: tags and blobs, files as tags, tiers and scores, persistence, pools, and the module chain.

Written for
Anyone configuring or scripting CLIO Core
Read time
8 minutes

Tags and blobs

The storage engine (the CTE) stores blobs: named byte ranges, up to any size. Blobs are grouped under tags, which are named containers. A tag is something like a directory. A blob is something like a file or a piece of one.

Both have names you choose, and both have numeric ids that the engine assigns. The Python and C++ APIs let you put and get blobs, list a tag's blobs, and query tags and blobs by regular expression.

Files are tags

The filesystem layer builds a POSIX namespace out of tags:

  • Every file, directory and symbolic link is a tag named by its path. The tag id doubles as the inode number.
  • A file's bytes are stored as page blobs of 1 MiB, named "0", "1", "2" and so on by offset.
  • Each of those tags also holds one small blob named ~i with the entry's attributes (mode, size, times). Directories have only that blob.

That is why the Python search APIs return a tag id and a blob name: together they mean "this page of this file". The filesystem design page explains the namespace, consistency and crash safety.

Tiers and scores

A storage tier (also called a target) is one place the engine can put bytes: a slice of RAM, a file on an NVMe drive, a file on a disk. Each tier has a capacity and a score from 0.0 (slowest) to 1.0 (fastest).

Blobs carry scores too. The data placement engine puts a new blob on the best tier that has room, using a policy you choose (max_bw, round_robin or random). Later, the organizer moves blobs whose score no longer matches their tier: hot data up, cold data down. With the frecency organizer, scores follow how recently and how often data is used.

Tier typebdev_typepath
DRAMramram::<name>
NVMe, SSD, HDD, parallel filesystemfileA file path on that device
GPU memoryhbm, pinnedA name; needs a GPU build
Object stores3, gcss3://bucket/prefix; needs a build with the cloud options
Latency testingnoopDiscards data

There is no separate NVMe or SSD type. The device is whatever the file path lives on.

Persistence levels

Each tier declares how long its contents last: volatile (lost at shutdown; the default), temporary (survives a restart) or long_term. The replication layer keeps its persistent copies only on non-volatile tiers. A configuration whose tiers are all volatile cannot keep data across a restart.

The metadata log (performance.metadata_log_path) records which blobs exist and where. Without it, a disk tier's bytes survive but nothing remembers what they belong to.

Pools and modules

Everything inside the runtime is a module (also called a ChiMod) running in a pool. A pool has a name, a module type and an id such as 512.0. The compose section of clio.yaml lists the pools to create at startup. clio_run compose start file.yaml adds more later.

Pool idModuleRole
301.0clio_bdevDefault RAM block device
400.0clio_cae_coreAssimilation (imports)
512.0clio_cte_coreStorage engine with your tiers
561.0clio_cte_replicationPersistent copies
562.0clio_cte_compressorOptional compression
563.0clio_cte_cacheNode-local copies; top of the chain
564.0clio_cte_indexerKeyword search index
565.0clio_cte_streamByte streams: page layout, file sizes, appends
560.0clio_cte_filesystemThe filesystem namespace

The module chain

Replication, caching, indexing and compression are separate modules that all speak the storage engine's interface. Each one forwards to the next through next_pool_id. In the default configuration, a write from the filesystem flows like this:

data path
filesystem 560.0 → cache 563.0 → indexer 564.0 → [compressor 562.0] → replication 561.0 → core 512.0

Any client can bind to any layer. Set CLIO_CTE_POOL=564.0, for example, and your Python client talks to the indexer, which serves search and forwards everything else. The interposition chain design page describes each layer's guarantees.