Tutorial · Compose
Compose the context filesystem from local storage.
Describe your RAM, NVMe, SSD and disks in the storage list of clio.yaml. CLIO Core places new data on the fastest tier with room, and keeps a persistent copy on the tier you mark as durable.
- Time
- 15 minutes
- You need
- A working install, a few GB on each device
- Result
- A four-tier filesystem
How a tier is described
Each entry under the CTE pool's storage list is one tier:
- path: "/mnt/nvme/clio/tier" # "ram::<name>" for DRAM, or a file path on the device
bdev_type: "file" # ram | file (hbm, pinned on GPU builds)
capacity_limit: "200GB" # required; "0g" on a ram tier means 80% of DRAM
score: 0.9 # 0.0 slowest ... 1.0 fastest; omit for automatic
persistence_level: "temporary" # volatile (default) | temporary | long_term| Key | Meaning |
|---|---|
path | Where the tier lives. NVMe, SSD and HDD are all file tiers; the device is whatever the path is on. ${HOME} and ${CLIO_STORAGE_ROOT} are expanded. CLIO Core appends _node<N> to the path for each node it places data on. |
capacity_limit | The most the tier may use. File tiers grow lazily, so this is a cap, not an upfront allocation. |
score | Tier speed, from 0 to 1. Placement prefers higher scores. The organizer moves blobs toward the tier that matches their own score. |
persistence_level | How long data lasts. Persistent copies only go to temporary or long_term tiers. A volatile-only setup loses everything on shutdown. |
1. Start from the default configuration
Copy the default file, so you keep the module chain (replication, indexer, cache, filesystem) and change only the tiers:
$ cp ~/.clio/clio.yaml ~/clio-tiers.yaml
$ sudo mkdir -p /mnt/nvme/clio /mnt/ssd/clio /mnt/hdd/clio
$ sudo chown $USER /mnt/nvme/clio /mnt/ssd/clio /mnt/hdd/clioUse your own mount points. On a laptop, two directories on the internal SSD are enough to follow along.
2. Describe the tiers
Open ~/clio-tiers.yaml and find the entry with mod_name: clio_cte_core. Replace its storage list:
- mod_name: clio_cte_core
pool_name: cte_main
pool_query: local
pool_id: "512.0"
storage:
- path: "ram::hot" # DRAM: the working set
bdev_type: "ram"
capacity_limit: "16GB"
score: 1.0
- path: "/mnt/nvme/clio/tier" # local NVMe: staged and recent data
bdev_type: "file"
capacity_limit: "200GB"
score: 0.9
- path: "/mnt/ssd/clio/tier" # SATA SSD: durable copies
bdev_type: "file"
capacity_limit: "500GB"
score: 0.5
persistence_level: "temporary"
- path: "/mnt/hdd/clio/tier" # disk: capacity
bdev_type: "file"
capacity_limit: "4TB"
score: 0.2
persistence_level: "long_term"
performance:
metadata_log_path: "${CLIO_STORAGE_ROOT}/cte_metadata_log"
transaction_log_capacity: "32MB"
dpe:
dpe_type: "max_bw" # max_bw | round_robin | randomKeep pool_id: "512.0": the FUSE daemon and every client look for the storage engine there. Keep the metadata_log_path too. Without it, the persistent tiers keep their bytes across a restart but nothing records which file they belong to.
Placement policies
dpe_type | Behavior | Use when |
|---|---|---|
max_bw | Fastest tier with room first | Almost always; the default |
round_robin | Rotates across tiers | Several devices of the same speed |
random | Random tier | Benchmarks and testing |
3. Point replication at your durable tier
The replication module writes a persistent copy of every file. It places those copies with replica_score, which defaults to 0.2. Set it to your durable tier's score. In this example that is the SSD:
- mod_name: clio_cte_replication
pool_name: clio_cte_replication
pool_query: local
pool_id: "561.0"
next_pool_id: "512.0"
num_replicas: 1
cache_score: 1.0
replica_score: 0.5 # matches the SSD tier above4. Start with the new tiers
Use a fresh storage root for the experiment, so the new layout does not meet metadata from the old one:
$ clio_run stop
$ export CLIO_SERVER_CONF=~/clio-tiers.yaml
$ clio_run start --disk /mnt/ssd/clio/state &
$ CLIO_WITH_RUNTIME=0 clio_cte_fuse ~/clio-mnt -f &If a tier is misconfigured, the runtime logs a Config error: line that names the key. Check for one before you mount.
5. Check the result
List the pools
shell$ clio_run compose listOpen the tier roster
Go to 127.0.0.1:8080, then Pools → clio_cte_core. You should see one row per tier, with its score, capacity and free space.
Write data and watch it land
shell$ dd if=/dev/urandom of=~/clio-mnt/sample.bin bs=1M count=2048The RAM row's free space drops first. The SSD row's bytes written grows a moment later, as persistent copies land. Once RAM is full, new pages go to NVMe.
Other layouts
Laptop
storage:
- path: "ram::hot"
bdev_type: "ram"
capacity_limit: "4GB"
score: 1.0
- path: "${CLIO_STORAGE_ROOT}/ssd_tier.dat"
bdev_type: "file"
capacity_limit: "50GB"
score: 0.2
persistence_level: "temporary"HPC node
storage:
- path: "ram::hot"
bdev_type: "ram"
capacity_limit: "64GB"
score: 1.0
- path: "/local/nvme/clio" # node-local burst buffer
bdev_type: "file"
capacity_limit: "1TB"
score: 0.9
- path: "/lustre/myproject/clio" # parallel filesystem
bdev_type: "file"
capacity_limit: "10TB"
score: 0.2
persistence_level: "long_term"RAM only
storage:
- path: "ram::scratch"
bdev_type: "ram"
capacity_limit: "0g" # 80% of DRAM
score: 1.0
# Volatile: everything is lost on stop. Keep the rest of the
# chain as it is; the runtime logs that nothing will survive
# a restart.Add a tier while running
You can register a tier without restarting. Open Pools → clio_cte_core in the dashboard and fill in the register form: a target name, a type and a capacity create a new block device and add it as a tier. A pool id in attach adds an existing block device pool instead. The form calls a REST route you can also use directly:
$ curl -s --data-urlencode name=/mnt/nvme/clio/extra.dat -d bdev_type=file -d capacity=100GB http://127.0.0.1:8080/api/mod/clio_cte_core/512.0/register_target
{"ok":true,"name":"/mnt/nvme/clio/extra.dat","bdev_pool_id":"900.0","attached":false}To attach a block device you composed yourself, compose it first (for example the tiers.yaml on the CLI page), then pass its pool id:
$ clio_run compose start tiers.yaml
$ curl -s --data-urlencode name=ram::scratch -d attach_pool_id=320.0 http://127.0.0.1:8080/api/mod/clio_cte_core/512.0/register_targetTiers added this way start with score 0, so placement uses them last, and they last until the next restart. Add them to clio.yaml to keep them.
Next: add Amazon S3 or Google Cloud Storage, or read how placement and tiering work.