Tutorial · Compose
Bring in S3 and Google Cloud.
Object stores connect to CLIO Core in two ways. The assimilation engine imports s3:// and gs:// objects. A build with the cloud options can also use a bucket as a storage tier.
- Time
- 20 minutes
- You need
- A source build with cloud options, and bucket credentials
- Status
- Not yet verified end to end
The commands on this page have not been run end to end against Amazon S3 or Google Cloud Storage. The other tutorials were checked on a live system; this one was checked only against the source code. Expect to adjust details, and please report what you find.
Which one do you need?
| Import (CAE) | Bucket as a tier | |
|---|---|---|
| What it does | Copies objects into CLIO Core once | Stores tiered data in a bucket |
| Data ends up | On your local tiers, searchable from Python | In the bucket |
| Build options | CAE_ENABLE_S3, CAE_ENABLE_GCS | CLIO_ENABLE_AMAZON_DRIVE, CLIO_ENABLE_GOOGLE_CLOUD |
| Status | Supported, with unit tests | Config parsing tested; no end-to-end test yet |
The pip wheel includes neither. Build from source with the options you need, or use Spack (spack install iowarp +s3 +gcs).
$ cmake --preset release-fuse \
-DCAE_ENABLE_S3=ON -DCAE_ENABLE_GCS=ON \
-DCLIO_ENABLE_AMAZON_DRIVE=ON -DCLIO_ENABLE_GOOGLE_CLOUD=ON
$ cmake --build build -j"$(nproc)" && sudo cmake --install buildS3 support needs the AWS SDK for C++ (aws-cpp-sdk-s3). The GCS importer needs google-cloud-cpp, and the GCS block device needs Poco (Net, NetSSL, Crypto, JSON). If one is missing, the build turns that option off with a warning.
Import objects
1. Provide credentials
Amazon S3
$ export AWS_ACCESS_KEY_ID=...
$ export AWS_SECRET_ACCESS_KEY=...
$ export AWS_DEFAULT_REGION=us-east-1
$ export S3_ENDPOINT=https://s3.us-east-1.amazonaws.com # or MinIO, Ceph, ...AWS_SESSION_TOKEN is read for temporary credentials. AWS_ENDPOINT_URL works in place of S3_ENDPOINT. Set these in the environment of clio_run start, because downloads run inside the runtime.
Google Cloud Storage
$ gcloud auth application-default login
# or, for a service account:
$ export GOOGLE_APPLICATION_CREDENTIALS=$HOME/keys/clio-sa.jsonThe importer uses Application Default Credentials: the variable above, gcloud's stored login, or the metadata server on Google Cloud VMs. GCS_ENDPOINT overrides the endpoint, for example for an emulator.
2. Import from Python
import os
os.environ["CLIO_WITH_RUNTIME"] = "0"
import clio_cee as cee
ctx = cee.ContextInterface()
rc = ctx.context_bundle([
cee.AssimilationCtx(src="s3://my-lab-bucket/runs/run42/output.bin",
dst="iowarp::run42", format="binary"),
cee.AssimilationCtx(src="gs://my-lab-bucket/papers/ice-2024.txt",
dst="iowarp::papers", format="binary"),
])
print("imported" if rc == 0 else f"failed: {rc}")Each object becomes data under the destination tag, stored on your local tiers. range_off and range_size import part of an object. Imports in one bundle can depend on each other through depends_on.
3. Or describe imports in YAML
The same imports can be written as an OMNI file, which suits repeatable pipelines:
transfers:
- src: s3://my-lab-bucket/runs/run42/output.bin
dst: iowarp::run42
format: binary
depends_on: ""
range_off: 0
range_size: 0 # 0 = the whole object
- src: gs://my-lab-bucket/papers/ice-2024.txt
dst: iowarp::papers
format: binaryImports create CTE tags, not filesystem entries, so they do not show up under the mount. Find them from Python with TagQuery, context_query or SemanticSearch. To make an object appear as a file, copy it into the mount instead: aws s3 cp s3://bucket/key ~/clio-mnt/.
Buckets as tiers
A build with CLIO_ENABLE_AMAZON_DRIVE or CLIO_ENABLE_GOOGLE_CLOUD can use a bucket as a storage tier. Give the tier bdev_type: "s3" or "gcs" and a path of the form s3://bucket/prefix or gcs://bucket/prefix. The prefix is required: CLIO Core appends _node<N> to the path for each node, and without a prefix that would change the bucket name.
1. Credentials for a bucket tier
Amazon S3
$ export AWS_DEFAULT_REGION=us-east-1
$ export AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=...Credentials come from the AWS SDK's usual sources: environment variables, ~/.aws/credentials, or an instance role.
Google Cloud Storage
$ export GCS_ACCESS_TOKEN=$(gcloud auth print-access-token)
$ export GCS_PROJECT_ID=my-project # used only when creating the bucketAccess tokens expire after about an hour, so refresh the token and restart the runtime for long runs.
2. Add the bucket to the tier list
In the clio_cte_core entry of clio.yaml, add the bucket below your local tiers:
storage:
- path: "ram::hot"
bdev_type: "ram"
capacity_limit: "16GB"
score: 1.0
- path: "${CLIO_STORAGE_ROOT}/cte_disk_tier.dat"
bdev_type: "file"
capacity_limit: "100GB"
score: 0.2
persistence_level: "temporary"
- path: "s3://my-lab-bucket/clio" # or gcs://my-lab-bucket/clio
bdev_type: "s3" # or gcs
capacity_limit: "1TB"
score: 0.05
persistence_level: "long_term"Give the bucket the lowest score, so only data the organizer demotes reaches it (see tiering). The CTE roster in the dashboard lists it as a tier.
If the runtime logs CLIO_ENABLE_AMAZON_DRIVE is not defined. Cannot use S3 bdev., the build lacks the S3 option. Reconfigure with -DCLIO_ENABLE_AMAZON_DRIVE=ON.