Tutorial · Compose

Bring in S3 and Google Cloud.

Object stores connect to CLIO Core in two ways. The assimilation engine imports s3:// and gs:// objects. A build with the cloud options can also use a bucket as a storage tier.

Time
20 minutes
You need
A source build with cloud options, and bucket credentials
Status
Not yet verified end to end
Not yet verified

The commands on this page have not been run end to end against Amazon S3 or Google Cloud Storage. The other tutorials were checked on a live system; this one was checked only against the source code. Expect to adjust details, and please report what you find.

Which one do you need?

Import (CAE)Bucket as a tier
What it doesCopies objects into CLIO Core onceStores tiered data in a bucket
Data ends upOn your local tiers, searchable from PythonIn the bucket
Build optionsCAE_ENABLE_S3, CAE_ENABLE_GCSCLIO_ENABLE_AMAZON_DRIVE, CLIO_ENABLE_GOOGLE_CLOUD
StatusSupported, with unit testsConfig parsing tested; no end-to-end test yet

The pip wheel includes neither. Build from source with the options you need, or use Spack (spack install iowarp +s3 +gcs).

shell
$ cmake --preset release-fuse \
    -DCAE_ENABLE_S3=ON -DCAE_ENABLE_GCS=ON \
    -DCLIO_ENABLE_AMAZON_DRIVE=ON -DCLIO_ENABLE_GOOGLE_CLOUD=ON
$ cmake --build build -j"$(nproc)" && sudo cmake --install build

S3 support needs the AWS SDK for C++ (aws-cpp-sdk-s3). The GCS importer needs google-cloud-cpp, and the GCS block device needs Poco (Net, NetSSL, Crypto, JSON). If one is missing, the build turns that option off with a warning.

Import objects

1. Provide credentials

Amazon S3

shell
$ export AWS_ACCESS_KEY_ID=...
$ export AWS_SECRET_ACCESS_KEY=...
$ export AWS_DEFAULT_REGION=us-east-1
$ export S3_ENDPOINT=https://s3.us-east-1.amazonaws.com   # or MinIO, Ceph, ...

AWS_SESSION_TOKEN is read for temporary credentials. AWS_ENDPOINT_URL works in place of S3_ENDPOINT. Set these in the environment of clio_run start, because downloads run inside the runtime.

Google Cloud Storage

shell
$ gcloud auth application-default login
# or, for a service account:
$ export GOOGLE_APPLICATION_CREDENTIALS=$HOME/keys/clio-sa.json

The importer uses Application Default Credentials: the variable above, gcloud's stored login, or the metadata server on Google Cloud VMs. GCS_ENDPOINT overrides the endpoint, for example for an emulator.

2. Import from Python

python
import os
os.environ["CLIO_WITH_RUNTIME"] = "0"
import clio_cee as cee

ctx = cee.ContextInterface()
rc = ctx.context_bundle([
    cee.AssimilationCtx(src="s3://my-lab-bucket/runs/run42/output.bin",
                        dst="iowarp::run42", format="binary"),
    cee.AssimilationCtx(src="gs://my-lab-bucket/papers/ice-2024.txt",
                        dst="iowarp::papers", format="binary"),
])
print("imported" if rc == 0 else f"failed: {rc}")

Each object becomes data under the destination tag, stored on your local tiers. range_off and range_size import part of an object. Imports in one bundle can depend on each other through depends_on.

3. Or describe imports in YAML

The same imports can be written as an OMNI file, which suits repeatable pipelines:

yaml · imports.yaml
transfers:
  - src: s3://my-lab-bucket/runs/run42/output.bin
    dst: iowarp::run42
    format: binary
    depends_on: ""
    range_off: 0
    range_size: 0         # 0 = the whole object
  - src: gs://my-lab-bucket/papers/ice-2024.txt
    dst: iowarp::papers
    format: binary
Imported data and the mount

Imports create CTE tags, not filesystem entries, so they do not show up under the mount. Find them from Python with TagQuery, context_query or SemanticSearch. To make an object appear as a file, copy it into the mount instead: aws s3 cp s3://bucket/key ~/clio-mnt/.

Buckets as tiers

A build with CLIO_ENABLE_AMAZON_DRIVE or CLIO_ENABLE_GOOGLE_CLOUD can use a bucket as a storage tier. Give the tier bdev_type: "s3" or "gcs" and a path of the form s3://bucket/prefix or gcs://bucket/prefix. The prefix is required: CLIO Core appends _node<N> to the path for each node, and without a prefix that would change the bucket name.

1. Credentials for a bucket tier

Amazon S3

shell
$ export AWS_DEFAULT_REGION=us-east-1
$ export AWS_ACCESS_KEY_ID=...  AWS_SECRET_ACCESS_KEY=...

Credentials come from the AWS SDK's usual sources: environment variables, ~/.aws/credentials, or an instance role.

Google Cloud Storage

shell
$ export GCS_ACCESS_TOKEN=$(gcloud auth print-access-token)
$ export GCS_PROJECT_ID=my-project        # used only when creating the bucket

Access tokens expire after about an hour, so refresh the token and restart the runtime for long runs.

2. Add the bucket to the tier list

In the clio_cte_core entry of clio.yaml, add the bucket below your local tiers:

yaml
    storage:
      - path: "ram::hot"
        bdev_type: "ram"
        capacity_limit: "16GB"
        score: 1.0
      - path: "${CLIO_STORAGE_ROOT}/cte_disk_tier.dat"
        bdev_type: "file"
        capacity_limit: "100GB"
        score: 0.2
        persistence_level: "temporary"
      - path: "s3://my-lab-bucket/clio"      # or gcs://my-lab-bucket/clio
        bdev_type: "s3"                      # or gcs
        capacity_limit: "1TB"
        score: 0.05
        persistence_level: "long_term"

Give the bucket the lowest score, so only data the organizer demotes reaches it (see tiering). The CTE roster in the dashboard lists it as a tier.

If the runtime logs CLIO_ENABLE_AMAZON_DRIVE is not defined. Cannot use S3 bdev., the build lacks the S3 option. Reconfigure with -DCLIO_ENABLE_AMAZON_DRIVE=ON.