Docs · Reference

The Python modules.

Two modules ship in the wheel. clio_cte_core_ext works directly with tags and blobs, including files in the mount. clio_cee is a higher-level API for bundling, querying and retrieving data.

Installed by
pip install iowarp-core
Releases the GIL
During every RPC

Connecting

Both modules attach to a running runtime when CLIO_WITH_RUNTIME=0. Set environment variables before the import, since the client reads them once when it connects.

python
import os
os.environ["CLIO_WITH_RUNTIME"] = "0"
os.environ["CLIO_CTE_POOL"] = "564.0"    # needed for SemanticSearch

import clio_cte_core_ext as cte
cte.clio_init(cte.RuntimeMode.kClient, False)
cte.initialize_cte("", cte.PoolQuery.Dynamic())
client = cte.get_cte_client()

clio_cte_core_ext

Tag

CallReturns
Tag(name)Opens the tag, creating it if needed. Tag(tag_id) wraps an existing id.
.PutBlob(name, data: bytes, off=0)Writes bytes into a blob
.GetBlob(name, size, off=0)bytes
.GetBlobSize(name)Size in bytes
.GetBlobScore(name)Placement score, 0.0 to 1.0
.GetContainedBlobs()List of blob names
.ReorganizeBlob(name, score)Moves a blob by giving it a new score
.GetTagId()TagId with major_ and minor_

Client

CallReturns
TagQuery(tag_regex, max_tags=0)List of tag names
BlobQuery(tag_regex, blob_regex, max_blobs=0)List of (tag_name, blob_name) pairs
SemanticSearch(tag_regex, blob_regex, query_text, k=10)List of SemanticSearchResult (tag_id, blob_name, score), best first
TemporalSearch(tag_regex, blob_regex, time_begin=0, time_end=0, max_entries=0)List of TemporalSearchResult (tag_id, blob_name, last_modified)
DelBlob(...), PollTelemetryLog(...)Delete a blob; read recent operations

Regular expressions must match the whole name (std::regex_match), so use .*part.* to match a substring. SemanticSearch ranks every blob that matches both patterns, including blobs with no query word (score 0). For files in the mount, use "[0-9]+" as the blob pattern to search pages and skip each file's ~i attribute blob.

Time values are not Unix time yet

last_modified currently comes from a steady clock that counts from boot, not from the Unix epoch, although the docstrings say epoch nanoseconds. A time_begin computed from time.time() matches nothing. Until this is fixed, treat last_modified as an ordering key within one run of the runtime, and do not filter by wall-clock time.

Every query also takes a pool_query argument: PoolQuery.Broadcast() (the default) asks every node, and PoolQuery.Local() asks only this one. Async versions return futures.

clio_cee

The Context Exploration Engine bundles data in, then queries and retrieves it. It suits documents you add from Python rather than files in the mount, because its query returns blob names only.

python
import clio_cee as cee

ctx = cee.ContextInterface()   # attaches to the running runtime
ctx.context_bundle([
    cee.AssimilationCtx(src="string::ice", dst="iowarp::climate_docs",
                        format="string",
                        src_data="Arctic sea ice extent fell to a record low."),
    cee.AssimilationCtx(src="file::/data/run42/output.bin", dst="iowarp::run42",
                        format="binary"),
])
names = ctx.context_query("climate_docs", ".*")     # ['ice']
payload = ctx.context_retrieve("climate_docs", ".*", max_results=5)
ctx.context_destroy(["climate_docs"])
MethodNotes
context_bundle(list[AssimilationCtx])Imports each source into the destination tag. Returns 0 on success.
context_query(tag_re, blob_re, max_results=0, prompt="", time_begin=0, time_end=0)List of blob names. A time bound selects temporal mode; else a prompt selects BM25 (top 10 if max_results is 0); else regex. BM25 over imported data needs imports routed through the indexer; see keyword search over imported data.
context_retrieve(tag_re, blob_re, max_results=1024, max_context_size=256 MiB, batch_size=32, prompt="", time_begin=0, time_end=0)List of strings, each holding matching blobs packed back to back
context_destroy(list[str])Deletes the named tags and their blobs

AssimilationCtx(src, dst, format, depends_on="", range_off=0, range_size=0, src_token="", dst_token="") describes one import. Sources use URL schemes: file:: (format binary), string:: (with src_data), hdf5:: in builds with HDF5, globus:// with Globus, and s3:// or gs:// in builds with the cloud options. Destinations are iowarp::<tag>.