Docs · Reference
The Python modules.
Two modules ship in the wheel. clio_cte_core_ext works directly with tags and blobs, including files in the mount. clio_cee is a higher-level API for bundling, querying and retrieving data.
- Installed by
- pip install iowarp-core
- Releases the GIL
- During every RPC
Connecting
Both modules attach to a running runtime when CLIO_WITH_RUNTIME=0. Set environment variables before the import, since the client reads them once when it connects.
import os
os.environ["CLIO_WITH_RUNTIME"] = "0"
os.environ["CLIO_CTE_POOL"] = "564.0" # needed for SemanticSearch
import clio_cte_core_ext as cte
cte.clio_init(cte.RuntimeMode.kClient, False)
cte.initialize_cte("", cte.PoolQuery.Dynamic())
client = cte.get_cte_client()clio_cte_core_ext
Tag
| Call | Returns |
|---|---|
Tag(name) | Opens the tag, creating it if needed. Tag(tag_id) wraps an existing id. |
.PutBlob(name, data: bytes, off=0) | Writes bytes into a blob |
.GetBlob(name, size, off=0) | bytes |
.GetBlobSize(name) | Size in bytes |
.GetBlobScore(name) | Placement score, 0.0 to 1.0 |
.GetContainedBlobs() | List of blob names |
.ReorganizeBlob(name, score) | Moves a blob by giving it a new score |
.GetTagId() | TagId with major_ and minor_ |
Client
| Call | Returns |
|---|---|
TagQuery(tag_regex, max_tags=0) | List of tag names |
BlobQuery(tag_regex, blob_regex, max_blobs=0) | List of (tag_name, blob_name) pairs |
SemanticSearch(tag_regex, blob_regex, query_text, k=10) | List of SemanticSearchResult (tag_id, blob_name, score), best first |
TemporalSearch(tag_regex, blob_regex, time_begin=0, time_end=0, max_entries=0) | List of TemporalSearchResult (tag_id, blob_name, last_modified) |
DelBlob(...), PollTelemetryLog(...) | Delete a blob; read recent operations |
Regular expressions must match the whole name (std::regex_match), so use .*part.* to match a substring. SemanticSearch ranks every blob that matches both patterns, including blobs with no query word (score 0). For files in the mount, use "[0-9]+" as the blob pattern to search pages and skip each file's ~i attribute blob.
last_modified currently comes from a steady clock that counts from boot, not from the Unix epoch, although the docstrings say epoch nanoseconds. A time_begin computed from time.time() matches nothing. Until this is fixed, treat last_modified as an ordering key within one run of the runtime, and do not filter by wall-clock time.
Every query also takes a pool_query argument: PoolQuery.Broadcast() (the default) asks every node, and PoolQuery.Local() asks only this one. Async versions return futures.
clio_cee
The Context Exploration Engine bundles data in, then queries and retrieves it. It suits documents you add from Python rather than files in the mount, because its query returns blob names only.
import clio_cee as cee
ctx = cee.ContextInterface() # attaches to the running runtime
ctx.context_bundle([
cee.AssimilationCtx(src="string::ice", dst="iowarp::climate_docs",
format="string",
src_data="Arctic sea ice extent fell to a record low."),
cee.AssimilationCtx(src="file::/data/run42/output.bin", dst="iowarp::run42",
format="binary"),
])
names = ctx.context_query("climate_docs", ".*") # ['ice']
payload = ctx.context_retrieve("climate_docs", ".*", max_results=5)
ctx.context_destroy(["climate_docs"])| Method | Notes |
|---|---|
context_bundle(list[AssimilationCtx]) | Imports each source into the destination tag. Returns 0 on success. |
context_query(tag_re, blob_re, max_results=0, prompt="", time_begin=0, time_end=0) | List of blob names. A time bound selects temporal mode; else a prompt selects BM25 (top 10 if max_results is 0); else regex. BM25 over imported data needs imports routed through the indexer; see keyword search over imported data. |
context_retrieve(tag_re, blob_re, max_results=1024, max_context_size=256 MiB, batch_size=32, prompt="", time_begin=0, time_end=0) | List of strings, each holding matching blobs packed back to back |
context_destroy(list[str]) | Deletes the named tags and their blobs |
AssimilationCtx(src, dst, format, depends_on="", range_off=0, range_size=0, src_token="", dst_token="") describes one import. Sources use URL schemes: file:: (format binary), string:: (with src_data), hdf5:: in builds with HDF5, globus:// with Globus, and s3:// or gs:// in builds with the cloud options. Destinations are iowarp::<tag>.