Tutorial · Explore

Query the filesystem from Python.

Files you write through the mount are indexed as they land. From Python you can rank them by keyword relevance, filter them by path, and read their bytes, all without going through the mount.

Time
15 minutes
Modules
clio_cte_core_ext, clio_cee
Search
BM25 over 1 MiB pages

What you will build

A small script, find.py, that takes a question such as "sea ice extent anomaly" and prints the files in your mount that best match it, with the matching text.

The script uses three facts from core concepts:

  • Each file is a tag named by its path inside the mount.
  • Its bytes are page blobs named "0", "1", … of 1 MiB each. A small blob named ~i holds its attributes, so search patterns use [0-9]+ to match pages only.
  • The indexer (pool 564.0) keeps a BM25 index of every page and serves search.

1. Put some text in the mount

shell
$ mkdir -p ~/clio-mnt/notes
$ echo "Arctic sea ice extent fell to a record low in September."  > ~/clio-mnt/notes/ice.md
$ echo "Ocean surface temperatures rose 0.3 C above the mean."      > ~/clio-mnt/notes/ocean.md
$ echo "Atmospheric CO2 reached 421 ppm at Mauna Loa."             > ~/clio-mnt/notes/co2.md
$ cp ~/papers/*.txt ~/clio-mnt/notes/          # any text you have

The indexer picks up new writes about every 100 ms, and a search waits for pending writes first. Results always include what you just wrote.

2. Connect

python · find.py
import os, sys
os.environ.setdefault("CLIO_WITH_RUNTIME", "0")   # attach to `clio_run start`
os.environ.setdefault("CLIO_CTE_POOL", "564.0")   # bind to the indexer layer

import clio_cte_core_ext as cte

cte.clio_init(cte.RuntimeMode.kClient, False)
cte.initialize_cte("", cte.PoolQuery.Dynamic())
client = cte.get_cte_client()

Binding to 564.0 is what makes SemanticSearch work: the storage core no longer serves search, and the indexer forwards every other call down the chain. Set both variables before the import.

3. List files by path

Tag and blob patterns are full-match regular expressions, so .*notes/.* matches every file under notes.

python
paths = client.TagQuery(".*notes/.*", 0)
print(paths)
# ['/notes/ice.md', '/notes/ocean.md', '/notes/co2.md', ...]

# Search results carry tag ids. Build a map back to paths.
by_id = {}
for p in paths:
    tid = cte.Tag(p).GetTagId()
    by_id[(tid.major_, tid.minor_)] = p

Use a pattern that does not depend on the exact prefix (.*notes/.* rather than /notes/.*). Your first TagQuery(".*", 0) shows how names look on your install. Subdirectories match too; they hold no pages, so they never appear in search results.

4. Rank files by relevance

python
question = " ".join(sys.argv[1:]) or "sea ice extent anomaly"
hits = [h for h in client.SemanticSearch(".*notes/.*", "[0-9]+", question, 5)
        if h.score > 0]                   # 0 means no query word appears

for h in hits:
    path = by_id.get((h.tag_id.major_, h.tag_id.minor_), "?")
    page = int(h.blob_name)
    print(f"{h.score:6.2f}  {path}  (MiB {page})")
example output
  2.41  /notes/ice.md  (MiB 0)
  0.37  /notes/ocean.md  (MiB 0)

Scores are BM25: higher means a better match, and they are only comparable within one query. Every page that matches the two patterns is ranked, so pages without any query word come back with a score of 0; the filter drops them. A large file can match on several pages, so group by path if you want one row per file. k caps the results; 0 means no cap.

5. Read the matching text

python
for h in hits[:3]:
    tag = cte.Tag(h.tag_id)
    size = tag.GetBlobSize(h.blob_name)
    text = tag.GetBlob(h.blob_name, size, 0).decode("utf-8", errors="replace")
    print("──", by_id.get((h.tag_id.major_, h.tag_id.minor_)))
    print(text[:300])

To read a whole file, read its pages in order:

python
def read_file(path):
    tag = cte.Tag(path)
    pages = sorted((b for b in tag.GetContainedBlobs() if b.isdigit()), key=int)
    return b"".join(tag.GetBlob(p, tag.GetBlobSize(p), 0) for p in pages)

This reads straight from the storage tiers and gives the same bytes as open() on the mount, without a FUSE round trip.

Narrow what gets indexed

By default the indexer tokenizes everything. For large binary datasets, restrict it in clio.yaml to the files worth searching:

yaml
  - mod_name: clio_cte_indexer
    pool_name: clio_cte_indexer
    pool_query: local
    pool_id: "564.0"
    next_pool_id: "561.0"
    index_log_path: "${CLIO_STORAGE_ROOT}/cte_indexer_index"
    tag_re: ".*\\.(md|txt|csv|json)$"    # only index text-like files
    blob_re: ".*"

The index persists under index_log_path, so a restart does not rescan your data.

The higher-level clio_cee API

For documents you add from Python rather than through the mount, clio_cee wraps import, listing, retrieval and deletion in four calls.

python
import os
os.environ.setdefault("CLIO_WITH_RUNTIME", "0")
import clio_cee as cee

ctx = cee.ContextInterface()
rc = ctx.context_bundle([
    cee.AssimilationCtx(src=f"string::{name}", dst="iowarp::climate_docs",
                        format="string", src_data=text)
    for name, text in [("ice", "Arctic sea ice extent fell to a record low."),
                       ("co2", "Atmospheric CO2 reached 421 ppm.")]
])
print(rc)                                          # 0 on success
print(ctx.context_query("climate_docs", ".*"))      # ['ice', 'co2'] in some order
print(ctx.context_retrieve("climate_docs", ".*"))   # the texts, packed into one string
ctx.context_destroy(["climate_docs"])

context_query returns blob names only, which is why the tutorial above uses clio_cte_core_ext for mounted files. The full signatures are in the Python API reference.

Keyword search over imported data

context_query(..., prompt="sea ice") ranks imported blobs by BM25. In the default configuration it returns nothing: the import engine (pool 400.0) writes straight to the storage core, below the indexer, so imports are never indexed. To index them, send imports through the top of the chain. In clio.yaml, move the clio_cae_core entry below the clio_cte_cache entry (an entry must come after the pool it forwards to) and change its next_pool_id:

yaml
  - mod_name: clio_cae_core
    pool_name: cae_main
    pool_query: local
    pool_id: "400.0"
    next_pool_id: "563.0"               # chain top (cache), so imports get indexed

Restart the runtime and import again. Then ctx.context_query("climate_docs", ".*", max_results=2, prompt="sea ice minimum") returns ['ice', 'co2'], best match first. As with SemanticSearch, blobs without any query word are still ranked, after the matches.

Next: watch the same data in the dashboard.