Tutorial · Explore
Query the filesystem from Python.
Files you write through the mount are indexed as they land. From Python you can rank them by keyword relevance, filter them by path, and read their bytes, all without going through the mount.
- Time
- 15 minutes
- Modules
- clio_cte_core_ext, clio_cee
- Search
- BM25 over 1 MiB pages
What you will build
A small script, find.py, that takes a question such as "sea ice extent anomaly" and prints the files in your mount that best match it, with the matching text.
The script uses three facts from core concepts:
- Each file is a tag named by its path inside the mount.
- Its bytes are page blobs named
"0","1", … of 1 MiB each. A small blob named~iholds its attributes, so search patterns use[0-9]+to match pages only. - The indexer (pool
564.0) keeps a BM25 index of every page and serves search.
1. Put some text in the mount
$ mkdir -p ~/clio-mnt/notes
$ echo "Arctic sea ice extent fell to a record low in September." > ~/clio-mnt/notes/ice.md
$ echo "Ocean surface temperatures rose 0.3 C above the mean." > ~/clio-mnt/notes/ocean.md
$ echo "Atmospheric CO2 reached 421 ppm at Mauna Loa." > ~/clio-mnt/notes/co2.md
$ cp ~/papers/*.txt ~/clio-mnt/notes/ # any text you haveThe indexer picks up new writes about every 100 ms, and a search waits for pending writes first. Results always include what you just wrote.
2. Connect
import os, sys
os.environ.setdefault("CLIO_WITH_RUNTIME", "0") # attach to `clio_run start`
os.environ.setdefault("CLIO_CTE_POOL", "564.0") # bind to the indexer layer
import clio_cte_core_ext as cte
cte.clio_init(cte.RuntimeMode.kClient, False)
cte.initialize_cte("", cte.PoolQuery.Dynamic())
client = cte.get_cte_client()Binding to 564.0 is what makes SemanticSearch work: the storage core no longer serves search, and the indexer forwards every other call down the chain. Set both variables before the import.
3. List files by path
Tag and blob patterns are full-match regular expressions, so .*notes/.* matches every file under notes.
paths = client.TagQuery(".*notes/.*", 0)
print(paths)
# ['/notes/ice.md', '/notes/ocean.md', '/notes/co2.md', ...]
# Search results carry tag ids. Build a map back to paths.
by_id = {}
for p in paths:
tid = cte.Tag(p).GetTagId()
by_id[(tid.major_, tid.minor_)] = pUse a pattern that does not depend on the exact prefix (.*notes/.* rather than /notes/.*). Your first TagQuery(".*", 0) shows how names look on your install. Subdirectories match too; they hold no pages, so they never appear in search results.
4. Rank files by relevance
question = " ".join(sys.argv[1:]) or "sea ice extent anomaly"
hits = [h for h in client.SemanticSearch(".*notes/.*", "[0-9]+", question, 5)
if h.score > 0] # 0 means no query word appears
for h in hits:
path = by_id.get((h.tag_id.major_, h.tag_id.minor_), "?")
page = int(h.blob_name)
print(f"{h.score:6.2f} {path} (MiB {page})") 2.41 /notes/ice.md (MiB 0)
0.37 /notes/ocean.md (MiB 0)Scores are BM25: higher means a better match, and they are only comparable within one query. Every page that matches the two patterns is ranked, so pages without any query word come back with a score of 0; the filter drops them. A large file can match on several pages, so group by path if you want one row per file. k caps the results; 0 means no cap.
5. Read the matching text
for h in hits[:3]:
tag = cte.Tag(h.tag_id)
size = tag.GetBlobSize(h.blob_name)
text = tag.GetBlob(h.blob_name, size, 0).decode("utf-8", errors="replace")
print("──", by_id.get((h.tag_id.major_, h.tag_id.minor_)))
print(text[:300])To read a whole file, read its pages in order:
def read_file(path):
tag = cte.Tag(path)
pages = sorted((b for b in tag.GetContainedBlobs() if b.isdigit()), key=int)
return b"".join(tag.GetBlob(p, tag.GetBlobSize(p), 0) for p in pages)This reads straight from the storage tiers and gives the same bytes as open() on the mount, without a FUSE round trip.
Narrow what gets indexed
By default the indexer tokenizes everything. For large binary datasets, restrict it in clio.yaml to the files worth searching:
- mod_name: clio_cte_indexer
pool_name: clio_cte_indexer
pool_query: local
pool_id: "564.0"
next_pool_id: "561.0"
index_log_path: "${CLIO_STORAGE_ROOT}/cte_indexer_index"
tag_re: ".*\\.(md|txt|csv|json)$" # only index text-like files
blob_re: ".*"The index persists under index_log_path, so a restart does not rescan your data.
The higher-level clio_cee API
For documents you add from Python rather than through the mount, clio_cee wraps import, listing, retrieval and deletion in four calls.
import os
os.environ.setdefault("CLIO_WITH_RUNTIME", "0")
import clio_cee as cee
ctx = cee.ContextInterface()
rc = ctx.context_bundle([
cee.AssimilationCtx(src=f"string::{name}", dst="iowarp::climate_docs",
format="string", src_data=text)
for name, text in [("ice", "Arctic sea ice extent fell to a record low."),
("co2", "Atmospheric CO2 reached 421 ppm.")]
])
print(rc) # 0 on success
print(ctx.context_query("climate_docs", ".*")) # ['ice', 'co2'] in some order
print(ctx.context_retrieve("climate_docs", ".*")) # the texts, packed into one string
ctx.context_destroy(["climate_docs"])context_query returns blob names only, which is why the tutorial above uses clio_cte_core_ext for mounted files. The full signatures are in the Python API reference.
Keyword search over imported data
context_query(..., prompt="sea ice") ranks imported blobs by BM25. In the default configuration it returns nothing: the import engine (pool 400.0) writes straight to the storage core, below the indexer, so imports are never indexed. To index them, send imports through the top of the chain. In clio.yaml, move the clio_cae_core entry below the clio_cte_cache entry (an entry must come after the pool it forwards to) and change its next_pool_id:
- mod_name: clio_cae_core
pool_name: cae_main
pool_query: local
pool_id: "400.0"
next_pool_id: "563.0" # chain top (cache), so imports get indexedRestart the runtime and import again. Then ctx.context_query("climate_docs", ".*", max_results=2, prompt="sea ice minimum") returns ['ice', 'co2'], best match first. As with SemanticSearch, blobs without any query word are still ranked, after the matches.