Python API
pip install recern-vector
import recern_vector as rv
The package ships type hints (py.typed). Requires CPython 3.11 or later.
Database
A database file and its collections. Everything is held in memory; changes reach the disk on save() or at the end of a with block.
| Member | Description |
|---|---|
rv.Database(path) |
Open an existing file. Raises OSError if it cannot be read, CorruptDatabaseError if it is not a valid database |
rv.Database.create(path) |
Create a new, empty file. Raises FileExistsError if it exists |
rv.Database.open_or_create(path) |
Open the file, or create it if missing |
db.path |
The file path, as a pathlib.Path |
db.create_collection(name, dim, metric="cosine", m=16, ef_construction=200, ef_search=64) |
Create a collection and return it. Raises ValueError if it exists or a setting is invalid. See Concepts |
db.collection(name), db[name] |
An existing collection. Raises KeyError if missing |
name in db |
Whether a collection exists |
db.collection_names() |
Names of all collections |
db.drop_collection(name) |
Delete a collection |
db.save() |
Write the database to disk atomically |
with db: |
Calls save() when the block exits without an exception. On an exception nothing is saved |
Collection
Returned by create_collection and db[name]. A collection object is a handle: it stays valid while the database is open, and sees every change.
| Member | Description |
|---|---|
c.name, c.dim, c.metric |
Settings |
len(c) |
Number of live records |
id in c |
Whether a record exists |
c.upsert(id, vector, metadata=None) |
Insert or replace one record |
c.upsert_many(ids, vectors, metadatas=None, *, threads=None) |
Insert or replace many records, building the index on threads cores (default: all). Atomic. Returns the number of records written |
c.delete(id) |
Remove a record. Returns whether it existed |
c.get(id) |
The Record, or None |
c.search(query, k=10, *, ef=None, exact=False, filter=None) |
The k nearest records, as a list of Hit sorted by distance |
c.explain(query, k=10, *, ef=None, exact=False, filter=None) |
A SearchReport: the hits and how the search ran |
c.stats() |
Index structure, reachability and memory, as a dict. See Inspecting and tuning |
c.estimate_recall(sample=100, k=10, ef_values=(16, 32, 64, 128, 256), seed=42) |
A RecallReport comparing the index with exact search |
c.compact() |
Rebuild without deleted records. Returns how many were removed |
Search arguments:
query: a vector of lengthdim.k: how many results, at least 1.ef: candidate list size for this search; defaults to the collection'sef_searchand is never smaller thank.exact: scan every record instead of using the index.filter: a metadata filter, see Filters.
Vectors
A vector can be a NumPy array, a list or tuple of numbers, or anything that supports the buffer protocol. float32 arrays are read directly; other types are converted.
For upsert_many, vectors is a 2-D array of shape (len(ids), dim) (fastest as contiguous float32) or a sequence of vectors. metadatas, if given, has one entry per id; an entry may be None.
Metadata
Metadata is converted to JSON: dict (string keys), list, tuple, str, int (up to 64 bits), float (finite), bool and None. Other types raise ValueError.
Result types
All result objects are read-only.
Hit: id: str, distance: float, metadata: dict | None. For cosine collections, similarity is 1 - distance.
Record: id: str, vector: list[float] (normalized for cosine collections), metadata.
SearchReport: hits: list[Hit], strategy: "hnsw" | "exact" | "filtered_exact", ef: int | None, visited: int, distance_computations: int, filter_selectivity: float | None, elapsed_ms: float.
RecallReport: k, sample, exact_p50_ms, points: list[RecallPoint].
RecallPoint: ef, recall (0 to 1), p50_ms, p95_ms.
Errors
| Exception | When |
|---|---|
ValueError |
Wrong dimension, NaN or infinite values, a zero vector in a cosine collection, an invalid setting or filter, unsupported metadata, a collection that already exists |
KeyError |
A collection that does not exist |
FileExistsError |
Database.create on an existing file |
OSError |
The file cannot be read or written |
rv.CorruptDatabaseError |
The file is not a Recern Vector database, is damaged, or was written by a newer format version. Subclass of rv.RecernVectorError |
rv.FORMAT_VERSION is the file format version this build writes.
Threads
Searches, get, stats, estimate_recall and save release the GIL and run in parallel from several threads. Writes wait for running reads and then run alone; upsert_many and compact use all cores themselves.
from concurrent.futures import ThreadPoolExecutor
with ThreadPoolExecutor(8) as pool:
results = list(pool.map(lambda q: collection.search(q, k=10), queries))