One file
A database is a single file you can copy, back up or ship with your application. Saves are atomic, and every file is checked with CRC32 when it opens.
Recern Vector is an embedded vector database in the spirit of SQLite. It runs inside your application, keeps everything in one file and shows you how every search ran. We are building it alongside the Recern workspace.
pip install recern-vectorRelease notes, v0.0.1 faiss HNSW throughput at recall 0.95 on SIFT1M
the same comparison on GloVe-100
to index one million vectors on 14 cores
median query latency at recall 0.983 on SIFT1M
A database is a single file you can copy, back up or ship with your application. Saves are atomic, and every file is checked with CRC32 when it opens.
Recern Vector is a library that runs in your process, with bindings for Rust and Python and a command-line tool. There is nothing to deploy, configure or keep running.
Any search can report its strategy, the nodes it visited, how selective the filter was and how long it took. Stats show the graph layers, memory use and any records a search cannot reach, and recall can be measured on your own data.
The same database from Python and from the command line. The output on the right is a real run on 20,000 records with a metadata filter.
import recern_vector as rv
db = rv.Database.open_or_create("docs.rvec")
docs = db.create_collection("docs", dim=384, metric="cosine")
docs.upsert_many(ids, embeddings, metadatas)
hits = docs.search(query, k=5, filter={"lang": "en"})
docs.explain(query, k=5) # strategy, visited nodes, time$ recern-vector query docs.rvec docs --like doc-42 -k 3 \
--filter '{"lang": "en", "year": {"$gte": 2020}}' \
--explain
# id distance metadata
1 doc-42 -0.00000 {"lang":"en","year":2024}
2 doc-18861 0.63155 {"lang":"en","year":2022}
3 doc-10899 0.63834 {"lang":"en","year":2024}
strategy hnsw
ef 64
selectivity ~19.9% of records match the filter
visited 7,244 nodes
distances 7,330 computed
time 529 µsPython bindings take NumPy float32 arrays without converting each element. Batch inserts build the index on all cores.
A graph index with per-collection m, ef_construction and ef_search, and an exact scan when you need ground truth.
Each collection has a fixed dimension and metric. Cosine vectors are normalized when they are stored.
MongoDB-style equality, $in and ranges. When a filter matches very few records, the query switches to an exact scan.
Batch inserts build the index on all cores. If one vector is invalid, nothing is written.
Replace or delete records by id, then compact to rebuild the collection without deleted records.
stats(), explain() and estimate_recall() are part of the API, not a separate tool.
A PyO3 wheel for CPython 3.11 and later, with type hints.
init, create-collection, insert, query, inspect, recall and compact.
One query at a time from Python on Apple M3 Max, with median latency beneath each figure. The full report also covers index build time, size on disk, method and limitations.
| Engine | SIFT1M | GloVe-100 |
|---|---|---|
| Recern VectorHNSW | 9,0410.11 ms | 6351.60 ms |
| faissHNSWFlat | 6,9610.15 ms | 3562.83 ms |
| LanceDBIVF_HNSW_SQ | 7311.34 ms | not reached |
| LanceDBIVF_PQ, tuned | 2454.03 ms | 1675.93 ms |
| sqlite-vecexact scan | 1283.79 ms | 7140.43 ms |
Today the whole database is held in memory, saving rewrites the file, and the file format may still change. Install it with pip install recern-vector or cargo add recern-vector; versions stay at 0.0.x until the format is stable.
Recern is a desktop workspace for understanding databases. It will also explain vector search in databases you already run, such as pgvector and Redis. Recern Vector is a separate companion project and is not part of the workspace Private Alpha.
Join the workspace Private Alpha