Faster than faiss HNSW at equal recall.
Recern Vector is an embedded vector database we are exploring alongside the Recern workspace: one file, no server, and every search can explain how it ran. We measured the Phase 1 prototype on the two standard ANN-Benchmarks datasets against faiss, LanceDB and sqlite-vec.
Source code and the benchmark harness: github.com/recerndata/recern-vector (bench/)
faiss HNSW throughput at recall ≥ 0.95 on SIFT1M
the same comparison on GloVe-100
to index 1,000,000 SIFT vectors on 14 cores
median query latency at recall 0.983 on SIFT1M
Queries per second at each recall level
Every mark is one search setting. Higher and further right is better: more queries per second at higher recall. Lines join the settings that no other setting of the same engine beats on both axes.
Fastest setting that reaches each recall level
SIFT1M · 1M × 128, Euclidean
| Engine | recall ≥ 0.90 | recall ≥ 0.95 | recall ≥ 0.99 |
|---|---|---|---|
| Recern VectorHNSW | 15,9730.06 ms | 9,0410.11 ms | 4,4090.22 ms |
| faissHNSWFlat | 13,0120.08 ms | 6,9610.14 ms | 3,5710.28 ms |
| LanceDBIVF_HNSW_SQ | 7651.28 ms | 7311.34 ms | not reached |
| LanceDBIVF_PQ, tuned | 2454.03 ms | 2454.03 ms | 2454.03 ms |
| sqlite-vecexact scan | 1283.78 ms | 1283.78 ms | 1283.78 ms |
GloVe-100 · 1.18M × 100, cosine
| Engine | recall ≥ 0.90 | recall ≥ 0.95 | recall ≥ 0.99 |
|---|---|---|---|
| Recern VectorHNSW | 2,3730.43 ms | 6351.59 ms | not reached |
| faissHNSWFlat | 1,6870.60 ms | 3562.83 ms | not reached |
| LanceDBIVF_HNSW_SQ | not reached | not reached | not reached |
| LanceDBIVF_PQ, tuned | 2713.68 ms | 1675.93 ms | 8112.14 ms |
| sqlite-vecexact scan | 7140.43 ms | 7140.43 ms | 7140.43 ms |
Queries per second, with median latency beneath. LanceDB's default IVF_PQ settings reach at most recall 0.36 on GloVe-100; the tuned configuration uses num_sub_vectors = dim / 4.
Building a million-vector index takes under a minute
Build time, HNSW m = 16, ef_construction = 200
| Dataset | Recern, 1 thread | Recern, 14 threads | faiss, 1 thread | faiss, 14 threads | LanceDB IVF_HNSW_SQ |
|---|---|---|---|---|---|
| SIFT1M | 345 s | 36 s | 284 s | 32 s | 28 s |
| GloVe-100 | 481 s | 47 s | 416 s | 43 s | 44 s |
Size on disk
| Engine | SIFT1M | GloVe-100 |
|---|---|---|
| Recern VectorHNSW | 631 MB | 620 MB |
| faissHNSWFlat | 626 MB | 614 MB |
| LanceDBIVF_HNSW_SQ | 789 MB | 769 MB |
| LanceDBIVF_PQ, tuned | 524 MB | 487 MB |
| sqlite-vecexact scan | 511 MB | 479 MB |
Recern's build scales about ten times on 14 cores and stays within 1.1–1.2× of faiss on the same hardware. LanceDB uses all cores; its HNSW index is partitioned and scalar-quantized, which speeds up the build and caps recall below 0.99.
What made the difference
Prefetching neighbor vectors
A graph search spends most of its time waiting for vectors to arrive from memory. Recern now gathers a node's unvisited neighbors, asks the CPU to start loading every cache line of their vectors, and only then computes distances, so the loads overlap. Median latency at ef = 80 on SIFT1M fell from 206 µs to 110 µs with identical results.
A parallel build that keeps every record findable
The build runs on all cores with one lock per graph node, as in hnswlib. Recern's built-in reachability check caught a race in an early version: up to 25 of 20,000 records could not be reached by any search. Inserts no longer descend through nodes another thread is still linking, and a final pass re-links any record the check still finds.
Denser neighbor lists
Recern fills each node's bottom-layer list to its 32-link limit with the best candidates the selection heuristic set aside. Each step costs more distance computations, but the search finds more true neighbors. A sparser graph built 37% faster in our tests and was slower at recall above 0.99, so the denser graph stays.
How we measured
- Datasets
- SIFT1M (1,000,000 × 128, Euclidean) and GloVe-100 (1,183,514 × 100, cosine) from ANN-Benchmarks, scored against their published ground-truth neighbors.
- Queries
- The first 1,000 test queries, top 10 results, recall@10. sqlite-vec was measured on the first 100 because exact scans are slow.
- Timing
- One query at a time from Python 3.13.15, after an untimed warm-up pass over 100 queries. Queries per second is the inverse of mean latency.
- Index settings
- HNSW with m = 16 and ef_construction = 200 for Recern, faiss and LanceDB IVF_HNSW_SQ. Search settings swept: ef from 10 to 1,280; nprobes and refine_factor for IVF_PQ.
- Threads
- Recern and faiss search on one thread. LanceDB builds and searches with its default thread pool.
- Hardware and versions
- Apple M3 Max, 36 GB, macOS 27.0.1. Recern Vector 0.0.1, faiss-cpu 1.15.1, LanceDB 0.40.0, sqlite-vec 0.1.9.
Read these numbers with care
- One machine and one run per setting. Build times varied by 10–15% between runs.
- Queries run one at a time from Python, which suits in-process libraries. LanceDB is designed for other workloads, such as data larger than memory, object storage and batched queries, and about 1 ms of its per-query time here is fixed overhead.
- sqlite-vec's stable release has no approximate index, so it appears as the exact-search baseline.
- faiss is the faiss-cpu wheel from PyPI. A build compiled for this CPU may be faster.
- Recern Vector is a prototype. The whole database is held in memory, saving rewrites the file, and the file format may still change. It is separate from the Recern workspace in Private Alpha.