Recern Vector · Benchmark report · October 2026

Faster than faiss HNSW at equal recall.

Recern Vector is an embedded vector database we are exploring alongside the Recern workspace: one file, no server, and every search can explain how it ran. We measured the Phase 1 prototype on the two standard ANN-Benchmarks datasets against faiss, LanceDB and sqlite-vec.

Source code and the benchmark harness: github.com/recerndata/recern-vector (bench/)

1.30×

faiss HNSW throughput at recall ≥ 0.95 on SIFT1M

1.78×

the same comparison on GloVe-100

35s

to index 1,000,000 SIFT vectors on 14 cores

0.11ms

median query latency at recall 0.983 on SIFT1M

Recall and speed

Queries per second at each recall level

Every mark is one search setting. Higher and further right is better: more queries per second at higher recall. Lines join the settings that no other setting of the same engine beats on both axes.

  • Recern Vector HNSW
  • faiss HNSWFlat
  • LanceDB IVF_HNSW_SQ
  • LanceDB IVF_PQ, tuned
  • sqlite-vec exact scan
0.50.80.90.950.990.999101001k10kRecall@10, stretched toward 1.0 →Queries per second (log) →faissRecern Vector
Recall@10 on the first 1,000 test queries; throughput of one query at a time from Python, on a log scale. The recall axis is stretched toward 1.0 so the high-recall region is readable; exact search sits at the right edge. Settings below recall 0.5 are not shown. The same numbers are in the tables below.
Throughput at a recall target

Fastest setting that reaches each recall level

SIFT1M · 1M × 128, Euclidean

Enginerecall ≥ 0.90recall ≥ 0.95recall ≥ 0.99
Recern VectorHNSW15,9730.06 ms9,0410.11 ms4,4090.22 ms
faissHNSWFlat13,0120.08 ms6,9610.14 ms3,5710.28 ms
LanceDBIVF_HNSW_SQ7651.28 ms7311.34 msnot reached
LanceDBIVF_PQ, tuned2454.03 ms2454.03 ms2454.03 ms
sqlite-vecexact scan1283.78 ms1283.78 ms1283.78 ms

GloVe-100 · 1.18M × 100, cosine

Enginerecall ≥ 0.90recall ≥ 0.95recall ≥ 0.99
Recern VectorHNSW2,3730.43 ms6351.59 msnot reached
faissHNSWFlat1,6870.60 ms3562.83 msnot reached
LanceDBIVF_HNSW_SQnot reachednot reachednot reached
LanceDBIVF_PQ, tuned2713.68 ms1675.93 ms8112.14 ms
sqlite-vecexact scan7140.43 ms7140.43 ms7140.43 ms

Queries per second, with median latency beneath. LanceDB's default IVF_PQ settings reach at most recall 0.36 on GloVe-100; the tuned configuration uses num_sub_vectors = dim / 4.

Index build and size

Building a million-vector index takes under a minute

Build time, HNSW m = 16, ef_construction = 200

DatasetRecern, 1 threadRecern, 14 threadsfaiss, 1 threadfaiss, 14 threadsLanceDB IVF_HNSW_SQ
SIFT1M345 s36 s284 s32 s28 s
GloVe-100481 s47 s416 s43 s44 s

Size on disk

EngineSIFT1MGloVe-100
Recern VectorHNSW631 MB620 MB
faissHNSWFlat626 MB614 MB
LanceDBIVF_HNSW_SQ789 MB769 MB
LanceDBIVF_PQ, tuned524 MB487 MB
sqlite-vecexact scan511 MB479 MB

Recern's build scales about ten times on 14 cores and stays within 1.1–1.2× of faiss on the same hardware. LanceDB uses all cores; its HNSW index is partitioned and scalar-quantized, which speeds up the build and caps recall below 0.99.

Engineering notes

What made the difference

119 → 60 ns per distance

Prefetching neighbor vectors

A graph search spends most of its time waiting for vectors to arrive from memory. Recern now gathers a node's unvisited neighbors, asks the CPU to start loading every cache line of their vectors, and only then computes distances, so the loads overlap. Median latency at ef = 80 on SIFT1M fell from 206 µs to 110 µs with identical results.

0 unreachable records

A parallel build that keeps every record findable

The build runs on all cores with one lock per graph node, as in hnswlib. Recern's built-in reachability check caught a race in an early version: up to 25 of 20,000 records could not be reached by any search. Inserts no longer descend through nodes another thread is still linking, and a final pass re-links any record the check still finds.

Higher recall at the same ef

Denser neighbor lists

Recern fills each node's bottom-layer list to its 32-link limit with the best candidates the selection heuristic set aside. Each step costs more distance computations, but the search finds more true neighbors. A sparser graph built 37% faster in our tests and was slower at recall above 0.99, so the denser graph stays.

Method

How we measured

Datasets
SIFT1M (1,000,000 × 128, Euclidean) and GloVe-100 (1,183,514 × 100, cosine) from ANN-Benchmarks, scored against their published ground-truth neighbors.
Queries
The first 1,000 test queries, top 10 results, recall@10. sqlite-vec was measured on the first 100 because exact scans are slow.
Timing
One query at a time from Python 3.13.15, after an untimed warm-up pass over 100 queries. Queries per second is the inverse of mean latency.
Index settings
HNSW with m = 16 and ef_construction = 200 for Recern, faiss and LanceDB IVF_HNSW_SQ. Search settings swept: ef from 10 to 1,280; nprobes and refine_factor for IVF_PQ.
Threads
Recern and faiss search on one thread. LanceDB builds and searches with its default thread pool.
Hardware and versions
Apple M3 Max, 36 GB, macOS 27.0.1. Recern Vector 0.0.1, faiss-cpu 1.15.1, LanceDB 0.40.0, sqlite-vec 0.1.9.
Limitations

Read these numbers with care

  • One machine and one run per setting. Build times varied by 10–15% between runs.
  • Queries run one at a time from Python, which suits in-process libraries. LanceDB is designed for other workloads, such as data larger than memory, object storage and batched queries, and about 1 ms of its per-query time here is fixed overhead.
  • sqlite-vec's stable release has no approximate index, so it appears as the exact-search baseline.
  • faiss is the faiss-cpu wheel from PyPI. A build compiled for this CPU may be faster.
  • Recern Vector is a prototype. The whole database is held in memory, saving rewrites the file, and the file format may still change. It is separate from the Recern workspace in Private Alpha.