Filters
A filter restricts a search to records whose metadata matches. Filters use a MongoDB-style JSON syntax in Python and on the command line, and can also be built in code in Rust.
collection.search(query, k=10, filter={
"lang": "en", # equality
"year": {"$gte": 2020, "$lt": 2025}, # range
"source.kind": {"$in": ["docs", "blog"]}, # one of several values, dotted path
})
Syntax
- A filter is a JSON object. Top-level keys are combined with AND. There is no OR or NOT yet; use
$infor "one of". - A key is a field name, or a dotted path into nested objects:
"source.kind"matches{"source": {"kind": "docs"}}. - A plain value means equality:
{"lang": "en"}is the same as{"lang": {"$eq": "en"}}. - An object whose keys all start with
$is a set of operators on that field, combined with AND.
| Operator | Matches when the field… | Argument |
|---|---|---|
$eq |
equals the value | any JSON value |
$in |
equals one of the values | array |
$gt, $gte |
is greater than (or equal to) the bound | number |
$lt, $lte |
is less than (or equal to) the bound | number |
Any other operator is an error.
Matching rules
- A record without the field never matches, including records without metadata.
- Numbers compare by value:
1equals1.0. - Ranges apply to numbers only. A string field never matches
$gtand friends; store dates as numbers (for example, a Unix timestamp or20240131). - Equality compares whole values. For
{"tags": ["a", "b"]}, the filter{"tags": "a"}does not match (unlike MongoDB);{"tags": ["a", "b"]}does. To filter by membership, store one record field per tag ({"tag_a": true}) or keep a single tag per record.
How filtered searches run
Before searching, Recern Vector estimates the filter's selectivity: the share of records that match, measured on a random sample of up to 512 live records.
- Selective filters (under about 2% of records) are answered by scanning the matching records exactly (
strategy="filtered_exact"). With few matches this is both faster than walking the graph and exact. - Other filters are applied while walking the HNSW graph (
strategy="hnsw"): the search visits nodes as usual and keeps only matching records, until it haskof them.
explain() shows which strategy ran, the estimated selectivity and how many nodes were visited:
report = collection.explain(query, k=10, filter={"tenant": "t7"})
print(report.strategy, report.filter_selectivity, report.visited)
# filtered_exact 0.006 100
A filtered search returns k results whenever at least k records match. With the graph strategy, recall can be lower than without a filter at the same ef, because fewer of the candidates the search looks at are eligible. In our tests with filters matching 4–10% of records, recall@10 was about 0.8 at ef=10 and 0.99 at ef=64. If filtered recall matters, compare against exact search on your data:
truth = {h.id for h in collection.search(query, k=10, exact=True, filter=f)}
found = {h.id for h in collection.search(query, k=10, ef=64, filter=f)}
print(len(truth & found) / 10)
estimate_recall() measures unfiltered searches only.
Rust
use recern_vector::Filter;
use serde_json::json;
// Parse the JSON syntax...
let f = Filter::from_json(&json!({"lang": "en", "year": {"$gte": 2020}}))?;
// ...or build it in code.
let f = Filter::And(vec![
Filter::eq("lang", "en"),
Filter::between("year", Some(2020.0), None), // inclusive bounds
Filter::is_in("source.kind", [json!("docs"), json!("blog")]),
]);
Filter::Range { field, gt, gte, lt, lte } gives exclusive bounds as well.