From 423efdb7c8db1cddd598b62f362c158947658994 Mon Sep 17 00:00:00 2001 From: Ahmed Metwally Date: Wed, 30 Sep 2026 21:28:57 -0700 Subject: [PATCH] Unify spars and spars_merge documentation README and DESIGN now describe both candidate generators as available on every inverted index type with every comparator, differing only in traversal. Align generator and sequence-comparator class comments with that wording. --- DESIGN.md | 49 ++++++++----------- README.md | 24 ++++----- .../comparator/BaseSequenceComparator.java | 9 ++-- .../inverted/generator/FilteredSearch.java | 4 +- .../index/inverted/generator/MergeSearch.java | 7 ++- 5 files changed, 42 insertions(+), 51 deletions(-) diff --git a/DESIGN.md b/DESIGN.md index b765579..ce567d0 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -783,7 +783,7 @@ the distinct terms of the row's multiset and carry how many times each occurs, which is what length and prefix filtering prune on. The comparator then verifies each surviving candidate against the ordered sequences, running the banded dynamic program under the budget the current `minSimilarity` allows. Both -`spars` and `spars_merge` follow that scoring path on `inverted_term`. They +`spars` and `spars_merge` follow that scoring path on the term index. They differ only in how the lists are traversed: key-major candidate generation (`spars`) versus row-major merge (`spars_merge`). @@ -910,14 +910,15 @@ they inserted them under. The three inverted index types build the same uni-sorted inverted lists but can traverse them with either of two candidate generators, selected per namespace with the `candidate_generator` index parameter. Both return identical results -and honor every metadata filtering strategy; they differ only in how much work -they do to get there. +and honor every metadata filtering strategy. Both work with every inverted index +type and every comparator; they differ only in how much work they do to get +there. Neither generator changes which comparators or index types a namespace +may configure. `spars` is the default and is key-major. Its vertical scan visits the query's keys cheapest first, each horizontal scan narrows that key's inverted list to the rows length filtering admits, and every surviving candidate is scored with -the comparator. Because it always scores through the comparator, it supports -every inverted index type and every supported comparator. +the comparator. `spars_merge` is row-major. One frontier spans all of the query's keys and advances them in step, so every inverted-list entry belonging to a candidate row @@ -932,30 +933,20 @@ currently holds. It trades a priority queue over the query's keys for the ability to abandon a row before all of its shared keys arrive, which pays off when a query has many keys and the minimum similarity rejects most rows early. -When the keys are the terms of a sparse record, the inverted lists also carry -the row's value at that key, so the accumulated conjunction is the row's exact -similarity and no further comparison is needed. Signature keys carry no usable -value, and a sequence's terms bound its similarity without determining it, so in -both cases the merge generator scores each surviving candidate through the -comparator. `inverted_hybrid` applies the generator -independently to each of the two, so its term index scores from the conjunction -while its signature index verifies. - -Pruning by partial conjunction requires a comparator that implements -the conjunction hooks and sparse term keys whose lists carry a value at each -key. On `inverted_term` with sparse records, the accumulated conjunction can be -the exact similarity. Jaccard, Ruzicka, and L2 implement the hooks, so -`spars_merge` applies partial-conjunction bounds there. Inverted-list values are -materialized only where those bounds read them. - -Sequence comparators implement prefix bounds for `spars` but not `ConjunctionScored` -hooks on ordered sequences. On `inverted_term`, `spars_merge` -still bounds rows with partial Ruzicka conjunction over the indexed term -multiset: the configured comparator supplies the minimum shared-key fraction, -and verification scores each surviving candidate through the edit distance. -Signature-keyed structures -and the signature half of `inverted_hybrid` still reject `spars_merge` with -sequence comparators, because that half requires a conjunction-scoring comparator. +When merge can treat the accumulated conjunction as the row's exact +similarity, it may skip a further comparison. That happens on term-keyed sparse +records under `candidates_and_verification`, where inverted lists carry the +value at each key. Otherwise merge bounds rows with a partial conjunction +appropriate to the comparator and key type, then verifies each surviving +candidate through the comparator. Signature keys always verify through the +comparator. Ordered sequences on the term index bound merge with partial +multiset conjunction and verify with the configured sequence comparator. +`inverted_hybrid` chooses the generator independently for its term index and +its signature index, so each half follows the rules of its keying. + +Sequence comparators implement prefix bounds for `spars`. Inverted-list posting +values are materialized only where `spars_merge` partial-conjunction bounds read +them. ## Discarding Popular Terms diff --git a/README.md b/README.md index e5dc0fc..7aca915 100644 --- a/README.md +++ b/README.md @@ -416,20 +416,16 @@ recall costs depends on how much similarity the discarded terms carried. The inverted index types can find candidates two ways, chosen with the `candidate_generator` index parameter. Both return identical results and honor -every metadata filtering strategy; they differ only in how much work they do. - -`spars`, the default, visits the query's keys cheapest first and scores every -surviving candidate. It works with every inverted index type and every -comparator. - -`spars_merge` advances all of the query's keys together, letting it abandon a -row as soon as no completion of it can reach the current minimum similarity. It -pays off when queries have many keys and the minimum similarity rejects most -rows early. It is available for `l2`, `jaccard`, and `ruzicka`, which implement -partial-conjunction bounds, and for `gld` and `ngld` on `inverted_term`, where -`spars_merge` bounds rows with partial Ruzicka conjunction over the indexed -multiset and verifies each surviving candidate with the configured sequence -comparator on the ordered terms. +every metadata filtering strategy. Both work with every inverted index type +and every comparator; they differ only in how much work they do to get there. + +`spars`, the default, is key-major. It visits the query's keys cheapest first +and scores every surviving candidate through the comparator. + +`spars_merge` is row-major. It advances all of the query's keys together and +abandons a row as soon as no completion of it can reach the current minimum +similarity. It pays off when queries have many keys and the minimum similarity +rejects most rows early. ## What USSI Does Not Do diff --git a/src/main/java/com/uber/ussi/comparator/BaseSequenceComparator.java b/src/main/java/com/uber/ussi/comparator/BaseSequenceComparator.java index f5de1dc..1feaa46 100644 --- a/src/main/java/com/uber/ussi/comparator/BaseSequenceComparator.java +++ b/src/main/java/com/uber/ussi/comparator/BaseSequenceComparator.java @@ -24,10 +24,11 @@ *

Shared indexed keys bound an order-sensitive distance without determining it, so candidate * generation verifies each surviving candidate against the ordered sequences. * - *

{@code spars} generates candidates with prefix bounds from this comparator over those indexed - * keys and scores ordered sequences. {@code spars_merge} on {@code inverted_term} bounds rows with - * partial Ruzicka conjunction over the indexed multiset, then scores each surviving candidate - * through this comparator on the ordered sequences. + *

{@code spars} and {@code spars_merge} are both available on every inverted index type with + * every comparator. {@code spars} generates candidates with prefix bounds from this comparator + * over the indexed keys and scores ordered sequences. {@code spars_merge} bounds rows with partial + * multiset conjunction over those indexed keys, then scores each surviving candidate through this + * comparator on the ordered sequences. */ abstract class BaseSequenceComparator extends Comparator implements KeyShareBounded { diff --git a/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/FilteredSearch.java b/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/FilteredSearch.java index e4a47a7..ae0081e 100644 --- a/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/FilteredSearch.java +++ b/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/FilteredSearch.java @@ -22,8 +22,8 @@ * Key-major candidate generation over uni-sorted inverted lists, implementing {@code spars}. * *

The query's keys are visited cheapest first, and each one's inverted list is narrowed to the - * rows that length filtering admits. Every candidate is then scored through the comparator, so this - * generator works for every inverted index type. + * rows that length filtering admits. Every candidate is then scored through the comparator. This + * generator works with every inverted index type and every comparator. * *

Public only for the sibling inverted index packages. */ diff --git a/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/MergeSearch.java b/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/MergeSearch.java index 618a8b3..2d2d2e0 100644 --- a/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/MergeSearch.java +++ b/src/main/java/com/uber/ussi/searchablestructure/index/inverted/generator/MergeSearch.java @@ -20,11 +20,14 @@ import javax.annotation.Nullable; /** - * Row-major merge candidate generation over uni-sorted inverted lists. + * Row-major merge candidate generation over uni-sorted inverted lists, implementing {@code + * spars_merge}. * *

One frontier spans the query's keys and advances them in step, so a candidate row's entries * all arrive together. The merge accumulates the row's conjunction as it goes and abandons the row - * once no completion of it can reach minSimilarity. + * once no completion of it can reach minSimilarity. When partial conjunction does not determine + * similarity, each surviving candidate is scored through the comparator. This generator works with + * every inverted index type and every comparator. * *

Public only for the sibling inverted index packages. */