Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 20 additions & 29 deletions DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -783,7 +783,7 @@ the distinct terms of the row's multiset and carry how many times each occurs,
which is what length and prefix filtering prune on. The comparator then verifies
each surviving candidate against the ordered sequences, running the banded
dynamic program under the budget the current `minSimilarity` allows. Both
`spars` and `spars_merge` follow that scoring path on `inverted_term`. They
`spars` and `spars_merge` follow that scoring path on the term index. They
differ only in how the lists are traversed: key-major candidate generation
(`spars`) versus row-major merge (`spars_merge`).

Expand Down Expand Up @@ -910,14 +910,15 @@ they inserted them under.
The three inverted index types build the same uni-sorted inverted lists but can
traverse them with either of two candidate generators, selected per namespace
with the `candidate_generator` index parameter. Both return identical results
and honor every metadata filtering strategy; they differ only in how much work
they do to get there.
and honor every metadata filtering strategy. Both work with every inverted index
type and every comparator; they differ only in how much work they do to get
there. Neither generator changes which comparators or index types a namespace
may configure.

`spars` is the default and is key-major. Its vertical scan visits the query's
keys cheapest first, each horizontal scan narrows that key's inverted list to
the rows length filtering admits, and every surviving candidate is scored with
the comparator. Because it always scores through the comparator, it supports
every inverted index type and every supported comparator.
the comparator.

`spars_merge` is row-major. One frontier spans all of the query's keys and
advances them in step, so every inverted-list entry belonging to a candidate row
Expand All @@ -932,30 +933,20 @@ currently holds. It trades a priority queue over the query's keys for the
ability to abandon a row before all of its shared keys arrive, which pays off
when a query has many keys and the minimum similarity rejects most rows early.

When the keys are the terms of a sparse record, the inverted lists also carry
the row's value at that key, so the accumulated conjunction is the row's exact
similarity and no further comparison is needed. Signature keys carry no usable
value, and a sequence's terms bound its similarity without determining it, so in
both cases the merge generator scores each surviving candidate through the
comparator. `inverted_hybrid` applies the generator
independently to each of the two, so its term index scores from the conjunction
while its signature index verifies.

Pruning by partial conjunction requires a comparator that implements
the conjunction hooks and sparse term keys whose lists carry a value at each
key. On `inverted_term` with sparse records, the accumulated conjunction can be
the exact similarity. Jaccard, Ruzicka, and L2 implement the hooks, so
`spars_merge` applies partial-conjunction bounds there. Inverted-list values are
materialized only where those bounds read them.

Sequence comparators implement prefix bounds for `spars` but not `ConjunctionScored`
hooks on ordered sequences. On `inverted_term`, `spars_merge`
still bounds rows with partial Ruzicka conjunction over the indexed term
multiset: the configured comparator supplies the minimum shared-key fraction,
and verification scores each surviving candidate through the edit distance.
Signature-keyed structures
and the signature half of `inverted_hybrid` still reject `spars_merge` with
sequence comparators, because that half requires a conjunction-scoring comparator.
When merge can treat the accumulated conjunction as the row's exact
similarity, it may skip a further comparison. That happens on term-keyed sparse
records under `candidates_and_verification`, where inverted lists carry the
value at each key. Otherwise merge bounds rows with a partial conjunction
appropriate to the comparator and key type, then verifies each surviving
candidate through the comparator. Signature keys always verify through the
comparator. Ordered sequences on the term index bound merge with partial
multiset conjunction and verify with the configured sequence comparator.
`inverted_hybrid` chooses the generator independently for its term index and
its signature index, so each half follows the rules of its keying.

Sequence comparators implement prefix bounds for `spars`. Inverted-list posting
values are materialized only where `spars_merge` partial-conjunction bounds read
them.

## Discarding Popular Terms

Expand Down
24 changes: 10 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -416,20 +416,16 @@ recall costs depends on how much similarity the discarded terms carried.

The inverted index types can find candidates two ways, chosen with the
`candidate_generator` index parameter. Both return identical results and honor
every metadata filtering strategy; they differ only in how much work they do.

`spars`, the default, visits the query's keys cheapest first and scores every
surviving candidate. It works with every inverted index type and every
comparator.

`spars_merge` advances all of the query's keys together, letting it abandon a
row as soon as no completion of it can reach the current minimum similarity. It
pays off when queries have many keys and the minimum similarity rejects most
rows early. It is available for `l2`, `jaccard`, and `ruzicka`, which implement
partial-conjunction bounds, and for `gld` and `ngld` on `inverted_term`, where
`spars_merge` bounds rows with partial Ruzicka conjunction over the indexed
multiset and verifies each surviving candidate with the configured sequence
comparator on the ordered terms.
every metadata filtering strategy. Both work with every inverted index type
and every comparator; they differ only in how much work they do to get there.

`spars`, the default, is key-major. It visits the query's keys cheapest first
and scores every surviving candidate through the comparator.

`spars_merge` is row-major. It advances all of the query's keys together and
abandons a row as soon as no completion of it can reach the current minimum
similarity. It pays off when queries have many keys and the minimum similarity
rejects most rows early.

## What USSI Does Not Do

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -24,10 +24,11 @@
* <p>Shared indexed keys bound an order-sensitive distance without determining it, so candidate
* generation verifies each surviving candidate against the ordered sequences.
*
* <p>{@code spars} generates candidates with prefix bounds from this comparator over those indexed
* keys and scores ordered sequences. {@code spars_merge} on {@code inverted_term} bounds rows with
* partial Ruzicka conjunction over the indexed multiset, then scores each surviving candidate
* through this comparator on the ordered sequences.
* <p>{@code spars} and {@code spars_merge} are both available on every inverted index type with
* every comparator. {@code spars} generates candidates with prefix bounds from this comparator
* over the indexed keys and scores ordered sequences. {@code spars_merge} bounds rows with partial
* multiset conjunction over those indexed keys, then scores each surviving candidate through this
* comparator on the ordered sequences.
*/
abstract class BaseSequenceComparator extends Comparator implements KeyShareBounded {

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -22,8 +22,8 @@
* Key-major candidate generation over uni-sorted inverted lists, implementing {@code spars}.
*
* <p>The query's keys are visited cheapest first, and each one's inverted list is narrowed to the
* rows that length filtering admits. Every candidate is then scored through the comparator, so this
* generator works for every inverted index type.
* rows that length filtering admits. Every candidate is then scored through the comparator. This
* generator works with every inverted index type and every comparator.
*
* <p>Public only for the sibling inverted index packages.
*/
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -20,11 +20,14 @@
import javax.annotation.Nullable;

/**
* Row-major merge candidate generation over uni-sorted inverted lists.
* Row-major merge candidate generation over uni-sorted inverted lists, implementing {@code
* spars_merge}.
*
* <p>One frontier spans the query's keys and advances them in step, so a candidate row's entries
* all arrive together. The merge accumulates the row's conjunction as it goes and abandons the row
* once no completion of it can reach minSimilarity.
* once no completion of it can reach minSimilarity. When partial conjunction does not determine
* similarity, each surviving candidate is scored through the comparator. This generator works with
* every inverted index type and every comparator.
*
* <p>Public only for the sibling inverted index packages.
*/
Expand Down
Loading