Query aware cost function - #6823
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2813fb45e0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| // Build requests for each index id | ||
| let jobs: Vec<SearchJob> = split_metadatas.iter().map(SearchJob::from).collect(); | ||
| // query_complexity_factor is 1.0 here (no effect) because list_fields doesn't execute the query |
There was a problem hiding this comment.
this seems to disagree with comments in with_priority.rs
/// Default tasks have zero priority and cost so short metadata operations, such as
/// list-fields processing, run before split searches.
There was a problem hiding this comment.
the job_cost here is used to place the job and is not related to cpu priority in with_priority.rs
should we use this job_cost for cpu priority too? my idea was this job should be quick on cpu so it can be 0 cost.
There was a problem hiding this comment.
oh okay, it wasn't clear to me that we had multiple notions of cost. Before i think we has the same cost function for placing and prioritization
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f2b5db780c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| self.remaining_job_cost = self | ||
| .remaining_job_cost | ||
| .saturating_sub(permit_request.job_cost); |
There was a problem hiding this comment.
Fail instead of masking remaining-cost underflow
If the remaining-cost invariant is violated—for example after the unchecked sum in from_task_metadata overflows in a release build—this saturating subtraction logs the error but continues with a zero-cost request, incorrectly promoting corrupted work to the front of the permit queue. Use checked arithmetic and treat an underflow as an invariant failure rather than allowing scheduling to proceed with fabricated state.
AGENTS.md reference: AGENTS.md:L19-L22
Useful? React with 👍 / 👎.
| /// Default tasks have zero priority and cost so short metadata operations, such as | ||
| /// list-fields processing, run before split searches. | ||
| impl Default for Priority { | ||
| fn default() -> Self { | ||
| Priority::Normal(0) | ||
| Priority::Normal { | ||
| priority: 0, | ||
| job_cost: 0, | ||
| } |
There was a problem hiding this comment.
Avoid zero-cost priority for unbounded list-fields work
When a broad list-fields request spans many splits, get_and_process_fields_metadata can enqueue up to 500 deserialization/filtering tasks at once through run_cpu_intensive, and every one now receives this zero cost. Since normal queue ordering prefers lower cost, all of those tasks run before every split search with the same request priority and a positive cost, allowing one metadata request to monopolize the shared search CPU pool. Give these per-split metadata tasks a nonzero size-based cost or otherwise bound their precedence rather than making all default work cheapest.
Useful? React with 👍 / 👎.
| let shape_cost = visitor.total.max(1.0); | ||
| let aggregation_cost = agg_cost(search_request.aggregation_request.as_deref())?; | ||
| Ok(shape_cost + aggregation_cost) |
There was a problem hiding this comment.
Include the requested hit count in query cost
A request for one hit and a request for the maximum hit window receive the same complexity factor because only the AST and aggregations are considered here. Per-split top-k collection and merging scale with max_hits, and the leaf request further adds start_offset, so a request with max_hits = 10_000 and start_offset = 10_000 can maintain a 20,000-entry candidate set per split while being queued like a one-hit query. Incorporate the effective leaf hit count into the estimate so large result windows cannot bypass the new scheduling protection.
Useful? React with 👍 / 👎.
| } | ||
| self.remaining_job_cost = self | ||
| .remaining_job_cost | ||
| .saturating_sub(permit_request.job_cost); |
There was a problem hiding this comment.
can we have a warn to know when we saturate so this isn't silent
Description
Large queries take up a lot of resources in the cluster and can slow down basic queries. We want to protect queries from large "poison pills"
hit_cost = 5 + num_docs / 100_000(constant overhead to open split etc + linear cost depending on split size)query_complexity_factor = shape_cost + agg_costHow was this PR tested?
Unit tests
Poison query experiment
2026-08-19T19:57:39Z)Before
With Cost Function