orenovadia opened a new pull request, #3160:
URL: https://github.com/apache/jackrabbit-oak/pull/3160

   ## Summary
   
   Follow-up to #3143, and a complement to #3159: this covers the exact-size 
count case that #3159 explicitly left for a follow-up PR ("Counting inside 
mongot (`$searchMeta`) ... are left for follow-up PRs").
   
   | | #3159 | This PR |
   |---|---|---|
   | Exact-size count | `$count` appended to the pipeline with 
`$sort`/`$set`/`$project` dropped | When the pipeline is a bare `[$search, 
$project]` pair, the count comes from the search operator's `count: {type: 
"total"}` option, evaluated inside mongot and returned as a single 
search-metadata document |
   | With post-search `$match`/`$sort` | `$count` over the truncated pipeline | 
Same as #3159 - the metadata count cannot see those stages, so the connector 
keeps the `$count` fallback |
   
   Why the metadata count is cheaper: appending `$count` - even to a truncated 
pipeline - makes mongod re-execute the `$search` and stream every matching 
document from mongot into mongod just to count it. The `count` option is 
answered by mongot from the index without fetching or sorting any documents. 
This is the same native feature the internal Atlas Search anti-pattern catalog 
recommends against the documented "$search followed by $match, $sort, $group, 
$count" pattern.
   
   The two PRs compose: once #3159 drops the redundant `jcr:score desc` 
`$sort`, a plain score-sorted query becomes a bare `[$search, $project]` pair, 
which is exactly the shape where the metadata count path applies.
   
   Local benchmark (20k docs, `mongodb/mongodb-atlas-local:8.3.3`, medians of 
10 interleaved samples; controls match within a few percent across arms):
   
   | scenario | #3143 (cursor drain) | #3159 base (`$count`) | this PR 
(`$searchMeta`) |
   |---|---:|---:|---:|
   | page of 10, no size | 19.4ms | 18.6ms | 16.6ms |
   | page of 10 after exact size | 161.6ms | 114.1ms | **30.2ms** |
   
   The remaining ~10-13ms over the control is the index-presence guard plus the 
count aggregation in the community build; a cached guard (part of #3159's 
scope) should close most of that.
   
   ## Reviewer focus
   
   - `searchOnlyPipeline()` guards the metadata path to a bare `[$search, 
$project]` pair: any post-search stage (property-filter `$match`, the 
sort-guard `$match`, `$sort`) makes the metadata count inexact or wrong, and 
the `$count` fallback applies instead. When #3159's sort-drop and eventual 
filter pushdown land, more queries take the metadata path automatically.
   - The `$searchMeta` aggregation runs through the same `aggregateCursor()` 
retry/readiness handling as result queries, so an index that is starting, 
missing, or `FAILED` fails exactly as before.
   - The count is exact for `type: "total"` up to mongot's count threshold, 
beyond which mongot reports a lower bound; matching documents are streamed in 
the result cursor in that regime anyway, so Oak's fallbacks apply.
   
   ## Not changed
   
   Adaptive batch sizes, moving property filters into the `$search` operator's 
`filter` clause (which would extend the metadata count to filtered queries), 
and a pre-existing inaccuracy both PRs inherit: `getSize(EXACT)` for a query 
whose property restriction Oak evaluates itself counts index rows before Oak's 
filtering (31 instead of 3 in #3159's test data).
   
   ## Testing
   
   - 
`MongotCoreQueryCompatibilityTest.exactSizeOfABareSearchCountsInsideMongotWithoutACountStage`:
 for a bare contains-query, an exact size request increments `$searchMeta` in 
the server's `aggStageCounters` and leaves `$count` unchanged; it fails against 
#3143's `$count`-pipeline code.
   - Full local ladder: `oak-search-mongot` connector suite green (29 tests 
across `MongotResultSizeTest`, `MongotFullTextIndexTest`, 
`MongotOrderByCommonTest`, plus the 10 in `MongotCoreQueryCompatibilityTest`) 
against atlas-local.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to