meticulous-dft opened a new pull request, #3159:
URL: https://github.com/apache/jackrabbit-oak/pull/3159

   ## Summary
   
   Follow-up to #3143. This removes three kinds of per-query work that never 
change results:
   
   | | Before | After |
   |---|---|---|
   | Search index check | `listSearchIndexes` before every `$search` 
aggregation, which is one extra round trip per query and two when an exact size 
is requested | Runs only when a `$search` returns no hits, or while retrying an 
index that is not yet queryable |
   | Exact-size count | `$count` appended to the full result pipeline, 
including the score `$set`, the `$sort` and the `$project` | Those stages are 
dropped before `$count`. A `$sort` blocks on every hit, and none of them change 
the count. |
   | `order by [jcr:score] desc` | `$set` of the score, then a blocking `$sort` 
of every hit | No sort. `$search` already returns hits by descending score, and 
the following `$match` stages keep that order. |
   
   The index check can't simply be removed: `$search` against a missing index 
returns no hits rather than an error. So a query that returns nothing still 
checks the index and fails with the same "is unavailable" error. A query that 
returns hits has proven the index exists without the lookup:
   
   ```text
   aggregateCursor(pipeline)
     cursor = aggregate(pipeline)
     if pipeline starts with $search and cursor has no first hit
       requireSearchIndex()        # missing or FAILED index -> 
IllegalStateException
     on "index not queryable yet" error
       requireSearchIndex()        # keep failing fast for a missing or FAILED 
index
       retry until the readiness timeout
   ```
   
   **Not changed.** Counting inside mongot (`$searchMeta`), adaptive batch 
sizes, and moving filters into `$search` are left for follow-up PRs. While 
testing I also saw that `getSize(EXACT)` for a query whose property restriction 
Oak evaluates itself counts index rows before Oak's filtering: 31 instead of 3 
in the test data. That predates #3143, which drained the same rows, and this PR 
doesn't change it.
   
   **Reviewer focus**
   - `MongotIndex.aggregateCursor`: check that `cursor.hasNext()` makes no 
extra round trip when the first batch is exhausted. Check also that the retry 
path still fails fast when the index is missing or `FAILED`, which the removed 
up-front check used to guarantee.
   - `MongotIndex.pipeline`: the sort is dropped only when the whole sort is 
`jcr:score` descending. Ascending score and mixed sort keys keep `$sort`.
   
   ## Testing
   
   - Added `exactSizeOfAnIndexSortedSearchNeedsNoIndexLookupOrSecondSort` and a 
descending-score check in `pathAndScoreOrderingComposeWithMongotResults`. 
Results are identical either way, so both count stages in the server's 
`aggStageCounters`, and both fail against the previous code.
   - Added `droppedSearchIndexFailsTheQueryInsteadOfReturningNoResults`, 
because the index check now runs only when a query has no hits; it fails if 
that check is removed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to