meticulous-dft opened a new pull request, #3159:
URL: https://github.com/apache/jackrabbit-oak/pull/3159
## Summary
Follow-up to #3143. This removes three kinds of per-query work that never
change results:
| | Before | After |
|---|---|---|
| Search index check | `listSearchIndexes` before every `$search`
aggregation, which is one extra round trip per query and two when an exact size
is requested | Runs only when a `$search` returns no hits, or while retrying an
index that is not yet queryable |
| Exact-size count | `$count` appended to the full result pipeline,
including the score `$set`, the `$sort` and the `$project` | Those stages are
dropped before `$count`. A `$sort` blocks on every hit, and none of them change
the count. |
| `order by [jcr:score] desc` | `$set` of the score, then a blocking `$sort`
of every hit | No sort. `$search` already returns hits by descending score, and
the following `$match` stages keep that order. |
The index check can't simply be removed: `$search` against a missing index
returns no hits rather than an error. So a query that returns nothing still
checks the index and fails with the same "is unavailable" error. A query that
returns hits has proven the index exists without the lookup:
```text
aggregateCursor(pipeline)
cursor = aggregate(pipeline)
if pipeline starts with $search and cursor has no first hit
requireSearchIndex() # missing or FAILED index ->
IllegalStateException
on "index not queryable yet" error
requireSearchIndex() # keep failing fast for a missing or FAILED
index
retry until the readiness timeout
```
**Not changed.** Counting inside mongot (`$searchMeta`), adaptive batch
sizes, and moving filters into `$search` are left for follow-up PRs. While
testing I also saw that `getSize(EXACT)` for a query whose property restriction
Oak evaluates itself counts index rows before Oak's filtering: 31 instead of 3
in the test data. That predates #3143, which drained the same rows, and this PR
doesn't change it.
**Reviewer focus**
- `MongotIndex.aggregateCursor`: check that `cursor.hasNext()` makes no
extra round trip when the first batch is exhausted. Check also that the retry
path still fails fast when the index is missing or `FAILED`, which the removed
up-front check used to guarantee.
- `MongotIndex.pipeline`: the sort is dropped only when the whole sort is
`jcr:score` descending. Ascending score and mixed sort keys keep `$sort`.
## Testing
- Added `exactSizeOfAnIndexSortedSearchNeedsNoIndexLookupOrSecondSort` and a
descending-score check in `pathAndScoreOrderingComposeWithMongotResults`.
Results are identical either way, so both count stages in the server's
`aggStageCounters`, and both fail against the previous code.
- Added `droppedSearchIndexFailsTheQueryInsteadOfReturningNoResults`,
because the index check now runs only when a query has no hits; it fails if
that check is removed.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]