meticulous-dft opened a new pull request, #3157:
URL: https://github.com/apache/jackrabbit-oak/pull/3157

   ## Summary
   
   Follow-up to #3143 from the MongoDB side, for the AEM search evaluation. 
#3143 cut the round trips spent reading results. The next per-query cost is on 
mongod: for every `$search` hit, mongod loads the full index document before 
running the path, type, property, sort and facet stages, and then keeps only 
`_path` and the score. Those documents carry the analyzed and full-text fields, 
including extracted binary text, so the cost grows with both the number of hits 
and the size of the text.
   
   This PR adds an opt-in `storedSource=true` index property. When set, mongot 
stores just the fields the pipeline reads after `$search`, and queries return 
those fields instead of full documents:
   
   ```diff
    $search (mongot)                         finds the hits
   -  mongod loads each hit's full document    includes _fulltext / extracted 
text
   +  mongot returns the stored fields         returnStoredSource: true
    $match  path / type / property filters   unchanged, now reads stored fields
    $sort, $project _path, _score, highlights, facet
   ```
   
   | Piece | Change | Why |
   |---|---|---|
   | Search index definition | `storedSource.include` lists every field read 
after `$search` (`MongoFieldNames.POST_SEARCH_FIELDS`), emitted sorted | Mongot 
reports the list back sorted. `MongotSearchIndexManager` compares lists in 
order, so an unsorted list would resubmit the definition, and trigger a 
rebuild, on every indexing cycle. |
   | `_fulltext` mapping | `store: false`, with the index's analyzer and search 
analyzer, unless a property sets `useInExcerpt` | Mongot reads a hit's stored 
fields together. With the large full-text field stored, stored-source queries 
were slower than document lookups for large documents (see measurements below). 
|
   | Queries | Regular `$search` sets `returnStoredSource`. Suggest and 
spellcheck don't, because they project fields that aren't stored. `_fulltext` 
is not highlighted when it isn't stored. | Mongot rejects a highlight request 
for a field indexed with `store: false`. |
   | Activation | Read from the stored `:index-definition`, so the flag takes 
effect only after a reindex | Mongot rejects `returnStoredSource` against a 
search index built without `storedSource`. A reindex builds a new collection 
generation and switches to it only once that index is ready. |
   
   **Measurements.** These are pipeline-level numbers, not AEM Cloud Service 
latency. They come from mongosh inside the `mongodb/mongodb-atlas-local:8.2.6` 
image on one machine: no network, all data in memory. The pipeline is the 
connector's shape: `$search` text on `_fulltext`, a `$match` on `_ancestors` 
and `_primaryType`, then `$project`. There are 2,000 hits, the batch size is 
1,000, and each figure is p50 / p95 ms over 100 runs after 10 warm-up runs. 
"Stored copy only" and "`store: false` only" isolate the two parts of this 
change.
   
   | Extracted text per document | Today | Stored copy only | `store: false` 
only | This PR |
   |---|---|---|---|---|
   | 1 KB | 18 / 26 | 10 / 16 | 15 / 18 | 8 / 12 |
   | 20 KB | 32 / 37 | 30 / 33 | 16 / 20 | 8 / 13 |
   | 64 KB | 30 / 34 | 74 / 82 | 18 / 20 | 8 / 15 |
   
   **Trade-off.** With `storedSource=true` and no `useInExcerpt` property, 
`rep:excerpt(.)` returns no excerpt, and the query no longer asks mongot for a 
highlight that would fail. This matches Lucene for property text. Lucene does 
still excerpt extracted binary text, which it always stores. When some property 
sets `useInExcerpt`, `_fulltext` stays stored, excerpts keep working, and the 
gain for large documents shrinks.
   
   **Out of scope.**
   - The property stays off by default. Existing definitions produce the same 
search index definition and are not rebuilt.
   - Moving filters, sorting, and counts or facets into mongot is left for 
follow-up PRs.
   - These numbers still need confirming on the Atlas test cluster before the 
flag is enabled there.
   
   **Reviewer focus**
   - `MongoFieldNames.POST_SEARCH_FIELDS`: this must list every field a 
post-`$search` stage reads. A missing field silently drops results instead of 
failing, which is what the stored-source compatibility suites guard against.
   - `MongotIndexDefinition`: check that `storedSource` is read from the stored 
definition, not the live one. `hasExcerptProperties()` covers both named and 
regex property definitions.
   - `MongotSearchIndexDefinitionBuilder`: the explicit `_fulltext` mapping 
must analyze exactly as the dynamic mapping did.
   
   ## Testing
   
   - Added stored-source subclasses of the core and advanced query 
compatibility tests that re-run every inherited contract against a 
stored-source index, because a field missing from the stored list only shows up 
as missing results (dropping `facet` from the list fails the inherited facet 
test).
   - Added `indexingCycleKeepsTheSearchIndexQueryable` to the core 
compatibility test, which runs in both modes, because resubmitting an unchanged 
definition makes mongot rebuild the index and stop serving it; it failed with 
the index `PENDING` before the include list was sorted.
   - Manually ran the full module suite with `storedSource` forced on (only the 
two unit tests that assert the default is off failed), and measured the 
pipeline as described above.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to