meticulous-dft opened a new pull request, #3157:
URL: https://github.com/apache/jackrabbit-oak/pull/3157
## Summary
Follow-up to #3143 from the MongoDB side, for the AEM search evaluation.
#3143 cut the round trips spent reading results. The next per-query cost is on
mongod: for every `$search` hit, mongod loads the full index document before
running the path, type, property, sort and facet stages, and then keeps only
`_path` and the score. Those documents carry the analyzed and full-text fields,
including extracted binary text, so the cost grows with both the number of hits
and the size of the text.
This PR adds an opt-in `storedSource=true` index property. When set, mongot
stores just the fields the pipeline reads after `$search`, and queries return
those fields instead of full documents:
```diff
$search (mongot) finds the hits
- mongod loads each hit's full document includes _fulltext / extracted
text
+ mongot returns the stored fields returnStoredSource: true
$match path / type / property filters unchanged, now reads stored fields
$sort, $project _path, _score, highlights, facet
```
| Piece | Change | Why |
|---|---|---|
| Search index definition | `storedSource.include` lists every field read
after `$search` (`MongoFieldNames.POST_SEARCH_FIELDS`), emitted sorted | Mongot
reports the list back sorted. `MongotSearchIndexManager` compares lists in
order, so an unsorted list would resubmit the definition, and trigger a
rebuild, on every indexing cycle. |
| `_fulltext` mapping | `store: false`, with the index's analyzer and search
analyzer, unless a property sets `useInExcerpt` | Mongot reads a hit's stored
fields together. With the large full-text field stored, stored-source queries
were slower than document lookups for large documents (see measurements below).
|
| Queries | Regular `$search` sets `returnStoredSource`. Suggest and
spellcheck don't, because they project fields that aren't stored. `_fulltext`
is not highlighted when it isn't stored. | Mongot rejects a highlight request
for a field indexed with `store: false`. |
| Activation | Read from the stored `:index-definition`, so the flag takes
effect only after a reindex | Mongot rejects `returnStoredSource` against a
search index built without `storedSource`. A reindex builds a new collection
generation and switches to it only once that index is ready. |
**Measurements.** These are pipeline-level numbers, not AEM Cloud Service
latency. They come from mongosh inside the `mongodb/mongodb-atlas-local:8.2.6`
image on one machine: no network, all data in memory. The pipeline is the
connector's shape: `$search` text on `_fulltext`, a `$match` on `_ancestors`
and `_primaryType`, then `$project`. There are 2,000 hits, the batch size is
1,000, and each figure is p50 / p95 ms over 100 runs after 10 warm-up runs.
"Stored copy only" and "`store: false` only" isolate the two parts of this
change.
| Extracted text per document | Today | Stored copy only | `store: false`
only | This PR |
|---|---|---|---|---|
| 1 KB | 18 / 26 | 10 / 16 | 15 / 18 | 8 / 12 |
| 20 KB | 32 / 37 | 30 / 33 | 16 / 20 | 8 / 13 |
| 64 KB | 30 / 34 | 74 / 82 | 18 / 20 | 8 / 15 |
**Trade-off.** With `storedSource=true` and no `useInExcerpt` property,
`rep:excerpt(.)` returns no excerpt, and the query no longer asks mongot for a
highlight that would fail. This matches Lucene for property text. Lucene does
still excerpt extracted binary text, which it always stores. When some property
sets `useInExcerpt`, `_fulltext` stays stored, excerpts keep working, and the
gain for large documents shrinks.
**Out of scope.**
- The property stays off by default. Existing definitions produce the same
search index definition and are not rebuilt.
- Moving filters, sorting, and counts or facets into mongot is left for
follow-up PRs.
- These numbers still need confirming on the Atlas test cluster before the
flag is enabled there.
**Reviewer focus**
- `MongoFieldNames.POST_SEARCH_FIELDS`: this must list every field a
post-`$search` stage reads. A missing field silently drops results instead of
failing, which is what the stored-source compatibility suites guard against.
- `MongotIndexDefinition`: check that `storedSource` is read from the stored
definition, not the live one. `hasExcerptProperties()` covers both named and
regex property definitions.
- `MongotSearchIndexDefinitionBuilder`: the explicit `_fulltext` mapping
must analyze exactly as the dynamic mapping did.
## Testing
- Added stored-source subclasses of the core and advanced query
compatibility tests that re-run every inherited contract against a
stored-source index, because a field missing from the stored list only shows up
as missing results (dropping `facet` from the list fails the inherited facet
test).
- Added `indexingCycleKeepsTheSearchIndexQueryable` to the core
compatibility test, which runs in both modes, because resubmitting an unchanged
definition makes mongot rebuild the index and stop serving it; it failed with
the index `PENDING` before the include list was sorted.
- Manually ran the full module suite with `storedSource` forced on (only the
two unit tests that assert the default is off failed), and measured the
pipeline as described above.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]