The GitHub Actions job "Build" on jackrabbit-oak.git/issue/OAK-12360 has failed.
Run started by GitHub user bhabegger (triggered by bhabegger).

Head commit for run:
74460fb7d643345550d7537cd855e7b0f561845c / Benjamin Habegger 
<[email protected]>
OAK-12360: Support per-property analyzers in Lucene and Elasticsearch index 
definitions

Today, an oak:index definition can declare a custom analyzer only under
analyzers/default. Any other analyzer configured under analyzers/<name> is
silently ignored - there's no way to apply a different tokenizer/stemmer/
stopword-set to individual properties (e.g. a language-specific analyzer
for one field while keeping the default for the rest of the index).

This adds a new optional analyzer property on a property definition:

    indexRules/<nodeType>/properties/<propName>/analyzer = "<name>"

referencing a sibling node analyzers/<name>, for both the Lucene and
Elasticsearch providers. Properties that don't set analyzer are completely
unaffected - fully backward compatible, no feature toggle needed since the
change is purely additive/opt-in.

Lucene (oak-search + oak-lucene):
- PropertyDefinition gains an analyzerName field, shared config read by
  both providers.
- LuceneIndexDefinition.createAnalyzer() builds a PerFieldAnalyzerWrapper
  entry for every analyzed property with a resolving analyzer reference,
  keyed by that property's actual Lucene field name (full:<pname> for
  index format V2+, <pname> for legacy V1) - the same name
  LuceneDocumentMaker writes documents under, so index-time and query-time
  stay consistent automatically.
- A dangling reference or a regexp property definition logs a warning and
  falls back to the default analyzer, never failing the index build.
- The aggregated :fulltext field (used by CONTAINS(*, ...)) always uses
  the single default analyzer regardless of per-property settings -
  documented as a known limitation, since every nodeScopeIndex property's
  raw text is funneled into one shared, single-analyzer field.

Elasticsearch (oak-search-elastic):
- ElasticCustomAnalyzer.buildCustomAnalyzers registers every named child
  under analyzers/*, not just default, with internal tokenizer/filter/
  char-filter keys namespaced per analyzer to avoid collisions between
  composed analyzers sharing one IndexSettingsAnalysis.Builder.
- ElasticIndexHelper.mapIndexRules picks the right per-property analyzer
  instead of always hardcoding "oak_analyzer", with the same warn+fallback
  treatment as Lucene for dangling references and regexp properties.
- Unlike Lucene, when a property is the sole analyzed nodeScopeIndex
  contributor, its own analyzer legitimately governs the aggregate
  jcr:contains(., ...) query too - Elasticsearch expands aggregate queries
  at query time to include each such property's own field, rather than
  persisting a copy into a fixed-analyzer aggregate field the way Lucene
  does. This is an intentional, documented difference from Lucene, not a
  limitation: the query term is analyzed the same way the content was,
  which is more correct in that case, not less.

Documentation for both providers (oak-doc/.../lucene.md and elastic.md) is
updated accordingly.

Report URL: https://github.com/apache/jackrabbit-oak/actions/runs/34366352594

With regards,
GitHub Actions via GitBox

Reply via email to