chengxis-mdb commented on PR #16394:
URL: https://github.com/apache/lucene/pull/16394#issuecomment-4972322766

   Benchmark results from luceneutil (wikimedium10m), run by my colleague 
Tianxiao Wei, including a new `ConstInDisj` task that issues pure-should 
disjunctions of constant-score term-set clauses (the shape described above):
   
   ```
                               TaskQPS baseline      StdDevQPS 
my_modified_version      StdDev                Pct diff p-value
                        ConstInDisj        8.28      (2.8%)       93.00     
(89.8%) 1023.9% ( 906% - 1148%) 0.000
   ```
   
   All other tasks are within noise (full run below). Note `ConstMSM2`, 
`OrHighHigh`, `AndHighHigh` etc. are unchanged — the gate keeps every 
non-disjunction-backed constant-score clause and all impact-scored clauses on 
their existing paths.
   
   <details><summary>Full luceneutil output</summary>
   
   ```
                               TaskQPS baseline      StdDevQPS 
my_modified_version      StdDev                Pct diff p-value
              HighTermDayOfYearSort      750.80      (6.2%)      725.13      
(4.2%)   -3.4% ( -13% -    7%) 0.042
                              range     6520.02      (5.3%)     6341.16      
(6.7%)   -2.7% ( -14% -    9%) 0.152
                           BM25MSM2      188.93      (4.7%)      184.19      
(9.4%)   -2.5% ( -15% -   12%) 0.288
                MedIntervalsOrdered     1235.81      (5.6%)     1206.70      
(5.6%)   -2.4% ( -12% -    9%) 0.184
                             Fuzzy1      268.96      (1.8%)      262.73      
(2.8%)   -2.3% (  -6% -    2%) 0.002
                            Prefix3     1639.66      (5.5%)     1603.03      
(3.9%)   -2.2% ( -11% -    7%) 0.138
              BrowseMonthSSDVFacets      204.59      (8.6%)      200.24      
(1.5%)   -2.1% ( -11% -    8%) 0.274
                        AndHighHigh     1162.49      (4.3%)     1138.39      
(4.8%)   -2.1% ( -10% -    7%) 0.151
                         OrHighHigh      994.12      (3.4%)      976.73      
(5.3%)   -1.7% ( -10% -    7%) 0.217
                    LowSloppyPhrase      756.00      (5.9%)      743.84      
(6.1%)   -1.6% ( -12% -   11%) 0.396
                   HighSloppyPhrase      443.93      (7.3%)      436.81      
(6.7%)   -1.6% ( -14% -   13%) 0.466
                        LowSpanNear      861.32      (4.4%)      847.66      
(4.3%)   -1.6% (  -9% -    7%) 0.253
                            Respell      186.44      (2.3%)      183.73      
(2.8%)   -1.5% (  -6% -    3%) 0.071
                         HighPhrase      473.57      (4.2%)      467.95      
(6.4%)   -1.2% ( -11% -    9%) 0.489
                        MedSpanNear      627.57      (3.9%)      620.20      
(6.0%)   -1.2% ( -10% -    9%) 0.462
                            MedTerm     3226.22      (3.7%)     3193.18      
(3.6%)   -1.0% (  -7% -    6%) 0.371
                          OrHighLow     1583.34      (2.9%)     1568.06      
(3.3%)   -1.0% (  -6% -    5%) 0.321
                LowIntervalsOrdered     1423.38      (3.1%)     1411.08      
(3.4%)   -0.9% (  -7% -    5%) 0.399
                    MedSloppyPhrase       97.97      (4.9%)       97.13      
(4.7%)   -0.9% (  -9% -    9%) 0.571
                           Wildcard      517.48      (2.8%)      513.67      
(3.8%)   -0.7% (  -7% -    6%) 0.484
                         AndHighMed     1582.62      (3.8%)     1572.13      
(5.7%)   -0.7% (  -9% -    9%) 0.664
                           PKLookup      467.78      (3.5%)      465.29      
(3.7%)   -0.5% (  -7% -    6%) 0.638
               BrowseDateSSDVFacets       37.58      (1.8%)       37.43      
(1.2%)   -0.4% (  -3% -    2%) 0.393
               HighIntervalsOrdered      269.55      (4.7%)      268.51      
(5.5%)   -0.4% ( -10% -   10%) 0.810
                          OrHighMed     1499.26      (2.5%)     1494.71      
(4.4%)   -0.3% (  -6% -    6%) 0.786
                             Fuzzy2      109.82      (1.8%)      109.51      
(3.0%)   -0.3% (  -4% -    4%) 0.711
                           HighTerm     2525.48      (4.3%)     2523.08      
(4.5%)   -0.1% (  -8% -    9%) 0.946
                          ConstMSM2     1746.68      (4.1%)     1745.22      
(5.9%)   -0.1% (  -9% -   10%) 0.959
              BrowseMonthTaxoFacets      138.20      (5.9%)      138.34      
(6.8%)    0.1% ( -11% -   13%) 0.961
                             IntNRQ     1982.74      (5.1%)     1984.90      
(4.5%)    0.1% (  -9% -   10%) 0.943
               BrowseDateTaxoFacets      160.59      (0.9%)      161.54      
(0.6%)    0.6% (   0% -    2%) 0.017
          BrowseDayOfYearTaxoFacets      160.14      (0.8%)      161.17      
(0.6%)    0.6% (   0% -    2%) 0.006
                            LowTerm     3684.93      (4.4%)     3708.81      
(6.5%)    0.6% (  -9% -   12%) 0.714
          BrowseDayOfYearSSDVFacets      178.71      (3.6%)      179.93      
(2.4%)    0.7% (  -5% -    6%) 0.487
        BrowseRandomLabelTaxoFacets      141.68      (2.7%)      142.71      
(2.0%)    0.7% (  -3% -    5%) 0.335
                       HighSpanNear      528.43     (11.0%)      533.07     
(11.5%)    0.9% ( -19% -   26%) 0.805
                          LowPhrase      939.27      (5.8%)      949.72      
(3.9%)    1.1% (  -8% -   11%) 0.474
                         AndHighLow     2840.67      (3.7%)     2878.15      
(4.3%)    1.3% (  -6% -    9%) 0.296
        BrowseRandomLabelSSDVFacets      127.32      (4.1%)      129.11      
(4.8%)    1.4% (  -7% -   10%) 0.322
                          MedPhrase     1046.02      (4.4%)     1060.90      
(4.0%)    1.4% (  -6% -   10%) 0.283
                  HighTermMonthSort     1434.77      (3.4%)     1455.43      
(4.7%)    1.4% (  -6% -    9%) 0.267
                        ConstInDisj        8.28      (2.8%)       93.00     
(89.8%) 1023.9% ( 906% - 1148%) 0.000
   ```
   
   </details>
   
   We also validated end-to-end on the production-like workload where we found 
the regression: top-k disjunctions of dense constant-score term-set clauses 
recover most of the slowdown introduced in 10.3 (avg latency 316ms → ~236ms 
across replicated runs), while equals/point-based and impact-scored shapes are 
unchanged. An earlier unconditional version of this patch (bulk path for every 
constant-score clause) measurably regressed single-postings and bit-set-backed 
clauses, which is what motivated the `DisjunctionDISIApproximation` gate.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to