FrankChen021 opened a new pull request, #20390:
URL: https://github.com/apache/druid/pull/20390

   ### Description
   
   This PR improves native expression vector processing by reducing null-vector 
propagation and allocation overhead in common numeric expression chains.
   
   The changes:
   
   - return `null` null-vectors from scalar and SIMD unary/binary processor 
bases when the output batch contains no nulls;
   - record scalar null presence only when the existing null branch is taken, 
avoiding a loop-carried accumulation on every row;
   - reuse per-processor buffers for numeric long/double casts instead of 
allocating a converted array for every evaluation;
   - replace stream-based fallback numeric conversions with simple loops;
   - use dedicated bound-constant processors with direct 
add/subtract/multiply/divide loops for same-type long and double arithmetic 
when the JDK Vector API path is not selected;
   - add correctness tests and a JMH benchmark covering long chains, mixed 
long/double chains, constants, nullability, vector sizes 128–2048, and Vector 
API on/off.
   
   SIMD selection continues to take precedence when the JDK Vector API is 
enabled. There are no configuration or expression-semantics changes.
   
   ### Benchmark
   
   The benchmark compares this branch against exact base commit 
`b2b9baedca090ac70e7293a239904a6d99fa6254` using the same JMH harness and JVM. 
Environment: arm64, Temurin OpenJDK 25. Results below use vector size 512, five 
warmup iterations, eight measurement iterations, three forks, and 300 ms per 
iteration.
   
   Lower `ns/op` is better. Speedup is calculated as `baseline / optimized - 
1`. The null-vector shapes are: **Absent**, where bindings return `null`; **All 
false**, where bindings return a non-null `boolean[]` with no null rows; and 
**Sparse**, where every sixteenth row is null. Each pair of rows is the same 
expression and null-vector shape with the Vector API disabled and enabled.
   
   | Workload | Expression | Null vector | Vector API | Master ns/op | 
Optimized ns/op | Speedup |
   |---|---|---:|---:|---:|---:|---:|
   | Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Absent | Off | 8,873 | 
7,052 | **25.8%** |
   | Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Absent | On | 1,813 | 
1,737 | **4.4%** |
   |  |  |  |  |  |  |  |
   | Long chain | `((((l1 + l2) * l3) - l2) + l1)` | All false | Off | 9,797 | 
9,274 | **5.6%** |
   | Long chain | `((((l1 + l2) * l3) - l2) + l1)` | All false | On | 2,435 | 
2,213 | **10.0%** |
   |  |  |  |  |  |  |  |
   | Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Sparse | Off | 10,308 | 
10,414 | **-1.0%** |
   | Long chain | `((((l1 + l2) * l3) - l2) + l1)` | Sparse | On | 2,418 | 
2,366 | **2.2%** |
   |  |  |  |  |  |  |  |
   | Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Absent | Off | 1,336 | 948 | 
**41.0%** |
   | Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Absent | On | 1,033 | 935 | 
**10.5%** |
   |  |  |  |  |  |  |  |
   | Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | All false | Off | 6,495 | 
1,404 | **362.6%** |
   | Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | All false | On | 1,420 | 
1,336 | **6.3%** |
   |  |  |  |  |  |  |  |
   | Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Sparse | Off | 6,883 | 2,722 
| **152.9%** |
   | Mixed numeric | `(((l1 + d1) * 2.5) - l2)` | Sparse | On | 1,419 | 1,393 | 
**1.9%** |
   |  |  |  |  |  |  |  |
   | Constant-heavy | `(((l1 + 7) * 3) - 11)` | Absent | Off | 5,171 | 496 | 
**943.6%** |
   | Constant-heavy | `(((l1 + 7) * 3) - 11)` | Absent | On | 745 | 710 | 
**4.8%** |
   |  |  |  |  |  |  |  |
   | Constant-heavy | `(((l1 + 7) * 3) - 11)` | All false | Off | 5,938 | 621 | 
**856.3%** |
   | Constant-heavy | `(((l1 + 7) * 3) - 11)` | All false | On | 776 | 791 | 
**-2.0%** |
   |  |  |  |  |  |  |  |
   | Constant-heavy | `(((l1 + 7) * 3) - 11)` | Sparse | Off | 5,711 | 873 | 
**553.9%** |
   | Constant-heavy | `(((l1 + 7) * 3) - 11)` | Sparse | On | 768 | 765 | 
**0.4%** |
   
   #### Constant-bound arithmetic
   
   `LongBivariateLongsConstantProcessor` and 
`DoubleBivariateDoublesConstantProcessor` specialize same-type arithmetic with 
one bound constant. The general bivariate processor materializes the constant 
as a second input vector and invokes an operation function for every row. The 
specialized processors retain the constant as a scalar, choose the operation 
and operand order once per batch, and execute direct primitive arithmetic 
loops. The double specialization exists for the same reason as the long 
version: it removes the redundant constant vector and keeps the hot 
double-arithmetic loop monomorphic and readily inlineable by HotSpot.
   
   For the constant-heavy expression with absent null vectors, the optimized 
non-Vector-API path is faster than the Vector API path (496 versus 710 ns/op). 
The scalar specialization reads only the varying input array and applies its 
scalar constants directly. SIMD selection still takes precedence when the 
Vector API is enabled, so that path uses the general bivariate processors, 
materializes each constant operand as a vector, and loads both operand arrays 
for every operation. At a 512-row batch and for this short arithmetic chain, 
that extra materialization and load/store overhead outweighs SIMD throughput. 
This is specific to the constant-heavy case: the longer all-variable chain 
remains substantially faster with the Vector API enabled.
   
   Two of the 18 baseline-to-optimized comparisons are small regressions: the 
sparse scalar long chain at -1.0%, and the all-false SIMD constant chain at 
-2.0%. The direct constant processors are not present in the first case, while 
the second case's 99.9% confidence intervals overlap the baseline. These small 
results are sensitive to run-to-run JVM/code-layout variation; neither offsets 
the gains in the corresponding expression family or null-vector shape.
   
   Geometric-mean speedups at vector size 512:
   
   - long chains: **7.5%**;
   - mixed numeric chains: **64.4%**;
   - constant-heavy chains: **196.0%**;
   - all 18 configurations: **73.6%**.
   
   ### Testing
   
   ```text
   mvn -ntp test -pl processing \
     
-Dtest='org.apache.druid.math.expr.VectorExprResultConsistencyTest,org.apache.druid.math.expr.VectorExprResultConsistencyVectorApiTest'
 \
     -Dsurefire.failIfNoSpecifiedTests=false \
     -Pskip-static-checks -Dweb.console.skip=true -T1C
   ```
   
   Result: 69 tests run, 0 failures, 0 errors, 0 skipped.
   
   The benchmark module also packages successfully with the new JMH benchmark.
   
   #### Release note
   
   Native vectorized numeric expressions reduce intermediate null-vector and 
numeric-conversion allocation overhead, improving representative 
expression-chain performance without changing query behavior or configuration.
   
   <hr>
   
   ##### Key changed/added classes in this PR
   
   * `CastToDoubleVectorProcessor`
   * `CastToLongVectorProcessor`
   * `LongBivariateLongsConstantProcessor`
   * `DoubleBivariateDoublesConstantProcessor`
   * `SimdProcessorUtils`
   * `ExpressionVectorProcessorBenchmark`
   
   <hr>
   
   This PR has:
   
   - [x] been self-reviewed.
   - [x] a release note entry in the PR description.
   - [x] added Javadocs for the new processor classes.
   - [x] added unit tests covering the new code paths.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to