Adds support for vectorized dot product on aarch64 (`sdot` and `udot`) through 
the Vector API

Dedicated dot product instructions have been shown to provide up to 10x 
throughput for lucene ([results](https://github.com/apache/lucene/pull/13572)) 
compared to vectorized multiply and add. This PR adds two new methods to 
`ByteVector` to leverage the aarch dot product instructions `sdot` and `udot` 
respectively.
- `IntVector dot(Vector<Byte> v, Vector<Integer> acc)`
- `IntVector dotUnsigned(Vector<Byte> v, Vector<Integer> acc)`

Each int lane of the result holds the dot product of the corresponding group of 
four bytes from the two operands, added to the matching accumulator lane. For 
example:

a = [a1, a2, a3, a4, ..., ..., a13, a14, a15, a16]
b = [b1, b2, b3, b4, ..., ..., b13, b14, b15, b16]
acc = [acc1, ..., ..., acc4]

a.dot(b, acc) -> [
    acc1 + a1 * b1 + a2 * b2 + a3 * b3 + a4 * b4, 
    ..., 
    ...,
    acc4 + a13 * b13 + a14 * b14 + a15 * b15 + a16 * b16
]


The equivalent instructions on x86 are `VPDPBSSD` and `VPDPBUUD` which perform 
the same operations and match the proposed new functions however this PR only 
includes aarch64.

Graviton 2

Benchmark                              Mode  Cnt      Score     Error   Units
VectorDotBenchmark.dotMulAdd          thrpt   25   1497.953 ±   2.734  ops/ms
VectorDotBenchmark.dotScalar          thrpt   25   2412.719 ±   0.014  ops/ms
VectorDotBenchmark.dotUnsignedMulAdd  thrpt   25   1498.898 ±   1.442  ops/ms
VectorDotBenchmark.dotUnsignedScalar  thrpt   25   2402.966 ±   0.157  ops/ms
VectorDotBenchmark.dotUnsignedVector  thrpt   25  23344.116 ± 222.150  ops/ms
VectorDotBenchmark.dotVector          thrpt   25  23243.016 ± 141.708  ops/ms


Graviton 3

Benchmark                              Mode  Cnt      Score     Error   Units
VectorDotBenchmark.dotMulAdd          thrpt   25   3257.311 ±   5.308  ops/ms
VectorDotBenchmark.dotScalar          thrpt   25   7946.950 ±   2.564  ops/ms
VectorDotBenchmark.dotUnsignedMulAdd  thrpt   25   3259.853 ±   3.511  ops/ms
VectorDotBenchmark.dotUnsignedScalar  thrpt   25   2536.268 ±   0.438  ops/ms
VectorDotBenchmark.dotUnsignedVector  thrpt   25  41263.821 ± 478.856  ops/ms
VectorDotBenchmark.dotVector          thrpt   25  41430.996 ± 688.936  ops/ms


---------
- [x] I confirm that I make this contribution in accordance with the [OpenJDK 
Interim AI Policy](https://openjdk.org/legal/ai).

-------------

Commit messages:
 - Clean up
 - Add tests plus minor fixes
 - Fix rebase
 - Implement Vector API dot product

Changes: https://git.openjdk.org/jdk/pull/32359/files
  Webrev: https://webrevs.openjdk.org/?repo=jdk&pr=32359&range=00
  Issue: https://bugs.openjdk.org/browse/JDK-8377386
  Stats: 1106 lines in 27 files changed: 1077 ins; 0 del; 29 mod
  Patch: https://git.openjdk.org/jdk/pull/32359.diff
  Fetch: git fetch https://git.openjdk.org/jdk.git pull/32359/head:pull/32359

PR: https://git.openjdk.org/jdk/pull/32359

Reply via email to