Adds support for vectorized dot product on aarch64 (`sdot` and `udot`) through the Vector API
Dedicated dot product instructions have been shown to provide up to 10x throughput for lucene ([results](https://github.com/apache/lucene/pull/13572)) compared to vectorized multiply and add. This PR adds two new methods to `ByteVector` to leverage the aarch dot product instructions `sdot` and `udot` respectively. - `IntVector dot(Vector<Byte> v, Vector<Integer> acc)` - `IntVector dotUnsigned(Vector<Byte> v, Vector<Integer> acc)` Each int lane of the result holds the dot product of the corresponding group of four bytes from the two operands, added to the matching accumulator lane. For example: a = [a1, a2, a3, a4, ..., ..., a13, a14, a15, a16] b = [b1, b2, b3, b4, ..., ..., b13, b14, b15, b16] acc = [acc1, ..., ..., acc4] a.dot(b, acc) -> [ acc1 + a1 * b1 + a2 * b2 + a3 * b3 + a4 * b4, ..., ..., acc4 + a13 * b13 + a14 * b14 + a15 * b15 + a16 * b16 ] The equivalent instructions on x86 are `VPDPBSSD` and `VPDPBUUD` which perform the same operations and match the proposed new functions however this PR only includes aarch64. Graviton 2 Benchmark Mode Cnt Score Error Units VectorDotBenchmark.dotMulAdd thrpt 25 1497.953 ± 2.734 ops/ms VectorDotBenchmark.dotScalar thrpt 25 2412.719 ± 0.014 ops/ms VectorDotBenchmark.dotUnsignedMulAdd thrpt 25 1498.898 ± 1.442 ops/ms VectorDotBenchmark.dotUnsignedScalar thrpt 25 2402.966 ± 0.157 ops/ms VectorDotBenchmark.dotUnsignedVector thrpt 25 23344.116 ± 222.150 ops/ms VectorDotBenchmark.dotVector thrpt 25 23243.016 ± 141.708 ops/ms Graviton 3 Benchmark Mode Cnt Score Error Units VectorDotBenchmark.dotMulAdd thrpt 25 3257.311 ± 5.308 ops/ms VectorDotBenchmark.dotScalar thrpt 25 7946.950 ± 2.564 ops/ms VectorDotBenchmark.dotUnsignedMulAdd thrpt 25 3259.853 ± 3.511 ops/ms VectorDotBenchmark.dotUnsignedScalar thrpt 25 2536.268 ± 0.438 ops/ms VectorDotBenchmark.dotUnsignedVector thrpt 25 41263.821 ± 478.856 ops/ms VectorDotBenchmark.dotVector thrpt 25 41430.996 ± 688.936 ops/ms --------- - [x] I confirm that I make this contribution in accordance with the [OpenJDK Interim AI Policy](https://openjdk.org/legal/ai). ------------- Commit messages: - Clean up - Add tests plus minor fixes - Fix rebase - Implement Vector API dot product Changes: https://git.openjdk.org/jdk/pull/32359/files Webrev: https://webrevs.openjdk.org/?repo=jdk&pr=32359&range=00 Issue: https://bugs.openjdk.org/browse/JDK-8377386 Stats: 1106 lines in 27 files changed: 1077 ins; 0 del; 29 mod Patch: https://git.openjdk.org/jdk/pull/32359.diff Fetch: git fetch https://git.openjdk.org/jdk.git pull/32359/head:pull/32359 PR: https://git.openjdk.org/jdk/pull/32359
