On Tue, 25 Aug 2026 06:01:40 GMT, Shawn Emery <[email protected]> wrote:
>> The existing arbitrary `conditionalSet()` benchmark is not necessarily >> indicative of susceptibility to a side-channel attack in production code due >> to a number reasons, including the use of non-constant input arguments and >> the dependencies of intermediate operations due to adding between successive >> `lookup()` calls when calculating point multiplication, as an example. As a >> result, the problem being addressed is a benchmark issue, as the current >> conditional set benchmark is nothing like what is utilized in EC production >> operations. This is why I lifted the existing `lookup()` logic as a form of >> measuring conditional set performance, which better emulates a production >> state. >> >> The above code uses flag arguments that are incorrect which would generate >> invalid bit masks, subsequently corrupting the associated limbs during >> conditional sets. Even with non-binary flags, because the pair-wise >> assignments use the same input, C2 could perform a bitwise OR of the two >> masks and use the resulting mask only once for the conditional assignment. >> As a result C2 may i) not need to reload registers for the second call, ii) >> not need to store the result to memory from the first call when the second >> call would overwrite the first, and iii) keep the intermediate limbs in >> registers. Correcting the flag arguments to binary would make it even >> easier for inlining as it could eliminate the the first call if the second >> call has a 1 flag. In aggregate, this likely contributes to the 18% drop in >> performance with intrinsics when using the above code, compared to a 52% >> performance gain when using intrinsics with the originally proposed >> solution. The goal is to try to reproduce the same shape that we see in production EC operations. > >> Yes, I think you are wrong. Callers of this routine (in the crypto code, not >> this test code) appear always to call it with a non-constant argument >> derived from input data. I don't think that is a contingent state of affairs >> but rather a consequence of what the crypto algorithms are doing. So, I >> don't see any danger of constant folding happening where it might actually >> matter for security. Perhaps @smemery can confirm whether that is the case >> (he did indicate that in an earlier comment). > > Yes, I confirmed this on a AArch64 (M3 Pro) system running a JVM based on > source last updated on 7/20/26, while testing both fallback and intrinsics > enabled: > > 1) using `pointMultiply()` on the P-256 curve with intrinsics disabled: > 1.a) all six conditional-selections (3 coordinates * (high and low nibbles)) > were kept by C2 > 1.b) selection was still masked-based - no constant folding > 1.c) C2 did not generate branching or selection-dependent memory accesses > > 2) with integer polynomial intrinsics (`intpoly_assign`) enabled: > 2.a) lookup selection was converted to a binary mask without branching > 2.b) generated stub that branches on public limb length check - > `IntegerPolynomialP256` uses 10 and Montgomery P-256 uses five > 2.c) conditional set only uses the XOR/AND/XOR sequence with no branching or > memory access based on flags or limb values > that's a general problem that we face irrespective of what we do with regard > to this test i.e. it's a problem for a follow-up RFE to fix, not this one. OK, I'll buy that. ------------- PR Review Comment: https://git.openjdk.org/jdk/pull/32047#discussion_r3851221852
