On Tue, 22 Sep 2026 14:43:52 GMT, Boris Ulasevich <[email protected]>
wrote:
>> A for-each loop over ArrayList.Itr carries two phases of one recurrence
>> through the loop: cursor (after the increment) and lastRet (before it). Both
>> are live at the same time, so the allocator puts a copying mov in the loop
>> body - on every iteration, for a value the loop never reads. This can be
>> dealt with at the Java level, see #32344, but this change is an attempt to
>> improve it at the level of C2 compilation.
>>
>> In a loop compiled by C2, lastRet is never written to memory: the iterator
>> is scalarized and the field lives in a register. In a plain for-each nobody
>> reads it (it is only used by set() and remove()) so its only consumers are
>> the Phi at the loop exit and the safepoint's debug info.
>>
>> C2 already knows how to get rid of a second index:
>> PhaseIdealLoop::replace_parallel_iv looks for a variable that walks
>> alongside the trip counter and rewrites its uses in terms of the counter.
>> It only recognizes PARALLEL shape:
>>
>>
>> PARALLEL: int a = 5; for (int iv = 0; iv < limit; iv++) { use(a); a
>> += 3; }
>> PREVIOUS: int prev = -1; for (int iv = 0; iv < limit; iv++) { use(prev);
>> prev = iv; }
>>
>>
>> The patch teaches the recognition step the second (PREVIOUS) shape: a phi in
>> the loop head whose backedge value is the trip counter plus a constant. Its
>> uses are then rewritten in terms of the trip counter, the phi becomes dead
>> code, and the copy in the loop body is no longer generated.
>>
>> With `-XX:LoopMaxUnroll=1` the resulting aarch64 assembly looks like the
>> following:
>>
>> @Benchmark
>> public void list_foreach(Blackhole bh) {
>> for (Object o : list) {
>> bh.consume(o);
>> }
>> }
>>
>> BEFORE AFTER
>> mov w14, w11 -
>> add x11, x10, w14, sxtw #2 add x12, x10, w16, sxtw #2
>> ldr w11, [x11, #0xc] ldr w12, [x12, #0xc]
>> lsl x11, x11, #3 add w16, w16, #0x1
>> add w11, w14, #0x1 lsl x12, x12, #3
>> cmp w11, w12 cmp w16, w14
>> b.lt #-0x18 b.lt #-0x14
>>
>>
>> Removing one instruction gives a speedup of up to 40% with
>> `-XX:LoopMaxUnroll=1`, while with default VM options the picture is mixed: a
>> 11–15% gain on some CPUs and nothing measurable on others.
>>
>> ---------
>> - [x] I confirm that I make this contribution in accordance with the
>> [OpenJDK Interim AI Policy](https://openjdk.org/legal/ai).
>
> Boris Ulasevich has updated the pull request incrementally with two
> additional commits since the last revision:
>
> - add a test case where prev=i+c
> - 1. splitting replace_parallel_iv: now calls replace_lagging_index and
> replace_independent_index
> 2. use a reference for is_constant_difference offset parameter
Really good, the change looks good to me otherwise, thanks a lot for working on
this.
src/hotspot/share/opto/loopnode.cpp line 4640:
> 4638: // long iv2 = ((long) iv * stride_con2 / stride_con) + (init2 -
> ((long) init * stride_con2 / stride_con))
> 4639: //
> 4640: void PhaseIdealLoop::replace_parallel_iv(IdealLoopTree *loop) {
You could move this comment to `PhaseIdealLoop::replace_independent_index`
src/hotspot/share/opto/loopnode.cpp line 4718:
> 4716: }
> 4717:
> 4718: replace_with_affine_index(loop, phi2, ratio_con, stride_con2_bt);
indentation
src/hotspot/share/opto/loopnode.cpp line 4727:
> 4725: Node* phi = cl->phi();
> 4726:
> 4727: Node* init2 = phi2->in(LoopNode::EntryControl);
indentation.
-------------
Marked as reviewed by qamai (Reviewer).
PR Review: https://git.openjdk.org/jdk/pull/32549#pullrequestreview-5279841217
PR Review Comment: https://git.openjdk.org/jdk/pull/32549#discussion_r4073088141
PR Review Comment: https://git.openjdk.org/jdk/pull/32549#discussion_r4073090351
PR Review Comment: https://git.openjdk.org/jdk/pull/32549#discussion_r4073103499