On Thu, 24 Sep 2026 16:02:31 GMT, Boris Ulasevich <[email protected]> 
wrote:

>> A for-each loop over ArrayList.Itr carries two phases of one recurrence 
>> through the loop: cursor (after the increment) and lastRet (before it). Both 
>> are live at the same time, so the allocator puts a copying mov in the loop 
>> body - on every iteration, for a value the loop never reads. This can be 
>> dealt with at the Java level, see #32344, but this change is an attempt to 
>> improve it at the level of C2 compilation.
>> 
>> In a loop compiled by C2, lastRet is never written to memory: the iterator 
>> is scalarized and the field lives in a register. In a plain for-each nobody 
>> reads it (it is only used by set() and remove()) so its only consumers are 
>> the Phi at the loop exit and the safepoint's debug info.
>> 
>> C2 already knows how to get rid of a second index: 
>> PhaseIdealLoop::replace_parallel_iv looks for a variable that walks 
>> alongside the trip counter and rewrites its uses in terms of the counter.
>> It only recognizes PARALLEL shape:
>> 
>> 
>> PARALLEL: int a = 5;     for (int iv = 0; iv < limit; iv++) { use(a);    a 
>> += 3;    }   
>> PREVIOUS: int prev = -1; for (int iv = 0; iv < limit; iv++) { use(prev); 
>> prev = iv; }
>> 
>> 
>> The patch teaches the recognition step the second (PREVIOUS) shape: a phi in 
>> the loop head whose backedge value is the trip counter plus a constant. Its 
>> uses are then rewritten in terms of the trip counter, the phi becomes dead 
>> code, and the copy in the loop body is no longer generated.
>> 
>> With `-XX:LoopMaxUnroll=1` the resulting aarch64 assembly looks like the 
>> following:
>> 
>> @Benchmark
>> public void list_foreach(Blackhole bh) {
>>     for (Object o : list) {
>>         bh.consume(o);
>>     }
>> }
>> 
>> BEFORE                               AFTER 
>>   mov  w14, w11                        -
>>   add  x11, x10, w14, sxtw #2          add  x12, x10, w16, sxtw #2
>>   ldr  w11, [x11, #0xc]                ldr  w12, [x12, #0xc]
>>   lsl  x11, x11, #3                    add  w16, w16, #0x1
>>   add  w11, w14, #0x1                  lsl  x12, x12, #3
>>   cmp  w11, w12                        cmp  w16, w14
>>   b.lt #-0x18                          b.lt #-0x14
>> 
>> 
>> Removing one instruction gives a speedup of up to 40% with 
>> `-XX:LoopMaxUnroll=1`, while with default VM options the picture is mixed: a 
>> 11–15% gain on some CPUs and nothing measurable on others.
>> 
>> ---------
>> - [x] I confirm that I make this contribution in accordance with the 
>> [OpenJDK Interim AI Policy](https://openjdk.org/legal/ai).
>
> Boris Ulasevich has updated the pull request incrementally with one 
> additional commit since the last revision:
> 
>   indent fix + move the comment

Looks good, let me submit the tests.

src/hotspot/share/opto/loopnode.cpp line 4605:

> 4603: 
> //---------------------------replace_parallel_iv-------------------------------
> 4604: // Replace parallel induction variable (parallel to trip counter)
> 4605: void PhaseIdealLoop::replace_parallel_iv(IdealLoopTree *loop) {

While you are at it, could you fix the style in this section, please? You may 
want to take a look at the diff by Github because it may include lines that are 
not technically touched by you, but since it's in the diff, it's better to fix 
the existing style.

-------------

PR Review: https://git.openjdk.org/jdk/pull/32549#pullrequestreview-5307158670
PR Review Comment: https://git.openjdk.org/jdk/pull/32549#discussion_r4095913414

Reply via email to