https://gcc.gnu.org/bugzilla/show_bug.cgi?id=124804
--- Comment #18 from Jakub Jelinek <jakub at gcc dot gnu.org> ---
So, I think it is clear what's wrong. In the no scheduling case, we have (insn
66 + 86 + 90):
lxvpx 44,30,2
...
xvbf16ger2pp 5,45,49
...
xvbf16ger2pp 5,44,48
which is correct, lxvpx loads a __vector_pair from the correct address to
$vs44/$vs45:
(gdb) x/4gx $r2+$r30
0x10020120 <a.2+160>: 0x3f943fa73fb73fa9 0x3f933f9e3f7a3f1a
0x10020130 <a.2+176>: 0x3f813f493f6a3fad 0x3fb03f663fba3f41
(gdb) stepi
0x0000000010000900 in BF16GEMV_T_MMA_8 ()
(gdb) p/x $vs44
$33 = {float128 = 0x3fb03f663fba3f413f813f493f6a3fad, uint128 =
0x3fb03f663fba3f413f813f493f6a3fad, v2_double = {0x3f813f493f6a3fad,
0x3fb03f663fba3f41}, v4_float = {0x3f6a3fad,
0x3f813f49, 0x3fba3f41, 0x3fb03f66}, v4_int32 = {0x3f6a3fad, 0x3f813f49,
0x3fba3f41, 0x3fb03f66}, v8_int16 = {0x3fad, 0x3f6a, 0x3f49, 0x3f81, 0x3f41,
0x3fba, 0x3f66, 0x3fb0},
v16_int8 = {0xad, 0x3f, 0x6a, 0x3f, 0x49, 0x3f, 0x81, 0x3f, 0x41, 0x3f, 0xba,
0x3f, 0x66, 0x3f, 0xb0, 0x3f}}
(gdb) p/x $vs45
$34 = {float128 = 0x3f933f9e3f7a3f1a3f943fa73fb73fa9, uint128 =
0x3f933f9e3f7a3f1a3f943fa73fb73fa9, v2_double = {0x3f943fa73fb73fa9,
0x3f933f9e3f7a3f1a}, v4_float = {0x3fb73fa9,
0x3f943fa7, 0x3f7a3f1a, 0x3f933f9e}, v4_int32 = {0x3fb73fa9, 0x3f943fa7,
0x3f7a3f1a, 0x3f933f9e}, v8_int16 = {0x3fa9, 0x3fb7, 0x3fa7, 0x3f94, 0x3f1a,
0x3f7a, 0x3f9e, 0x3f93},
v16_int8 = {0xa9, 0x3f, 0xb7, 0x3f, 0xa7, 0x3f, 0x94, 0x3f, 0x1a, 0x3f, 0x7a,
0x3f, 0x9e, 0x3f, 0x93, 0x3f}}
so $vs44 contains the 16 bytes at $r2+$r30+16 and $vs45 contains the 16 bytes
at $r2+$r30+0 and the latter is used in the middle operand of the third
xvbf16ger2pp into $a5 and
the former in the fourth xvbf16ger2pp into $a5.
Now, with early scheduling, RA decided not to load it early, so does instead
lxvpx 32,28,2
xvbf16ger2pp 5,33,63
...
lxvx 32,28,2
xvbf16ger2pp 5,32,62
$r2+$r28 is the correct address to find the 32 bytes we need:
(gdb) x/4gx $r2+$r28
0x10020120 <a.2+160>: 0x3f943fa73fb73fa9 0x3f933f9e3f7a3f1a
0x10020130 <a.2+176>: 0x3f813f493f6a3fad 0x3fb03f663fba3f41
and lxvpx loads 32 bytes and then uses $vs33, so on little endian the first 16
bytes from that (aka what lxvx 33,28,2 would load too).
But what's wrong is lxvx 32,28,2 is wrong, we want at that point to use the 16
bytes at $r2+$r28+16, but get instead what is in $r2+$28+0.
So, I think there is some misunderstanding between the backend and RA on
endianity, what actually
(subreg:V16QI (reg:OO) 16)
and
(subreg:V16QI (reg:OO) 0)
mean for ppc64le and what it means on
(subreg:V16QI (mem:OO) 16)
and
(subreg:V16QI (mem:OO) 0)