On Fri, Jul 17, 2026 at 12:13 PM Andrew Stubbs <[email protected]> wrote: > > On 17/07/2026 09:13, Richard Biener wrote: > > On Wed, Jul 8, 2026 at 4:48 PM Andrew Stubbs <[email protected]> wrote: > >> > >> Hi all, > >> > >> This patch series implements the "MEMs with address vectors" feature > >> that I want to use in AMD GCN, and have previously discussed on gcc@. > >> The motivation is to simplify a new back-end optimization pass I'm > >> working on and plan to post in the coming weeks. > >> > >> The first patch, in this series, is the middle-end enablement and > >> documentation changes. > >> > >> The second patch converts the amdgcn back-end to use the new scheme, > >> without adding any real new features. It depends on the > >> INSN_BASE_REG_CLASS patch I posted last week (still in review) to run > >> correctly. > >> > >> The third and final patch demonstrates how I can use the new MEM to > >> extend my machine description without writing whole new insns for the > >> vector modes, just for them to evaluate to the same assembler code in > >> the end. > >> > >> I have not yet finished testing the patches, so these might not be quite > >> the final form, but I'm not expecting big changes to the middle-end > >> portion. (I can self-approve the amdgcn parts, once the other is > >> committed, but comments are welcome.) > >> > >> OK for mainline? (Pending test results.) > > > > I think this exposes what GCN does in a natural way. I'm not sure > > the advantage (exposing former gather/scatter as move) is not > > also its disadvantage (moves are special!), and I somewhat question > > the benefit over using some UNSPEC for this. For example I would > > assume that you side-step RTL alias/dependence analysis, meaning these > > are barriers for optimization and scheduling? > > This is intended to be an enabling technology for my real project, in > which I want to transform scalar operations into vector operations in > the backend, and the special casing for memory accesses was getting out > of hand. By expressing (unmasked) vector operations identically to the > scalar equivalent I can simplify things greatly, and the other compiler > passes can do (some of) their special MEM handling as they normally would.
I see. I guess you ruled out a ([parallel (set (mem:<scalar> ...) (...)) (set (mem:<scalar> ...) (...))]) etc. with vector extracts for the address. At least for "transform scalar ops to vectors" the actual representation of the MEM could be a black box as long as you have addresses in vectors (which is possible already). > > The proposed semantics (overlap has undefined behavior) does not > > match up with usual vector ISA scatter semantics (left-to-right), so > > the representation isn't universally useable (if we solve the masking > > issue). > > Unfortunately, this is the semantics of the hardware; specifying any > other assumptions in the compiler might mean that it wouldn't be safe to > use it for GCN. (It might be the case that the hardware is deterministic > in practice, but the details are not documented in the ISA manuals; it > just says that it will "write-combine data across the work-items".) Yeah, I think your proposed semantic is fine, it's just unfortunate we cannot make use of it elsewhere. > We can say that "the compiler assumes the UB doesn't happen", and trust > that the algorithm is constructed such that this is true. (It certainly > should be the case for the code I'm generating.) I'm not familiar with > the parts of the compiler where these assumptions are made. I don't > believe I have specifically disabled any such analysis, but if something > somewhere has a "return false" rather than a fatal assertion then I > suppose I wouldn't have seen it. > > > That said, my overall hunch is that we're extending general RTL > > semantics in a somewhat ad-hoc (and incomplete?) way. Not > > that I have a better idea here. > > I prefer "incremental" to "incomplete", and as ad-hoc designs go it > feels fairly natural, given what I want to represent. If time and money > were no issue I could certainly spend a lot longer designing and > implementing this feature, but it's a sub-project of a task order that > comes with a budget and a deadline, so I've implemented as much of it as > is necessary and I can test without constructing testcases for > hypothetical use-cases. Sure. Richard. > > > > Thanks, > > Richard. > > > >> > >> Thanks > >> > >> Andrew Stubbs (3): > >> rtl: Allow "(mem:<vecmode> (reg:<vecmode>))" > >> amdgcn: Implement "(mem (reg:<vecmode>))" > >> amdgcn: Add vector atomics > >> > >> gcc/config/gcn/constraints.md | 49 ++- > >> gcc/config/gcn/gcn-protos.h | 10 +- > >> gcc/config/gcn/gcn-valu.md | 745 ++++++++++++---------------------- > >> gcc/config/gcn/gcn.cc | 529 +++++++++++++----------- > >> gcc/config/gcn/gcn.h | 1 + > >> gcc/config/gcn/gcn.md | 123 +++--- > >> gcc/doc/rtl.texi | 8 + > >> gcc/doc/tm.texi | 13 +- > >> gcc/emit-rtl.cc | 30 +- > >> gcc/explow.cc | 12 +- > >> gcc/recog.cc | 4 +- > >> gcc/rtl.h | 2 +- > >> gcc/rtlanal.cc | 4 +- > >> gcc/simplify-rtx.cc | 5 +- > >> gcc/target.def | 13 +- > >> gcc/targhooks.cc | 4 +- > >> gcc/targhooks.h | 3 +- > >> 17 files changed, 768 insertions(+), 787 deletions(-) > >> > >> -- > >> 2.54.0 > >> >
