On Wed, Sep 02, 2026 at 12:00:25PM +0200, Christian König wrote: > On 9/2/26 11:53, Leon Romanovsky wrote: > > On Wed, Sep 02, 2026 at 10:44:59AM +0200, Christian König wrote: > >> On 9/2/26 10:32, Leon Romanovsky wrote: > >>> On Wed, Sep 02, 2026 at 09:56:06AM +0200, Christian König wrote: > >>>> On 9/2/26 09:39, Leon Romanovsky wrote: > >>>>> On Wed, Sep 02, 2026 at 09:00:46AM +0200, Christian König wrote: > >>>>>> On 9/1/26 19:08, David Hu wrote: > >>>>>>> From: David Hu <[email protected]> > >>>>>>> > >>>>>>> This series address two related issues in scatter-gather mapping, > >>>>>>> specifically for the MMIO based dma-buf mapping. The fixes ensure > >>>>>>> sgt mapping is correct, and proper for large MMIO regions. > >>>>>>> > >>>>>>> Patch 1 fixes a silent integer overflow for mapping length exceeding > >>>>>>> 4G > >>>>>>> (Previously submitted as [PATCH v7] dma-buf: Fix silent overflow for > >>>>>>> phys vec to sgt) > >>>>>>> https://lore.kernel.org/all/[email protected]/ > >>>>>>> > >>>>>>> Patch 2 Splits sgl by largest page aligned chunk > >>>>>>> (Previously submitted as [PATCH v3] dma-buf: Split sgl by largest > >>>>>>> page-aligned chunk) > >>>>>>> https://lore.kernel.org/all/[email protected]/ > >>>>>> > >>>>>> *sigh* such issues are exactly the reason why I didn't wanted the > >>>>>> dma-mapping stuff inside DMA-buf. That clearly doesn't belong here. > >>>>> > >>>>> And this is why so many in the kernel community want to get rid of SG > >>>>> lists. It would be great if DMA-BUF could also eliminate the need to > >>>>> convert to an SGL, like Jason proposed. > >>>>> > >>>>> The DMA layer no longer needs SGL. These bugs belong to the DMA-BUF > >>>>> layer, > >>>>> which is the one that depends on it. > >>>> > >>>> I'm all fine using an array/xarray of dma_addr_t in DMA-buf, just > >>>> phys_vec is a clear no-go. > >>> > >>> You are proposing the same thing as an SGL, just in a different format. > >> > >> Yes, because that is the right thing todo as far as I can see. > >> > >>> It does not address the issue that dma_addr_t is expected to hold a DMA > >>> address, while that is not always the case. For example, in the P2P case, > >>> the addresses are not DMA addresses. > >> > >> Yes they are. They must be DMA addresses because that is the only thing > >> the importer needs to do it's DMA. > > > > They can perform DMA, but that still does not make them suitable for the > > dma_addr_t type. For the PCI_P2PDMA_MAP_BUS_ADDR flow, these addresses > > follow completely different rules: they are not unmapped, require no cache > > synchronization, are valid only for peer access, and require separate error > > handling. > > The PCI_P2PDMA_MAP_BUS_ADDR is not supported by DMA-buf and as far as I can > see is a complete dead end. > > > All of this information is lost if only the dma_addr_t is stored. > > Yes and that is fully intentional. > > DMA-buf handles that cleanly on the buffer object level and not like > PCI_P2PDMA_MAP_BUS_ADDR as a completely broken design on a per address/page > basis. > > Technical background is that the PCI_P2PDMA_MAP_BUS_ADDR approach can only be > handled by a very very small subset of HW. >
With the rise of accelerators, the P2PDMA_MAP_BUS_ADDR is going to be increasingly more common where accelerators directly transfer data to NICs, storage devices, other accelerators etc. With VFIO gaining a DMABUF exporter the use has already spread to RDMA & NVMe devices (w/ SPDK). In fact, we found these bugs while trying to map large BAR regions for RDMA. Thus, it would be great if we could find alignment here. AFAICT, I foresee the use of dmabufs to only increase for PCI_P2PDMA_MAP_BUS_ADDR. IIRC when the network stack moved to net_iovs to support dmabufs, the SGL became a primary concern and partly the reason why we have "unreadable" skbs for memory we can indeed access if mapped correctly. Thanks, Praan
