On Fri, 2 Oct 2026 16:09:25 -0400 Md Rayhanul Islam <[email protected]> wrote:
> This RFC reworks the BCM2711 DMA fix around a per-device DMA context and > an EAL sync API, after the feedback on the first posting [1]. > > On Raspberry Pi 4 / Compute Module 4 the PCIe host bridge translates DMA > addresses and is not cache coherent, so a device such as the Intel I210 > cannot DMA at all without both being handled. > > What changed since the first posting: > > * The translation and the coherency are read per device from its own > bridge, not kept as one process-wide offset. EAL no longer rewrites > IOVAs; the driver translates where it programs the hardware. > * A driver opts in with RTE_PCI_DRV_DMA_NONCOHERENT, and the bus refuses > such a device to any driver that has not. That replaces the refusal > written into em by hand. > * The cache maintenance is an experimental EAL interface with explicit > directions, and the loops end in DSB SY. > * Completion comes from the Done bits, not the head register. > * Both environment variables are gone. > > Tested between two Compute Module 4 boards with Intel I210s, both running > this series: testpmd txonly holds 1.42 Mpps with 64-byte frames, and a > 512 MB UDP transfer with DPDK at both ends arrives byte-identical. Built > with GCC and with clang 14 -Werror. More detailed review with AI, I had to direct it to ignore using AF_XDP in this case. Native PMDs on Raspberry Pi class boards are a reasonable goal, and the analysis of descriptor line sharing is good. Needs rework before a non-RFC version. Direction - Split the series: (1) detection and refusing to probe, which is a fix on its own since igb on a CM4 today DMAs to bus address 0; (2) the EAL cache maintenance API; (3) driver support. - Coherency detection should not depend on the bridge allowlist, only window parsing does. On arm64 DT, no dma-coherent means non-coherent (same as of_dma_is_coherent()). At least warn for any DT host bridge without it. - igb will not be the last PMD (igc is next). Line ownership rules for rings, mempool validation against the window and burst variant selection belong in common code, not copied per driver. - An address limit is a property of memory. Enforce it at hugepage allocation (extend rte_mem_set_dma_mask(); a bit mask can't express 3 GB) and keep only the offset add in the datapath. - Document that this needs uio_pci_generic or vfio-noiommu. VFIO with an IOMMU refuses non-coherent devices. Translation - BCM2711 does not translate: pcie0 dma-ranges is 1:1 with a 3 GB limit. - BCM2712 (Pi 5, CM5) does (PCIe 0x10_0000_0000 maps to CPU 0), but the parser rejects it: 64-bit memory space code, more than one dma-ranges entry (MSI window, plus a 32-bit window on pcie2), and brcm,bcm2712-pcie is not in the table. The offset path has never run. Bus and EAL - PCI_LOG uses dev->name before pci_common_set() sets it. - --iova-mode=va overrides the forced PA. Probe must check rte_eal_iova_mode() == RTE_IOVA_PA. - The DT walk runs for every PCI device on every platform. Skip it when /sys/firmware/devicetree does not exist. - rte_pci_dma_info is driver-only API; it goes in bus_pci_driver.h. - rte_mem_sync.h: arch code goes in lib/eal/arm/include behind a generic/ header. "Empty on coherent platforms" is wrong (it is empty on non-arm64). Read CTR_EL0 once at init, since it can trap. Stride by the CTR_EL0 line size, not RTE_CACHE_LINE_MIN_SIZE. for_cpu is clean+invalidate, so RTE_ASSERT alignment. Add a test in app/test. igb - Per-packet validation (igb_tx_pkt_reachable, igb_rx_buf_usable) belongs at queue setup or in tx_prepare. TX silently stops the burst, so retrying applications spin. RX stalls and counts as mbuf allocation failure. - Context descriptors are written before the segment loop, so their lines are never acquired. - The acquire narrows the DD loss window but does not close it. Correctness rests on the last_rs fallback; document that. - last_rs is updated per packet, so the in-burst fallback checks a descriptor not yet posted. Snapshot it at burst entry and reset it in igb_reset_tx_queue(). - The dma_per_line fallback to 1 is dead code, and if ever reached it would reintroduce the bug. Make it an error. - The new fields at the head of the queue structs push the coherent path's hot fields down. Move them to the end and show pahole. - TXDCTL: if OR-ing into WTHRESH is wrong, it is wrong on every platform. Fix it separately with a Fixes: tag, and explain forcing wthresh to 0. - Nits: limits.h and unistd.h are unused, igb_dma_set_burst() in the igbvf path is dead, and the error on 32-bit Arm is misleading. Testing - 1.42 Mpps is not 64 byte line rate (1.488). What frame size? - Need RX and io/mac fwd numbers, runs with checksum, VLAN insert and TSO offloads, and a Pi 5. - Board RAM size? On a 4 or 8 GB CM4, hugepages can land above 3 GB.

