On Fri,  2 Oct 2026 16:09:25 -0400
Md Rayhanul Islam <[email protected]> wrote:

> This RFC reworks the BCM2711 DMA fix around a per-device DMA context and
> an EAL sync API, after the feedback on the first posting [1].
> 
> On Raspberry Pi 4 / Compute Module 4 the PCIe host bridge translates DMA
> addresses and is not cache coherent, so a device such as the Intel I210
> cannot DMA at all without both being handled.
> 
> What changed since the first posting:
> 
> * The translation and the coherency are read per device from its own
>   bridge, not kept as one process-wide offset.  EAL no longer rewrites
>   IOVAs; the driver translates where it programs the hardware.
> * A driver opts in with RTE_PCI_DRV_DMA_NONCOHERENT, and the bus refuses
>   such a device to any driver that has not.  That replaces the refusal
>   written into em by hand.
> * The cache maintenance is an experimental EAL interface with explicit
>   directions, and the loops end in DSB SY.
> * Completion comes from the Done bits, not the head register.
> * Both environment variables are gone.
> 
> Tested between two Compute Module 4 boards with Intel I210s, both running
> this series: testpmd txonly holds 1.42 Mpps with 64-byte frames, and a
> 512 MB UDP transfer with DPDK at both ends arrives byte-identical.  Built
> with GCC and with clang 14 -Werror.

More detailed review with AI, I had to direct it to ignore using AF_XDP
in this case.

Native PMDs on Raspberry Pi class boards are a reasonable goal, and the
analysis of descriptor line sharing is good. Needs rework before a
non-RFC version.

Direction
- Split the series: (1) detection and refusing to probe, which is a fix
  on its own since igb on a CM4 today DMAs to bus address 0; (2) the EAL
  cache maintenance API; (3) driver support.
- Coherency detection should not depend on the bridge allowlist, only
  window parsing does. On arm64 DT, no dma-coherent means non-coherent
  (same as of_dma_is_coherent()). At least warn for any DT host bridge
  without it.
- igb will not be the last PMD (igc is next). Line ownership rules for
  rings, mempool validation against the window and burst variant
  selection belong in common code, not copied per driver.
- An address limit is a property of memory. Enforce it at hugepage
  allocation (extend rte_mem_set_dma_mask(); a bit mask can't express
  3 GB) and keep only the offset add in the datapath.
- Document that this needs uio_pci_generic or vfio-noiommu. VFIO with an
  IOMMU refuses non-coherent devices.

Translation
- BCM2711 does not translate: pcie0 dma-ranges is 1:1 with a 3 GB limit.
- BCM2712 (Pi 5, CM5) does (PCIe 0x10_0000_0000 maps to CPU 0), but the
  parser rejects it: 64-bit memory space code, more than one dma-ranges
  entry (MSI window, plus a 32-bit window on pcie2), and
  brcm,bcm2712-pcie is not in the table. The offset path has never run.

Bus and EAL
- PCI_LOG uses dev->name before pci_common_set() sets it.
- --iova-mode=va overrides the forced PA. Probe must check
  rte_eal_iova_mode() == RTE_IOVA_PA.
- The DT walk runs for every PCI device on every platform. Skip it when
  /sys/firmware/devicetree does not exist.
- rte_pci_dma_info is driver-only API; it goes in bus_pci_driver.h.
- rte_mem_sync.h: arch code goes in lib/eal/arm/include behind a
  generic/ header. "Empty on coherent platforms" is wrong (it is empty
  on non-arm64). Read CTR_EL0 once at init, since it can trap. Stride by
  the CTR_EL0 line size, not RTE_CACHE_LINE_MIN_SIZE. for_cpu is
  clean+invalidate, so RTE_ASSERT alignment. Add a test in app/test.

igb
- Per-packet validation (igb_tx_pkt_reachable, igb_rx_buf_usable)
  belongs at queue setup or in tx_prepare. TX silently stops the burst,
  so retrying applications spin. RX stalls and counts as mbuf allocation
  failure.
- Context descriptors are written before the segment loop, so their
  lines are never acquired.
- The acquire narrows the DD loss window but does not close it.
  Correctness rests on the last_rs fallback; document that.
- last_rs is updated per packet, so the in-burst fallback checks a
  descriptor not yet posted. Snapshot it at burst entry and reset it in
  igb_reset_tx_queue().
- The dma_per_line fallback to 1 is dead code, and if ever reached it
  would reintroduce the bug. Make it an error.
- The new fields at the head of the queue structs push the coherent
  path's hot fields down. Move them to the end and show pahole.
- TXDCTL: if OR-ing into WTHRESH is wrong, it is wrong on every
  platform. Fix it separately with a Fixes: tag, and explain forcing
  wthresh to 0.
- Nits: limits.h and unistd.h are unused, igb_dma_set_burst() in the
  igbvf path is dead, and the error on 32-bit Arm is misleading.

Testing
- 1.42 Mpps is not 64 byte line rate (1.488). What frame size?
- Need RX and io/mac fwd numbers, runs with checksum, VLAN insert and
  TSO offloads, and a Pi 5.
- Board RAM size? On a 4 or 8 GB CM4, hugepages can land above 3 GB.

Reply via email to