On Fri, Oct 02, 2026 at 09:36:47PM -0400, Hamza Mahfooz wrote:
> Commit 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers
> instead of full pages to improve memory efficiency.") started handing
> out RX buffers with zero headroom so that two buffers fit into one page
> at the default MTU.
> 
> The MANA TX path, however, stores the per scatter-gather entry DMA
> mappings in `struct mana_skb_head` at skb->head, and mana_start_xmit()
> therefore calls skb_cow_head(skb, MANA_HEADROOM). The port advertises
> this requirement as ndev->needed_headroom = MANA_HEADROOM.
> 
> As a result every packet that is received and then forwarded out of a
> MANA port fails the skb_cow() in ip_forward() and gets reallocated and
> copied by pskb_expand_head(). This is invisible to a plain RX or TX
> workload, but it puts a full skb reallocation plus memcpy on the hot
> path of every single forwarded packet, which is exactly what a
> router/NVA workload does.
> 
> Restore the headroom. Note that reserving MANA_HEADROOM (232) is not
> enough: ip_forward() asks for LL_RESERVED_SPACE(dev), which rounds
> hard_header_len + needed_headroom up to HH_DATA_MOD and is 256 bytes on
> ethernet. Use LL_RESERVED_SPACE() directly so the value keeps tracking
> both constants. Also, since LL_RESERVED_SPACE() tracks MANA_HEADROOM,
> it grows with MAX_SKB_FRAGS and for MAX_SKB_FRAGS >= 19 it is greater
> than 256, so we have to account for that by using the headroom the RX
> queue actually uses (instead of assuming XDP_PACKET_HEADROOM) and
> turning MANA_XDP_MTU_MAX into MANA_XDP_MTU_MAX(ndev) (note that at the
> default CONFIG_MAX_SKB_FRAGS=17 they are equivalent).
> 
> At the default MTU on a 4K page this means a buffer no longer fits twice
> into a page (SKB_DATA_ALIGN(1500 + MANA_RXBUF_PAD + 256) = 2112), so the
> frag-vs-single decision is now made by computing the real buffer size
> instead of comparing the MTU against PAGE_SIZE / 2. The page_pool
> fragment path is still used wherever at least two buffers genuinely fit,
> e.g. on 16K and 64K page sizes.
> 
> Measured on an Azure VM with a MANA NIC acting as a forwarding NVA (UDP,
> 1400 byte payload, 4 streams, 8 Gbps offered, only the forwarding
> node's kernel differs), 8 runs each, median:
> 
>                   forwarded pps    throughput
>   before              272,830       3.06 Gbps
>   after               390,560       4.37 Gbps   (+43%)
> 
> perf on the forwarding node, same workload:
> 
>                   memset_orig   __pi_memcpy   pskb_expand_head
>   before             10.07%         3.96%         present
>   after               0.94%         0.64%         gone
> 
> Cc: [email protected]
> Fixes: 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers 
> instead of full pages to improve memory efficiency.")
> Signed-off-by: Hamza Mahfooz <[email protected]>
> ---
> v2:
> - Fix the XDP headroom mismatch with CONFIG_MAX_SKB_FRAGS >= 19,
>   by passing rxq->headroom to xdp_prepare_buff().
>   mana_build_skb() then picks up the correct offset via
>   xdp->data - xdp->data_hard_start. Also, turn MANA_XDP_MTU_MAX
>   into MANA_XDP_MTU_MAX(ndev) to account for the headroom,
>   since it is no longer a constant. (Narcisa, Sashiko)
> - Use the new mana_single_rxbuf_per_page_forced() helper in
>   mana_set_priv_flags(). (Sashiko)
> - Trim the comment above mana_get_rxbuf_headroom() and drop the stale
>   "XDP headroom" wording from the comment above mana_get_rxbuf_cfg().
>   (Narcisa, Sashiko)

Reviewed-by: Simon Horman <[email protected]>


Reply via email to