Stephen,
I’m thinking, perhaps the best path forward for now, is to continue using a 
private mbuf layout and carry the collateral internally — at least for now. The 
changes are getting well beyond the initial desired impacts and has potential 
make things become unstable.

I would love to see some sort of solution along these lines in the future so we 
can discontinue the use of a private patch. To be clear, the solution is not 
built around cnxk, cnxk is just impacted — we have several other platform 
builds that do not use cnxk but we still carry the metadata in the mbuf itself 
for performance reasons.

WRT the skip functions, we fit under the 384 max for wqe but there’s no room 
for additional growth beyond that with how it’s laid out. The other 2 are not 
as sensitive to the size but still have valid limits which are well beyond our 
use case. Today we carry 256 bytes which caps us on wqe.

That said, we plan to discuss internally next week on how to move forward as we 
prep for our v25.11 upgrade work. I want to thank you and Morten, however, for 
your time and support on this work.
Regards,
-rt

From: Stephen Hemminger <[email protected]>
Date: Friday, October 9, 2026 at 12:14 PM
To: Randy Tice (rtice) <[email protected]>
Cc: [email protected] <[email protected]>; Morten Brørup <[email protected]>; 
Nithin Dabilpuram <[email protected]>; Harman Kalra <[email protected]>
Subject: Re: [PATCH v4 0/1] mbuf: add runtime metadata dynamic-field storage

On Thu, 8 Oct 2026 16:15:26 +0000
"Randy Tice (rtice)" <[email protected]> wrote:

> I did a brief look at this and still not convinced about the exact use case.
> What is the problem this is trying to solve and why can't it be done
> by using existing API’s.
>
> #RT:
>   Our implementation currently requires an additional 256 bytes of metadata 
> per
>   mbuf. That does not fit in the existing dynamic-field storage, so today we
>   carry a private patch to extend struct rte_mbuf. The goal of this work is to
>   replace that private struct change with a supported upstream mechanism.
>
>   We did look at using mbuf private data first. It can work when the 
> application
>   owns all mbuf pool creation, but it is not sufficient for our case without
>   another global/base reservation mechanism. Some mbuf pools are created by
>   drivers or libraries rather than directly by the application; the CNXK 
> inline
>   IPsec/OOP meta pool is one example (NIX_INL_META_POOL, created through
>   cnxk_nix_inl_meta_pool_cb()). To make private data work there, we had to add
>   an EAL argument that reserved a base private size for all pktmbufs, 
> including
>   PMD-created pools.
>
>   Private data also lacks a central layout registry. If multiple modules use
>   private data, they must coordinate offsets out of band to avoid overlaying
>   each other. The dynamic-field registry solves that coordination problem, but
>   the existing copied dynamic-field area is too small and has copy/clone
>   semantics that are wrong for this metadata.
>
>   That is why this version uses a globally configured per-mbuf metadata area
>   managed by the dynamic-field registry, with explicit metadata fields that 
> are
>   not copied by generic copy/clone/attach paths. This direction came out of 
> the
>   prior discussion with you, Morten, and me: avoid a Cisco-private mbuf struct
>   patch, avoid per-pool private-data layout coordination, and keep 
> sizeof(struct
>   rte_mbuf) fixed.

You uncovered a design flaw in the CNXK driver and the
proposed solution is wider than it needs to be.

All mbufs visible to application must come from mempools controlled by the 
application.
The design of CNXK driver is wrong, it shouldn't be using a private hidden pool.
It is ok for drivers to have hidden mempools that are used for non-visible 
things,
an example is the packet capture mempool where the mbufs only go into capture 
stream.

More long winded AI description:

On Thu, 8 Oct 2026 16:15:26 +0000
"Randy Tice (rtice)" <[email protected]> wrote:

> We did look at using mbuf private data first. It can work when the
> application owns all mbuf pool creation, but it is not sufficient for
> our case without another global/base reservation mechanism. Some mbuf
> pools are created by drivers or libraries rather than directly by the
> application; the CNXK inline IPsec/OOP meta pool is one example

That is the real bug. A driver should not hand the application mbufs
from a pool the application did not create. The Rx queue API already
takes the pool from the application, and rte_eth_rxconf can carry
more than one (rx_mempools). If cnxk needs a meta pool whose mbufs
reach the application, it should get it from the application, or at
minimum create it with the same priv_size as the Rx queue pool.

Internal pools are fine when the mbufs never cross the API, e.g.
dumpcap or the bonding LACP pool.

With that fixed, the existing private area does what you need. It is
per pool, sized by the application, and is not copied by copy, clone
or attach. Offset coordination inside it is the application's job,
and if a registry is wanted it can be layered on top of priv without
changing the mbuf layout.

So my answer to option 1 vs 2 is neither. I don't want a global
layout knob that puts a load in every inline helper and needs
per-driver range checks, to work around one driver.

What I would take:
 - cnxk: validate wqe_skip/later_skip/first_skip against priv_size
   and headroom. This is a bug today, independent of this series.
 - cnxk: meta pool comes from, or matches, the application pool.
 - ethdev: under RTE_ETHDEV_DEBUG_RX, check that received mbufs
   belong to a pool configured on that queue.

Nithin, Harman: do mbufs from NIX_INL_META_POOL get returned to the
application by rx_burst, or are they only consumed internally?

Reply via email to