On Fri, Aug 28, 2026 at 02:03:01PM -0300, Jason Gunthorpe wrote: > For something complex like this, if you can't concretetly tie the HW > to a net namespace, and follow the net namespace rules for visibility, > then it is going to be a painful choice. I speak from alot of rdma > experiance where net namespaces have been consistently challenging.
Agreed. The series does not provide meaningful namespace semantics; it is host-global and confined to init_net. I would rather make that scope explicit than invent per-netns fabric semantics that the hardware does not have. Generic Netlink was a practical starting point, and drm_ras established the pattern in DRM, but that is a simpler problem and does not settle the choice here. A device-fd model like Jiri's may fit visibility and access control better, since a global family has no descriptor to check against. I would not switch transports based on an early pre-RFC, but I no longer consider the transport choice settled. > net/ is mainly focused on IP networking, it is where you should be > putting the ethernet layer at the bottom of the ua link over ethernet, > SUE, or whatever. Agreed. Ethernet PHY, MAC, packet processing and switch routing belong in the networking stack, not in DRM or an accelerator driver. Scale-up can reuse those layers where they fit. What differs is the accelerator-facing semantics, even when the underlying transport reuses Ethernet. From the accelerator side, a scale-up endpoint behaves more like another core in a tightly coupled system than an ordinary NIC. This series does not duplicate the networking layers. It models endpoints, ports, direct adjacency, membership and local state: no netdev, packet processing, Ethernet PHY management, route computation or switch forwarding. > There are so many variations of these "scale up" fabrics now, it would > probably be appropriate to have one subsystem that aims to work with > all of them. It is almost rdma but different enough it probably > wouldn't fit well. Worth testing. I would rather approach that through the minimum common object model than start by defining a complete subsystem. DRM is the current location because the initial providers are DRM and accel drivers, not because DRM should own the transport or must be the final home. The objects do not depend on GEM, scheduling, display or DRM memory semantics. Switches are the difficult case. A UALink switch is not a DRM device, and complete Pod topology may involve information owned by the Pod Controller and switch-management plane, not only accelerator drivers. If that information has to be represented as first-class objects, DRM may not be the right final home. There is no established common Linux answer for this class of fabric yet. I would rather start with the minimum topology objects, test them against direct-link and switched implementations, and review the model with multiple vendors before deciding its final scope and home. The LPC BoF is where I want the accelerator, networking and RDMA sides in the same room: https://lpc.events/event/20/contributions/2412/ Thanks, Konstantin
