On Wed, Jul 22, 2026 at 11:40:06AM +0800, Gregory Price wrote:
> On Wed, Jul 22, 2026 at 10:20:00PM +0800, Richard Cheng wrote:
> > On Mon, Jul 20, 2026 at 03:33:54PM +0800, Gregory Price wrote:
> > Hi Gregory,
> > 
> > I applied your series on mm-next and give a quick review,
> > not thoroughly, still I have some questions below regarding the design.
> >  
> 
> Hi Richard,
> 
> Thank you for the read.
> 
> Just a heads up, this was based on mm-new because of some recent work
> on Brendan's page_alloc.h cleanup, but if you managed to get it applied
> then maybe that's made its way forward already.
> 
> > > The goal here is to flip that dynamic, isolate by default and then
> > > opt-in to specific services that the device says is safe.
> > > 
> > > Isolation at the NUMA/Zonelist layer provides a powerful mechanism for
> > > memory hosted on accelerators - re-use of the kernel mm/ code.
> > > 
> > >  - Accelerators (GPUs) can use demotion, numactl, and reclaim.
> > >  - Special memory devices (Compressed RAM) with special access controls
> > >    (promote-on-write) can have generic services written for them.
> > >  - Network devices with large memory regions intended for ring buffers
> > >    can use the buddy and standard networking stack.
> > >  - Slow, disaggregated memory pools which aren't suitable as general
> > >    purpose memory get cleaner interfaces (no need to re-write the buddy
> > >    in userland, can use migration interface, etc).
> > >  - Per-workload dedicated memory nodes (disaggregated VM memory)
> > > 
> > 
> > For accelerator part, take CXL Type-2 as example, it has its own protocol
> > CXL.cache, CXL.mem which is the rule they need to obey during memory 
> > transaction,
> > Adding more rules for them since they are NUMA node confused me, I wonder 
> > the reason ?
> > 
> 
> CXL protocols simply state how to do the memory transaction at a
> physical / transport level.  It does not make any statement on how the
> operating systems are to make sense of these devices or what constructs
> / abstractions to build to actually manage the devices themselves.
>

Ahhh thanks this clears my question.
Then this abstraction makes sense.

> This series is not necessarily attached to CXL in particular, you could
> just as easily carve out memory from the general pool onto a private
> node to ensure it only gets used for a particular use.
> 
> In fact, that is how I have been testing this with dax/kmem:
> https://github.com/gourryinverse/linux/commit/f279e741d9c3643597525bfab916b29b95cb635b
> 
> > Because NUMA node, at least for me, representing topology/locality rather 
> > than
> > something with ownership or capability or rules.
> > 
> 
> The NUMA abstraction representing topology/locality is a construct we
> (the OS developers) have decided on historically - but there's nothing
> that dictates we can never create new useful abstractions with it.
> 
> Consider:
>     N_NORMAL_MEMORY,        /* The node has regular memory */
>     N_HIGH_MEMORY,          /* The node has regular or high memory */
> 
> These have nothing to do with topology or locality, they are node states
> that only have meaning in the context of linux mm/.
> 

Ok, I see.

> 
> To make something clear - there is no *requirement* for any particular
> device to use this abstraction.  It simply enables a cleaner way for
> devices carrying memory to re-utilize mm/ services while ensuring their
> memory does not silently get used under system pressure (or some vagrant
> in userspace doing `numactl --interleave all`).
> 
> Ignoring all the CAP bits entirely, if you just took the base series
> your driver could re-use the buddy allocator without any special logic
> AND have confidence that your driver is the only possible user of that
> memory (barring some truly obscene bug).
> 
> > >  - isolation via a dedicated zonelist:
> > >      - private nodes are omitted from FALLBACK/NOFALLBACK
> > >      - added ZONELIST_PRIVATE(_NOFALLBACK)
> > 
> > I saw the reply in patch 5, so in fact there's not only one
> > dedicate zonelist, but numerous ?
> > I raise the question because zonelist was supposed to be a
> > global, unbypasssable thing in MM design, but now what you
> > are trying to do is to seperate the whole global list into several
> > parts ?
> > 
> > Can you explain why in current design you don't consider to support
> > something like ZONELIST_PRIVATE[n]={0, .. ,n-1} ? that's my imagination
> > of what a global zonelist should look like, no matter private or non.
> > 
> 
> First let me say that there's nothing that prevents us from doing this,
> and we *could* make this the default case - but this decision was
> intentional by me.
> 
> Having them all present in each-other's zonelists by default creates a
> number of implicit opt-ins that are unclear:
> 
>   - any direct zonelist iterator now iterates all zones on all private
>     nodes, even those nodes do not opt into the same services
> 
>     e.g. zonelist iteration in reclaim that targets Private Node A would
>     attempt to reclaim Private Node B as well.  That would require and
>     extra explicit filter.
> 
>     I ask: Why do this? Just isolate in the zonelist, and if there is
>            a desire for intersections - make it explicit, not implicit.
> 

Fair enough, thanks for the headup.

>   - fallback allocations can now occur across private nodes, even
>     if those nodes are intended for different purposes.
> 
>     Obviously you can use nodemask to tighten the allocation target, but
>     I use `numactl --interleave --all` as an example of a clear case
>     where intersected nodelists may not give you the behavior you want.
> 
> 
> So for consistency - everything including zonelists have full isolation.
> This way all interactions with a private node must be explicit - always.
> 

Good choice I think, avoiding those implicit opt-in assumption will make
future developers more aware of things and do not just stumble into them.

> 
> This is actually one of the problems with ZONE_DEVICE - and you can see
> it in this patch set.  Some of the hooks in mm/ for zone_device only
> apply to PTE cases, and are absent from PMD cases - only because PMD
> mappings in ZONE_DEVICE aren't supported.
> 
> That kind of implicit behavior is quite bad.
> 
> That said, future improvements could include something like:
> 
> for_reclaimable_zone(ZONELIST_PRIVATE) {}
> 
> where we loosen this isolation, and formalize a filter, but I would
> like to see the usecase for it first.  Loosening the isolation defeats
> the entire purpose of the series, so there should be a strong reason to
> do so.

Totally agree.

> 
> > > Allocation Isolation
> > > ====================
> > > page_alloc presently controls whether a node's memory can be allocated
> > > on a given call by 4 things (in order of authority)
> > > 
> > >   1) ZONELIST membership
> > >      If a node is not in the walked zonelist, it's unreachable.
> > >
> > 
> > So a device gets hotplugged in the system will get a dedicated zonelist 
> > here ?
> 
> Yes.
> 
> > And make sure it obeys the device's own protocol if it has one ?
> > 
> 
> I'm not sure i follow this question, can you help me understand?
> 

Nevermind, you just answer this part above, thanks alot.

> If by protocol you mean the CAP bits (opting into reclaim, demotion,
> etc), then yes.  If you mean something else (CXL) then I think that's
> orthogonal and unrelated.
> 
> > > Private nodes:
> > >   1) Never appear in any ZONELIST_FALLBACK
> > >   2) Have an empty ZONELIST_NOFALLBACK
> > >   3) Only appear in their own ZONELIST_PRIVATE(_NOFALLBACK)
> > > 
> > > 1 & 2 mean all existing in-tree callers to page_alloc can NEVER
> > > accidentally allocate from a private node.
> > > 
> > > An allocation must explicitly ask via a zonelist and a nodemask.
> > > 
> > >   alloc_flags |= ALLOC_ZONELIST_PRIVATE; /* use ZONELIST_PRIVATE */
> > >   __alloc_pages(..., nodemask);     /* with the private node set */
> > > 
> > 
> > As I stated above, NUMA concept was quite naive at first glance for my
> > limited knowledge.
> > 
> > This is quite alot to add for NUMA node concept, I'll want to see
> > more explanation in v6 and learn from it, thanks.
> > 
> 
> Sure, I can expand on it.  I think it's not as much to add as you think
> though - it's simply adding the concept of isolation to a NUMA node.
> 
> Some more explanation below, but if you think there is something i
> should explicit spell out in v6 cover, please let me know.
> 
> ---
> 
> I agree that NUMA as a concept was best-effort for, comically enough,
> somewhat *Uniform* memory access - instead of *Non*-Uniform memory access.
> 
> The current abstraction quite nicely handles the case where all memory
> on the system is roughly of the same calibre (DDR4, DDR5, etc) and
> roughly for the same purpose (general system memory).
> 
> But it is quite incapable of handling truly heterogeneous memory systems
> (precious HBM attached to the CPU, GPU's with HBM over a coherent link,
> hardware-compressed memory expansion, network devices w/ memory, etc).
> 
> If you look at the history of ZONE_DEVICE, what it fundamentally does is
> slaps an isolation mechanism on top of NUMA nodes because the NUMA
> abstraction doesn't provide one. 
> 
> The problem with that approach is now you have to reason about a node
> having both fungible and non-fungible memory.  It creates the need for
> something like migrate_device.c when a properly isolated NUMA node could
> just use migrate.c directly (with a coherent link).
> 
> 
> One example of what isolation on the node enables:
> 
> A Private Node can hotplug memory in ZONE_NORMAL - which means it
> can be GUP pinned.  That means driver support for GPU direct storage
> is simply `alloc_pages_node() + pin()`. The driver doesn't have to
> worry about something like SLAB accidentally using that same memory.
> 
> All without having to rewrite a bunch of mm/ in a driver, and with
> having to do some kind of heroics with ZONE_DEVICE that causes even
> more special mm/ interactions.
> 
> That only comes from adding an isolation primitive to a NUMA node.
> 
> 
> > > Isolating private node folios from kernel services
> > > ==================================================
> > > We implement filter points in mm/ to prevent operations on
> > > private node memory.  Where possible, we even re-use existing
> > > filter points from ZONE_DEVICE.
> > > 
> > > Most filter points are one or two lines of code:
> > > 
> > > Combining ZONE_DEVICE and N_MEMORY_PRIVATE opt-out spots:
> > >   -     if (folio_is_zone_device(folio))
> > >   +     if (unlikely(folio_is_private_managed(folio)))
> > > 
> > > Disabling a service:
> > >   +  if (!node_is_private(nid)) {
> > >   +      kswapd_run(nid);
> > >   +      kcompactd_run(nid);
> > >   +  }
> > > 
> > > Disallowing a uapi interaction:
> > >   +  if (node_state(nid, N_MEMORY_PRIVATE))
> > >   +      return -EINVAL;
> > > 
> > 
> > I'm not sure of why do we re-implement more filter and basically
> > doing the same thing ? Any unavoidable scenario ?
> > 
> > re-using the exisintg filter would be nice if that's possible.
> > 
> 
> We re-use (combine) existing filters where possible.
> 
> We add new ones where ZONE_DEVICE did not implement filters due to some
> *implicit* filter already existing.
> 
> Two clear examples:
>    - reclaim does not target ZONE_DEVICE
>    - ZONE_DEVICE does not support PMD
> 
> Both cases result in private-node filters that otherwise would have to
> exist for ZONE_DEVICE (and if ZONE_DEVICE ever grows PMD support, those
> filters will have to be added).
> 
> > > Bonus Configuration: HBM device memory tiering
> > > ==============================================
> > > echo 1 > dax0.0/private       # make the node private
> > > echo 0 > dax0.0/adistance     # highest tier
> > > echo 1 > dax0.0/reclaim       # reclaim active
> > > echo 1 > dax0.0/demotion      # may demote from the node
> > > echo 1 > dax0.0/user_numa     # mbind()
> > > echo online_movable > dax0.0/state
> > > echo 1 > numa/demotion_enabled
> > > 
> > > This is an HBM device which is treated as the top-tier in the
> > > system but for which memory can only enter via explicit mbind().
> > > 
> > > It can be overcommitted because it can be reclaimed (demotions
> > > go to CPU DRAM, and reclaim can swap from it).
> > > 
> > > If the HBM is managed by an accelerator (GPU), the mmu_notifier
> > > allows it to know when reclaim is moving memory out to do
> > > device-mmu invalidation prior to migration.
> > > 
> > > Prereqs, base commit, references
> > > ================================
> > > akpm/mm-new - for Brendan Jackman's mm/page_alloc.h work[3]
> > > 
> > 
> > Still thanks for the work, I learn alot from your work as well, thanks.
> > 
> 
> Of course, and thank you for reading.
> 
> If nothing else I hope this series helps folks understand the page
> allocator better (it certainly has helped me).  If there is anything
> you think I can improve or better explain, I am happy to discuss.
> 
> I am planning a larger publication of all my research sometime in the
> future, but I think code is more impactful and useful - so I am
> prioritizing that for now :].
> 
> ~Gregory

That would be helpful for me, thanks alot for all the explanation again !

--Richard


Reply via email to