On Wed, Jul 22, 2026 at 11:40:06AM +0800, Gregory Price wrote: > On Wed, Jul 22, 2026 at 10:20:00PM +0800, Richard Cheng wrote: > > On Mon, Jul 20, 2026 at 03:33:54PM +0800, Gregory Price wrote: > > Hi Gregory, > > > > I applied your series on mm-next and give a quick review, > > not thoroughly, still I have some questions below regarding the design. > > > > Hi Richard, > > Thank you for the read. > > Just a heads up, this was based on mm-new because of some recent work > on Brendan's page_alloc.h cleanup, but if you managed to get it applied > then maybe that's made its way forward already. > > > > The goal here is to flip that dynamic, isolate by default and then > > > opt-in to specific services that the device says is safe. > > > > > > Isolation at the NUMA/Zonelist layer provides a powerful mechanism for > > > memory hosted on accelerators - re-use of the kernel mm/ code. > > > > > > - Accelerators (GPUs) can use demotion, numactl, and reclaim. > > > - Special memory devices (Compressed RAM) with special access controls > > > (promote-on-write) can have generic services written for them. > > > - Network devices with large memory regions intended for ring buffers > > > can use the buddy and standard networking stack. > > > - Slow, disaggregated memory pools which aren't suitable as general > > > purpose memory get cleaner interfaces (no need to re-write the buddy > > > in userland, can use migration interface, etc). > > > - Per-workload dedicated memory nodes (disaggregated VM memory) > > > > > > > For accelerator part, take CXL Type-2 as example, it has its own protocol > > CXL.cache, CXL.mem which is the rule they need to obey during memory > > transaction, > > Adding more rules for them since they are NUMA node confused me, I wonder > > the reason ? > > > > CXL protocols simply state how to do the memory transaction at a > physical / transport level. It does not make any statement on how the > operating systems are to make sense of these devices or what constructs > / abstractions to build to actually manage the devices themselves. >
Ahhh thanks this clears my question. Then this abstraction makes sense. > This series is not necessarily attached to CXL in particular, you could > just as easily carve out memory from the general pool onto a private > node to ensure it only gets used for a particular use. > > In fact, that is how I have been testing this with dax/kmem: > https://github.com/gourryinverse/linux/commit/f279e741d9c3643597525bfab916b29b95cb635b > > > Because NUMA node, at least for me, representing topology/locality rather > > than > > something with ownership or capability or rules. > > > > The NUMA abstraction representing topology/locality is a construct we > (the OS developers) have decided on historically - but there's nothing > that dictates we can never create new useful abstractions with it. > > Consider: > N_NORMAL_MEMORY, /* The node has regular memory */ > N_HIGH_MEMORY, /* The node has regular or high memory */ > > These have nothing to do with topology or locality, they are node states > that only have meaning in the context of linux mm/. > Ok, I see. > > To make something clear - there is no *requirement* for any particular > device to use this abstraction. It simply enables a cleaner way for > devices carrying memory to re-utilize mm/ services while ensuring their > memory does not silently get used under system pressure (or some vagrant > in userspace doing `numactl --interleave all`). > > Ignoring all the CAP bits entirely, if you just took the base series > your driver could re-use the buddy allocator without any special logic > AND have confidence that your driver is the only possible user of that > memory (barring some truly obscene bug). > > > > - isolation via a dedicated zonelist: > > > - private nodes are omitted from FALLBACK/NOFALLBACK > > > - added ZONELIST_PRIVATE(_NOFALLBACK) > > > > I saw the reply in patch 5, so in fact there's not only one > > dedicate zonelist, but numerous ? > > I raise the question because zonelist was supposed to be a > > global, unbypasssable thing in MM design, but now what you > > are trying to do is to seperate the whole global list into several > > parts ? > > > > Can you explain why in current design you don't consider to support > > something like ZONELIST_PRIVATE[n]={0, .. ,n-1} ? that's my imagination > > of what a global zonelist should look like, no matter private or non. > > > > First let me say that there's nothing that prevents us from doing this, > and we *could* make this the default case - but this decision was > intentional by me. > > Having them all present in each-other's zonelists by default creates a > number of implicit opt-ins that are unclear: > > - any direct zonelist iterator now iterates all zones on all private > nodes, even those nodes do not opt into the same services > > e.g. zonelist iteration in reclaim that targets Private Node A would > attempt to reclaim Private Node B as well. That would require and > extra explicit filter. > > I ask: Why do this? Just isolate in the zonelist, and if there is > a desire for intersections - make it explicit, not implicit. > Fair enough, thanks for the headup. > - fallback allocations can now occur across private nodes, even > if those nodes are intended for different purposes. > > Obviously you can use nodemask to tighten the allocation target, but > I use `numactl --interleave --all` as an example of a clear case > where intersected nodelists may not give you the behavior you want. > > > So for consistency - everything including zonelists have full isolation. > This way all interactions with a private node must be explicit - always. > Good choice I think, avoiding those implicit opt-in assumption will make future developers more aware of things and do not just stumble into them. > > This is actually one of the problems with ZONE_DEVICE - and you can see > it in this patch set. Some of the hooks in mm/ for zone_device only > apply to PTE cases, and are absent from PMD cases - only because PMD > mappings in ZONE_DEVICE aren't supported. > > That kind of implicit behavior is quite bad. > > That said, future improvements could include something like: > > for_reclaimable_zone(ZONELIST_PRIVATE) {} > > where we loosen this isolation, and formalize a filter, but I would > like to see the usecase for it first. Loosening the isolation defeats > the entire purpose of the series, so there should be a strong reason to > do so. Totally agree. > > > > Allocation Isolation > > > ==================== > > > page_alloc presently controls whether a node's memory can be allocated > > > on a given call by 4 things (in order of authority) > > > > > > 1) ZONELIST membership > > > If a node is not in the walked zonelist, it's unreachable. > > > > > > > So a device gets hotplugged in the system will get a dedicated zonelist > > here ? > > Yes. > > > And make sure it obeys the device's own protocol if it has one ? > > > > I'm not sure i follow this question, can you help me understand? > Nevermind, you just answer this part above, thanks alot. > If by protocol you mean the CAP bits (opting into reclaim, demotion, > etc), then yes. If you mean something else (CXL) then I think that's > orthogonal and unrelated. > > > > Private nodes: > > > 1) Never appear in any ZONELIST_FALLBACK > > > 2) Have an empty ZONELIST_NOFALLBACK > > > 3) Only appear in their own ZONELIST_PRIVATE(_NOFALLBACK) > > > > > > 1 & 2 mean all existing in-tree callers to page_alloc can NEVER > > > accidentally allocate from a private node. > > > > > > An allocation must explicitly ask via a zonelist and a nodemask. > > > > > > alloc_flags |= ALLOC_ZONELIST_PRIVATE; /* use ZONELIST_PRIVATE */ > > > __alloc_pages(..., nodemask); /* with the private node set */ > > > > > > > As I stated above, NUMA concept was quite naive at first glance for my > > limited knowledge. > > > > This is quite alot to add for NUMA node concept, I'll want to see > > more explanation in v6 and learn from it, thanks. > > > > Sure, I can expand on it. I think it's not as much to add as you think > though - it's simply adding the concept of isolation to a NUMA node. > > Some more explanation below, but if you think there is something i > should explicit spell out in v6 cover, please let me know. > > --- > > I agree that NUMA as a concept was best-effort for, comically enough, > somewhat *Uniform* memory access - instead of *Non*-Uniform memory access. > > The current abstraction quite nicely handles the case where all memory > on the system is roughly of the same calibre (DDR4, DDR5, etc) and > roughly for the same purpose (general system memory). > > But it is quite incapable of handling truly heterogeneous memory systems > (precious HBM attached to the CPU, GPU's with HBM over a coherent link, > hardware-compressed memory expansion, network devices w/ memory, etc). > > If you look at the history of ZONE_DEVICE, what it fundamentally does is > slaps an isolation mechanism on top of NUMA nodes because the NUMA > abstraction doesn't provide one. > > The problem with that approach is now you have to reason about a node > having both fungible and non-fungible memory. It creates the need for > something like migrate_device.c when a properly isolated NUMA node could > just use migrate.c directly (with a coherent link). > > > One example of what isolation on the node enables: > > A Private Node can hotplug memory in ZONE_NORMAL - which means it > can be GUP pinned. That means driver support for GPU direct storage > is simply `alloc_pages_node() + pin()`. The driver doesn't have to > worry about something like SLAB accidentally using that same memory. > > All without having to rewrite a bunch of mm/ in a driver, and with > having to do some kind of heroics with ZONE_DEVICE that causes even > more special mm/ interactions. > > That only comes from adding an isolation primitive to a NUMA node. > > > > > Isolating private node folios from kernel services > > > ================================================== > > > We implement filter points in mm/ to prevent operations on > > > private node memory. Where possible, we even re-use existing > > > filter points from ZONE_DEVICE. > > > > > > Most filter points are one or two lines of code: > > > > > > Combining ZONE_DEVICE and N_MEMORY_PRIVATE opt-out spots: > > > - if (folio_is_zone_device(folio)) > > > + if (unlikely(folio_is_private_managed(folio))) > > > > > > Disabling a service: > > > + if (!node_is_private(nid)) { > > > + kswapd_run(nid); > > > + kcompactd_run(nid); > > > + } > > > > > > Disallowing a uapi interaction: > > > + if (node_state(nid, N_MEMORY_PRIVATE)) > > > + return -EINVAL; > > > > > > > I'm not sure of why do we re-implement more filter and basically > > doing the same thing ? Any unavoidable scenario ? > > > > re-using the exisintg filter would be nice if that's possible. > > > > We re-use (combine) existing filters where possible. > > We add new ones where ZONE_DEVICE did not implement filters due to some > *implicit* filter already existing. > > Two clear examples: > - reclaim does not target ZONE_DEVICE > - ZONE_DEVICE does not support PMD > > Both cases result in private-node filters that otherwise would have to > exist for ZONE_DEVICE (and if ZONE_DEVICE ever grows PMD support, those > filters will have to be added). > > > > Bonus Configuration: HBM device memory tiering > > > ============================================== > > > echo 1 > dax0.0/private # make the node private > > > echo 0 > dax0.0/adistance # highest tier > > > echo 1 > dax0.0/reclaim # reclaim active > > > echo 1 > dax0.0/demotion # may demote from the node > > > echo 1 > dax0.0/user_numa # mbind() > > > echo online_movable > dax0.0/state > > > echo 1 > numa/demotion_enabled > > > > > > This is an HBM device which is treated as the top-tier in the > > > system but for which memory can only enter via explicit mbind(). > > > > > > It can be overcommitted because it can be reclaimed (demotions > > > go to CPU DRAM, and reclaim can swap from it). > > > > > > If the HBM is managed by an accelerator (GPU), the mmu_notifier > > > allows it to know when reclaim is moving memory out to do > > > device-mmu invalidation prior to migration. > > > > > > Prereqs, base commit, references > > > ================================ > > > akpm/mm-new - for Brendan Jackman's mm/page_alloc.h work[3] > > > > > > > Still thanks for the work, I learn alot from your work as well, thanks. > > > > Of course, and thank you for reading. > > If nothing else I hope this series helps folks understand the page > allocator better (it certainly has helped me). If there is anything > you think I can improve or better explain, I am happy to discuss. > > I am planning a larger publication of all my research sometime in the > future, but I think code is more impactful and useful - so I am > prioritizing that for now :]. > > ~Gregory That would be helpful for me, thanks alot for all the explanation again ! --Richard

