Sean Christopherson <[email protected]> writes: > On Tue, Aug 11, 2026, Pratyush Yadav wrote: >> On Mon, Aug 10 2026, Sean Christopherson wrote: >> >> > On Tue, Jul 28, 2026, Tarun Sahu wrote: >> >> Register a Live Update Orchestrator (LUO) file handler for KVM VM files >> >> to serialize and deserialize VM state across kexec live updates. >> >> >> >> Currently, Only VM type (e.g. arch.vm_type on x86) is preserved as part >> >> of VM preservation. >> > >> > Why? >> > >> >> On retrieval, kvm_luo_retrieve() recreates the KVM VM file via >> >> kvm_create_vm_file() and use an atomically incremented ID for the internal >> >> fdname, as the final fdname assigned by userspace is not yet known during >> >> retrieval. As this fdname is only used in debugfs infra, This will not >> >> break >> >> any UAPI. >> >> >> >> This infrastructure establishes the foundation for preserving guest_memfd >> >> instances across live updates, and can be expanded in the future to >> >> preserve additional VM state. >> > >> > Uh, why guest_memfd? As much as I want to push guest_memfd adoption, it >> > seems >> > guest_memfd should be the _last_ thing we support, not the first. As >> > evidenced >> > by the last two decades, it's very doable to have KVM VMs without >> > guest_memfd, >> > but it's rather hard to have VMs without vCPUs. >> >> You _can_ preserve vCPUs today using KVM_{GET,SET}_REGS, they just won't >> run in the background during the reboot. > > What about x86 CoCo VMs? Which are quite literally _the_ reason guest_memfd > was > created in the first place. > >> This series can save you from dumping VM memory to disk if it is backed by >> guest_memfd. > > Or to word it another way, one _can_ save guest_memfd, it's just > slower. > > > My point is that this series needs to provide a _lot_ more information about > the > bigger KVM picture.
Yes, I agree. I will try to layout the plan. End goal of this series is to preserve VM memory which is backed by guest_memfd. Not all VM are backed by guest_memfd. So preservation of KVM (vm_file) is independent of preservation of guest_memfd. But guest_memefd can not be preserved without preserving the KVM (vm_file). So I agree to your suggestion: diff --git virt/kvm/Makefile.kvm virt/kvm/Makefile.kvm index d047d4cf58c9..e6f098498795 100644 --- virt/kvm/Makefile.kvm +++ virt/kvm/Makefile.kvm @@ -13,3 +13,8 @@ kvm-$(CONFIG_HAVE_KVM_IRQ_ROUTING) += $(KVM)/irqchip.o kvm-$(CONFIG_HAVE_KVM_DIRTY_RING) += $(KVM)/dirty_ring.o kvm-$(CONFIG_HAVE_KVM_PFNCACHE) += $(KVM)/pfncache.o kvm-$(CONFIG_KVM_GUEST_MEMFD) += $(KVM)/guest_memfd.o + +ifdef CONFIG_LIVEUPDATE +kvm-y += $(KVM)/kvm_luo.o +kvm-$(CONFIG_KVM_GUEST_MEMFD) += $(KVM)/guest_memfd_luo.o +endif Guest_memfd can't be created alone without struct kvm (kvm_gmem_create() and KVM_GMEM_CREATE ioctl). This is how guest_memfd has been designed. I don't want to break this design which has been accepted upstream after lot of discussion. So while LUO preserve guest_memfd's data, It just preserve the PFNs value, few flags that belongs to the guest_memfd. On restore side in new kernel, These PFNs, flags will be populated to a _newly_ created guest_memfd. that is how preservation and retrieval works. So To create this new guest_memfd, We need struct kvm (vm_file). this vm_file must be the same VM which had this guest_memfd in old kernel, So we preserve the vm_file, get its TOKEN preserve with guest_memfd and during restore, we create the guest_memfd with the same VM (vm_file/struct kvm). I agree, I did not do good job explaining things in commit message. I will make sure to update them in next revision. > For those of us that are on the very fringes of live update, > it's practically impossible to review because, to us, it seems very arbitrary. > > The part that's especially confusing is the saving of the VM type. That comes > straight from userspace, so it's super bizarre to automatically save/restore > that, > but nothing else. Like, guest_memfd needs struct kvm to create itself. struct kvm (vm_file) needs vm_type to create itself (kvm_create_vm() or KVM_CREATE_VM IOCTL). LUO does not provde functionality to pass any subsystem specific arguments. So vm_type needs to be preserved even though, userspace is aware about it. So Why do we preserve only vm_type: To keep things simple for this series, As target is guest_memfd. Currently I dont have discreet plan on what else will ,in future, be needed to be preserved. Which I agree not a absoulute right way to approach. I will layout a rough plan on KVM side preservation. Having this need for backward compatiblity: I responded here: https://lore.kernel.org/all/[email protected]/ > >> >> +KVM LIVE UPDATE >> >> +M: Pasha Tatashin <[email protected]> >> >> +M: Mike Rapoport <[email protected]> >> >> +M: Pratyush Yadav <[email protected]> >> >> +R: Tarun Sahu <[email protected]> >> >> +L: [email protected] >> >> +L: [email protected] >> >> +S: Maintained >> >> +T: git >> >> git://git.kernel.org/pub/scm/linux/kernel/git/liveupdate/linux.git >> > >> > NAK on taking changes through a different tree. This is KVM code, period. >> > >> > In general, I'm skeptical of the dedicated MAINTAINERS entry. It's >> > extremely >> > difficult to tell since this series is little more than a skeleton (either >> > that >> > or liveupdate is way simpler that I was expecting), but I suspect that >> > maintaining > > ... > >> > E.g. the LUO APIs seem pretty straightforward; I assume the bulk of the >> > complexity >> > is going to be in knowing what to save/restore, and how, which is much >> > more about >> > KVM than it is about liveupdate. >> >> I think it is fine if you want to take these changes through the KVM >> tree, but I would like live update maintainers to be listed as reviewers >> at least. > > Why not simply add a file pattern match to the LIVE UPDATE entry? > > diff --git MAINTAINERS MAINTAINERS > index 8014b9f8253e..2eb57b22c37f 100644 > --- MAINTAINERS > +++ MAINTAINERS > @@ -15052,8 +15052,8 @@ F: include/linux/liveupdate.h > F: include/uapi/linux/liveupdate.h > F: kernel/liveupdate/ > F: lib/tests/liveupdate.c > -F: mm/memfd_luo.c > F: tools/testing/selftests/liveupdate/ > +N: [^a-z]luo > > LLC (802.2) > L: [email protected] > >> At the same time, I also keep being (pleasantly) >> surprised at preservation being relatively simple. For example, the code >> to preserve a shmem file (via memfd) is roughly 600 lines, a big chunk >> of which is comments. The code of course has some limitations, but it is >> good enough for use in production. >> >> For one, we care about ABI breakages and versioning. > > Which is amusing to me because that implies KVM does not, and I would hazard > to > guess that KVM has the biggest ABI surface of any subsystem in the kernel by a > country mile (though I'm probably wildly underestimating the effective ABI > surface > of filesystems). > >> The serialized state is a part of live update ABI and changes to it should be >> ACKed by us. > > Meh, "Don't break userspace" is a universal rule in the kernel, I genuinely > don't > see why liveupdate needs special treatment. > >> For another, how the file handlers interact with their dependencies can >> affect the behaviour that VMMs observe. Those changes should also pass by >> some live update eyes. > > Perhaps in the short term, but IMO, that's not a winning strategy in the long > term. From my perspective, that like saying the PAGE CACHE maintainers should > review every usage of the filemap APIs, because how the APIs are used impacts > the page cache and affects userspace-visible behavior. There are myriad > analogies > like that throughout the kernel. > > Yes, liveupdate is new and shiny, but IMO for it to be successful and > maintainable, > it needs to be treated like any other core infrastructure in the kernel, not a > special snowflake whose details are known only by a handful of people. > Because > I think it's likely liveupdate goes one of two ways: either liveupdate > becomes a > very niche thing that is used sparingly throughout the kernel, or it becomes a > broadly used feature that is supported by many filesystems and subsystems. > > If liveupdate is relegated to niche status, then it probably isn't going to > see > a significant amount of ongoing development, at which point the folks working > on > liveupdate will naturally migrate to other projects, and maintenance will > largely > be left to subsystem maintainers. > > If liveupdate is broadly used, then having a single group of people maintain > every subsystem's usage won't scale, and maintenance will again largely fall > on > the shoulder of subsystem maintainers. Which is totally fine and working as > intended, because that's exactly what subystem maintainers are signing up for > by merging support for liveupdate.

