On Thu, Oct 1, 2026 at 8:00 PM Borislav Petkov <[email protected]> wrote: > > On Tue, Sep 29, 2026 at 01:25:09PM -0400, Zack Rusin wrote: > > Partly. So steal time tells Linux how much time a virtual CPU spent > > ready to run but waiting for the hypervisor to schedule it. > > > > For example, during a 100 ms interval, the vCPU might execute for 70 > > ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps > > kernel: > > - account for CPU usage accurately: avoid charging applications for > > time when the hypervisor wasn't running their vCPU. > > - make fairer scheduling decisions: base task execution accounting on > > the CPU time tasks actually received. > > - expose host contention: the st field in top and counters in > > /proc/stat help explain why a VM is slow even though its applications > > don't appear to consume all available CPU time. > > Aaaha, IOW, that's the "st" column here: > > $ vmstat > procs -----------memory---------- ---swap-- -----io---- -system-- > -------cpu------- > r b swpd free buff cache si so bi bo in cs us sy id wa > st gu > 1 0 0 1714852 9748 85300 0 0 3793 51 1445 0 0 1 99 > 0 0 0 > > In any case, you could keep that helpful explanation in yout 0th message. :)
Yeah, that's the one. Sounds good, and thanks for the links to the tip and submitting-patches docs, I'll follow them for v3. > > If options are binary then "no" :) It's not for monitoring tools, it's > > for the kernel's own steal-time accounting above, which currently > > doesn't work in confidential VMware guests. So the issue is that in a > > confidential guest (AMD SEV SNP, Intel TDX) memory is private by > > default, so the hypervisor can't read or write it. Any page the > > hypervisor has to write must first be converted to shared by the > > guest. Linux gives us the address of the steal-time buffer without > > converting it. The hypervisor then tries to write into a private page, > > which doesn't work. Currently ESXi deliberately powers off the VM when > > that happens. That only affects VMs with the steal clock enabled, > > which ESXi leaves off by default except for Photon guests. I'll > > probably change ESXi to just disable steal time when a guest gives us > > a private page, but that only avoids the power-off; steal time still > > won't work in these guests without this series. > > > > KVM has the same need for three of its per-CPU buffers (steal time, > > async page faults and PV EOI). It already converts them, but only on > > AMD and in its own loop. Kiryl asked on v1 for one common place that > > converts all such per CPU buffers early, on both AMD and Intel, > > Right, why early? Because on SMP kernels both KVM and VMware hand the boot CPU's buffer address to the hypervisor in smp_prepare_boot_cpu(), right after setup_per_cpu_areas() and well before mm_core_init(). Sharing a page usually means splitting the direct map around it, and set_memory_decrypted() needs the page allocator for the new page table. To keep that existing registration point, the common code needed its own early sharing path, which is most of what patches 4 and 5 are. > I mean, I am still trying to see the justification for this diffstat > > 14 files changed, 247 insertions(+), 51 deletions(-) > > and whether it is really worth it. For VMware alone it isn't. But you guys have to decide based on what your long-term plan here is. I'd be happiest if we stayed within the confines of the VMware code. That's what v1 of this did (besides the small guard fix outside) and I'll only stray from this if I'm explicitly told that a patch/series won't be accepted without some cleanups/re-architecture somewhere else. I'm here because internally we've had multiple teams needing various bugs/features in the kernel and looking at some parts of our Linux kernel code was making me sad so I decided to collect the changes and use this opportunity to also cleanup some of our code. (I have three other series in my tree, they all touch arch/x86/kernel/cpu/vmware.c so I'm trying to push them out in a semi-coherent order.) But this is very much a spare time thing. I can guarantee I'll find enough time to always review the VMware bits, maintain them and answer questions about them, but I can't do the same for other parts of the kernel. So the more focused we can keep those patches the easier this is for me. > > instead of each hypervisor driver doing it. That's patches 2-6. > > > > Patch 1 fixes an old layout bug in uniprocessor kernels, where these > > buffers can share a page with unrelated data. The series also fixes > > SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n, > > which currently hang at boot (we reproduced the hang). > > That should tell you how much we care about UP. We would even love to make SMP > the default. hah :) The UP handling was a good chunk of the complexity (patch 1, the kernel-image aliases and their kexec teardown), and the simpler approach below doesn't need any of it. > In talking to Peter about it, we were wondering whether this can be made > simpler. Like do not touch perCPU but do a normal page for each CPU's steal > time gunk and thus do not split the large page and then that early > enc/decrypting of memory I don't like either. > > Perhaps we should start with the simplest approach first. Sounds good to me, that's pretty much the smaller fix I'd offered on v1. v3 would then be VMware only: if the steal clock is available, allocate the steal-time structs for all possible CPUs once the page allocator is up, share them with set_memory_decrypted(), and register the boot CPU from an early initcall. The existing hotplug callback registers the other CPUs as they come up. The memory comes from the page allocator, so it's in the direct map and TDX's vmalloc problem doesn't come up. No percpu, linker-script or early-conversion changes, and KVM stays as it is. This would also move registration for unencrypted guests to the initcall. One small tweak: I'd pack all CPUs into one allocation (64 bytes each) rather than use a page per CPU. Sharing a page can still split the large direct map mapping around it, so a page per CPU could split one large mapping per CPU, while a single allocation concentrates those splits in that one allocation. And that only happens in confidential guests with the steal clock enabled. Kiryl, I know this is the opposite of what you asked for on v1, but it sounds like Boris and Peter would rather start with the VMware-only fix. Shout if you object. z
smime.p7s
Description: S/MIME Cryptographic Signature

