Hi, Boris. On Tue, Sep 29, 2026 at 1:40 AM Borislav Petkov <[email protected]> wrote: > On Tue, Sep 29, 2026 at 12:02:49AM -0400, Zack Rusin wrote: > > VMware publishes its per-CPU steal-time GPA without first sharing the > > storage in encrypted guests. > > So this one liner is the only explanation why this patchset exists, AFAICT. So > I asked AI. I'm pasting what it said below, at the end. > > Is any of it true?
Partly. So steal time tells Linux how much time a virtual CPU spent ready to run but waiting for the hypervisor to schedule it. For example, during a 100 ms interval, the vCPU might execute for 70 ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps kernel: - account for CPU usage accurately: avoid charging applications for time when the hypervisor wasn't running their vCPU. - make fairer scheduling decisions: base task execution accounting on the CPU time tasks actually received. - expose host contention: the st field in top and counters in /proc/stat help explain why a VM is slow even though its applications don't appear to consume all available CPU time. The benefit is greatest on busy or overcommitted hosts. On a host where vCPUs rarely wait, steal time stays near zero. To make this work the hypervisor writes a running total into a small per-CPU buffer in guest memory, and the guest kernel reads it. More below... > If so, how much and is that the reason this patchset exists? > > I.e., you want for vmware monitoring tools to work with CoCo guests. > > Yes, no? Anything else? If options are binary then "no" :) It's not for monitoring tools, it's for the kernel's own steal-time accounting above, which currently doesn't work in confidential VMware guests. So the issue is that in a confidential guest (AMD SEV SNP, Intel TDX) memory is private by default, so the hypervisor can't read or write it. Any page the hypervisor has to write must first be converted to shared by the guest. Linux gives us the address of the steal-time buffer without converting it. The hypervisor then tries to write into a private page, which doesn't work. Currently ESXi deliberately powers off the VM when that happens. That only affects VMs with the steal clock enabled, which ESXi leaves off by default except for Photon guests. I'll probably change ESXi to just disable steal time when a guest gives us a private page, but that only avoids the power-off; steal time still won't work in these guests without this series. KVM has the same need for three of its per-CPU buffers (steal time, async page faults and PV EOI). It already converts them, but only on AMD and in its own loop. Kiryl asked on v1 for one common place that converts all such per CPU buffers early, on both AMD and Intel, instead of each hypervisor driver doing it. That's patches 2-6. Patch 1 fixes an old layout bug in uniprocessor kernels, where these buffers can share a page with unrelated data. The series also fixes SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n, which currently hang at boot (we reproduced the hang). > And looking at your commit messages, they don't really talk about why the > patches exist. > > And I don't think you know your audience - you're sending a bunch of patches > touching arch/x86/ and you throw all that virt gibberish around like it ain't > no tomorrow and you're thinking that tip people can follow. And I think we can > follow only bits and pieces but all of us will be mostly head-scratching here > and probably ignore the whole set. > > So perhaps you should lay off your virt hat for a minute and try to explain in > an approachable way what you're trying to do so that even non-virt people can > follow. Fair enough. I'll rework the cover letter and the commit messages so each one starts with the problem. The series originally wasn't really touching x86 core parts and I haven't updated it for a larger crowd. Would you like an explanation of steal time, like the above, in the cover as well? z
smime.p7s
Description: S/MIME Cryptographic Signature

