On Tue, Sep 29, 2026 at 01:25:09PM -0400, Zack Rusin wrote: > Partly. So steal time tells Linux how much time a virtual CPU spent > ready to run but waiting for the hypervisor to schedule it. > > For example, during a 100 ms interval, the vCPU might execute for 70 > ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps > kernel: > - account for CPU usage accurately: avoid charging applications for > time when the hypervisor wasn't running their vCPU. > - make fairer scheduling decisions: base task execution accounting on > the CPU time tasks actually received. > - expose host contention: the st field in top and counters in > /proc/stat help explain why a VM is slow even though its applications > don't appear to consume all available CPU time.
Aaaha, IOW, that's the "st" column here: $ vmstat procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu------- r b swpd free buff cache si so bi bo in cs us sy id wa st gu 1 0 0 1714852 9748 85300 0 0 3793 51 1445 0 0 1 99 0 0 0 In any case, you could keep that helpful explanation in yout 0th message. :) > If options are binary then "no" :) It's not for monitoring tools, it's > for the kernel's own steal-time accounting above, which currently > doesn't work in confidential VMware guests. So the issue is that in a > confidential guest (AMD SEV SNP, Intel TDX) memory is private by > default, so the hypervisor can't read or write it. Any page the > hypervisor has to write must first be converted to shared by the > guest. Linux gives us the address of the steal-time buffer without > converting it. The hypervisor then tries to write into a private page, > which doesn't work. Currently ESXi deliberately powers off the VM when > that happens. That only affects VMs with the steal clock enabled, > which ESXi leaves off by default except for Photon guests. I'll > probably change ESXi to just disable steal time when a guest gives us > a private page, but that only avoids the power-off; steal time still > won't work in these guests without this series. > > KVM has the same need for three of its per-CPU buffers (steal time, > async page faults and PV EOI). It already converts them, but only on > AMD and in its own loop. Kiryl asked on v1 for one common place that > converts all such per CPU buffers early, on both AMD and Intel, Right, why early? I mean, I am still trying to see the justification for this diffstat 14 files changed, 247 insertions(+), 51 deletions(-) and whether it is really worth it. > instead of each hypervisor driver doing it. That's patches 2-6. > > Patch 1 fixes an old layout bug in uniprocessor kernels, where these > buffers can share a page with unrelated data. The series also fixes > SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n, > which currently hang at boot (we reproduced the hang). That should tell you how much we care about UP. We would even love to make SMP the default. > Fair enough. I'll rework the cover letter and the commit messages so > each one starts with the problem. The series originally wasn't really > touching x86 core parts and I haven't updated it for a larger crowd. > Would you like an explanation of steal time, like the above, in the > cover as well? Yes please. Also, we have some blurb about how to write those: https://docs.kernel.org/process/maintainer-tip.html#patch-subject and https://docs.kernel.org/process/submitting-patches.html In talking to Peter about it, we were wondering whether this can be made simpler. Like do not touch perCPU but do a normal page for each CPU's steal time gunk and thus do not split the large page and then that early enc/decrypting of memory I don't like either. Perhaps we should start with the simplest approach first. Thx. -- Regards/Gruss, Boris. https://people.kernel.org/tglx/notes-about-netiquette

