On Tue, Sep 29, 2026 at 01:25:09PM -0400, Zack Rusin wrote:
> Partly. So steal time tells Linux how much time a virtual CPU spent
> ready to run but waiting for the hypervisor to schedule it.
> 
> For example, during a 100 ms interval, the vCPU might execute for 70
> ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps
> kernel:
> - account for CPU usage accurately: avoid charging applications for
> time when the hypervisor wasn't running their vCPU.
> - make fairer scheduling decisions: base task execution accounting on
> the CPU time tasks actually received.
> - expose host contention: the st field in top and counters in
> /proc/stat help explain why a VM is slow even though its applications
> don't appear to consume all available CPU time.

Aaaha, IOW, that's the "st" column here:

$ vmstat
procs -----------memory---------- ---swap-- -----io---- -system-- 
-------cpu-------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa 
st gu
 1  0      0 1714852   9748  85300    0    0  3793    51 1445    0  0  1 99  0  
0  0

In any case, you could keep that helpful explanation in yout 0th message. :)

> If options are binary then "no" :) It's not for monitoring tools, it's
> for the kernel's own steal-time accounting above, which currently
> doesn't work in confidential VMware guests. So the issue is that in a
> confidential guest (AMD SEV SNP, Intel TDX) memory is private by
> default, so the hypervisor can't read or write it. Any page the
> hypervisor has to write must first be converted to shared by the
> guest. Linux gives us the address of the steal-time buffer without
> converting it. The hypervisor then tries to write into a private page,
> which doesn't work. Currently ESXi deliberately powers off the VM when
> that happens. That only affects VMs with the steal clock enabled,
> which ESXi leaves off by default except for Photon guests. I'll
> probably change ESXi to just disable steal time when a guest gives us
> a private page, but that only avoids the power-off; steal time still
> won't work in these guests without this series.
> 
> KVM has the same need for three of its per-CPU buffers (steal time,
> async page faults and PV EOI). It already converts them, but only on
> AMD and in its own loop. Kiryl asked on v1 for one common place that
> converts all such per CPU buffers early, on both AMD and Intel,

Right, why early?

I mean, I am still trying to see the justification for this diffstat

 14 files changed, 247 insertions(+), 51 deletions(-)

and whether it is really worth it.

> instead of each hypervisor driver doing it. That's patches 2-6.
> 
> Patch 1 fixes an old layout bug in uniprocessor kernels, where these
> buffers can share a page with unrelated data. The series also fixes
> SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n,
> which currently hang at boot (we reproduced the hang).

That should tell you how much we care about UP. We would even love to make SMP
the default.

> Fair enough. I'll rework the cover letter and the commit messages so
> each one starts with the problem. The series originally wasn't really
> touching x86 core parts and I haven't updated it for a larger crowd.
> Would you like an explanation of steal time, like the above, in the
> cover as well?

Yes please.

Also, we have some blurb about how to write those:

https://docs.kernel.org/process/maintainer-tip.html#patch-subject

and

https://docs.kernel.org/process/submitting-patches.html

In talking to Peter about it, we were wondering whether this can be made
simpler. Like do not touch perCPU but do a normal page for each CPU's steal
time gunk and thus do not split the large page and then that early
enc/decrypting of memory I don't like either.

Perhaps we should start with the simplest approach first.

Thx.

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette

Reply via email to