Hi, Boris.

On Tue, Sep 29, 2026 at 1:40 AM Borislav Petkov <[email protected]> wrote:
> On Tue, Sep 29, 2026 at 12:02:49AM -0400, Zack Rusin wrote:
> > VMware publishes its per-CPU steal-time GPA without first sharing the
> > storage in encrypted guests.
>
> So this one liner is the only explanation why this patchset exists, AFAICT. So
> I asked AI. I'm pasting what it said below, at the end.
>
> Is any of it true?

Partly. So steal time tells Linux how much time a virtual CPU spent
ready to run but waiting for the hypervisor to schedule it.

For example, during a 100 ms interval, the vCPU might execute for 70
ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps
kernel:
- account for CPU usage accurately: avoid charging applications for
time when the hypervisor wasn't running their vCPU.
- make fairer scheduling decisions: base task execution accounting on
the CPU time tasks actually received.
- expose host contention: the st field in top and counters in
/proc/stat help explain why a VM is slow even though its applications
don't appear to consume all available CPU time.

The benefit is greatest on busy or overcommitted hosts. On a host
where vCPUs rarely wait, steal time stays near zero.

To make this work the hypervisor writes a running total into a small
per-CPU buffer in guest memory, and the guest kernel reads it. More
below...

> If so, how much and is that the reason this patchset exists?
>
> I.e., you want for vmware monitoring tools to work with CoCo guests.
>
> Yes, no? Anything else?

If options are binary then "no" :) It's not for monitoring tools, it's
for the kernel's own steal-time accounting above, which currently
doesn't work in confidential VMware guests. So the issue is that in a
confidential guest (AMD SEV SNP, Intel TDX) memory is private by
default, so the hypervisor can't read or write it. Any page the
hypervisor has to write must first be converted to shared by the
guest. Linux gives us the address of the steal-time buffer without
converting it. The hypervisor then tries to write into a private page,
which doesn't work. Currently ESXi deliberately powers off the VM when
that happens. That only affects VMs with the steal clock enabled,
which ESXi leaves off by default except for Photon guests. I'll
probably change ESXi to just disable steal time when a guest gives us
a private page, but that only avoids the power-off; steal time still
won't work in these guests without this series.

KVM has the same need for three of its per-CPU buffers (steal time,
async page faults and PV EOI). It already converts them, but only on
AMD and in its own loop. Kiryl asked on v1 for one common place that
converts all such per CPU buffers early, on both AMD and Intel,
instead of each hypervisor driver doing it. That's patches 2-6.

Patch 1 fixes an old layout bug in uniprocessor kernels, where these
buffers can share a page with unrelated data. The series also fixes
SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n,
which currently hang at boot (we reproduced the hang).

> And looking at your commit messages, they don't really talk about why the
> patches exist.
>
> And I don't think you know your audience - you're sending a bunch of patches
> touching arch/x86/ and you throw all that virt gibberish around like it ain't
> no tomorrow and you're thinking that tip people can follow. And I think we can
> follow only bits and pieces but all of us will be mostly head-scratching here
> and probably ignore the whole set.
>
> So perhaps you should lay off your virt hat for a minute and try to explain in
> an approachable way what you're trying to do so that even non-virt people can
> follow.

Fair enough. I'll rework the cover letter and the commit messages so
each one starts with the problem. The series originally wasn't really
touching x86 core parts and I haven't updated it for a larger crowd.
Would you like an explanation of steal time, like the above, in the
cover as well?

z

Attachment: smime.p7s
Description: S/MIME Cryptographic Signature

Reply via email to