On Tue, Sep 29, 2026 at 3:54 AM Peter Zijlstra <[email protected]> wrote: > On Tue, Sep 29, 2026 at 12:02:51AM -0400, Zack Rusin wrote: > > @@ -1022,11 +1024,13 @@ > > * Note: We use a separate section so that only this section gets > > * decrypted to avoid exposing more than we wish. > > */ > > -#ifdef CONFIG_AMD_MEM_ENCRYPT > > +#if defined(CONFIG_X86_MEM_ENCRYPT) && defined(CONFIG_SMP) > > #define PERCPU_DECRYPTED_SECTION \ > > . = ALIGN(PAGE_SIZE); \ > > + __start_percpu_decrypted = .; \ > > *(.data..percpu..decrypted) \ > > - . = ALIGN(PAGE_SIZE); > > + . = ALIGN(PAGE_SIZE); \ > > + __end_percpu_decrypted = .; > > #else > > #define PERCPU_DECRYPTED_SECTION > > #endif > > So you're page aligning something that will get different protection and > will thus shatter large pages? > > That is somewhat uncool. We like large pages, large pages are good.
I'm a huge large pages fan too. I have a poster in my office and
everything... And I never do things that are uncool. I have a wwcdd
(what would cool dude do) bracelet and I look at it every time I make
a decision, which given that I'm old gets on my wife's nerves a lot.
That is to say I'm handsome and innocent.
The alignment isn't new. It has been there since ac26963a1175
("percpu: Introduce DEFINE_PER_CPU_DECRYPTED") under AMD_MEM_ENCRYPT,
which common distro kernels enable. This patch extends it to TDX-only
builds and adds the boundary symbols. On bare metal and in ordinary
guests nothing changes protection, so no large pages are split.
In confidential guests, yes, converting each CPU's page splits the
direct map around it. The alignment doesn't cause that: a large
mapping can't be part private and part shared, so a converted page
can't stay inside one. The alignment only keeps unrelated per-CPU data
out of the shared page. Because that page sits in every CPU's per-CPU
unit, each 2M mapping that holds one of them becomes 4K mappings,
including the ordinary per-CPU data around it. In the layouts we've
tested that's most of the first per-CPU chunk.
SMP KVM SEV guests already pay this today. sev_map_percpu_data()
converts the same pages for every possible CPU, at the same
granularity. What v2 adds is TDX guests and guests on other
hypervisors.
Where you're right is that v2 converts in every encrypted guest apart
from Hyper-V vTOM, even when nothing uses the pages. VMware only needs
them when the steal clock is enabled.
If this is a deal breaker for you for v3 I could convert only when the
hypervisor's platform setup says it will use them: VMware when the
steal clock is available, KVM when async page faults, steal time or PV
EOI are enabled. That decision would be made before the conversion
(for KVM that means making it in kvm_init_platform() instead of
kvm_guest_init()), so there's still no readiness state, and other
confidential guests don't lose large pages to this.
If you'd also like to keep large pages while the steal clock is in
use, the alternative is to move these buffers out of the per-CPU area
into one tightly packed allocation. That concentrates the splits in
that one allocation, so the rest of the per-CPU area keeps its large
pages. Kiryl preferred the per-CPU section when we discussed v1, so
I'd like to hear from both of you before going that way. fwiw, I'd
vastly prefer to first fix the steal-time bug for VMware before trying
to do this though (whether it's with this series or a smaller
VMware-only fix). x86 kernel team is notoriously opinionated, which
isn't a bad thing, but it means that a rewrite of core x86 bits like
that will take a long time and it's starting to feel a little like I
stepped into something and you're all behind blast doors yelling "cut
all the wires you coward!!!" and I'm just confused going "I just came
here to get some blueberries".
z
smime.p7s
Description: S/MIME Cryptographic Signature

