enter_smm() clears CR0.PG and EFER.LMA/LME but leaves CR3 as-is, so a
64-bit guest whose page tables live above 4GiB enters SMM with
CR3[63:32] != 0 while outside of long mode.  Bare-metal SVM accepts that
state, but Hyper-V's emulation of VMRUN for a nested hypervisor rejects
it as an invalid VMCB, and the vCPU dies on the very first instruction
of the SMI handler:

  KVM: entry failed, hardware error 0xffffffff
  EIP=00008000 EFL=00000002 [-------] CPL=0 II=0 A20=1 SMM=1 HLT=0
  CS =f900 7bff9000 ffffffff 00809300
  CR0=00050032 CR2=2a91044a CR3=77681000 CR4=00000000
  EFER=0000000000000000

QEMU prints only CR3[31:0] here; the CR3 saved in SMRAM for this vCPU
was 0x277681000.  All vCPUs that failed had a CR3 above 4GiB, while the
one vCPU whose CR3 was below 4GiB entered SMM without issue.

This reproduces reliably when booting a Windows 11 guest (8GiB of RAM)
with Secure Boot OVMF, i.e. with SMM enabled, in KVM on WSL2 on an AMD
host, with Windows 11 25H2 build 26200.9457, the latest released build.

With paging disabled outside of long mode, only CR3[31:0] is
reachable: a MOV to CR3 can only write 32 bits, and the SMI handler
must load its own CR3 before enabling paging.  RSM restores the full
CR3 from the SMRAM state-save area, which was written before this point.
Clear the upper 32 bits when KVM runs on Hyper-V, so that the SMM entry
state passes Hyper-V's VMRUN consistency checks.  Leave the behavior on
other hosts unchanged, as the problem belongs to Hyper-V and should be
fixed there.

Signed-off-by: Qiliang Yuan <[email protected]>
---
V1 -> V2:
- Only clear CR3[63:32] when KVM runs on Hyper-V, leave other hosts
  unchanged (Sean)
- Note in the changelog that the latest released Windows 11 build
  (25H2, 26200.9457) is still affected
- Cc Hyper-V maintainers and linux-hyperv

v1: 
https://lore.kernel.org/r/[email protected]
---
 arch/x86/kvm/smm.c | 15 +++++++++++++++
 1 file changed, 15 insertions(+)

diff --git a/arch/x86/kvm/smm.c b/arch/x86/kvm/smm.c
index f623c5986119..34c16b8c2c0b 100644
--- a/arch/x86/kvm/smm.c
+++ b/arch/x86/kvm/smm.c
@@ -2,6 +2,7 @@
 #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
 
 #include <linux/kvm_host.h>
+#include <asm/hypervisor.h>
 #include "x86.h"
 #include "kvm_cache_regs.h"
 #include "kvm_emulate.h"
@@ -361,6 +362,20 @@ void enter_smm(struct kvm_vcpu *vcpu)
        if (guest_cpu_cap_has(vcpu, X86_FEATURE_LM))
                if (kvm_x86_call(set_efer)(vcpu, 0))
                        goto error;
+
+       /*
+        * Work around Hyper-V rejecting VMRUN for a nested hypervisor when
+        * CR3[63:32] != 0 with EFER.LMA=0, which is exactly the state left
+        * behind by entering SMM with page tables above 4GiB.  CR3 is
+        * unmodified on SMM entry, but with paging disabled and outside of
+        * long mode bits 63:32 are unreachable, and RSM restores the full
+        * value from SMRAM, so clearing them is invisible to the guest.
+        */
+       if (hypervisor_is_type(X86_HYPER_MS_HYPERV) &&
+           kvm_read_cr3(vcpu) >> 32) {
+               vcpu->arch.cr3 = (u32)vcpu->arch.cr3;
+               kvm_register_mark_dirty(vcpu, VCPU_EXREG_CR3);
+       }
 #endif
 
        vcpu->arch.cpuid_dynamic_bits_dirty = true;

---
base-commit: eb3f4b7426cfd2b79d65b7d37155480b32259a11
change-id: 20260928-kvm-smm-cr3-upper-bits-9819a1a2f337

Best regards,
-- 
Qiliang Yuan <[email protected]>


Reply via email to