"Lorenzo Stoakes (ARM)" <[email protected]> writes:

> On Tue, Sep 22, 2026 at 06:37:04PM +0530, Aneesh Kumar K.V wrote:
>> "Lorenzo Stoakes (ARM)" <[email protected]> writes:
>>
>> > On Tue, Sep 22, 2026 at 03:28:33PM +0530, Aneesh Kumar K.V wrote:
>> >> > This is because pKVM instantiates vCPUs upon run,
>> >> >
>> >>
>> >> Can pKVM instantiate the hyp vCPU during pre-faulting ?
>> >
>> > It would be an unusual and unexpected thing to do - suddenly a pre-fault
>> > operation is initialising a vCPU explicitly for pKVM.
>> >
>> > A caller is not going to reasonably expect this and might treat a failure 
>> > to
>> > pre-alloc as fine to carry on whereas in fact it was a failure to 
>> > initailised a
>> > pKVM vCPU.
>> >
>> > It'd also require significant changes to how pKVM is set up, right now it's
>> > hardcoded to be done unconditionally at run via 
>> > kvm_arch_vcpu_run_pid_change()
>> > -> pkvm_create_hyp_vcpu(), so all that would have to change and be checked 
>> > and
>> > tested and... that'd be really out of scope I think :)
>> >
>> > And pre-faulting really makes most sense BEFORE you run a VM. It doesn't 
>> > make so
>> > much sense mid-run.
>> >
>> > But more fundamentally, the stage 2 page tables, as I understand it, are 
>> > owned
>> > by pKVM and so aren't really available to be pre-faulted.
>> >
>> > Maybe unprotected-under-pKVM VMs but then it's questionable as to how 
>> > useful
>> > that would be given that it would be confusing to users vs. how it works 
>> > for
>> > other VMs.
>> >
>> > So in general, no I don't think it's a good idea.
>> >
>> > And even if we wanted to pursue some version of this, it's _definitely_ out
>> > of scope for the initial pre-faulting bring-up series.
>> >
>> >>
>> >>
>> >> > but pre-faulting is  typically performed before a vCPU is run.  It 
>> >> > would be confusing and
>> >> > inconsistent to error out on non-running vCPUs but to pre-fault running
>> >> > ones.
>> >>
>> >>
>> >> I use KVM pre-faulting when transitioning pages from shared to private
>> >
>> > You mean you'd prefer to use? Or you are using it on another arch?
>> >
>> >> with CoCo guest. This ensures that a trusted device can DMA to private
>> >> memory before the guest accesses it.
>> >
>> > Hm what do you mean by private memory?
>> >
>> > I see:
>> >
>> > #ifndef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES
>> > static inline bool kvm_arch_has_private_mem(struct kvm *kvm)
>> > {
>> >    return false;
>> > }
>> > #endif
>> >
>> > And only x86 selects KVM_GENERIC_MEMORY_ATTRIBUTES?
>> >
>> > Do you mean something else?
>> >
>>
>> I am using this with ARM CCA-DA, based on the patch series from Jack Thomson 
>> <[email protected]>.
>>
>> https://gitlab.arm.com/linux-arm/kvmtool-cca/-/commit/80e7aad61c5639de2f0cb4a5525dad0c96156428
>>
>> We do this while the VM is running.
>
> Right, that's a non-mainline kernel I guess? Which presumably implements 
> private memory.
>
> Jack himself experienced a panic with his pKVM code, so the code you're using 
> is
> not upstreamable, unfortunately. And he'd already shelved pKVM support AFAICT.
>
> And reviewers pointed out actually implementing the pKVM stuff properly would 
> be
> quite involved, even if you wanted to do that (hence follow-up).
>
> Also you end up stuck with the same problems as I mentioned above - you can't
> sanely bring the vCPU pre-run, so now you have extremely weird behaviour - 
> only
> pre-faults if vCPU initialised, running, and unprotected pKVM.
>

IIUC, pkvm_pgtable_stage2_map() only uses the hyp vCPU's pKVM memcache
(&vcpu->vcpu.arch.pkvm_memcache). I agree that this does not need to be
addressed in this series. However, there is also a desire to keep the CCA
and pKVM code paths similar by using helpers such as kvm_vm_is_protected().
Since an RMM can create a stage-2 mapping without a REC (vCPU), rejecting
pre-faulting for all protected VM configurations may be incorrect. I can
handle this in my series though.

I also want to point out that there are use cases for pre-faulting while
the VM is running.

>
> Protected pKVM support is a whole other layer of complexity and it's not 
> obvious
> that you're really achieving what pre-fault is supposed to.
>
> In any case Oliver literally just asked me to _simplify_ weird edge cases for
> this series :) so I am not sure something like that is going to be accepted.
>
> Are you sure you're actually running in pKVM mode btw? CCA doesn't AFAICT? In
> which case this series _should_ work fine for you.
>


It is not pKVM; it runs in Realm mode.

>
> Anyway, if we really do need to add something for pKVM it needs to be a follow
> up. Let's get the basics working first :)
>

sure.

> (Note that kvmtool will need to be updated to retry pre-fault on -EAGAIN, 
> -EINTR
>  as this series can, albeit unlikely, return -EAGAIN.)
>

I will check this when I rebase my kernel onto this series.

-aneesh

Reply via email to