在 2026/7/10 16:44, Bibo Mao 写道:
> 
> 
> On 2026/7/10 下午4:09, Tao Cui wrote:
>>
>>
>> 在 2026/7/10 10:58, Bibo Mao 写道:
>>>
>>>
>>> On 2026/7/10 上午8:50, Tao Cui wrote:
>>>> From: Tao Cui <[email protected]>
>>>>
>>>> kvm_set_pv_features() programs the KVM_FEATURE cpucfg attribute, which is a
>>>> per-vCPU setting. It was called from kvm_arch_put_registers() under a
>>>> function-local static guard, so it ran only once for the whole VM: only the
>>>> first vCPU got its pv features pushed to KVM, and on SMP guests the others
>>>> never saw KVM_FEATURE_IPI / KVM_FEATURE_STEAL_TIME.
>>>>
>>>> Drop the static guard and push pv features per vCPU under
>>>> KVM_PUT_FULL_STATE, the same gate kvm_set_stealtime() already uses. Host
>>>> feature detection stays in kvm_arch_init_vcpu(); the per-vCPU state write
>>>> belongs in kvm_arch_put_registers().
>>>>
>>>> Signed-off-by: Tao Cui <[email protected]>
>>>> ---
>>>>    target/loongarch/kvm/kvm.c | 15 ++++++---------
>>>>    1 file changed, 6 insertions(+), 9 deletions(-)
>>>>
>>>> diff --git a/target/loongarch/kvm/kvm.c b/target/loongarch/kvm/kvm.c
>>>> index d6539c12ac..c557ee3c3d 100644
>>>> --- a/target/loongarch/kvm/kvm.c
>>>> +++ b/target/loongarch/kvm/kvm.c
>>>> @@ -816,7 +816,6 @@ int kvm_arch_get_registers(CPUState *cs, Error **errp)
>>>>    int kvm_arch_put_registers(CPUState *cs, KvmPutState level, Error 
>>>> **errp)
>>>>    {
>>>>        int ret;
>>>> -    static int once;
>>>>          ret = kvm_loongarch_put_regs_core(cs);
>>>>        if (ret) {
>>>> @@ -843,19 +842,17 @@ int kvm_arch_put_registers(CPUState *cs, KvmPutState 
>>>> level, Error **errp)
>>>>            return ret;
>>>>        }
>>>>    -    if (!once) {
>>>> +    if (level >= KVM_PUT_FULL_STATE) {
>>>> +        /*
>>>> +         * pv_features and steal time are per-vCPU state. Push them on
>>>> +         * full-state sync so every vCPU gets its own settings; the kernel
>>>> +         * clears the steal-time guest_addr on KVM_PUT_RESET_STATE.
>>>> +         */
>>>>            ret = kvm_set_pv_features(cs);
>>> pv feature is a little different from steal-time. steal-time guest_addr is 
>>> created from guest OS, pv feature is created from VMM at beginning. 
>>> steal-time guest_addr can be set for many times, and there is bit 
>>> KVM_STEAL_PHYS_VALID checking with steal-time guest_addr, however pv 
>>> feature can be set only once with existing method.
>>>
>>
>> Hi Bibo,
>>
>> Thanks for catching this — I hadn't fully considered the double-call path.
>>
>> KVM_PUT_FULL_STATE fires both at realize (cpu_synchronize_post_init) and
>> on incoming migration load (cpu_synchronize_all_post_init), so
>> kvm_set_pv_features() runs twice on the destination.
>>
>>> Although I do not understand flow of VM migration, with KVM_PUT_FULL_STATE 
>>> state changing, there are at least two places where this state is set, one 
>>> is from cpu_common_realizefn() which calls cpu_synchronize_post_init(), the 
>>> other is qemu_loadvm_state()/qemu_loadvm_state_main() which calls 
>>> cpu_synchronize_all_post_init().
>>>
>>> It seems that VM will fail to migrate since kvm_set_pv_features is called 
>>> twice at least here. Do you test VM migration with this patch?
>>
>> I did test migration (virt-11.2 -> virt-11.2, virt-11.1 -> virt-11.1) and
>> it passed, but that was on a single host — source and destination computed
>> the same pv_features. The kernel (kvm_loongarch_cpucfg_set_attr) only
>> rejects a re-set when the value differs:
>>
>>      if ((kvm->arch.pv_features & LOONGARCH_PV_FEAT_UPDATED) &&
>>          ((kvm->arch.pv_features & valid) != val))
>>          return -EINVAL;
>>
>> So a cross-host migration where the two sides compute different pv_features
>> would indeed fail on the second set.
>>
>> I'll add a per-vCPU guard so the push happens exactly once (at the first
>> FULL_STATE sync); subsequent syncs are skipped and the destination keeps
>> advertising the features its own host supports. Does that sound like the
>> right direction?
> I do not consider pv feature migration supports before. There are two points 
> from my side when VM migrates from hostA to hostB, maybe there are other 
> requirements out of my knowledge :(
>   1. the pv feature value got from running vCPUs on hostA should be supported 
> on hostB.
>   2. if VM is rebooted or one vCPU is hot added, the pv feature value got 
> from new added vCPU on hostB should be the same with boot CPU0, which is got 
> from hostA.
> 

Thanks, those are good points — my set-once guard doesn't handle them
correctly.

For Point 2, the root issue is that kvm->arch.pv_features is VM-level:
the first vCPU locks the value, and all others (including hot-added) must
match. So after migrating hostA -> hostB, a hot-added vCPU on hostB must
still use hostA's value, not re-compute from hostB.

I've been thinking about how to handle the push. The challenge is that
there's no single QEMU hook that fires exactly once with the final
pv_features value for both fresh boot and migration:
  - post_load only fires on migration load, not on fresh boot;
  - KVM_PUT_FULL_STATE fires at both realize and load (the double-call).

One option that seems to work: push from the first KVM_PUT_RUNTIME_STATE
sync (when the guest first runs), gated by a per-vCPU flag. By that point
env->pv_features holds the correct value for both cases — host-detected on
fresh boot, or the migrated value after post_load. The flag ensures the
push happens only once per vCPU.

This still doesn't solve hot-add after cross-host migration: a newly
created vCPU would re-compute from hostB via check_pv_features, getting a
different value from CPU0. For that I think pv_features needs to be stored
at the VM level (matching kvm->arch.pv_features), so hot-added vCPUs
inherit it rather than re-detecting.

I've also checked a few other migration paths: savevm/snapshot restore
goes through the same post_load, and multi-hop migration keeps the
original host's value — the runtime-state push should cover both.
During testing I observed that cross-machine-type migration (virt-11.1 <->
virt-11.2) fails at QEMU's configuration check; I haven't investigated
whether pv_features adds further constraints there.

I also looked at how other architectures handle this. On x86, CPUID
(including KVM feature leaves) is set via KVM_SET_CPUID2 in
kvm_arch_init_vcpu(), not in put_registers(); the kernel allows
re-setting CPUID and validates the features against the host on each SET.
On arm, KVM_ARM_VCPU_INIT and device attributes (PMU, PVTIME IPA, etc.)
are set in init_vcpu / machine init; the kernel likewise permits
re-setting device attributes. Neither architecture pushes feature
configuration through put_registers.

On LoongArch, the kernel currently locks pv_features after the first SET
and rejects any different value. This is what makes the migration case
hard — the destination cannot override the host-detected value with the
migrated one. If the kernel could relax this (allow overriding while
validating that the new value is a subset of what the host supports,
similar to x86's KVM_SET_CPUID2), the QEMU side would be simpler: push
once in init_vcpu (host-detected), then override from post_load on
migration. I think adjusting both QEMU and KVM together would be simpler
than working around it entirely in QEMU — would that work?

I'm still going through the overall flow to check if there are other
edge cases I've missed.

Tao

> Regards
> Bibo Mao
>>
>> Thanks,
>> Tao
>>
>>>
>>> Regards
>>> Bibo Mao
>>>>            if (ret) {
>>>>                return ret;
>>>>            }
>>>> -        once = 1;
>>>> -    }
>>>>    -    if (level >= KVM_PUT_FULL_STATE) {
>>>> -        /*
>>>> -         * only KVM_PUT_FULL_STATE is required, kvm kernel will clear
>>>> -         * guest_addr for KVM_PUT_RESET_STATE
>>>> -         */
>>>>            ret = kvm_set_stealtime(cs);
>>>>            if (ret) {
>>>>                return ret;
>>>>
>>>
> 


Reply via email to