On 9/11/26 16:23, Denis V. Lunev wrote:
> Restoring a big Windows guest spends most of its destination-side time in
> post_load hooks which do nothing but move memory regions around. Each one
> ends a memory transaction, and a transaction commit re-renders every
> flatview it touches at a cost which grows with the number of regions in
> the machine. A hook which runs once per vCPU therefore pays that render
> once per vCPU, and the machine gets slower to migrate the bigger it is.
>
> The Hyper-V SynIC is the case that hurts: restoring the synthetic
> interrupt controller maps a message page and an event page per vCPU, so a
> 64-vCPU guest forces 128 remaps, each with its own rebuild, while the rest
> of the stream is still being read.
>
> Patch 1 adds post_load_deferrable. A vmsd which sets it has its hook
> queued during the load and run once the stream has been consumed, in the
> order the hooks would have fired, with the whole drain sharing one memory
> transaction. Patch 2 sets it on the SynIC subsection.
>
> It is opt-in rather than automatic, and the three preconditions are
> spelled out on the field: the hook must not fail, nothing later in the
> load may depend on what it does, and it must not read guest memory or
> resolve an address space. A hook which breaks the first is fatal rather
> than silently reported, because by drain time the source may already have
> been told the migration succeeded.
>
> Deferring is not free in general, which is the other reason it is opt-in.
> Deferring the APIC post_load, whose cost is a synchronous run_on_cpu per
> vCPU rather than a memory remap, moves 3 ms out of the section walk and
> pays about 9 ms of drain for it. Deferral helps a hook which repeats
> topology work; it makes a hook which does cross-thread work worse.
>
> Measurements
> ------------
>
> Destination-side non-iterable load, ie. the sum of vmstate_downtime_load
> over non-iterable sections, on a guest which has actually programmed its
> Hyper-V state. Five interleaved rounds per point on an otherwise idle
> host, twice; medians, with the spread across all ten rounds.
>
>   upstream                          377 ms   (363-400)
>   + pci mapping transactions        247 ms   (242-255)
>   + this series                      96 ms   (86-97)
>
> Two postings against one problem, so the whole ladder is shown. The first
> step is a pci pair which batches a device's BAR and bridge window updates
> into a single transaction, posted separately and now queued in Michael's
> tree:
>
>   https://lore.kernel.org/qemu-devel/[email protected]/
>
> Those two are listed because they change what a rebuild costs, and so
> change what this series is worth. Together the postings take the load
> from 377 ms to 96 ms; this series is the 247 ms to 96 ms step.
>
> Where it goes: the cpu sections fall from 154 ms to 1.3 ms. The deferred
> hooks themselves cost 77 us at the drain, so the work is removed rather
> than moved somewhere the per-section metric cannot see.
>
> The saving scales with vCPU count, since that is how many times the remap
> repeats, and with the number of memory regions in the machine, since that
> is what a rebuild costs.
>
> Guest under test
> ----------------
>
> Windows Server 2022, installed unattended, idle at the console:
>
>   -machine q35,accel=kvm
>   -cpu host,hv-synic,hv-stimer,hv-stimer-direct,hv-vapic,hv-runtime,
>        hv-time,hv-ipi,hv-crash,hv-reset,hv-frequencies,hv-vpindex,
>        hv-spinlocks=0x1fff
>   -smp 64,sockets=2,cores=32,threads=1
>   -m 4G
>   65 pcie-root-ports, 9 virtio devices behind them, qxl
>
> Host: AMD EPYC 7443P, 24 cores / 48 threads.
>
> hv-synic is the flag that matters. Without it the guest never programs
> the SynIC pages and the effect under test does not exist.
>
> Measured with a save/restore harness rather than a live migration: a
> restore from a captured stream walks the same qemu_loadvm_state_main()
> path a destination does, which removes libvirt, the network and the
> second host from the measurement.
>
> CC: Peter Xu <[email protected]>
> CC: Fabiano Rosas <[email protected]>
> CC: Paolo Bonzini <[email protected]>
> CC: Zhao Liu <[email protected]>
> Signed-off-by: Denis V. Lunev <[email protected]>
>
> Denis V. Lunev (2):
>   migration: let a vmstate defer its post_load to end of stream
>   target/i386: defer the Hyper-V SynIC post_load
>
>  include/migration/vmstate.h |  40 +++++++++++++
>  migration/savevm.c          |  23 ++++++++
>  migration/vmstate.c         |  67 +++++++++++++++++++++-
>  target/i386/machine.c       |   1 +
>  tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
>  5 files changed, 240 insertions(+), 2 deletions(-)
>
please disregard this for now. I'll come back with v2 which
will put as around 30 ms.

Flatview rebuild should be optimized.

Sorry for noise,
    Den

Reply via email to