This series improves RDMA migration throughput during iterative precopy and the iterable stop-copy phase by allowing multiple writes to be in flight for the same registration chunk.
RDMA tracks RAM in registration chunks. During the final iterable transfer, workloads can leave many scattered 4 KiB dirty pages, with multiple pages falling into the same chunk. The current code waits for an earlier write to a chunk to complete before posting the next one. These serial completion waits can leave send queue capacity unused and limit RDMA bandwidth, increasing downtime. A ram_save_iterate() call can stop due to rate limiting before the RAM scanner completes a full round. Therefore, it is not the correct boundary for ordering page versions. Add a return-valued notifier at the point where the RAM scanner wraps to the first RAMBlock. Convert multifd's existing per-round synchronization to use this notifier, and register an RDMA notifier that flushes buffered writes and drains outstanding completions at the same boundary. This provides the explicit completion barrier before a newer version of a page can be sent. The per-chunk wait is therefore unnecessary and is removed, together with the now-unused transit bitmap. The test configuration is: RDMA chunk size: 32 MiB Guest: 32 vCPUs, 128 GiB RAM Test runs: 10 Maximum HCA bandwidth: 100 Gbps Workload: idle Average result(pin-all=true) Before After pin-all is true Downtime 396.7 ms 345.7 ms ~12.9% improve Final iterable bandwidth 9952.0 MiB/s 11639.10 MiB/s ~17.0% improve Average results(pin-all=false) Before After pin-all is false Downtime 248.7 ms 191.1 ms ~23.2% improve Final iterable bandwidth 5605.1 MiB/s 9826.0 MiB/s ~75.3% improve "Final iterable bandwidth" is the bandwidth measured during qemu_savevm_state_complete_precopy_iterable(). Non-iterable data is not transferred using RDMA Write. Changes in v3: - Optimze comments about RAMRoundNotifyData in patch1 - Collect Reviewed-by tags Changes in v2 (Suggested by Peter): - Add a generic RAM round synchronization notifier. - Drain RDMA writes at the RAM scan-round boundary. - Remove the per-chunk wait and the unused transit bitmap. Yanfei Xu (4): migration/ram: Add RAM round synchronization notifiers migration/rdma: Drain writes at RAM round boundaries migration/rdma: Allow multiple in-flight writes per chunk migration/rdma: Remove unused transit bitmap migration/multifd-nocomp.c | 24 ++++++++++ migration/ram.c | 45 +++++++++++++++--- migration/ram.h | 12 +++++ migration/rdma.c | 97 ++++++++++++++++++++------------------ migration/trace-events | 5 +- 5 files changed, 128 insertions(+), 55 deletions(-) -- 2.20.1
