On Tue, Jul 14, 2026 at 07:45:36PM +0530, Aadeshveer Singh wrote: > This series implements a "fast snapshot load" mechanism to > significantly reduce the perceived resume time of a VM from a snapshot > file. > > Currently, resuming a VM from a snapshot file requires loading all RAM > pages into the QEMU instance before execution begins. This extension > allows the user to run the VM nearly instantly by loading only the > required device states up front and loading RAM pages lazily, by > trapping access to pages that have not yet been loaded. > > Using the Linux userfaultfd syscall, a fault thread catches all page > faults caused by the guest and loads in the pages required to keep > the VM running. Concurrently, an eager background thread iteratively > loads all remaining pages into RAM so the guest does not have to > depend on the fault thread indefinitely. > > Much of code is reused from postcopy for fault handling and precopy > for reading mapped ram file. Implementation revolves around two > threads named the fault thread and eager load thread. Fault thread as > name suggests catches page faults by the guest and serves them using > userfaultfd. Postcopy fault thread is reused but instead of requesting > source for a page it loads the page directly by reading from file. In > order to remove the dependency of guest on fault thread indefinitely > the eager load thread loads in the entire RAM sequentially, and after > iterating through the entire RAM signals fault thread to exit and > calls cleanup. > > In order to prevent the case of a page being loaded twice(in the > case when eager load thread is loading it and fault thread also > tries to serve fault on same page) a bitmap called pending_bmap is > used to track pages which are pending and not being loaded by any > thread. Atomic operations on this bitmap allows coordination between > threads to prevent any unwanted behaviours > > Future direction: > - Add support for multifd > - Add support for vhost-user
Please always still at least have a section for testing. It's an important piece of information on what you have tested (and what you have not, which can also be included when it may matter). For example, I know you should have tested at least normal and some huge page setup, you can list them there. You can also mention if you only tested x86_64, or if you have played with anything else; logically this should also work elsewhere whenever all facilities are available. It's fine to only test on x86_64, though. Thanks, > > --- > v1 -> v2 > > - Structure Change: Expanded from 7 to 11 patches > - Move test removal patch from first to post implementation as > suggested by Peter > - Patch 1 to 5 are meant to be non functional changes meant to > modernize parts of code and prepare for feature Implementation > - Patch 10 performs a smoke test for implemented feature > - Patch 11 adds a documentation page for fast snapshot load > - Added explicit fails for features not supported/incompatible with > fast snapshot load in compatibility check in options.c as suggested > by Peter > - Passed errp object to qemu_get_buffer_at() to improve on error > reporting as suggested by Peter > - Size of pending_bmap is updated to have a bit for every guest page > according to guest page size and not host page size as pointed > by Peter > - New function(ramblock_file_bitmap_page_is_nonzero()) has been > introduced to test for zero pages while loading essentially enableing > support for hugepages > - Updated the documentation comment over postcopy_mapped_ram_load_page() > to add errp as pointed by Peter > - Added a new function try_mark_postcopy_blocktime_begin() to call > mark_postcopy_blocktime_begin() taking care of assumptions like > skipping if page is already received to improve on the readability > as pointed by Peter > - Fixed bug, that sent non page alined addresses to > mark_postcopy_blocktime_begin() as pointed by Peter > - Separate the postcopy_listen_thread_bh() rename patch as suggested > by Peter > - Add comment explaining changes in process_incoming_migration_co() > as suggested by Peter > - Make qemu_loadvm_run_fast_snapshot_load() and > postcopy_ram_eager_load_setup() return void as these functions > never fail as pointed by Peter > > Aadeshveer Singh (11): > migration: Propagate error in postcopy setup functions > migration: Extract blocktime marking helper > migration: Rename postcopy_listen_thread_bh > migration: Use file_bmap for RAMBlock during incoming file load > migration: Make qemu_get_buffer_at() thread-safe > migration: add RAMBlock field and helper for fast snapshot load > migration: add support for fault thread to load pages from disk > migration: add eager load thread and setup for fast snapshot load > migration/tests: remove capability conflict test > postcopy-ram+mapped-ram > migration/tests: Add test for fast snapshot load > docs/migration: Add documentation for fast snapshot load feature > > docs/devel/migration/fast-snapshot-load.rst | 81 ++++++ > docs/devel/migration/features.rst | 1 + > include/system/ramblock.h | 6 + > migration/migration.c | 60 +++-- > migration/migration.h | 5 + > migration/options.c | 14 +- > migration/postcopy-ram.c | 266 ++++++++++++++++---- > migration/postcopy-ram.h | 7 +- > migration/qemu-file.c | 9 +- > migration/qemu-file.h | 4 +- > migration/ram.c | 103 +++++++- > migration/ram.h | 2 + > migration/savevm.c | 16 ++ > migration/savevm.h | 2 + > migration/trace-events | 2 + > tests/qtest/migration/file-tests.c | 24 ++ > tests/qtest/migration/misc-tests.c | 52 ---- > 17 files changed, 511 insertions(+), 143 deletions(-) > create mode 100644 docs/devel/migration/fast-snapshot-load.rst > > -- > 2.55.0 > -- Peter Xu
