On Thu, Jul 16, 2026 at 03:30:26AM +0300, Vladimir Sementsov-Ogievskiy wrote:
> On 16.07.26 01:45, Michael S. Tsirkin wrote:
> > On Thu, Jul 16, 2026 at 01:11:05AM +0300, Vladimir Sementsov-Ogievskiy
> > wrote:
> > > On 15.07.26 23:32, Michael S. Tsirkin wrote:
> > > > On Wed, Jul 15, 2026 at 11:21:23PM +0300, Vladimir Sementsov-Ogievskiy
> > > > wrote:
> > > > > On 15.07.26 19:00, Peter Xu wrote:
> > > > > > On Wed, Jul 15, 2026 at 04:52:48PM +0100, Daniel P. Berrangé wrote:
> > > > > > > On Wed, Jul 15, 2026 at 11:48:40AM -0400, Michael S. Tsirkin
> > > > > > > wrote:
> > > > > > > > On Wed, Jul 15, 2026 at 06:01:39PM +0300, Michael Tokarev wrote:
> > > > > > > > > On 14.07.2026 18:42, Vladimir Sementsov-Ogievskiy wrote:
> > > > > > > > > > Hi all!
> > > > > > > > > >
> > > > > > > > > > Here is a migration for TAP net backend, including its
> > > > > > > > > > properties and
> > > > > > > > > > open fds.
> > > > > > > > > >
> > > > > > > > > > With this new feature, management software doesn't need to
> > > > > > > > > > initialize
> > > > > > > > > > new TAP and do a switch to it. Nothing should be done around
> > > > > > > > > > virtio-net in local migration: it just migrates and
> > > > > > > > > > continues to use
> > > > > > > > > > same TAP device. So we avoid extra logic in management
> > > > > > > > > > software, extra
> > > > > > > > > > allocations in kernel (for new TAP), and corresponding
> > > > > > > > > > extra delay in
> > > > > > > > > > migration downtime.
> > > > > > > > >
> > > > > > > > > This is quite a big patch set, - is this really worth the
> > > > > > > > > effort to do
> > > > > > > > > all this just for *local* migration? What's the possible use
> > > > > > > > > case for
> > > > > > > > > this in real, - am I right this is just about upgrading the
> > > > > > > > > host qemu?
> > > > > > > > > And with that in mind, isn't it sufficient to use what we
> > > > > > > > > already have
> > > > > > > > > (namely, create new tap, start new qemu instance, and migrate
> > > > > > > > > the usual
> > > > > > > > > way), and tolerate some very minor downtime while the
> > > > > > > > > networking code
> > > > > > > > > learns the new network topology (isn't it happening almost
> > > > > > > > > instantly
> > > > > > > > > anyway, and if not, the management can help by sending
> > > > > > > > > gratitious ARP)?
> > > > > > > > >
> > > > > > > > > I wonder what's the use for this at yandex?
> > > > > > > > >
> > > > > > > > > Thanks,
> > > > > > > > >
> > > > > > > > > /mjt
> > > > > > > >
> > > > > > > > Well just theoretically, imagine a big VM, reserving twice the
> > > > > > > > amount
> > > > > > > > of memory just to migrate is not nice at all.
> > > > > > >
> > > > > > > Don't we already have the ability to skip memory transfer by
> > > > > > > setting
> > > > > > > the "x-ignore-shared" capability, assuming the VM RAM has a shared
> > > > > > > memory backing.
> > > > > >
> > > > > > Right, IIUC all similar single-host migrations like this series or
> > > > > > CPR (or
> > > > > > anything else...) should always need to enable x-ignore-shared in
> > > > > > the first
> > > > > > place. That's almost always the starting point of optimizing local
> > > > > > migrations.. no matter how the memory will be shared (by the same
> > > > > > pool of
> > > > > > page cache, or persisted over kexec, etc.).
> > > > > >
> > > > >
> > > > > Yes, sharing RAM between source and target + enabling x-ignore-shared
> > > > > is a first thing to do.
> > > > >
> > > > > This series optimizes TAP recreating. Not only skip recreating but
> > > > > also allow to
> > > > > exclude cloud-networking component form live-update entirely, making
> > > > > the process
> > > > > simpler (less components involved), and as I already said, reducing
> > > > > corresponding
> > > > > downtime.
> > > > >
> > > > > --
> > > > > Best regards,
> > > > > Vladimir
> > > >
> > > > So can you explain, how is this better than
> > > > 1. a persistent tap
> > >
> > > You mean open same tap device both on source and target?
> > >
> > > This will require some additional steps anyway, to avoid packet loss,
> > > like keeping queues disabled on target until post_load.
> > >
> > > Another thing is MAX_TAP_QUEUES=256 in kernel: this is a problem, if you
> > > have more
> > > than 128 queues already opened on source. Seems cleaner just pass already
> > > opened
> > > queues to the target.
> > >
> > > And finally, on hosts with many CPUs, TAP queue allocate noticeable amount
> > > of RAM, so having x2 queues during migration would be an overhead.
> >
> > ah I forgot. yes queues do get
> >
> > > > 2. a non persistent tap that some server gets a hold of
> > >
> > > You mean, just pass FDs externally, instead of using QEMUs migration
> > > stream?
> > > That's possible. But requires mgmt tool to store (or get from source) and
> > > pass
> > > these FDs. Requires mgmt to even know about these FDs.
> >
> > e.g. libvirt already does, right? it creates them?
>
> Not sure. But even if it creates and passes them to QEMU. Does it store them?
> Libvirt may restart, and it shouldn't break further migration.
>
> So, we'll need also an API to get FDs from QEMU. To get FDs from QEMU just
> to pass them to target QEMU - I don't think it's more clear solution.
>
> >
> > > But why? QEMU already
> > > can pass FDs through migration for vfio devices (CPR), why is TAP worse?
> >
> > it's not that it's worse. it's that we are growing bespoke mechanisms so
> > far.
> > so if qemu gets tap fd on command line then what? how does that
> > interact?
>
> Do you mean how this series handles that case? No problem: given on source
> from command line
> FDs will migrate to the target the same way. Or I don't understand the
> question.
But right now target gives its own FDs.
> >
> > > Like with persistent tap, it will require some changes in Qemu anyway, to
> > > avoid packet loss
> > > (like patch 12/15 here).
> >
> > that one is more like a bugfix.
> >
> > > >
> > > > IOW why does qemu need to bother.
> > >
> > > QEMU owns the TAP fd and has full knowledge of its state. Pushing this
> > > responsibility
> > > to an external tool means the tool needs to understand QEMU internals
> > > just to pass
> > > an fd that QEMU already has. I think, that's a worse separation of
> > > concerns.
> > >
> > > Of course, there are other ways to do TAP local migration. But looking at
> > > wider
> > > picture, where we want to migrate not only TAP, but also vfio devices
> > > (already
> > > implemented as CPR migration, but may be updated to use similar approach
> > > as in
> > > this series, to use one migration channel), vhost-user-blk (my another
> > > series
> > > in flight) and vhost-user-fs (not yet published), vhost-vsock migration
> > > ("[PATCH v3 0/7] migration/cpr: support vhost-vsock devices" in flight
> > > series),
> > > it seems a good generic approach: simply pass backends (including open
> > > FDs) to
> > > the target, not involving mgmt. Qemu has full knowledge about these FDs
> > > and
> > > owns the whole state. Migrating them in QEMU looks correct for me.
> >
> > if you find a way to generalize things and reuse them for your
> > purposes without intrusive changes all over qemu, fine.
>
> That's not only my purposes (i.e., as I understand, not only Yandex use case).
> Originally the work was started by Oracle (CPR), and I see interest from other
> companies too.
>
> >
> > But this one is poking as far as frontend code even.
> >
>
> You mean "[PATCH v19 11/15] virtio-net: support local migration of backend" ?
>
> CPR-approach avoids it, because FDs are passed _before_ any devices created
> on target. Is that better? In this way I don't really like the fact that we
> start to use (call some ioctls) in target QEMU the FDs, which are still
> actively used in running source QEMU . I don't like it, but that works,
> and I started working on live-migration on TAP from CPR tap series by
> Steve. Still, such approach doesn't work with vhost-user-blk, as trying
> to share (even for initialization) the fd between source and running
> target may just break things. So I decided to move to similar approach for
> both
> net and storage. Do you think it's better to return to CPR-approach for TAP
> (accepting that it won't generalize to vhost-user) like in
> "[PATCH v4 0/8] Live update: tap and vhost" -
> https://lore.kernel.org/qemu-devel/[email protected]/
> -
> where both approaches were compared and mine was preferred?
Hmm. Can CPR use yours?
>
> --
> Best regards,
> Vladimir