On 16.07.26 01:45, Michael S. Tsirkin wrote:
On Thu, Jul 16, 2026 at 01:11:05AM +0300, Vladimir Sementsov-Ogievskiy wrote:
On 15.07.26 23:32, Michael S. Tsirkin wrote:
On Wed, Jul 15, 2026 at 11:21:23PM +0300, Vladimir Sementsov-Ogievskiy wrote:
On 15.07.26 19:00, Peter Xu wrote:
On Wed, Jul 15, 2026 at 04:52:48PM +0100, Daniel P. Berrangé wrote:
On Wed, Jul 15, 2026 at 11:48:40AM -0400, Michael S. Tsirkin wrote:
On Wed, Jul 15, 2026 at 06:01:39PM +0300, Michael Tokarev wrote:
On 14.07.2026 18:42, Vladimir Sementsov-Ogievskiy wrote:
Hi all!

Here is a migration for TAP net backend, including its properties and
open fds.

With this new feature, management software doesn't need to initialize
new TAP and do a switch to it. Nothing should be done around
virtio-net in local migration: it just migrates and continues to use
same TAP device. So we avoid extra logic in management software, extra
allocations in kernel (for new TAP), and corresponding extra delay in
migration downtime.

This is quite a big patch set, - is this really worth the effort to do
all this just for *local* migration?  What's the possible use case for
this in real, - am I right this is just about upgrading the host qemu?
And with that in mind, isn't it sufficient to use what we already have
(namely, create new tap, start new qemu instance, and migrate the usual
way), and tolerate some very minor downtime while the networking code
learns the new network topology (isn't it happening almost instantly
anyway, and if not, the management can help by sending gratitious ARP)?

I wonder what's the use for this at yandex?

Thanks,

/mjt

Well just theoretically, imagine a big VM, reserving twice the amount
of memory just to migrate is not nice at all.

Don't we already have the ability to skip memory transfer by setting
the "x-ignore-shared" capability, assuming the VM RAM has a shared
memory backing.

Right, IIUC all similar single-host migrations like this series or CPR (or
anything else...) should always need to enable x-ignore-shared in the first
place.  That's almost always the starting point of optimizing local
migrations.. no matter how the memory will be shared (by the same pool of
page cache, or persisted over kexec, etc.).


Yes, sharing RAM between source and target + enabling x-ignore-shared is a 
first thing to do.

This series optimizes TAP recreating. Not only skip recreating but also allow to
exclude cloud-networking component form live-update entirely, making the process
simpler (less components involved), and as I already said, reducing 
corresponding
downtime.

--
Best regards,
Vladimir

So can you explain, how is this better than
1. a persistent tap

You mean open same tap device both on source and target?

This will require some additional steps anyway, to avoid packet loss,
like keeping queues disabled on target until post_load.

Another thing is MAX_TAP_QUEUES=256 in kernel: this is a problem, if you have 
more
than 128 queues already opened on source. Seems cleaner just pass already opened
queues to the target.

And finally, on hosts with many CPUs, TAP queue allocate noticeable amount
of RAM, so having x2 queues during migration would be an overhead.

ah I forgot. yes queues do get

2. a non persistent tap that some server gets a hold of

You mean, just pass FDs externally, instead of using QEMUs migration stream?
That's possible. But requires mgmt tool to store (or get from source) and pass
these FDs. Requires mgmt to even know about these FDs.

e.g. libvirt already does, right? it creates them?

Not sure. But even if it creates and passes them to QEMU. Does it store them?
Libvirt may restart, and it shouldn't break further migration.

So, we'll need also an API to get FDs from QEMU. To get FDs from QEMU just
to pass them to target QEMU - I don't think it's more clear solution.


But why? QEMU already
can pass FDs through migration for vfio devices (CPR), why is TAP worse?

it's not that it's worse. it's that we are growing bespoke mechanisms so far.
so if qemu gets tap fd on command line then what? how does that
interact?

Do you mean how this series handles that case? No problem: given on source from 
command line
FDs will migrate to the target the same way. Or I don't understand the question.


Like with persistent tap, it will require some changes in Qemu anyway, to avoid 
packet loss
(like patch 12/15 here).

that one is more like a bugfix.


IOW why does qemu need to bother.

QEMU owns the TAP fd and has full knowledge of its state. Pushing this 
responsibility
to an external tool means the tool needs to understand QEMU internals just to 
pass
an fd that QEMU already has. I think, that's a worse separation of concerns.

Of course, there are other ways to do TAP local migration. But looking at wider
picture, where we want to migrate not only TAP, but also vfio devices (already
implemented as CPR migration, but may be updated to use similar approach as in
this series, to use one migration channel), vhost-user-blk (my another series
in flight) and vhost-user-fs (not yet published), vhost-vsock migration
("[PATCH v3 0/7] migration/cpr: support vhost-vsock devices" in flight series),
it seems a good generic approach: simply pass backends (including open FDs) to
the target, not involving mgmt. Qemu has full knowledge about these FDs and
owns the whole state. Migrating them in QEMU looks correct for me.

if you find a way to generalize things and reuse them for your
purposes without intrusive changes all over qemu, fine.

That's not only my purposes (i.e., as I understand, not only Yandex use case).
Originally the work was started by Oracle (CPR), and I see interest from other
companies too.


But this one is poking as far as frontend code even.


You mean "[PATCH v19 11/15] virtio-net: support local migration of backend" ?

CPR-approach avoids it, because FDs are passed _before_ any devices created
on target. Is that better? In this way I don't really like the fact that we
start to use (call some ioctls) in target QEMU the FDs, which are still
actively used in running source QEMU . I don't like it, but that works,
and I started working on live-migration on TAP from CPR tap series by
Steve. Still, such approach doesn't work with vhost-user-blk, as trying
to share (even for initialization) the fd between source and running
target may just break things. So I decided to move to similar approach for both
net and storage. Do you think it's better to return to CPR-approach for TAP
(accepting that it won't generalize to vhost-user) like in
"[PATCH v4 0/8] Live update: tap and vhost" -
https://lore.kernel.org/qemu-devel/[email protected]/
 -
where both approaches were compared and mine was preferred?


--
Best regards,
Vladimir

Reply via email to