On 15.07.26 23:20, Michael S. Tsirkin wrote:
On Wed, Jul 15, 2026 at 11:10:28PM +0300, Vladimir Sementsov-Ogievskiy wrote:
On 15.07.26 18:01, Michael Tokarev wrote:
On 14.07.2026 18:42, Vladimir Sementsov-Ogievskiy wrote:
Hi all!
Here is a migration for TAP net backend, including its properties and
open fds.
With this new feature, management software doesn't need to initialize
new TAP and do a switch to it. Nothing should be done around
virtio-net in local migration: it just migrates and continues to use
same TAP device. So we avoid extra logic in management software, extra
allocations in kernel (for new TAP), and corresponding extra delay in
migration downtime.
This is quite a big patch set,
You probably haven't been following Steve's work (about 5 years) on
CPR Live Update :) My patch set is small.
- is this really worth the effort to do
all this just for *local* migration? What's the possible use case for
this in real, - am I right this is just about upgrading the host qemu?
Exactly. You say "just for", but actually local migration is a lot more
"massive" operation for us. Remote migrations are done mostly to release the
physical server for service/redeploy (you should first migrate all the
vms to other servers). That's a relatively seldom operation.
On the contrary, when updating to new QEMU version, you migrate _all_
vms on _all_ servers. These (minimal) downtimes affects all the
customers, some are sensitive to freezes.
And with that in mind, isn't it sufficient to use what we already have
(namely, create new tap, start new qemu instance, and migrate the usual
way), and tolerate some very minor downtime while the networking code
learns the new network topology (isn't it happening almost instantly
anyway, and if not, the management can help by sending gratitious ARP)?
I wonder what's the use for this at yandex?
This series together with similar about vhost-user-blk local migration
gives several times win in freeze-time. It's significant for us.
--
Best regards,
Vladimir
I believe it is.
But let's focus on what exactly are we saving here.
It is not that clear.
Is it the overhead of tap losing state such as
bridge forgetting the tap mac. Or what?
I don't have a full expertize here, but the overhead is in switching
to a new TAP in our cloud virtual networking component. Could this
switching be optimized, to always work very fast? That's a good
question that I don't have an answer to. Probably yes.
We decided to go another way: drop extra components from the process,
and avoid any extra switching at all. Both cloud networking component, and
storage (with vhost-usr-blk live-udpate series) don't even know, that
QEMU is migrating. That seems to be a clear design for local migration,
and shows good results.
--
Best regards,
Vladimir