On Tue, Sep 08, 2026 at 08:24:56PM +0100, Daniel P. Berrangé wrote:
> On Tue, Sep 08, 2026 at 02:12:53PM -0400, Peter Xu wrote:
> > On Mon, Sep 07, 2026 at 03:51:12PM +0100, Daniel P. Berrangé wrote:
> > > The migration code caches vmstate/ram writes into an iovec and
> > > flushes this every 128kb.
> > 
> > I am just curious how did this 128K came from. Perhaps this?
> > 
> >   MAX_IOV_SIZE / 2 * 4K
> > 
> > Where QEMU has:
> > 
> > #define MAX_IOV_SIZE MIN_CONST(IOV_MAX, 64)
> > 
> > And it needs to divides 2 because we always push save_page_header() first,
> > which is 8B (in reality, maybe that'll also include some footers ahead from
> > the last page..), then another 4K following it.  Then in average when
> > hitting 64 io vectors there're 32 pages, coming up to be that.
> > 
> > I think that is right math for bulk ram phase, but maybe worth spelling out
> > a bit.. because if above holds it's not very obvious..
> 
> Tracing the qio_channel_writev() calls yet again, I think my
> mention of 128k is wrong. I'm now actually seeing alot of
> 1/2 MB writes. eg

The 1/2 MB writes are likely from multifd senders.

> 
> Writev 64 (nvio=1)
> Writev 64 (nvio=1)
> Writev 64 (nvio=1)
> Writev 64 (nvio=1)

These are likely, MultiFDInit_t, and maybe there're just 4 multifd
channels?

> Writev 1344 (nvio=1)
> Writev 1344 (nvio=1)
> Writev 1344 (nvio=1)
> Writev 1344 (nvio=1)
> Writev 281 (nvio=1)
> Writev 8 (nvio=1)
> Writev 525632 (nvio=129)
> Writev 525632 (nvio=129)
> Writev 13632 (nvio=4)
> Writev 173376 (nvio=43)
> Writev 525632 (nvio=129)
> Writev 525632 (nvio=129)
> Writev 525632 (nvio=129)
> Writev 525632 (nvio=129)
> Writev 525632 (nvio=129)
> Writev 525632 (nvio=129)
> 
> 
> > 
> > > 
> > > The QIOChannelTLS receives the iovec, but size GNUTLS cannot
> > > accept iovec data, it iterates calling send for each element.
> > > 
> > > As a result of the migration data pattern, this results in
> > > GNUTLS putting writes on the wire that alternate between about
> > > 4k and 30 bytes.
> > > 
> > > This is triggering the nagle algorithm on migration-test for
> > > many of the TLS test cases, resulting in a "go slow" for I/O
> > > that eventually hits the migration timeout configured by the
> > > test.
> > 
> > Worth spell out the qio_channel_set_delay() experiment?
> > 
> > Frankly, even knowing qio_channel_set_delay(NO_DELAY) on all channels would
> > fix it too, I don't think I fully get why the hang happened.
> 
> Note, it was never technically a "hang", it was just a "go slow".
> The src was still sending and the dst was still receiving but it
> was pathologically slow, a few KBs per second, instead of 100s or
> 1000s of MBs.

Ah OK, yes "hang" isn't accurate.  IMHO it would be nice to mention the
bandwidth measured in the commit log.

> 
> > Nagle, if my understanding is correct.. should be something trying to
> > accumulate small writes only, it means write can be slightly delayed, but
> > it didn't further explain why even if we push writting to it, it didn't
> > flush properly.
> 
> The nagle algorithm influences the TCP window size. The src cannot
> send more data, until the dst has acknowledged packets already sent.

OK, so it's TLS specific behavior (within gnutls)?

> 
> IIUC, normally if you send large volumes of data the window size will
> grow large quite quickly. If you send lots of small packets, nagle
> can keep the window size small and thus delay pending writes.
> 
> Migration with large iovec arrays was causnig alot of small writes,
> so I think that meant the window size did not grow enough to get
> a high speed.
> 
> > Say, I understand TLS is special now with its io_writev(), being
> > qio_channel_tls_writev(), split the iov into multiple calls to
> > qcrypto_tls_session_write(), which is likely why the problem existed, but I
> > don't think I know why multiple qcrypto_tls_session_write() (and I believe
> > ultimately, assuming small but continuous write()s to the socket fd) will
> > cause a hang.  Any clue?
> 
> What I can't explain is why only certain contributors ever saw this
> as a problem ?

Me too.  I think the NODELAY test at least proved it is relevant to how ACK
happens, and if that ACK delay behaves differently on different host, it
may explain.

> 
> > > This patch thus queries the max TLS record size and then
> > > flattens the iovec into buffers of this size. If the
> > > iovec only contains a single element, bounce buffering
> > > is skipped to avoid the redundant copy.
> > 
> > I saw there is also gnutls_record_cork() and the uncork(), which seems to
> > resolve the same issue (I tried to look at gnutls git history but I didn't
> > find any mention of why the API introduced.. though).
> > 
> > Any thoughts on why not relying on that, say, would it work too if cork()
> > at start of qio_channel_tls_writev(), loop, then uncork()?
> 
> Yes, relying on gnutls_record_cork is something I can try - it would
> certainly be nice to avoid the bounce buffering, as that's significant
> overhead when we're talking about iovecs with 1/2 MB of data at a
> time.

I had a quick look at v2, when looking into the cork() a bit more, I found
that gnutls is doing the caching before encryption not after, so I think
there's still a bounce buffer..

Said that, I wonder if using cork() is still a good approach, not only if
that solves the current problem, but also because it trades "memcpy" with
"less syscalls" too as side effect: IIUC we used to write() too frequently,
in case of RAM headers maybe one write on a few bytes worst case, but now
it's one shot, and IIUC the size should be the same as qemufile caching.

What I plan to do is I want to do a simple perf test tomorrow with TLS
migration, single threaded as start, to see if v2 would improve performance
(ignoring the fact it would fix the nodelay issue).

Another thing I can report early is v2 fails to compile when gnutls-devel
isn't available.

Thanks,

> 
> I'll prepare a v2.
> 
> With regards,
> Daniel
> -- 
> |: https://berrange.com       ~~        https://hachyderm.io/@berrange :|
> |: https://libvirt.org          ~~          https://entangle-photo.org :|
> |: https://pixelfed.art/berrange   ~~    https://fstop138.berrange.com :|
> 

-- 
Peter Xu


Reply via email to