Hi,

We've been chasing a slow QUIC connection/memory leak on a production
HTTP/3 edge (running 3.2.21) and traced it down to a specific ordering
issue in the mux_quic / quic_conn interaction. Confirmed the same code
path is still present in v3.2.23 and on current haproxy-3.2 branch HEAD,
so writing it up here.

If a client's first post-handshake flight on a QUIC connection contains
RESET_STREAM and STOP_SENDING on stream 0 without ever completing an
HTTP/3 HEADERS frame, and this becomes visible to HAProxy before the
mux's very first qcc_io_cb() run for that connection, the resulting
quic_conn is never freed. On one of our production nodes this had
accumulated to about 1.22M stuck connections over 55 days, consuming
around 80GB in the quic_conn_rxbuf pool alone (a fixed 64KiB buffer held
per connection for its whole lifetime). timeout client and timeout
http-request were both configured with finite values (65s and 10s) on
the affected frontend, which rules out the more obvious "no timeout
configured" explanation.

The mechanism, as far as we can tell: qcc_recv_reset_stream() creates
the qcs via qcc_get_qcs()->h3_attach()->qcs_wait_http_req(), which sets
nb_hreq=1. It then calls qcs_close_remote() but never calls
qcc_refresh_timeout() anywhere in the function, unlike its sibling
qcc_recv_stop_sending() which does (guarded by !qcc->nb_hreq). Since the
local side of the stream never closes, qcs_is_completed() never becomes
true and nb_hreq stays at 1 forever. If the connection-level error is
already visible by the time qcc_io_cb() runs for the first time,
qcc_io_send() bails out before qcc_app_init() ever runs, so total stays
0 for that pass and qcc_io_cb()'s own "if (total) qcc_refresh_timeout()"
is skipped too. qcc->task gets a finite expire computed once at
qmux_init() time but qmux_init() never calls task_queue() on it, and
since qcc_refresh_timeout() is the only place that does, the task simply
never gets inserted into the scheduler's wait queue. Meanwhile
qcc_is_dead() checks nb_hreq before the transport-error flags:

    static inline int qcc_is_dead(const struct qcc *qcc)
    {
            /* Maintain connection if there is still request streams active. */
            if (qcc->nb_hreq)
                    return 0;

            if (qcc->flags & (QC_CF_ERR_CONN|QC_CF_ERRL_DONE) ||
                !qcc->task) {
                    return 1;
            }

            return 0;
    }

so with nb_hreq stuck at 1 this always returns 0 even once
QC_CF_ERR_CONN is set. And quic_conn_release() refuses to run at all
while the mux is still considered alive (BUG_ON(qc->mux_state ==
QC_MUX_READY) in src/quic_conn.c), which with the above never changes
for these connections.

We attached gdb (read-only, using a symbol-bearing debug build as the
symbol source against the running production binary) and confirmed the
whole chain field for field on a sample of the stuck connections: all of
them sitting in the per-thread quic_conns_clo list rather than the
active quic_conns list, qc->idle_timer_task == NULL, qc->mux_state ==
QC_MUX_READY, qcc->nb_sc == 0, qcc->nb_hreq == 1 (checked 2000
connections from the closing list, 2000/2000 had nb_hreq == 1, so this
isn't a rare edge case), qcc->flags == QC_CF_ERR_CONN only, qcc->app_st
== QCC_APP_ST_NULL (qcc_app_init() never ran, which is also why there
was exactly one qcs per connection instead of the usual control stream
plus request stream), qcc->task non-NULL but not present in either the
wait queue or run queue, and the single qcs on each connection sitting
at id 0, state QC_SS_HREM, flags RECV_RESET|HREQ_RECV|TO_RESET|
SIZE_KNOWN, with rx.offset_max == 0 (zero bytes of stream data ever
arrived).

We reproduced the same state locally, though we deliberately avoided
building a client that forces the datagram-coalescing timing on the
wire, since that would double as a working remote resource-exhaustion
trigger against any QUIC-enabled HAProxy. Instead we forced the
equivalent ordering from the server side in an instrumented debug build,
gating qcc_io_cb()'s first pass behind a short delay while qcc->app_st
was still QCC_APP_ST_NULL, and drove it with an ordinary, unmodified
aioquic client that just cancels stream 0 via the library's normal
reset/stop_sending API before ever sending a complete request. That
reproduced the identical state (matched production on 14 of 16 checked
fields directly; the remaining 2 we could only infer since they're
behind a struct that's private to h3.c and not visible from mux_quic.c
instrumentation). We suspect, based on qc->idle_start and qcs->start
being identical on the production samples, that on the wire this
ordering happens naturally whenever a client's abort frames land in the
same UDP datagram as its Handshake-epoch Finished message, but we
haven't confirmed that with an actual packet capture.

The fix we validated against the reproduced state is to stop letting
nb_hreq override an already-confirmed transport error in qcc_is_dead():

    static inline int qcc_is_dead(const struct qcc *qcc)
    {
    +       /* A connection already dead at transport level can never
    +        * complete any request, so nb_hreq must not keep it alive.
    +        */
    +       if (qcc->flags & (QC_CF_ERR_CONN|QC_CF_ERRL_DONE))
    +               return 1;
    +
            /* Maintain connection if there is still request streams active. */
            if (qcc->nb_hreq)
                    return 0;

            if (qcc->flags & (QC_CF_ERR_CONN|QC_CF_ERRL_DONE) ||
                !qcc->task) {
                    return 1;
            }

            return 0;
    }

With just this, the connection is reclaimed on the very next tasklet
pass, no timer involved, and a normal completed H3 GET tears down
identically with or without the patch. We also tried adding a
qcc_refresh_timeout() call inside qcc_recv_reset_stream() itself,
mirroring what qcc_recv_stop_sending() already does. That works too,
but only reclaims the connection once timeout http-request/timeout
client elapses, and does nothing if neither is configured, so we'd
suggest the qcc_is_dead() change as the primary fix. The
qcc_recv_reset_stream() addition still seems worth doing on its own
merits though, since it brings it in line with its sibling function
regardless of this specific bug.

Happy to share the instrumented diff or the gdb session in more detail
if useful. We don't have a packet capture confirming the exact trigger
in the wild yet, production traffic is TLS end to end and we didn't have
key logging enabled at the time, so the client-side behavior producing
this is still inferred rather than directly observed.

Reply via email to