On Wed, Sep 30, 2026 at 12:31:53PM +0100, Zsolt Parragi wrote:
> The question is, what would be good enough proof?

I got to the same conclusion independently, from the client side:
clients on Windows get "connection reset" instead of the server's
FATAL, and there's nothing a client can do about it, because by the
time it reads, the data is already gone.  Here are my results, in case
another machine helps.

Setup: Windows 11 Pro 25H2 (build 26200), MSVC 19.50, meson, OpenSSL
3.6.2, master at 9510a826e4a, all over localhost.  "Patched" means
master plus the revert of 29992a6a509b.

1. wait_cleanup with your pg_sleep(0.1) after pg_terminate_backend(),
   and the rest of the injection_points isolation suite, 20 runs each:

     unpatched: 20/20 lose the FATAL ("PQconsumeInput failed: server
                closed the connection unexpectedly"), as in Andrey's CI
     patched:   20/20 get the FATAL, and the suite passes every time

2. Startup FATAL.  The attached script sends a startup packet for a
   role that doesn't exist and sleeps before reading; 50 attempts per
   delay:

     delay before read   unpatched   patched
     0 ms                        50/50           50/50
     50 ms                      4/50             50/50
     300 ms                    0/50             50/50

3. The tests that hung in 2022 (commit_ts/002_standby,
   commit_ts/003_standby_2, recovery/001_stream_rep), patched, 20 runs
   each: all 60 passed, no hangs.

4. SSL, patched: ssl/001-004, 20 runs each, all passed.  Alexander
   reported in [1] that the revoked-client-cert case in 001_ssltests.pl
   sometimes got "Software caused connection abort" with the earlier
   patch set, so I also looped just that case 2000 times on both builds.
   It reported "certificate revoked" every time on each.

5. A full meson test run, patched, with PG_TEST_EXTRA=ssl but without
   injection points (those are covered by 1).  Everything passed except
   pg_test_timing/001_basic and psql/001_basic, which fail here because
   of the Polish locale's decimal comma, not because of the patch.

Not tested: a connection that isn't over localhost, and the back
branches.

Two more pieces of history that may help: Thomas already proposed
re-committing 6051857fc on master in March 2025 [2], and nobody
objected, but it didn't happen.  And there's a remaining walreceiver
hang on WSAECONNRESET that a8458f508 doesn't cover [3], but it happens
without the revert too, so I don't think it's related.

So +1 for trying the revert on master.

[1] https://postgr.es/m/[email protected]
[2] 
https://postgr.es/m/CA+hUKGKJSOAdAukP4QTkR3-jFws39+8C197XC-a97dgYr=c...@mail.gmail.com
[3] https://postgr.es/m/[email protected]

--
Kacper Kuras

Attachment: repro_startup_fatal.pl
Description: repro_startup_fatal.pl

Reply via email to