Hello Fujii-san,
16.09.2026 13:30, Fujii Masao wrote:
On Mon, Sep 7, 2026 at 1:34 AM Ayush Tiwari<[email protected]> wrote:
Hmm, I think we can just disable the autovacuum completely.
I revised 0002 to disable autovacuum for the test node instead. With your
reproducer, the revised patch passes all 18 tests. The regular test passes
as well.
Thoughts?
LGTM. So barring any objections, I will commit the patches.
It looks like skink asks for the fix, e.g. as below [3]:
69/325 recovery - postgresql:recovery/031_recovery_conflict ERROR
1013.92s (exit status 255 or 0xff)
[11:14:59.084](0.080s) ok 14 - startup deadlock: lock acquisition is waiting
Waiting for replication conn standby's replay_lsn to pass 0/3461410 on primary
done
timed out waiting for match: (?^:User transaction caused buffer deadlock with recovery.) at
/home/bf/bf-build/skink/REL_17_STABLE/pgsql/src/test/recovery/t/031_recovery_conflict.pl line 318.
This is a new failure mode, but it's still caused by autovacuum, as far as
I can see.
I've reproduced this locally when running 20 tests in parallel, in a VM
with 10% CPU limit.
16 [10:21:35] t/031_recovery_conflict.pl ..
16 Dubious, test returned 255 (wstat 65280, 0xff00)
16 All 14 subtests passed
16 [10:21:35]
16
16 Test Summary Report
16 -------------------
16 t/031_recovery_conflict.pl (Wstat: 65280 (exited 255) Tests: 14 Failed:
0)
16 Non-zero exit status: 255
16 Parse errors: No plan found in TAP output
16 Files=1, Tests=14, 510 wallclock secs ( 0.06 usr 0.05 sys + 2.28 cusr
18.31 csys = 20.70 CPU)
16 Result: FAIL
16 # Tests were run but no plan was declared and done_testing() was not
seen.
16 # Looks like your test exited with 255 just after 14.
16 make: *** [Makefile:28: check] Error 1
With log_autovacuum_min_duration = 0:
/031_recovery_conflict_primary.log:
2026-10-01 10:17:23.078 UTC [193321][client backend][4/8:0] LOG: statement:
BEGIN;
2026-10-01 10:17:23.255 UTC [193321][client backend][4/8:0] LOG: statement: INSERT INTO test_recovery_conflict_table1(a)
SELECT generate_series(1, 100) i;
2026-10-01 10:17:23.364 UTC [193321][client backend][4/8:767] LOG: statement:
ROLLBACK;
...
2026-10-01 10:17:27.804 UTC [193405][autovacuum worker][11/7:0] LOG: automatic vacuum of table
"test_db.public.test_recovery_conflict_table1": index scans: 0
pages: 0 removed, 1 remain, 1 scanned (100.00% of total), 0 eagerly
scanned
tuples: 100 removed, 2 remain, 0 are dead but not yet removable
...
2026-10-01 10:17:33.901 UTC [193477][client backend][7/5:0] LOG: connection authorized: user=vagrant database=test_db
application_name=031_recovery_conflict.pl
2026-10-01 10:17:36.171 UTC [193477][client backend][7/6:0] LOG: statement:
VACUUM FREEZE test_recovery_conflict_table1;
2026-10-01 10:17:36.996 UTC [193477][client backend][:0] LOG: disconnection: session time: 0:00:03.096 user=vagrant
database=test_db host=[local]
...
As a result, there is no "WAL redo at X/XXXX for Heap2/PRUNE_VACUUM_SCAN"
after 10:17:36 and "User transaction caused buffer deadlock with recovery"
in 031_recovery_conflict_standby.log.
[1]
https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=skink&dt=2026-09-28%2023%3A37%3A34
[2]
https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=skink&dt=2026-09-29%2010%3A39%3A34
[3]
https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=skink&dt=2026-09-30%2008%3A38%3A16
Best regards,
Alexander