Hi hackers, pg_resetwal does not look for backup_label today. On a restored base backup, pg_control says "in production", so pg_resetwal asks for -f, and with -f it goes ahead. The server then fails with "could not locate required checkpoint record" and a hint to remove backup_label. Removing it at that point throws away what the backup needs to become consistent.
The attached patch makes pg_resetwal refuse when backup_label exists, even with -f. This follows the postmaster.pid check, which -f does not override either. It came up in the CF 4997 thread [1], and David preferred that -f not bypass it [2]. On a pg_basebackup copy it now says: pg_resetwal: error: backup label file "backup_label" exists pg_resetwal: hint: If you are restoring from a backup, configure recovery instead. If you are not restoring from a backup, delete the backup label file and try again. -n is refused too. A dry run should show what a real run would do, and the real run refuses. pg_controldata still works for reading the control file. Robert worried in [3] that people with a bad backup will just run pg_resetwal, which is worse than starting from the wrong checkpoint. This patch targets that step. A backup that still has its backup_label is the case where the user has everything needed for a correct restore, and pg_resetwal is the wrong tool. Now it says so before it does any damage. A user can still delete the file and rerun, but that is a second deliberate step, and the docs now say when that is safe. 0002 adds TAP tests and is optional. I think this is master only, since it changes what an existing command accepts. [1] https://postgr.es/m/[email protected] [2] https://postgr.es/m/[email protected] [3] https://postgr.es/m/ca+tgmozkdrwyd7kipfhajbg+tjm3ufrqobk1etg3ntvs97-...@mail.gmail.com Thanks, Shihao
v1-0002-Test-pg_resetwal-backup_label-check.patch
Description: Binary data
v1-0001-pg_resetwal-Refuse-to-run-when-backup_label-exist.patch
Description: Binary data
