On Mon, 27 Jul 2026 at 02:36, Michael Paquier <[email protected]> wrote: > > On Mon, Jun 01, 2026 at 03:30:58PM +0900, Kyotaro Horiguchi wrote: > > I wonder whether strengthening the history-based matching would be > > sufficient instead. If timelines with the same TLI but different > > histories can be treated as distinct and pg_rewind continues walking > > the history chain until it finds a common ancestor, that seems like a > > fairly natural fit with the existing timeline model. > > Yeah, perhaps there is something that we could do here better in > pg_rewind in terms of TLI history. A second software layer to ensure > TLI unicity feels just like a shortcut: we already have a LSN.
The LSN can easily match for two promotions. > > UUIDs would certainly make identification straightforward, although > > they would also introduce longer identifiers that are a bit less > > convenient for humans to work with. My initial thought is that it may > > be worth exploring how far we can get with the existing history > > information before introducing a new identifier. > > This. For this reason, I am not convinced that this proposal is > neither acceptable or something that we need to do at all. > > Why should we bear the burden of introducing a second level of > timeline identification knowing that by design we have to rely on a > *single* archive location for a new timeline selection when a standby > is triggered for promotion? My opinion is that this is trying to fix > a problem for something that it not actually a problem. If you play > with HA scenarios where there is a risk of two standbys reusing the > same timeline number, just don't do that. We've historically claimed > that the archive strategy is wrong if your deployments cannot > guarantee a unique TLI assignment. A new sub-identification system > will not provide more guarantees. The trouble is that with the current archiving model it is impossible to ensure a unique TLI assignment. Restore command checks for existence, the first identifier that gives an error will get picked, the promotion finished and then later it is asynchronously archived. If there is a failure between promotion and archiving you will get multiple nodes on the same timeline. There's also the issue that restore command has no way of distinguishing whether the archive is unavailable or the requested file is missing. Having two failures so close by might sound improbable, but my experience shows that failures are very often correlated. E.g. a gray failure of networked storage can cause systems to become unresponsive to the extent of being indistinguishable from dead, and this often happens in a rolling manner. Primary fails, secondary promotes and then immediately fails itself. I have seen this happen almost daily when someone misconfigures their backups to happen at the same time across their fleet. However I am not sure if this proposal is going far enough. The goal here is to uniquely identify transaction log to originate from a specific postgres instance running as a primary. It is reasonable to expect that HA systems ensure that for each instance running as primary there exists a single promote call. But not really much else. This proposal only handles the history file uniqueness in pg_rewind. Unless the archive command is very carefully managed and the archive itself is fully consistent, it is also possible to get WAL files in the archive from different primaries but with the same TLI. If things go well this might just result in replication failing, but if things align in the wrong way it can also cause silent data corruption. This is still too fragile for my taste. One straightforward solution is to replace the whole TLI with a UUID and invent some new method of determining "latest" timeline. This has the obvious problem of breaking almost all of the existing tooling. But a less invasive way to at least detect issues and give backup tools ways to solve the problem would be to use this UUID proposal and also include it in XLogLongPageHeaderData. Recovery can then validate that the WAL file is from the correct timeline. Regards, Ants Aasma
