Thanks for your reply. Those are great tips, however this was a freak firmware bug.
"Degraded RAID array" was the wrong wording since that has a precise meaning for RAID. We were not in an ordinary degraded RAID mode with, say, a disk failure in the array. The array itself (or driver, or controller?) was misbehaving. Reads were as far as I can tell, actually unbounded: never returning and not erroring. I don't expect it to recur after firmware updates. It may, however, have revealed a bug or design flaw in Pacemaker as this is exactly the kind of hardware fault I would expect to trigger fencing and promotion of the hot standby. Only manual intervention by an operator (as the controller was stalled) led to a successful STONITH (we ran the script by hand, a DC election then occurred and Pacemaker correctly moved the primary database resource and the resource agent promoted the hot standby). I have a draft simple reproduction if it helps maintainers triage this or alternatively help me understand that this was expected behavior and why. https://github.com/cosgroveb/pacemaker-shm-timer-repro/blob/main/repro.sh On Tue, Aug 11, 2026 at 8:33 AM Windl, Ulrich via Users < [email protected]> wrote: > Hi! > > > > A non-pacemaker answer: I can imagine two solutions for your problem: > > 1. Limit the reconstruction rate of the RAID > 2. Limit the amount of (dirty) filesystem cache (we had read stalls > when someone backed up the database and most of the RAM had been filled > with dirty buffers to write out) > > > > Kind regards, > > Ulrich Windl > > > > *From:* Users <[email protected]> *On Behalf Of *Brian > Cosgrove > *Sent:* Monday, August 10, 2026 10:03 PM > *To:* [email protected] > *Subject:* [EXT] [EXT] [ClusterLabs] Can stalled libqb SHM I/O delay > Pacemaker's operation timer? > > > > Could Pacemaker’s controller block on a stalled block-device read while > faulting in a swapped- > out libqb SHM page, before it receives the executor reply and starts the > operation timer? > > > > Would that behavior be expected? > > > > I'm investigating a lockup where the PostgreSQL primary and Pacemaker DC > were on the same node when a degraded RAID array caused reads to appear > unbounded without returning errors. Pacemaker did not stop the PostgreSQL > resource or promote a standby. crm_mon showed the DC as standby (with > active resources). > > > > Testing that simulates the incident reproduces that production Pacemaker > behavior. In that setup a libqb SHM page is swapped out and controller > blocks in shmem_fault() before the executor reply and operation timer. The > surviving production logs (such as the crm_mon output) match. Corosync > membership remained intact in both the incident and the test. > > > > > > -- > > Brian Cosgrove > _______________________________________________ > Manage your subscription: > https://lists.clusterlabs.org/mailman/listinfo/users > > ClusterLabs home: https://www.clusterlabs.org/ > -- Brian Cosgrove
_______________________________________________ Manage your subscription: https://lists.clusterlabs.org/mailman/listinfo/users ClusterLabs home: https://www.clusterlabs.org/
