Thanks for your reply. Those are great tips, however this was a freak
firmware bug.

"Degraded RAID array" was the wrong wording since that has a precise
meaning for RAID. We were not in an ordinary degraded RAID mode with, say,
a disk failure in the array. The array itself (or driver, or controller?)
was misbehaving. Reads were as far as I can tell, actually unbounded: never
returning and not erroring. I don't expect it to recur after firmware
updates.

It may, however, have revealed a bug or design flaw in Pacemaker as this is
exactly the kind of hardware fault I would expect to trigger fencing and
promotion of the hot standby.

Only manual intervention by an operator (as the controller was stalled) led
to a successful STONITH (we ran the script by hand, a DC election then
occurred and Pacemaker correctly moved the primary database resource and
the resource agent promoted the hot standby).

I have a draft simple reproduction if it helps maintainers triage this or
alternatively help me understand that this was expected behavior and why.

https://github.com/cosgroveb/pacemaker-shm-timer-repro/blob/main/repro.sh

On Tue, Aug 11, 2026 at 8:33 AM Windl, Ulrich via Users <
[email protected]> wrote:

> Hi!
>
>
>
> A non-pacemaker answer: I can imagine two solutions for your problem:
>
>    1. Limit the reconstruction rate of the RAID
>    2. Limit the amount of (dirty) filesystem cache (we had read stalls
>    when someone backed up the database and most of the RAM had been filled
>    with dirty buffers to write out)
>
>
>
> Kind regards,
>
> Ulrich Windl
>
>
>
> *From:* Users <[email protected]> *On Behalf Of *Brian
> Cosgrove
> *Sent:* Monday, August 10, 2026 10:03 PM
> *To:* [email protected]
> *Subject:* [EXT] [EXT] [ClusterLabs] Can stalled libqb SHM I/O delay
> Pacemaker's operation timer?
>
>
>
> Could Pacemaker’s controller block on a stalled block-device read while
> faulting in a swapped-
> out libqb SHM page, before it receives the executor reply and starts the
> operation timer?
>
>
>
> Would that behavior be expected?
>
>
>
> I'm investigating a lockup where the PostgreSQL primary and Pacemaker DC
> were on the same node when a degraded RAID array caused reads to appear
> unbounded without returning errors. Pacemaker did not stop the PostgreSQL
> resource or promote a standby. crm_mon showed the DC as standby (with
> active resources).
>
>
>
> Testing that simulates the incident reproduces that production Pacemaker
> behavior. In that setup a libqb SHM page is swapped out and controller
> blocks in shmem_fault() before the executor reply and operation timer. The
> surviving production logs (such as the crm_mon output) match. Corosync
> membership remained intact in both the incident and the test.
>
>
>
>
>
> --
>
> Brian Cosgrove
> _______________________________________________
> Manage your subscription:
> https://lists.clusterlabs.org/mailman/listinfo/users
>
> ClusterLabs home: https://www.clusterlabs.org/
>


-- 
Brian Cosgrove
_______________________________________________
Manage your subscription:
https://lists.clusterlabs.org/mailman/listinfo/users

ClusterLabs home: https://www.clusterlabs.org/

Reply via email to