Could Pacemaker’s controller block on a stalled block-device read while
faulting in a swapped-
out libqb SHM page, before it receives the executor reply and starts the
operation timer?

Would that behavior be expected?

I'm investigating a lockup where the PostgreSQL primary and Pacemaker DC
were on the same node when a degraded RAID array caused reads to appear
unbounded without returning errors. Pacemaker did not stop the PostgreSQL
resource or promote a standby. crm_mon showed the DC as standby (with
active resources).

Testing that simulates the incident reproduces that production Pacemaker
behavior. In that setup a libqb SHM page is swapped out and controller
blocks in shmem_fault() before the executor reply and operation timer. The
surviving production logs (such as the crm_mon output) match. Corosync
membership remained intact in both the incident and the test.


-- 
Brian Cosgrove
_______________________________________________
Manage your subscription:
https://lists.clusterlabs.org/mailman/listinfo/users

ClusterLabs home: https://www.clusterlabs.org/

Reply via email to