Could Pacemaker’s controller block on a stalled block-device read while faulting in a swapped- out libqb SHM page, before it receives the executor reply and starts the operation timer?
Would that behavior be expected? I'm investigating a lockup where the PostgreSQL primary and Pacemaker DC were on the same node when a degraded RAID array caused reads to appear unbounded without returning errors. Pacemaker did not stop the PostgreSQL resource or promote a standby. crm_mon showed the DC as standby (with active resources). Testing that simulates the incident reproduces that production Pacemaker behavior. In that setup a libqb SHM page is swapped out and controller blocks in shmem_fault() before the executor reply and operation timer. The surviving production logs (such as the crm_mon output) match. Corosync membership remained intact in both the incident and the test. -- Brian Cosgrove
_______________________________________________ Manage your subscription: https://lists.clusterlabs.org/mailman/listinfo/users ClusterLabs home: https://www.clusterlabs.org/
