Hi Chris,

Thank you for your comment.

We are currently investigating this issue.

Although we have been unable to reproduce it—and thus cannot attach detailed 
logs—I would like to respond to your comment.

>> Our users have configured a 5-node cluster using Pacemaker 2.1.9, which is 
>> included with RHEL 9.6.
>Is this something you just started noticing with pacemaker-2.1.9?

Yes.
This is the first time we have implemented this configuration in version 2.1.9.

>> After successfully building and starting the 5-node cluster, we induced a 
>> failure in the fourth node (test4), triggered STONITH, and then rejoined 
>> test4 to the cluster.
>> When we subsequently ran `pcs status --full` from the fourth node (test4), 
>> we observed an issue where resources actually running on nodes test2 and 
>> test3 were displayed as "Stopped."
>Just to make sure this is definitely a pacemaker problem, can you try 
>using the crm_mon command instead of pcs?  So in this case, use
>`crm_mon -1r` to get the same output.

Okay!
Since the actual environment is our user environment, I will ask them to verify 
this by running crm_mon.


>Also, am I understanding the situation:
>* On the node that was fenced and rejoined, the output is incorrect.
>* On all other nodes, the output is correct.
>Does the node have to be fenced to cause the problem, or is it enough to 
>just bring the node down and back up (either by stopping the cluster on 
>that node, putting it in maintainence mode, etc.)?

It appears that the issue is not directly related to STONITH.
According to additional information from our users, the issue can be reproduced 
simply by stopping and restarting the node after a resource migration.

We have also been informed that resource statuses are displayed correctly on 
nodes other than the restart node.


>> Furthermore, the display did not correct itself even after leaving the 
>> system for some time.
>I imagine this is because there haven't been any CIB changes, so the 
>fenced node hasn't received an update.  You could test this by doing 
>something that would trigger a CIB update.

Just to be safe, after the issue occurs for the user, I will have them 
tentatively update the attribute using `attrd_updater` and verify the display.


>> Are you aware of any issues where the display becomes distorted in the 
>> version of Pacemaker bundled with RHEL 9.6?
>> Alternatively, is there any record of such an issue being resolved in a 
>> newer version of Pacemaker?
>I'm not aware of anything like this, but I haven't skimmed the bug 
>backlog to see if there are any similar reports.  It doesn't sound 
>familiar, at least.

Okay!
Thanks.


We are attempting to reproduce the issue using five nodes with dummy resources, 
but we have not yet succeeded in reproducing it.

There is one point that concerns us.
In the user's environment, when the issue occurs, the following log appears on 
the node that has restarted.
It appears that the `num_updates` of the restarted node(somvgtw1) has advanced 
beyond that of the DC node(somvgtw3) prior to the cluster configuration.
---
(snip)
Sep 03 14:44:42.784 somvgtw1 pacemaker-based     [3422470] 
(cib_process_replace)        info: Digest matched on replace from somvgtw3: 
8aec5935cf9ab4569e64d3fe357f1f1b
Sep 03 14:44:42.784 somvgtw1 pacemaker-based     [3422470] 
(cib_process_replace)        info: Replacement 0.13.9 from somvgtw3 not applied 
to 0.13.20: current num_updates is greater than the replacement
(snip)
---
We suspect that the CIB synchronization from the DC node is not completing 
successfully.

Many thanks,
Hideo Yamauchi.

_______________________________________________
Manage your subscription:
https://lists.clusterlabs.org/mailman/listinfo/users

ClusterLabs home: https://www.clusterlabs.org/

Reply via email to