This article explains an issue where the MNHA cluster fails over from one node to another continuously.
You will notice that the active and backup roles for SRG1+ shift from node0 to node1 and vice versa continuously after every few minutes. This will cause traffic interruption as well. You can use below given command to monitor the cluster status:
user@srx>show chassis high-availability information
When you configure BFD monitoring for SRG1+ with aggressive timers in MNHA setup, there is a possibility that one of the devices might not be able to handle the load resulting in the BFD flaps randomly which will trigger the cluster failover again and time. In such scenarios it is recommended to try and configure BFD with less aggressive timers for the stability of the cluster.
You can use below given command to monitor the BFD sessions:
user@srx> show bfd session