This article explains the reason for the alarm 'Loss of communication with Backup RE' on the MX104 device and how to recover.
When the master RE lost it's keep-alives messages from the backup, we will see the alarm ‘Loss of communication with Backup RE0’ and similar messages would appear:
Mastership:
Nov 29 09:16:34 event = E_NO_IPC, state = master, param = 0x0x0
Nov 29 09:16:34 currentAction = A_WARN1
Nov 29 09:16:34 No response from the other routing engine for the last 2 seconds.
Nov 29 09:16:34 Currentstate master NextState master reason_code 1
Nov 29 09:16:34 new state = master
Nov 29 09:16:38 mcontrol_send_re_info: length is now 672 bytes for RE info
Nov 29 09:16:40 mcontrol_ore_alive_set: other RE is not alive
Nov 29 09:16:40 ORE not alive, keepalive loss 2 above threshold of 1
Nov 29 09:16:43 ORE not alive, keepalive loss 3 above threshold of 1
Nov 29 09:16:46 ORE not alive, keepalive loss 4 above threshold of 1
Nov 29 09:16:49 ORE not alive, keepalive loss 5 above threshold of 1
Nov 29 09:16:52 ORE not alive, keepalive loss 6 above threshold of 1
Nov 29 09:16:54 failed to receive keepalives from other RE for the last 20 sec
Nov 29 09:16:55 ORE not alive, keepalive loss 7 above threshold of 1
Messages:
Nov 29 09:16:43.323 MX104 alarmd[3132]: Alarm set: RE id=1761673229, color=YELLOW, class=CHASSIS, reason=Loss of communication with Backup RE
Nov 29 09:16:43.328 MX104 craftd[2075]: Minor alarm set, Loss of communication with Backup RE
The loss of communication can be usually be caused by any of the following scenarios:
Workaround is to reboot the RE during the MW time.
If the issue persists, then we may need to check the CB and by connecting a spare RE (if available to cross check, if the hardware is faulty).
KB71237 [juniper.net]: https://supportportal.juniper.net/s/article/RE-not-Active-at-GSJKTN001