Description

This article explains the reason for the alarm 'Loss of communication with Backup RE' on the MX104 device and how to recover.

Symptoms

When the master RE lost it's keep-alives messages from the backup, we will see the alarm ‘Loss of communication with Backup RE0’ and similar messages would appear:


Mastership:

 

Nov 29 09:16:34 event = E_NO_IPC, state = master, param = 0x0x0

Nov 29 09:16:34 currentAction = A_WARN1

Nov 29 09:16:34 No response from the other routing engine for the last 2 seconds.

 

Nov 29 09:16:34 Currentstate master NextState master reason_code 1

Nov 29 09:16:34 new state = master

Nov 29 09:16:38 mcontrol_send_re_info: length is now 672 bytes for RE info

Nov 29 09:16:40 mcontrol_ore_alive_set: other RE is not alive

Nov 29 09:16:40 ORE not alive, keepalive loss 2 above threshold of 1

Nov 29 09:16:43 ORE not alive, keepalive loss 3 above threshold of 1

Nov 29 09:16:46 ORE not alive, keepalive loss 4 above threshold of 1

Nov 29 09:16:49 ORE not alive, keepalive loss 5 above threshold of 1

Nov 29 09:16:52 ORE not alive, keepalive loss 6 above threshold of 1

Nov 29 09:16:54 failed to receive keepalives from other RE for the last 20 sec

Nov 29 09:16:55 ORE not alive, keepalive loss 7 above threshold of 1

 

Messages:


Nov 29 09:16:43.323 MX104 alarmd[3132]: Alarm set: RE id=1761673229, color=YELLOW, class=CHASSIS, reason=Loss of communication with Backup RE

Nov 29 09:16:43.328 MX104 craftd[2075]: Minor alarm set, Loss of communication with Backup RE

Solution

The loss of communication can be usually be caused by any of the following scenarios:

  • Backup Routing Engine experienced an issue that prevented it from sending out keepalive messages, however the backup Routing Engine was otherwise operating as expected. 
  • Backup Routing Engine rebooted due to either a software or a hardware issue.
  • Backup Routing Engine keeps rebooting due to either a software or a hardware issue.
  • Backup Routing Engine failed.
  • CB (control board) that hosts the backup Routing Engine failed.
  • Master Routing Engine and backup Routing Engine were experiencing some sort of unidentifiable communication problem, but each Routing Engine was operating properly in all other respects.


Workaround is to reboot the RE during the MW time.

If the issue persists, then we may need to check the CB and by connecting a spare RE (if available to cross check, if the hardware is faulty).

Modification History

2024-12-04 : Article Created

Related Information

KB71237 [juniper.net]https://supportportal.juniper.net/s/article/RE-not-Active-at-GSJKTN001