Description

 

This article describes the issue with RE switchover when RE0 went offline, but mastership switchover was not switched to backup RE (RE1).

Symptoms

  • Customer reported, there was no ping response to MX2010, leading to affected traffic.
  • Consequently, a technician was dispatched to the site to establish a direct console connection.
  • Upon arrival, RE1 was online while RE0 was offline with its LED lights off.
  • During RE0's downtime, RE1 failed to attain mastership, resulting in router inaccessibility.

 

Logs:

 

RE1 –

 

May 1 14:02:37.097  snmpd[23688]: LIBJSNMP_NS_LOG_INFO: INFO: send_trap: RE currently not a global-master, discarding trap 

May 1 14:02:37.097  chassisd[25601]: CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 9, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 0, jnxFruType 6, jnxFruSlot0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)

 

Mastership - 

 

 

May 1 14:02:51 ORE not alive, keepalive loss 1 above threshold of 0

May 1 14:02:51 mcontrol_ore_alive_set: other RE is alive              >>>>>>>> RE1 detects RE0 is still alive 

May 1 14:02:51 event = E_ORE_M, state = master, param = 0x0x99b7d08

May 1 14:02:51 currentAction = A_CKM1

May 1 14:02:51 Duplicate Master Routing Engine                >>>>>>>>> Duplicate mastership detected as RE0 failed to relinquish its mastership

May 1 14:02:51 The local routing engine gives up mastership.

May 1 14:02:51 mcontrol_shutdown

 

 

 

RE0 –

 

May 1 14:02:05.524  snmpd[29418]: SNMPD_AUTH_FAILURE: nsa_log_community: unauthorized SNMP community from 10.20.48.102 to 10.140.198.67 (b9hFAPAHa)

May 1 14:02:05.524  snmpd[29418]: SNMPD_AUTH_FAILURE: nsa_log_community: unauthorized SNMP community from 10.20.48.102 to 10.140.198.67 (73ujF5Xm0KzwdgPqavAlInxMoOHs1VN8)

May 1 14:02:07.000  /var/run/scripts/jet/agentjunos[78367]: time="2024-05-01T14:02:07.882960655Z" level=info msg="Received GET request=/version, remoteAddr=100.88.5.160:57520, path=/version" 

May 1 14:02:07.000  /var/run/scripts/jet/agentjunos[78367]: time="2024-05-01T14:02:07.883037598Z" level=info msg="Agent version is 2.4.1-8239f2e" 

May 1 16:11:26.236  eventd: sendto: No route to host

May 1 16:11:26.234 eventd[18830]: SYSTEM_ABNORMAL_SHUTDOWN: System abnormally shut down

May 1 16:11:26.242  eventd[18830]: SYSTEM_OPERATIONAL: System is operational

 

Mastership - 

 

 

Mar 14 19:57:57 ORE not alive, keepalive loss 1 above threshold of 0

Mar 14 19:57:57 mcontrol_ore_alive_set: other RE is alive

Mar 14 19:58:52 mcontrol_ore_alive_set: other RE is not alive

Mar 14 19:58:52 ORE not alive, keepalive loss 1 above threshold of 0

Mar 14 19:58:53 mcontrol_ore_alive_set: other RE is alive

May 1 16:11:26 CHASSISD release 20.3X75-D34.6 built by builder on 2022-07-28 22:25:27 UTC

May 1 16:11:34 *** mcontrol init V01 ***

May 1 16:11:34 soft-restart: is not a master

May 1 16:11:34 initial vc state: initializing

May 1 16:11:34 vc vccb init, mid=255 sn= slots=8

May 1 16:11:34 mcontrol_ore_alive_set: other RE is not alive

May 1 16:11:34 Socket = 0x00000035

May 1 16:11:34 mcontrol hipri thread created

May 1 16:11:34 *** re_priority is 1***

May 1 16:11:34 init master recon flag TRUE

May 1 16:11:34 event = E_CFG_M, state = init, param = 0x0x0

May 1 16:11:34 currentAction = A_REQC

 

Solution

 

Both RE chassis and mastership logs need to be verified to ensure any hardware issues are observed on the affected RE.

 

From Chassid at the issue time, RE0/CB0 was in a bad state where it was not completely down but was not functioning correctly.

 

Logs:

 



May 1 14:02:37 CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 9, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 0, jnxFruType 6, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)

 

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Disable Cause 0x00000000

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Disable Cause 0x00000000

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Disable Cause 0x00000044

 

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003

May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003

 

May 1 14:02:39 CB#0 power not verified on in 650 ms

 

May 1 14:02:39 CHASSISD_SNMP_TRAP10: SNMP trap generated: FRU power on (jnxFruContentsIndex 12, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName CB 0, jnxFruType 5, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn -1240977574)

 

 

From the logs we can see at this time there was a RE switchover attempted and RE1 tried to take mastership, but it seems that even when RE1 became master the other RE (RE0) did not correctly relinquish the mastership due to which Duplicate master logs were reported and RE1 deferred the mastership and became backup again. After this it was unable to acquire the mastership again even after consecutive keepalives were missed with other RE until RE0/CB0 were manually reseated.

 

 

Modification History

2024-05-24 : Article Created