This article describes the issue with RE switchover when RE0 went offline, but mastership switchover was not switched to backup RE (RE1).
Logs:
RE1 – May 1 14:02:37.097 snmpd[23688]: LIBJSNMP_NS_LOG_INFO: INFO: send_trap: RE currently not a global-master, discarding trap May 1 14:02:37.097 chassisd[25601]: CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 9, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 0, jnxFruType 6, jnxFruSlot0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0) Mastership - May 1 14:02:51 ORE not alive, keepalive loss 1 above threshold of 0 May 1 14:02:51 mcontrol_ore_alive_set: other RE is alive >>>>>>>> RE1 detects RE0 is still alive May 1 14:02:51 event = E_ORE_M, state = master, param = 0x0x99b7d08 May 1 14:02:51 currentAction = A_CKM1 May 1 14:02:51 Duplicate Master Routing Engine >>>>>>>>> Duplicate mastership detected as RE0 failed to relinquish its mastership May 1 14:02:51 The local routing engine gives up mastership. May 1 14:02:51 mcontrol_shutdown RE0 – May 1 14:02:05.524 snmpd[29418]: SNMPD_AUTH_FAILURE: nsa_log_community: unauthorized SNMP community from 10.20.48.102 to 10.140.198.67 (b9hFAPAHa) May 1 14:02:05.524 snmpd[29418]: SNMPD_AUTH_FAILURE: nsa_log_community: unauthorized SNMP community from 10.20.48.102 to 10.140.198.67 (73ujF5Xm0KzwdgPqavAlInxMoOHs1VN8) May 1 14:02:07.000 /var/run/scripts/jet/agentjunos[78367]: time="2024-05-01T14:02:07.882960655Z" level=info msg="Received GET request=/version, remoteAddr=100.88.5.160:57520, path=/version" May 1 14:02:07.000 /var/run/scripts/jet/agentjunos[78367]: time="2024-05-01T14:02:07.883037598Z" level=info msg="Agent version is 2.4.1-8239f2e" May 1 16:11:26.236 eventd: sendto: No route to host May 1 16:11:26.234 eventd[18830]: SYSTEM_ABNORMAL_SHUTDOWN: System abnormally shut down May 1 16:11:26.242 eventd[18830]: SYSTEM_OPERATIONAL: System is operational Mastership - Mar 14 19:57:57 ORE not alive, keepalive loss 1 above threshold of 0 Mar 14 19:57:57 mcontrol_ore_alive_set: other RE is alive Mar 14 19:58:52 mcontrol_ore_alive_set: other RE is not alive Mar 14 19:58:52 ORE not alive, keepalive loss 1 above threshold of 0 Mar 14 19:58:53 mcontrol_ore_alive_set: other RE is alive May 1 16:11:26 CHASSISD release 20.3X75-D34.6 built by builder on 2022-07-28 22:25:27 UTC May 1 16:11:34 *** mcontrol init V01 *** May 1 16:11:34 soft-restart: is not a master May 1 16:11:34 initial vc state: initializing May 1 16:11:34 vc vccb init, mid=255 sn= slots=8 May 1 16:11:34 mcontrol_ore_alive_set: other RE is not alive May 1 16:11:34 Socket = 0x00000035 May 1 16:11:34 mcontrol hipri thread created May 1 16:11:34 *** re_priority is 1*** May 1 16:11:34 init master recon flag TRUE May 1 16:11:34 event = E_CFG_M, state = init, param = 0x0x0 May 1 16:11:34 currentAction = A_REQC
Both RE chassis and mastership logs need to be verified to ensure any hardware issues are observed on the affected RE.
From Chassid at the issue time, RE0/CB0 was in a bad state where it was not completely down but was not functioning correctly.
Logs: May 1 14:02:37 CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 9, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 0, jnxFruType 6, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0) May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Disable Cause 0x00000000 May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Disable Cause 0x00000000 May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Disable Cause 0x00000044 May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power Up State 0x0000002c May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 1 Cause 0x0000007f May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003 May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003 May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003 May 1 14:02:38 tiny_i2cs_power_ok: CB-0 Power VFail 2 Cause 0x00000003 May 1 14:02:39 CB#0 power not verified on in 650 ms May 1 14:02:39 CHASSISD_SNMP_TRAP10: SNMP trap generated: FRU power on (jnxFruContentsIndex 12, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName CB 0, jnxFruType 5, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn -1240977574)
From the logs we can see at this time there was a RE switchover attempted and RE1 tried to take mastership, but it seems that even when RE1 became master the other RE (RE0) did not correctly relinquish the mastership due to which Duplicate master logs were reported and RE1 deferred the mastership and became backup again. After this it was unable to acquire the mastership again even after consecutive keepalives were missed with other RE until RE0/CB0 were manually reseated.