Description

This article explains the one of reason behind all FPCs resetting during a backup RE reset and provides steps to determine which component causes this issue.

Symptoms

In this scenario, RE0 is the master and RE1 is the backup. Mastership was toggled to make RE1 the master. After the switch, RE0, now the backup, was restarted. Following this restart, all the FPCs on the box also restarted. Ideally, when the backup RE is rebooted, all FPC connections to the master RE should remain unaffected, and no restart should occur

 

  • RE0 is the master and RE1 is the backup.

Routing Engine status:
 Slot 0:
  Current state         Master
  Election priority       Master

Routing Engine status:
 Slot 1:
  Current state         Backup
  Election priority       Backup

 

  • Toggle mastership to make RE1 the master

Jul 29 06:29:11 Router mgd[90522]: %INTERACT-6-UI_CMDLINE_READ_LINE: User 'user', command 'request chassis routing-engine master acquire ' 

 

  • RE0 was rebooted

Jul 29 06:29:59 Router mgd[73740]: %INTERACT-6-UI_CMDLINE_READ_LINE: User 'user', command 'request system reboot '

 

  • After reboot, all the FPCs on the box restarted.

Jul 29 06:29:28 Router chassisd[21462]: %DAEMON-5-CHASSISD_SNMP_TRAP10: SNMP trap generated: redundancy switchover (jnxRedundancyContentsIndex 9, jnxRedundancyL1Index 2, jnxRedundancyL2Index 0, jnxRedundancyL3Index 0, jnxRedundancyDescr Routing Engine 1, jnxRedundancyConfig 3, jnxRedundancyState 2, jnxRedundancySwitchoverCount 3, jnxRedundancySwitchoverTime 1744537406, jnxRedundancySwitchoverReason 4)
Jul 29 06:30:09 Router chassisd[21462]: %DAEMON-5-CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 9, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 0, jnxFruType 6, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)
Jul 29 06:30:24 Router chassisd[21462]: %DAEMON-5-CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 7, jnxFruL1Index 2, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName FPC: DPCE 40x 1GE X @ 1/*/*, jnxFruType 3, jnxFruSlot 1, jnxFruOfflineReason 23, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)
Jul 29 06:30:24 Router chassisd[21462]: %DAEMON-5-CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 7, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName FPC: MPC 3D 16x 10GE @ 0/*/*, jnxFruType 3, jnxFruSlot 0, jnxFruOfflineReason 23, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)
Jul 29 06:30:25 Router chassisd[21462]: %DAEMON-5-CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 7, jnxFruL1Index 9, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName FPC: DPCE 40x 1GE X @ 8/*/*, jnxFruType 3, jnxFruSlot 8, jnxFruOfflineReason 23, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)
Jul 29 06:30:26 Router chassisd[21462]: %DAEMON-5-CHASSISD_SNMP_TRAP10: SNMP trap generated: Fru Offline (jnxFruContentsIndex 7, jnxFruL1Index 8, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName FPC: MPC 3D 16x 10GE @ 7/*/*, jnxFruType 3, jnxFruSlot 7, jnxFruOfflineReason 23, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)

Solution

 

  • Check how FPC are connected to master RE ie is it on EM0 or EM1 interface. In this case, logs on RE1(master) shows that all the FPC connectivity to RE1 was via EM1 interface, which is incorrect. As EM1 is the backup interface used when normal EM0 path is having some issue. Ideally when the RE1 is master, all the FPCs should connect RE1 via EM0 which is not the case here. 

 

  • Also after mastership was changed, you could see RE1 Chassisd log indicating RE1 to FPC0/1/7/8 is via em1 - backup RE red alarm being generated due to this. The alarm also indicates that connectivity to FPCs and RE1 was via EM1 not via EM0. 

 
(mastership was toggled and RE1 was master)
Jul 29 06:29:11 Router mgd[90522]: %INTERACT-6-UI_CMDLINE_READ_LINE: User 'user', command 'request chassis routing-engine master acquire '
 
(after RE1 become master, TNP entry for 0x00000017(FPC7), 0x00000010(FPC0), 0x00000018(FPC8) and 0x00000011(FPC1) were added to interface EM1 with hopcount of 1)
Jul 29 06:29:18 Router kernel: %KERN-6: TNPv3: adding remote neighbor entry 0x00000017 hopcount 1 to interface em1. 
Jul 29 06:29:18 Router kernel: %KERN-6: TNPv3: adding remote neighbor entry 0x00000010 hopcount 1 to interface em1.
Jul 29 06:29:20 Router kernel: %KERN-6: TNPv3: adding remote neighbor entry 0x00000018 hopcount 1 to interface em1
Jul 29 06:29:20 Router kernel: %KERN-6: TNPv3: adding remote neighbor entry 0x00000011 hopcount 1 to interface em1.
 
(Red alarm set on RE1 for all the FPCs)
 
Jul 29 06:29:28 send: yellow alarm set, device Routing Engine 1, reason Backup RE Active
Jul 29 06:29:28 send: yellow alarm set, device Routing Engine 1, reason Backup RE Active
Jul 29 06:29:33 ch_alarm_fpc_route_intf_slot: RE1 to FPC0 is via em1 - backup RE
Jul 29 06:29:33 send: red alarm set, device Routing Engine 1, reason RE1 to one or many FPCs is via em1: Backup RE
Jul 29 06:29:33 ch_alarm_fpc_route_intf_slot: RE1 to FPC1 is via em1 - backup RE
Jul 29 06:29:33 send: red alarm set, device Routing Engine 1, reason RE1 to one or many FPCs is via em1: Backup RE
Jul 29 06:29:33 ch_alarm_fpc_route_intf_slot: RE1 to FPC7 is via em1 - backup RE
Jul 29 06:29:33 send: red alarm set, device Routing Engine 1, reason RE1 to one or many FPCs is via em1: Backup RE
Jul 29 06:29:33 ch_alarm_fpc_route_intf_slot: RE1 to FPC8 is via em1 - backup RE
Jul 29 06:29:33 send: red alarm set, device Routing Engine 1, reason RE1 to one or many FPCs is via em1: Backup RE
 

  • So now, when the backup RE was rebooted, all the FPCs' chassisd connections are expected to drop because all the FPCs' connections are via backup RE/CB on the EM1 interface on RE1.

 

  • After the mastership switch, since all the FPC connections to Master is via EM1, we need further isolate if there is an issue with master RE's EM0 interface or with its CB.

 

  • Check if the EM0 interface on master RE is up and passing two way traffic, if yes then the EM0 interface is good.

At each RE, collect the following command outputs:
(several times with 'set cli timestamp' enabled) 'show interfaces em* extensive'

Collect the Ethernet switch outputs  At each RE, collect:
‘show chassis ethernet-switch’;
(several times with 'set cli timestamp' enabled) 'show chassis ethernet-switch errors';
(several times with 'set cli timestamp' enabled) 'show chassis ethernet-switch statistics';
'test chassis ethernet-switch shell-cmd "ps"';
(several times with 'set cli timestamp' enabled) 'test chassis ethernet-switch shell-cmd "show counters all"';

In this case, we saw Port 13 on SCB1, which connects to EM0 on master RE1, has RX byte counters that are not increasing. This indicates an issue with the CB1 switch.

user@Router> show chassis ethernet-switch statistics | no-more
Jul 30 22:28:41
Statistics for port 13 connected to device: RE-GigE
    TX Byte Counter                 1725395298
    RX Byte Counter                 2131618639

user@Router> show chassis ethernet-switch statistics | no-more   
Jul 30 22:29:02
Statistics for port 13 connected to device: RE-GigE
    TX Byte Counter                 1725401570
    RX Byte Counter                 2131618639


user@Router> show chassis ethernet-switch statistics | no-more   
Jul 30 22:29:21
Statistics for port 13 connected to device: RE-GigE
    TX Byte Counter                 1725407906
    RX Byte Counter                 2131618639


 

  • After replacing the CB1, no more FPC restarts occurred when the Backup RE was rebooted.

 

 

Modification History

2024-08-01 : Article Created

2026-06-15 : Modified hostname & username

2026-06-15: Mark to be external