Customer complained that they were intermittently observing "RE1 offline" alarm for the backup RE.
Observed below log message on the master RE0:
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: Watchdog timeout Queue[0]-- resetting
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - Interface is RUNNING and ACTIVE
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: TX Queue 0 ------
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: hw tdh = 0, hw tdt = 0
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: Tx Queue Status = -2147483648
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: TX descriptors avail = 118
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: Tx Descriptors avail failure = 18833
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: RX Queue 0 ------
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: hw rdh = 0, hw rdt = 0
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: RX discarded packets = 0
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: RX Next to Check = 0
<2>1 2023-12-14T12:50:49.550Z MX960-RE0 kernel - - - em1: RX Next to Refresh = 0
<2>1 2023-12-14T12:50:53.655Z MX960-RE0 kernel - - - em1: Hardware Initialization Failed
This could also lead to both routing-engines showing different time as the backup RE is unable to sync with the master RE.
RE1 messages will also intermittently show RE0 offline messages:
<29>1 2023-12-14T14:52:19.254Z MX960-RE0 chassisd 4888 CHASSISD_SNMP_TRAP10 [[email protected] trap="Fru Offline" argument1="jnxFruContentsIndex" value1="9" argument2="jnxFruL1Index" value2="1" argument3="jnxFruL2Index" value3="0" argument4="jnxFruL3Index" value4="0" argument5="jnxFruName" value5="Routing Engine 0" argument6="jnxFruType" value6="6" argument7="jnxFruSlot" value7="0" argument8="jnxFruOfflineReason" value8="2" argument9="jnxFruLastPowerOff" value9="0" argument10="jnxFruLastPowerOn" value10="0"] SNMP trap generated: Fru Offline (jnxFruContentsIndex 9, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 0, jnxFruType 6, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)
Similar FRU offline messages will also appear on the RE0 for the backup RE1.
Collect below outputs:
show chassis ethernet-switch
show chassis ethernet-switch statistics
show chassis ethernet-switch port-state
show chassis ethernet-switch errors
show interfaces em0 extensive
show interfaces em1 extensive
show system virtual-memory
show system queues
show tnp addresses
Checked statistics for internal interfaces - em0 and em1 on both RE0 and RE1:
Physical interface: em1, Enabled, Physical link is Up
Input errors:
Errors: 0, Drops: 0, Framing errors: 0, Runts: 0, Giants: 0, Policed discards: 0, Resource errors: 0
Output errors:
Carrier transitions: 0, Errors: 20002, Drops: 0, MTU errors: 0, Resource errors: 0
The output errors on the Master RE0 em1 interface was continuously incrementing.
The em0 interface connects the routing-engine to the ethernet switch on the local Switch Control Board (SCB).
The em1 interface connects the routing-engine to the ethernet switch on the other Switch Control Board (SCB).
Please refer the article KB33394 [juniper.net] for more information on em interfaces on MX routers.
Since the errors were in the outgoing direction on RE0 em1 interface, suggested to restart the RE0.
If these errors were in ingress direction on the em1 interface then the remote SCB needs to be restarted (CB1 in this case).
If RE restart does not clear the issue, we need to proceed with RE replacement.
If the issue still persists, then the remote CB needs to be replaced.
Please contact JTAC, if the issue does not resolve with above steps, to determine any possible issues with the chassis midplane.