Description

This article provides a brief guide to help troubleshooting scenarios in which one node of a chassis cluster becomes unresponsive/unreachable with no apparent reason.

Symptoms

  • Cluster instability.
  • Control link / fabric link losses.
  • Node unresponsive or not passing traffic.

Solution

If the device is completely unresponsive to console, SSH, management, etc., do a hard reboot and wait for the SRX to come back up.

 

Refer to the following KBs if issue persists:

 

KB6508 [juniper.net]: How to determine if an RMA is required

KB12790 [juniper.net]: How to recover when device fails to boot with the message 'can't load '/kernel'

 

If the device is recovered, collect RSI and /var/log and attempt to do an RCA:

 

  • Check on both nodes jsrpd folder to determine when the cluster started reporting control link losses, jitter, schedule slips if any. Most likely there will be failures seen from before the node went unresponsive.
  • Check on chassisd and messages folder logs of the affected node, check if any coredump attempts were logged, FPC failures, panic reboots, etc.

Modification History

2024-07-18 : Article Created