Description

The secondary control links went down, and could not recover by changing the port transceivers or cables. We need to change the FPC to identify further 

Symptoms

The secondary control links went down, and could not recover by changing the port transceivers or cables. We need to change the FPC to identify a hardware/software failure with FPC.
 

Solution

Here lists the detailed steps on how to change the FPC on SRX5K series with minimum traffic loss: 
Note: traffic loss would be expected in step 11, please schedule a downtime to proceed 


1. Delete interface-monitor and preempt configurations on both nodes

# delete chassis cluster redundancy-group X interface-monitor
# delete chassis cluster redundancy-group X preempt

2. Enable the following features to avoid traffic loss after isolating the nodes

# set security flow tcp-session no-syn-check
# set security flow tcp-session no-sequence-check
# set security flow tcp-session no-syn-check-in-tunnel

3. Migrate all traffic from node 0 to node 1

> request chassis cluster failover redundancy-group X node 1
> request chassis cluster failover reset redundancy-group X

Note: please failover RG0 in the last as it takes the longest time to complete 

4. Perform a traffic check to ensure no issue with traffic flow on Node 1.

5. Remove/Unplug all traffic cables from Node 0.

6. Double check the traffic flow on Node 1.

7. Remove/unplug the HA and FAB cables on Node 0.

8. Power off Node 0(request vmhost halt/unplug the power cable)

9. Replace/change the FPC Card in a different slot accordingly 

10. Power on and perform health check on Node 0 without connecting cables.


> show chassis hardware 
> show chassis fpc 
> show chassis fpc pic-status 
> show chassis alarms 
> show chassis environment 
> show chassis routine-engine 
> show system core-dumps
> show system processes extensive 

11. Remove traffic cables from Node 1, and connect the traffic cables on Node 0 to physically failover the traffic from Node 1 to Node 0

Note: Please expect traffic loss here. We could not  predict the downtime period. 

12. Check the traffic flow session on Node 0 to ensure the correct traffic handling

13. Power off Node 1(request vmhost halt/unplug the power cable), and change the FPC slot accordingly

14. Power on node 1 and perform Health Check as Step 10

15. Reconnect the HA primary cable alone between Node 0 and Node 1 and both FAB cables

16. Check cluster status, and make sure they are under Primary-SecondaRy status 


> show chassis cluster interfaces
> show chassis cluster status 

17. Configure the secondary control link if it's needed

# set chassis cluster control-ports fpc X port x
# set chassis cluster control-ports fpc X port x

18. Reconnect the secondary control link between node 0 and node 1 if it's needed, and check the interface status 

> show chassis cluster interfaces

19. Reverse the changes made in Step 1&2 


Note: please mind the changes on FPC, they should match the new ones 

20. Confirm the device/link status to be good, then reconnect traffic cable on Node 1

Related guides for reference: 


- For the replacement/reseating of the FPC cards on the SRX5800 Please follow the below.
https://www.juniper.net/documentation/us/en/hardware/srx5800/topics/topic-map/srx5800-maintaining-line-cards-modules.html#id-maintaining-interface-cards-and-spcs-on-the-srx5800-services-gateway

- For the replacement/reseating of the FPC cards on the SRX5600 Please follow the below.
https://www.juniper.net/documentation/us/en/hardware/srx5600/topics/topic-map/srx5600-maintaining-line-cards-modules.html

- For the replacement/reseating of the FPC cards on the SRX5400 Please follow the below.
https://www.juniper.net/documentation/us/en/hardware/srx5400/topics/topic-map/srx5400-maintaining-line-cards-modules.html
 

Modification History

2024-03-18 : Article Created

Related Information

[SRX] FPC Card Replacement/ Reseat procedure for the SRX5k series. (juniper.net)