Description

This article describes the issue of the flapping of the fabric links on both of the nodes in the SRX1400 chassis cluster.

Symptoms

If SRX1400 is running any of the 11.2R1 to R4 code, there is a possibility of this issue occurring.

To identify the issue, perform the following procedure:


  1. Find out if the device is showing any vmcores. For example:
    SRX1400> show system core-dumps no-forwarding
    
    -rw------- 1 root wheel 206791753 Dec 16 02:18 /var/crash/core-vmcore-spu10.core.0.20111216.0213.gz
    -rw------- 1 root wheel 205615386 Dec 16 06:27 /var/crash/core-vmcore-spu10.core.0.20111216.0621.gz
    -rw------- 1 root wheel 206755423 Dec 16 09:11 /var/crash/core-vmcore-spu10.core.0.20111216.0906.gz
    -rw------- 1 root wheel 205070725 Dec 16 10:34 /var/crash/core-vmcore-spu10.core.0.20111216.1028.gz
    

    These cores are very frequent and occur on both nodes.
  2. Verify if the following log messages exist in the JSRPD log:
    Dec 15 16:24:10 SPU failure detected. suspending fabric monitoring, clearing coldsync status local copy
    and kernel ssam blob and clearing ifstate download status of CP
    

    You can also notice the following logs on JSRPD:
    Dec 15 16:30:32 Detected 1 NPC down out of 1 total NPCs. Setting npc-mon-weight to 255
    Dec 15 16:30:32 fabric timer never started dead timer not stopped
    Dec 15 16:31:43 fab1 marked down
    Dec 15 16:31:43 fab1 process
    Dec 15 16:31:43 jsrpd_ifd_msg_handler: Interface fab1 is going down
    Dec 15 16:31:43 stop monitoring for fab1 child ge-4/0/8 is down
    
  3. Chassisd logs show the following errors:

    Dec 16 10:28:33 LCC: fpc_a40_recv_pic_status : FPC 1 PIC 0: report error - SPU core dump -
    Dec 16 10:28:33 LCC: setting CP state to Down
    Dec 16 10:28:33 LCC: fpc_a40_recv_pic_status: On LCC sending pic_status with error 38 to SCC
    Dec 16 10:28:38 LCC: rcv: ch_ipc_dispatch() null ipc read for args 0x2356da0 pipe 0x2355b00
    Dec 16 10:28:42 LCC: fpc_a40_recv_pic_status : FPC 1 PIC 0: report error - SPU core dump -
    Dec 16 10:28:42 LCC: setting CP state to Down
    Dec 16 10:34:08 LCC: send: red alarm set, device FPC 1 PIC 0, reason FPC 1 PIC 0 SPU core dump complete
    Dec 16 10:34:08 CHASSISD_SNMP_TRAP7: SNMP trap generated: Fru Failed (jnxFruContentsIndex 7, jnxFruL1Index 6, 
    jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName node1 FPC: SRX1k Dual Wide NPC+SPC Support Card @ 1/*/*, jnxFruType 3, 
    jnxFruSlot 5)
    Dec 16 10:34:08 CHASSISD_PIC_OFFLINE_NOTICE: Taking PIC 0 in FPC 1 offline: SPU core dump complete
    Dec 16 10:34:08 LCC: fpc_a10_recv_pic_offline_req : SPU(FPC 1 PIC 0) report error: SPU core dump complete,
     all FPCs will be reset
    Dec 16 10:34:08 LCC: fpc_a10_power_cycle_all_fpcs : taking all FPCs offline and reset CPP
    Dec 16 10:34:08 CHASSISD_FRU_OFFLINE_NOTICE: Taking FPC 0 offline: Soft-reset on SPC/SPU failure
    Dec 16 10:34:08 LCC: I2CS write cmd to FPC#0 [0x0], reg 0x12, cmd 0x44
    Dec 16 10:34:08 LCC: fpc_down slot 0 reason Soft-reset on SPC/SPU failure cargs 0x2356d20
    Dec 16 10:34:08 LCC: fpc_a10_disconnect slot is 0
    


    If all of the above matches, then you have the same issue.

Solution


  • This issue is fixed in 11.2R5.
  • It is recommended to upgrade to 11.2R5 or 11.4R1 or later.
  • The final option is to return to 11.1.