Description

Customer reported flapping of afeb in their device leading to restart of all FPC

Symptoms

We could observe an ASIC error was detected on FPC0, resulting in a major alarm on FPC0


Jan 29 22:44:22.379 labroot chassisd[2083]: CHASSISD_FPC_ASIC_ERROR: <FPC 0> ASIC Error detected errorno 0x0003000b Restart action performed 


Jan 29 22:44:22.382 labroot alarmd[3046]: Alarm set: FPC id=1744830473, color=RED, class=CHASSIS, reason=FPC 0 Major Errors


Jan 29 22:44:22.383 labroot craftd[2086]: Major alarm set, FPC 0 Major Errors

Solution

After analyzing the logs, we observed that an ASIC error was detected on FPC0, triggering a major alarm and causing the interfaces to go down which subsequently led to the AFEB restart.


AFEB restart resulted in the restart of all FPCs.


Jan 29 22:44:19.362 labroot afeb0 CMError: /fpc/0/pfe/0/cm/0/MQ_Chip(0)/0/CMERROR_MQ_DDRIF_CHKSUM_ERR_MAJOR (0x3000b), scope: pfe, category: functional, severity: major, module: MQ Chip(0), type: DDRIF: Checksum errors


Jan 29 22:44:20.058 labroot afeb0 Performing action get-state for error /fpc/0/pfe/0/cm/0/MQ_Chip(0)/0/CMERROR_MQ_DDRIF_CHKSUM_ERR_MAJOR (0x3000b) in module: MQ Chip(0) with scope: pfe category: functional level: major


Jan 29 22:44:20.539 labroot afeb0 MQCHIP(0) DDRIF FO0 Checksum Error


Jan 29 22:44:20.818 labroot afeb0 MQCHIP(0) DDRIF Chksum Cnts Current 84, Total 385


Jan 29 22:44:22.278 labroot afeb0 Performing action cmalarm for error /fpc/0/pfe/0/cm/0/MQ_Chip(0)/0/CMERROR_MQ_DDRIF_CHKSUM_ERR_MAJOR (0x3000b) in module: MQ Chip(0) with scope: pfe category: functional level: major


Jan 29 22:44:22.379 labroot chassisd[2083]: CHASSISD_FPC_ASIC_ERROR: <FPC 0> ASIC Error detected errorno 0x0003000b Restart action performed 


Jan 29 22:44:22.382 labroot alarmd[3046]: Alarm set: FPC id=1744830473, color=RED, class=CHASSIS, reason=FPC 0 Major Errors


Jan 29 22:44:22.383 labroot craftd[2086]: Major alarm set, FPC 0 Major Errors


Jan 29 22:44:22.905 labroot chassisd[2083]: CHASSISD_IFDEV_DETACH_FPC: ifdev_detach_fpc(0)


Jan 29 22:44:23.446 labroot rpd[2813]: RPD_IFL_NOTIFICATION: EVENT [UpDown] ge-0/0/5.942 index 495 [Broadcast Multicast] address #0 c.86.10.bc.87.f2


Jan 29 22:44:23.943 labroot chassisd[2083]: CHASSISD_SNMP_TRAP10: SNMP trap generated: FRU power off (jnxFruContentsIndex 6, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName AFEB MX104, jnxFruType 5, jnxFruSlot 0, jnxFruOfflineReason 2, jnxFruLastPowerOff 0, jnxFruLastPowerOn 0)


Jan 29 22:44:31.587 labroot chassisd[2083]: CHASSISD_SNMP_TRAP7: SNMP trap generated: FRU insertion (jnxFruContentsIndex 6, jnxFruL1Index 1, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName AFEB, jnxFruType 5, jnxFruSlot 0)


Jan 29 22:44:34.268 labroot ppmd[2094]: PPMD_PFE_SHUTDWN: PPMD: Connection Shutdown/Closed with PFE: afeb0


This issue is caused by a memory read parity error condition, where the AFEB reported MQ-chip instruction memory parity error.


Typically, this is a transient hardware issue related to the MQ-chip instruction memory. To correct these errors, the AFEB restarted automatically , which subsequently led to the restart of all FPCs.  Customer has been advised to monitor the device for 48 hours. Device remains stable.



Modification History

2025-02-03 : Article Created