This article explains the meaning of HMC Fatal Error syslog messages and the action to take.
As per the syslog messages, the issue is on HMC7 (link2,3) on PECHIP1.
fpc0 INTR: throttle 3600sec PECHIP[2]:pe.irw.intr.status:do_not_transmit_in(0): (Count:11162172) fpc0 pechip_cmerror_set_error:2518: level = Fatal, cmerror_code = 0x2105ca, recover_error = 0, recovery_counter = 0, fh_message = 0x0 fpc0 INTR: throttle 60sec PECHIP[1]:pe.hmcif.link[2].intr.status:hmc_err(0): (Count:1) fpc0 INTR: throttle 60sec PECHIP[1]:pe.hmcif.link[3].intr.status:hmc_err(0): (Count:1) fpc0 PE Chip:PE-1[1]: HMCIF: Link2: HMC Fatal Error cmd:62 lng:1 ltag:1 dinv:0 errstat:127 err_cnt:0x40000000 fpc0 PE Chip:PE-1[1]: HMCIF: Link3: HMC Fatal Error cmd:62 lng:1 ltag:1 dinv:0 errstat:127 err_cnt:0x40000000 alarmd[4158]: Alarm set: FPC color=RED, class=CHASSIS, reason=FPC 0 Major Errors alarmd[4158]: Alarm set: FPC color=RED, class=CHASSIS, reason=FPC 0 Major Errors craftd[4159]: Receive FX craftd set alarm message: color: 1 class: 100 object: 104 slot: 0 silent: 0 short_reason=FPC 0 Major Errors long_reason=FPC 0 Major Errors id=150995048 reason=150994944 craftd[4159]: Receive FX craftd set alarm message: color: 1 class: 100 object: 104 slot: 0 silent: 0 short_reason=FPC 0 Major Errors long_reason=FPC 0 Major Errors id=150995048 reason=150994944 craftd[4159]: Major alarm set, FPC 0 Major Errors
This is a soft error caused by a single event in a portion of the HMC memory module. Sometimes, traffic blackholing can happen with this issue; other times, there is no explicit impact to traffic.
Collect the following output to investigate the cause of these errors:
HMC Error Logs
request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 0 hmc_err_log_0" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 1 hmc_err_log_0" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 2 hmc_err_log_0" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 3 hmc_err_log_0" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 4 hmc_err_log_0" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 5 hmc_err_log_0"
HMC Error Counters
request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 0 hmc_err_cnt" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 1 hmc_err_cnt" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 2 hmc_err_cnt" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 3 hmc_err_cnt" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 4 hmc_err_cnt" request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 5 hmc_err_cnt"
request pfe execute target fpc0 command "show hmc asic"
Examine if there are any errors in the pechip register:
XXXXXXXXXXXXXXXXXX> request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 0 hmc_err_log_0" SENT: Ukern command: bringup jspec read pechip[1] register hmcif link 0 hmc_err_log_0 GOT: GOT: GOT: 0x01ee0060 pe.hmcif.link[0].hmc_err_log_0 00000000 GOT: cmd[29:24] : 0x0 GOT: lng[23:20] : 0x0 GOT: ltag[16:8] : 0x0 GOT: dinv[7:7] : 0x0 GOT: errstat[6:0] : 0x0 LOCAL: End of file {master:0} XXXXXXXXXXXXXXXXXX> request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 1 hmc_err_log_0" SENT: Ukern command: bringup jspec read pechip[1] register hmcif link 1 hmc_err_log_0 GOT: GOT: GOT: 0x01ee4060 pe.hmcif.link[1].hmc_err_log_0 00000000 GOT: cmd[29:24] : 0x0 GOT: lng[23:20] : 0x0 GOT: ltag[16:8] : 0x0 GOT: dinv[7:7] : 0x0 GOT: errstat[6:0] : 0x0 LOCAL: End of file {master:0} XXXXXXXXXXXXXXXXXX> request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 2 hmc_err_log_0" SENT: Ukern command: bringup jspec read pechip[1] register hmcif link 2 hmc_err_log_0 GOT: GOT: GOT: 0x01ee8060 pe.hmcif.link[2].hmc_err_log_0 3E10017F GOT: cmd[29:24] : 0x3e GOT: lng[23:20] : 0x1 GOT: ltag[16:8] : 0x1 GOT: dinv[7:7] : 0x0 GOT: errstat[6:0] : 0x7f LOCAL: End of file {master:0} XXXXXXXXXXXXXXXXXX> request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 3 hmc_err_log_0" SENT: Ukern command: bringup jspec read pechip[1] register hmcif link 3 hmc_err_log_0 GOT: GOT: GOT: 0x01eec060 pe.hmcif.link[3].hmc_err_log_0 3E10017F GOT: cmd[29:24] : 0x3e GOT: lng[23:20] : 0x1 GOT: ltag[16:8] : 0x1 GOT: dinv[7:7] : 0x0 GOT: errstat[6:0] : 0x7f LOCAL: End of file {master:0} XXXXXXXXXXXXXXXXXX> request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 4 hmc_err_log_0" SENT: Ukern command: bringup jspec read pechip[1] register hmcif link 4 hmc_err_log_0 GOT: GOT: GOT: 0x01ef0060 pe.hmcif.link[4].hmc_err_log_0 00000000 GOT: cmd[29:24] : 0x0 GOT: lng[23:20] : 0x0 GOT: ltag[16:8] : 0x0 GOT: dinv[7:7] : 0x0 GOT: errstat[6:0] : 0x0 LOCAL: End of file {master:0} XXXXXXXXXXXXXXXXXX> request pfe execute target fpc0 command "bringup jspec read pechip[1] register hmcif link 5 hmc_err_log_0" SENT: Ukern command: bringup jspec read pechip[1] register hmcif link 5 hmc_err_log_0 GOT: GOT: GOT: 0x01ef4060 pe.hmcif.link[5].hmc_err_log_0 00000000 GOT: cmd[29:24] : 0x0 GOT: lng[23:20] : 0x0 GOT: ltag[16:8] : 0x0 GOT: dinv[7:7] : 0x0 GOT: errstat[6:0] : 0x0 LOCAL: End of file
As expected, 0x7f error code is seen on link 2 and 3, which is mapping to HMC_7.
A small percentage of QFX10002 and QFX10008 line card may experience soft error in Hybrid Memory Cubic (HMC) memory module. The soft error was caused by a single event upset (SEU) in a portion of the HMC logic die and lead to a fatal error. Traffic null route on the particular FPC may happen during the event. Single Event Upsets due to soft errors is common to all memory devices. It does not signify that the memory is bad. SW resiliency is the method to mitigate. The failed system must be power-cycled in order to recover. If it boots, recovers, and continue to function normally, you can continue to use it.
FPC reset is needed to clear out the HMC error. As QFX10002 is not chassis based, the switch needs to be rebooted as a whole.
In some cases, HMC errors are caused by corrupt memory module. In such cases, you will see unexpected behavior post reboot and from logs we can see HW failure, or the FPC stay functional post reboot but the HMC error resurface after certain time. In such situations, please contact your JTAC Representative for a replacement.