Description

This article explains the meaning of the " fpc6 kernel: EDAC MPC85xx MC0: Faulty ECC bit: 0" error logs messages on the QFX10016/PTX Series devices.

Symptoms

Users may see the following log messages in QFX10016/PTX Series device.
 

kernel: EDAC MPC85xx MC0: Faulty ECC bit: 0
 

When you review the complete log, you can obtain more details about the message as shown below:
 

kernel: EDAC MPC85xx MC0: Err Detect Register: 0x00000004
kernel: EDAC MPC85xx MC0: Faulty ECC bit: 0
kernel: EDAC MPC85xx MC0: Expected Data / ECC: 0x80000099_0c383000 / 0x7f
kernel: EDAC MPC85xx MC0: Captured Data / ECC: 0x00000099_0c383000 / 0x7e
kernel: EDAC MPC85xx MC0: Err addr: 0x42fc1680
kernel: EDAC MPC85xx MC0: PFN: 0x00042fc1

kernel: EDAC MC0: 1 CE mpc85xx_mc_err on mc#0csrow#0channel#0 (csrow:0 channel:0 page:0x42fc1 offset:0x680 grain:8 syndrome:0x7e)


 

Solution

As you can see in the logs, its indicates that there was an error in the memory subsystem of your Juniper QFX switch, The error was caused by faulty ECC bit.

The error occurred in the EDAC MPC85xx MC0 Module, which is responsible for detecting and correcting error in the memory.

Workaround:

- If the error logs occur only once, you can ignore it. If it is being reported frequently on the same FPC card, reboot the card once during the MW because these are transient errors that are usually get rectified after a reboot.

- However, if the error continues even after reboot, the module or card where is error is being reported may need replacement.

Modification History

2024-01-08 : Article Created

Related Information

https://www.juniper.net/documentation/en_US/release-independent/junos/topics/concept/routing-engine-memory-failure-overview-nog.html