Description

This article explains the solution for the "kernel: mce: [Hardware Error]: Machine check events logged" message seen on the host-log of an RE which went for an abnormal reboot

Symptoms

The customer reported the “Backup RE Active” Alarm

The Master RE rebooted with the last reboot reason as “0x80000:CPU Reset”

The host logs of the specific RE show "kernel: mce: [Hardware Error]: Machine check events logged" 

 

Snap

====

 

From “show chassis routing-engine” output:

root@re1> show chassis routing-engine no-forwarding

Routing Engine status:

 Slot 0:

  Current state         Backup

  Temperature        28 degrees C / 82 degrees F

  DRAM           11671 MB (12288 MB installed)

  Memory utilization     13 percent

  5 sec CPU utilization:

   User           1 percent

   Background        0 percent

   Kernel          4 percent

   Interrupt         1 percent

   Idle           95 percent

  Model             RE-QFX10016

  Serial ID           BCAE2968

  Start time          2023-03-24 05:07:54 UTC

  Uptime            21 minutes, 36 seconds

  Last reboot reason      0x80000:CPU Reset

  Load averages:        1 minute 5 minute 15 minute

                    0.25   0.43   0.39

 

From host-logs of RE0[previous master RE]:

 

2023-03-24T04:49:33.536633+00:00 localhost monit[8951]: 'mcelog' trying to restart

2023-03-24T04:49:33.536650+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service

2023-03-24T04:49:35.562021+00:00 localhost logger: MCE:      0      0      0      0      0      0      0      0   Machine check exceptions

2023-03-24T04:50:05.699211+00:00 localhost monit[8951]: 'mcelog' trying to restart

2023-03-24T04:50:05.699231+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service

2023-03-24T04:50:37.858875+00:00 localhost monit[8951]: 'mcelog' trying to restart

2023-03-24T04:50:37.858890+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service

2023-03-24T04:51:10.019018+00:00 localhost monit[8951]: 'mcelog' trying to restart

2023-03-24T04:51:10.019034+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service

2023-03-24T05:09:05.867226+00:00 msr01b monit[8951]: 'mcelog' trying to restart

2023-03-24T05:09:05.867265+00:00 msr01b monit[8951]: 'mcelog' start: /sbin/service

2023-03-24T05:10:55.423989+00:00 msr01b kernel: mce: [Hardware Error]: Machine check events logged

Solution

-data corruption detected in the CPU caches, in main memory by an integrated memory controller,

-data transfer errors on the front side bus or CPU interconnect or other internal errors

-As the RE reboots after this error and works fine, Monitor the RE for the same issue. If the same issue is seen on the same RE, We may need to replace the RE.

Modification History

Revision 1