This article explains the solution for the "kernel: mce: [Hardware Error]: Machine check events logged" message seen on the host-log of an RE which went for an abnormal reboot
The customer reported the “Backup RE Active” Alarm
The Master RE rebooted with the last reboot reason as “0x80000:CPU Reset”
The host logs of the specific RE show "kernel: mce: [Hardware Error]: Machine check events logged"
Snap
====
From “show chassis routing-engine” output:
root@re1> show chassis routing-engine no-forwarding
Routing Engine status:
Slot 0:
Current state Backup
Temperature 28 degrees C / 82 degrees F
DRAM 11671 MB (12288 MB installed)
Memory utilization 13 percent
5 sec CPU utilization:
User 1 percent
Background 0 percent
Kernel 4 percent
Interrupt 1 percent
Idle 95 percent
Model RE-QFX10016
Serial ID BCAE2968
Start time 2023-03-24 05:07:54 UTC
Uptime 21 minutes, 36 seconds
Last reboot reason 0x80000:CPU Reset
Load averages: 1 minute 5 minute 15 minute
0.25 0.43 0.39
From host-logs of RE0[previous master RE]:
2023-03-24T04:49:33.536633+00:00 localhost monit[8951]: 'mcelog' trying to restart
2023-03-24T04:49:33.536650+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service
2023-03-24T04:49:35.562021+00:00 localhost logger: MCE: 0 0 0 0 0 0 0 0 Machine check exceptions
2023-03-24T04:50:05.699211+00:00 localhost monit[8951]: 'mcelog' trying to restart
2023-03-24T04:50:05.699231+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service
2023-03-24T04:50:37.858875+00:00 localhost monit[8951]: 'mcelog' trying to restart
2023-03-24T04:50:37.858890+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service
2023-03-24T04:51:10.019018+00:00 localhost monit[8951]: 'mcelog' trying to restart
2023-03-24T04:51:10.019034+00:00 localhost monit[8951]: 'mcelog' start: /sbin/service
2023-03-24T05:09:05.867226+00:00 msr01b monit[8951]: 'mcelog' trying to restart
2023-03-24T05:09:05.867265+00:00 msr01b monit[8951]: 'mcelog' start: /sbin/service
2023-03-24T05:10:55.423989+00:00 msr01b kernel: mce: [Hardware Error]: Machine check events logged
-data corruption detected in the CPU caches, in main memory by an integrated memory controller,-data transfer errors on the front side bus or CPU interconnect or other internal errors-As the RE reboots after this error and works fine, Monitor the RE for the same issue. If the same issue is seen on the same RE, We may need to replace the RE.
Revision 1