Description

This is article explains the reason behind MX304 going into Complete Hung state with Transient Memory related Issue

Symptoms

2024-05-25T07:32:08.131828+00:00 router-node kernel: mce: Uncorrected hardware memory error in user-access at 1769260000

2024-05-25T07:32:08.131829+00:00 router-node kernel: Memory failure: 0x1769260: Killing qemu-system-x86:11178 due to hardware memory corruption

2024-05-25T07:32:08.131829+00:00 router-node kernel: Memory failure: 0x1769260: Killing vhost-11178:11187 due to hardware memory corruption

2024-05-25T07:32:08.131833+00:00 router-node kernel: Memory failure: 0x1769260: huge page still referenced by 511 users

2024-05-25T07:32:08.131833+00:00 router-node kernel: Memory failure: 0x1769260: recovery action for huge page: Failed

Solution

When you observed a MX304 is not responding to SSH or Console, Need to engage site technician to power cycle the device and share the RSI and /var/logs from the device to JTAC.

Please share the vmhost logsb with JTAC to identify the root cause.
Follow the KB to collect the vmhost Logs : 
https://supportportal.juniper.net/s/article/MX-PTX-How-to-collect-var-log-files-from-Next-Generation-Routing-Engine-NG-RE

Modification History

2024-07-03 : Article Created