Description

Users may see the following alarms on Trio-based Flexible PIC Concentrators (FPCs) with a queuing chip (QX-chip) in MX-Series routers:

"Error retrieving q-node depth max details" and "QXCHIP(n): QX-chip[n]: drop [x] - SRAM single bit ECC error"

Symptoms

You may observe the following alarms:

May 17 12:13:36 MX960_RE0 alarmd[8177]: Alarm set: FPC color=YELLOW, class=CHASSIS, reason=FPC 1 Minor Errors
May 17 12:13:36 MX960_RE0 craftd[7613]: Minor alarm set, FPC 1 Minor Errors 
May 17 18:43:57 MX960_RE0 fpc1 COS_HALP(cos_halp_update_q_depth_values:1292): cos_halp_update_q_depth_values:Error retrieving q-node depth max details
May 17 18:43:57 MX960_RE0 fpc1 COS_HALP(cos_halp_update_q_stats:1098): cos_halp_update_q_stats : Error updating queue depth: scheduler 4, queue 6, cchip_id 1
May 17 18:43:58 MX960_RE0 fpc1 QX-chip(1): drop 0 - SRAM single bit ECC error
May 17 18:43:58 MX960_RE0 fpc1 Cmerror Op Set: QXCHIP(1): QX-chip[1]: drop 0 - SRAM single bit ECC error
May 17 18:44:05 MX960_RE0 fpc1 trinity_pio: 1 PIO errors occurred
May 17 18:44:05 MX960_RE0 fpc1 trinity_pio: Last error: 9 QXM-1 Trinity PCI 0x0014e0fc Read PCIe 0

Solution

This issue happened due to the SRAM single bit ECC error, which is a transient hardware issue. While retrieving the max-queue depth information, when an SRAM parity error is error is encountered, a syslog error is reported every few seconds.

The issue may be seen on Trio-based line cards (FPC) with a queuing chip (that is, QX, XQ, and XQSS chips, which are used by MX-MPC2E-3D-EQ and others). The problem is tracked in PR1479240.

There is no operational impact due to this issue and the severity of this syslog message has been moved from error to info level.

This kind of transient SRAM single bit ECC error normally can be cleared by a reset of the FPC.

If the issue persists after resetting the FPC, please open a case with Support to investigate further.

Modification History

Version 1.0