Description

This article explains why a Major Error alarm with the syslog message "MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR" might be seen on MX Series routers and what should be done to resolve the error.

Symptoms

MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR indicates a transient error reported due to a parity error in SRAM memory.  Following logs may be seen during the issue period:

[Apr 4 00:27:29.573 LOG: Err] eachip_hmcif_rx_intr_handler(7239): EA[1:0]: HMCIF Rx: Checksum error detected on WO response - Chunk Address 0x19a55a

[Apr 4 00:29:09.793 LOG: Err] MQSS(1): WO: Packet Error - Error Packets 1, Connection 4

[Apr 4 00:29:09.793 LOG: Err] eachip_hmcif_rx_intr_handler(7239): EA[1:0]: HMCIF Rx: Checksum error detected on WO response - Chunk Address 0x59b3b6

[Apr 4 00:30:12.992 LOG: Info] MQSS(1): CHMAC0: Detected Ethernet MAC Remote Fault Delta Event for Port 6 (xe-0/1/1:2)

[Apr 4 00:30:12.992 LOG: Info] MQSS(1): CHMAC0: Cleared Ethernet MAC Remote Fault Delta Event for Port 6 (xe-0/1/1:2)

Apr 4 00:46:58.769 LOG: Notice] Performing action cmalarm for error /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQPTR (0x2203b3) in module: MQSS(1) with scope: pfe category: functional level: minor

[Apr 4 00:46:58.769 LOG: Debug] cmerror_take_action_helper: performing action 4 for scope 1 category 0 level 0 err_id /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQPTR(0x2203b3) module id 26

[Apr 4 00:46:58.769 LOG: Err] Cmerror Op Set: MQSS(1): MQSS(1): DRD: RORD0 Protect: Parity error detected for PQ pointer memory - data32_log_err 0x1, data32_log_address 0x74c

(URI: /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQP[Apr 4 00:46:59.548 LOG: Err] MQSS(1): WO: Packet Error - Error Packets 1, Connection 4

[Apr 4 00:46:59.548 LOG: Err] eachip_hmcif_tx_intr_handler(3531): EA[1:0]: HMCIF Tx: Out-of-range request detected - Client ID 21, Request Type 1, Request Size 7, Link Number 0, Request Address 0xbaccaf0

[Apr 4 00:46:59.549 LOG: Err] eachip_hmcif_rx_intr_handler(7239): EA[1:0]: HMCIF Rx: Checksum error detected on WO response - Chunk Address 0xbaa33a

[Apr 4 00:46:59.549 LOG: Debug] Cmerror: Draining ASIC error message queue

[Apr 4 00:46:59.549 LOG: Debug] cmerror_process_queue: module = MQSS(1)

[Apr 4 00:46:59.549 LOG: Debug] Cmerror: processing the task op_type 1 for scope 1 and category 0 level 1 level_count 0 occur_count 4 clear_count 4 level_threshold 1 level_action 0x7

item errid 2228301 item_threshold 1 item_count 0 item_sub_err_state 0 sub_item errid 0 sub_[Apr 4 00:46:59.549 LOG: Debug] Cmerror: Level 1 count increment 1 occur_count 5 clear_count 4

[Apr 4 00:46:59.550 LOG: Notice] CMError: /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR (0x22004d), scope: pfe, category: functional, severity: major, module: MQSS(1), type: BCMF ICM: MCIF I/F response with error during TAIL data write

[Apr 4 00:46:59.550 LOG: Debug] Cmerror: Level 1 count 1 (occur_count 5 clear_count 4)crossed threshold 1 action 0x7

[Apr 4 00:46:59.550 LOG: Debug] is_managed_external 1, category 0

[Apr 4 00:46:59.550 LOG: Debug] alarm count for level 1 is incremented and set to 1

[Apr 4 00:46:59.550 LOG: Debug] is_managed_external 1, category 0

[Apr 4 00:46:59.550 LOG: Debug] Cmerror category is NOT internal

[Apr 4 00:46:59.550 LOG: Notice] Performing action log for error /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR (0x22004d) in module: MQSS(1) with scope: pfe category: functional level: major

[Apr 4 00:46:59.550 LOG: Debug] cmerror_take_action_helper: performing action 1 for scope 1 category 0 level 1 err_id /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR(0x22004d) module id 26

[Apr 4 00:46:59.550 LOG: Notice] Performing action get-state for error /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR (0x22004d) in module: MQSS(1) with scope: pfe category: functional level: major

[Apr 4 00:46:59.550 LOG: Debug] cmerror_take_action_helper: performing action 2 for scope 1 category 0 level 1 err_id /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR(0x22004d) module id 26

 

Indications:

  • Usually, this error is a chain reaction to an original/underlying issue, so always check the preceding logs
  • This error is traffic impacting  and PFE may be disabled to avoid traffic blackhole
  • A log will be registered to notify the condition 

 

Solution

MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR error is after the effect of the DRD block SRAM parity error "MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQPTR". HMCIF (Hybrid Cube/Controller Interface) reports that the BCMF's (Buffer Cell Manager Fabric) ICM(Input Cell Manager) encountered an error in the DRD (Dispatch and Reorder] block while writing the packet.

 

Perform these steps to determine the cause and resolve the problem (if any). 

  • Collect the show command output

show log messages
show log chassisd
start shell network pfe <fpc#.0>
show nvram
show syslog messages
show cmerror modules brief ---> check for MQSS id [In current example error is in "MQSS(1)"] 
 
For instance:
FMPC7(bng-jtac-mx960-r2035-re0 vty)# sh cmerror module brief  
-------------------------------------------------------------------------
Module Name       Active Errors PFE   Callback  ModuleData
                     Specific Function       
-------------------------------------------------------------------------
26    MQSS(1)     0       Yes   0x149140800 0x3750258664


show cmerror module 26 error  <id> ---> replace id with error id [for instance 0x22004d from above logs]
 

SMPC0(router vty)# sh cmerror module 26 error 0x22004d

Error-id       : 0x22004d

Error Name        : MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR

Identifier        : /fpc/0/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMF_ICM_FI_INT_REG_HMCIF_TAIL_WRACK_ERR

Description        : BCMF ICM: MCIF I/F response with error during TAIL data write

State           : enabled

Scope           : pfe

Category         : functional

PFE            : 1

Configured Level     : Major

Default Level       : Major

Count           : 2

Threshold         : 1

Error Limit        : 0

Occur Count        : 2

Clear Count        : 0

OverItemThres Occur Count : 2

OverItemThres Clear Count : 0

Last-occurred(ms ago) : 3354698

Logs:

----------------------------------------------------------

Index Time         Sub-Err  State  Description

----------------------------------------------------------

0   04/04/25 00:47:11  0     Set   unknown

1   04/04/25 00:47:11  0     Set   unknown

 

  • Analyze the show command output.
  • In the 'show log messages', review the events that occurred at or just before the appearance of the error message. Frequently, these events help identify the cause.
    • No RMA is required
    • If the error is seen repeatedly, an FPC restart  during a maintenance window should clear this error
Contact your technical support representative if this issue transcends FPC restart
 
Tip: When looking at an event in the logs, it is important to focus on the first error message in a collection of syslog messages. The first error message is usually the cause of all the follow-on error messages. The follow-on collateral damage error messages can be ignored.

 

 

Modification History

2025-04-04 : Article Created