Description

This article explains the meaning of the "MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET" alarm that is seen on MX Series routers and indicates if any actions need to be taken.

 

Symptoms

The following log messages are seen on the Flexible PIC Concentrator (FPC):

 

Jun 1 22:27:46 TEST-MX : %PFE-3: fpc9 MQSS(1): BCMW CBUF: SRAM Protect 1: Multiple Errors 0x4 

Jun 1 22:27:46 TEST-MX : %PFE-5: fpc9 Error: /fpc/9/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM2 (0x22012f), scope: pfe, category: functional, severity: major, module: MQSS(1), type: BCMW_CBUF_SRAM_PAR1_PROTECT: Detected: Bank 0, Sub-

Jun 1 22:27:46 TEST-MX : %PFE-5: fpc9 Performing action get-state for error /fpc/9/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM2 (0x22012f) in module: MQSS(1) with scope: pfe category: functional level: major 

Jun 1 22:27:46 TEST-MX tftpd[44823]: %FTP-6: Filename: '/var/tmp/pfe_debug_commands'

Jun 1 22:27:46 TEST-MX tftpd[44823]: %FTP-6: Mode: 'octet'

Jun 1 22:27:46 TEST-MX tftpd[44823]: %FTP-6: 128.0.0.25: read request for /var/tmp/pfe_debug_commands: success

Jun 1 22:27:46 TEST-MX inetd[7012]: %DAEMON-4: Number of tftp connections at max limit (1)

Jun 1 22:27:46 TEST-MX tftpd[44825]: %FTP-6: Filename: '/var/tmp/pfe_debug_info_SMPC9'

Jun 1 22:27:46 TEST-MX tftpd[44825]: %FTP-6: Mode: 'octet'

Jun 1 22:27:46 TEST-MX tftpd[44825]: %FTP-6: 128.0.0.25: write request for /var/tmp/pfe_debug_info_SMPC9: success

Jun 1 22:27:47 TEST-MX : %PFE-3: fpc9 MQSS(1): BCMW CBUF: SRAM Protect 1: Multiple Errors 0x4 

Jun 1 22:27:57 TEST-MX last message repeated 10 times

Jun 1 22:27:58 TEST-MX : %PFE-5: fpc9 Performing action cmalarm for error /fpc/9/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM2 (0x22012f) in module: MQSS(1) with scope: pfe category: functional level: major 

Jun 1 22:27:58 TEST-MX : %PFE-5: fpc9 Performing action disable-pfe for error /fpc/9/pfe/0/cm/0/MQSS(1)/1/MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM2 (0x22012f) in module: MQSS(1) with scope: pfe category: functional level: major 

Jun 1 22:27:58 TEST-MX : %PFE-5: fpc9 PFE 1: 'PFE Disable' action performed. Bringing down ifd et-9/1/2 224 

Jun 1 22:27:58 TEST-MX : %PFE-6: fpc9 MQSS(1): CMACPCS0: Detected Ethernet MAC Local Fault Delta Event (et-9/1/2) 

Jun 1 22:27:58 TEST-MX : %PFE-6: fpc9 MQSS(1): CMACPCS0: Ethernet PCS Multilane Alignment Not Done Delta Event (et-9/1/2) 

Jun 1 22:27:58 TEST-MX : %PFE-6: fpc9 MQSS(1): CMACPCS0: Ethernet PCS per lane Block Not Locked Delta Event (et-9/1/2) - port_block_lock 0x0 

Jun 1 22:27:58 TEST-MX : %PFE-6: fpc9 ifp et-9/1/2 ifd_mdown 

Jun 1 22:27:58 TEST-MX : %PFE-5: fpc9 PFE 1: 'PFE Disable' action performed. Bringing down ifd et-9/1/5 225 

 

Multiple continuous occurrences indicate persistent underlying issues.

 

Same errors with similar tags:

 

  0x22012d  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM0
  0x22012e  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM1
  0x22012f  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM2
  0x220130  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK0_MEM3
  
  0x220131  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK1_MEM0
  0x220132  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK1_MEM1
  0x220133  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK1_MEM2
  0x220134  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK1_MEM3
  
  0x220135  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK2_MEM0
  0x220136  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK2_MEM1
  0x220137  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK2_MEM2
  0x220138  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK2_MEM3
  
  0x220139  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK3_MEM0
  0x22013a  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK3_MEM1
  0x22013b  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK3_MEM2
  0x22013c  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET1_REG_DETECTED_BNK3_MEM3
  
  0x22013d  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK4_MEM0
  0x22013e  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK4_MEM1
  0x22013f  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK4_MEM2
  0x220140  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK4_MEM3
  
  0x220141  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK5_MEM0
  0x220142  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK5_MEM1
  0x220143  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK5_MEM2
  0x220144  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK5_MEM3
  
  0x220145  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK6_MEM0
  0x220146  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK6_MEM1
  0x220147  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK6_MEM2
  0x220148  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK6_MEM3
  
  0x220149  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK7_MEM0
  0x22014a  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK7_MEM1
  0x22014b  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK7_MEM2
  0x22014c  MQSS_CMERROR_BCMW_CBUF_WI_SRAM_PAR_PROTECT_FSET2_REG_DETECTED_BNK7_MEM3

Solution

This is due to a transient hardware SRAM error with EACHIP.

 

The "Parity error corrected for buffer free list memory for bank" message reports a transient hardware error which was automatically corrected.

 

Since this is a major error "PFE Disable" action will be performed and the affected PFE will disabled causing the interfaces under that PFE to go down.

 

Perform these steps to determine the cause and resolve the problem (if any). Continue through each step until the problem is resolved.

 

1) Collect the show command output.

Capture the output to a file (in case you have to open a technical support case). To do this, configure each SSH client/terminal emulator to log your session.

 

show log messages

show log chassisd

show system errors active detail fpc <fpc>

start shell network pfe

show nvram

show syslog messages

show cmerror module brief  

show cmerror module <id> error 0x22012f <<<<<< replace id with module id from above output

exit

 

 

2) Analyze the show command output.

In the 'show log messages', review the events that occurred at or just before the appearance of the "Parity error corrected for buffer free list memory for bank" message. Frequently these events help identify the cause.

 

 

3) To clear the error condition and enable the affected PFE, need to reboot the FPC.

Take the FPC offline and bring it online again or restart the FPC during a maintenance window, to clear a possible improper state on the chip:

 

user@mx> request chassis fpc slot <> offline

Wait for a few minutes.

user@mx> request chassis fpc slot <> online

 

OR

 

user@mx> request chassis fpc slot <> restart

 

Please contact JTAC if the major errors alarm return even after the FPC reset.

Modification History

2024-06-18 : Article Created