This article explains why a Major Error alarm with the syslog message "ZTCHIP_MQSS_CMERROR_MALLOC_INT_REG_OCFREELIST_OVF" might be seen on MX Series routers and what should be done to resolve the error.
ZTCHIP_MQSS_CMERROR_MALLOC_INT_REG_OCFREELIST_OVF log message represents an effect rather than the cause of a issue condition. Check the preceding logs to determine the issue cause.Below logs outline a scenario where FPC 0 was encountering WO interrupt register packet error in the ZTCHIP MQSS block due to packet integrity failure. Eventually this lead to On-chip memory allocation [Malloc] signature error during indirect read in one or more chunk descriptors and overflow of On-chip memory freelist, thereby exceeding the available pointer thresholds.Aug 13 22:30:01 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): WO: Packet Error - Error Packets 255, Connection 32 (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_WO_INT_REG_PKT_ERR)Aug 13 22:30:11 lab_mx960 : %PFE-3: fpc1 MQSS(0): FI: Error cell sent to reorder engine - Stream 0, Count 4Aug 13 22:30:13 lab_mx960 : %PFE-3: fpc0 MQSS(0): FO: Packet Error - Error Packets 15, Stream 4, Port 2Aug 13 22:30:13 lab_mx960 : %PFE-3: fpc0 ZT[0:0]: MCIF Rx: Checksum error detected on WO response - Chunk Address 0x3f82c79Aug 13 22:30:13 lab_mx960 : %PFE-3: fpc0 ZT[0:0]: MCIF Rx: Injected checksum error detected on WO response - Chunk Address 0x290c3c9Aug 13 22:30:13 lab_mx960 : %PFE-3: fpc0 ZT[0:0]: MCIF Tx1: Detected out-of-range request - data64_err_log 0x38a430410400Aug 13 22:30:13 lab_mx960 : %PFE-3: fpc0 MQSS(0): DRD: UNROLL0: MC chunk length error in stage 4 - Chunk Address: 0x3f80c62Aug 13 22:30:13 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): WO: Packet Error - Error Packets 255, Connection 12 (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_WO_INT_REG_PKT_ERR)Aug 13 22:32:03 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): WO: Packet Error - Error Packets 255, Connection 10 (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_WO_INT_REG_PKT_ERR)Aug 13 22:32:03 lab_mx960 : %PFE-3: fpc1 MQSS(1): FI: Error cell sent to reorder engine - Stream 0, Count 6Aug 13 22:32:03 lab_mx960 : %PFE-3: fpc0 PFE 1: Possible Remote PFE fault detected, total 1 times since up.Aug 13 22:32:03 lab_mx960 : %PFE-3: fpc0 PFE 1: exceeding aggr threshold 100 curr/last err_cell (317/217)Aug 13 22:32:03 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: CM[1]: MPC fabric remote PFE error (aggregate based) (1) exceed raising threshold (1) occurrance (2) for module/pfe (10:1) (URI: /fpc/0/pfe/0/cm/0/CM[1]/1/CM_CMERROR_FABRIC_REMOTE_PFE_AGGR)Aug 13 22:32:05 lab_mx960 alarmd[14380]: %DAEMON-4: Alarm set: FPC id=167772264, color=YELLOW, class=CHASSIS, reason=FPC 0 Minor ErrorsAug 13 22:32:05 lab_mx960 craftd[13159]: %DAEMON-4: Minor alarm set, FPC 0 Minor Errors Aug 13 22:32:57 lab_mx960 : %PFE-5: fpc0 CMError: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQPTR (0x227f79), scope: pfe, category: functional, severity: minor, module: MQSS(0), type: DRD_RORD: Detected: PQ pointer memory, oc_category: defaultAug 13 22:32:57 lab_mx960 : %PFE-5: fpc0 Performing action log for error /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQPTR (0x227f79) in module: MQSS(0) with scope: pfe category: functional level: minor, oc_category: defaultAug 13 22:32:57 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): DRD: RORD0 Protect: Parity error detected for PQ pointer memory - data32_log_err 0x1, data32_log_address 0x6c1 (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_DRD_RORD_ENG_SRAM_PAR_PROTECT_FSET_REG_DETECTED_PQPTR)Aug 13 22:34:36 lab_mx960 : %PFE-3: fpc0 user.err ztchip-luss: dispatch_event_handler(1085): ZT[0:0].slice[0].disp[2] TARGET_ZONE_INACTIVE.Aug 13 22:37:41 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): WO: Packet Error - Error Packets 255, Connection 32 (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_WO_INT_REG_PKT_ERR)Aug 13 22:37:41 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): MALLOC: OC deallocation error when free list FIFO is full (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_MALLOC_INT_REG_OCFREELIST_OVF)Aug 13 22:37:41 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: XQSS(0): XQ-SS[0]: CPQW Freelist Manager receives more free pointers than size of free list (URI: /fpc/0/pfe/0/cm/0/XQSS(0)/0/XQSS_CMERROR_CPQW_ERR_INT_FSET_FL_OVERFLOW_ERR)Aug 13 22:37:41 lab_mx960 : %PFE-3: fpc1 MQSS(0): FI: Error cell sent to reorder engine - Stream 0, Count 532Aug 13 22:37:43 lab_mx960 : %PFE-3: fpc0 MQSS(0): MALLOC: OC signature error during indirect read in one or more chunk descriptors - Unroll 0, MCIF 0, Signature 1, Address 0x3fa7fbbAug 13 22:37:43 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: MQSS(0): MQSS(0): MALLOC: OC deallocation error when free list FIFO is full (URI: /fpc/0/pfe/0/cm/0/MQSS(0)/0/ZTCHIP_MQSS_CMERROR_MALLOC_INT_REG_OCFREELIST_OVF)Aug 13 22:37:43 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: XQSS(0): XQ-SS[0]: CPQW Freelist Manager receives more free pointers than size of free list (URI: /fpc/0/pfe/0/cm/0/XQSS(0)/0/XQSS_CMERROR_CPQW_ERR_INT_FSET_FL_OVERFLOW_ERR)Aug 13 22:37:44 lab_mx960 : %PFE-3: fpc0 user.err packetio: [Error] Wedge-Detect : Host Loopback Wedge Detected: PFE: 0Aug 13 22:37:44 lab_mx960 : %PFE-3: fpc0 user.err packetio: [Error] Wedge-Detect : Host Loopback Wedge error block: error counter: HostAug 13 22:37:44 lab_mx960 : %PFE-3: fpc0 Cmerror Op Set: HstLbk:pfe:0: Host loopback timeout for pfe: 0 (URI: /fpc/0/pfe/0/cm/0/HstLbk:pfe:0/0/CMERROR_HOST_LOOPBACK_DECLARE_WEDGE) ---> Finally resulting in a wedge condition which would blackhole traffic
The "ZTCHIP_MQSS_CMERROR_MALLOC_INT_REG_OCFREELIST_OVF" message reports an On-chip free list overflow condition. It is an internal software state issue in the ZTCHIP related to pointer management. The reported error "ZTCHIP_MQSS_CMERROR_MALLOC_INT_REG_OCFREELIST_OVF" is a functional issue impacting the PFE and is usually seen amid other error logs which could mean that PFE is already in an unstable/issue state.
This memory-related issue is transient and can be rectified by resetting the affected PFE/FPC.
Perform these steps to determine the cause and resolve the problem (if any).
Collect the show command output.
Analyze the show command output.