Description

This KB explains about the parity error seen in MMU block of PFE on QFX5k platforms and provides a way to troubleshoot the same.

Symptoms

Below error logs will be seen on the device when parity error in MMU are generated :-

 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 unit 0 MMU SC MEM PAR mem error interrupt count: 0 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 MMU SC MEM PAR 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 unit 0 MMU SC MEM PAR 1bit error at address 0x0d90211d 

Feb 19 06:44:07 device : %PFE-3: fpc0 ECC errors generated 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:soc_ser_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 SER_CORRECTION: reg/mem:11489 btype:11 sblk:4 at:-1 stage:0 addr:0x0d90211d port: 0 index: 8477 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 mem: 11489=MMU_REPL_LIST_TBL_PIPE0 blkoffset:43 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_entry_restore: 

Feb 19 06:44:07 device : %PFE-3: fpc0 CLEAR_RESTORE: MMU_REPL_LIST_TBL_PIPE0[11489] blk: mmu_sc0 index: 8477 : [4][d90211d] 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 MMU SC MEM PAR 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 unit 0 MMU SC MEM PAR 1bit error at address 0x0d902121 

Feb 19 06:44:07 device : %PFE-3: fpc0 ECC errors generated 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:soc_ser_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 SER_CORRECTION: reg/mem:11489 btype:11 sblk:4 at:-1 stage:0 addr:0x0d902121 port: 0 index: 8481 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 mem: 11489=MMU_REPL_LIST_TBL_PIPE0 blkoffset:43 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_entry_restore: 

Feb 19 06:44:07 device : %PFE-3: fpc0 CLEAR_RESTORE: MMU_REPL_LIST_TBL_PIPE0[11489] blk: mmu_sc0 index: 8481 : [4][d902121] 

Solution

Parity errors may occur due to:

  • Random radiation or particle hits the silicon dice 
  • EMI disturbances caused by high energy transmitters. The source can be either internal to the building or external radio towers, power lines, etc.
  • Power supply spikes caused by bad power regulation and variations in power. This can inject into the silicon through the power connections.
  • Parity errors are usually non permanent damage of the internal memories and registers and are correctable.​

 

These error can occur in different blocks of the PFE pipeline and this KB explains about the errors seen in Memory management unit of the PFE which is responsible for deep packet buffering etc.

 

Below error logs are seen :-

 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 unit 0 MMU SC MEM PAR mem error interrupt count: 0 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 MMU SC MEM PAR 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 unit 0 MMU SC MEM PAR 1bit error at address 0x0d90211d 

Feb 19 06:44:07 device : %PFE-3: fpc0 ECC errors generated >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> Parity Error occured

Feb 19 06:44:07 device : %PFE-3: fpc0 0:soc_ser_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 SER_CORRECTION: reg/mem:11489 btype:11 sblk:4 at:-1 stage:0 addr:0x0d90211d port: 0 index: 8477  >>>>>>>>>>>>>>>>>>>>>>> Parity error corrected

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 mem: 11489=MMU_REPL_LIST_TBL_PIPE0 blkoffset:43 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_entry_restore: 

Feb 19 06:44:07 device : %PFE-3: fpc0 CLEAR_RESTORE: MMU_REPL_LIST_TBL_PIPE0[11489] blk: mmu_sc0 index: 8477 : [4][d90211d] 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 MMU SC MEM PAR 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_tomahawk_ser_process_mmu_err: 

Feb 19 06:44:07 device : %PFE-3: fpc0 unit 0 MMU SC MEM PAR 1bit error at address 0x0d902121 

Feb 19 06:44:07 device : %PFE-3: fpc0 ECC errors generated 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:soc_ser_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 SER_CORRECTION: reg/mem:11489 btype:11 sblk:4 at:-1 stage:0 addr:0x0d902121 port: 0 index: 8481 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_correction: 

Feb 19 06:44:07 device : %PFE-3: fpc0 mem: 11489=MMU_REPL_LIST_TBL_PIPE0 blkoffset:43 

Feb 19 06:44:07 device : %PFE-3: fpc0 0:_soc_ser_mem_entry_restore: 

Feb 19 06:44:07 device : %PFE-3: fpc0 CLEAR_RESTORE: MMU_REPL_LIST_TBL_PIPE0[11489] blk: mmu_sc0 index: 8481 : [4][d902121] 

 

 

On QFX5k platforms automatic detection and correction of parity error is enabled so as we can see in above logs post detection of parity error in MMU,it was auto corrected by soft error recovery [SER] .However below steps must be followed when these error logs are seen :-

 

 

  • Collect below output from the device :-

 

user@device> request pfe execute command "set dcbcm bcmshell 'soc'" target fpc0

SENT: Ukern command: set dcbcm bcmshell 'soc'

 

 

HW (unit 0)

Unit 0 Driver Control Structure:

Chip=BCM56960_B1 Rev=0x12 Driver=BCM56960_A0

Flags=0x40107: attached initialized link-scan mem-clear-use-dma; board type 0x0

CM: Base=0x0

Disabled: reg_flags=0x100 mem_flags=0x0

SchanOps=-2023576587 MMUdbg=0 LinkPause=0

Counter: int=500000us per=2170000us dmaBuf=0x9354b030

EvictInvalidPoolIdCount=0

Timeout: Schan=0(300000us) MIIM=0(1000000us)

Intr: Total=258292427 Sc=0 ScErr=0 MMU/ARLErr=0

LinkStat=139 PCIfatal=0 PCIparity=0

ARLdrop=0 ARLmbuf=0 ARLxfer=0 ARLcnt0=0

TableDMA=0 TSLAM-DMA=0 CCM-DMA=0 SW=0

MemCmd[BSE]=0 MemCmd[CSE]=0 MemCmd[HSE]=0

ChipFunc[0]=0 ChipFunc[1]=0 ChipFunc[2]=0

ChipFunc[3]=0 ChipFunc[4]=0

FifoDma[0]=32 FifoDma[1]=18372 FifoDma[2]=0 FifoDma[3]=0

I2C=0 MII=0 StatsDMA=0 Desc=215758329 Chain=113090081 PciTimeOut=0

Error: SDRAM=0 CFAP=0 Fcell=0 MmuSR=0 l2FifoDMA=0

SER events(mem=3 reg=0 nak=0 stat=0 ecc=2 direct=1 fifo=1 tcam=0)

SER corrections(fix=0 clear=3 restore=0 special=0 err:0)

PKT DMA: dcb=t32 tpkt=443254286 tbyt=546006093 rpkt=512308929 rbyt=2685666666

DV: List: max-q=256 cur-tq=256 cur-rq=0 dv-size=256

DV: Statistics: allocs=81070834 frees=81070774 alloc-q=77198762

Mem cache (count=365 size=17742592 vmap size=150592 errmap size=8552)

Reg cache (count=30 size=3223200)

dma-ch-0 TX Idle Queue=0 (0x0) default intr mbm

dma-ch-1 RX Active Queue=20 (0xa870ecc0) default intr mbm

dma-ch-2 RX Active Queue=20 (0xa87d8880) intr mbm

dma-ch-3 RX Active Queue=20 (0xa88a36d8) intr mbm

dma-ch-4 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-5 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-6 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-7 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-8 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-9 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-10 -- Idle Queue=0 (0x0) intr no-mbm

dma-ch-11 -- Idle Queue=0 (0x0) intr no-mbm

 

 

Focus on the highlighted part in above output which provides a count of error and also the count of error fixed,corrected or cleared.

 

If the count of error generated  which is denoted by SER events and corrected which is denoted by SER corrections is same then that means errors were cleared by the SER logic .

SER events(mem=3 reg=0 nak=0 stat=0 ecc=2 direct=1 fifo=1 tcam=0)
SER corrections(fix=0 clear=3 restore=0 special=0 err:0)

 

  • In case SER is unable to correct the errors then power drain the device or perform hypervisor reboot and check the same SOC output again.

 

  • If errors are still seen then please open a Case with JTAC.

 

Modification History

2024-02-28 : Article Created