Description

The PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR indicates a CRC error in the Egress Packet Write (EPW) block of the PE chip, where every egress packet undergoes a CRC check. Even a single CRC failure triggers a log, and the error clears automatically if no new errors occur within 60 seconds. These events are often linked to transient issues on SIB-to-FPC links and may not always indicate a persistent hardware fault.

 

There will be random PE Chip CRC error set and clear messages

 

 

Symptoms

Details of these events ---

  • Indicates a CRC error in the EPW block of the PE chipEPW is the egress packet write block within the pechip. 
  • Error /fpc/<slot>/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR.
  • CRC check will be performed on every packet in the EPW block on the egress path.
  • A single CRC failure triggers a Cmerror Op Set log. If the check fails, then a CRC error will be reported.
  • For the EPW CRC error, even if we see a single CRC error on EPW, it will be reported.
  • These events will be cleared if no new CRC errors are reported in the next 60 seconds in the EPW block. 

Aug 1 09:30:26.238 user-host fpc0 Cmerror Op Clear: PE Chip: PE0[0]: EPW: clear crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 09:35:25.299 user-host fpc0 Cmerror Op Set: PE Chip: PE0[0]: EPW: crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 09:36:46.659 user-host fpc0 Cmerror Op Clear: PE Chip: PE0[0]: EPW: clear crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 09:39:14.279 user-host fpc9 Cmerror Op Set: PE Chip: PE0[0]: EPW: crc error (URI: /fpc/9/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 09:40:36.452 user-host fpc9 Cmerror Op Clear: PE Chip: PE0[0]: EPW: clear crc error (URI: /fpc/9/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 09:49:08.273 user-host fpc1 Cmerror Op Set: PE Chip: PE0[0]: EPW: crc error (URI: /fpc/1/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 09:50:29.223 user-host fpc1 Cmerror Op Clear: PE Chip: PE0[0]: EPW: clear crc error (URI: /fpc/1/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 10:01:24.643 user-host fpc0 Cmerror Op Set: PE Chip: PE0[0]: EPW: crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 10:02:45.825 user-host fpc0 Cmerror Op Clear: PE Chip: PE0[0]: EPW: clear crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 10:07:47.075 user-host fpc11 Cmerror Op Set: PE Chip: PE4[4]: EPW: crc error (URI: /fpc/11/pfe/0/cm/0/PE_Chip/4/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 10:09:08.718 user-host fpc11 Cmerror Op Clear: PE Chip: PE4[4]: EPW: clear crc error (URI: /fpc/11/pfe/0/cm/0/PE_Chip/4/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 10:45:03.145 user-host fpc0 Cmerror Op Set: PE Chip: PE0[0]: EPW: crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 
Aug 1 10:46:23.813 user-host fpc0 Cmerror Op Clear: PE Chip: PE0[0]: EPW: clear crc error (URI: /fpc/0/pfe/0/cm/0/PE_Chip/0/PECHIP_CMERROR_EPW_MISC_INT_EVENTS_CRC_ERR) 

 


  • Possible reason for these events:

    • Most of the time, it's caused by transient intermittent issues on SIB-to-FPC links OR SPMB to FPC.
    • Errors may appear and clear randomly across different FPCs and SPMB to SIB. 
  • Troubleshooting Steps:

    • Run fabric health check scripts to trace and verify link health.
    • Check if errors are isolated to specific SIBs or FPCs.
    • If errors persist or impact traffic, escalate to JTAC for deeper analysis.

 

 

PEChip Errors: 6

--------------

PEChip 0

IGP errs received from IQM/DBM : 1

 

 

PEChip 2

IGP errs received from IQM/DBM : 2

 

 

Solution

The issue was resolved after performing the steps mentioned below, and the errors have not been seen again.

 

Pre-checks and Initial Data Collection

  1. Gather all pre-checks and the specified output listed below - RE0 data before reboot
  2. Then, proceed with a VMHost reboot of the Routing Engine RE1 - RE1 Reboot
  3. Switchover to Backup Routing Engine
    1. Initiate a switchover to the backup Routing Engine (RE1) - Mastership switchover to RE1
    2. Once all interfaces are fully populated and operational, collect the same set of data as mentioned earlier - RE1 Data after reboot
  4. Monitoring and RE0 VMHost Reboot
    1. Monitor both Routing Engines for one hour - Look for errors logs using "show log messages | match Cmerror"
    2. After confirming stability, proceed with the VMHost reboot of the Routing Engine RE0 - RE0 Reboot
  5. Mastership Restoration
    1. Continue monitoring both Routing Engines for another hour - Look for errors logs using "show log messages | match Cmerror"
    2. Then restore mastership back to RE0 - Mastership switchover to RE0
  6. Final Data Collection and Case Update
    1. Once RE0 has regained mastership and all interfaces are populated and operational, collect the same set of data again and upload it to the case - RE0 data after reboot

 

 

*************************************
If the customer is ready to collect data and investigate further, then below fabric health check plan can be used. It consists of an off-box PYTHON script triggered from the jump server to collect data. 

Fabric Health Check Plan ---

To perform a detailed analysis from a fabric health perspective, please follow the steps below:

1. Script Execution:

  • Use the provided Python script "get_fabric_debugs.py" OFF-box with ROOT access privileges.
  • Ensure execution is done under an approved planned maintenance window.
  • The script is non-intrusive and does not contain any system or process-impacting CLI commands.
  • Despite this, it is recommended to obtain prior approval before execution.

2. Routing Engine Switchover:

  • During the same maintenance window:
  • Perform a switchover to the backup Routing Engine (RE1).
  • Execute the same script with RE1 in the master state.
  • After data collection, restore mastership back to RE0.

3. Data Collection:

  • Collect script output three times:
  • First: With the current mastership (RE0).
  • Second: After switching mastership to RE1.
  • Third: After switching back to RE0.

4. Data Sharing:

  • Share all collected data with clear file names indicating the Routing Engine used during each collection.
  • Include observations and findings along with the data.

 

Using the OFF box script to collect data ---

Please use the script and YAML configuration file (attached to the case) to collect fabric debug data from the impacted device. The script needs to be run on a JUMP server where Python3 is installed and which can access the target router, and not from the target itself.

 

1. get_fabric_debugs.py (link to download compressed file of both files LINK TO DOWNLOAD script and yaml )

2. debug_config.yaml

 

Instructions to run the script:

1) Edit the YAML file "debug_config.yaml" with the ROOT ACCESS, username, password and RE hostnames of the device.

2) The script requires Python 3. Please replace the following line in the script with the path to your python3 binary

  #!/homes/sanjayh/opt/python-3.8.1/bin/python3

3) From the shell on the server, set the path to the YAML file as shown below

  export FABRIC_DEBUG_CONFIG_FILE=debug_config.yaml

4) The script can be run as shown below.

5) Depending on the number of fpcs and sibs, the script will take a while to complete.

 

./get_fabric_debugs.py CollectDebugData.test_get_basic > fabric_debug_data.txt

./get_fabric_debugs.py CollectDebugData.test_check_ccl_link_errors > fabric_ccl_link_test_result.txt

./get_fabric_debugs.py CollectDebugData.test_get_ccl_stats >> fabric_debug_data.txt

./get_fabric_debugs.py CollectDebugData.test_get_fpc_stats >> fabric_debug_data.txt

./get_fabric_debugs.py CollectDebugData.test_get_spmb_stats >> fabric_debug_data.txt

 

6) If the script complains about any missing Python modules, please use "pip3 install <module-name>" to install the missing modules. Or else, talk to your system administrator for help.

7) At the end of the test, please share the files "fabric_debug_data.txt" and "fabric_ccl_link_test_result.txt"

8) Please note that since the script collects data remotely one command at a time, it takes a while to complete. Please allow it to run to completion.
*************************************


 

Modification History

2025-10-13 : Article Created