Description

CGNAT stopped forwarding traffic, SPC3 flowed core

Nov 28 14:43:08 xxxx mib2d[16710]: SNMP_TRAP_LINK_DOWN: ifIndex 677, ifAdminStatus down(2), ifOperStatus down(2), ifName vms-0/1/0.32384

Nov 28 14:43:08 xxxx mib2d[16710]: SNMP_TRAP_LINK_DOWN: ifIndex 678, ifAdminStatus down(2), ifOperStatus down(2), ifName vms-0/1/0.32385

Nov 28 14:43:08 xxxx mib2d[16710]: SNMP_TRAP_LINK_DOWN: ifIndex 676, ifAdminStatus down(2), ifOperStatus down(2), ifName vms-0/1/0

 

Nov 28 14:42:41 xxxx (FPC Slot 0, PIC Slot 1) Global CPU zone change GREEN=>RED (99.99 %)

Nov 28 14:42:41 xxxx (FPC Slot 0, PIC Slot 1) CPU zone change GREEN=>RED (99.99 %)

Nov 28 14:42:41 xxxx (FPC Slot 0, PIC Slot 1) CPU utilization (99.99 percent) exceeded threshold 

Nov 28 14:42:41 xxxx rtlogd[16699]: RTLOGD_CPU_LIMIT_TRAP: FPC0:PIC1 CPU zone change GREEN=>RED (99.00 %)

Nov 28 14:42:59 xxxx fpc0 talus_intx_interrupt_handler Talus(1) common.int_status is 0x80 

Nov 28 14:42:59 xxxx fpc0 pic1: TALUS(1) PCIe(3) DMA RX interrupt received. Queue stuck status 0x1000000 

Nov 28 14:43:00 xxxx fpc0 talus_intx_interrupt_handler Talus(1) common.int_status is 0x80 

Nov 28 14:43:00 xxxx fpc0 pic1: TALUS(1) PCIe(3) DMA RX interrupt received. Queue stuck status 0x1000000 

Nov 28 14:43:01 xxxx fpc0 talus_intx_interrupt_handler Talus(1) common.int_status is 0x80 

Nov 28 14:43:01 xxxx fpc0 pic1: TALUS(1) PCIe(3) DMA RX interrupt received. Queue stuck status 0x1000000 

Nov 28 14:43:02 xxxx fpc0 talus_intx_interrupt_handler Talus(1) common.int_status is 0x80 

Nov 28 14:43:02 xxxx fpc0 pic1: TALUS(1) PCIe(3) DMA RX interrupt received. Queue stuck status 0x1000000 

Nov 28 14:43:03 xxxx fpc0 spu_spc_x86_64_check_heartbeat: fpc 0 pic1 flowd lost heartbeat for 5 consecutive seconds, 1 time detected, longest 5. 

Nov 28 14:43:03 xxxx fpc0 talus_intx_interrupt_handler Talus(1) common.int_status is 0x80 

Nov 28 14:43:03 xxxx fpc0 pic1: TALUS(1) PCIe(3) DMA RX interrupt received. Queue stuck status 0x1000000 

Nov 28 14:43:03 xxxx fpc0 pic1: TALUS(1) PCIe(3) DMA RX stuck for 5 seconds. Queue stuck status 0x1000000 

Nov 28 14:43:04 xxxx (FPC Slot 0, PIC Slot 1) user.notice root: PIC_REBOOT: Internal SA not enabled

Nov 28 14:43:04 xxxx (FPC Slot 0, PIC Slot 1) user.notice root: PIC_REBOOT: SPC3_PIC_SOFT_COREDUMP

Nov 28 14:43:04 xxxx (FPC Slot 0, PIC Slot 1) daemon.emerg srxpfe[4038]: --------------------------------------

Nov 28 14:43:04 xxxx (FPC Slot 0, PIC Slot 1) daemon.emerg srxpfe[4038]: ABORT! 

Nov 28 14:43:04 xxxx (FPC Slot 0, PIC Slot 1) 1 2024-11-28T14:43:04.599130+13:00 xxxx_re0-fpc0.pic1 - - [timeQuality tzKnown="1" isSynced="0"] Coredump started for xxxx_re0-fpc0.pic1-flowd_spc3.elf

Nov 28 14:43:06 xxxx (FPC Slot 0, PIC Slot 1) user.notice root: /usr/bin/pfe-app-wrapper: Stopping /etc/init.d/flowd_spc3

:

Nov 28 14:59:19 xxxx (FPC Slot 0, PIC Slot 1) 1 2024-11-28T14:59:19.552764+13:00 xxxx_re0-fpc0.pic1 - - [timeQuality tzKnown="1" isSynced="0"] Coredump has been created on FPC base-os. xxxx_re0-fpc0.pic1-flowd_spc3.elf

Nov 28 15:04:00 xxxx (FPC Slot 0, PIC Slot 1) 1 2024-11-28T15:04:00.253796+13:00 xxxx_re0-fpc0.pic1 - - [timeQuality tzKnown="1" isSynced="0"] Coredump has been transferred to RE. xxxx_re0-fpc0.pic1-flowd_spc3.elf

Symptoms

CGNAT stopped forwarding traffic, SPC3 flowed core

-rw-r--r-- 1 root wheel 3607867324 Nov 28 15:04 /var/crash/core-xxxx-fpc0.pic1-flowd_spc3.elf.0.tgz

Solution

The issue is fixed on PR-1841859:

 

https://prsearch.juniper.net/PR1841859


Release Notes:
On all MX Junos devices with SPC3, on receiving bursty traffic PIC may go down and cause network disruption.



Modification History

2025-01-22 : Article Created