Description

We have just experienced a pair of SRX345s fail to failover while node0 was unresponsive. Node0 was unresponsive via the terminal server while the outage occurred. Node1 was responsive and healthy. The cluster did not failover from node0 to node1 automatically. A manual failover of redundancy groups to node1 via the terminal server was required.

Symptoms

  • FPCs are flapped due a High memory consumption and not being able to connect to the device via SSH,
  • Not able to log on to the SRX via SSH.

Solution

 

 

 

The issue is that FPCs is flapped due a high memory consumption and not able to connect to the device via SSH, The FPC flapped at time 2023-03-20 18:13:01, We can't know when the issue started, as the recent messages file not covered before Aug 27 05:30:01,

 

Important Messages log,

Aug 27 05:30:01 ER1.GDC_SRX.node0 /kernel: rt_pfe_veto: Memory over consumed. Op 8 err 12, rtsm_id 0:-1, msg type 10, veto simulation: 0

Aug 27 05:30:01 ER1.GDC_SRX.node0 /kernel: rt_pfe_veto: free kmem_map memory = (20963328) curproc = rpd

Aug 27 05:32:46 ER1.GDC_SRX.node0 /kernel: rt_pfe_veto: Memory over consumed. Op 8 err 12, rtsm_id 0:-1, msg type 10, veto simulation: 0

Aug 27 05:32:46 ER1.GDC_SRX.node0 /kernel: rt_pfe_veto: free kmem_map memory = (20963328) curproc = rpd

Aug 27 05:32:46 ER1.GDC_SRX.node0 /kernel: rt_pfe_veto: Possible slowest client is fwdd0. States processed - 791174963. States to be processed - 3

Aug 27 05:32:46 ER1.GDC_SRX.node0 /kernel: rt_pfe_veto: Possible second slowest client is fwdd1. States processed - 791174963. States to be processed - 3

Aug 27 06:29:09 ER1.GDC_SRX.node0 /kernel: tcp_timer_keep: Dropping socket connection due to keepalive timer expiration, idle/intvl/cnt: 7200000/75000/8

Aug 27 06:29:09 ER1.GDC_SRX.node0 /kernel: tcp_timer_keep:Local(0x81100001:53860) Foreign(0x8f100001:33010)

Aug 27 06:29:46 ER1.GDC_SRX.node0 /kernel: kmem type ifstat using 313537K, exceeding limit 286720K

Aug 27 06:30:29 ER1.GDC_SRX.node0 /kernel: tcp_timer_keep: Dropping socket connection due to keepalive timer expiration, idle/intvl/cnt: 7200000/75000/8

Aug 27 06:30:29 ER1.GDC_SRX.node0 /kernel: tcp_timer_keep:Local(0x81100001:55390) Foreign(0x8f100001:33010)

Aug 27 06:30:46 ER1.GDC_SRX.node0 /kernel: kmem type ifstat using 313538K, exceeding limit 286720K

Aug 27 06:31:46 ER1.GDC_SRX.node0 /kernel: kmem type ifstat using 313538K, exceeding limit 286720K

Aug 27 06:31:49 ER1.GDC_SRX.node0 /kernel: tcp_timer_keep: Dropping socket connection due to keepalive timer expiration, idle/intvl/cnt: 7200000/75000/8

Aug 27 06:31:49 ER1.GDC_SRX.node0 /kernel: tcp_timer_keep:Local(0x81100001:50126) Foreign(0x8f100001:33010)

Aug 27 06:32:47 ER1.GDC_SRX.node0 /kernel: kmem type ifstat using 313538K, exceeding limit 286720K

 

All the above messages are repeated on all Message files,

 

This is a software issue related to PR1528605 which you can view on the following link, https://prsearch.juniper.net/problemreport/PR1528605, Your current software version is JUNOS [20.2R1.10] and this software issue is Resolved in below software versions,

 

junos:17.3R3-S11 junos:17.4R3-S5 junos:18.1R3-S13 junos:18.2R3-S7 junos:18.2R3-S8 junos:18.2X75-D34 

junos:18.2X75-D436 junos:18.2X75-D56 junos:18.2X75-D61 junos:18.2X75-D67 junos:18.3R3-S4 junos:18.4R2-S7 

junos:18.4R3-S6 junos:19.1R3-S4 junos:19.2R1-S6 junos:19.3R2-S6 junos:19.3R3-S1 junos:19.4R1-S4 

junos:19.4R2-S4 junos:19.4R3-S1 junos:20.1R2 junos:20.1R3 junos:20.2R2-S2 junos:20.2R3 junos:20.3R1-S2 

junos:20.3R2 junos:20.3X75-D10 junos:20.4R1 junos:21.1R1

 

 

Modification History

First Version
2023-10-19:  Article reviewed for accuracy; Minor non-technical changes needed.