This article explains an RG0 failover failure scenario observed on mid-range and high-end SRX series devices.
This issue has been addressed in PR1904267 and is fixed in junos:20.2R3-S11, junos:21.4R3-S12, junos:22.4R3-S9, junos:23.2R2-S6, junos:23.4R2-S7, junos:24.2R2-S4, junos:24.4R2-S2, junos:25.2R1-S2, junos:25.2R2, junos:25.4R1, junos:26.1R1.
On all Junos OS SRX except branch SRX platforms, in a high availability (HA) SRX cluster scenario, when the Routing Engine (RE) of redundancy group 0 (RG0) switches over to a new primary RE, it waits for 16 seconds for each pfeman process (which acts as a PFE peer) in every Service Processing Unit (SPU) to reconnect to the RE on the backup node. If a pfeman (Packet Forwarding Engine Management) process fails to reconnect within that 16-second window, the RE closes the TCP connection to it. This results in the srxpfe (SRX Packet Forwarding Engine) failing to reconnect with the new primary RE (node 1) after the switchover from node 0 (previous primary) to node 1 (new primary).
2025-11-05 : Article Created
2026-08-20: Modified the article information and visibility