Description

Minimal down time upgrade failed, customer mentioned the upgrade steps were followed correctly.

Symptoms

When the secondary node was rebooted during the "request software add path no-validate reboot" command, the primary node stopped passing traffic.

Traffic was restored once node1 was promoted to primary and it's interfaces were enabled.

 

The following logs are found post upgrade:

 

Jun 19 08:06:37 ch_get_chassis_ops_srxtvp: Failed sysctl for hw.product.pvi.config.mac_address.from_interface.

Jun 19 08:06:37 fpc_pepsi_init_power_limits slot 0

Jun 19 08:06:37 FPC absent. Nothing to do

 

These indicates the nodes were not correctly isolated during the upgrade process.

Solution

Please bear in mind during minimal down time upgrade, the nodes must be isolated either disconnecting the cables or through the configuration.

Disable the network interfaces on the backup device. This is performed to isolate the unit from the network so that it will not impact traffic when the upgrade procedure is in progress.

Break control and fabric link communication paths by using configuration adjustments or physical cable removals to ensure that the nodes do not communicate with one another while on different Junos OS versions.

Please refer to the KB17947 [juniper.net] for the full steps on how to proceed with the upgrade correctly:

[SRX] How to upgrade an SRX cluster with minimal down time?

 

Modification History

2024-07-04 : Article Created