Description

During a planned maintenance to replace Legacy RE with NGRE, customer noticed FPC's were offlined due to reduced fabric BW condition and the box was unreachable. RE 1 was powered off and CB 1 was taken offline post which NGRE was installed. 

RE1 powered off

 command 'request system power-off other-routing-engine '

CB1 offlined 

Jul 10 09:10:08 acb_take_offline: CB 1, offlines, offlining planes

Jul 10 09:15:46 chassisd[19339]: %DAEMON-5-CHASSISD_SNMP_TRAP7: SNMP trap generated: Fru Online (jnxFruContentsIndex 9, jnxFruL1Index 2, jnxFruL2Index 0, jnxFruL3Index 0, jnxFruName Routing Engine 1, jnxFruType 6, jnxFruSlot 1)

-Immediately after taking CB1 offline, all FPC’s online, reported degraded fabric condition and after 10 seconds, FPC’s were taken offline.

Jul 10 09:10:14 send: red alarm set, device FPC 0, reason FPC 0 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 1, reason FPC 1 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 7, reason FPC 7 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 8, reason FPC 8 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 9, reason FPC 9 degraded fabric condition detected

Symptoms


Jul 10 09:10:12 FM: fm_hsl2_info.nplanes: 6, fm_hsl2_get_num_active_planes: 4
Jul 10 09:10:12 FM: Plane Sate: 1 1 0 0 1 2 0 0; staggered_pmask: 15 2a 00 00 00 00 00 00
Jul 10 09:10:12 FM: fm_hsl2_info.nplanes: 6, fm_hsl2_get_num_active_planes: 3
Jul 10 09:10:13 FM: Plane Sate: 1 1 0 0 1 4 0 0; staggered_pmask: 15 2a 00 00 00 00 00 00          -> plane 5 was still activating
Jul 10 09:10:13 FM: fm_hsl2_info.nplanes: 6, fm_hsl2_get_num_active_planes: 3
Jul 10 09:10:14 FM: Plane Sate: 1 1 0 0 1 4 0 0; staggered_pmask: 15 2a 00 00 00 00 00 00        -> plane 5 was still activating
Jul 10 09:10:14 FM: fm_hsl2_info.nplanes: 6, fm_hsl2_get_num_active_planes: 3
Jul 10 09:10:30 FM: Plane Sate: 1 1 0 0 1 1 0 0; staggered_pmask: 15 2a 00 00 00 00 00 00          -> plane 5 got activated but it was too late/post 'degraded fabric condition detected'
Jul 10 09:10:30 FM: fm_hsl2_info.nplanes: 6, fm_hsl2_get_num_active_planes: 4
Jul 10 09:10:31 FM: Plane Sate: 1 1 0 0 1 1 0 0; staggered_pmask: 15 2a 00 00 00 00 00 00

Jul 10 09:10:14 send: red alarm set, device FPC 0, reason FPC 0 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 1, reason FPC 1 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 7, reason FPC 7 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 8, reason FPC 8 degraded fabric condition detected
Jul 10 09:10:14 send: red alarm set, device FPC 9, reason FPC 9 degraded fabric condition detected
Jul 10 09:10:14 FH: Detected that few/all of "offline-on-fabric-bandwidth-reduction" or "bandwidth-degradation" knob configured FPC(s) are running in reduced fabric BW
Jul 10 09:10:14 FH: Action will be taken after 10 sec

Solution

-From the session logs, once NGRE was installed, CB0 and CB1 was showing Online Master/Standby with RE 1 as Standby. And since the FPC’s started reporting fabric degradation immediately after CB1 was taken offline, planes 4 and 5 from CB 2 had to switch Online, which was also confirmed from the chassisd logs.

-Now from the time CB 1 was taken offline at 09:10:08 till the time FPC’s started reporting degraded fabric condition at Jul 10 09:10:14, and planes 4 and 5 being activated, we shouldn’t have seen this condition. So, on further inspecting the configuration and logs, the direct reason why FPCs got offlined when CB1 was brought down should be due to below config knob.

fpc 0 {
       pic 0 {
           inline-services {
               bandwidth 1g;
           }
       }
       offline-on-fabric-bandwidth-reduction;             <==
       sampling-instance sample-fpc0;
       inline-services {
           flow-table-size {
               ipv4-flow-table-size 14;
               ipv6-flow-table-size 1;
           }
       }
   }

Recommendation for any upcoming maintenance windows for the said RE upgrade activity is to remove the knob “offline-on-fabric-bandwidth-reduction” and enable post replacment.

Modification History

2025-07-14 : Article Created

2026-06-14: Typo correction and mark as external