Description

On MX routers (such as the MX240 MX480 & MX960) that are fitted with MX SCB or Enhanced MX SCB Control Boards, during certain time intervals, a bogus Control Board Failure will be reported via a major alarm being raised.

From the moment this alarm is reported, subsequent chassis related activities, such as insertion/removal of certain FRUs or online/offline actions performed by button press, will not be correctly detected by the system.

Symptoms

Additional logic was introduced within chassisd (in Junos 10.4R11, 11.4R4, 12.1R3, 12,2R1, and 12.3R1), to detect when too many interrupts occur. This change affects MX routers (such as the MX240, MX480, and MX960) that are fitted with MX SCB or Enhanced MX SCB Control Boards. 

This logic is designed to detect when two or more interrupts occur within one second. When this event occurs, a Major Alarm is raised and further interrupts are disregarded by the system. This logic is intended to detect known interrupt storm conditions and prevent other undesired behavior, which can be triggered by these excessive interrupts. Hardware interrupts are generated, when any of the following events occur:

  • A Button is pressed on the front panel of the router.
  • The RE online/offline button is pressed (this button is located on the RE).
  • The FPM ACO (Alarm Cut Off) button is pressed.
  • An FRU in the following list of FRUs is inserted into ore removed from the router:

    • PEM
    • FAN TRAY
    • CB
    • RE
    • FPC
A flaw in this element of software, at certain router uptime periods, indicates that the system can no longer effectively detect the advancement of time. This flaw means that the system, specifically the chassisd process, incorrectly detects excessive interrupts and this prevents further interrupts from being correctly handled.

The time periods, when this condition and the bogus alarms are noticed, are at intervals of 24.85 days. Interrupts are correctly and incorrectly handled at these intervals, based on routing-engine uptime:

Time Interrupt handling
0 to 24.85 days Interrupts are handled normally
24.85 to 49.7 days Interrupts may be incorrectly handled
49.7 to 74.55 days Interrupts are handled normally
74.55 to 99.4 days Interrupts may be incorrectly handled
This sequence of 24.85 periods of correct interrupt handling, followed by periods of incorrect interrupt handling, is indefinitely repeated.

During these blackout periods (when interrupts may be incorrectly handled), a series of up to 3 successive interrupts (as detailed above), will trigger a major CB alarm. From this trigger point onwards, the router no longer reacts correctly to subsequent valid interrupts. This means that no more hardware insertions, hardware removals, online button presses, or offline button presses will be processed by the router. If three such interrupts occur in the blackout periods, the following signatures will be observed.

The following log is reported in the message log file, when the bogus alarm is triggered:
Jan 14 09:48:04 router chassisd[20697]: fpm_atlas_acb_storm_state_change: acb storm active
Jan 14 09:48:04 router alarmd[1444]: Alarm set: CB color=RED, class=CHASSIS, reason=CB 0 Failure
Jan 14 09:48:04 router craftd[1445]: Major alarm set, CB 0 Failure
The following alarm is reported by the system, when the bogus alarm is triggered:
username@router> show chassis alarms
1 alarm currently active
Alarm time Class Description
2013-01-14 09:48:04 CET Major CB 0 Failure
The following additional log message is reported in the chassisd log, when the bogus alarm is triggered:
Jan 14 09:48:04 send: red alarm set, device CB 0, reason CB 0 Failure
To trigger the above Major CB alarm, interrupts as described above are necessary. Here are some examples of logs reported in the chassisd log file for various sorts of interrupt. Logs similar to the following are reported in the chassisd log when a Button Press Offline interrupt is handled by chassisd.
Jan 8 15:55:20 fpm_atlas_acb_intr acb_ints_pending 0x1000020
Jan 8 15:55:20 fpm_atlas_I2CS_INT live_int 0x8000
Jan 8 15:55:20 fpm_atlas_int_I2CS_FPD handling FPM interrupt
Jan 8 15:55:20 fpm_atlas_int_I2CS_FPD: taking FPC 4 offline
Jan 8 15:55:20 CHASSISD_FRU_OFFLINE_NOTICE: Taking FPC 4 offline: Offlined by button press
Logs similar to the following are reported in the chassisd log, when a DPC insert interrupt is handled by chassisd:
Jan 9 18:15:54 fpm_atlas_acb_intr acb_ints_pending 0x10020
Jan 9 18:15:54 fpm_atlas_CH_PRS_CHG handling CH_PRS interrupt
Jan 9 18:15:54 fpm_atlas_acb_intr re-enabling interrupts (0x00010000)
Jan 9 18:15:54 re_kontron_gpio_intr: Button events pending 0x00000000
Jan 9 18:15:54 exit fpm_atlas_acb_intr ...
Jan 9 18:15:55 FPC 4 added
Logs similar to the following are reported in the chassisd log, when a DPC remove interrupt is handled by chassisd:
Jan 4 15:24:25 fpm_atlas_CH_PRS_CHG handling CH_PRS interrupt
Jan 4 15:24:25 fpm_atlas_acb_intr re-enabling interrupts (0x00010000)
Jan 4 15:24:25 re_kontron_gpio_intr: Button events pending 0x00000000
Jan 4 15:24:25 exit fpm_atlas_acb_intr ...
Jan 4 15:24:25 FPC 2 removed
The following examples, which could lead to undesired behavior, are of when the above symptom of bogus alarm is noticed and the major alarm for CB failure is active:

As FRU installation relies on interrupts, if a new FRU is installed, this FRU will not be recognized by the system and it will not be possible to bring it online; when installed.

Similarly, if an FRU upgrade is being performed, for example one FPC is being replaced by another of a different type, the system will not recognize the initial FRU removal and when another FRU of a differing type is installed, it will not be recognized; as it's predecessor's removal was never registered. When the different FPC type is brought online, it may lead to fabric drops; as different FPCs have different fabric interconnects. In the worst case scenario, this can trigger a chassisd restart, which affects all FPCs.

Another scenario could be the replacement of a defective component. When the defective FRU is removed, in a certain time period, this might trigger the above alarm and then no further interrupts will be recognized by the system. When this part is subsequently replaced, the alarm generated by the previously defective component may persist. Leading the operator to suspect the replacement FRU is also defective. If the condition is cleared, the component failure alarm and the bogus CB major failure alarm will both be cleared.

The solution for this bug is described in the Solution section that follows; but in the interim, there are two possible short term mitigation options:

You can clear the alarm and restore normal chassis interrupt handling by using the following command:
restart chassis-control immediately
This will restart chassisd , irrespective of Graceful Routing Engine Switchover (GRES) being enabled or not. FPCs should reconnect to the new instance of chassisd. The Major CB failure alarm will be cleared. The router will then react normally to the next one interrupt. If further interrupts occur and the system uptime is still within the blackout periods, then the router will encounter the same issue again. There is a small risk that this reconnect could fail, if the restart of the chassis takes longer than expected, which could lead to the FPC restarting and subsequent traffic loss.

If Graceful Routing Engine Switchover (GRES) is configured, you could mitigate the issue by adhering to another approach. After the Major CB Failure alarm is raised, restart the backup routing engine and switch mastership; so that this newly restarted routing engine becomes the primary. As the uptime of this routing engine, which has newly become the primary, is not in the blackout period, it will correctly detect and react to further interrupts; until the RE enters the blackout period after a further 24.85 days.

It is not advised to restart the chassisd process with any argument, other than immediately . Do not use the following commands to recover from this issue, as these will not clear the condition and some of them may cause FPCs to restart; which leads to an impact on traffic through the router.
restart chassis control
restart chassis-control soft
restart chassis-control gracefully

Solution

The incorrect interrupt handling in the blackout periods was an error in software, which is specifically related to measuring time intervals between successive interrupts.

This issue is tracked via PR823969 and PSN-2013-01-811 . It has been resolved in 10.4R13, 11.4R7, 12.1R6, 12.2R4, and 12.3R1 or higher.

Related Information