Description

While performing JunOS Selective Update - RPKI 19.4R2-S3-J9.2 on MX240, CB1 gave major error after swaping mastership over to RE1, after moving mastership back to RE0 the error cleared. Protocols appear to be working, but needing health check of MX and potential root cause.

Symptoms

CB1 alarm during upgrade:

Apr 19 00:20:52 2024 re1 alarmd[81782]: Alarm set: CB id=16777590, color=RED, class=CHASSIS, reason=CB 1 Failure

 

Apr 19 00:47:29 2024 re1 alarmd[81782]: Alarm cleared: RE id=83886441, color=YELLOW, class=CHASSIS, reason=Backup RE Active

Apr 19 00:47:29 2024 re1 craftd[18054]: Minor alarm cleared, Backup RE Active

 

The RE's didn't go down during this:

 

hostname> show chassis routing-engine no-forwarding

 

Routing Engine status:

 Slot 0:

  Current state         Master

  Election priority       Master (default)

  Temperature         33 degrees C / 91 degrees F

  CPU temperature       33 degrees C / 91 degrees F

  DRAM           32704 MB (32768 MB installed)

  Memory utilization     21 percent

  Model             RE-S-1800x4

  Serial ID           9016290902

  Start time           2022-11-30 09:09:47 CST

  Uptime             505 days, 15 hours, 3 minutes, 34 seconds

  Last reboot reason       0x2:watchdog

  Load averages:         1 minute  5 minute 15 minute

                    0.17    0.26    0.41

Routing Engine status:

 Slot 1:

  Current state         Backup

  Election priority       Backup (default)

  Temperature         34 degrees C / 93 degrees F

  CPU temperature       32 degrees C / 89 degrees F

  DRAM           32704 MB (32768 MB installed)

  Memory utilization     20 percent

  Model             RE-S-1800x4

  Serial ID           9016290993

  Start time           2022-11-30 09:09:48 CST

  Uptime             505 days, 15 hours, 3 minutes, 27 seconds

  Last reboot reason       0x2:watchdog

  Load averages:         1 minute  5 minute 15 minute

                    0.16    0.20    0.38

 

Solution

The recent logs and RSI showed no further errors. The device appears healthy in the investigation. Advised monitoring for a week to confirm the error was transient. We monitored and no new alarms appeared.

Modification History

2024-04-30 : Article Created