Description

This KB article provides fix details for the below continuous logs for memory threshold breaches and LAG flaps.


May 20 21:34:41 switch /kernel: %KERN-4: Percentage memory available(19)less than threshold(20 %)- 27619
May 20 21:35:36 switch /kernel: %KERN-4: Percentage memory available(19)less than threshold(20 %)- 27620
May 20 21:36:36 switch /kernel: %KERN-4: Percentage memory available(19)less than threshold(20 %)- 27621

May 20 21:36:42 switch /kernel: %KERN-4: lag_bundlestate_ifd_change: bundle ae7 is now Up. uplinks 48 >= min_links 48
May 20 21:36:45 switch /kernel: %KERN-4: lag_bundlestate_ifd_change: bundle ae132: bundle IFD minimum bandwidth or minimum links not met, Bandwidth (Current : Required) 0 : 40000000000 Number of links (Current : Required) 0 : 1
May 20 21:36:45 switch /kernel: %KERN-4: lag_bundlestate_ifd_change: bundle ae143: bundle IFD minimum bandwidth or minimum links not met, Bandwidth (Current : Required) 0 : 40000000000 Number of links (Current : Required) 0 : 1
May 20 21:36:45 switch /kernel: %KERN-4: lag_bundlestate_ifd_change: bundle ae123: bundle IFD minimum bandwidth or minimum links not met, Bandwidth (Current : Required) 0 : 40000000000 Number of links (Current : Required) 0 : 1
 

Symptoms

Memory threshold breach and LAG flapping-related log messages are seen without any physical interface flap.

Solution

We are hitting a couple of PRs here and these issues are fixed already. Please see the details below:

 

1. High CPU issue:

=========

The high CPU issue is seen because of the bug reported in the "PR:1554340:QFX5100 FPC CPU with a lot of spike on 18.4X16.7 compare with 14.1X53-D28.17"

 

  • This issue is seen on QFX platforms using Broadcom chipsets like Qfx5100/EX4600.
  • Triggers of this issue are:

 

1. On all QFX platforms using Broadcom chip.

2. QSFP implemented on the device.

3. 18.4X16.7 or later X release running on the device.

 

  • A high CPU is seen which might affect device performance.

Customer View: https://prsearch.juniper.net/problemreport/PR1554340

Resolved-In

junos:18.4R2-S9 junos:19.4R3-S8 junos:20.1R3-S4 junos:20.2R3-S4 junos:20.3R3-S3 junos:20.3X75-D43 junos:20.4R3 junos:21.1R2 junos:21.1R3 junos:21.2R1 junos:21.3R1 junos:21.4R3-S4

 

2. LAG flap issue (even the physical interface is stable)

================================

 

The customer is facing the LAG flap on the device running with a high CPU.

A similar issue of the LAG flap was reported and fixed in "PR 1630201 - NETCONF connections trigger LACP timeout."

External-Title

LACP timeout might be observed during high CPU utilization

Customer view: https://prsearch.juniper.net/problemreport/PR1630201

 

Release-Note On QFX5100 switches, when the CPU utilization (Routing Engine and FPC) is high and there are multiple Network Configuration Protocol (NETCONF) sessions running or SNMP polling is happening over multiple sessions simultaneously, the LACP session configured in fast mode might timeout.

 

We have the fix for these issues available in the below release:

 

Resolved-In

evo:22.2R1-EVO junos:18.4R2-S10 junos:20.2R3-S5 junos:20.3R3-S5 junos:20.3X75-D40 junos:20.4R3-S4 junos:21.1R3-S3 junos:21.2R3 junos:21.3R2 junos:21.3R3 junos:21.4R2 junos:22.1R1 junos:22.2R1

 

  • There is no workaround available for this issue.
  • Please check internally with the accounts team and you can upgrade the device to the latest qualified Junos version to avoid these issues.

 

Both PR:-1554340 and 1554340 are fixed in 18.4X29.3.

Modification History

2023-07-27: Sanitised[Removed Customer names]
2023-06-07: Initial publication