Description

When a link flaps repeatedly within a very short time frame on a multirate interface on MPC10E line card, on supported platforms (MX240,MX480,MX960,MX2010,MX2020), traffic stops egressing the affected interface and report syslog messages during link down event. When there are continuous flaps and if those flaps are very fast under 1 second and continuous then this issue will be seen.

 

You will observe in the logs:

"XQSS_CMERROR_DSTAT_INT_REG_DROP0_QDEPTH_UNDRN" or "mqss_wo_coreif_conn_credits_wait_for_init_value" during link down event

 

Symptoms

LACP (Link Aggregation Control) will not come up even after the interfaces recovered from flapping on all platforms which supports MPC10E cards. Trigger for this issue to occur is continuous and fast link flaps. It may occur over the hours or days or sometimes weeks. Once this issue is hit then PFE wedges and thus all the ports on that particular PFE stop forwarding the traffic. Since MPC10 has 4 ports per PFE thus 4 ports may get affected at a time. "Show lacp interfaces <interface name>" can be used to check status of LACP links.

 

Once this issue is hit then pfe wedges and thus all the ports on that particular pfe stop forwarding the traffic. Since MPC10 has 4 ports per pfe thus 4 ports may get affected.

 

Triggers:

This issue might be seen if the following conditions are met:

* On all Junos Platforms with MPC10E Line card(MX240,MX480,MX960,MX2010,MX2020)

* LACP is configured

* Continuous and Fast Link flaps of bundle interfaces with 100ms interval

Solution

 

Upgrade to a fixed release on:

https://prsearch.juniper.net/problemreport/PR1719682

This PR is specific to MPC10 platform

Modification History

2024-05-03 : Article Created

Related Information

https://prsearch.juniper.net/problemreport/PR1719682