ISSUE: Unable to reach servers under same subnet
ISSUE TRIGGER: Upgrade of switch on Sep 16th
ISSUE IMPACT: Application traffic impacted
ENVIRONMENT: Production
DEVICES:
Spines - Model: qfx5200 Code: 22.2R3-S1.9
Leafs - Model: qfx5200 Code: 22.2R3-S1.9
TOPOLOGY:
SRC: SERVER1 < --- > leaf1/2 < --- > spines 1/2/3/4 (l3 gateway) < --- > leaf1/2 < --- > DST: SERVER2
TROUBLESHOOTING SUMMARY:
1. Customer is running EVPN/VXLAN in CRB setup where are seeing connectivity issues between servers in the same VLAN. Layer 3 gateway for the servers are on spines 1 tp 4 irb.
Ping from source to destination fails when both the lag links to the server are enabled on leaf switches. When one of the eth link is disabled from server end, ping between source and destination works.
2. Based on the logs, looks like the issue started after upgrade of leaf1 on Sep 16th.
3. ARP, evpn, ethernet-switching and route tables were reflecting data as expected. No programming issue was seen.
4. Customer initiated flood between source and destination and we tried to correlate packet drops. No packet drop correlation seen for RX or TX. There were some correlation for this ING_NIV_RX_FRAMES_VL.xe83 queue but this was incrementing even for the working sw2 so we disregarded this as unrelated.
5. We applied firewall filters in leaf and spines. We didn't see the icmp reply counter incrementing.
6. Tried clearing ethernet-switching table and mac-ip-table for the affected server IP's in leaf and spines but issue persisted.
7. Upon further troubleshooting, we isolated the problem to the spines, specifically spine2, which was dropping packets with the reason:
"sw.mflt.dmac_dflt_drop - Mac Filter DMAC default drop".
Some of the servers are unreachable in the same subnet
The spines are currently running 22.2R3-S1 software version.
Several PRs addressing related issues have been resolved in versions between 22.2R3-S1 and the latest 22.2R3-S5. Customer will be upgrading the switches to 22.2R3-S(recent S release) and engage JTAC if the issue persists after upgrade