Description

This article documents an interop scenario with Cisco where if GR is enabled for BGP on Cisco and later disabled, the time when BGP flaps between Cisco-Juniper may still lead the Juniper node to retain routes from the Cisco peer for whatever restart time was requested by Cisco initially.

Symptoms

If GR was enabled initially on Cisco and then disabled later, Juniper would still show the stale timer in 'show bgp neighbors'. This retention of routes on Juniper would lead to a traffic drop for the restart time which was initially requested by Cisco (while GR was enabled on it)

root@Juniper> show bgp neighbor 10.10.10.2
Peer: 10.10.10.2+19821 AS 65001 Local: 10.10.10.1+179 AS 65000
 Description: Cisco-peering
 Group: REMOTE_SITE          Routing-Instance: master
 Type: External   State: Established   Flags: <Sync RSync>
 Last State: EstabSync    Last Event: RecvKeepAlive
 Last Error: Hold Timer Expired Error
 Options: <Multihop Preference LocalAddress HoldTime AuthKey PeerAS LocalAS Refresh>
 Authentication key is configured
 Local Address: 10.161.80.3 Holdtime: 10 Preference: 170 Local AS: 65000 Local System AS: 0
 Number of flaps: 366
 Last flap event: Restart
 Error: 'Hold Timer Expired Error' Sent: 144 Recv: 173
 Error: 'Cease' Sent: 2 Recv: 0
 Peer ID: 10.10.10.2   Local ID: 10.10.10.1       Active Holdtime: 10
 Keepalive Interval: 3         Group index: 1   Peer index: 39
 BFD: disabled, down
 NLRI for restart configured on peer: inet-unicast
 NLRI advertised by peer: inet-unicast
 NLRI for this session: inet-unicast
 Peer supports Refresh capability (2)
 Stale routes from peer are kept for: 300
 Restart time requested by this peer: 120
 NLRI that peer supports restart for: inet-unicast

Solution

image.png
 

1. Consider an eBGP peering between Juniper and Cisco. Lets assume BGP GR configuration is default on Juniper. By default, Juniper only supports helper mode for BGP, without any additional GR configuration. Lets consider GR restart capability is configured on Cisco. So at this point, here is how the 'show BGP neighbor' output will look like on Juniper for the Cisco BGP peer:

root@Juniper> show bgp neighbor 10.10.10.1
Peer: 10.10.10.2+19821 AS 65001 Local: 10.10.10.1+179 AS 65000
 Description: Cisco-peering
 Group: REMOTE_SITE          Routing-Instance: master
 Type: External   State: Established   Flags: <Sync RSync>
 Last State: EstabSync    Last Event: RecvKeepAlive
 Last Error: Hold Timer Expired Error
 Options: <Multihop Preference LocalAddress HoldTime AuthKey PeerAS LocalAS Refresh>
 Authentication key is configured
 Local Address: 10.161.80.3 Holdtime: 10 Preference: 170 Local AS: 65000 Local System AS: 0
 Number of flaps: 366
 Last flap event: Restart
 Error: 'Hold Timer Expired Error' Sent: 144 Recv: 173
 Error: 'Cease' Sent: 2 Recv: 0
 Peer ID: 10.10.10.2   Local ID: 10.10.10.1       Active Holdtime: 10
 Keepalive Interval: 3         Group index: 1   Peer index: 39
 BFD: disabled, down
 NLRI for restart configured on peer: inet-unicast
 NLRI advertised by peer: inet-unicast
 NLRI for this session: inet-unicast
 Peer supports Refresh capability (2)
 Stale routes from peer are kept for: 300
 Restart time requested by this peer: 120
 NLRI that peer supports restart for: inet-unicast


2. Now, lets assume that the GR capability is completely disabled on Cisco. After this change, Cisco does not tear down the BGP session and as a result does re-initiate the BGP session with new capabilities which excludes GR. So, Juniper router does not have a way to understand if the GR capability has now been removed on Cisco. Juniper assumes the same capability as we can see in the 'show bgp neighbor' output on Juniper:

root@Juniper> show bgp neighbor 10.10.10.1
Peer: 10.10.10.2+19821 AS 65001 Local: 10.10.10.1+179 AS 65000
 Description: Cisco-peering
 Group: REMOTE_SITE          Routing-Instance: master
 Type: External   State: Established   Flags: <Sync RSync>
 Last State: EstabSync    Last Event: RecvKeepAlive
 Last Error: Hold Timer Expired Error
 Options: <Multihop Preference LocalAddress HoldTime AuthKey PeerAS LocalAS Refresh>
 Authentication key is configured
 Local Address: 10.161.80.3 Holdtime: 10 Preference: 170 Local AS: 65000 Local System AS: 0
 Number of flaps: 366
 Last flap event: Restart
 Error: 'Hold Timer Expired Error' Sent: 144 Recv: 173
 Error: 'Cease' Sent: 2 Recv: 0
 Peer ID: 10.10.10.2   Local ID: 10.10.10.1       Active Holdtime: 10
 Keepalive Interval: 3         Group index: 1   Peer index: 39
 BFD: disabled, down
 NLRI for restart configured on peer: inet-unicast
 NLRI advertised by peer: inet-unicast
 NLRI for this session: inet-unicast
 Peer supports Refresh capability (2)
 Stale routes from peer are kept for: 300
 Restart time requested by this peer: 120
 NLRI that peer supports restart for: inet-unicast


3. In this scenario, if the BGP between Cisco-Juniper flaps, Juniper would still retain these routes from Cisco for the restart time (120 seconds in this particular case) and would keep advertising these routes to upstream neighbors attracting traffic towards Cisco. This may lead to a traffic drop for at least 2 mins until an alternate BGP peer is preferred for these impacted prefixes

4. In the scenario where BGP-GR capabilities are deleted on Cisco, the user can manually tear down the BGP once so that the peering comes up with new capabilities

5. For a BGP peering between Juniper-to-Juniper, whenever there is a change in restart capability on any one end, the session is torn down automatically and re-established  so that the new capabilities can be negotiated, as reflected in below output:

root@Juniper> show bgp summary 
Threading mode: BGP I/O
Default eBGP mode: advertise - accept, receive - accept
Groups: 2 Peers: 2 Down peers: 1
Table          Tot Paths  Act Paths Suppressed    History Damp State    Pending
inet.0               
                       1          1          0          0          0          2
Peer                     AS      InPkt     OutPkt    OutQ   Flaps Last Up/Dwn State|#Active/Received/Accepted/Damped...
5.0.0.1               65001          5          2       0       2           7 Establ


root@Juniper>  show bgp neighbor    
Peer: 5.0.0.1+57135 AS 65001   Local: 1.0.0.1+179 AS 65001
  Group: IBGP                  Routing-Instance: master
  Forwarding routing-instance: master  
  Type: Internal    State: Established    Flags: <Sync InboundConvergencePending>
  Last State: OpenConfirm   Last Event: RecvKeepAlive
  Last Error: None
  Options: <Preference Damping AddressFamily Multipath Refresh>
  Options: <GracefulShutdownRcv>
  Address families configured: inet-unicast
  Holdtime: 90 Preference: 170
  Graceful Shutdown Receiver local-preference: 0
  Number of flaps: 2
  Last flap event: RecvNotify
  Error: 'Cease' Sent: 0 Recv: 6
  Peer ID: 5.0.0.1         Local ID: 1.0.0.1           Active Holdtime: 90
  Keepalive Interval: 30         Group index: 0    Peer index: 3    SNMP index: 0     
  I/O Session Thread: bgpio-0 State: Enabled
  BFD: disabled, down
  NLRI for restart configured on peer: inet-unicast
  NLRI advertised by peer: inet-unicast
  NLRI for this session: inet-unicast
  Peer supports Refresh capability (2)
  Stale routes from peer are kept for: 300
  Peer does not support Restarter functionality
  Peer does not support Receiver functionality
  Peer does not support LLGR Restarter or Receiver functionality


root@Juniper>  show log syslog | match 5.0.0.1 
Aug  5 09:31:10  Juniper rpd[8406]: bgp_handle_notify:4464: NOTIFICATION received from 5.0.0.1 (Internal AS 65001): code 6 (Cease) subcode 9 (Hard Reset) [code 6 (Cease) subcode 3 (Peer Unconfigured)] 
       
>> As seen in this log message, the peer torn down and re-initiated the BGP session due to BPG-GR capability change on it.

Modification History

2022-09-12: Initial publication.