Juniper Apstra software product version 5.1.0 is available to licensed, registered Juniper customers from the Juniper Apstra software download site.Documentation for Juniper Apstra 5.1.0 is available from the Juniper Apstra documentation site.
N/A
Modular Chassis Profile for Arista DCS-7308X3 with DCS-7300X3-32C-LC line card.
With this new UI feature, now you can click on the tooltips to learn more about how to leverage and use the different features within Apstra.
Additional RBAC lock capabilities to prevent multiple users from overwriting or applying conflicting changes. If a lock is enabled, other users are warned or prevented from making changes depending on permissions.
A new GUI for creating and managing Apstra Edge instances to connect to Apstra Cloud Services (ACS) is now available. This GUI provides easy management and visualization of the containers.
You can now use SONiC based devices as access-switches including access-switch pairs with ESI.
You can now use SONiC based devices in Collapsed-Fabric blueprints
Prior to this feature, you would require full permissions to a blueprint in order to enable/disable interfaces. Extending upon the capability to enable/disable interfaces this feature adds the granular permissions under "Manage Racks and Links" and "Manage Generic Systems" when assigning blueprint specific permissions to users. Assigning users either of these permissions will let them enable or disable the interface as part of Fabric Front End(FFE) operations. This is particularly helpful when you have a remote data center team that you want to be able to bring up interfaces when installing new servers but restrict access to other parts of the blueprint they don't need to modify.
Prior to this feature, when upgrading the Network Operating System (NOS) of your switch, you had 1200 seconds(20 min) for Apstra to fetch the NOS image and download it to the switch. In some customer environments this process would exceed the 1200 seconds and timeout. With the timeout value being hardcoded, users had to either stage the image on a server in closer proximity to Apstra or undeploy the switch from the blueprint and perform a traditional NOS upgrade outside of Apstra. Now, with this feature the timeout value is configurable, but still 1200 seconds by default. In order to modify the timeout value navigate to Advanced Settings under Managed Devices, and you can choose a customized timeout value suitable for your environment.
When modifying a staging blueprint as a non-administrator user with admin permissions on the BP to edit and overwrite other users' staged changes, a banner warning message appears, informing the user that the blueprint has been locked and by whom, so that the user is warned before overwriting changes.
Anomalies History introduced as Tech-Preview in 5.0.0 is now GA in 5.1.0.
Optical Transceivers probe now has support for Cisco devices.
When adding Apstra Flow server details into the Apstra UI, a connectivity check is performed to verify the correct IP/address and username/password is provided to SSH into the Flow VM.
This release supports upgrade paths from previous Apstra 5.0.X releases.
Users must use VM-VM upgrades from Apstra 5.0.X releases. See the Apstra user guide for more information on Apstra upgrades.
Users can set up SAML Single Sign-On (SSO) to authenticate into Apstra. This integration enhances user experience by allowing seamless authentication through corporate identity providers (IdPs) using the SAML 2.0 protocol.
A new option is available under Platform -> Technical Support to allow users to input their SSRN to help expedite support assistance with JTAC.
The API endpoints for anomalies now include a new 'anomalous_node_id' field to provide a mapping between built-in anomalies and graph node IDs. This is applicable for the following API Endpoints:
GET /api/anomalies -> Get all anomalies.
GET /api/blueprints/{blueprint_id}/anomalies -> Get all anomalies for a given blueprint.
GET /api/systems/{system_id}/anomalies -> Get all anomalies for device.
The mapping is as follow:
BGP anomaly -> Mapped to interface node.
Cabling anomaly -> Mapped to link node.
Interface anomaly -> Mapped to interface node.
Hostname anomaly -> Mapped to system node.
LAG anomaly -> Mapped to interface node, with if_type ='port_channel'.
MLAG anomaly -> Mapped to domain node (with domain_type = 'mlag') + interface node (with if_type = 'port_channel') for each to_generic dual-attached link.
Liveness anomaly -> Mapped to system node.
Route anomaly -> Mapped to system node.
Config anomaly -> Mapped to system node.
Deployment anomaly -> Mapped to system node.
As an Apstra admin you have the option to enforce a customized login banner which your Apstra users must accept prior to being able to login to Apstra. To enable, modify, or update the login banner a user will need the write permission "Login Banner" under Platform. This is particularly helpful for organizations that require users to accept a notice and consent policy prior to logging in to their systems. This lets you enforce USG login banner requirements.
You can now deploy the Edgecore/Accton AS9736-64D switch in Apstra when running SONiC.
You can now deploy the Edgecore/Accton AS9726-32DB switch in Apstra when running SONiC.
The in-product API documentation now has a change log to display all API changes for any new release. This allows you to keep track of all changes in a simple to simplify maintaining and updating any application code using Apstra's APIs. API Endpoint life-cycle will have four states:
'New' for new features.
'Changed' for changes in existing functionality.
'Deprecated' for soon-to-be removed features.
'Removed' for now removed features.Full implementation of 'Deprecated' and 'Removed' states will be done in 5.2.0.
Inter-port constraints have been added to the following Cisco devices
Cisco_N9K_C93600CD_GX
Inter-port constraints have been added to the following Arista devicesFixed Form factor:
Arista_DCS-7050QX-32S
Arista_DCS-7050SX3-48YC12
Arista_DCS-7280CR3-32P4
Arista_DCS-7280CR3K-32D4
Arista_DCS-7280CR3MK-32P4S
Arista_DCS-7280QRA-C36SModular:
Arista_DCS-7500R3_36CQ
When you try to modify your Logical Devices or Device Profiles to add or delete ports, you could inadvertently delete all the current LD Port Groups or DP Ports configurations. With this change an informational message is presented to inform users that resizing, adding, or removing panels will reset Display IDs to their original value. Recommended practice is to update the panel layout and selection should be first to prevent data to be overridden.
The following updates have been made for switch operating systems qualified for the Apstra 5.1.0 release.
Juniper Networks:Junos (All roles)21.4R322.2R322.4R323.4R2-S4
Junos Evolved for IP-Forwarder role (Spines in EVPN or any role in an IP-Fabric): 22.2R3-EVO22.4R3-EVO23.4R2-S3-EVO
Junos Evolved for EVPN leaf roles: 22.2R3-EVO22.4R3-EVO23.4R2-S3-EVO
Junos Interconnect Gateway Leaf: 22.4R3 (minimum) 23.4R2-S4
Junos Evolved Interconnect Gateway Leaf: 22.4R3-EVO (minimum) 23.4R2-S3-EVO
Cisco Systems:9.3(13)10.2(6)10.3(4a)
Arista Networks:4.24.5M 4.28.7.1M 4.30.3M
Dell EMC & Edgecore:Enterprise SONiC 4.1.2Enterprise SONiC Edge Standard 4.1.2Enterprise SONiC 4.2.1 Enterprise SONiC Edge Standard 4.2.1
With new sorting and filtering capabilities, you can easily find the exact rack type you're looking for. You can sort on any column and filter on all rack attributes.
This UI revamp focuses on enhancing the day 2 experience for our users by surfacing useful information already available in the product. It is all about making that useful data easier to see and turn into action. The changes are based on feedback gathered from customers and field teams.
You now have additional Graph queries in the predefined catalog focussing on Virtual Networks and their footprint.
For the Interface Flapping (Specific Interfaces) probe, interface selection can now be done using System or Interface Tags instead of individually selecting specific systems and their interfaces. This makes it easier to bulk define the scope of devices and interfaces to monitor, and the probe can automatically react to changes in Tags without the need for manual updates. Additional context has been added, such as "Remote Interface Name" and "Remote System Label" Historical retention has now been extended to 30 days.
QSFP-DD800 optics in QFX5240-64OD is now qualified and supported by the Optical transceivers IBA probe.
The Telegraf input plugin for Apstra has been refactored to now use the external plugin model. In this model, the plugin is compiled separately and provided as an executable binary file to run inside the Telegraf's execd input plugin.
You now have drop-down menus in the Custom Collectors user workflow anytime you need to target a system, that is in "Execute command", "Test Query" or "Validate schema". That menu lets you choose a device either by its hostname, Hardware Model or IP Address, making it easier to pinpoint a system without having to remember its IP Address. A search field is also available.
You now have a a visual representation of the IBA pipeline that optimizes node placement and minimize edge crossings for better readability of the probe's pipeline.
You can now select multiple custom collectors to request batch delete instead of having to delete them one by one. You also can perform this batch deletion at the service level, in which case you will be notified with the underlying collector(s) deletion which will result from the service(s) deletion.
You now have support for auto-completion of the expressions when you author a custom collector helping you to write a complex expression on the fly by automatically pulling accessor names (with their field name, xPath and type) and functions (with their arguments and descriptions).
BGP Monitoring probe, which looks after BGP Flaps is now available for all NOSes including Cisco NXOS and SONiC,
Support for the ACX7100 has graduated from tech preview to general availability. You can now use the ACX7100 with Junos 23.4R2 as DCI gateway for integrated DCI use cases. This applies to both the ACX7100-32C and ACX7100-48L variants.
Note: When using ACX as a DCI gateway, the L3 VNI for DCI must match between all fabrics.
The flow collector now parses additional fields in the RDMA header for RoCE v2 traffic, which are displayed in the various dashboards.
Additional health checks in the metrics endpoint for VM CPU and memory, collector version, Apstra connection status, and SNMP configuration status.
Additional help text and a README to help guide users when using the cluster setup script to understand the different cluster roles better.
Tech Previews give you the ability to test functionality and provide feedback during the development process of innovations that are not final production features. The goal of a Tech Preview is for the feature to gain wider exposure and potential full support in a future release. Customers are encouraged to provide feedback and functionality suggestions for a Technology Preview feature before it becomes fully supported.
Tech Previews may not be functionally complete, may have functional alterations in future releases, or may get dropped under changing markets or unexpected conditions, at Juniper’s sole discretion. Juniper recommends that you use Tech Preview features in non-production environments only.
Juniper considers feedback to add and improve future iterations of the general availability of the innovations. Your feedback does not assert any intellectual property claim, and Juniper may implement your feedback without violating your or any other party's rights.
These features are "as is" and voluntary use. Support Services will attempt to resolve any issues that customers experience when using these features and create bug reports on behalf of support cases. However, Juniper may not provide comprehensive support services to Tech Preview features. Certain features may have reduced or modified security, accessibility, availability, and reliability standards relative to General Availability software. Tech Preview is not supported under existing service agreements, SLAs, or support service.
For additional details, please contact Juniper Support or your local account team.
Optical Transceivers probe now has Tech-Preview support for SONiC devices.
As an Apstra admin you now have the option to enforce a customized login banner which your Apstra users must accept prior to being able to login to Apstra. This is particularly helpful for organizations that require users to accept a notice and consent policy prior to logging in to their systems. This lets you enforce USG login banner requirements.
Users may encounter duplication of Interface Maps (IM) when adding a new access switch to the leaf in a collapsed fabric blueprint, specifically with Juniper EX series devices. This issue is due to a device profile configuration mismatch: the backend incorrectly generates an additional IM with a duplicate label because of inconsistencies in connector type information (RJ45 vs. rj45). Users can view these duplications in the Interface Maps section by navigating to Staged -> Catalog -> Interface Maps.
AOS show tech collection failed due to the increased size of the folder /var/tmp/show_tech_tmp and the AOS is transitioning into read-only operation mode.
/var/tmp/show_tech_tmp
Show tech functionality has been enhanced by implementing the following improvements 1) cleanup of temporary directories is Implemented upon early exit to free up disk space 2) Tel files are Compressed to optimize the disk size.
Apstra ZTP versions 5.0.x or below may fail to start the dhcp container service if a custom dhcp config is used where multiline /* */ or # bash comments are used. It may also fail to start where config file includes are used <? include "filehere" ?>. The dhcp init.sh container startup script fails to parse out comments and file includes so that socket_name path can be determined and created at startup.
/* */
#
<? include "filehere" ?>
A fix to the init script triggers failure with an error that suggests the user's invalid updates (not supported comment formats such as multiline). User must use ZTP UI to change DHCP configuration instead of manual changes. Enhancement to cover diverse comments formats will be addressed in the future release by RFE-3447.
RFE-3447
The Juniper EX4400 platform does not have a dedicated Locator/ID LED. The Junos "show chassis beacon" command will always return ON. All ports or connected ports will be lit in "GREEN" depending on the explicit beacon command. Also, it uses a 5-minute default timer, and CLI supports between 1 and 120 minutes. After a predefined time, the beacon status changes back to the default state in CLI. The switch port status is not changing based on Junos "request chassis beacon" command.
"show chassis beacon"
"GREEN"
"request chassis beacon"
In the case of a VLAN with DHCPv6 relay(s) configured in Arista EOS devices, changing the VRF may result in stale DHCPv6 relay commands remaining in the VLAN configuration. This is the result of an incorrect negation performed while the VRF was being changed.
Apstra 5.1.0 will be correcting this problem and will also be including an upgrade plugin to clean up VLANs of stale DHCPv6 relay entries. The plugin will act upon virtual networks that are IPv6-enabled at the time of upgrade to Apstra 5.1.0.
In the extremely narrow scenario where stale DHCPv6 relay entries have been left to a virtual network that used to be IPv6-enabled but no longer is, users are advised to either negate the stale entries manually or apply a full config (which is traffic affecting) to the device.
After a CT (connectivity template) with dynamic BGP peering and BGP Prefix Dynamic Neighbor information is assigned to the SVI interface for a system, if the system is removed from the virtual network later, the CT becomes unassigned status, which allows the user to delete the CT. After the CT is removed later, protocol_session becomes orphaned from the associated CT. it can lead to failure in deleting the routing zone.
Application point type has been changed for Dynamic BGP prefix peering ("Dynamic BGP Peering" primitive type with any of "IPv4 Subnet for BGP Prefix Dynamic Neighbors" or "IPv6 Subnet for BGP Prefix Dynamic Neighbors" fields populated). Previously the primitive was applied to SVI or Subinterface, after the fix it is applied to Loopback in the corresponding Routing Zone. This change does not affect created BGP session as it was attached to Loopback even before the fix.
However, the fix influenced the CT structure. As Subinterface is not a valid application point for Dynamic BGP prefix peering any more, all user-defined CTs, which have Dynamic BGP prefix peering on top of Logical Link will be automatically modified during the upgrade to Apstra 5.1.0. For each "Logical Link" primitive all its child "Dynamic BGP" primitives with prefix peering configured will be split to a separate CT (together with all child primitives). "Dynamic BGP" primitives without prefix peering will stay as is.
The change affects only CT structure and its application points. It does not affect device configuration created from the CT.
The subsequent steps to gather pristine configuration based on the newly upgraded NOS and push full service configuration would fail, causing a service impact, even though the device NOS is upgraded during the NOS upgrade operation.
Apstra System Agent added some interlocks in the NOS upgrade job for Cisco Device to wait until it is ready post-booting up with new OS image
The JUNOS/EVO device uses GRPC for the MAC Telemetry service. During the GRPC processing, Apstra Controller uses device's credential information (username and password) to populate GRPC meta data. If the password includes non-printable ASCII characters, a validation error for invalid characters can lead DeviceTelemetryAgent to fail with a crash.
Apstra 5.1.0 includes a fix to allow non-printable ASCII characters in the password so that DeviceTelemetryAgent doesn't crash continuously. However, device NOS also needs to support non-printable ASCII characters in the password.
Apstra introduced custom telemetry services in the 4.2.0 release. Users define a service schema to structure and store data, based on key and value from the CLI output. The UI doesn't allow hyphens in telemetry key and value names. However, the API allows them. If a telemetry service registry entry with a hyphen is created via the API, the upgrade to Apstra 5.x may fail with validation error.
File "/usr/local/lib/python3.10/dist-packages/lollipop/errors.py", line 182, in raise_errors raise ValidationError(self.errors) lollipop.errors.ValidationError: Invalid data: {'key': 'Unable to identify "key" from schema'}
In Apstra 5.1.0, telemetry key/value names with hyphens are now disallowed.
Apstra 5.1.0 Upgrade validates key and value formats before importing data from the old controller to the upgraded controller.
Assigning a role with a space in the role name to a user and then removing it, while keeping the other role still assigned to the user, causes the "Event Log" page in the UI to display as a blank page.
The Apstra support team will provide the hotpatch and implementation process to be performed on the AOS server
If the gRPC probe used to check the device's gRPC health status keeps failing, Apstra gRPC Telemetry services remain in a failed state on the JUNOS device running 23.4R2-S3. When the device is restarted, the issue becomes apparent. Even if a gRPC probe request is received, the device doesn't respond with the data. Because Apstra gRPC probe doesn't have a timeout mechanism for the gRPC probe request, Apstra gRPC continues to use the same TCP connection, which may have issues.
Apstra 5.1.0 enhanced the gRPC probe with a timeout feature that terminates the old TCP connection and establishes a new TCP connection when the timeout occurs.
Customers may observe servers on a collapsed fabric failing to PXEboot where interface is rendered with a large hold time for up event as part of the collapsed fabric reference design
Customer still needs to use configlet with preferred customized value until future enhancement (interface policy) will be implemented.
Apstra 5.0.0 has added to some Juniper Device Profiles and Linecard Profiles a mechanism (using the validations field in the interface setting) to describe port group constraints and alert the customer of possible port breakout combinations that are prohibited. Because the schema validation of these constraints is not correctly enforced, a customer may be able to create a custom device profile with incorrect or incomplete port group constraints, which could result in an aborted rendering of the Junos device configuration.
Next Apstra release 5.1.0 will be enforcing proper schemata for the port group constraints in Junos devices and other NOS families gaining interport constraint support (NXOS, EOS).
Despite the configlet being applied to some nodes based on tags, there were some link changes in the logical diff section. The logical diff tab continued to display the changes even after the configlet was reverted, and there was nothing to commit in the uncommitted tab section.
Even if it is shown in the Device> Managed Devices> Telemetry> Anomalies tab, the MAC anomaly was not showing in the active > Anomalies tab. It was identified that UI request to backend doesn't include anomaly type for MAC.
MAC querying might have crashed because it added all the systems to the index, although some systems might lack system_id if their deploy_mode is not deployed. Therefore, it led to KeyError as None (system_id) is not a valid value for the "Tac::String" field
Fix makes MAC querying take into account system's deploy mode for security zones by analogy with the existing handling of virtual networks
When draining one MLAG leaf in Maintenance Mode, MLAG telemetry expectations are not correct. No anomalies should be seen even though AOS will shutdown the server facing links and port-channels. However, anomalies are seen because of incorrect expectations.
After installing a new image during the NOS upgrade process, a reboot command with the newly installed image is executed on the device running Arista EOS 4.30.3M. The device running 4.30.3M remains online for a significantly longer period of time without rebooting than the other release. As a result, the system agent in charge of the NOS upgrade job makes the assumption that the device has successfully rebooted with a new image. The NOS upgrade fails version comparison validation since the device still runs the old version without rebooting.
Any custom probe that utilizes the Extensible Service Collector processor must now have explicitly string types for its property key value types due to additional validation constraints. If this is not done, the UI will display a validation error pointing to the offending property key.
When Apstra streaming reports interface counter metric data, the streamed data is already in per second format (for example, tx_bytes is TX bytes per second); however, packets per second and bit per second data are calculated incorrectly by additionally dividing by delta_seconds.
Correct calculation logic to stream the right calculated data
Every time a config apply happens in a SONiC device managed by Apstra, the FRR daemon configuration is gracefully reloaded by the frr-reload.py script inside the bgp container. The output of that script is directed to the file /var/log/frr/frr-reload.log inside the same container. The size of that log file is not expected to ever become a concern, unless a customer performs many thousands of config apply operations with a rather large FRR configuration.
The command `docker exec bgp logrotate --verbose /etc/logrotate.d/frr` can be placed in a cronjob to activate the logrotate job inside `/etc/logrotate.d/frr` in regular intervals.
The expected routes for the loopback address of the leaf node in the spine node are calculated by using pod_label to determine whether the target leaf node and the current spine node are in the same pod. When the pod label is updated by UI, the ExpectationRenderer Agent may not update the new pod label information into all spine and leaf nodes, resulting in including the wrong nexthops into superspine nodes for leaf loopback address, even if the leaf node is directly connected from spine node.
A part_of_pod attribute of a device instance is used to determine pod membership, and it affects route expectation. Now the value is changed from pod label to pod ID to make sure that pod label changes would not affect pod membership changes.
In some cases, the frontend UI generates incorrect queries (without specifying per-metric aggregation) to retrieve time series data for stages. When this happens, the backend selects the default aggregation method based on the value type, which means that instead of no aggregation, average aggregation is used by default. As a result, the output does not match the aggregation method used.
In some cases (e.g. for EVPN Host Flapping per System stage of EVPN Host Flapping probe with enabled context data) the chosen aggregation method is not applied
In some cases where a VRF has been created, used, and then deleted, the FRR bgpd daemon may still indicate the existence of that VRF. Example, the Vrf-PURPLE in this vtysh output:
leaf2# show vrf vrf Vrf-PURPLE inactive vrf Vrf-blue id 120 table 1001 (configured) vrf Vrf-red id 122 table 1002 (configured) vrf mgmt id 47 table 5000
The existence of Vrf-PURPLE confuses the Apstra BGP route collector, causing it to crash. In such a case, the BGP route telemetry will stop working.
The Apstra BGP collector will be made to ignore such stale VRFs in Apstra 5.1.0 and later versions.
This is a rare case in which a device reboots during a gRPC session, resulting in stale polling timers on the Apstra Agent side. When gRPC restarts, the stale timers are replaced by new timers, which trigger the handling timer for collection, resulting in an agent crash. After the agent restarts from the crash, the system functions normally without any further crashes.
Floating-point precision discrepancies can cause problems in the integration between IBA and metricdb when configuring IBA probes. To be more precise, the live data that was obtained from IBA is queried using metricdb using a trie-based matcher. However, minor variations in floating-point values (such as 0.1 being read as 0.10000000149) could cause metricdb to fail to match the desired keys. This can cause probes to miss crucial data when querying specific values.
Updated data type handling to improve accuracy and ensure proper compatibility with MetricDb. This resolution will be included in the upcoming 5.1.0 release.
Apstra enables wide mode in any Trident 4 chipset for any SONiC devices (Z9432F-ON, S5448F-ON) that support ESI, as long as the CLOS topology uses ESI. If the topology is collapsed rather than CLOS, the Apstra reference design contains an error that prevents wide mode from being automatically enabled as needed. This can occasionally cause traffic drops for packets entering the fabric.
Apstra 5.1.0 will improve the logic to enable Trident 4 wide mode automatically, including for ESI collapsed fabric topologies. Customers using an ESI collapsed fabric with Trident 4 SONiC devices must still apply a full configuration after upgrading. Such customers should schedule a downtime window and run the full configuration apply on all devices that meet those criteria.
When displaying racks, Apstra generates statistical information for the rack by sorting through each leaf's position data in ESI or MLAG cases. When a rack is updated by inserting a generic system into one of the leaf nodes that make up ESI/MLAG, Apstra uses sorting criteria based on the label from the leaf nodes. This inconsistent sorting criteria leads to calculation statistics referring to non-existent keys, resulting in errors. This issue only arises when a generic system is connected to one leaf node of an ESI/MLAG pair via a single attachment for a specific speed and the other leaf node lacks an interface for the same speed.
When the rack is updated, change the sorting criteria from the leaf's label to the leaf's position data.
When 10 Mbps speed is selected via Update Link Speed, the user interface (UI) disables the update button so that it cannot be applied, even though the interface supports 10 Mbps.
UI allows 10 Mbps in the Update Link Speed to be updated when interface support 10 Mbps
The upgrade to 5.0.0 validates the device's pristine configuration. When the pristine configuration for the JUNOS device contains system login message or announcement with # characters, upgrade validation fails with an error, resulting in the upgrade failing. When performing a NOS upgrade or device agent upgrade in 5.0.0, the same validation error may occur.
When a MAC entry is learned via Virtual Network, the VNI column in the MAC Address Table of MAC Monitor Probe includes a hyperlink to show details for the associated virtual network. In the case of a VLAN type Virtual Network, the VNI value is displayed as 0 rather than NA (Not Available), and this is used to generate the incorrect filter for the Active->Virtual->Virtual Network Tab, which matches Virtual Network with VNI value with 0. As a result, the tab does not display any associated virtual networks.
When a rack is built with leaf devices and generic systems, group labels for generic systems can have the same value as leaf or access switches' target_switch_label, contrary to the expectation that the group label should not be the same value as target_switch_label inside the rack. Any changes to the rack, such as adding a generic system or deleting an existing generic system, would fail due to the validation error caused by not meeting the above expectation.
During the EOS upgrade from 4.28.7.1M to 4.30.3M in the Arista device MLAG pair, the device that would be upgraded first may exhibit incorrect ARP behaviour and miss receiving ARP entries from its MLAG peer. Stochastic traffic failures may also be observed between the two MLAG members.
It should be noted that using an MLAG pair of EOS devices from different versions is not supported and will result in traffic loss. For more information, please refer to the vendor release notes. It is recommended to update both switches in an MLAG pair in quick succession.
The device telemetry health probe for the Junos and EVO devices running >= 23. 4R3-S3 or >=22. 3R2-S2 release shows abnormalities associated with the "gRPC connection reset" The Apstra telemetry collector is no longer compatible with gRPC due to a fix for XPATH in the specific NOS versions, which causes anomalies in the telemetry health probe.
>= 23. 4R3-S3
>=22. 3R2-S2
When operating Juniper ACX devices, you might encounter a situation where layer-3 packets are inaccurately identified as layer-2 packets. This can result in incomplete packets being exported to the flow collector.
Junos QFX5K and 10K devices running 23.4R2-S1/S2/S3 don't support the XPath /interfaces/interface/subinterfaces/subinterface/state/admin-status which makes Interface telemetry services fail
/interfaces/interface/subinterfaces/subinterface/state/admin-status
Junos QFX 5K and 10K devices running 23.4R2-S1/S2/S3 will be automatically switched into polling mechanism for interface telemetry service
On Juniper QFX52xx platforms running the Junos-EVO image, the L3 sub-interface configurations applied via connectivity template (CT) are not rendered correctly over the aggregate Ethernet and regular physical interfaces. This results in deployment failures in affected blueprints.
This issue is tied to hardware constraints on Broadcom TH3, TH4, and TH5 ASIC platforms. The flexible-vlan-tagging feature required to support L3 sub-interfaces is not officially supported on the QFX52xx series (QFX5220, QFX5230, and QFX5240). While this command may have appeared in earlier Junos-EVO versions such as 22.4R2, its presence was unintentional and not validated for use. In Junos-EVO 23.4R2, the option is intentionally hidden, and any attempt to configure it will result in a commit error, reflecting Junipers enforcement of this limitation.
Apstra enforces this hardware limitation through validation preventing user getting into deployment error, explicitly blocking sub-interface configurations for affected platforms. Apstra returns the following error when validation fails:
"System {system} OS Family {os_family} with ASIC {asic} does not support subinterfaces due to hardware constraints. This error can be relaxed within validation policy in the event of a vendor-supplied OS update adding support."
will introduce changes in a future release to enable sub-interface configurations on Juniper TH3, TH4, and TH5 ASIC-based platforms (QFX5220, QFX5230, QFX5240). This support will not rely on the flexible-vlan-tagging feature, which is not officially supported on these platforms in any Junos-EVO release. Instead, Apstra uses vlan-tagging as per Junos Evo platform recommendation.
PTX on EVO version 23.4R2 will drop the packets towards the vtep addresses x.x.x.0/32 (i.e 10.0.0.0), which has "0" in the last octet.
When both untagged and tagged VLANs are defined over the same port for L2 generic systems, traffic drop over tagged VLANs is observed on QFX 5230/5240 Junos-Evo platforms.
When TLS for Apstra API (EF_JUNIPER_APSTRA_API_TLS_ENABLE: "true" in the /etc/juniper/flowcoll.yml) is enabled, the flowcoll process might fail at startup time.
Flow image 7.5.3.1 fixed the issue. Upgrade flow image into 7.5.3.1
Attempting to pull the configlet preview for specific blueprint device ("On Device Configlet Preview"), by clicking on the device label under the general configlet preview page might fail with a slightly misleading error, if the device is unassigned. Certainly trying to get a preview for a device which is not assigned is bound to cause an error, as a real preview for a device that doesn't exist isn't possible. However, the error emitted is slightly confusing.
MetricQueryManagerAgent handles large historical data to serve the '/blueprints//anomalies-history' endpoint. Depending on the amount of data, the agent's memory footprint may increase significantly. The benchmark environment recorded a memory footprint of up to 2.2Gb for the agent. The memory footprint settles after the initial bump.If the system administrator is concerned about the MetricQueryManagerAgent footprint's impact on the system's available memory, the following workaround is recommended.
Restart MetricQueryManagerAgent and avoid using the 'Time Series' query.
When AnomalyGenerator tries writing to MetricDB and taking SysDB snapshots if vlanid is missing from the primary key it will cause the AnomalyGernator to crash wih the below trace information
Unique Index (PrimaryKeyIndex) violation newRow index matches that at row: 65536 python3.10: /Project/leblon/infra/TableTop.tin:206: void Aos::DoubleLinkListHelper::addToList(Aos::RowIndexHelper&, U32): Assertion `false && "!row"' failed. Process 22875 died with signal 6 (SIGABRT) errno 0 code -6 (unknown)
Add the following to disable all anomaly logging into the /etc/aos/aos.conf file and then restart the aos service.
[anomaly_metric_logging]enable = 0
The scenario change-device-password CLI command securely updates device credentials by performing tasks like SSH checks, configlet staging, blueprint commits, and agent password updates. In the current Apstra CLI version, the system agent check has a 60 second timeout, while configlet staging is limited to just 20 seconds. If these operations take longer than expected, the command may fail with errors like:
Failure 1: Task Stage creation of Configlet for password change may fail with: AssertionError: Timeout waiting for Wait that last task status is succeeded Failure 2: Task Check System agent status may fail with: 409 Conflict: Agent is already running a job (check)
These failures occur when backend tasks exceed the current timeout settings, which is particularly noticeable in Apstra 4.2.x and later versions, where performance issues with configlet and configuration rendering are known. Additionally, longer durations in the check job can result from changes in the customer’s environment.
Manual intervention is required to cleanup and proceed as follows:
1. For System Agent Check Failure (409 Conflict): * Update the pristine configuration to reflect the new encrypted password for the user. * Remove the temporary configlets named `change_pass_<...>_junos`. * Manually commit the blueprint. * Re-run agent check jobs to confirm the password update. 2. For Background Task Timeout (e.g., Configlet Staging Failure): * Revert the blueprint to the previous working version.
After completing the steps above, use the latest Apstra CLI image for the Apstra release (4.2.2, 5.0.X, 5.1.0, 6.0.0) and retry the scenario change-device-password command. The most recent Apstra CLI image can be obtained by contacting Apstra Support.
When switching between routing engines in a Junos Evolved System (such as the PTX10008) with dual routing engines, the Apstra Onbox agent may report an incorrect management IP, management interface, or management MAC address.
This bug does not affect Offbox agent.
It is recommended that the Apstra Onbox agent be restarted after a switchover using the command "request system application app aos restart node reX", where reX is the Master routing engine.
Apstra customers 4.2.x and greater may encounter an issue where the MTU value for IP Links to Generic Systems cannot be set to 9216 via the UI. Although the UI states that only even values in the range 1280-9216 are accepted, the input of 9216 is incorrectly rejected, while 9214 is accepted.
Important Note: 1. This issue is applicable only to customers upgrading from Apstra 4.1.x to 4.2.x or later, where Fabric MTU remains disabled post-upgrade and customers who wish to continue without enabling the Granular MTU feature. 2. This issue is not applicable to customers with fresh 4.2.x deployments, where Fabric MTU is enabled by default, activating the Granular MTU feature.
Although the UI blocks 9216, the backend API does accept this value. As a workaround, users can update the MTU value via the REST API:
1. Navigate to Platform → Developers → REST API Explorer. 2. Use the PATCH /api/blueprints/{blueprint_id}/fabric-settings endpoint with the following payload: { "external_router_mtu": 9216 } 3. Verify the update by performing a GET on the same endpoint: GET /api/blueprints/{blueprint_id}/fabric-settings 4. Navigate to Blueprints → Blueprint Name → Uncommited and check the diff 5. Commit the Blueprint
If further assistance is needed, please contact Apstra Support.
When setting the reservation mode from the configurator using the different options except None and the reservation mode flag is set in the DHCP configuration, it causes the static IP address to be assigned from the pool rather than the static IP address mapped against the MAC address in the configurator.
User can follow the below steps to get the static IP to the host using ZTP1. If Reservation Mode is set to None, users can configure static IPs by navigating from ZTP UI to DHCPv4 Configurator, setting Reservation Mode to None, and toggling Reservations-Global.2. If Reservation Mode is enabled, users must move the hosts inside Reservations to the subnet level based on the required IP address while keeping the Reservation Mode set to Global.
When monitoring Apstra ZTP device status in the Apstra UI under "ZTP Status" / "Devices", there may be duplicate entries for Junos devices. Apstra ZTP will try to ensure the physical management interface for the Junos device is used instead of any virtual management interface (e.g. "vme" interface). Junos may use the virtual interface when ZTP starts but cannot be added to the required "mgmt_junos" routing-instance. This is done as the first step in ZTP in order to ensure that the management IP address does not change during the rest of the steps involved in ZTP (especially those involving connectivity to Apstra). Enabling a different management interface will cause the DHCP server to give out a new lease. Also, the vendor class identifier for the new management interface is cleared so that the DHCP server does not give out vendor-specific options to this interface, which may re-trigger a new ZTP session while the current session is active. This is expecetd behavior.
The timestamp field of AosMessage in the streaming has used uint64 format for both millisecond and microsecond timestamp information. The current timestamp uin64 field will be changed in release 6.1.0 to the google.protobuf.Timestamp format, which is incompatible with the previous version, in order to provide consistent, accurate timestamp information. The streaming receiver side must be modified to accommodate this incompatible modification.
An event involving BGP flaps from BGP peers configured on IRB/Loopback interfaces for Juniper EVO device has been reported during the commit with incremental changes (deleting an IRB/Loopback interface in the same VRF). The BGP flaps happen when the device's configuration is committed in override mode, which was Apstra's default setting prior to 6.1.0. However, using load update mode did not reveal the issue.
If the EVO device is running as an off-box agent, please add key load_mode with value update into open options in the edit agent menu. Otherwise, recommend upgrading Apstra to 6.1.X, which uses load update as the default mode for commit.
When Dashboard for blueprint is displayed via clicking Blueprint in the UI, Apstra UI failed to load the blueprint dashboard because the preference information of the dashboard is configured with an empty string value, not the right value for preference.
The problem can be fixed via issuing a REST API POST call for /api/aaa/users/(target_user_id]/preferences with the below payload information. After the POST call is successful, the affected user must log out and back into the UI for it to be effective.
{ "key":"dashboard", "value": {} }
or user need to download the UI hotpatch ( https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HDePeIAL ) and apply to the controller VM as follows
admin@aos-server:~$ sudo su [sudo] password for admin: root@aos-server:/home/admin# gzip -d aos-web-ui-5.1.0-esc539.run.gz root@aos-server:/home/admin# chmod 755 aos-web-ui-5.1.0-esc539.run root@aos-server:/home/admin# sudo ./aos-web-ui-5.1.0-esc539.run Verifying archive integrity... All good. Uncompressing AOS WebUI installer 100% ### Backing up existing AOS WebUI into /opt/aos/frontend/snapshot/2025-03-26_23-50-19 ... ### Copying AOS WebUI file into aos_controller_1 ... ### Initializing new AOS WebUI ... ### Done!
When multiple show-tech jobs are triggered simultaneously, the /var/log partition can reach high utilization, causing the controller to enter read-only mode. This may result in incomplete job status updates, leaving several jobs stuck in in-progress or pending states. In a corner case, this condition can also lead to repeated crashes of SystemAgentManager, preventing automatic recovery even after disk space is reclaimed. Avoid triggering bulk show-tech collection on a large number of devices. Ensure sufficient disk space is available before running show-tech.
If the issue occurs, reclaim space in /var/log and restart AOS services. In most cases, SystemAgentManager recovers automatically, clearing stuck jobs and completing or failing pending ones. If SystemAgentManager crash persist and jobs remain stuck, manual cleanup via Acons is required. It is recommended to contact Apstra Support for assistance.
When attempting to change the keepalive IP address of an MCLAG pair of EOS devices, the configuration deployment of the devices belonging to the pair may fail.
Changing the keepalive address would likely result in traffic disruption even if the deployment were eventually successful. It is recommended to execute a full config apply to both devices after the configuration deployment of them fails.
If the rendered device configuration for the VLAN description contains newline characters populated from the virtual network's description field, the commit operation fails because JUNOS and JUNOS-EVO devices do not support multi-line string for the VLAN description.
Please use one-line formatted string in the description field of Virtual Network rather than multi-line string.
When making use of the built-in jinja function on a custom configlet, {{ function.merge_vlans_to_list(interface_model["allowed_vlans"]) }} fails when used within a configlet. An error will be seen in rendered configuration previews "TypeError: unsupported operand type(s) for -: 'int' and 'str'"
In the configlet, map the string values to integers, such as {{ function.merge_vlans_to_list(interface_model["allowed_vlans"] | map("int") }}
While configlet is being applied to the SONiC device, if the device agent restarts after being disconnected from the controller, the agent executes any remaining changes and collects the running configuration as golden configuration to monitor for configuration anomalies. Because the process of applying configlet changes is still running independently of the agent, it introduces changes into the running configuration even when the golden configuration is collected by the agent. The following changes from the process cause configuration anomalies in the SONiC device.
After reviewing the running configuration on the SONiC device, if all the changes from the configlet are correctly applied, the customer can safely accept changes to avoid further configuration anomalies.
Adding multiple AAA servers in the blueprint through Staged > Catalog > AAA Servers leads to a configuration load error in the JUNOS and EVO device during commit check or commit.
Recommend using configlet instead of using UI (Stage > Catalog > AAA Server) when multiple AAA servers needs to configured.
When system resources are scarce, Docker (dockerd) may fail to respond within the expected timeframe (60 seconds), resulting in a connection timeout error. This keeps the ContainerLauncherAgent from successfully launching offbox containers. The logs show that high CPU load and late clock events contributed to this failure. As a result, the agent becomes stuck, with multiple "Launch action already pending" warnings appearing in the logs. This prevents certain containers from starting and leaves them in a "absent" state.
2025-02-05 11:45:35,039 INFO aos.cluster.container_launcher:Update task container aos-offbox-10_42_29_205-f status: state='absent', error=Not found 2025-02-05 11:47:06,379 WARNING aos.cluster.container_launcher:Launch action already pending for container 'aos-offbox-10_42_84_192-f' requests.exceptions.ReadTimeout: UnixHTTPConnectionPool(host='localhost', port=None): Read timed out. (read timeout=60)
The ContainerLauncherAgent currently operates on two threads. The first thread identifies containers that need to be scheduled and sends them to the second thread, which interacts with Docker. If a Docker connection timeout occurs, the second thread fails, preventing the agent from completing its tasks. Because there is no automatic retry mechanism in place, the container goes missing.
Customers can workaround this issue by restarting the ContainerLauncherAgent. It will cause the containers relaunched correctly.
To resolve the issue, the user can log into the controller VM and restart the ContainerLauncherAgent using the steps listed below.
1. Log in to the controller VM. 2. Execute "ps -ef | grep -i ContainerLauncherAgent" to check the process ID (PID) of the ContainerLauncherAgent. 3. Run "sudo pkill -f ContainerLauncherAgent" to kill the ContainerLauncherAgent process. 4. Confirm that the process has restarted by running "ps -ef | grep -i ContainerLauncherAgent" again.
When attempting to convert leaf switches from ESI to MLAG within the same blueprint, the operation fails because Apstra does not allow mixing ESI and MLAG redundancy models at the rack level. This restriction is enforced starting in Apstra 4.2.0 and is expected behavior. During the operation, users may see the following error in the UI or REST API Explorer:
"Combining MLAG and ESI leaf pairs not supported"
Currently, converting ESI racks to MLAG racks or vice versa requires replacing all racks within a single FE operation using the REST API. However, due to limitation, this conversion can cause BuilderAgent to fail when links are present between external generic system and leaf switches. In this occurs, please revert the changes and follow the steps outlined in workaround section.
To successfully convert the rack type using modify-racks API, the following workaround can be used:
1. Identify the leaf switches connected to the external generic system 2. Identify the Connectivity Templates (CTs) associated with the external generic interfaces and unassign the corresponding application points 3. Delete the links between the leaf switches and the external generic system 4. Undeploy and unassign the leaf devices from the blueprint 5. Unassign interface maps from the blueprint 6. Use the REST API /api/blueprints/{blueprint_id}/modify-racks to convert the racks from ESI to MLAG 7. Import MLAG-compatible interface maps and assign them to the switches 8. Recreate the links to the external generic system 9. Reassign the endpoints to the appropriate Connectivity Templates
For additional guidance, please contact Apstra Technical Support.
While viewing/editing CT (Connectivity Template)s across blueprints, it's possible that the CT may incorrectly display empty field values.
By clicking the browser refresh button, CT would display the correct data.
Users experience confusion when committing changes for specific devices because the Dashboard shows 'Pending Service Config' for all devices, which can mislead them into thinking other devices are being updated as well. This is a known behavior in Apstra's current design. When a commit is made, all devices temporarily enter a 'Pending' state while the system determines which devices require changes. Even devices that don't need updates briefly show as pending, which can create the false impression that changes are being made. Additionally, as the number of devices in a blueprint increases, the delay becomes more noticeable because Apstra processes each device sequentially. This raises concerns about performance and efficiency when managing larger blueprints.
There is no immediate workaround. The behavior is aligned with the current system design.
In version 4.2.0, Apstra introduces the capability for users to forcibly delete a Virtual Network, even if it has active endpoints. Apstra will initially display the interfaces to which the Virtual Network (VN) is currently allocated and prompt the user to confirm the deletion. It's important to note a limitation in the current design: if a user deletes a VN assigned in a CT where Multiple VLANs are present, all active endpoints will be unassigned.
User should manually remove the specific VLAN from the CT before proceeding to delete it from the Staged > Virtual Networks section.
Starting with EOS 4.30+, Arista EOS introduces a stricter hardware validation mechanism through a new default system l1 configuration block. This change causes Apstra deployment to fail if speed is configured on ports without compatible transceivers, even if the ports are unused.
system l1 unsupported speed action error unsupported error-correction action error
To maintain compatibility, system agent has been updated to detect this configuration and automatically change the unsupported speed action error to unsupported speed action warning during agent installation. This ensures that systems with unused or unpopulated interfaces do not fail deployment due to this stricter validation. The updated system agent logic safely applies this change only for EOS 4.30.x and later.
When the device is deployed or drained, Apstra showed a noticeably longer delay in finishing the operation than the Apstra 4.1.X release. The problem was linked to the significantly increased delay in the Jinja configuration rendering area following Apstra's migration from Python version 2 to version 3. Additionally, it affects the rendering configuration for the blueprint's configlet processing.
Recommend upgrading to the Apstra 6.0.0 release, which addressed the issue. In case of 4.2.X customer, 2-step upgrade (4.2.X -> 5.0.1 -> 6.0.0) is required
DeviceTelemetryAgent.{pid}.log files in /var/log/aos/ in the offbox agents become large and can fill up the disk
The following Python script can be added to run via crontab on an hourly basis, which will clean up older log files. This workaround needs to be applied to controller VM and worker VMs where offbox agents are running (nodes with offbox tags in the Platform/Apstra Cluster/Nodes).
copy from next line
# Copyright 2024-present, Apstra, Inc. All rights reserved. # # This source code is licensed under End User License Agreement found in the # LICENSE file at http://apstra.com/eula import configparser import json import os import re import shutil import subprocess import traceback SystemIdPattern = re.compile(r'AOS_SYSTEM_ID=offbox,(.+),(.+)') def update_aos_conf(task_id): aos_config = os.path.join( '/var/lib/aos/conf.d/task/offbox/', task_id, 'aos.conf', ) parser = configparser.ConfigParser() if os.path.isfile(aos_config): parser.read(aos_config) if not parser.has_section('logrotate'): parser.add_section('logrotate') if 'max_kept_backups' not in parser.options('logrotate'): parser.set('logrotate', 'max_kept_backups', '1') staging_file = aos_config + '.staging' with open(staging_file, 'w') as f: parser.write(f) shutil.move(staging_file, aos_config) return True return False def refresh_logging_infra(container_id): subprocess.check_output([ 'docker', 'exec', container_id, 'pkill', '-HUP', 'DeviceKeeperAge' ]) def get_offbox_containers(): containers = subprocess.check_output([ 'docker', 'ps', '-q', '--filter', 'label=AOS_CLUSTER_APPLICATION=offbox', ]).decode() return containers.splitlines() def get_task_ids(): def extract_info(container_env): try: envs = json.loads(container_env) except ValueError: return None, None for env in envs: matched = SystemIdPattern.match(env) if matched: return matched.group(1), matched.group(2) return None, None containers = get_offbox_containers() if not containers: return containers_env = subprocess.check_output([ 'docker', 'inspect', '--format', '{{json .Config.Env}}', *containers, ]).decode() for line in containers_env.splitlines(): task_id, container_id = extract_info(line) if not task_id: print('Failed to extract task id from container env: {}'.format(line)) continue yield task_id, container_id def main(): for task_id, container_id in get_task_ids(): try: if update_aos_conf(task_id): refresh_logging_infra(container_id) except: print('Failed to update aos.conf for container: {}'.format(container_id)) traceback.print_exc() main()
script finishes here. don't copy this line
Even if the above workaround is applied, there is a chance of filling up partition. The below command can be executd with root permission to clean up logs quickly in the controller VM and worker VMs.
find /var/log/aos/task -name "*.log" -size +10M -print | grep "DeviceTelemetry" | xargs -I {} sudo cp /dev/null {}
Execute CLI commands in the Juniper device supported only show and request chassis beacon commands in the Apstra < 6.1.0 environment. Additional commands (ping and traceroute) are introduced in Execute CLI commands in the Apstra >= 6.1.0 environment for easier troubleshooting environments.
Apstra gRPC probing of JUNOS devices is causing large ephemeral database files which may fill the disk and cause issues accessing the device via SSH.
jtac-QFX5120-48Y-8C-r011 mgd[14126]: UI_EPHEMERAL_COMMIT: User 'root' has requested commit on 'junos-analytics' ephemeral databasejtac-QFX5120-48Y-8C-r011 mgd[14126]: UI_EPHEMERAL_COMMIT_COMPLETED: commit complete on 'junos-analytics' ephemeral database
Add the following to JUNOS devices' pristine config or to the configlet to reduce the number of stored versions within the ephemeral database to reduce its size. Even if JUNOS 23.2R1 introduced the recycling feature of the ephemeral database, the issue can still happen in the Apstra Qualified JUNOS (not EVO) version, 23.4R2-S4.
set system configuration-database ephemeral purge-on-version 30
Because the JUNOS device might not send the full snapshot of interface-related data per reporting interval (default interval = 120 seconds) after initial synchronization, Device Telemetry Health detects continuous anomalies for gRPC Periodic Response Timeout in the interface telemetry service in a high-scale environment. The problem still exists even if the reporting interval is extended.
It is advised to disable gRPC globally, which forces all gRPC-related telemetry services (interface, MAC) to switch to polling mode rather than gRPC, since the problem may occur at random on several devices.To disable the gRPC service globally, change grpc_enabled = 0 in the /etc/aos/aos.conf file and then restart AOS service in the Apstra Controller.
[telemetry_global_config] # Python multithreading enable/disable knob for telemetry collection multithreading_config = 1 # Execution timeout for extensible telemetry collectors command_timeout = 120 # Knob to enable/disable gRPC based service collectors grpc_enabled = 0
As a health check for the device's gRPC server, Apstra 5.1.0 incorporated a new gRPC probe. This results in excessive logging on Junos devices, which causes log messages to roll over more quickly than intended.
Mar 7 12:19:00 leaf2 mosquitto[17106]: New connection from 128.0.0.4 on port 21883. Mar 7 12:19:00 leaf2 mosquitto[17106]: New client connected from 128.0.0.4 as client-4-re0-NA_event_subscriber-re0 (c1, k0). Mar 7 12:19:00 leaf2 mosquitto[17106]: New connection from 128.0.0.4 on port 21883. Mar 7 12:19:00 leaf2 mosquitto[17106]: New client connected from 128.0.0.4 as client-4-re0-NA_periodic_subscriber-re0 (c1, k0). Mar 7 12:19:01 leaf2 mgd[69155]: UI_EPHEMERAL_COMMIT: User 'root' has requested commit on 'junos-analytics' ephemeral database Mar 7 12:19:01 leaf2 mgd[69155]: UI_EPHEMERAL_COMMIT_COMPLETED: commit complete on 'junos-analytics' ephemeral database Mar 7 12:19:05 leaf2 mgd[69187]: UI_EPHEMERAL_COMMIT: User 'root' has requested commit on 'junos-analytics' ephemeral database Mar 7 12:19:05 leaf2 mgd[69187]: UI_EPHEMERAL_COMMIT_COMPLETED: commit complete on 'junos-analytics' ephemeral database
A filter can be applied via configlet for both the messages log files on the device and syslog hosts like the below example.
set system syslog file messages match "!(.*INTERACT-4-UI_EPHEMERAL_COMMIT.*|.*INTERACT-4-UI_EPHEMERAL_DATABASE_PURGED.*|.*mosquitto\[.*|.*AUTH-5-JADE_AUTH_SUCCESS.*|.*jsd\[.*DAEMON-3-UI_CHILD_SIGNALED.*|.*sshd\[.*exited.*)" set system syslog host x.x.x.x match "!(.*INTERACT-4-UI_EPHEMERAL_COMMIT.*|.*INTERACT-4-UI_EPHEMERAL_DATABASE_PURGED.*|.*mosquitto\[.*|.*AUTH-5-JADE_AUTH_SUCCESS.*|.*jsd\[.*DAEMON-3-UI_CHILD_SIGNALED.*|.*sshd\[.*exited.*)"
Because the MAC telemetry service uses the gRPCOnChange mode, the device only sends updates after initial synchronization. When the JUNOS device subscribes to the PATH (/network-instances/network-instance/mac-table/entries/entry) for MAC telemetry service, Apstra's gRPC client (Apstra) receives the first full data from two processes (l2ald, l2aldTM). During the initial synchronization, these processes use their own sequence number range (duplicate range), which makes the gRPC client think that the gRPC packets may be dropped internally. Granular sequence number handling will be introduced in 6.1.0 to address the existing sequence overrun issue.
The default collection period for IBA interface flapping is only 60 seconds and the flapping threshold is 5 times. But, the default collection period for interface service is 2 minutes, the interface will receive updates every 2 minutes only.
So, the default anomaly window is too small to capture 5 interface flaps. For 5 flaps it should be at least 10 minutes.
When a configlet action (import/delete) occurs in a blueprint with a large number of configlets, it can fail with a timeout, and the BlueprintDiffProducerAgent process can crash due to a heartbeat timeout. The problem occurs when the blueprint has a configlet with an incorrect Jinja expression via configlet import/delete actions. Failure of rendering configuration with incorrect Jina expressions can cause all configlets to be re-evaluated for all eligible devices, potentially resulting in much longer configlet processing.
The below steps can be applied as a workaround. If further assistance is needed, please contact the Apstra support team.
1. In order to recover from the crash of BlueprintDiffProducerAgent, please increase the heartbeat_period to 1200 secs in agent_management section of aos.conf file (<=6.1.0: /etc/aos/aos.conf, >=6.1.1: /user/root/etc/aos/aos.conf) and restart AOS service (service aos restart).
[agent_management] # Override the default heartbeat timeout for agents spawned dynamically by # AgentManager. The value must be a non-negative number. The unit is seconds. # The value 0 is used to turn off heartbeat-based agent timeouts and restarts. # The minimum non-0 value allowed is 60. If not provided, then the default # timeout value (600 seconds) is used. heartbeat_period = 1200
2. Please check each configlet in the blueprint and delete the configlet with incorrect Jinja expression from the blueprint.
When a Connectivity Template (CT) includes a single Virtual Network (VN) primitive, multiple BGP peering primitives, and distinct routing policies, incremental configuration changes can occur during endpoint assignment and unassignment. These operations may unexpectedly swap or alter import/export routing policies, potentially disrupting routing configurations and causing traffic interruptions on commit.
In CTs with multiple BGP peering primitives, all BGP primitives are managed under a Batch policy. Apstra does not guarantee the execution order within a Batch Policy, especially during unassign and assign operations, where resources are allocated from a pool, and assignment order is unpredictable. This can lead to the unexpected swapping of routing policies.
It's important to note that this issue occurs only when the CT contains multiple BGP peering primitives with distinct routing policies. CTs with a single VN primitive and a single BGP peering primitive (and associated routing policies) do not experience this behavior.
To mitigate this issue, the following workaround is recommended:
1. Delete the distinct routing policies from the Virtual Network primitive of the affected Connectivity Template (CT). 2. Create separate Connectivity Templates (CTs) for each routing policy, ensuring that each CT corresponds to one routing policy. 3. Assign each CT to the appropriate protocol endpoints.
This workaround will help avoid the unexpected swapping of routing policies during endpoint assignment and reassignment. Please contact Apstra Support Team for more information
Junos RFC5549 BGP peer sessions were rendering the same route-map twice on import/export statements. This only applies to 'ipv6-only' bgp peers.
Configurations were rendering:
neighbor a05:fab:192:168:50::254 { description "facing_leaf1-generic"; local-address a05:fab:192:168:50::1; peer-as 65510; family inet { unicast { extended-nexthop; } } family inet6 { unicast; } import ( RoutesFromExt-default-Default_immutable && RoutesFromExt-default-Default_immutable ); export ( RoutesToExt-default-Default_immutable && RoutesToExt-default-Default_immutable ); }
This has been addressed in 6.1.0, where the import and export statements will contain that route-map entry only once.
This may result in a service disruption on upgrade for those BGP peers as Junos generically may reset the peer when it detects any import/export policy reference change.
This will also apply to ipv6-only, non-EVPN blueprints for route-maps such as "LEAF_TO_SPINE_FABRIC_OUT" between all superspine/spine/leaves.
The JUNOS 23.4R2-S5-EVO on-box agent experiences a 2-minute timeout for the NETCONF operation with a silent XML parsing error message without raising an explicit error when it performs a get-configuration operation to compare against the golden configuration or a lock configuration operation to update any incoming changes via NETCONF. In an attempt to restore after the timeout, the agent starts the reconnect logic; however, this fails because it fails to include the proper binding step.
When the EVO on-box agent's service configuration is stuck in the deploy state, the user must restart the AOS services. It is advised to upgrade the on-box agent after updating Apstra Controller to version 6.0.0 before device NOS is upgraded to 23.4R2-S5-EVO.
The license information from the previous version of the Apstra controller is not transferred to the upgraded controller when it is upgraded. As a result, license information is missed in the upgraded Apstra controller.
License files need to be copied from the old controller to the upgraded controller via the below command, and then restart the license container in the upgraded controller.
sudo scp -r admin@{old_controller_ip}:/var/lib/aos/.license /var/lib/aos/ docker restart aos_license_1
In ESI-based EVPN deployments, the MAC Monitor probe may incorrectly report MAC addresses as missing even though they are present on the device. This can cause Virtual Networks Containing Systems With Missing MAC Addresses anomalies to remain active even after the underlying network issue has been resolved or maintenance has been completed.
Edit the MAC Monitor probe and disable the Raise Anomaly option. If required, disable and then re-enable the MAC Monitor probe to clear the existing anomaly state. Keep Raise Anomaly disabled until upgrading to Apstra 6.2.0, as the issue may recur before the fix is applied.
Upgrade script in 4.2.2 and following releases introduced a regression that manifests itself in multi-step migration scenarios.Step-1. Upgrade from 4.2.1 -> X (eg. 4.2.2 ) - causes the permission on '/var/lib/aos/metricdb/iba' folder and it's sub-directories to be 700.Step-2. Upgrade from X (eg. 4.2.2) -> Y (eg. 5.0.1) - causes the silent failure in step that copies '/var/lib/aos/metricdb' folder and ALL its sub-directories to new VM.
The impact of this failure is that the 'Audit', 'IBA stage history' and 'Aos cluster health history' data is lost in the final upgraded AOS instance. The data from the previous release will be lost subsequently if there are further migration steps involved.This issue affects all the releases starting upgrade from 4.2.2. If your upgrade source Apstra is at least 4.2.2, please apply the workaround suggested BEFORE performing the upgrade.
Apply workaround fix(aos_54413_fix_metricdb_permissions.run: https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000Gc0kCIAR) to the old version Apstra Controller Node *BEFORE* every upgrade, following the below steps.
1. Copy the bundle aos_54413_fix_metricdb_permissions.run to the source (old) Apstra controller node. The tool expects Apstra service to be running because it needs to get cluster node information from Sysdb.
2. Make it as executable and execute the bundle as sudo
admin@aos-server:~$ chmod 755 ./aos_54413_fix_metricdb_permissions.run admin@aos-server:~$ sudo ./aos_54413_fix_metricdb_permissions.run Verifying archive integrity... All good. Uncompressing Fix for AOS-54413 for AOS >= 4.2.2 100% AOS[2025-05-25_19:36:22]: Fixing controller node AOS[2025-05-25_19:36:23]: Getting cluster node metadata AOS[2025-05-25_19:36:24]: Fixing worker node: 10.28.75.6 Logs have been collected at: /home/admin/aos_54413_fix_logs_20250525_193623.tar.gz
3. The absence of any errors means that the issue has been fixed. In case of errors during execution, please reach out Juniper Apstra Support Team.
If AOS instances upgraded without work-around and if old apstra VM is preserved, Contact Juniper Apstra support team to help with migrating MetricDB data.
A banner configured in the Cisco NXOS device must be single line. Multiline banner (motd or exec) is not supported.
Configure single line banner (motd or exec)
The NOS upgrade procedure must parse the pristine configuration in order to determine which ports must be disabled when the Skip Shutting Down Interface During Upgrade option is not checked in the Advanced Settings of Managed Device. The NOS upgrade would fail if the configuration line in the pristine configuration extended into multiple lines because parsing the command line misses the end-of-command-line character (.
Recommend re-onboarding of the device with a clean, pristine configuration.
When the Junos EVPN Next-hop and Interface count maximums parameter in the staged->Fabric settings->Fabric-policy is enabled, Apstra introduced modifying the default hardware settings for VXLAN routing's resource (next-hop and interfaces) for QFX5110, QFX5120, EX4650, and EX4400 devices in the rendered configuration () starting with version 4.2.0. Whenever configuration changes in VXLAN routing's resource, JUNOS triggers PFE automatic restarts to reflect new changes with service impact. The typical scenarios would be when the device becomes deployed, undeployed, or the device is in NOS upgrade. To prevent unnecessary PFE restarts in those scenarios, the configuration for VXLAN routing's resource needs to be included in the pristine configuration.
If the Junos EVPN Next-hop and Interface count maximums parameter in the staged->Fabric settings->Fabric-policy is enabled, add the below configuration into the device's pristine configuration.QFX5120 and EX4650 VXLAN routing's resource
forwarding-options { vxlan-routing { next-hop 45056; interface-num 8192; overlay-ecmp; } }
QFX5110 VXLAN routing's resource
forwarding-options { vxlan-routing { next-hop 32768; interface-num 8192; overlay-ecmp; } }
EX4400 VXLAN routing's resource (add overlay-ecmp if Junos EX-Series Overlay ECMP is also enabled)
forwarding-options { vxlan-routing { next-hop 16384; interface-num 6144; overlay-ecmp; } }
Customers may encounter the following Server-side Validation Error in the Web UI when the pristine configuration contains multiple system stanzas which is not a expected behvavior:
"Cannot parse config: system already parsed."
According to ScotchInventoryAgent logs, POST requests to update the pristine configuration failed with a 422 Unprocessable Entity error, indicating a validation issue:
2025-02-17 23:31:18,730 680:INFO:aos.scotch.libs.scotch_flask:request: POST /api/systems/AN10555621/pristine-config HTTP/1.0 34074 bytes 2025-02-17 23:31:18,737 680:INFO:aos.scotch.libs.scotch_flask:response: 422 55 bytes 0.007347 seconds
Background of the issue:
1. In Apstra 4.2.x, gRPC was introduced to support Telemetry Streaming, and as a result, having two system blocks in the pristine configuration was expected in that release. 2. Starting from Apstra 5.0.0, enhancements were made to automatically merge multiple system stanzas in the pristine configuration during the NOS upgrade process. 3. If a customer chooses to remain on their current NOS version for an extended period, multiple system stanzas can exist in the pristine configuration without causing issues.
If a customer chooses to remain on their current NOS version for an extended period and needs to forcefully update the pristine configuration, they should manually merge the system stanzas within the pristine configuration using the UI and then perform a Force Update.
For further assistance, please contact Juniper Apstra Support.
When Junos devices become unreachable, triggering a commit check leads to a rejected state after approximately 600 seconds. However, devices in this rejected commit check state are put into commitCheckinProgress state after an AOS restart or Full Config Push.The problem occurs because Apstra does not treat devices in the commitCheckInProgress state as pending. This leads to a misleading UI/UX experience (some devices may still be in the commitCheckInProgress state, but the main dashboard may show zero pending devices).
There is no immediate workaround for this issue. Needs to check device status explicitly by clicking the pending icon in the service config section of the dashboard, even if the pending icon shows 0. However, once management connectivity is restored, the Deployment Status updates correctly as expected.
The user reported that an IP conflict occurred while stretching the new VLAN to switches that did not have the routing zone, which caused the build error. When the routing zone is extended into the switch with a static IP address for the loopback IP address, the IP address may already be in use in another routing zone dynamically because the IP pool is shared by those routing zones. IP conflict error messages can occur because the same IP address is assigned statically in one routing zone and dynamically in another routing zone.
There are two workarounds available to cater the issue 1. moving to only dynamic addresses for VRF loopback addresses by clearing static IP address or2. Disabling VRF optimization in the blueprint to keep all static IP addresses being allocated in all time.
When a user collects a controller Show Tech by UI with the include backup option enabled, the generated Show Tech also includes a backup of the aos-edge-auth.json file. This file contains sensitive information such as the registration_key, org_id, and other credentials used to authenticate with JCloud.
If this backup is restored in a different environment, the apstra_edge container in that environment may use these credentials to connect to JCloud. This can unintentionally trigger an Edge re-registration event, disrupting the original environment by recreating its active Edge instance.
The Apstra Device Profile for the Dell S5448F-ON is not enforcing the device constraints for the total number of logical ports per pipeline. Customer should be aware of the constraints around pipelines in the specific model and should take care to not exceed 18 logical ports per pipeline.
For further information, please refer to the documentation of the vendor.
Rebooting the device or restarting FRR in SONiC may cause the FRR running configuration sections (related with route-map) to be rearranged. The rearranging of sections will typically show a configuration deviation even if the running configuration is exactly the same as before.
The anomaly can be eliminated by the user reviewing the deviation and accepting the changes. No further action is necessary.
The device console may display cosmetic Python traceback error messages indicating an unreachable network when Sonic ZTP is provisioning a device. In the event that the device modifies the VRF of its management interface and is unable to send logging messages to the ZTP server, this may occur.
If the device completes the ZTP process, these messages in the console can safely be ignored.
When a port group in the vCenter is configured with a private vlan mode, the VLAN specification contains the pvlanid property instead of the vlanid property. However, the vlanid property is always expected from the port group's vlan specification by the Device Telemetry Agent's collector if it is not trunk mode. Anomalies could be reported if the collector's execution fails due to a reference to the vlanid property, which is nonexistent.
Recommend not using port group with private vlan in the Virtual Distributed Switch.
Sutatined Optical Threshold anomalies are frequently observed over disconnected (not connected) interfaces in the JUNOS device. The main reason for the problem is that the JUNOS device sometimes reports the received average power as a very low value or - Inf value when the interface is disconnected, and Apstra does not correctly parse the - Inf value. When a very low power value falls below the warning level, Apstra creates an anomaly for the low receive power port. But when the same interface reports a - Inf value that is later incorrectly parsed, Apstra eliminates the interface from the list of interfaces with an optical transceiver, thereby resolving the raised anomaly falsely. This is the reason anomalies are frequently raised and cleared over the same interface. No workaround is available for this issue
None
Apstra uses application weight information for container types (iba, offbox, and apstra_edge) prior to launching a container. The upgrade process for only upgraded environments (5.0.X to 5.1.0) fails to include application weight information for the apstra_edge container type, causing the TaskScheduler agent to crash and restart. When AOS is restarted, the situation may worsen (all offbox and iba agents remain down). Please apply a workaround to resolve the issue. This issue would not exist in a 5.1.0 clean deployment environment.
Run REST API call with PUT operation to /api/cluster/application-weight using the below payload information, then restart aos (service aos restart).
/api/cluster/application-weight
{ "iba": 1000, "offbox": 250, "apstra_edge": 500 }
Upon navigating from the Stage tab to the Uncommitted tab, the page does not load and continues to display the loading spinner. The UI is continuously polling the API endpoint /api/blueprints//tasks?mode=full, which forces the backend to attach the complete task payload to every response. Since these payloads can contain large data information, the API responses become significantly heavier. Because the UI repeatedly polls this endpoint, it increases backend processing time and causes the page loading issue, particularly when tasks have large payloads.
When a worker node with an offbox tag is isolated from a controller node due to a network failure, offbox containers running on the worker node are relocated to the other eligible nodes (including controller) with an offbox tag. Even if the isolated worker node recovers from the failure, the UI still displays the launched error status for the offbox containers in the relocated node (the other worker node or controller) rather than the recovered worker node. The duplicate offbox containers are also observed from both nodes (the relocated node and the restored worker node) during the period. The long delay in convergence after the failed worker node is recovered from network failure is due to an increasing retrial interval for connecting to the cluster from the isolated worker node as a retrial for connecting fails with several attempts. Restarting the AOS service in the worker VM as a workaround allows the worker node to converge quickly after the worker node is reconnected to the cluster.
The user needs to restart the AOS services on the worker node, which reconnects to the controller
When the Device Profile allows multiple transformations for the same speed and the IM (Interface Map) needs mixed transformation for the same speed across ports (for example, Port 1: 1X100G, Port 2: 4X100G, etc.), it is not possible to create an IM using the UI.
Recommend using the same transformation for the same speed across ports in the IM (Interface Map) when IM is created in the UI
The current default timeout for collecting show tech via UI is 20 minutes (1200 seconds) in the controller_timeout value of the show_tech section of aos configuration file /etc/aos/aos.conf. The collection contains operations for exporting all binary log files from all running agents across containers into printable format files. Depending on the number of accumulated log files, the translation processing time may exceed the default timeout for collecting showtech, resulting in timeout failures for the UI showtech collection task.
/etc/aos/aos.conf
The timeout value for controller_timeout under the show_tech section of /etc/aos/aos.conf file should be increased to a higher value, such as 1 hour and 30 minutes, followed by restarting aos service (sudo service aos restart). Otherwise, a manual command (aos_show_tech) can be used to collect controller show tech.
Users will not be able to create a Datacenter Blueprint using the EVPN reference design with IPv6 RFC:5549 as the Spine to Leaf Links Underlay Type, as this configuration is not currently supported. This limitation is tracked in RFE-1364 which is planned for release in 6.1.0. If the user attempts to create an unsupported blueprint using the steps below, the UI will throw a critical error message - EVPN is not allowed with spine-leaf IPv6 links addressing policy.
Steps to Reproduce: 1. Create a template with the overlay control protocol set to MP-EBGP EVPN. 2. Navigate to the Blueprint Creation page. 3. Select Datacenter as the reference design. 4. Set Spine-to-Leaf Links Underlay Type to IPv6 RFC:5549. 5. Attempt to create the blueprint.
However, the blueprint becomes locked and cannot be deleted, even with admin privileges. Deletion can only be performed via REST API Explorer or CLI. This issue was addressed in 6.0.0, ensuring that the delete function works correctly.
Please follow the steps below to delete the failed and locked blueprint:
1. Navigate to Platform -> Developers -> REST API Explorer. 2. In the REST API Explorer, go to blueprint -> {blueprint_id} -> DELETE. 3. Select the locked blueprint ID and click Try It. 4. Expect a success response code to appear on the same page.
Please contact Apstra support for assistance if you encounter any issues during this process.
Integrated DCI feature(vxlan stitching) was introduced in Apstra 4.2.0. In versions 4.2.x and later, customers using this feature may encounter an issue where all VTEP loopback addresses from the fabric including those from non-border leaf devices are being advertised to external routers over BGP in the default routing zone.
This affects only VXLAN DCI Stitching deployments(Stitching requirement for VTEP loopbacks for only border leaf nodes vs OTT requirement for all VTEP loopbacks in the fabric). Even when customers configure routing policies to export only loopback of border leaf nodes, Apstra backend logic automatically includes all VTEP loopbacks. Due to the current design, Apstra does not differentiate between border and non-border leaf roles in this context, resulting in the unintended advertisement of all fabric loopbacks to external peers.
Engineering has confirmed this as a bug. The expected behavior is to advertise only the loopback addresses of border leaf switches to external routers in the default routing zone. There is no official workaround to modify this behavior through standard configuration. Engineering is actively working on a fix to address this issue in a future release. The only option is to use a custom configlet to override Apstra default export logic. Please reach out to Apstra Technical Support for assistance.
User-defined import / export route-targets with ":0" are rejected with validation errors.
{ "rt_policy": { "import_RTs": { "0": "Type 0 RD X:Y must be in format 2-byte ASN:4-byte value. Provided value: \"65500:0\"" } }
When the virtual infra manager is removed from the Apstra controller, Apstra should have cleared any data related to the virtual infra manager. Because it's not cleared, when the same virtual infra manager is added back to Apstra later, old data is still used together with the new collected data from the virtual infra manager's collector. In some scenarios, when old, uncleaned data has an error condition, it can trigger continuous error even if newly collected data doesn't have an error condition.
If the virtual infrastructure manager requires re-onboarding (removing and then adding back) from Apstra, the user must take the actions listed below.1. Remove virtual infra manager from Apstra Controller (External Systems/Virtual Infra Managers).2. Restart the AOS service.3. Add the virtual infra manager back to to the Apstra
Since the UI misses polling of the node detail information to the Apstra backend, the Virtual Network Endpoints view of Generic System Node (Staged > Physical > Topology > Virtual Networks Endpoints) shows empty information.
Refresh Web page in the browser to make the UI send requests explicitly to collect data
Apstra creates a relationship between the PNIC and the Link Discovery Policy (which determines which discovery protocol is used) configured in the VDS when a PNIC is assigned to a VDS (Virtual Distributed Switch). One PNIC may inadvertently become linked to two relationships without clearing out the previous relationship when a user moves a PNIC directly from one VDS to another VDS. An error message below appears when the PNIC becomes unassigned from VDS because there is more than one relationship between the PNIC and Link Discovery Policy that is invalid.
virtual_infra failed to collect data, plugin raised exception: {'item_iter': <aos.sdk.graph.graph.RelationshipIterator object at 0x7f60c03f9570>, 'items': [df57a2f0-969c-4dee-9831-a53526bd7d5a-[:policy]->4c8c4c31-df9a-4188-933f-6b5d67703a1f, df57a2f0-969c-4dee-9831-a53526bd7d5a-[:policy]->1787d662-4a41-4b8b-9b77-acac5508e771]}
Instead of performing one direct migration action from one VDS to another, the problem can be avoided by two actions: unassigning the PNIC from the old VDS and then assigning it to the new VDS.
Procedures for fixing the errors as a workaround1. Remove the virtual infra manager from not only the blueprint but also the External Systems/Virtual Infra Managers.2. Restart the AOS service.3. Add the virtual infra manager back to the External Systems/Virtual Infra Managers and then blueprint.
When multiple vNICs from a single VM are assigned to the same vNET (port group or VDS) in the Virtual Infra, the VMs Without Fabric Configured VLANs probe raises anomalies in the Analytics->Anomalies because the graph query in the VMs backed by Fabric VLANs processor treats those vNICs as identical.
Please clone existing VMs Without Fabric Configured VLANs probe with a different name, and then modify graph query in the VMs backed by Fabric VLANs processor to include vnic into the existing distinct statement like the below. If further assistance is needed, please contact Juniper Apstra Support Team.
Graph Query: match( node('system', name='server', role='generic', management_level='unmanaged', external=False) .out('hosted_interfaces') .node('interface', name='server_intf') .out('hosted_vn_endpoints') .node('vn_endpoint', name='vn_endpoint') .in_('member_endpoints') .node('virtual_network') .out('instantiated_by') .node('vn_instance', name='vn_instance') .having( node(name='vn_instance') .in_('hosted_vn_instances') .node('system', system_id=not_none(), deploy_mode='deploy') .out('hosted_interfaces') .node('interface') .out('link') .node('link') .in_('link') .node('interface') .in_('hosted_interfaces') .node(name='server'), at_least=1 ), node(name='server') .in_('is_realized_by') .node('hypervisor', name='hv'), node(name='hv') .out('hosts') .node('vm', name='vm') .out('has') .node('vnic', name='vnic') .out('part_of') .node('vnet', vn_type='vlan', name='vnet') ) .distinct(['server', 'vn_endpoint', 'vm', 'vnic']) .where(lambda vnet, vn_instance, vn_endpoint: (vnet.vlans == [0] and vn_endpoint.tag_type == 'untagged' or vn_instance.vlan_id in vnet.vlans))
When a user attempts to open the racks or pods tab in the Staged->Physical section, rack schema validation logic will be executed for every rack in the blueprint. In order to avoid incompatibilities with the underlying database, non-ASCII characters cannot be included in the link group label, which might be explicitly supplied when adding a generic system to leaf or access nodes.
When adding generic systems, make sure that the link group label contains just ASCII characters to prevent this problem. The workaround requires using an API call to remove non-ASCII characters from the link group label in the rack. Please get in touch with the Juniper Apstra Support team for more help.
With current cluster design, worker nodes may fail to recover their configuration state after a temporary loss of SSH connectivity to the controller. This condition can be observed in the UI by navigating to Platform > Apstra Cluster > Nodes > Worker, where the following error may be displayed:
Configuration Error: ssh: connect to host 10.28.17.4 port 22: Connection refused
When the controller (ClusterManagerAgent) attempts to push configuration to worker nodes, it retries SSH connections up to three times. If all attempts fail, due to connection refused or no route to host, the node is marked with a FAILED Configuration State. Once this state is set, the system does not automatically retry configuration, even if SSH connectivity is later restored. Although worker nodes may resume sending keepalives and transition back to an active operational state, the configuration state remains in failed state indefinitely. Due to this the overall node state may continue to appear FAILED despite restored connectivity.
Manually trigger a configuration synchronization using one of the following methods:
1. Navigate to Platform > Developers > REST API Explorer 2. Execute REST API: POST /api/cluster/worker/sync [OR] 1. Restart AOS from the controller VM: systemctl restart aos
When sFlow collector is configured with mgmt_instance, the configuration will be ignored in the ACX platform with a warning such as the below example.sflow {polling-interval 10;sample-rate {ingress 10000;egress 10000;}source-ip 10.217.6.15;collector 10.217.0.165 {udp-port 6343;#### Warning: statement ignored: unsupported platform (ACX7024X)##routing-instance mgmt_junos;
##
## Warning: statement ignored: unsupported platform (ACX7024X)
Please use a non-management instance for exporting sFlow until the ACX platform supports a management instance for sFlow export.
Anomalies are raised due to mismatch in the operational status of interfaces due to interface status showing "unknown" on Juniper EX4400-48T devices running Junos 22.4R3.
Restart the Apstra AOS service to collect the right interface status information. This issue is not observed in higher JUNOS versions. Recommend upgrading to an Apstra-qualified higher JUNOS version (>=23.4R2-S4).
When Apstra Edge container is configured with Apstra Flow together in the Apstra Controller VM, Apstra Edge shows 50% CPU utilization, which may affect the performance of the Apstra Controller and other containers. The issue is observed when the Apstra Edge container 0.5.0 image is used. In the latest version, 0.16.5, the issue is not shown anymore.
Please follow the instructions for upgrading the Apstra Cloud Services Edge image based on https://www.juniper.net/documentation/us/en/software/apstra5.1/apstra-user-guide/topics/task/auto-updating-edge-5.1.html after downloading the most recent Juniper Support image for Apstra Cloud Services Edge from Apstra 5.1.0
When the Apstra commit was executed, the JUNOS device with the off-box agent reported BFD underlay flaps between the committed device and the other devices. Apstra currently uses the load override option as the default action for device commits, which may result in high CPU utilization, preventing time-sensitive daemons from acquiring CPU time slices and triggering unexpected events such as BFD timeouts or writing EEPROM errors. Starting with 6.1.0, Apstra intends to use load update as the default commit action for JUNOS and EVO devices.
The workaround to use load update can be applied only to the JUNOS/EVO offbox agent. Please follow the below step to apply workaround1. Navigate into Devices->Managed Devices->{DEVICE_IP}-> Agent2. Click Edit button to edit Agent3. Add an option into Open Options with the key as load_mode and the value as update in the Edit Offbox System Agent(s) window.4. Click Update button
The NXOS rollback feature on the N9K-C93600CD-GX device has significant limitations when the devices' interfaces are broken out.Ports 1-24 in the model are organized into four-port groups: (1, 2, 3, 4), (5, 6, 7, 8), (9, 10, 11, 12), (13, 14, 15, 16), (17, 18, 19, 20), and (21, 22, 23, 24). When port 1 is broken out as 4x10G or 4x25G, port 3 is automatically broken out in the same mode, and vice versa. When any port in the quadruple is split into 2x50G, all four ports are automatically split in the same mode. Similarly, ports 26-28 are organized in pairs of two, i.e. (25, 26) and (27, 28). Both ports in the pair must operate in the same breakout mode.
In most cases where a breakout (or more than one) exists, rollback fails to generate a working rollback patch. The reason for this is that the breakouts cannot be reversed if the remaining broken-out interfaces in the same port group have not been shutdown first. For example, to negate the breakout of port 1, the broken-out interfaces of port 3 must be shutdown, and vice versa. It appears that the rollback logic shuts down the interfaces associated with the port whose breakout is being reverted (port 1 in the previous example), but fails to shut down other broken-out ports in the same port group (port 3).
The safer way for the N9K-C93600CD-GX to be used with AOS is for the customer to avoid using breakouts altogether on the device.No issue with rollback when ports 29-36 have been broken out has been observed. Breakouts on these ports can be rolled back In the case that the last interface of a port-group is the only one used and broken out, would the nxos rollback feature (and rollback to pristine) be successful. However this is highly discouragedIn any other case the only way to reverting to pristine would be to manually shudtown all broken down interfaces before reverting to pristine (or using the rollback to a pristine config)
Apstra may encounter gRPC sequence number overrun for MAC telemetry service, recognized as losing data and initiates re-subscription for service. This sequence can continue repeatedly in a highly loaded environment (>=100K entries), making the JSD to handle continuous subscription and cancel subscription requests with memory leaking. This continous accumulation of memory leaking may lead into process crash by OOM and then triggering device reboot.
The workaround is to disable gRPC in the Apstra. If the customer wants to continue to use gRPC in the Apstra, recommed upgrading to the latest 6.1.X release (which includes fixes for the sequence overrun misleading issue) so that meory leaking can be prevented. Please reach out Juniper Apstra Support team for the further assistance.
gRPC server reset count anomalies are observed in the JUNOS-EVO platform when gRPC Max Client connection limit error occurs in the device due to the problem that gRPC stalled connections are not cleared. gRPC keepalive is not enabled by default on the JUNOS-EVO platform running 22.2R3 or 22.4R3, which is the cause of the problem. gRPC keepalive is enabled for 300 seconds in the >=23.4R2-EVO release to avoid a build-up of stalled gRPC connections.
gRPC Max Client connection limit error
>=23.4R2-EVO release
In JUNOS-EVO device running 22.2R3 or 22.4R3, apply the below configuration via configlet into the device to enable gRPC keepalive or upgrade the device to >=23.4R2-EVO. For further assistance, please contact the Juniper Apstra Support Team.
>=23.4R2-EVO
set system services extension-service request-response grpc grpc-keep-alive 300
When the JUNOS-EVO device running <=23.4R2-S3-EVO is loaded with high-scale configuration, the gRPC subscription for Apstra's MAC telemetry service may fail with a timeout because the l2ald agent process is in hang status. While the device is in the status, the interface telemetry service may fail concurrently. The Device Telemetry Health probe would log anomalies resulting from continuous telemetry service failures.
After the device is back to the stable normal status after internal service restart, the gRPC will be subscribed to the device correctly, and anomalies from the probe will be cleared. Polling can be used instead of GRPC for telemetry services to prevent issues from long delay by GRPC subscription.
When Juniper EVO device hosts DHCP servers in a border leaf role with DHCP relay configuration, DHCP may not work as intended due to an unresolved bug in Junos EVO which prevents DHCP packets from being processed correctly. Please refer to the following KB for dhcp relay limitations: https://supportportal.juniper.net/s/article/Juniper-Apstra-Support-for-Stateless-DHCP-Relay?language=en_US
Using Apstra's configlet feature, create configlet to remove the rendered DHCP relay configurations and apply it to the Juniper EVO border leaf device.
Due to an outstanding bug in all available versions of Junos EVO, when two virtual networks are hosted on the same set of ESI leafs, one with DHCP enabled and one with DHCP disabled, the virtual network with DHCP disabled will also have DHCP requests forwarded.
When Aptra registers a subscription for a specific xpath into Junos EVO device, Junos EVO device may keep multiple subscriptions for the same path and cause telemetry sequence overruns in telemetry data.
When available, upgrade to Junos EVO 23.4R2-S4 or greater
A new default limit for the number of IP addresses per MAC per bridge domain in EVPN (mac-ip-limit) was added in Junos and Junos Evolved Releases 24.2R1, 23.4R2, and 23.2R2. 200 IPs per MAC is the default setting.
Clients who use EVPN fabrics, such as those with MAC-VRF deployments in fabrics managed by Apstra, might observe that MACs linked to more than 200 IP addresses cease to learn new IPs in the EVPN MAC-IP table. The Junos software release introduced this expected behavior.
The fabric can support more IPs per MAC while preserving per-bridge-domain enforcement by setting mac-ip-limit globally using an Apstra Configlet. No software fix is required.
To adjust the limit in an Apstra-managed fabric, create a Configlet in Apstra with the below command, specifying the desired limit:
set protocols evpn mac-ip-limit <desired-value>
Import the Configlet into the blueprint and apply it to the relevant switches.
Note: Although the command is global in Junos, the limit is enforced per MAC per bridge domain, including inside MAC-VRFs.
In the Apsta 4.2 reference design change for MAC-VRF, the Junos "forwarding-options evpn-vxlan shared-tunnels" configuration is added via the Apstra rendered configuration. However, this command requires a device reboot to take effect with the Junos warning "Config: forwarding-options evpn-vxlan shared-tunnels has changed. A system reboot is mandatory". A user doing a Junos upgrade with Apstra may re-experience this issue after the device is upgraded.
To avoid the need to a additional, manual reboot after a device Junos upgrade, the user can add the following configuration to the Apstra device system-agent pristine-configuration.
forwarding-options { evpn-vxlan { shared-tunnels; } }
This can be done in the "Decvices / Managed Devices / Pristine Configuration" Apstra UI or using the Apstra-CLI "system pristine_config_append" command.
A netconf/mgd bug in QFX device running JUNOS >= 23.4R2 and <= 23.4R2-S3 could result in the loss of management access, including SSH, netconf, and gRPC, which Apstra uses to manage the device.
If the device is running the exposed version, recommend upgrading to JUNOS >= 23.4R2-S4
A member leaf in a Sonic MLAG leaf pair does not show the local MAC in its mac table, but because its peer learns it as DYNAMIC, the probe generates expectations that this MAC will be present across the fabric.
When MLAG port-channel across dual nodes is unbundled into 2 x single connected port-channels, Apstra removes the "VPC" configuration on the interface port-channel. This change causes Cisco NXOS to trigger interfaces that remain down with the "IntFailErrDis" state as wrong behavior.
Perform the below steps to recover 1. Remove the port-channel configuration 2. Bounce the error-disabled interface.3. configure port-channel configuration again
For power supplies, Apstra primarily uses the output of the show chassis environment pem or show chassis environment psm command. On the QFX10002-36Q, both PEMs are reported, however, only one includes the XML tag that designates the component class as Power. The power supply information is further augmented using the show chassis environment command. Since the tag is missing for one PEM, Apstra does not recognize it as a valid power supply component. This is a known issue in Junos. Below is an example of the XML output from the show chassis environment without the class tag:
<environment-item> <name>FPC 0 Power Supply 1</name> <status>Present</status> </environment-item>
Due to inconsistent Junos behavior, the Power Supply State Check processor of Apstra does not evaluate the affected PEM, and no anomaly is raised. No workaround is available for this issue. This issue is expected to be addressed in newer Junos versions from 24.4R2, and the corrected behavior is expected to be present in supported versions for Apstra 6.1.0 and later.
If a SONiC device removes and then re-adds an IP address, the device may fail to add a static route involving that address to the kernel routing table, even if the static route configuration exists. The show ip route output in vtysh in the SONiC device experiencing this issue may include the following output lines:
show ip route
vtysh
S>r 198.51.100.2/32 [1/0] via 192.168.0.9, Po1.4, weight 1, 01:59:25 B * 198.51.100.2/32 [20/0] via 10.0.0.3, Vlan201 onlink, weight 1, 01:59:25
Above, the "S" static route has been rejected by the kernel and was not installed.
A full config apply will restore normal operation, in case such a problem occurs.
IPv6 BFD sessions experience flappings in the Cisco C9504 chassis product running NXOS 9.3.13 or 10.2.6 when default echo mode is enabled.
Please disable BFD echo mode under the interface which runs IPV6 BFD using configletExample configuration)
interface Ethernet1/2 no bfd echo
This issue was discovered in the Juniper QFX 5230/5240 Juniper EVO device for non-EVPN blueprints (Pure IP Fabric) running the 23.4R3-S3-EVO version. When an interface is assigned both untagged and tagged vlans and then the untagged vlan is removed, the traffic from the leaf device will have the incorrect VLAN ID, causing network connectivity issues.
Please unassign the interface from CTs (Connectibity Template) with the tagged vlans, followed by a commit, and then reassign CTs (Connectibity Template) to the interface. followed by commit.
2026-07-23: Added AOS-62738
2026-05-26: Added AOS-61246
2026-05-15: Updated AOS-54864
2026-05-14: Updated AOS-42437
2026-04-03: Added AOS-55780
2026-03-23: Added AOS-60167
2026-03-18: Removed AOS-51029,Added AOS-59821,AOS-59842,AOS-60114
2026-03-03: Added AOS-47430,AOS-59726
2026-02-18: Added AOS-54971
2026-02-03: Added AOS-59198
2026-01-08: Added AOS-48525
2025-12-11: Added AOS-58680
2025-11-18: Added AOS-58200
2025-11-10: Updated AOS-57490
2025-10-02: Added AOS-56571, AOS-57297, AOS-57315, AOS-57490
2025-09-17: Added AOS-57099
2025-09-08: Added AOS-52744, AOS-56498
2025-08-26: Updated Download link for AOS-53355, AOS-54413
2025-08-20: Updated AOS-52519, Added AOS-56175, AOS-56282
2025-08-13: Added 40023, AOS-43348, AOS-45139
2025-08-04: Updated AOS-53355
2025-07-29: Added AOS-52519 and Updated AOS-51846
2025-06-26: Added AOS-54006
2025-06-25: Added AOS-54864
2025-06-09: Added AOS-54497, AOS-54513
2025-05-27: Added AOS-54413
2025-05-22: Added AOS-50584, AOS-54391
2025-05-19: Added AOS-51083, AOS-54198
2025-05-08: Added AOS-44623, AOS-49957, AOS-52731, AOS-53919, AOS-54088
2025-05-01: Updated AOS-51964. Added AOS-53940, AOS-53627, AOS-53602.
2025-04-17 : Added AOS-53526,AOS-53534,AOS-53783
2025-04-12 : Updated AOS-47861
2025-04-04 : Added AOS-53537,AOS-53538
2025-04-01 : Updated RFE with Feature Category
2025-03-27 : Updated RFE-3241(JUNOS Qualified Version: 23.4R2-S4, removed 23.4R2, 23.4R2-S3), Added AOS-50120, AOS-50216, AOS-52766, AOS-52774, AOS-52892, AOS-52927, AOS-53355
2025-03-21 : Added AOS-52788, AOS-53023, AOS-53220, and Removed AOS-50283
2025-03-14 : Added AOS-52300
2025-03-13 : Updated AOS-43808
2025-02-17 : Added AOS-52519, AOS-52484, AOS-52431, AOS-52339, AOS-51579, AOS-51029
2025-01-31 : Added AOS-52217
2025-01-29 : Updated AOS-51846
2025-01-21 : Updated AOS-49934, added AOS-50710,AOS-50912, AOS-51846, AOS-51964
2025-01-09 : Updated AOSEXT-21
2025-01-07: Updated AOS-42437 and added AOS-51549, AOS-50112.
2025-01-02: updated
2024-12-20: initial published
2024-12-12: initial creation