Juniper Apstra software product version 6.0.0 is available to licensed, registered Juniper customers from the Juniper Apstra software download site.Documentation for Juniper Apstra 6.0.0 is available from the Juniper Apstra documentation site.
NA
You now can use in your AI Blueprints the predefined configlet for DCQCN, which can be found in the catalog under the label "Datacenter QoS Congestion Notification". Use the property-set with the same label to adjust the configuration parameters such as the drop profile's fill level.
You can now create a Rail-Aligned Rack-type in a Collapsed design. Use the "Rail-Collapsed" fabric connectivity design in the Create Rack type menu. This creates a Rail-only rack structure, without spines, useful when the server count can be accommodated within a single Stripe. By eliminating the need for spine switches, you can achieve significant cost reductions. However, it's important to note that migration from a Rail-collapsed to a Clos design is not possible. To ensure optimal performance, verify that the GPU node is compatible with PXN for NVIDIA or its equivalent for other vendors. PXN enables NVIDIA NVSwitch connectivity between GPUs within the node, allowing data to move to a GPU on the same rail as the destination before being sent to the destination without crossing rails. The Simplified AI Template Designer also supports the Rail-Collapsed design.
You can now create a Rack-Type with Rail support as a first-class citizen component for constructing your AI blueprints. The Rail-Index is automatically designated, but you have the option to manually specify it if you prefer. You can choose multiple generics and a single leaf to form interfaces within the same Rail. Additionally, you have the flexibility to merge rail-aligned and traditional server connectivity within a single rack.
You can create in your AI Blueprints a Load Balancing Policy and choose between Standard DLB (Dynamic Load Balancing) or GLB (Global Load Balancing). You can choose between Flowlet or Per Packet mode and customize all critical configuration parameters, including Inactivity Timer, Sampling Rate and Egress Quantization parameters. Once defined, apply the load-Balancing policy selectively or in bulk on the target device(s). The system automatically prevents you from deploying unsupported policies, such as an attempt to deploy GLB on a topology where not all devices are QFX5240 (mandatory requirement for GLB to operate).
You can now use the simplified AI template designer, which takes 5 inputs to generate a template that can then be deployed as a blueprint. These inputs are: Total Number of Servers, GPUs per Server, Servers Per stripe, GPU NIC speed and Oversubscription ratio. This allows you to quickly and easily create a validated and optimized Apstra template tailored to your specific resource requirements. You can also export these racks from the template to use them for day 2 operations. The Template designer uses by default the Juniper recommended devices, including the following Device Profiles: QFX5220, QFX5230-64CD, QFX5240, and PTX10004/8/16 with 36CD Line card. You can use the 'Advanced Calculator Settings' section to customize this list by adding/excluding specific models. The Template designer will select Generic Systems Logical Devices, which can then be mapped to NVIDIA DGX Device Profiles.
You can now include in your AI Blueprints an automatic pre-calculation and pre-provision of the required Virtual Networks, along with their associated Connectivity Template definition and Connectivity Template assignments. Under Physical -> Racks, you will find a new "Rails" tab, which will list any missing Virtual Network in the form of a warning and allow you to bulk provision them for every server in every rail in the blueprint. This streamlines the process of creating and configuring Virtual Networks for Rail-aligned blueprints by eliminating the mental burden of managing a large number of VLAN assignments.
You can now install Onbox system agents on NVIDIA DGX servers, including the following models: A100, H100, and H200. Use the Create On-box agent in Telemetry-Only mode and provide an IP address allowing Out-Of-Band access to the server, and use a user with sudo privileges and passwordless access to the device. Acknowledge the systems after the successful agent install and follow the standard procedure for Serial Number assignment in the blueprint. The following Telemetry services will be enabled on that device:
LLDP,
Hostname,
Interface,
Interface_Counters,
Resource_Util,
Disk_Util,
Gpu_Hardware_Counters
Gpu_Infiniband_Dev_To_Interface. With Hostname and LLDP, you now have validation and expectations for the Server-to-Leaf cabling. You can use the existing "Fetch Discovered LLDP Data" capability to bulk read the configured hostnames and bulk update the blueprint with that information; that simple two-step process will keep the source of truth synched, and so the accuracy of the cabling validations. Other telemetry services are used in different IBA Probes.
You can now have in your AI Blueprints a Per Stripe dashboard auto-created for every new Stripe. This Dashboard leverages data from the predefined and auto-enabled "Stripe and Rail Traffic" probe and provides the following information: Per-Rail aggregated traffic view both live and historical (last 7 days) views as well Rail imbalance detection with anomaly raising where the imbalance threshold is tunable from the predefined probe menu. Additionally, you now have a blueprint wide dashboard labelled "All Stripes Traffic Summary", leveraging the same IBA probe as a source of source and gives a broader view with a break-down of Inter-Stripe traffic vs. Inter-Stripe traffic both live and historical (7 days) on a per Stripe basis. You also can have Stripe level imbalance with tunable threshold.
Now you can collect and analyze Nvidia GPU interface counters to help with congestion and traffic monitoring in real-time.
You can now see in your AI Blueprints a new section in the blueprint dashboard view labelled GPU NIC utilization. This provides a honeycomb style of visualization providing Transmitted and Received network utilization for each GPU's NIC. Various filters are available for drill-down as well as grouping options. You can drill-down on a per Stripe, Rail or GPU Server. You can also group per Rail or per Stripe for more aggregate summary. Selecting any item (cell) provides you with several hyperlink redirects for more topological context.
You can now validate in your AI Blueprints the Ethernet lossless service by leveraging the predefined and auto-enabled "Interface Queue Stats" probe along with its predefined and auto-enabled "Queue Stats Monitoring" Dashboard. The probe collects the following key congestion control metrics: Transmitted ECN, Transmitted and Received PFC, Ingress and Egress Buffer Utilization and Packet drop ratio. The data collection is done on all server-facing interfaces for queues 0, 3, 4 which corresponds respectively to Best-Effort, CNP and NO-LOSS forwarding classes. If you need to monitor additional queues, you can specify it through the "Queue Specific Thresholds" section of the predefined probe menu. Per-Queue anomaly thresholds for ECN, PFC, Buffer occupancy and Packet discard can be adjusted through the same menu. The probe keeps data retention for 14 days.
You can now have in your AI Blueprints monitoring the GPU's NIC ROCE level statistics through the predefined and auto-enabled "GPU Hardware Traffic Monitoring" probe. This probe provides you with visibility over 30 NVIDIA Linux Hardware counters, documented on the NVIDIA website https://enterprise-support.nvidia.com/s/article/understanding-mlx5-linux-counters-and-status-parameters. The probe will raise anomalies for Received CNP, Received and Detected Out-Of-Sequence packets. These metrics help you understand whether your Load Balancing and DCQCN parameters are correctly set or if they need fine-tuning. The probe keeps data retention for 14 days.
This release supports upgrade paths from previous Apstra 5.0.X and 5.1.X releases.
Users must use VM-VM upgrades from Apstra 5.0.X and 5.1.X releases. See the Apstra Installation and Upgrade guide for more information on Apstra upgrades.
The following updates have been made for switch operating systems qualified for the Apstra 6.0.0 release.
Juniper Networks:Junos Evolved for AIDC:23.4x100-D20
Junos (All roles):21.4R322.2R322.4R323.4R2-S4
Junos Evolved for IP-Forwarder role (Spines in EVPN or any role in an IP-Fabric):22.2R3-EVO22.4R3-EVO23.4R2-S4-EVO
Junos Evolved for EVPN leaf roles:22.2R3-EVO22.4R3-EVO23.4R2-S4-EVO
Junos Interconnect Gateway Leaf:22.4R3 (minimum)23.4R2-S4
Junos Evolved Interconnect Gateway Leaf:22.4R3-EVO (minimum)23.4R2-S4-EVO
Cisco Systems:9.3(13)10.2(6)10.3(4a)
Arista Networks:4.24.5M4.28.7.1M4.30.3M
Dell EMC & Edgecore:Enterprise SONiC 4.1.2Enterprise SONiC Edge Standard 4.1.2Enterprise SONiC 4.2.1Enterprise SONiC Edge Standard 4.2.1Enterprise SONiC 4.2.3Enterprise SONiC Edge Standard 4.2.3Enterprise SONiC 4.4.2Enterprise SONiC Edge Standard 4.4.2
You can now monitor CPU, Memory, and Disk usage of your GPU nodes in the predefined "Device System Health IBA" probe. The predefined probe menu has been revised to display differentiated thresholds per system type for switches and servers. The predefined Dashboard widgets have also been expanded to reflect the new collected metrics and alerts.
You can now use TLS when you define a telemetry streaming receiver. You can upload a certificate to the Apstra server (local store). The certificate must be a valid x509 certificate. This enables you to ensure that confidential telemetry information remains protected during transfer, thereby further enhancing your security posture and facilitating compliance with regulatory obligations.
You can now see the AOS-SDK version in the pip show aos-sdk command output.
The Rack-Type Builder method for creating Rack-Type has been discontinued in favor of retaining only the Rack-Type Designer. The Designer has been enhanced with several new functionalities to closely match the existing capabilities of the Builder.
Application point type has been changed for Dynamic BGP prefix peering ("Dynamic BGP Peering" primitive type with any of "IPv4 Subnet for BGP Prefix Dynamic Neighbors" or "IPv6 Subnet for BGP Prefix Dynamic Neighbors" fields populated) in Apstra release 5.1.0. Previously the primitive was applied to SVI or Subinterface, after the fix it is applied to Loopback in the corresponding Routing Zone. As a result, Connectivity Template with "IP Link" and prefix peering "Dynamic BGP Peering" primitives can be stacked in the UI, but can never be applied, because "IP Link" primitive produces no valid application points for "Dynamic BGP Peering" primitive with prefix peering.
When AnomalyGenerator tries writing to MetricDB and taking SysDB snapshots if vlanid is missing from the primary key it will cause the AnomalyGernator to crash wih the below trace information
Unique Index (PrimaryKeyIndex) violation newRow index matches that at row: 65536 python3.10: /Project/leblon/infra/TableTop.tin:206: void Aos::DoubleLinkListHelper::addToList(Aos::RowIndexHelper&, U32): Assertion `false && "!row"' failed. Process 22875 died with signal 6 (SIGABRT) errno 0 code -6 (unknown)
The scenario change-device-password CLI command securely updates device credentials by performing tasks like SSH checks, configlet staging, blueprint commits, and agent password updates. In the current Apstra CLI version, the system agent check has a 60 second timeout, while configlet staging is limited to just 20 seconds. If these operations take longer than expected, the command may fail with errors like:
Failure 1: Task Stage creation of Configlet for password change may fail with: AssertionError: Timeout waiting for Wait that last task status is succeeded Failure 2: Task Check System agent status may fail with: 409 Conflict: Agent is already running a job (check)
These failures occur when backend tasks exceed the current timeout settings, which is particularly noticeable in Apstra 4.2.x and later versions, where performance issues with configlet and configuration rendering are known. Additionally, longer durations in the check job can result from changes in the customer’s environment.
The issue is resolved in Apstra CLI versions 4.2.2, 4.2.2.1, 5.0.0, 5.1.0, and 6.0.0, which increase the timeout thresholds for background tasks such as configlet staging and agent status checks. These updates prevent premature task failures caused by longer processing times in earlier versions.
Apstra customers 4.2.x and greater may encounter an issue where the MTU value for IP Links to Generic Systems cannot be set to 9216 via the UI. Although the UI states that only even values in the range 1280-9216 are accepted, the input of 9216 is incorrectly rejected, while 9214 is accepted.
Important Note: 1. This issue is applicable only to customers upgrading from Apstra 4.1.x to 4.2.x or later, where Fabric MTU remains disabled post-upgrade and customers who wish to continue without enabling the Granular MTU feature. 2. This issue is not applicable to customers with fresh 4.2.x deployments, where Fabric MTU is enabled by default, activating the Granular MTU feature.
Apstra addressed the UI validation bug in 6.0.0 to allow users to set the IP Links to Generic Systems MTU value to 9216 under Fabric Settings.
When Dashboard for blueprint is displayed via clicking Blueprint in the UI, Apstra UI failed to load the blueprint dashboard because the preference information of the dashboard is configured with an empty string value, not the right value for preference.
Apstra Web UI validates the empty string for preference and translates it as the right value.
When attempting to change the keepalive IP address of an MCLAG pair of EOS devices, the configuration deployment of the devices belonging to the pair may fail.
When making use of the built-in jinja function on a custom configlet, {{ function.merge_vlans_to_list(interface_model["allowed_vlans"]) }} fails when used within a configlet. An error will be seen in rendered configuration previews "TypeError: unsupported operand type(s) for -: 'int' and 'str'"
AOS 6.0.0 adds support for both integers and strings within function.merge_vlans_to_list()
When system resources are scarce, Docker (dockerd) may fail to respond within the expected timeframe (60 seconds), resulting in a connection timeout error. This keeps the ContainerLauncherAgent from successfully launching offbox containers. The logs show that high CPU load and late clock events contributed to this failure. As a result, the agent becomes stuck, with multiple "Launch action already pending" warnings appearing in the logs. This prevents certain containers from starting and leaves them in a "absent" state.
2025-02-05 11:45:35,039 INFO aos.cluster.container_launcher:Update task container aos-offbox-10_42_29_205-f status: state='absent', error=Not found 2025-02-05 11:47:06,379 WARNING aos.cluster.container_launcher:Launch action already pending for container 'aos-offbox-10_42_84_192-f' requests.exceptions.ReadTimeout: UnixHTTPConnectionPool(host='localhost', port=None): Read timed out. (read timeout=60)
The ContainerLauncherAgent currently operates on two threads. The first thread identifies containers that need to be scheduled and sends them to the second thread, which interacts with Docker. If a Docker connection timeout occurs, the second thread fails, preventing the agent from completing its tasks. Because there is no automatic retry mechanism in place, the container goes missing.
Customers can workaround this issue by restarting the ContainerLauncherAgent. It will cause the containers relaunched correctly.
Starting with EOS 4.30+, Arista EOS introduces a stricter hardware validation mechanism through a new default system l1 configuration block. This change causes Apstra deployment to fail if speed is configured on ports without compatible transceivers, even if the ports are unused.
system l1 unsupported speed action error unsupported error-correction action error
To maintain compatibility, system agent has been updated to detect this configuration and automatically change the unsupported speed action error to unsupported speed action warning during agent installation. This ensures that systems with unused or unpopulated interfaces do not fail deployment due to this stricter validation. The updated system agent logic safely applies this change only for EOS 4.30.x and later.
Apstra system agent has been updated to detect system l1 configuration and automatically change unsupported speed action error to unsupported speed action warn during agent installation. This ensures that systems with unused or unpopulated interfaces do not fail deployment due to this stricter validation. The updated system agent logic safely applies this change only for EOS 4.30.x and later.
When the device is deployed or drained, Apstra showed a noticeably longer delay in finishing the operation than the Apstra 4.1.X release. The problem was linked to the significantly increased delay in the Jinja configuration rendering area following Apstra's migration from Python version 2 to version 3. Additionally, it affects the rendering configuration for the blueprint's configlet processing.
In high-scale environments, such as 16K GPU AI/ML topologies, several containers occupy CPU resources for much longer periods of time, reducing CPU resource availability and causing heartbeat timeouts.
Recommend increasing CPU power for the Apstra Controller VM by doubling the number of vCPUs (8 -> 16)
Apstra gRPC probing of JUNOS devices is causing large ephemeral database files which may fill the disk and cause issues accessing the device via SSH.
jtac-QFX5120-48Y-8C-r011 mgd[14126]: UI_EPHEMERAL_COMMIT: User 'root' has requested commit on 'junos-analytics' ephemeral databasejtac-QFX5120-48Y-8C-r011 mgd[14126]: UI_EPHEMERAL_COMMIT_COMPLETED: commit complete on 'junos-analytics' ephemeral database
Apstra 6.0.0 introduced a different mechanism using an invalid XPATH query to verify the health of the gRPC service on the device, which doesn't introduce ephemeral DB changes. However, manual ephemeral purge configuration via configlet would still be recommended because the ephemeral DB might reach the full condition.
set system configuration-database ephemeral purge-on-version 30
The JUNOS 23.4R2-S5-EVO on-box agent experiences a 2-minute timeout for the NETCONF operation with a silent XML parsing error message without raising an explicit error when it performs a get-configuration operation to compare against the golden configuration or a lock configuration operation to update any incoming changes via NETCONF. In an attempt to restore after the timeout, the agent starts the reconnect logic; however, this fails because it fails to include the proper binding step.
In 6.0.0, the Onbox agent in the 23.4R2-S5-EVO (not qualified NOS) doesn't yield a silent XML parsing error, not triggering a timeout for NETCONF operations.
The license information from the previous version of the Apstra controller is not transferred to the upgraded controller when it is upgraded. As a result, license information is missed in the upgraded Apstra controller.
License files will be transferred from the old controller to the upgraded controller as part of the upgrade process.
A banner configured in the Cisco NXOS device must be single line. Multiline banner (motd or exec) is not supported.
The Rack Designer does not have Port channel ID Range min and max value fields while adding a generic system to the leaf. This range was added, allowing for more than a single PortChannel to a single Generic System. These fields must follow the following rules when assigning PortChannel IDs to Generics.
1. They must not overlap in scope of single switch.2. The ranges are continuous (cannot have "holes"), e.g. Range1 [0..100] and Range2 [10..12] are considered to be overlapping even if only 1 PCID from the first range is used.
Min and Max Port Channel fields were added to Rack Designer to allow for more than a single PortChannel assignment for a Generic System.
When Junos devices become unreachable, triggering a commit check leads to a rejected state after approximately 600 seconds. However, devices in this rejected commit check state are put into commitCheckinProgress state after an AOS restart or Full Config Push.The problem occurs because Apstra does not treat devices in the commitCheckInProgress state as pending. This leads to a misleading UI/UX experience (some devices may still be in the commitCheckInProgress state, but the main dashboard may show zero pending devices).
Junos devices in the rejected commit check state are correctly reflected as pending in the deployment status of the blueprint dashboard.
When a user collects a controller Show Tech by UI with the include backup option enabled, the generated Show Tech also includes a backup of the aos-edge-auth.json file. This file contains sensitive information such as the registration_key, org_id, and other credentials used to authenticate with JCloud.
If this backup is restored in a different environment, the apstra_edge container in that environment may use these credentials to connect to JCloud. This can unintentionally trigger an Edge re-registration event, disrupting the original environment by recreating its active Edge instance.
In Apstra 6.0.0, Show Tech backups taken from the GUI no longer include sensitive files like aos-edge-auth.json, keeping credential information secure. However, when a backup is taken through the CLI, the authentication file is included, since these backups are meant for full system recovery during failures or outages.
Rebooting the device or restarting FRR in SONiC may cause the FRR running configuration sections (related with route-map) to be rearranged. The rearranging of sections will typically show a configuration deviation even if the running configuration is exactly the same as before.
Apstra uses application weight information for container types (iba, offbox, and apstra_edge) prior to launching a container. The upgrade process for only upgraded environments (5.0.X to 5.1.0) fails to include application weight information for the apstra_edge container type, causing the TaskScheduler agent to crash and restart. When AOS is restarted, the situation may worsen (all offbox and iba agents remain down). Please apply a workaround to resolve the issue. This issue would not exist in a 5.1.0 clean deployment environment.
The fix is to to update the application weight using API call “/api/cluster/application-weight�.. PUT on this
{"iba": 1000,"offbox": 250,"apstra_edge": 500}
After that AOS restart is needed.
Upon navigating from the Stage tab to the Uncommitted tab, the page does not load and continues to display the loading spinner. The UI is continuously polling the API endpoint /api/blueprints//tasks?mode=full, which forces the backend to attach the complete task payload to every response. Since these payloads can contain large data information, the API responses become significantly heavier. Because the UI repeatedly polls this endpoint, it increases backend processing time and causes the page loading issue, particularly when tasks have large payloads.
This polling behavior has been optimized in AOS version 6.0.0 to avoid unnecessary full payload retrieval during regular status checks. Recommend upgrading to Apstra version 6.0.0 to benefit from this performance improvement.
When the Device Profile allows multiple transformations for the same speed and the IM (Interface Map) needs mixed transformation for the same speed across ports (for example, Port 1: 1X100G, Port 2: 4X100G, etc.), it is not possible to create an IM using the UI.
The current default timeout for collecting show tech via UI is 20 minutes (1200 seconds) in the controller_timeout value of the show_tech section of aos configuration file /etc/aos/aos.conf. The collection contains operations for exporting all binary log files from all running agents across containers into printable format files. Depending on the number of accumulated log files, the translation processing time may exceed the default timeout for collecting showtech, resulting in timeout failures for the UI showtech collection task.
/etc/aos/aos.conf
Users will not be able to create a Datacenter Blueprint using the EVPN reference design with IPv6 RFC:5549 as the Spine to Leaf Links Underlay Type, as this configuration is not currently supported. This limitation is tracked in RFE-1364 which is planned for release in 6.1.0. If the user attempts to create an unsupported blueprint using the steps below, the UI will throw a critical error message - EVPN is not allowed with spine-leaf IPv6 links addressing policy.
Steps to Reproduce: 1. Create a template with the overlay control protocol set to MP-EBGP EVPN. 2. Navigate to the Blueprint Creation page. 3. Select Datacenter as the reference design. 4. Set Spine-to-Leaf Links Underlay Type to IPv6 RFC:5549. 5. Attempt to create the blueprint.
However, the blueprint becomes locked and cannot be deleted, even with admin privileges. Deletion can only be performed via REST API Explorer or CLI. This issue was addressed in 6.0.0, ensuring that the delete function works correctly.
This issue will be addressed in 6.0.0, ensuring that the delete function works correctly.
After editing a VN (Virtual Network) configuration, if the same Virtual Network is open, none of the changes appear until the VN is opened again in the UI.
VN changes now properly reflect after the saving and reopening the same VN which changes were made.
When a rack is built with leaf devices and generic systems, group labels for generic systems can have the same value as leaf or access switches' target_switch_label, contrary to the expectation that the group label should not be the same value as target_switch_label inside the rack. Any changes to the rack, such as adding a generic system or deleting an existing generic system, would fail due to the validation error caused by not meeting the above expectation.
When Aptra registers a subscription for a specific xpath into Junos EVO device, Junos EVO device may keep multiple subscriptions for the same path and cause telemetry sequence overruns in telemetry data.
When available, upgrade to Junos EVO 23.4R2-S4 or greater
If a SONiC device removes and then re-adds an IP address, the device may fail to add a static route involving that address to the kernel routing table, even if the static route configuration exists. The show ip route output in vtysh in the SONiC device experiencing this issue may include the following output lines:
show ip route
vtysh
S>r 198.51.100.2/32 [1/0] via 192.168.0.9, Po1.4, weight 1, 01:59:25 B * 198.51.100.2/32 [20/0] via 10.0.0.3, Vlan201 onlink, weight 1, 01:59:25
Above, the "S" static route has been rejected by the kernel and was not installed.
This is an FRR bug in SONiC 4.1.2 and 4.2.1. A fix for this defect is included in SONiC 4.4.x and is verified to work. Please refer to vendor issue SONIC-89523(Next Hop Group fails to install in the kernel).
This issue was discovered in the Juniper QFX 5230/5240 Juniper EVO device for non-EVPN blueprints (Pure IP Fabric) running the 23.4R3-S3-EVO version. When an interface is assigned both untagged and tagged vlans and then the untagged vlan is removed, the traffic from the leaf device will have the incorrect VLAN ID, causing network connectivity issues.
The state_check processor outputs a state column for each series. However, users may not be able to filter series in the output stage using the per series state column when querying stage data. Users attempting to use a filter such as properties.state = 'false, true' may see no results, even if matching data is present. This is a known limitation in how filtering works on series-level data in the state_check processors output. The filtering behavior may not align with user expectations when attempting to use conditions on the state column. Engineering evaluating potential improvements to allow filtering on state column in state processor for a future release.
MetricQueryManagerAgent handles large historical data to serve the '/blueprints//anomalies-history' endpoint. Depending on the amount of data, the agent's memory footprint may increase significantly. The benchmark environment recorded a memory footprint of up to 2.2Gb for the agent. The memory footprint settles after the initial bump.If the system administrator is concerned about the MetricQueryManagerAgent footprint's impact on the system's available memory, the following workaround is recommended.
Restart MetricQueryManagerAgent and avoid using the 'Time Series' query.
In Rack-Type Designer, the ability to specify an Access switch count which was available in Rack Builder is currently not supported. The Rack-Type Designer, introduced in Apstra 4.2.0, replaces the traditional Rack-Type Builder to provide a more intuitive and user-friendly experience. As of Apstra 6.0.0, Rack Builder is deprecated and no longer available.
While Rack Designer does not include all functionalities of the Rack Builder, such as the ability to add multiple access switches and logical connections at once, the new interface offers a superior UX, improved workflow, and new capabilities. Some features may require different steps, while others have been redesigned or omitted for usability improvements.
Users can utilize the Clone functionality in Rack Designer to replicate multiple access switches as an alternative to the count feature.
For further improvement, feature requests may be required via the sales account team.
In the AI Cluster Template workflow, selecting "Manually Selected" under Device Models incorrectly shows "4 selected" even when no models are chosen. The count should reflect actual selections or indicate default behavior clearly.
When instantiating predefined probes such as "VMs Without Fabric Configured VLANs" Apstra UI may fail to display the probe with a type error, making the probe unusable.
To workaround the issue, user need to download the UI hotpatch (https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HD1BNIA1) and apply into the controller VM as follows. Fixes for AOS-55184 and AOS-55286 are included in the UI hotpatch for AOS-55673:
AOS-55184
AOS-55286
AOS-55673
admin@aos-server:~$ sudo su [sudo] password for admin: root@aos-server:/home/admin# gzip -d aos-web-ui-6.0.0-aos55673.run.gz root@aos-server:/home/admin# chmod 755 aos-web-ui-6.0.0-aos55673.run root@aos-server:/home/admin# bash aos-web-ui-6.0.0-aos55673.run Verifying archive integrity... All good. Uncompressing AOS WebUI installer 100% ### Backing up existing AOS WebUI into /opt/aos/frontend/snapshot/2025-07-23_20-31-02 ... Successfully copied 226MB to /opt/aos/frontend/snapshot/2025-07-23_20-31-02 ### Copying AOS WebUI file into aos_controller_1 ... Successfully copied 226MB to aos_controller_1:/opt/aos/frontend_images/ ### Initializing new AOS WebUI ... ### Done! root@aos-server:/home/admin# exit
In Apstra 6.0, the Main Topology View (Blueprint > Staged > Physical > Topology) displays ESI/MLAG pairs with leaf2 positioned on the left and leaf1 on the right, which is the reverse of the layout seen in prior versions. In previous versions, the UI consistently rendered ESI/MLAG pairs with leaf1 on the left and leaf2 on the right, aligning with user expectations. The change in Apstra 6.0 is purely cosmetic and does not affect system behavior or configuration.
To workaround the issue, user need to download the UI hotpatch (https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HCCgNIAX) and apply into the controller VM as follows:
admin@aos-server:~$ sudo su [sudo] password for admin: root@aos-server:/home/admin# gzip -d aos-web-ui-6.0.0-aos55286.run.gz root@aos-server:/home/admin# chmod 755 aos-web-ui-6.0.0-aos55286.run root@aos-server:/home/admin# bash aos-web-ui-6.0.0-aos55286.run Verifying archive integrity... All good. Uncompressing AOS WebUI installer 100% ### Backing up existing AOS WebUI into /opt/aos/frontend/snapshot/2025-07-09_09-59-37 ... Successfully copied 226MB to /opt/aos/frontend/snapshot/2025-07-09_09-59-37 ### Copying AOS WebUI file into aos_controller_1 ... Successfully copied 226MB to aos_controller_1:/opt/aos/frontend_images/ ### Initializing new AOS WebUI ... ### Done! root@aos-server:/home/admin# exit
When setting the reservation mode from the configurator using the different options except None and the reservation mode flag is set in the DHCP configuration, it causes the static IP address to be assigned from the pool rather than the static IP address mapped against the MAC address in the configurator.
User can follow the below steps to get the static IP to the host using ZTP1. If Reservation Mode is set to None, users can configure static IPs by navigating from ZTP UI to DHCPv4 Configurator, setting Reservation Mode to None, and toggling Reservations-Global.2. If Reservation Mode is enabled, users must move the hosts inside Reservations to the subnet level based on the required IP address while keeping the Reservation Mode set to Global.
When monitoring Apstra ZTP device status in the Apstra UI under "ZTP Status" / "Devices", there may be duplicate entries for Junos devices. Apstra ZTP will try to ensure the physical management interface for the Junos device is used instead of any virtual management interface (e.g. "vme" interface). Junos may use the virtual interface when ZTP starts but cannot be added to the required "mgmt_junos" routing-instance. This is done as the first step in ZTP in order to ensure that the management IP address does not change during the rest of the steps involved in ZTP (especially those involving connectivity to Apstra). Enabling a different management interface will cause the DHCP server to give out a new lease. Also, the vendor class identifier for the new management interface is cleared so that the DHCP server does not give out vendor-specific options to this interface, which may re-trigger a new ZTP session while the current session is active. This is expecetd behavior.
The timestamp field of AosMessage in the streaming has used uint64 format for both millisecond and microsecond timestamp information. The current timestamp uin64 field will be changed in release 6.1.0 to the google.protobuf.Timestamp format, which is incompatible with the previous version, in order to provide consistent, accurate timestamp information. The streaming receiver side must be modified to accommodate this incompatible modification.
An event involving BGP flaps from BGP peers configured on IRB/Loopback interfaces for Juniper EVO device has been reported during the commit with incremental changes (deleting an IRB/Loopback interface in the same VRF). The BGP flaps happen when the device's configuration is committed in override mode, which was Apstra's default setting prior to 6.1.0. However, using load update mode did not reveal the issue.
If the EVO device is running as an off-box agent, please add key load_mode with value update into open options in the edit agent menu. Otherwise, recommend upgrading Apstra to 6.1.X, which uses load update as the default mode for commit.
When multiple show-tech jobs are triggered simultaneously, the /var/log partition can reach high utilization, causing the controller to enter read-only mode. This may result in incomplete job status updates, leaving several jobs stuck in in-progress or pending states. In a corner case, this condition can also lead to repeated crashes of SystemAgentManager, preventing automatic recovery even after disk space is reclaimed. Avoid triggering bulk show-tech collection on a large number of devices. Ensure sufficient disk space is available before running show-tech.
If the issue occurs, reclaim space in /var/log and restart AOS services. In most cases, SystemAgentManager recovers automatically, clearing stuck jobs and completing or failing pending ones. If SystemAgentManager crash persist and jobs remain stuck, manual cleanup via Acons is required. It is recommended to contact Apstra Support for assistance.
User feedback was added in Apstra 6.0.0 to enable users to share their Apstra experiences. After the release of Apstra 6.0.0, the link to the feedback form was changed. It results in an error message about a missing form.
To workaround the issue, the user needs to download the UI hotpatch (https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HCCgNIAX) and apply it to the controller VM as below. The hotpatch addresses both the AOS-55184 and AOS-55286 issues.
If the rendered device configuration for the VLAN description contains newline characters populated from the virtual network's description field, the commit operation fails because JUNOS and JUNOS-EVO devices do not support multi-line string for the VLAN description.
Please use one-line formatted string in the description field of Virtual Network rather than multi-line string.
While configlet is being applied to the SONiC device, if the device agent restarts after being disconnected from the controller, the agent executes any remaining changes and collects the running configuration as golden configuration to monitor for configuration anomalies. Because the process of applying configlet changes is still running independently of the agent, it introduces changes into the running configuration even when the golden configuration is collected by the agent. The following changes from the process cause configuration anomalies in the SONiC device.
After reviewing the running configuration on the SONiC device, if all the changes from the configlet are correctly applied, the customer can safely accept changes to avoid further configuration anomalies.
Adding multiple AAA servers in the blueprint through Staged > Catalog > AAA Servers leads to a configuration load error in the JUNOS and EVO device during commit check or commit.
Recommend using configlet instead of using UI (Stage > Catalog > AAA Server) when multiple AAA servers needs to configured.
When attempting to convert leaf switches from ESI to MLAG within the same blueprint, the operation fails because Apstra does not allow mixing ESI and MLAG redundancy models at the rack level. This restriction is enforced starting in Apstra 4.2.0 and is expected behavior. During the operation, users may see the following error in the UI or REST API Explorer:
"Combining MLAG and ESI leaf pairs not supported"
Currently, converting ESI racks to MLAG racks or vice versa requires replacing all racks within a single FE operation using the REST API. However, due to limitation, this conversion can cause BuilderAgent to fail when links are present between external generic system and leaf switches. In this occurs, please revert the changes and follow the steps outlined in workaround section.
To successfully convert the rack type using modify-racks API, the following workaround can be used:
1. Identify the leaf switches connected to the external generic system 2. Identify the Connectivity Templates (CTs) associated with the external generic interfaces and unassign the corresponding application points 3. Delete the links between the leaf switches and the external generic system 4. Undeploy and unassign the leaf devices from the blueprint 5. Unassign interface maps from the blueprint 6. Use the REST API /api/blueprints/{blueprint_id}/modify-racks to convert the racks from ESI to MLAG 7. Import MLAG-compatible interface maps and assign them to the switches 8. Recreate the links to the external generic system 9. Reassign the endpoints to the appropriate Connectivity Templates
For additional guidance, please contact Apstra Technical Support.
While viewing/editing CT (Connectivity Template)s across blueprints, it's possible that the CT may incorrectly display empty field values.
By clicking the browser refresh button, CT would display the correct data.
In version 4.2.0, Apstra introduces the capability for users to forcibly delete a Virtual Network, even if it has active endpoints. Apstra will initially display the interfaces to which the Virtual Network (VN) is currently allocated and prompt the user to confirm the deletion. It's important to note a limitation in the current design: if a user deletes a VN assigned in a CT where Multiple VLANs are present, all active endpoints will be unassigned.
User should manually remove the specific VLAN from the CT before proceeding to delete it from the Staged > Virtual Networks section.
DeviceTelemetryAgent.{pid}.log files in /var/log/aos/ in the offbox agents become large and can fill up the disk
The following Python script can be added to run via crontab on an hourly basis, which will clean up older log files. This workaround needs to be applied to controller VM and worker VMs where offbox agents are running (nodes with offbox tags in the Platform/Apstra Cluster/Nodes).
copy from next line
# Copyright 2024-present, Apstra, Inc. All rights reserved. # # This source code is licensed under End User License Agreement found in the # LICENSE file at http://apstra.com/eula import configparser import json import os import re import shutil import subprocess import traceback SystemIdPattern = re.compile(r'AOS_SYSTEM_ID=offbox,(.+),(.+)') def update_aos_conf(task_id): aos_config = os.path.join( '/var/lib/aos/conf.d/task/offbox/', task_id, 'aos.conf', ) parser = configparser.ConfigParser() if os.path.isfile(aos_config): parser.read(aos_config) if not parser.has_section('logrotate'): parser.add_section('logrotate') if 'max_kept_backups' not in parser.options('logrotate'): parser.set('logrotate', 'max_kept_backups', '1') staging_file = aos_config + '.staging' with open(staging_file, 'w') as f: parser.write(f) shutil.move(staging_file, aos_config) return True return False def refresh_logging_infra(container_id): subprocess.check_output([ 'docker', 'exec', container_id, 'pkill', '-HUP', 'DeviceKeeperAge' ]) def get_offbox_containers(): containers = subprocess.check_output([ 'docker', 'ps', '-q', '--filter', 'label=AOS_CLUSTER_APPLICATION=offbox', ]).decode() return containers.splitlines() def get_task_ids(): def extract_info(container_env): try: envs = json.loads(container_env) except ValueError: return None, None for env in envs: matched = SystemIdPattern.match(env) if matched: return matched.group(1), matched.group(2) return None, None containers = get_offbox_containers() if not containers: return containers_env = subprocess.check_output([ 'docker', 'inspect', '--format', '{{json .Config.Env}}', *containers, ]).decode() for line in containers_env.splitlines(): task_id, container_id = extract_info(line) if not task_id: print('Failed to extract task id from container env: {}'.format(line)) continue yield task_id, container_id def main(): for task_id, container_id in get_task_ids(): try: if update_aos_conf(task_id): refresh_logging_infra(container_id) except: print('Failed to update aos.conf for container: {}'.format(container_id)) traceback.print_exc() main()
script finishes here. don't copy this line
Even if the above workaround is applied, there is a chance of filling up partition. The below command can be executd with root permission to clean up logs quickly in the controller VM and worker VMs.
find /var/log/aos/task -name "*.log" -size +10M -print | grep "DeviceTelemetry" | xargs -I {} sudo cp /dev/null {}
ECN marked packet information is not currently visible in the dashboard widget, although it is correctly displayed in the staged view in the probe.
Depending on the data, we can apply sorting to the keys to guarantee consistent ordering of the series and modify the predefined dashboard parameters through the Edit menu in the user interface (UI) so that the dashboard displays N rows. But this method only lets us work with a portion of the original series. Modifications must be made on the Metric DB side in order to correctly represent the Top N series across all data. A partial solution is offered by the UI-based workaround until those changes are made, but it does not ensure that we are showing the actual top N series; rather, it only shows a filtered and sorted subset according to the current stage.
Execute CLI commands in the Juniper device supported only show and request chassis beacon commands in the Apstra < 6.1.0 environment. Additional commands (ping and traceroute) are introduced in Execute CLI commands in the Apstra >= 6.1.0 environment for easier troubleshooting environments.
The current Radius client in the Apstra Controller adheres to standard RFC 2865 functionality, inherently not FIPS-compliant because it relies on weak, non-compliant algorithms (MD5/MD4) for password hashing and authentication.
In Rack-Type Designer, the ability to specify a Generic System (GS) count, which was available in Rack Builder, is currently not supported. The Rack-Type Designer, introduced in Apstra 4.2.0, replaces the traditional Rack-Type Builder to offer a more intuitive and user-friendly experience. As of Apstra 6.0.0, Rack Builder is deprecated and no longer available.
While Rack Designer does not include all the functionalities of Rack Builder, such as the ability to add multiple generic systems and logical links at once, the new interface provides a superior UX, improved workflow, and new capabilities. Some features may require additional steps, while others have been redesigned or omitted to improve usability.
Users can utilize the Clone functionality in Rack Designer to replicate multiple generic systems as an alternative to the GS count feature.
Because the JUNOS device might not send the full snapshot of interface-related data per reporting interval (default interval = 120 seconds) after initial synchronization, Device Telemetry Health detects continuous anomalies for gRPC Periodic Response Timeout in the interface telemetry service in a high-scale environment. The problem still exists even if the reporting interval is extended.
It is advised to disable gRPC globally, which forces all gRPC-related telemetry services (interface, MAC) to switch to polling mode rather than gRPC, since the problem may occur at random on several devices.To disable the gRPC service globally, change grpc_enabled = 0 in the /etc/aos/aos.conf file and then restart AOS service in the Apstra Controller.
[telemetry_global_config] # Python multithreading enable/disable knob for telemetry collection multithreading_config = 1 # Execution timeout for extensible telemetry collectors command_timeout = 120 # Knob to enable/disable gRPC based service collectors grpc_enabled = 0
Because the MAC telemetry service uses the gRPCOnChange mode, the device only sends updates after initial synchronization. When the JUNOS device subscribes to the PATH (/network-instances/network-instance/mac-table/entries/entry) for MAC telemetry service, Apstra's gRPC client (Apstra) receives the first full data from two processes (l2ald, l2aldTM). During the initial synchronization, these processes use their own sequence number range (duplicate range), which makes the gRPC client think that the gRPC packets may be dropped internally. Granular sequence number handling will be introduced in 6.1.0 to address the existing sequence overrun issue.
The default collection period for IBA interface flapping is only 60 seconds and the flapping threshold is 5 times. But, the default collection period for interface service is 2 minutes, the interface will receive updates every 2 minutes only.
So, the default anomaly window is too small to capture 5 interface flaps. For 5 flaps it should be at least 10 minutes.
When a configlet action (import/delete) occurs in a blueprint with a large number of configlets, it can fail with a timeout, and the BlueprintDiffProducerAgent process can crash due to a heartbeat timeout. The problem occurs when the blueprint has a configlet with an incorrect Jinja expression via configlet import/delete actions. Failure of rendering configuration with incorrect Jina expressions can cause all configlets to be re-evaluated for all eligible devices, potentially resulting in much longer configlet processing.
The below steps can be applied as a workaround. If further assistance is needed, please contact the Apstra support team.
1. In order to recover from the crash of BlueprintDiffProducerAgent, please increase the heartbeat_period to 1200 secs in agent_management section of aos.conf file (<=6.1.0: /etc/aos/aos.conf, >=6.1.1: /user/root/etc/aos/aos.conf) and restart AOS service (service aos restart).
[agent_management] # Override the default heartbeat timeout for agents spawned dynamically by # AgentManager. The value must be a non-negative number. The unit is seconds. # The value 0 is used to turn off heartbeat-based agent timeouts and restarts. # The minimum non-0 value allowed is 60. If not provided, then the default # timeout value (600 seconds) is used. heartbeat_period = 1200
2. Please check each configlet in the blueprint and delete the configlet with incorrect Jinja expression from the blueprint.
Using the system parameters (go to the staging tab of the freeform blueprint, choose the target switch, enter the edit mode for the S/N field, then click the Reset value button) will not allow the Apstra UI to unassign a serial number for a system. The failure of "deploy_mode" will result in the message "Value is required." because the Apstra backend expects the deploy_mode value to be one of the values ("deploy", "ready", "drain", or "undeploy"), but the Apstra UI sends the deploy_mode value as null when the serial number is unassigned.
Go to the staged -> Physical -> Systems, select the target switch, and then click the 'Change System IDs assignments' icon. When the Assignment System dialog pops up, click the 'remove assignment' icon and check Deploy Mode as undeploy.
interface 25g speed configuration was not properly rendered over ports 4-27 when 25g transformation is used in the Juniper ACX7024 device profile
Add a 25g transformation for ports 4-27 to include interface speed 25g setting.
Junos RFC5549 BGP peer sessions were rendering the same route-map twice on import/export statements. This only applies to 'ipv6-only' bgp peers.
Configurations were rendering:
neighbor a05:fab:192:168:50::254 { description "facing_leaf1-generic"; local-address a05:fab:192:168:50::1; peer-as 65510; family inet { unicast { extended-nexthop; } } family inet6 { unicast; } import ( RoutesFromExt-default-Default_immutable && RoutesFromExt-default-Default_immutable ); export ( RoutesToExt-default-Default_immutable && RoutesToExt-default-Default_immutable ); }
This has been addressed in 6.1.0, where the import and export statements will contain that route-map entry only once.
This may result in a service disruption on upgrade for those BGP peers as Junos generically may reset the peer when it detects any import/export policy reference change.
This will also apply to ipv6-only, non-EVPN blueprints for route-maps such as "LEAF_TO_SPINE_FABRIC_OUT" between all superspine/spine/leaves.
When rack types are exported, modified, and then re-imported into a blueprint, generic systems in the rack may be unexpectedly renamed, resulting in some servers losing their original user-defined labels even though only rack parameters were changed.
Customers should avoid editing the rack type in the Global Catalog UI when they need to preserve generic system names.(1) Export the rack type from the blueprint to the Global Catalog.(2) use the API PUT /api/design/rack-types/{rack_type_id} to update only the link_per_spine_speed (e.g., from 25 to 100) directly in the rack type JSON without changing the generic system group_label or count structure.(3) re-import the updated rack type into the blueprint.
The Last Modified field of the interface telemetry service, which uses gRPC periodic mode, is inadvertently updated according to the interface telemetry service's default interval even if no data has been collected from the device. This problem is noticed when gRPC periodic mode is used for Interface telemetry service. The Last Modified field should be updated if a status change is observed, and the Last Fetched field should be updated if the device reports telemetry data.
After configuring the GCM ciphers [email protected] and [email protected] on the Junos device, the check job began failing because these ciphers are not supported in Apstra 6.0.0. As a result, the cipher exchange between the client and server couldn't be handshaken between the device and Apstra, leading to an SSH negotiation failure leading to connection failure.
Apstra 6.1.1 supports GCM ciphers. The user needs to upgrade to Apstra version 6.1.1 to use GCM ciphers between JUNOS devices and Apstra.
In ESI-based EVPN deployments, the MAC Monitor probe may incorrectly report MAC addresses as missing even though they are present on the device. This can cause Virtual Networks Containing Systems With Missing MAC Addresses anomalies to remain active even after the underlying network issue has been resolved or maintenance has been completed.
Edit the MAC Monitor probe and disable the Raise Anomaly option. If required, disable and then re-enable the MAC Monitor probe to clear the existing anomaly state. Keep Raise Anomaly disabled until upgrading to Apstra 6.2.0, as the issue may recur before the fix is applied.
The Router MAC (also known as Master Bridge MAC or bridge MAC) on a SONiC device is a unique MAC address assigned to the switch and used as the source MAC address for packets originating from the switch itself, such as those generated by the virtual router or bridge interfaces. The current Mac Monitor probe does not report Router MAC on the originating switch (for example, if leaf1 has MasterBridgeMac as 52:54:00:4e:c6:60, this MAC will show up as a missing MAC on leaf1 for all VNs).This is a bug related to analytics only; it does not have any network operational impact.
Upgrade script in 4.2.2 and following releases introduced a regression that manifests itself in multi-step migration scenarios.Step-1. Upgrade from 4.2.1 -> X (eg. 4.2.2 ) - causes the permission on '/var/lib/aos/metricdb/iba' folder and it's sub-directories to be 700.Step-2. Upgrade from X (eg. 4.2.2) -> Y (eg. 5.0.1) - causes the silent failure in step that copies '/var/lib/aos/metricdb' folder and ALL its sub-directories to new VM.
The impact of this failure is that the 'Audit', 'IBA stage history' and 'Aos cluster health history' data is lost in the final upgraded AOS instance. The data from the previous release will be lost subsequently if there are further migration steps involved.This issue affects all the releases starting upgrade from 4.2.2. If your upgrade source Apstra is at least 4.2.2, please apply the workaround suggested BEFORE performing the upgrade.
Apply workaround fix(aos_54413_fix_metricdb_permissions.run: https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000Gc0kCIAR) to the old version Apstra Controller Node *BEFORE* every upgrade, following the below steps.
1. Copy the bundle aos_54413_fix_metricdb_permissions.run to the source (old) Apstra controller node. The tool expects Apstra service to be running because it needs to get cluster node information from Sysdb.
2. Make it as executable and execute the bundle as sudo
admin@aos-server:~$ chmod 755 ./aos_54413_fix_metricdb_permissions.run admin@aos-server:~$ sudo ./aos_54413_fix_metricdb_permissions.run Verifying archive integrity... All good. Uncompressing Fix for AOS-54413 for AOS >= 4.2.2 100% AOS[2025-05-25_19:36:22]: Fixing controller node AOS[2025-05-25_19:36:23]: Getting cluster node metadata AOS[2025-05-25_19:36:24]: Fixing worker node: 10.28.75.6 Logs have been collected at: /home/admin/aos_54413_fix_logs_20250525_193623.tar.gz
3. The absence of any errors means that the issue has been fixed. In case of errors during execution, please reach out Juniper Apstra Support Team.
If AOS instances upgraded without work-around and if old apstra VM is preserved, Contact Juniper Apstra support team to help with migrating MetricDB data.
On ACX7100-32C devices, configuring a port with 10GE speed (non-channelized) could result in the port failing to link up, accompanied by "Invalid Port Speed Configuration" and "Optics does not support configured speed" alarms. This occurred because the built-in ACX7100-32C device profile did not automatically generate the required "unused" configuration for the adjacent port within the same port group(for example 0 and 1 are in the same group for 10GE)
Clone the existing built-in ACX7100-32C Device Profile and update Transformation #7 (10GE non-channelized) to include an unused_port_list configuration for the adjacent interface in the same port group. If you need further assistance, please reach out HPE Apstra Support Team.
Steps:1. Clone the ACX7100-32C Device Profile to create a custom copy.2. In the cloned profile, edit port setting for Transformation #7 (10GE non-channelized port) to add unused_interface_list entries for the other port in the same port group (e.g., port 0 and port 1 share a group, port 2 and port 3 share a group, etc.).
Example) Port 0 setting
Old setting:
{"global": {"breakout": false, "fpc": 0, "pic": 0, "port": 0, "speed": ""}, "interface": {"speed": "10g"}, "validations": [{"constraint": "1x25or1x10", "port_group": "P0_P1"}, {"constraint": "no_constraint", "port_group": "P0_P1_P2_P3"}]}
New setting: add "unused_interfaces_list" key with value ["et-0/0/1"] for adjacent port.
{"global": {"breakout": false, "fpc": 0, "pic": 0, "port": 0, "speed": ""}, "interface": {"speed": "10g","unused_interfaces_list": ["et-0/0/1"]}, "validations": [{"constraint": "1x25or1x10", "port_group": "P0_P1"}, {"constraint": "no_constraint", "port_group": "P0_P1_P2_P3"}]}
3. Create new Interface Map with the updated Device Profile.
4. Assign the new cloned device profile into the managed devices and import the new IM into blueprint5. Assign the new IM into the deployed devices
The NOS upgrade procedure must parse the pristine configuration in order to determine which ports must be disabled when the Skip Shutting Down Interface During Upgrade option is not checked in the Advanced Settings of Managed Device. The NOS upgrade would fail if the configuration line in the pristine configuration extended into multiple lines because parsing the command line misses the end-of-command-line character (.
Recommend re-onboarding of the device with a clean, pristine configuration.
Apstra anticipates that the NOS upgrade will result in a version change after the NOS device image installation and device reboot with the new image. However, JSU (Junos Selective Update) installs only selected packages without changing the version, followed by a restart process rather than a device reboot. Therefore, NOS upgrade by Apstra using JSU image will fail with the post-validation check (pre-install version != post-install version and image filename must include version information).
Please use the CLI to upgrade JSU instead of utilizing Apstra's NOS upgrade, or run Apstra NOS upgrade (make sure the JSU image filename includes version information) and then execute a full push configuration when the NOS upgrade fails due to post-validation check.
When the Optical Transceivers Probe is turned on, the optical_xcvr telemetry service is enabled for all systems with on-box agents (including GPU servers) or off-box agents. Because the optical_xcvr telemetry service is not designed to run on GPU servers, it fails to collect optical transceiver information and returns an error message.
The issue can be resolved by correcting the graph query of the Optical Xcvr Stats process (adding system_type='switch') in the Optical Transceivers Probe to prevent the telemetry service from running on the GPU servers. Please modify the graph query as shown below.
node("device_profile", name="device_profile") .in_("device_profile") .node("interface_map") .in_("interface_map") .node("system", system_id=not_none(),system_type='switch' , deploy_mode=is_in(["deploy", "drain"]), name="system")
When the Junos EVPN Next-hop and Interface count maximums parameter in the staged->Fabric settings->Fabric-policy is enabled, Apstra introduced modifying the default hardware settings for VXLAN routing's resource (next-hop and interfaces) for QFX5110, QFX5120, EX4650, and EX4400 devices in the rendered configuration () starting with version 4.2.0. Whenever configuration changes in VXLAN routing's resource, JUNOS triggers PFE automatic restarts to reflect new changes with service impact. The typical scenarios would be when the device becomes deployed, undeployed, or the device is in NOS upgrade. To prevent unnecessary PFE restarts in those scenarios, the configuration for VXLAN routing's resource needs to be included in the pristine configuration.
If the Junos EVPN Next-hop and Interface count maximums parameter in the staged->Fabric settings->Fabric-policy is enabled, add the below configuration into the device's pristine configuration.QFX5120 and EX4650 VXLAN routing's resource
forwarding-options { vxlan-routing { next-hop 45056; interface-num 8192; overlay-ecmp; } }
QFX5110 VXLAN routing's resource
forwarding-options { vxlan-routing { next-hop 32768; interface-num 8192; overlay-ecmp; } }
EX4400 VXLAN routing's resource (add overlay-ecmp if Junos EX-Series Overlay ECMP is also enabled)
forwarding-options { vxlan-routing { next-hop 16384; interface-num 6144; overlay-ecmp; } }
Customers may encounter the following Server-side Validation Error in the Web UI when the pristine configuration contains multiple system stanzas which is not a expected behvavior:
"Cannot parse config: system already parsed."
According to ScotchInventoryAgent logs, POST requests to update the pristine configuration failed with a 422 Unprocessable Entity error, indicating a validation issue:
2025-02-17 23:31:18,730 680:INFO:aos.scotch.libs.scotch_flask:request: POST /api/systems/AN10555621/pristine-config HTTP/1.0 34074 bytes 2025-02-17 23:31:18,737 680:INFO:aos.scotch.libs.scotch_flask:response: 422 55 bytes 0.007347 seconds
Background of the issue:
1. In Apstra 4.2.x, gRPC was introduced to support Telemetry Streaming, and as a result, having two system blocks in the pristine configuration was expected in that release. 2. Starting from Apstra 5.0.0, enhancements were made to automatically merge multiple system stanzas in the pristine configuration during the NOS upgrade process. 3. If a customer chooses to remain on their current NOS version for an extended period, multiple system stanzas can exist in the pristine configuration without causing issues.
If a customer chooses to remain on their current NOS version for an extended period and needs to forcefully update the pristine configuration, they should manually merge the system stanzas within the pristine configuration using the UI and then perform a Force Update.
For further assistance, please contact Juniper Apstra Support.
Receivers may report a significant number of errors in their statistics due to the possibility of high streaming data being dropped in high-scale environments, such as 8K/16K GPU AI/ML topologies.
There is a configuration parameter in the /etc/aos/aos.conf file to control the number of messages for streaming that can be queued before sending them to the TCP socket. Please adjust the value for tcp_queue_size in the streaming section in the /etc/aos/aos.conf file to mitigate the issue and then restart aos service by executing sudo service aos restart.
[streaming] tcp_queue_size = 40000. # default value is 20000.
When a port group in the vCenter is configured with a private vlan mode, the VLAN specification contains the pvlanid property instead of the vlanid property. However, the vlanid property is always expected from the port group's vlan specification by the Device Telemetry Agent's collector if it is not trunk mode. Anomalies could be reported if the collector's execution fails due to a reference to the vlanid property, which is nonexistent.
Recommend not using port group with private vlan in the Virtual Distributed Switch.
Sutatined Optical Threshold anomalies are frequently observed over disconnected (not connected) interfaces in the JUNOS device. The main reason for the problem is that the JUNOS device sometimes reports the received average power as a very low value or - Inf value when the interface is disconnected, and Apstra does not correctly parse the - Inf value. When a very low power value falls below the warning level, Apstra creates an anomaly for the low receive power port. But when the same interface reports a - Inf value that is later incorrectly parsed, Apstra eliminates the interface from the list of interfaces with an optical transceiver, thereby resolving the raised anomaly falsely. This is the reason anomalies are frequently raised and cleared over the same interface. No workaround is available for this issue
None
Integrated DCI feature(vxlan stitching) was introduced in Apstra 4.2.0. In versions 4.2.x and later, customers using this feature may encounter an issue where all VTEP loopback addresses from the fabric including those from non-border leaf devices are being advertised to external routers over BGP in the default routing zone.
This affects only VXLAN DCI Stitching deployments(Stitching requirement for VTEP loopbacks for only border leaf nodes vs OTT requirement for all VTEP loopbacks in the fabric). Even when customers configure routing policies to export only loopback of border leaf nodes, Apstra backend logic automatically includes all VTEP loopbacks. Due to the current design, Apstra does not differentiate between border and non-border leaf roles in this context, resulting in the unintended advertisement of all fabric loopbacks to external peers.
Engineering has confirmed this as a bug. The expected behavior is to advertise only the loopback addresses of border leaf switches to external routers in the default routing zone. There is no official workaround to modify this behavior through standard configuration. Engineering is actively working on a fix to address this issue in a future release. The only option is to use a custom configlet to override Apstra default export logic. Please reach out to Apstra Technical Support for assistance.
User-defined import / export route-targets with ":0" are rejected with validation errors.
{ "rt_policy": { "import_RTs": { "0": "Type 0 RD X:Y must be in format 2-byte ASN:4-byte value. Provided value: \"65500:0\"" } }
When the virtual infra manager is removed from the Apstra controller, Apstra should have cleared any data related to the virtual infra manager. Because it's not cleared, when the same virtual infra manager is added back to Apstra later, old data is still used together with the new collected data from the virtual infra manager's collector. In some scenarios, when old, uncleaned data has an error condition, it can trigger continuous error even if newly collected data doesn't have an error condition.
If the virtual infrastructure manager requires re-onboarding (removing and then adding back) from Apstra, the user must take the actions listed below.1. Remove virtual infra manager from Apstra Controller (External Systems/Virtual Infra Managers).2. Restart the AOS service.3. Add the virtual infra manager back to to the Apstra
Since the UI misses polling of the node detail information to the Apstra backend, the Virtual Network Endpoints view of Generic System Node (Staged > Physical > Topology > Virtual Networks Endpoints) shows empty information.
Refresh Web page in the browser to make the UI send requests explicitly to collect data
Apstra creates a relationship between the PNIC and the Link Discovery Policy (which determines which discovery protocol is used) configured in the VDS when a PNIC is assigned to a VDS (Virtual Distributed Switch). One PNIC may inadvertently become linked to two relationships without clearing out the previous relationship when a user moves a PNIC directly from one VDS to another VDS. An error message below appears when the PNIC becomes unassigned from VDS because there is more than one relationship between the PNIC and Link Discovery Policy that is invalid.
virtual_infra failed to collect data, plugin raised exception: {'item_iter': <aos.sdk.graph.graph.RelationshipIterator object at 0x7f60c03f9570>, 'items': [df57a2f0-969c-4dee-9831-a53526bd7d5a-[:policy]->4c8c4c31-df9a-4188-933f-6b5d67703a1f, df57a2f0-969c-4dee-9831-a53526bd7d5a-[:policy]->1787d662-4a41-4b8b-9b77-acac5508e771]}
Instead of performing one direct migration action from one VDS to another, the problem can be avoided by two actions: unassigning the PNIC from the old VDS and then assigning it to the new VDS.
Procedures for fixing the errors as a workaround1. Remove the virtual infra manager from not only the blueprint but also the External Systems/Virtual Infra Managers.2. Restart the AOS service.3. Add the virtual infra manager back to the External Systems/Virtual Infra Managers and then blueprint.
When multiple vNICs from a single VM are assigned to the same vNET (port group or VDS) in the Virtual Infra, the VMs Without Fabric Configured VLANs probe raises anomalies in the Analytics->Anomalies because the graph query in the VMs backed by Fabric VLANs processor treats those vNICs as identical.
Please clone existing VMs Without Fabric Configured VLANs probe with a different name, and then modify graph query in the VMs backed by Fabric VLANs processor to include vnic into the existing distinct statement like the below. If further assistance is needed, please contact Juniper Apstra Support Team.
Graph Query: match( node('system', name='server', role='generic', management_level='unmanaged', external=False) .out('hosted_interfaces') .node('interface', name='server_intf') .out('hosted_vn_endpoints') .node('vn_endpoint', name='vn_endpoint') .in_('member_endpoints') .node('virtual_network') .out('instantiated_by') .node('vn_instance', name='vn_instance') .having( node(name='vn_instance') .in_('hosted_vn_instances') .node('system', system_id=not_none(), deploy_mode='deploy') .out('hosted_interfaces') .node('interface') .out('link') .node('link') .in_('link') .node('interface') .in_('hosted_interfaces') .node(name='server'), at_least=1 ), node(name='server') .in_('is_realized_by') .node('hypervisor', name='hv'), node(name='hv') .out('hosts') .node('vm', name='vm') .out('has') .node('vnic', name='vnic') .out('part_of') .node('vnet', vn_type='vlan', name='vnet') ) .distinct(['server', 'vn_endpoint', 'vm', 'vnic']) .where(lambda vnet, vn_instance, vn_endpoint: (vnet.vlans == [0] and vn_endpoint.tag_type == 'untagged' or vn_instance.vlan_id in vnet.vlans))
Because the Jinja comment is not recognized as a valid Jinja expression during rendering, it is submitted as normal commands to the device, causing the JUNOS/EVO device to reject the invalid commands and deployment to fail.
To make the template "Jinja-aware", the user needs to include a no-op Jinja construct at the top of the configlet. This will bypass the stringent line-by-line set/delete checks. The workaround entails introducing a dummy control block or expression directly after the Jinja comments.
configlet example with error
{# v1.0 - 22 Jan 2026 - Author: ... - Initial version #} {# Objective: Example of non-working configlet #} set system time-zone Europe/Luxembourg
configlet example with workaround
{# v1.0 - 22 Jan 2026 - Author: ... - Initial version #} {# Objective: Example of working configlet #} {% if 1 > 0 %}{% endif %} set system time-zone Europe/Luxembourg
With current cluster design, worker nodes may fail to recover their configuration state after a temporary loss of SSH connectivity to the controller. This condition can be observed in the UI by navigating to Platform > Apstra Cluster > Nodes > Worker, where the following error may be displayed:
Configuration Error: ssh: connect to host 10.28.17.4 port 22: Connection refused
When the controller (ClusterManagerAgent) attempts to push configuration to worker nodes, it retries SSH connections up to three times. If all attempts fail, due to connection refused or no route to host, the node is marked with a FAILED Configuration State. Once this state is set, the system does not automatically retry configuration, even if SSH connectivity is later restored. Although worker nodes may resume sending keepalives and transition back to an active operational state, the configuration state remains in failed state indefinitely. Due to this the overall node state may continue to appear FAILED despite restored connectivity.
Manually trigger a configuration synchronization using one of the following methods:
1. Navigate to Platform > Developers > REST API Explorer 2. Execute REST API: POST /api/cluster/worker/sync [OR] 1. Restart AOS from the controller VM: systemctl restart aos
When sFlow collector is configured with mgmt_instance, the configuration will be ignored in the ACX platform with a warning such as the below example.sflow {polling-interval 10;sample-rate {ingress 10000;egress 10000;}source-ip 10.217.6.15;collector 10.217.0.165 {udp-port 6343;#### Warning: statement ignored: unsupported platform (ACX7024X)##routing-instance mgmt_junos;
##
## Warning: statement ignored: unsupported platform (ACX7024X)
Please use a non-management instance for exporting sFlow until the ACX platform supports a management instance for sFlow export.
Anomalies are raised due to mismatch in the operational status of interfaces due to interface status showing "unknown" on Juniper EX4400-48T devices running Junos 22.4R3.
Restart the Apstra AOS service to collect the right interface status information. This issue is not observed in higher JUNOS versions. Recommend upgrading to an Apstra-qualified higher JUNOS version (>=23.4R2-S4).
When the Apstra commit was executed, the JUNOS device with the off-box agent reported BFD underlay flaps between the committed device and the other devices. Apstra currently uses the load override option as the default action for device commits, which may result in high CPU utilization, preventing time-sensitive daemons from acquiring CPU time slices and triggering unexpected events such as BFD timeouts or writing EEPROM errors. Starting with 6.1.0, Apstra intends to use load update as the default commit action for JUNOS and EVO devices.
The workaround to use load update can be applied only to the JUNOS/EVO offbox agent. Please follow the below step to apply workaround1. Navigate into Devices->Managed Devices->{DEVICE_IP}-> Agent2. Click Edit button to edit Agent3. Add an option into Open Options with the key as load_mode and the value as update in the Edit Offbox System Agent(s) window.4. Click Update button
The NXOS rollback feature on the N9K-C93600CD-GX device has significant limitations when the devices' interfaces are broken out.Ports 1-24 in the model are organized into four-port groups: (1, 2, 3, 4), (5, 6, 7, 8), (9, 10, 11, 12), (13, 14, 15, 16), (17, 18, 19, 20), and (21, 22, 23, 24). When port 1 is broken out as 4x10G or 4x25G, port 3 is automatically broken out in the same mode, and vice versa. When any port in the quadruple is split into 2x50G, all four ports are automatically split in the same mode. Similarly, ports 26-28 are organized in pairs of two, i.e. (25, 26) and (27, 28). Both ports in the pair must operate in the same breakout mode.
In most cases where a breakout (or more than one) exists, rollback fails to generate a working rollback patch. The reason for this is that the breakouts cannot be reversed if the remaining broken-out interfaces in the same port group have not been shutdown first. For example, to negate the breakout of port 1, the broken-out interfaces of port 3 must be shutdown, and vice versa. It appears that the rollback logic shuts down the interfaces associated with the port whose breakout is being reverted (port 1 in the previous example), but fails to shut down other broken-out ports in the same port group (port 3).
The safer way for the N9K-C93600CD-GX to be used with AOS is for the customer to avoid using breakouts altogether on the device.No issue with rollback when ports 29-36 have been broken out has been observed. Breakouts on these ports can be rolled back In the case that the last interface of a port-group is the only one used and broken out, would the nxos rollback feature (and rollback to pristine) be successful. However this is highly discouragedIn any other case the only way to reverting to pristine would be to manually shudtown all broken down interfaces before reverting to pristine (or using the rollback to a pristine config)
Apstra may encounter gRPC sequence number overrun for MAC telemetry service, recognized as losing data and initiates re-subscription for service. This sequence can continue repeatedly in a highly loaded environment (>=100K entries), making the JSD to handle continuous subscription and cancel subscription requests with memory leaking. This continous accumulation of memory leaking may lead into process crash by OOM and then triggering device reboot.
The workaround is to disable gRPC in the Apstra. If the customer wants to continue to use gRPC in the Apstra, recommed upgrading to the latest 6.1.X release (which includes fixes for the sequence overrun misleading issue) so that meory leaking can be prevented. Please reach out Juniper Apstra Support team for the further assistance.
gRPC server reset count anomalies are observed in the JUNOS-EVO platform when gRPC Max Client connection limit error occurs in the device due to the problem that gRPC stalled connections are not cleared. gRPC keepalive is not enabled by default on the JUNOS-EVO platform running 22.2R3 or 22.4R3, which is the cause of the problem. gRPC keepalive is enabled for 300 seconds in the >=23.4R2-EVO release to avoid a build-up of stalled gRPC connections.
gRPC Max Client connection limit error
>=23.4R2-EVO release
In JUNOS-EVO device running 22.2R3 or 22.4R3, apply the below configuration via configlet into the device to enable gRPC keepalive or upgrade the device to >=23.4R2-EVO. For further assistance, please contact the Juniper Apstra Support Team.
>=23.4R2-EVO
set system services extension-service request-response grpc grpc-keep-alive 300
Apstra 6.0.0 introduced a new IBA probe, Interface Queue Stats, which provides detailed insights into ingress and egress buffer utilization on a per-queue basis, along with other relevant metrics. This IBA probe is designed for AI/ML-based fabrics, particularly those using rail-based blueprints. However, customers using Junos EVO versions earlier than 23.4R2.X100-D31 will notice that the ingress buffer utilization is reported as 0. This issue arises from a bug in EVO devices, where ingress buffer utilization is not exported by default. The bug affecting ingress buffer utilization is resolved in the 23.4R2.X100-D31 release of Junos EVO.
To enable correct reporting of ingress buffer utilization, customers need to create a configlet with the below configuration for each line card slot used in their fabric.
set chassis fpc <fpc-id> traffic-manager buffer-monitor-enable
When Juniper EVO device hosts DHCP servers in a border leaf role with DHCP relay configuration, DHCP may not work as intended due to an unresolved bug in Junos EVO which prevents DHCP packets from being processed correctly. Please refer to the following KB for dhcp relay limitations: https://supportportal.juniper.net/s/article/Juniper-Apstra-Support-for-Stateless-DHCP-Relay?language=en_US
Using Apstra's configlet feature, create configlet to remove the rendered DHCP relay configurations and apply it to the Juniper EVO border leaf device.
A new default limit for the number of IP addresses per MAC per bridge domain in EVPN (mac-ip-limit) was added in Junos and Junos Evolved Releases 24.2R1, 23.4R2, and 23.2R2. 200 IPs per MAC is the default setting.
Clients who use EVPN fabrics, such as those with MAC-VRF deployments in fabrics managed by Apstra, might observe that MACs linked to more than 200 IP addresses cease to learn new IPs in the EVPN MAC-IP table. The Junos software release introduced this expected behavior.
The fabric can support more IPs per MAC while preserving per-bridge-domain enforcement by setting mac-ip-limit globally using an Apstra Configlet. No software fix is required.
To adjust the limit in an Apstra-managed fabric, create a Configlet in Apstra with the below command, specifying the desired limit:
set protocols evpn mac-ip-limit <desired-value>
Import the Configlet into the blueprint and apply it to the relevant switches.
Note: Although the command is global in Junos, the limit is enforced per MAC per bridge domain, including inside MAC-VRFs.
In the Apsta 4.2 reference design change for MAC-VRF, the Junos "forwarding-options evpn-vxlan shared-tunnels" configuration is added via the Apstra rendered configuration. However, this command requires a device reboot to take effect with the Junos warning "Config: forwarding-options evpn-vxlan shared-tunnels has changed. A system reboot is mandatory". A user doing a Junos upgrade with Apstra may re-experience this issue after the device is upgraded.
To avoid the need to a additional, manual reboot after a device Junos upgrade, the user can add the following configuration to the Apstra device system-agent pristine-configuration.
forwarding-options { evpn-vxlan { shared-tunnels; } }
This can be done in the "Decvices / Managed Devices / Pristine Configuration" Apstra UI or using the Apstra-CLI "system pristine_config_append" command.
For power supplies, Apstra primarily uses the output of the show chassis environment pem or show chassis environment psm command. On the QFX10002-36Q, both PEMs are reported, however, only one includes the XML tag that designates the component class as Power. The power supply information is further augmented using the show chassis environment command. Since the tag is missing for one PEM, Apstra does not recognize it as a valid power supply component. This is a known issue in Junos. Below is an example of the XML output from the show chassis environment without the class tag:
<environment-item> <name>FPC 0 Power Supply 1</name> <status>Present</status> </environment-item>
Due to inconsistent Junos behavior, the Power Supply State Check processor of Apstra does not evaluate the affected PEM, and no anomaly is raised. No workaround is available for this issue. This issue is expected to be addressed in newer Junos versions from 24.4R2, and the corrected behavior is expected to be present in supported versions for Apstra 6.1.0 and later.
In Junos EVO version 23.4R2-S5-EVO and 23.4R2-S6-EVO, devices do not export sFlow packets through the management interface. This issue affects all QFX device models.
You can resolve this issue using one of the following approaches:
1. If you require sFlow export through the management interface, please use the Apstra-qualified Junos EVO release 23.4R2-S4-EVO instead of 23.4R2-S5-EVO.
2. If you are using Junos EVO release 23.4R2-S5-EVO, configure sFlow to export through a non-management (in-band revenue) interface instead of the management interface.
2026-07-23: Added AOS-62738
2026-06-08: Added AOS-61624
2026-05-26: Updated AOS-59416, Added AOS-61246
2026-05-15: Updated AOS-54864
2026-04-21: Added AOS-52788
2026-04-14: Added AOS-60612
2026-04-13: Added AOS-59984
2026-04-03: Added AOS-55780
2026-03-23: Added AOS-60167,AOS-60183
2026-03-18: Removed AOS-51029,Updated AOS-59821,Added AOS-59842,AOS-60114
2026-03-05: AOS-59821,AOS-59931
2026-03-03: AOS-47430,AOS-59726,AOS-59869,AOS-59895
2026-02-18: Added AOS-54971
2026-02-11: Added AOS-59269
2026-02-03: Added AOS-59198,A0S-59416
2026-01-08: Added AOS-48525,AOS-58505
2025-12-11: Added AOS-58568,AOS-58680
2025-12-05: Added AOS-44623,AOS-57025
2025-11-19: Updated AOS-57740
2025-11-18: Updated AOS-57740, Added AOS-58200
2025-11-10: Updated AOS-57490
2025-10-28: Added AOS-57740
2025-10-02: Added AOS-56571, AOS-57297, AOS-57315, AOS-57490
2025-09-17: Added AOS-57099
2025-09-08: Added AOS-56498
2025-08-26: Updated download link for AOS-54413, AOS-55184, AOS-55286, AOS-55673. Updated AOS-55673
2025-08-20: Added AOS-56175, AOS-56282
2025-08-13: Added 40023, AOS-43348, AOS-45139
2025-07-29: Added AOS-52519
2025-07-23: Added AOS-55673
2025-07-10: Updated RFE-3429 with SONiC 4.4.2, Added AOS-55184
2025-07-09: Added AOS-55286
2025-06-26: Added AOS-54006
2025-06-25: Added AOS-54607,AOS-54767,AOS-54864
2025-06-09: Added AOS-54497, AOS-54513
2025-05-27: Initial publishing