Alert Type

SRN - Software Release Notification
Low/NotificationNA
Low/NotificationNA

Product Affected

Juniper Apstra

Alert Description

Juniper Apstra software product version 6.0.0 is available to licensed, registered Juniper customers from the Juniper Apstra software download site.

Documentation for Juniper Apstra 6.0.0 is available from the Juniper Apstra documentation site.

Junos Selective Update (JSU) feasible

Not applicable

Call to Action

NA

Solution

Juniper Apstra Version 6.0.0 Release Notes



New Features

 

Predefined DCQCN configlet for AI Clusters (RFE-3519)

Feature Category: Device Operating Systems

You now can use in your AI Blueprints the predefined configlet for DCQCN, which can be found in the catalog under the label "Datacenter QoS Congestion Notification". Use the property-set with the same label to adjust the configuration parameters such as the drop profile's fill level.



Support of Rail-Collapsed Rack and Templates designs (RFE-3433)

Feature Category: Design, Build, Operate

You can now create a Rail-Aligned Rack-type in a Collapsed design. Use the "Rail-Collapsed" fabric connectivity design in the Create Rack type menu. This creates a Rail-only rack structure, without spines, useful when the server count can be accommodated within a single Stripe. By eliminating the need for spine switches, you can achieve significant cost reductions. However, it's important to note that migration from a Rail-collapsed to a Clos design is not possible. To ensure optimal performance, verify that the GPU node is compatible with PXN for NVIDIA or its equivalent for other vendors. PXN enables NVIDIA NVSwitch connectivity between GPUs within the node, allowing data to move to a GPU on the same rail as the destination before being sent to the destination without crossing rails. The Simplified AI Template Designer also supports the Rail-Collapsed design.



Support for Rail-Aligned designs in Rack-Types (RFE-3432)

Feature Category: Design, Build, Operate

You can now create a Rack-Type with Rail support as a first-class citizen component for constructing your AI blueprints. The Rail-Index is automatically designated, but you have the option to manually specify it if you prefer. You can choose multiple generics and a single leaf to form interfaces within the same Rail. Additionally, you have the flexibility to merge rail-aligned and traditional server connectivity within a single rack.



Support for Load Balancing Policies (RFE-3376)

Feature Category: Design, Build, Operate

You can create in your AI Blueprints a Load Balancing Policy and choose between Standard DLB (Dynamic Load Balancing) or GLB (Global Load Balancing). You can choose between Flowlet or Per Packet mode and customize all critical configuration parameters, including Inactivity Timer, Sampling Rate and Egress Quantization parameters. Once defined, apply the load-Balancing policy selectively or in bulk on the target device(s). The system automatically prevents you from deploying unsupported policies, such as an attempt to deploy GLB on a topology where not all devices are QFX5240 (mandatory requirement for GLB to operate).



Simplified AI Template Designer (RFE-3394)

Feature Category: Design, Build, Operate

You can now use the simplified AI template designer, which takes 5 inputs to generate a template that can then be deployed as a blueprint. These inputs are: Total Number of Servers, GPUs per Server, Servers Per stripe, GPU NIC speed and Oversubscription ratio. This allows you to quickly and easily create a validated and optimized Apstra template tailored to your specific resource requirements. You can also export these racks from the template to use them for day 2 operations. The Template designer uses by default the Juniper recommended devices, including the following Device Profiles: QFX5220, QFX5230-64CD, QFX5240, and PTX10004/8/16 with 36CD Line card. You can use the 'Advanced Calculator Settings' section to customize this list by adding/excluding specific models. The Template designer will select Generic Systems Logical Devices, which can then be mapped to NVIDIA DGX Device Profiles.



Automatic Virtual Networks provisionning in rail-aligned blueprints (RFE-3518)

Feature Category: Design, Build, Operate

You can now include in your AI Blueprints an automatic pre-calculation and pre-provision of the required Virtual Networks, along with their associated Connectivity Template definition and Connectivity Template assignments. Under Physical -> Racks, you will find a new "Rails" tab, which will list any missing Virtual Network in the form of a warning and allow you to bulk provision them for every server in every rail in the blueprint. This streamlines the process of creating and configuring Virtual Networks for Rail-aligned blueprints by eliminating the mental burden of managing a large number of VLAN assignments.



Support for Onbox agent for NVIDIA DGX servers (A100, H100, H200) (RFE-3409)

Feature Category: Telemetry and Analytics

You can now install Onbox system agents on NVIDIA DGX servers, including the following models: A100, H100, and H200. Use the Create On-box agent in Telemetry-Only mode and provide an IP address allowing Out-Of-Band access to the server, and use a user with sudo privileges and passwordless access to the device. Acknowledge the systems after the successful agent install and follow the standard procedure for Serial Number assignment in the blueprint.
The following Telemetry services will be enabled on that device:

  • LLDP,

  • Hostname,

  • Interface,

  • Interface_Counters,

  • Resource_Util,

  • Disk_Util,

  • Gpu_Hardware_Counters

  • Gpu_Infiniband_Dev_To_Interface.
    With Hostname and LLDP, you now have validation and expectations for the Server-to-Leaf cabling. You can use the existing "Fetch Discovered LLDP Data" capability to bulk read the configured hostnames and bulk update the blueprint with that information; that simple two-step process will keep the source of truth synched, and so the accuracy of the cabling validations. Other telemetry services are used in different IBA Probes.



Stripe and Rail Traffic Probe and Dashboard (RFE-3520)

Feature Category: Telemetry and Analytics

You can now have in your AI Blueprints a Per Stripe dashboard auto-created for every new Stripe. This Dashboard leverages data from the predefined and auto-enabled "Stripe and Rail Traffic" probe and provides the following information: Per-Rail aggregated traffic view both live and historical (last 7 days) views as well Rail imbalance detection with anomaly raising where the imbalance threshold is tunable from the predefined probe menu.Â
Additionally, you now have a blueprint wide dashboard labelled "All Stripes Traffic Summary", leveraging the same IBA probe as a source of source and gives a broader view with a break-down of Inter-Stripe traffic vs. Inter-Stripe traffic both live and historical (7 days) on a per Stripe basis. You also can have Stripe level imbalance with tunable threshold.



Interface counter collectors for Nvidia GPU platforms (RFE-3337)

Feature Category: Telemetry and Analytics

Now you can collect and analyze Nvidia GPU interface counters to help with congestion and traffic monitoring in real-time.



GPU NIC traffic statistics (RFE-3522)

Feature Category: Telemetry and Analytics

You can now see in your AI Blueprints a new section in the blueprint dashboard view labelled GPU NIC utilization. This provides a honeycomb style of visualization providing Transmitted and Received network utilization for each GPU's NIC. Various filters are available for drill-down as well as grouping options. You can drill-down on a per Stripe, Rail or GPU Server. You can also group per Rail or per Stripe for more aggregate summary. Selecting any item (cell) provides you with several hyperlink redirects for more topological context.



"Interface Queue Stats Monitoring" probe and dashboard (RFE-3345)

Feature Category: Telemetry and Analytics

You can now validate in your AI Blueprints the Ethernet lossless service by leveraging the predefined and auto-enabled "Interface Queue Stats" probe along with its predefined and auto-enabled "Queue Stats Monitoring" Dashboard. The probe collects the following key congestion control metrics: Transmitted ECN, Transmitted and Received PFC, Ingress and Egress Buffer Utilization and Packet drop ratio. The data collection is done on all server-facing interfaces for queues 0, 3, 4 which corresponds respectively to Best-Effort, CNP and NO-LOSS forwarding classes. If you need to monitor additional queues, you can specify it through the "Queue Specific Thresholds" section of the predefined probe menu. Per-Queue anomaly thresholds for ECN, PFC, Buffer occupancy and Packet discard can be adjusted through the same menu. The probe keeps data retention for 14 days.



"GPU Hardware Traffic Monitoring" probe and dashboard (RFE-3521)

Feature Category: Telemetry and Analytics

You can now have in your AI Blueprints monitoring the GPU's NIC ROCE level statistics through the predefined and auto-enabled "GPU Hardware Traffic Monitoring" probe. This probe provides you with visibility over 30 NVIDIA Linux Hardware counters, documented on the NVIDIA website https://enterprise-support.nvidia.com/s/article/understanding-mlx5-linux-counters-and-status-parameters. The probe will raise anomalies for Received CNP, Received and Detected Out-Of-Sequence packets. These metrics help you understand whether your Load Balancing and DCQCN parameters are correctly set or if they need fine-tuning. The probe keeps data retention for 14 days.



Upgrade Paths to Apstra 6.0.0 (RFE-3384)

Feature Category: Platform

This release supports upgrade paths from previous Apstra 5.0.X and 5.1.X releases.

Users must use VM-VM upgrades from Apstra 5.0.X and 5.1.X releases. See the Apstra Installation and Upgrade guide for more information on Apstra upgrades.



Changed Features

 

Qualified switch operating systems with Apstra 6.0.0 (RFE-3429)

Feature Category: Device Operating Systems

The following updates have been made for switch operating systems qualified for the Apstra 6.0.0 release.

Juniper Networks:
Junos Evolved for AIDC:
23.4x100-D20

Junos (All roles):
21.4R3
22.2R3
22.4R3
23.4R2-S4

Junos Evolved for IP-Forwarder role (Spines in EVPN or any role in an IP-Fabric):
22.2R3-EVO
22.4R3-EVO
23.4R2-S4-EVO

Junos Evolved for EVPN leaf roles:
22.2R3-EVO
22.4R3-EVO
23.4R2-S4-EVO

Junos Interconnect Gateway Leaf:
22.4R3 (minimum)
23.4R2-S4

Junos Evolved Interconnect Gateway Leaf:
22.4R3-EVO (minimum)
23.4R2-S4-EVO

Cisco Systems:
9.3(13)
10.2(6)
10.3(4a)

Arista Networks:
4.24.5M
4.28.7.1M
4.30.3M

Dell EMC & Edgecore:
Enterprise SONiC 4.1.2
Enterprise SONiC Edge Standard 4.1.2
Enterprise SONiC 4.2.1
Enterprise SONiC Edge Standard 4.2.1
Enterprise SONiC 4.2.3
Enterprise SONiC Edge Standard 4.2.3
Enterprise SONiC 4.4.2
Enterprise SONiC Edge Standard 4.4.2



GPU nodes support in "Device System Health" probe (RFE-3370)

Feature Category: Telemetry and Analytics

You can now monitor CPU, Memory, and Disk usage of your GPU nodes in the predefined "Device System Health IBA" probe. The predefined probe menu has been revised to display differentiated thresholds per system type for switches and servers. The predefined Dashboard widgets have also been expanded to reflect the new collected metrics and alerts.



TLS support for Streaming receivers (RFE-3373)

Feature Category: Platform

You can now use TLS when you define a telemetry streaming receiver.
You can upload a certificate to the Apstra server (local store). The certificate must be a valid x509 certificate.
This enables you to ensure that confidential telemetry information remains protected during transfer, thereby further enhancing your security posture and facilitating compliance with regulatory obligations.



Correct display of AOS-SDK version in pip list / pip show commands (RFE-3234)

Feature Category: Platform

You can now see the AOS-SDK version in the pip show aos-sdk command output.



Removed Features

 

Deprecation of Rack-Type builder (RFE-3425)

Feature Category: Design, Build, Operate

The Rack-Type Builder method for creating Rack-Type has been discontinued in favor of retaining only the Rack-Type Designer. The Designer has been enhanced with several new functionalities to closely match the existing capabilities of the Builder.



Fixed Apstra General Issues

 

After upgrade, composite CTs that include dynamic prefix peering with IP link are not split into separate CTs when the CT is not attached to any interface in the source VM (AOS-53211)

Application point type has been changed for Dynamic BGP prefix peering ("Dynamic BGP Peering" primitive type with any of "IPv4 Subnet for BGP Prefix Dynamic Neighbors" or "IPv6 Subnet for BGP Prefix Dynamic Neighbors" fields populated) in Apstra release 5.1.0. Previously the primitive was applied to SVI or Subinterface, after the fix it is applied to Loopback in the corresponding Routing Zone. As a result, Connectivity Template with "IP Link" and prefix peering "Dynamic BGP Peering" primitives can be stacked in the UI, but can never be applied, because "IP Link" primitive produces no valid application points for "Dynamic BGP Peering" primitive with prefix peering.



Apstra AnomalyGenerator crash with PrimaryKeyIndex violation newRow (AOS-52744)

When AnomalyGenerator tries writing to MetricDB and taking SysDB snapshots if vlanid is missing from the primary key it will cause the AnomalyGernator to crash wih the below trace information


Unique Index (PrimaryKeyIndex) violation newRow index matches that at row: 65536 python3.10: /Project/leblon/infra/TableTop.tin:206: void Aos::DoubleLinkListHelper::addToList(Aos::RowIndexHelper&, U32): Assertion `false && "!row"' failed. Process 22875 died with signal 6 (SIGABRT) errno 0 code -6 (unknown)


Apstra CLI device password change may fail due to task timeout (AOS-54971)

The scenario change-device-password CLI command securely updates device credentials by performing tasks like SSH checks, configlet staging, blueprint commits, and agent password updates. In the current Apstra CLI version, the system agent check has a 60 second timeout, while configlet staging is limited to just 20 seconds. If these operations take longer than expected, the command may fail with errors like:


Failure 1: Task Stage creation of Configlet for password change may fail with: AssertionError: Timeout waiting for Wait that last task status is succeeded Failure 2: Task Check System agent status may fail with: 409 Conflict: Agent is already running a job (check)

These failures occur when backend tasks exceed the current timeout settings, which is particularly noticeable in Apstra 4.2.x and later versions, where performance issues with configlet and configuration rendering are known. Additionally, longer durations in the check job can result from changes in the customer’s environment.

Resolution

The issue is resolved in Apstra CLI versions 4.2.2, 4.2.2.1, 5.0.0, 5.1.0, and 6.0.0, which increase the timeout thresholds for background tasks such as configlet staging and agent status checks. These updates prevent premature task failures caused by longer processing times in earlier versions.



Apstra UI prevents setting "IP Links to Generic Systems MTU" to 9216 under fabric settings (AOS-53627)

Apstra customers 4.2.x and greater may encounter an issue where the MTU value for IP Links to Generic Systems cannot be set to 9216 via the UI. Although the UI states that only even values in the range 1280-9216 are accepted, the input of 9216 is incorrectly rejected, while 9214 is accepted.


Important Note: 1. This issue is applicable only to customers upgrading from Apstra 4.1.x to 4.2.x or later, where Fabric MTU remains disabled post-upgrade and customers who wish to continue without enabling the Granular MTU feature. 2. This issue is not applicable to customers with fresh 4.2.x deployments, where Fabric MTU is enabled by default, activating the Granular MTU feature.
Resolution

Apstra addressed the UI validation bug in 6.0.0 to allow users to set the IP Links to Generic Systems MTU value to 9216 under Fabric Settings.



Blueprint Dashboard page is rendering a blank page when Blueprint is clicked on the Apstra Web UI (AOS-53355)

When Dashboard for blueprint is displayed via clicking Blueprint in the UI, Apstra UI failed to load the blueprint dashboard because the preference information of the dashboard is configured with an empty string value, not the right value for preference.

Resolution

Apstra Web UI validates the empty string for preference and translates it as the right value.



Changing peer keepalive IP address of MLAG generates a deploy error in Arista EOS device (AOS-52892)

When attempting to change the keepalive IP address of an MCLAG pair of EOS devices, the configuration deployment of the devices belonging to the pair may fail.



Configlets using Jinja custom function.merge_vlans_to_list() fails with TypeError on datacenter interface device model allowed_vlans (AOS-52484)

When making use of the built-in jinja function on a custom configlet, {{ function.merge_vlans_to_list(interface_model["allowed_vlans"]) }} fails when used within a configlet. An error will be seen in rendered configuration previews "TypeError: unsupported operand type(s) for -: 'int' and 'str'"

Resolution

AOS 6.0.0 adds support for both integers and strings within function.merge_vlans_to_list()



ContainerLauncherAgent does not retry launching containers after a Docker client timeout, which can result in missing offbox containers (AOS-52431)

When system resources are scarce, Docker (dockerd) may fail to respond within the expected timeframe (60 seconds), resulting in a connection timeout error. This keeps the ContainerLauncherAgent from successfully launching offbox containers. The logs show that high CPU load and late clock events contributed to this failure. As a result, the agent becomes stuck, with multiple "Launch action already pending" warnings appearing in the logs. This prevents certain containers from starting and leaves them in a "absent" state.


2025-02-05 11:45:35,039 INFO aos.cluster.container_launcher:Update task container aos-offbox-10_42_29_205-f status: state='absent', error=Not found 2025-02-05 11:47:06,379 WARNING aos.cluster.container_launcher:Launch action already pending for container 'aos-offbox-10_42_84_192-f' requests.exceptions.ReadTimeout: UnixHTTPConnectionPool(host='localhost', port=None): Read timed out. (read timeout=60)

The ContainerLauncherAgent currently operates on two threads. The first thread identifies containers that need to be scheduled and sends them to the second thread, which interacts with Docker. If a Docker connection timeout occurs, the second thread fails, preventing the agent from completing its tasks. Because there is no automatic retry mechanism in place, the container goes missing.

Customers can workaround this issue by restarting the ContainerLauncherAgent. It will cause the containers relaunched correctly.



Deployment may fail on EOS 4.30+ due to stricter transceiver checks (AOS-49957)

Starting with EOS 4.30+, Arista EOS introduces a stricter hardware validation mechanism through a new default system l1 configuration block. This change causes Apstra deployment to fail if speed is configured on ports without compatible transceivers, even if the ports are unused.


system l1 unsupported speed action error unsupported error-correction action error

To maintain compatibility, system agent has been updated to detect this configuration and automatically change the unsupported speed action error to unsupported speed action warning during agent installation. This ensures that systems with unused or unpopulated interfaces do not fail deployment due to this stricter validation. The updated system agent logic safely applies this change only for EOS 4.30.x and later.

Resolution

Apstra system agent has been updated to detect system l1 configuration and automatically change unsupported speed action error to unsupported speed action warn during agent installation. This ensures that systems with unused or unpopulated interfaces do not fail deployment due to this stricter validation. The updated system agent logic safely applies this change only for EOS 4.30.x and later.



Deployment Performance degrades when draining or deploying (AOS-54006)

When the device is deployed or drained, Apstra showed a noticeably longer delay in finishing the operation than the Apstra 4.1.X release. The problem was linked to the significantly increased delay in the Jinja configuration rendering area following Apstra's migration from Python version 2 to version 3. Additionally, it affects the rendering configuration for the blueprint's configlet processing.



For 16K GPUs AI/ML high scaled topologies, Apstra controller VM needs 16 vCPUs (AOS-53979)

In high-scale environments, such as 16K GPU AI/ML topologies, several containers occupy CPU resources for much longer periods of time, reducing CPU resource availability and causing heartbeat timeouts.

Resolution

Recommend increasing CPU power for the Apstra Controller VM by doubling the number of vCPUs (8 -> 16)



gRPC junos ephemeral database UI_EPHEMERAL_COMMIT filling partition (AOS-52519)

Apstra gRPC probing of JUNOS devices is causing large ephemeral database files which may fill the disk and cause issues accessing the device via SSH.

jtac-QFX5120-48Y-8C-r011 mgd[14126]: UI_EPHEMERAL_COMMIT: User 'root' has requested commit on 'junos-analytics' ephemeral database
jtac-QFX5120-48Y-8C-r011 mgd[14126]: UI_EPHEMERAL_COMMIT_COMPLETED: commit complete on 'junos-analytics' ephemeral database

Resolution

Apstra 6.0.0 introduced a different mechanism using an invalid XPATH query to verify the health of the gRPC service on the device, which doesn't introduce ephemeral DB changes. However, manual ephemeral purge configuration via configlet would still be recommended because the ephemeral DB might reach the full condition.

set system configuration-database ephemeral purge-on-version 30



JUNOS EVO on-box agent running 23.4R2-S5-EVO experienced silent NETCONF XML parsing error during configuration retrieval and update, resulting in a timeout and reconnect attempts (AOS-57315)

The JUNOS 23.4R2-S5-EVO on-box agent experiences a 2-minute timeout for the NETCONF operation with a silent XML parsing error message without raising an explicit error when it performs a get-configuration operation to compare against the golden configuration or a lock configuration operation to update any incoming changes via NETCONF. In an attempt to restore after the timeout, the agent starts the reconnect logic; however, this fails because it fails to include the proper binding step.

Resolution

In 6.0.0, the Onbox agent in the 23.4R2-S5-EVO (not qualified NOS) doesn't yield a silent XML parsing error, not triggering a timeout for NETCONF operations.



License information is not transferred during Apstra controller upgrade (AOS-52927)

The license information from the previous version of the Apstra controller is not transferred to the upgraded controller when it is upgraded. As a result, license information is missed in the upgraded Apstra controller.

Resolution

License files will be transferred from the old controller to the upgraded controller as part of the upgrade process.



Multiline banner motd or exec is not supported in the Cisco NXOS Device (AOS-40278)

A banner configured in the Cisco NXOS device must be single line. Multiline banner (motd or exec) is not supported.



Rack type designer doesn't allow more than a single Port channel ID for generics (AOS-52299)

The Rack Designer does not have Port channel ID Range min and max value fields while adding a generic system to the leaf. This range was added, allowing for more than a single PortChannel to a single Generic System. These fields must follow the following rules when assigning PortChannel IDs to Generics.

1. They must not overlap in scope of single switch.
2. The ranges are continuous (cannot have "holes"), e.g. Range1 [0..100] and Range2 [10..12] are considered to be overlapping even if only 1 PCID from the first range is used.

Resolution

Min and Max Port Channel fields were added to Rack Designer to allow for more than a single PortChannel assignment for a Generic System.



Rejected commit check state in the Junos Device not reflected correctly into Dashboard Service Config Status (AOS-53220)

When Junos devices become unreachable, triggering a commit check leads to a rejected state after approximately 600 seconds. However, devices in this rejected commit check state are put into commitCheckinProgress state after an AOS restart or Full Config Push.
The problem occurs because Apstra does not treat devices in the commitCheckInProgress state as pending. This leads to a misleading UI/UX experience (some devices may still be in the commitCheckInProgress state, but the main dashboard may show zero pending devices).

Resolution

Junos devices in the rejected commit check state are correctly reflected as pending in the deployment status of the blueprint dashboard.



Show Tech Backups by UI includes sensitive Apstra Edge credentials (AOS-53919)

When a user collects a controller Show Tech by UI with the include backup option enabled, the generated Show Tech also includes a backup of the aos-edge-auth.json file. This file contains sensitive information such as the registration_key, org_id, and other credentials used to authenticate with JCloud.

If this backup is restored in a different environment, the apstra_edge container in that environment may use these credentials to connect to JCloud. This can unintentionally trigger an Edge re-registration event, disrupting the original environment by recreating its active Edge instance.

Resolution

In Apstra 6.0.0, Show Tech backups taken from the GUI no longer include sensitive files like aos-edge-auth.json, keeping credential information secure. However, when a backup is taken through the CLI, the authentication file is included, since these backups are meant for full system recovery during failures or outages.



SONiC FRR restart or device reboot may cause configuration anomaly from rearrangement of FRR running configuration sections (AOS-49906)

Rebooting the device or restarting FRR in SONiC may cause the FRR running configuration sections (related with route-map) to be rearranged. The rearranging of sections will typically show a configuration deviation even if the running configuration is exactly the same as before.



TaskScheduler agent's continuous restart because of missing application weight for Apstra Edge container (AOS-52217)

Apstra uses application weight information for container types (iba, offbox, and apstra_edge) prior to launching a container. The upgrade process for only upgraded environments (5.0.X to 5.1.0) fails to include application weight information for the apstra_edge container type, causing the TaskScheduler agent to crash and restart. When AOS is restarted, the situation may worsen (all offbox and iba agents remain down). Please apply a workaround to resolve the issue. This issue would not exist in a 5.1.0 clean deployment environment.

Resolution

The fix is to to update the application weight using API call “/api/cluster/application-weight�.. PUT on this

{
"iba": 1000,
"offbox": 250,
"apstra_edge": 500
}

After that AOS restart is needed.



The UI page remains stuck in a loading state when navigating from the Stage tab to the Uncommitted tab (AOS-47430)

Upon navigating from the Stage tab to the Uncommitted tab, the page does not load and continues to display the loading spinner. The UI is continuously polling the API endpoint /api/blueprints//tasks?mode=full, which forces the backend to attach the complete task payload to every response. Since these payloads can contain large data information, the API responses become significantly heavier. Because the UI repeatedly polls this endpoint, it increases backend processing time and causes the page loading issue, particularly when tasks have large payloads.

Resolution

This polling behavior has been optimized in AOS version 6.0.0 to avoid unnecessary full payload retrieval during regular status checks. Recommend upgrading to Apstra version 6.0.0 to benefit from this performance improvement.



UI doesn't allow IM (Interface Map) creation with mixed transformations for the same speed (AOS-47654)

When the Device Profile allows multiple transformations for the same speed and the IM (Interface Map) needs mixed transformation for the same speed across ports (for example, Port 1: 1X100G, Port 2: 4X100G, etc.), it is not possible to create an IM using the UI.



UI show tech collection for controller fails with "Show tech execution failed on target apstra_vm" (AOS-51964)

The current default timeout for collecting show tech via UI is 20 minutes (1200 seconds) in the controller_timeout value of the show_tech section of aos configuration file /etc/aos/aos.conf. The collection contains operations for exporting all binary log files from all running agents across containers into printable format files. Depending on the number of accumulated log files, the translation processing time may exceed the default timeout for collecting showtech, resulting in timeout failures for the UI showtech collection task.



Unable to delete a Datacenter Blueprint entry from the Blueprint page when Apstra fails to create an unsupported blueprint (AOS-52300)

Users will not be able to create a Datacenter Blueprint using the EVPN reference design with IPv6 RFC:5549 as the Spine to Leaf Links Underlay Type, as this configuration is not currently supported. This limitation is tracked in RFE-1364 which is planned for release in 6.1.0. If the user attempts to create an unsupported blueprint using the steps below, the UI will throw a critical error message - EVPN is not allowed with spine-leaf IPv6 links addressing policy.


Steps to Reproduce: 1. Create a template with the overlay control protocol set to MP-EBGP EVPN. 2. Navigate to the Blueprint Creation page. 3. Select Datacenter as the reference design. 4. Set Spine-to-Leaf Links Underlay Type to IPv6 RFC:5549. 5. Attempt to create the blueprint.

However, the blueprint becomes locked and cannot be deleted, even with admin privileges. Deletion can only be performed via REST API Explorer or CLI. This issue was addressed in 6.0.0, ensuring that the delete function works correctly.

Resolution

This issue will be addressed in 6.0.0, ensuring that the delete function works correctly.



Virtual Network configuration changes do not properly reflect as changed after making a change in Apstra (AOS-40852)

After editing a VN (Virtual Network) configuration, if the same Virtual Network is open, none of the changes appear until the VN is opened again in the UI.

Resolution

VN changes now properly reflect after the saving and reopening the same VN which changes were made.



When Generic System is added/deleted from leaf device in the rack, fail with an error not enough ports on leaf to connect (AOS-51116)

When a rack is built with leaf devices and generic systems, group labels for generic systems can have the same value as leaf or access switches' target_switch_label, contrary to the expectation that the group label should not be the same value as target_switch_label inside the rack. Any changes to the rack, such as adding a generic system or deleting an existing generic system, would fail due to the validation error caused by not meeting the above expectation.



Fixed Third-Party Issues

 

Junos EVO gRPC Sequence Number Overruns (AOS-50857)

When Aptra registers a subscription for a specific xpath into Junos EVO device, Junos EVO device may keep multiple subscriptions for the same path and cause telemetry sequence overruns in telemetry data.

Resolution

When available, upgrade to Junos EVO 23.4R2-S4 or greater



Static route for loopback address of external router not installed in the SONiC device (AOS-45557)

If a SONiC device removes and then re-adds an IP address, the device may fail to add a static route involving that address to the kernel routing table, even if the static route configuration exists. The show ip route output in vtysh in the SONiC device experiencing this issue may include the following output lines:


S>r 198.51.100.2/32 [1/0] via 192.168.0.9, Po1.4, weight 1, 01:59:25 B * 198.51.100.2/32 [20/0] via 10.0.0.3, Vlan201 onlink, weight 1, 01:59:25

Above, the "S" static route has been rejected by the kernel and was not installed.

Resolution

This is an FRR bug in SONiC 4.1.2 and 4.2.1. A fix for this defect is included in SONiC 4.4.x and is verified to work. Please refer to vendor issue SONIC-89523(Next Hop Group fails to install in the kernel).



Traffic from leaf carries with wrong vlan id when VLAN configuration is changed from mixed untagged and tagged vlans to only tagged vlans (AOS-51183)

This issue was discovered in the Juniper QFX 5230/5240 Juniper EVO device for non-EVPN blueprints (Pure IP Fabric) running the 23.4R3-S3-EVO version. When an interface is assigned both untagged and tagged vlans and then the untagged vlan is removed, the traffic from the leaf device will have the incorrect VLAN ID, causing network connectivity issues.



Known Apstra General Issues

 

[IBA] Filtering by the state column in the output of the state processor does not work as expected (AOS-51538)

The state_check processor outputs a state column for each series. However, users may not be able to filter series in the output stage using the per series state column when querying stage data. Users attempting to use a filter such as properties.state = 'false, true' may see no results, even if matching data is present. This is a known limitation in how filtering works on series-level data in the state_check processors output. The filtering behavior may not align with user expectations when attempting to use conditions on the state column. Engineering evaluating potential improvements to allow filtering on state column in state processor for a future release.



A large bump in the memory footprint for MetricQueryManagerAgent is seen when the 'Time Series' data option is selected on the Active->Anomalies tab (AOS-48770)

MetricQueryManagerAgent handles large historical data to serve the '/blueprints//anomalies-history' endpoint. Depending on the amount of data, the agent's memory footprint may increase significantly. The benchmark environment recorded a memory footprint of up to 2.2Gb for the agent. The memory footprint settles after the initial bump.
If the system administrator is concerned about the MetricQueryManagerAgent footprint's impact on the system's available memory, the following workaround is recommended.

Workaround

Restart MetricQueryManagerAgent and avoid using the 'Time Series' query.



Access switch count option is unavailable in Rack Designer (AOS-52278)

In Rack-Type Designer, the ability to specify an Access switch count which was available in Rack Builder is currently not supported. The Rack-Type Designer, introduced in Apstra 4.2.0, replaces the traditional Rack-Type Builder to provide a more intuitive and user-friendly experience. As of Apstra 6.0.0, Rack Builder is deprecated and no longer available.

While Rack Designer does not include all functionalities of the Rack Builder, such as the ability to add multiple access switches and logical connections at once, the new interface offers a superior UX, improved workflow, and new capabilities. Some features may require different steps, while others have been redesigned or omitted for usability improvements.

Users can utilize the Clone functionality in Rack Designer to replicate multiple access switches as an alternative to the count feature.

For further improvement, feature requests may be required via the sales account team.

Workaround

Users can utilize the Clone functionality in Rack Designer to replicate multiple access switches as an alternative to the count feature.



AI Cluster Template - Incorrect selection count is displayed under 'Device profile to Consider' when the 'Manually Selected' option is used (AOS-54767)

In the AI Cluster Template workflow, selecting "Manually Selected" under Device Models incorrectly shows "4 selected" even when no models are chosen. The count should reflect actual selections or indicate default behavior clearly.



Apstra 6.0 UI rendering issue can cause predefined probes to error with TypeError: Cannot read properties of undefined (reading 'types') (AOS-55673)

When instantiating predefined probes such as "VMs Without Fabric Configured VLANs" Apstra UI may fail to display the probe with a type error, making the probe unusable.

Workaround

To workaround the issue, user need to download the UI hotpatch (https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HD1BNIA1) and apply into the controller VM as follows. Fixes for AOS-55184 and AOS-55286 are included in the UI hotpatch for AOS-55673:


admin@aos-server:~$ sudo su [sudo] password for admin: root@aos-server:/home/admin# gzip -d aos-web-ui-6.0.0-aos55673.run.gz root@aos-server:/home/admin# chmod 755 aos-web-ui-6.0.0-aos55673.run root@aos-server:/home/admin# bash aos-web-ui-6.0.0-aos55673.run Verifying archive integrity... All good. Uncompressing AOS WebUI installer 100% ### Backing up existing AOS WebUI into /opt/aos/frontend/snapshot/2025-07-23_20-31-02 ... Successfully copied 226MB to /opt/aos/frontend/snapshot/2025-07-23_20-31-02 ### Copying AOS WebUI file into aos_controller_1 ... Successfully copied 226MB to aos_controller_1:/opt/aos/frontend_images/ ### Initializing new AOS WebUI ... ### Done! root@aos-server:/home/admin# exit


Apstra 6.0 UI rendering issue causes ESI/MLAG leaf switches to appear in the wrong order (AOS-55286)

In Apstra 6.0, the Main Topology View (Blueprint > Staged > Physical > Topology) displays ESI/MLAG pairs with leaf2 positioned on the left and leaf1 on the right, which is the reverse of the layout seen in prior versions. In previous versions, the UI consistently rendered ESI/MLAG pairs with leaf1 on the left and leaf2 on the right, aligning with user expectations. The change in Apstra 6.0 is purely cosmetic and does not affect system behavior or configuration.

Workaround

To workaround the issue, user need to download the UI hotpatch (https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HCCgNIAX) and apply into the controller VM as follows:


admin@aos-server:~$ sudo su [sudo] password for admin: root@aos-server:/home/admin# gzip -d aos-web-ui-6.0.0-aos55286.run.gz root@aos-server:/home/admin# chmod 755 aos-web-ui-6.0.0-aos55286.run root@aos-server:/home/admin# bash aos-web-ui-6.0.0-aos55286.run Verifying archive integrity... All good. Uncompressing AOS WebUI installer 100% ### Backing up existing AOS WebUI into /opt/aos/frontend/snapshot/2025-07-09_09-59-37 ... Successfully copied 226MB to /opt/aos/frontend/snapshot/2025-07-09_09-59-37 ### Copying AOS WebUI file into aos_controller_1 ... Successfully copied 226MB to aos_controller_1:/opt/aos/frontend_images/ ### Initializing new AOS WebUI ... ### Done! root@aos-server:/home/admin# exit


Apstra ZTP DHCP server not honoring static IP address assignment reservation (AOS-52731)

When setting the reservation mode from the configurator using the different options except None and the reservation mode flag is set in the DHCP configuration, it causes the static IP address to be assigned from the pool rather than the static IP address mapped against the MAC address in the configurator.

Workaround

User can follow the below steps to get the static IP to the host using ZTP
1. If Reservation Mode is set to None, users can configure static IPs by navigating from ZTP UI to DHCPv4 Configurator, setting Reservation Mode to None, and toggling Reservations-Global.
2. If Reservation Mode is enabled, users must move the hosts inside Reservations to the subnet level based on the required IP address while keeping the Reservation Mode set to Global.



Apstra ZTP Duplicate Entries for Junos Devices (AOS-40023)

When monitoring Apstra ZTP device status in the Apstra UI under "ZTP Status" / "Devices", there may be duplicate entries for Junos devices. Apstra ZTP will try to ensure the physical management interface for the Junos device is used instead of any virtual management interface (e.g. "vme" interface). Junos may use the virtual interface when ZTP starts but cannot be added to the required "mgmt_junos" routing-instance. This is done as the first step in ZTP in order to ensure that the management IP address does not change during the rest of the steps involved in ZTP (especially those involving connectivity to Apstra). Enabling a different management interface will cause the DHCP server to give out a new lease. Also, the vendor class identifier for the new management interface is cleared so that the DHCP server does not give out vendor-specific options to this interface, which may re-trigger a new ZTP session while the current session is active. This is expecetd behavior.



Backward incompatible change for the AosMessage.timestamp field from uint64 to google.protobuf.Timestamp (AOS-50584)

The timestamp field of AosMessage in the streaming has used uint64 format for both millisecond and microsecond timestamp information. The current timestamp uin64 field will be changed in release 6.1.0 to the google.protobuf.Timestamp format, which is incompatible with the previous version, in order to provide consistent, accurate timestamp information. The streaming receiver side must be modified to accommodate this incompatible modification.



BGP peers configured on IRB/Loopback interfaces may flap on Juniper EVO device during commit (AOS-59821)

An event involving BGP flaps from BGP peers configured on IRB/Loopback interfaces for Juniper EVO device has been reported during the commit with incremental changes (deleting an IRB/Loopback interface in the same VRF). The BGP flaps happen when the device's configuration is committed in override mode, which was Apstra's default setting prior to 6.1.0. However, using load update mode did not reveal the issue.

Workaround

If the EVO device is running as an off-box agent, please add key load_mode with value update into open options in the edit agent menu. Otherwise, recommend upgrading Apstra to 6.1.X, which uses load update as the default mode for commit.



Bulk show-tech collection on multiple devices may lead to high disk utilization on the controller, causing job failures, stuck tasks and repeated SystemAgentManager crash (AOS-60114)

When multiple show-tech jobs are triggered simultaneously, the /var/log partition can reach high utilization, causing the controller to enter read-only mode. This may result in incomplete job status updates, leaving several jobs stuck in in-progress or pending states. In a corner case, this condition can also lead to repeated crashes of SystemAgentManager, preventing automatic recovery even after disk space is reclaimed. Avoid triggering bulk show-tech collection on a large number of devices. Ensure sufficient disk space is available before running show-tech.

Workaround

If the issue occurs, reclaim space in /var/log and restart AOS services. In most cases, SystemAgentManager recovers automatically, clearing stuck jobs and completing or failing pending ones. If SystemAgentManager crash persist and jobs remain stuck, manual cleanup via Acons is required. It is recommended to contact Apstra Support for assistance.



Click on Give Feedback icon in UI causes "form doesn't exist." error message (AOS-55184)

User feedback was added in Apstra 6.0.0 to enable users to share their Apstra experiences. After the release of Apstra 6.0.0, the link to the feedback form was changed. It results in an error message about a missing form.

Workaround

To workaround the issue, the user needs to download the UI hotpatch (https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000HCCgNIAX) and apply it to the controller VM as below. The hotpatch addresses both the AOS-55184 and AOS-55286 issues.


admin@aos-server:~$ sudo su [sudo] password for admin: root@aos-server:/home/admin# gzip -d aos-web-ui-6.0.0-aos55286.run.gz root@aos-server:/home/admin# chmod 755 aos-web-ui-6.0.0-aos55286.run root@aos-server:/home/admin# bash aos-web-ui-6.0.0-aos55286.run Verifying archive integrity... All good. Uncompressing AOS WebUI installer 100% ### Backing up existing AOS WebUI into /opt/aos/frontend/snapshot/2025-07-09_09-59-37 ... Successfully copied 226MB to /opt/aos/frontend/snapshot/2025-07-09_09-59-37 ### Copying AOS WebUI file into aos_controller_1 ... Successfully copied 226MB to aos_controller_1:/opt/aos/frontend_images/ ### Initializing new AOS WebUI ... ### Done! root@aos-server:/home/admin# exit


Commit failures in the JUNOS and EVO devices are caused by newline characters in the virtual network description field (AOS-54198)

If the rendered device configuration for the VLAN description contains newline characters populated from the virtual network's description field, the commit operation fails because JUNOS and JUNOS-EVO devices do not support multi-line string for the VLAN description.

Workaround

Please use one-line formatted string in the description field of Virtual Network rather than multi-line string.



Configuration anomalies in the SONiC device caused by the configlet when the device agent restarted after losing connection to the controller (AOS-50752)

While configlet is being applied to the SONiC device, if the device agent restarts after being disconnected from the controller, the agent executes any remaining changes and collects the running configuration as golden configuration to monitor for configuration anomalies. Because the process of applying configlet changes is still running independently of the agent, it introduces changes into the running configuration even when the golden configuration is collected by the agent. The following changes from the process cause configuration anomalies in the SONiC device.

Workaround

After reviewing the running configuration on the SONiC device, if all the changes from the configlet are correctly applied, the customer can safely accept changes to avoid further configuration anomalies.



Configuring more than one AAA server via the UI results in a configuration load error in the JUNOS and EVO device during commit check or commit (AOS-60183)

Adding multiple AAA servers in the blueprint through Staged > Catalog > AAA Servers leads to a configuration load error in the JUNOS and EVO device during commit check or commit.

Workaround

Recommend using configlet instead of using UI (Stage > Catalog > AAA Server) when multiple AAA servers needs to configured.



Conversion of leaf switches from ESI-based to MLAG-based redundancy using APIs within an existing blueprint fails when links to an external generic system are present (AOS-59198)

When attempting to convert leaf switches from ESI to MLAG within the same blueprint, the operation fails because Apstra does not allow mixing ESI and MLAG redundancy models at the rack level. This restriction is enforced starting in Apstra 4.2.0 and is expected behavior. During the operation, users may see the following error in the UI or REST API Explorer:


"Combining MLAG and ESI leaf pairs not supported"

Currently, converting ESI racks to MLAG racks or vice versa requires replacing all racks within a single FE operation using the REST API. However, due to limitation, this conversion can cause BuilderAgent to fail when links are present between external generic system and leaf switches. In this occurs, please revert the changes and follow the steps outlined in workaround section.

Workaround

To successfully convert the rack type using modify-racks API, the following workaround can be used:


1. Identify the leaf switches connected to the external generic system 2. Identify the Connectivity Templates (CTs) associated with the external generic interfaces and unassign the corresponding application points 3. Delete the links between the leaf switches and the external generic system 4. Undeploy and unassign the leaf devices from the blueprint 5. Unassign interface maps from the blueprint 6. Use the REST API /api/blueprints/{blueprint_id}/modify-racks to convert the racks from ESI to MLAG 7. Import MLAG-compatible interface maps and assign them to the switches 8. Recreate the links to the external generic system 9. Reassign the endpoints to the appropriate Connectivity Templates

For additional guidance, please contact Apstra Technical Support.



CT(Connectivity Template) field may appear with empty value (AOS-56498)

While viewing/editing CT (Connectivity Template)s across blueprints, it's possible that the CT may incorrectly display empty field values.

Workaround

By clicking the browser refresh button, CT would display the correct data.



Deleting Virtual Networks in CTs with Multiple VLANs - All Active Endpoints Unassigned (AOS-44623)

In version 4.2.0, Apstra introduces the capability for users to forcibly delete a Virtual Network, even if it has active endpoints. Apstra will initially display the interfaces to which the Virtual Network (VN) is currently allocated and prompt the user to confirm the deletion. It's important to note a limitation in the current design: if a user deletes a VN assigned in a CT where Multiple VLANs are present, all active endpoints will be unassigned.

Workaround

User should manually remove the specific VLAN from the CT before proceeding to delete it from the Staged > Virtual Networks section.



DeviceTelemetryAgent.{pid}.log by gRPC trace logs filling up disk (AOS-51846)

DeviceTelemetryAgent.{pid}.log files in /var/log/aos/ in the offbox agents become large and can fill up the disk

Workaround

The following Python script can be added to run via crontab on an hourly basis, which will clean up older log files. This workaround needs to be applied to controller VM and worker VMs where offbox agents are running (nodes with offbox tags in the Platform/Apstra Cluster/Nodes).

      1. copy from next line


# Copyright 2024-present, Apstra, Inc. All rights reserved. # # This source code is licensed under End User License Agreement found in the # LICENSE file at http://apstra.com/eula import configparser import json import os import re import shutil import subprocess import traceback SystemIdPattern = re.compile(r'AOS_SYSTEM_ID=offbox,(.+),(.+)') def update_aos_conf(task_id): aos_config = os.path.join( '/var/lib/aos/conf.d/task/offbox/', task_id, 'aos.conf', ) parser = configparser.ConfigParser() if os.path.isfile(aos_config): parser.read(aos_config) if not parser.has_section('logrotate'): parser.add_section('logrotate') if 'max_kept_backups' not in parser.options('logrotate'): parser.set('logrotate', 'max_kept_backups', '1') staging_file = aos_config + '.staging' with open(staging_file, 'w') as f: parser.write(f) shutil.move(staging_file, aos_config) return True return False def refresh_logging_infra(container_id): subprocess.check_output([ 'docker', 'exec', container_id, 'pkill', '-HUP', 'DeviceKeeperAge' ]) def get_offbox_containers(): containers = subprocess.check_output([ 'docker', 'ps', '-q', '--filter', 'label=AOS_CLUSTER_APPLICATION=offbox', ]).decode() return containers.splitlines() def get_task_ids(): def extract_info(container_env): try: envs = json.loads(container_env) except ValueError: return None, None for env in envs: matched = SystemIdPattern.match(env) if matched: return matched.group(1), matched.group(2) return None, None containers = get_offbox_containers() if not containers: return containers_env = subprocess.check_output([ 'docker', 'inspect', '--format', '{{json .Config.Env}}', *containers, ]).decode() for line in containers_env.splitlines(): task_id, container_id = extract_info(line) if not task_id: print('Failed to extract task id from container env: {}'.format(line)) continue yield task_id, container_id def main(): for task_id, container_id in get_task_ids(): try: if update_aos_conf(task_id): refresh_logging_infra(container_id) except: print('Failed to update aos.conf for container: {}'.format(container_id)) traceback.print_exc() main()
      1. script finishes here. don't copy this line

Even if the above workaround is applied, there is a chance of filling up partition. The below command can be executd with root permission to clean up logs quickly in the controller VM and worker VMs.


find /var/log/aos/task -name "*.log" -size +10M -print | grep "DeviceTelemetry" | xargs -I {} sudo cp /dev/null {}

 



ECN Marked Packet Information is not visible on the Dashboard (AOS-54234)

ECN marked packet information is not currently visible in the dashboard widget, although it is correctly displayed in the staged view in the probe.

Workaround

Depending on the data, we can apply sorting to the keys to guarantee consistent ordering of the series and modify the predefined dashboard parameters through the Edit menu in the user interface (UI) so that the dashboard displays N rows. But this method only lets us work with a portion of the original series. Modifications must be made on the Metric DB side in order to correctly represent the Top N series across all data. A partial solution is offered by the UI-based workaround until those changes are made, but it does not ensure that we are showing the actual top N series; rather, it only shows a filtered and sorted subset according to the current stage.



Execute CLI Command in the Juniper Device doesn't support ping and traceroute command (AOS-55780)

Execute CLI commands in the Juniper device supported only show and request chassis beacon commands in the Apstra < 6.1.0 environment. Additional commands (ping and traceroute) are introduced in Execute CLI commands in the Apstra >= 6.1.0 environment for easier troubleshooting environments.



External Radius Provider doesn't work in the Apstra Controller when FIPS is enabled (AOS-60612)

The current Radius client in the Apstra Controller adheres to standard RFC 2865 functionality, inherently not FIPS-compliant because it relies on weak, non-compliant algorithms (MD5/MD4) for password hashing and authentication.



Generic system count option is unavailable in Rack Designer (AOS-52227)

In Rack-Type Designer, the ability to specify a Generic System (GS) count, which was available in Rack Builder, is currently not supported. The Rack-Type Designer, introduced in Apstra 4.2.0, replaces the traditional Rack-Type Builder to offer a more intuitive and user-friendly experience. As of Apstra 6.0.0, Rack Builder is deprecated and no longer available.

While Rack Designer does not include all the functionalities of Rack Builder, such as the ability to add multiple generic systems and logical links at once, the new interface provides a superior UX, improved workflow, and new capabilities. Some features may require additional steps, while others have been redesigned or omitted to improve usability.

Users can utilize the Clone functionality in Rack Designer to replicate multiple generic systems as an alternative to the GS count feature.

For further improvement, feature requests may be required via the sales account team.

Workaround

Users can utilize the Clone functionality in Rack Designer to replicate multiple generic systems as an alternative to the GS count feature.



gRPC Periodic Response Timeouts in the Interface Telemetry Service (AOS-56175)

Because the JUNOS device might not send the full snapshot of interface-related data per reporting interval (default interval = 120 seconds) after initial synchronization, Device Telemetry Health detects continuous anomalies for gRPC Periodic Response Timeout in the interface telemetry service in a high-scale environment. The problem still exists even if the reporting interval is extended.

Workaround

It is advised to disable gRPC globally, which forces all gRPC-related telemetry services (interface, MAC) to switch to polling mode rather than gRPC, since the problem may occur at random on several devices.
To disable the gRPC service globally, change grpc_enabled = 0 in the /etc/aos/aos.conf file and then restart AOS service in the Apstra Controller.


[telemetry_global_config] # Python multithreading enable/disable knob for telemetry collection multithreading_config = 1 # Execution timeout for extensible telemetry collectors command_timeout = 120 # Knob to enable/disable gRPC based service collectors grpc_enabled = 0


gRPC Sequence Number Overrun in the MAC Telemetry Service (AOS-56282)

Because the MAC telemetry service uses the gRPCOnChange mode, the device only sends updates after initial synchronization. When the JUNOS device subscribes to the PATH (/network-instances/network-instance/mac-table/entries/entry) for MAC telemetry service, Apstra's gRPC client (Apstra) receives the first full data from two processes (l2ald, l2aldTM). During the initial synchronization, these processes use their own sequence number range (duplicate range), which makes the gRPC client think that the gRPC packets may be dropped internally. Granular sequence number handling will be introduced in 6.1.0 to address the existing sequence overrun issue.

Workaround

It is advised to disable gRPC globally, which forces all gRPC-related telemetry services (interface, MAC) to switch to polling mode rather than gRPC, since the problem may occur at random on several devices.
To disable the gRPC service globally, change grpc_enabled = 0 in the /etc/aos/aos.conf file and then restart AOS service in the Apstra Controller.


[telemetry_global_config] # Python multithreading enable/disable knob for telemetry collection multithreading_config = 1 # Execution timeout for extensible telemetry collectors command_timeout = 120 # Knob to enable/disable gRPC based service collectors grpc_enabled = 0


IBA interface fllapping probe default parameters produce no anomalies (AOS-48525)

The default collection period for IBA interface flapping is only 60 seconds and the flapping threshold is 5 times. But, the default collection period for interface service is 2 minutes, the interface will receive updates every 2 minutes only.

So, the default anomaly window is too small to capture 5 interface flaps. For 5 flaps it should be at least 10 minutes.



Import/Delete configlet task action timed out with BlueprintDiffProducerAgent crash (AOS-59726)

When a configlet action (import/delete) occurs in a blueprint with a large number of configlets, it can fail with a timeout, and the BlueprintDiffProducerAgent process can crash due to a heartbeat timeout. The problem occurs when the blueprint has a configlet with an incorrect Jinja expression via configlet import/delete actions. Failure of rendering configuration with incorrect Jina expressions can cause all configlets to be re-evaluated for all eligible devices, potentially resulting in much longer configlet processing.

Workaround

The below steps can be applied as a workaround. If further assistance is needed, please contact the Apstra support team.

1. In order to recover from the crash of BlueprintDiffProducerAgent, please increase the heartbeat_period to 1200 secs in agent_management section of aos.conf file (<=6.1.0: /etc/aos/aos.conf, >=6.1.1: /user/root/etc/aos/aos.conf) and restart AOS service (service aos restart).


[agent_management] # Override the default heartbeat timeout for agents spawned dynamically by # AgentManager. The value must be a non-negative number. The unit is seconds. # The value 0 is used to turn off heartbeat-based agent timeouts and restarts. # The minimum non-0 value allowed is 60. If not provided, then the default # timeout value (600 seconds) is used. heartbeat_period = 1200

2. Please check each configlet in the blueprint and delete the configlet with incorrect Jinja expression from the blueprint.



In the unassignment serial number of the freeform blueprint, a system fails with the error message "Value is required." (AOS-58505)

Using the system parameters (go to the staging tab of the freeform blueprint, choose the target switch, enter the edit mode for the S/N field, then click the Reset value button) will not allow the Apstra UI to unassign a serial number for a system. The failure of "deploy_mode" will result in the message "Value is required." because the Apstra backend expects the deploy_mode value to be one of the values ("deploy", "ready", "drain", or "undeploy"), but the Apstra UI sends the deploy_mode value as null when the serial number is unassigned.

Workaround

Go to the staged -> Physical -> Systems, select the target switch, and then click the 'Change System IDs assignments' icon. When the Assignment System dialog pops up, click the 'remove assignment' icon and check Deploy Mode as undeploy.



Interface 25G speed configuration not properly rendered for 25g transformation in the Juniper ACX7024 device profile (AOS-57025)

interface 25g speed configuration was not properly rendered over ports 4-27 when 25g transformation is used in the Juniper ACX7024 device profile

Workaround

Add a 25g transformation for ports 4-27 to include interface speed 25g setting.



Junos - BGP CTs for RFC5549 (ipv6 + ipv4-over-ipv6) export route-maps applied twice (AOS-57297)

Junos RFC5549 BGP peer sessions were rendering the same route-map twice on import/export statements. This only applies to 'ipv6-only' bgp peers.

Configurations were rendering:


neighbor a05:fab:192:168:50::254 { description "facing_leaf1-generic"; local-address a05:fab:192:168:50::1; peer-as 65510; family inet { unicast { extended-nexthop; } } family inet6 { unicast; } import ( RoutesFromExt-default-Default_immutable && RoutesFromExt-default-Default_immutable ); export ( RoutesToExt-default-Default_immutable && RoutesToExt-default-Default_immutable ); }

This has been addressed in 6.1.0, where the import and export statements will contain that route-map entry only once.

This may result in a service disruption on upgrade for those BGP peers as Junos generically may reset the peer when it detects any import/export policy reference change.

This will also apply to ipv6-only, non-EVPN blueprints for route-maps such as "LEAF_TO_SPINE_FABRIC_OUT" between all superspine/spine/leaves.



Labels for generic systems may unintentionally change after editing rack operation (AOS-59269)

When rack types are exported, modified, and then re-imported into a blueprint, generic systems in the rack may be unexpectedly renamed, resulting in some servers losing their original user-defined labels even though only rack parameters were changed.

Workaround

Customers should avoid editing the rack type in the Global Catalog UI when they need to preserve generic system names.
(1) Export the rack type from the blueprint to the Global Catalog.
(2) use the API PUT /api/design/rack-types/{rack_type_id} to update only the link_per_spine_speed (e.g., from 25 to 100) directly in the rack type JSON without changing the generic system group_label or count structure.
(3) re-import the updated rack type into the blueprint.



Last Fetched and Last Modified fields of The JUNOS/EVO interface telemetry service, which use gRPC periodic mode, are updated incorrectly (AOS-59984)

The Last Modified field of the interface telemetry service, which uses gRPC periodic mode, is inadvertently updated according to the interface telemetry service's default interval even if no data has been collected from the device. This problem is noticed when gRPC periodic mode is used for Interface telemetry service. The Last Modified field should be updated if a status change is observed, and the Last Fetched field should be updated if the device reports telemetry data.



Liveness anomalies were observed after configuring GCM ciphers on the Junos device (AOS-59895)

After configuring the GCM ciphers [email protected] and [email protected] on the Junos device, the check job began failing because these ciphers are not supported in Apstra 6.0.0. As a result, the cipher exchange between the client and server couldn't be handshaken between the device and Apstra, leading to an SSH negotiation failure leading to connection failure.

Workaround

Apstra 6.1.1 supports GCM ciphers. The user needs to upgrade to Apstra version 6.1.1 to use GCM ciphers between JUNOS devices and Apstra.



MAC Monitor probe could incorrectly report persistent Missing MAC Address anomalies (AOS-62738)

In ESI-based EVPN deployments, the MAC Monitor probe may incorrectly report MAC addresses as missing even though they are present on the device. This can cause Virtual Networks Containing Systems With Missing MAC Addresses anomalies to remain active even after the underlying network issue has been resolved or maintenance has been completed.

Workaround

Edit the MAC Monitor probe and disable the Raise Anomaly option. If required, disable and then re-enable the MAC Monitor probe to clear the existing anomaly state. Keep Raise Anomaly disabled until upgrading to Apstra 6.2.0, as the issue may recur before the fix is applied.



Mac Monitor Probe may show missing MAC count for Router Mac on all VNs in SONiC device (AOS-54607)

The Router MAC (also known as Master Bridge MAC or bridge MAC) on a SONiC device is a unique MAC address assigned to the switch and used as the source MAC address for packets originating from the switch itself, such as those generated by the virtual router or bridge interfaces. The current Mac Monitor probe does not report Router MAC on the originating switch (for example, if leaf1 has MasterBridgeMac as 52:54:00:4e:c6:60, this MAC will show up as a missing MAC on leaf1 for all VNs).
This is a bug related to analytics only; it does not have any network operational impact.



MetricDb migration fails silently for multi-step migration. Eg. 4.2.1 -> X -> Y, data is lost for release Y (AOS-54413)

Upgrade script in 4.2.2 and following releases introduced a regression that manifests itself in multi-step migration scenarios.
Step-1. Upgrade from 4.2.1 -> X (eg. 4.2.2 ) - causes the permission on '/var/lib/aos/metricdb/iba' folder and it's sub-directories to be 700.
Step-2. Upgrade from X (eg. 4.2.2) -> Y (eg. 5.0.1) - causes the silent failure in step that copies '/var/lib/aos/metricdb' folder and ALL its sub-directories to new VM.

The impact of this failure is that the 'Audit', 'IBA stage history' and 'Aos cluster health history' data is lost in the final upgraded AOS instance. The data from the previous release will be lost subsequently if there are further migration steps involved.
This issue affects all the releases starting upgrade from 4.2.2. If your upgrade source Apstra is at least 4.2.2, please apply the workaround suggested BEFORE performing the upgrade.

Workaround

Apply workaround fix(aos_54413_fix_metricdb_permissions.run: https://supportportal.juniper.net/sfc/servlet.shepherd/document/download/069Dp00000Gc0kCIAR) to the old version Apstra Controller Node *BEFORE* every upgrade, following the below steps.

1. Copy the bundle aos_54413_fix_metricdb_permissions.run to the source (old) Apstra controller node. The tool expects Apstra service to be running because it needs to get cluster node information from Sysdb.

2. Make it as executable and execute the bundle as sudo


admin@aos-server:~$ chmod 755 ./aos_54413_fix_metricdb_permissions.run admin@aos-server:~$ sudo ./aos_54413_fix_metricdb_permissions.run Verifying archive integrity... All good. Uncompressing Fix for AOS-54413 for AOS >= 4.2.2 100% AOS[2025-05-25_19:36:22]: Fixing controller node AOS[2025-05-25_19:36:23]: Getting cluster node metadata AOS[2025-05-25_19:36:24]: Fixing worker node: 10.28.75.6 Logs have been collected at: /home/admin/aos_54413_fix_logs_20250525_193623.tar.gz

3. The absence of any errors means that the issue has been fixed. In case of errors during execution, please reach out Juniper Apstra Support Team.

If AOS instances upgraded without work-around and if old apstra VM is preserved, Contact Juniper Apstra support team to help with migrating MetricDB data.



Non-channalized 10GE port link in the ACX7100-32C doesn't come up when adjacent port in the same port group is not explicitly configured as unused (AOS-61624)

On ACX7100-32C devices, configuring a port with 10GE speed (non-channelized) could result in the port failing to link up, accompanied by "Invalid Port Speed Configuration" and "Optics does not support configured speed" alarms. This occurred because the built-in ACX7100-32C device profile did not automatically generate the required "unused" configuration for the adjacent port within the same port group(for example 0 and 1 are in the same group for 10GE)

Workaround

Clone the existing built-in ACX7100-32C Device Profile and update Transformation #7 (10GE non-channelized) to include an unused_port_list configuration for the adjacent interface in the same port group. If you need further assistance, please reach out HPE Apstra Support Team.

Steps:
1. Clone the ACX7100-32C Device Profile to create a custom copy.
2. In the cloned profile, edit port setting for Transformation #7 (10GE non-channelized port) to add unused_interface_list entries for the other port in the same port group (e.g., port 0 and port 1 share a group, port 2 and port 3 share a group, etc.).

Example) Port 0 setting

Old setting:


{"global": {"breakout": false, "fpc": 0, "pic": 0, "port": 0, "speed": ""}, "interface": {"speed": "10g"}, "validations": [{"constraint": "1x25or1x10", "port_group": "P0_P1"}, {"constraint": "no_constraint", "port_group": "P0_P1_P2_P3"}]}

New setting: add "unused_interfaces_list" key with value ["et-0/0/1"] for adjacent port.


{"global": {"breakout": false, "fpc": 0, "pic": 0, "port": 0, "speed": ""}, "interface": {"speed": "10g","unused_interfaces_list": ["et-0/0/1"]}, "validations": [{"constraint": "1x25or1x10", "port_group": "P0_P1"}, {"constraint": "no_constraint", "port_group": "P0_P1_P2_P3"}]}

3. Create new Interface Map with the updated Device Profile.

4. Assign the new cloned device profile into the managed devices and import the new IM into blueprint
5. Assign the new IM into the deployed devices



NOS Upgrade for Juniper device fails when the configuration line in the pristine configuration extends into more than one line (AOS-53602)

The NOS upgrade procedure must parse the pristine configuration in order to determine which ports must be disabled when the Skip Shutting Down Interface During Upgrade option is not checked in the Advanced Settings of Managed Device. The NOS upgrade would fail if the configuration line in the pristine configuration extended into multiple lines because parsing the command line misses the end-of-command-line character (.

Workaround

Recommend re-onboarding of the device with a clean, pristine configuration.



NOS upgrade with JSU(Junos Selective Update) image for JUNOS and EVO device fails with device in pristine configuration status (AOS-59869)

Apstra anticipates that the NOS upgrade will result in a version change after the NOS device image installation and device reboot with the new image. However, JSU (Junos Selective Update) installs only selected packages without changing the version, followed by a restart process rather than a device reboot. Therefore, NOS upgrade by Apstra using JSU image will fail with the post-validation check (pre-install version != post-install version and image filename must include version information).

Workaround

Please use the CLI to upgrade JSU instead of utilizing Apstra's NOS upgrade, or run Apstra NOS upgrade (make sure the JSU image filename includes version information) and then execute a full push configuration when the NOS upgrade fails due to post-validation check.



optical_xcvr Telemery service fails with "show error" on GPU Systems (AOS-58568)

When the Optical Transceivers Probe is turned on, the optical_xcvr telemetry service is enabled for all systems with on-box agents (including GPU servers) or off-box agents. Because the optical_xcvr telemetry service is not designed to run on GPU servers, it fails to collect optical transceiver information and returns an error message.

Workaround

The issue can be resolved by correcting the graph query of the Optical Xcvr Stats process (adding system_type='switch') in the Optical Transceivers Probe to prevent the telemetry service from running on the GPU servers. Please modify the graph query as shown below.


node("device_profile", name="device_profile") .in_("device_profile") .node("interface_map") .in_("interface_map") .node("system", system_id=not_none(),system_type='switch' , deploy_mode=is_in(["deploy", "drain"]), name="system")


PFE(Packet Forwarding Engine) in the QFX5120 platform restarts during NOS upgrade (AOS-57490)

When the Junos EVPN Next-hop and Interface count maximums parameter in the staged->Fabric settings->Fabric-policy is enabled, Apstra introduced modifying the default hardware settings for VXLAN routing's resource (next-hop and interfaces) for QFX5110, QFX5120, EX4650, and EX4400 devices in the rendered configuration () starting with version 4.2.0. Whenever configuration changes in VXLAN routing's resource, JUNOS triggers PFE automatic restarts to reflect new changes with service impact. The typical scenarios would be when the device becomes deployed, undeployed, or the device is in NOS upgrade. To prevent unnecessary PFE restarts in those scenarios, the configuration for VXLAN routing's resource needs to be included in the pristine configuration.

Workaround

If the Junos EVPN Next-hop and Interface count maximums parameter in the staged->Fabric settings->Fabric-policy is enabled, add the below configuration into the device's pristine configuration.
QFX5120 and EX4650 VXLAN routing's resource


forwarding-options { vxlan-routing { next-hop 45056; interface-num 8192; overlay-ecmp; } }

QFX5110 VXLAN routing's resource


forwarding-options { vxlan-routing { next-hop 32768; interface-num 8192; overlay-ecmp; } }

EX4400 VXLAN routing's resource (add overlay-ecmp if Junos EX-Series Overlay ECMP is also enabled)


forwarding-options { vxlan-routing { next-hop 16384; interface-num 6144; overlay-ecmp; } }


Pristine Config Update Fails with "System Already Parsed" Error for Junos Devices (AOS-52788)

Customers may encounter the following Server-side Validation Error in the Web UI when the pristine configuration contains multiple system stanzas which is not a expected behvavior:


"Cannot parse config: system already parsed."

According to ScotchInventoryAgent logs, POST requests to update the pristine configuration failed with a 422 Unprocessable Entity error, indicating a validation issue:


2025-02-17 23:31:18,730 680:INFO:aos.scotch.libs.scotch_flask:request: POST /api/systems/AN10555621/pristine-config HTTP/1.0 34074 bytes 2025-02-17 23:31:18,737 680:INFO:aos.scotch.libs.scotch_flask:response: 422 55 bytes 0.007347 seconds

Background of the issue:


1. In Apstra 4.2.x, gRPC was introduced to support Telemetry Streaming, and as a result, having two system blocks in the pristine configuration was expected in that release. 2. Starting from Apstra 5.0.0, enhancements were made to automatically merge multiple system stanzas in the pristine configuration during the NOS upgrade process. 3. If a customer chooses to remain on their current NOS version for an extended period, multiple system stanzas can exist in the pristine configuration without causing issues.
Workaround

If a customer chooses to remain on their current NOS version for an extended period and needs to forcefully update the pristine configuration, they should manually merge the system stanzas within the pristine configuration using the UI and then perform a Force Update.

For further assistance, please contact Juniper Apstra Support.



Streaming doesn't work correctly on large scale topologies, resulting in the high rate of dropped messages (AOS-54117)

Receivers may report a significant number of errors in their statistics due to the possibility of high streaming data being dropped in high-scale environments, such as 8K/16K GPU AI/ML topologies.

Workaround

There is a configuration parameter in the /etc/aos/aos.conf file to control the number of messages for streaming that can be queued before sending them to the TCP socket. Please adjust the value for tcp_queue_size in the streaming section in the /etc/aos/aos.conf file to mitigate the issue and then restart aos service by executing sudo service aos restart.


[streaming] tcp_queue_size = 40000. # default value is 20000.


Sustained Execution Failure Anomalies in Device Telemetry Health for Virtual Infra (AOS-54391)

When a port group in the vCenter is configured with a private vlan mode, the VLAN specification contains the pvlanid property instead of the vlanid property. However, the vlanid property is always expected from the port group's vlan specification by the Device Telemetry Agent's collector if it is not trunk mode. Anomalies could be reported if the collector's execution fails due to a reference to the vlanid property, which is nonexistent.

Workaround

Recommend not using port group with private vlan in the Virtual Distributed Switch.



Sustained Optical Threshold Anomaly reoccurrence in the JUNOS device (AOS-54497)

Sutatined Optical Threshold anomalies are frequently observed over disconnected (not connected) interfaces in the JUNOS device. The main reason for the problem is that the JUNOS device sometimes reports the received average power as a very low value or - Inf value when the interface is disconnected, and Apstra does not correctly parse the - Inf value. When a very low power value falls below the warning level, Apstra creates an anomaly for the low receive power port. But when the same interface reports a - Inf value that is later incorrectly parsed, Apstra eliminates the interface from the list of interfaces with an optical transceiver, thereby resolving the raised anomaly falsely. This is the reason anomalies are frequently raised and cleared over the same interface. No workaround is available for this issue

Workaround

None



Unintended advertisement of all fabric VTEP loopbacks to external routers in default routing zone in VXLAN DCI environment (AOS-54864)

Integrated DCI feature(vxlan stitching) was introduced in Apstra 4.2.0. In versions 4.2.x and later, customers using this feature may encounter an issue where all VTEP loopback addresses from the fabric including those from non-border leaf devices are being advertised to external routers over BGP in the default routing zone.

This affects only VXLAN DCI Stitching deployments(Stitching requirement for VTEP loopbacks for only border leaf nodes vs OTT requirement for all VTEP loopbacks in the fabric). Even when customers configure routing policies to export only loopback of border leaf nodes, Apstra backend logic automatically includes all VTEP loopbacks. Due to the current design, Apstra does not differentiate between border and non-border leaf roles in this context, resulting in the unintended advertisement of all fabric loopbacks to external peers.

Engineering has confirmed this as a bug. The expected behavior is to advertise only the loopback addresses of border leaf switches to external routers in the default routing zone. There is no official workaround to modify this behavior through standard configuration. Engineering is actively working on a fix to address this issue in a future release. The only option is to use a custom configlet to override Apstra default export logic. Please reach out to Apstra Technical Support for assistance.



User-defined import/export route-targets raise validation errors for ":0" suffixes (AOS-59842)

User-defined import / export route-targets with ":0" are rejected with validation errors.


{ "rt_policy": { "import_RTs": { "0": "Type 0 RD X:Y must be in format 2-byte ASN:4-byte value. Provided value: \"65500:0\"" } }


Virtual Infra manager's information not cleared even if virtual infra manager is removed from Apstra Controller (AOS-53538)

When the virtual infra manager is removed from the Apstra controller, Apstra should have cleared any data related to the virtual infra manager. Because it's not cleared, when the same virtual infra manager is added back to Apstra later, old data is still used together with the new collected data from the virtual infra manager's collector. In some scenarios, when old, uncleaned data has an error condition, it can trigger continuous error even if newly collected data doesn't have an error condition.

Workaround

If the virtual infrastructure manager requires re-onboarding (removing and then adding back) from Apstra, the user must take the actions listed below.
1. Remove virtual infra manager from Apstra Controller (External Systems/Virtual Infra Managers).
2. Restart the AOS service.
3. Add the virtual infra manager back to to the Apstra



Virtual Network Endpoint View in the UI showing empty information (AOS-53783)

Since the UI misses polling of the node detail information to the Apstra backend, the Virtual Network Endpoints view of Generic System Node (Staged > Physical > Topology > Virtual Networks Endpoints) shows empty information.

Workaround

Refresh Web page in the browser to make the UI send requests explicitly to collect data



VirtualInfra telemetry service shows error message when a PNIC is unassigned from the VDS (Virtual Distributed Switch) (AOS-53537)

Apstra creates a relationship between the PNIC and the Link Discovery Policy (which determines which discovery protocol is used) configured in the VDS when a PNIC is assigned to a VDS (Virtual Distributed Switch). One PNIC may inadvertently become linked to two relationships without clearing out the previous relationship when a user moves a PNIC directly from one VDS to another VDS. An error message below appears when the PNIC becomes unassigned from VDS because there is more than one relationship between the PNIC and Link Discovery Policy that is invalid.


virtual_infra failed to collect data, plugin raised exception: {'item_iter': <aos.sdk.graph.graph.RelationshipIterator object at 0x7f60c03f9570>, 'items': [df57a2f0-969c-4dee-9831-a53526bd7d5a-[:policy]->4c8c4c31-df9a-4188-933f-6b5d67703a1f, df57a2f0-969c-4dee-9831-a53526bd7d5a-[:policy]->1787d662-4a41-4b8b-9b77-acac5508e771]}
Workaround

Instead of performing one direct migration action from one VDS to another, the problem can be avoided by two actions: unassigning the PNIC from the old VDS and then assigning it to the new VDS.

Procedures for fixing the errors as a workaround
1. Remove the virtual infra manager from not only the blueprint but also the External Systems/Virtual Infra Managers.
2. Restart the AOS service.
3. Add the virtual infra manager back to the External Systems/Virtual Infra Managers and then blueprint.



VMs Without Fabric Configured VLANs probe raise anomalies when multiple vNICs from a VM are assigned to the same vNET (AOS-54088)

When multiple vNICs from a single VM are assigned to the same vNET (port group or VDS) in the Virtual Infra, the VMs Without Fabric Configured VLANs probe raises anomalies in the Analytics->Anomalies because the graph query in the VMs backed by Fabric VLANs processor treats those vNICs as identical.

Workaround

Please clone existing VMs Without Fabric Configured VLANs probe with a different name, and then modify graph query in the VMs backed by Fabric VLANs processor to include vnic into the existing distinct statement like the below. If further assistance is needed, please contact Juniper Apstra Support Team.


Graph Query: match( node('system', name='server', role='generic', management_level='unmanaged', external=False) .out('hosted_interfaces') .node('interface', name='server_intf') .out('hosted_vn_endpoints') .node('vn_endpoint', name='vn_endpoint') .in_('member_endpoints') .node('virtual_network') .out('instantiated_by') .node('vn_instance', name='vn_instance') .having( node(name='vn_instance') .in_('hosted_vn_instances') .node('system', system_id=not_none(), deploy_mode='deploy') .out('hosted_interfaces') .node('interface') .out('link') .node('link') .in_('link') .node('interface') .in_('hosted_interfaces') .node(name='server'), at_least=1 ), node(name='server') .in_('is_realized_by') .node('hypervisor', name='hv'), node(name='hv') .out('hosts') .node('vm', name='vm') .out('has') .node('vnic', name='vnic') .out('part_of') .node('vnet', vn_type='vlan', name='vnet') ) .distinct(['server', 'vn_endpoint', 'vm', 'vnic']) .where(lambda vnet, vn_instance, vn_endpoint: (vnet.vlans == [0] and vn_endpoint.tag_type == 'untagged' or vn_instance.vlan_id in vnet.vlans))


When JUNOS configlet for Set/Delete has jinja comments, rendering configlet fails with message, Junos-based cli configuration commands must start with either \"set\" or \"delete\" (AOS-59416)

Because the Jinja comment is not recognized as a valid Jinja expression during rendering, it is submitted as normal commands to the device, causing the JUNOS/EVO device to reject the invalid commands and deployment to fail.

Workaround

To make the template "Jinja-aware", the user needs to include a no-op Jinja construct at the top of the configlet. This will bypass the stringent line-by-line set/delete checks. The workaround entails introducing a dummy control block or expression directly after the Jinja comments.

configlet example with error


{# v1.0 - 22 Jan 2026 - Author: ... - Initial version #} {# Objective: Example of non-working configlet #} set system time-zone Europe/Luxembourg

configlet example with workaround


{# v1.0 - 22 Jan 2026 - Author: ... - Initial version #} {# Objective: Example of working configlet #} {% if 1 > 0 %}{% endif %} set system time-zone Europe/Luxembourg


Worker nodes may remain in a failed configuration state after connectivity issues, requiring manual intervention for recovery (AOS-60167)

With current cluster design, worker nodes may fail to recover their configuration state after a temporary loss of SSH connectivity to the controller. This condition can be observed in the UI by navigating to Platform > Apstra Cluster > Nodes > Worker, where the following error may be displayed:


Configuration Error: ssh: connect to host 10.28.17.4 port 22: Connection refused

When the controller (ClusterManagerAgent) attempts to push configuration to worker nodes, it retries SSH connections up to three times. If all attempts fail, due to connection refused or no route to host, the node is marked with a FAILED Configuration State. Once this state is set, the system does not automatically retry configuration, even if SSH connectivity is later restored. Although worker nodes may resume sending keepalives and transition back to an active operational state, the configuration state remains in failed state indefinitely. Due to this the overall node state may continue to appear FAILED despite restored connectivity.

Workaround

Manually trigger a configuration synchronization using one of the following methods:


1. Navigate to Platform > Developers > REST API Explorer 2. Execute REST API: POST /api/cluster/worker/sync [OR] 1. Restart AOS from the controller VM: systemctl restart aos



Known Third-Party Issues

 

ACX platform doesn't support export sflow over mgmt instance (AOS-58680)

When sFlow collector is configured with mgmt_instance, the configuration will be ignored in the ACX platform with a warning such as the below example.

sflow {
polling-interval 10;
sample-rate {
ingress 10000;
egress 10000;
}
source-ip 10.217.6.15;
collector 10.217.0.165 {
udp-port 6343;
##
## Warning: statement ignored: unsupported platform (ACX7024X)
##
routing-instance mgmt_junos;

Workaround

Please use a non-management instance for exporting sFlow until the ACX platform supports a management instance for sFlow export.



Anomalies are rasied for interfaces on Juniper EX4400-48T devices running JUNOS 22.4R3 (AOS-56571)

Anomalies are raised due to mismatch in the operational status of interfaces due to interface status showing "unknown" on Juniper EX4400-48T devices running Junos 22.4R3.

Workaround

Restart the Apstra AOS service to collect the right interface status information. This issue is not observed in higher JUNOS versions. Recommend upgrading to an Apstra-qualified higher JUNOS version (>=23.4R2-S4).



BFD underlay flap in the JUNOS device after commit (AOS-58200)

When the Apstra commit was executed, the JUNOS device with the off-box agent reported BFD underlay flaps between the committed device and the other devices. Apstra currently uses the load override option as the default action for device commits, which may result in high CPU utilization, preventing time-sensitive daemons from acquiring CPU time slices and triggering unexpected events such as BFD timeouts or writing EEPROM errors. Starting with 6.1.0, Apstra intends to use load update as the default commit action for JUNOS and EVO devices.

Workaround

The workaround to use load update can be applied only to the JUNOS/EVO offbox agent. Please follow the below step to apply workaround
1. Navigate into Devices->Managed Devices->{DEVICE_IP}-> Agent
2. Click Edit button to edit Agent
3. Add an option into Open Options with the key as load_mode and the value as update in the Edit Offbox System Agent(s) window.
4. Click Update button



Cisco N9K-C93600CD-GX rollback fails when using breakouts (AOS-47891)

The NXOS rollback feature on the N9K-C93600CD-GX device has significant limitations when the devices' interfaces are broken out.
Ports 1-24 in the model are organized into four-port groups: (1, 2, 3, 4), (5, 6, 7, 8), (9, 10, 11, 12), (13, 14, 15, 16), (17, 18, 19, 20), and (21, 22, 23, 24). When port 1 is broken out as 4x10G or 4x25G, port 3 is automatically broken out in the same mode, and vice versa. When any port in the quadruple is split into 2x50G, all four ports are automatically split in the same mode. Similarly, ports 26-28 are organized in pairs of two, i.e. (25, 26) and (27, 28). Both ports in the pair must operate in the same breakout mode.

In most cases where a breakout (or more than one) exists, rollback fails to generate a working rollback patch. The reason for this is that the breakouts cannot be reversed if the remaining broken-out interfaces in the same port group have not been shutdown first. For example, to negate the breakout of port 1, the broken-out interfaces of port 3 must be shutdown, and vice versa. It appears that the rollback logic shuts down the interfaces associated with the port whose breakout is being reverted (port 1 in the previous example), but fails to shut down other broken-out ports in the same port group (port 3).

Workaround

The safer way for the N9K-C93600CD-GX to be used with AOS is for the customer to avoid using breakouts altogether on the device.
No issue with rollback when ports 29-36 have been broken out has been observed. Breakouts on these ports can be rolled back
In the case that the last interface of a port-group is the only one used and broken out, would the nxos rollback feature (and rollback to pristine) be successful. However this is highly discouraged
In any other case the only way to reverting to pristine would be to manually shudtown all broken down interfaces before reverting to pristine (or using the rollback to a pristine config)



gRPC Sequence Overrun may causes JSD OOM(Out Of Memory) crashes and device reboots on Junos and Junos EVO Devices (AOS-61246)

Apstra may encounter gRPC sequence number overrun for MAC telemetry service, recognized as losing data and initiates re-subscription for service. This sequence can continue repeatedly in a highly loaded environment (>=100K entries), making the JSD to handle continuous subscription and cancel subscription requests with memory leaking. This continous accumulation of memory leaking may lead into process crash by OOM and then triggering device reboot.

Workaround

The workaround is to disable gRPC in the Apstra. If the customer wants to continue to use gRPC in the Apstra, recommed upgrading to the latest 6.1.X release (which includes fixes for the sequence overrun misleading issue) so that meory leaking can be prevented. Please reach out Juniper Apstra Support team for the further assistance.



gRPC server reset count anomalies in the JUNOS-EVO platform (AOS-53526)

gRPC server reset count anomalies are observed in the JUNOS-EVO platform when gRPC Max Client connection limit error occurs in the device due to the problem that gRPC stalled connections are not cleared. gRPC keepalive is not enabled by default on the JUNOS-EVO platform running 22.2R3 or 22.4R3, which is the cause of the problem. gRPC keepalive is enabled for 300 seconds in the >=23.4R2-EVO release to avoid a build-up of stalled gRPC connections.

Workaround

In JUNOS-EVO device running 22.2R3 or 22.4R3, apply the below configuration via configlet into the device to enable gRPC keepalive or upgrade the device to >=23.4R2-EVO. For further assistance, please contact the Juniper Apstra Support Team.


set system services extension-service request-response grpc grpc-keep-alive 300


IBA Probe Interface Queue Stats not reporting correct ingress utilization (AOS-52617)

Apstra 6.0.0 introduced a new IBA probe, Interface Queue Stats, which provides detailed insights into ingress and egress buffer utilization on a per-queue basis, along with other relevant metrics. This IBA probe is designed for AI/ML-based fabrics, particularly those using rail-based blueprints. However, customers using Junos EVO versions earlier than 23.4R2.X100-D31 will notice that the ingress buffer utilization is reported as 0. This issue arises from a bug in EVO devices, where ingress buffer utilization is not exported by default. The bug affecting ingress buffer utilization is resolved in the 23.4R2.X100-D31 release of Junos EVO.

Workaround

To enable correct reporting of ingress buffer utilization, customers need to create a configlet with the below configuration for each line card slot used in their fabric.


set chassis fpc <fpc-id> traffic-manager buffer-monitor-enable


Juniper EVO When Configured With DHCP Relay as Border Leaf Role, DHCP Packets May Be Discarded (AOS-43348)

When Juniper EVO device hosts DHCP servers in a border leaf role with DHCP relay configuration, DHCP may not work as intended due to an unresolved bug in Junos EVO which prevents DHCP packets from being processed correctly. Please refer to the following KB for dhcp relay limitations: https://supportportal.juniper.net/s/article/Juniper-Apstra-Support-for-Stateless-DHCP-Relay?language=en_US

Workaround

Using Apstra's configlet feature, create configlet to remove the rendered DHCP relay configurations and apply it to the Juniper EVO border leaf device.



JUNOS: EVPN MACs are limited to 200 IPs per MAC for bridge domain by default (AOS-57099)

A new default limit for the number of IP addresses per MAC per bridge domain in EVPN (mac-ip-limit) was added in Junos and Junos Evolved Releases 24.2R1, 23.4R2, and 23.2R2. 200 IPs per MAC is the default setting.

Clients who use EVPN fabrics, such as those with MAC-VRF deployments in fabrics managed by Apstra, might observe that MACs linked to more than 200 IP addresses cease to learn new IPs in the EVPN MAC-IP table. The Junos software release introduced this expected behavior.

The fabric can support more IPs per MAC while preserving per-bridge-domain enforcement by setting mac-ip-limit globally using an Apstra Configlet. No software fix is required.

Workaround

To adjust the limit in an Apstra-managed fabric, create a Configlet in Apstra with the below command, specifying the desired limit:


set protocols evpn mac-ip-limit <desired-value>

Import the Configlet into the blueprint and apply it to the relevant switches.

Note: Although the command is global in Junos, the limit is enforced per MAC per bridge domain, including inside MAC-VRFs.



Manual Reboot Required for "shared-tunnels" Configuration Following Junos Upgrade (AOS-45139)

In the Apsta 4.2 reference design change for MAC-VRF, the Junos "forwarding-options evpn-vxlan shared-tunnels" configuration is added via the Apstra rendered configuration. However, this command requires a device reboot to take effect with the Junos warning "Config: forwarding-options evpn-vxlan shared-tunnels has changed. A system reboot is mandatory". A user doing a Junos upgrade with Apstra may re-experience this issue after the device is upgraded.

Workaround

To avoid the need to a additional, manual reboot after a device Junos upgrade, the user can add the following configuration to the Apstra device system-agent pristine-configuration.


forwarding-options { evpn-vxlan { shared-tunnels; } }

This can be done in the "Decvices / Managed Devices / Pristine Configuration" Apstra UI or using the Apstra-CLI "system pristine_config_append" command.



QFX10002-36Q devices may not raise power supply anomalies due to inconsistent information from Junos (AOS-54513)

For power supplies, Apstra primarily uses the output of the show chassis environment pem or show chassis environment psm command. On the QFX10002-36Q, both PEMs are reported, however, only one includes the XML tag that designates the component class as Power. The power supply information is further augmented using the show chassis environment command. Since the tag is missing for one PEM, Apstra does not recognize it as a valid power supply component. This is a known issue in Junos. Below is an example of the XML output from the show chassis environment without the class tag:


<environment-item> <name>FPC 0 Power Supply 1</name> <status>Present</status> </environment-item>

Due to inconsistent Junos behavior, the Power Supply State Check processor of Apstra does not evaluate the affected PEM, and no anomaly is raised. No workaround is available for this issue. This issue is expected to be addressed in newer Junos versions from 24.4R2, and the corrected behavior is expected to be present in supported versions for Apstra 6.1.0 and later.

Workaround

None



sFlow export through the management interface does not work in Junos EVO version 23.4R2-S5-EVO (AOS-57740)

In Junos EVO version 23.4R2-S5-EVO and 23.4R2-S6-EVO, devices do not export sFlow packets through the management interface. This issue affects all QFX device models.

Workaround

You can resolve this issue using one of the following approaches:

1. If you require sFlow export through the management interface, please use the Apstra-qualified Junos EVO release 23.4R2-S4-EVO instead of 23.4R2-S5-EVO.

2. If you are using Junos EVO release 23.4R2-S5-EVO, configure sFlow to export through a non-management (in-band revenue) interface instead of the management interface.

Modification History

2026-07-23: Added AOS-62738

2026-06-08: Added AOS-61624

2026-05-26: Updated AOS-59416, Added AOS-61246

2026-05-15: Updated AOS-54864

2026-04-21: Added AOS-52788

2026-04-14: Added AOS-60612

2026-04-13: Added AOS-59984

2026-04-03: Added AOS-55780

2026-03-23: Added AOS-60167,AOS-60183

2026-03-18: Removed AOS-51029,Updated AOS-59821,Added AOS-59842,AOS-60114

2026-03-05: AOS-59821,AOS-59931

2026-03-03: AOS-47430,AOS-59726,AOS-59869,AOS-59895

2026-02-18: Added AOS-54971

2026-02-11: Added AOS-59269

2026-02-03: Added AOS-59198,A0S-59416

2026-01-08: Added AOS-48525,AOS-58505

2025-12-11: Added AOS-58568,AOS-58680

2025-12-05: Added AOS-44623,AOS-57025

2025-11-19: Updated AOS-57740

2025-11-18: Updated AOS-57740, Added AOS-58200

2025-11-10: Updated AOS-57490

2025-10-28: Added AOS-57740

2025-10-02: Added AOS-56571, AOS-57297, AOS-57315, AOS-57490

2025-09-17: Added AOS-57099

2025-09-08: Added AOS-56498

2025-08-26: Updated download link for AOS-54413, AOS-55184, AOS-55286, AOS-55673. Updated AOS-55673

2025-08-20: Added AOS-56175, AOS-56282

2025-08-13: Added 40023, AOS-43348, AOS-45139

2025-07-29: Added AOS-52519

2025-07-23: Added AOS-55673

2025-07-10: Updated RFE-3429 with SONiC 4.4.2, Added AOS-55184

2025-07-09: Added AOS-55286

2025-06-26: Added AOS-54006

2025-06-25: Added AOS-54607,AOS-54767,AOS-54864

2025-06-09: Added AOS-54497, AOS-54513

2025-05-27: Initial publishing