Description

> The customer attempted to upgrade the switch to 22.2R3, but it remained unresponsive for about 3 hours.

> No syslog logs were generated during this time, and there was no console output.


Symptoms

> From the host logs, it appears that the switch could have been in a possible loop after the switch upgrade was initiated.


025-06-18T09:10:16.233349+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_setup: Connecting to RE 192.168.1.2 asynchronously

2025-06-18T09:10:19.345233+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:10:22.417214+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:10:28.497120+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:10:31.569118+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:10:31.569277+00:00 RBIB14ER01-node vehostd[4392]: tcp_client_start_retry_timer: Failed to connect to the server after 3 retries

2025-06-18T09:10:31.569413+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_event_handler: Connection failed with status 1. Retrying to inform RE 192.168.1.2

2025-06-18T09:10:31.569541+00:00 RBIB14ER01-node vehostd[4392]: vehostd_retrier_retry: This is 20 th try. Firing retry after 24 seconds

2025-06-18T09:10:55.593163+00:00 RBIB14ER01-node vehostd[4392]: connect_retry: Retrying node alive connection to RE 192.168.1.2

2025-06-18T09:10:55.593413+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_setup: Connecting to RE 192.168.1.2 asynchronously

2025-06-18T09:10:58.705113+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:01.194437+00:00 RBIB14ER01-node CROND[32423]: (root) CMD (/etc/cron.daily/logrotate)

2025-06-18T09:11:01.198504+00:00 RBIB14ER01-node crond[537]: sendmail: account default not found: no configuration file available

2025-06-18T09:11:01.209430+00:00 RBIB14ER01-node CROND[32422]: (root) MAIL (mailed 88 bytes of output but got status 0x004e#012)

2025-06-18T09:11:01.777217+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:07.857140+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:10.929116+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:10.929278+00:00 RBIB14ER01-node vehostd[4392]: tcp_client_start_retry_timer: Failed to connect to the server after 3 retries

2025-06-18T09:11:10.929415+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_event_handler: Connection failed with status 1. Retrying to inform RE 192.168.1.2

2025-06-18T09:11:10.929543+00:00 RBIB14ER01-node vehostd[4392]: vehostd_retrier_retry: This is 21 th try. Firing retry after 24 seconds

2025-06-18T09:11:34.953068+00:00 RBIB14ER01-node vehostd[4392]: connect_retry: Retrying node alive connection to RE 192.168.1.2

2025-06-18T09:11:34.953329+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_setup: Connecting to RE 192.168.1.2 asynchronously

2025-06-18T09:11:38.065117+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:41.137104+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:47.217220+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:50.289147+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:11:50.289303+00:00 RBIB14ER01-node vehostd[4392]: tcp_client_start_retry_timer: Failed to connect to the server after 3 retries

2025-06-18T09:11:50.289438+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_event_handler: Connection failed with status 1. Retrying to inform RE 192.168.1.2

2025-06-18T09:11:50.289565+00:00 RBIB14ER01-node vehostd[4392]: vehostd_retrier_retry: This is 22 th try. Firing retry after 24 seconds

2025-06-18T09:12:01.210497+00:00 RBIB14ER01-node CROND[32432]: (root) CMD (/usr/sbin/logrotate /etc/logrotate.conf > /dev/null 2>&1)

2025-06-18T09:12:01.210911+00:00 RBIB14ER01-node CROND[32433]: (root) CMD (/etc/cron.daily/logrotate)

2025-06-18T09:12:01.216330+00:00 RBIB14ER01-node crond[537]: sendmail: account default not found: no configuration file available

2025-06-18T09:12:01.237396+00:00 RBIB14ER01-node logrotate: ALERT exited abnormally with [1]

2025-06-18T09:12:01.237654+00:00 RBIB14ER01-node CROND[32431]: (root) MAIL (mailed 156 bytes of output but got status 0x004e#012)

2025-06-18T09:12:14.313229+00:00 RBIB14ER01-node vehostd[4392]: connect_retry: Retrying node alive connection to RE 192.168.1.2

2025-06-18T09:12:14.313669+00:00 RBIB14ER01-node vehostd[4392]: node_alive_connect_setup: Connecting to RE 192.168.1.2 asynchronously

2025-06-18T09:12:17.425120+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:12:20.497153+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:12:26.577224+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

2025-06-18T09:12:29.649209+00:00 RBIB14ER01-node vehostd[4392]: Connection failed. Will retry after a while

 

Which was then auto-cleared after the VM-HOST had rebooted.

 

2025-06-18T12:15:59.452263+00:00 RBIB14ER01-node vehostd[4392]: save_reboot_reason: vm_reset_reason change to 48

2025-06-18T12:15:59.452382+00:00 RBIB14ER01-node vehostd[4392]: vjunos_pending: [vjunos0] Rebooting system...

2025-06-18T12:16:00.451611+00:00 RBIB14ER01-node vehostd[4392]: Rebooting Syestem from Hypervisor

2025-06-18T12:16:00.460242+00:00 RBIB14ER01-node systemd[1]: Stopping SNTPC service...

2025-06-18T12:16:00.460585+00:00 RBIB14ER01-node systemd[1]: Stopped target Timers.

2025-06-18T12:16:00.460815+00:00 RBIB14ER01-node systemd[1]: Stopped Daily Cleanup of Temporary Directories.

2025-06-18T12:16:00.461103+00:00 RBIB14ER01-node systemd[1]: Stopped target System Time Synchronized.

2025-06-18T12:16:00.461439+00:00 RBIB14ER01-node systemd[1]: Stopping Virtual Machine and Container Registration Service...

2025-06-18T12:16:00.472278+00:00 RBIB14ER01-node systemd[1]: Stopping Virtual machine log manager...

2025-06-18T12:16:00.472645+00:00 RBIB14ER01-node systemd[1]: Removed slice system-serial\x2dgetty.slice.

2025-06-18T12:16:00.472922+00:00 RBIB14ER01-node systemd[1]: Stopping vehostd service...

 

> The fault resolved itself following the VM reboot, without any user intervention. The switch has been stable since.

Solution

No action is required, if the switch remains stable.

Modification History

2025-06-25 : Article Created