Description

A QFX10002 that is experiencing SSD errors may reboot when collecting the RSI or any PFE-related commands. 

Symptoms

SSD errors are logged on the console and the Host OS is inaccessible. 

 

The below SSD error logs will be observed on the console:

end_request: I/O error, dev sda, sector 8970517

end_request: I/O error, dev sda, sector 8970517

end_request: I/O error, dev sda, sector 8970517

 

The HOST OS will be inaccessible but reachable:

% vhclient -s

rcmd: 192.168.1.1: Connection reset by peer

 

> ping 192.168.1.1 interface em2.32768    

PING 192.168.1.1 (192.168.1.1): 56 data bytes

64 bytes from 192.168.1.1: icmp_seq=0 ttl=64 time=0.304 ms

64 bytes from 192.168.1.1: icmp_seq=1 ttl=64 time=0.139 ms

 

Attempting to collect any stats or logs from PFE may result in the FPC reload:

Mar 6 18:44:15 xcr01.mil01 kernel: peer_input_pending_internal:[5256] VKS1 for peer type 27 indx 0 reported a sb_state 32 = SBS_CANTRCVMORE

Mar 6 18:44:15 xcr01.mil01 kernel: peer_input_pending_internal:[5256] VKS1 for peer type 27 indx 0 reported a sb_state 32 = SBS_CANTRCVMORE

Solution

A power cycle is required to clear this potentially transient error state. 

Modification History

1st version