+cc Bobby and Michal who touched this code recently On Tue, Sep 29, 2026 at 06:33:04PM +0200, Jerome Mohm via B4 Relay wrote:
From: Jerome Mohm <[email protected]>vsock_bpf_recvmsg() takes lock_sock(sk) and holds it across the whole receive loop, including vsock_msg_wait_data(), which sleeps in wait_woken() without dropping the lock. When data arrives, the transport's delivery context (the vsock-loopback worker, or the virtio/vhost rx path) calls virtio_transport_recv_pkt() -> lock_sock() on the same socket and blocks. That context is the only one that would queue the skb and call sk_data_ready() to wake the reader, so neither side makes progress. A blocking recv() hangs until a signal arrives (a finite SO_RCVTIMEO also breaks it); while it lasts the shared delivery worker is stalled, so all vsock rx on that transport stops, not only the affected socket. The hung-task watchdog reports the worker blocked in D state: INFO: task kworker/1:3:107 blocked for more than 5 seconds. task:kworker/1:3 state:D Workqueue: vsock-loopback vsock_loopback_work Call Trace: __schedule schedule __lock_sock lock_sock_nested virtio_transport_recv_pkt vsock_loopback_work The reader holds the same sk_lock-AF_VSOCK it is waiting on, taken in vsock_bpf_recvmsg(). This code is based on net/unix/unix_bpf.c, whose unix_msg_wait_data() drops u->iolock around wait_woken() and re-takes it afterwards. The vsock port substituted lock_sock() for that serialisation lock but omitted the unlock and relock. tcp_bpf and the native vsock_connectible_wait_data() both drop the lock across the wait; vsock_bpf is the only one that does not. Release the socket lock around the wait and re-acquire it before re-checking for data, so the caller's locking is unchanged. Fixes: 634f1a7110b4 ("vsock: support sockmap") Cc: [email protected] Assisted-by: LLM Signed-off-by: Jerome Mohm <[email protected]> --- net/vmw_vsock/vsock_bpf.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/net/vmw_vsock/vsock_bpf.c b/net/vmw_vsock/vsock_bpf.c index 9049d2648646..bb7d81a95baa 100644 --- a/net/vmw_vsock/vsock_bpf.c +++ b/net/vmw_vsock/vsock_bpf.c @@ -50,7 +50,9 @@ static bool vsock_msg_wait_data(struct sock *sk, struct sk_psock *psock, long ti sk_set_bit(SOCKWQ_ASYNC_WAITDATA, sk); ret = vsock_has_data(sk, psock); if (!ret) { + release_sock(sk); wait_woken(&wait, TASK_INTERRUPTIBLE, timeo); + lock_sock(sk); ret = vsock_has_data(sk, psock); } sk_clear_bit(SOCKWQ_ASYNC_WAITDATA, sk); --- base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e change-id: 20260929-kbh3-1-022-fix-340aa3f7f071 Best regards, -- Jerome Mohm <[email protected]>
LGTM, but I'd like also Bobby and Michal opinion: Acked-by: Stefano Garzarella <[email protected]> Thanks, Stefano

