Netdev List
 help / color / mirror / Atom feed
From: Stefano Garzarella <sgarzare@redhat.com>
To: Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
	 virtualization@lists.linux.dev, netdev@vger.kernel.org,
	mst@redhat.com, stefanha@redhat.com,  dongli.zhang@oracle.com,
	maciej.szmigiero@oracle.com, bchaney@akamai.com,
	 mark.kanda@oracle.com, ptikhomirov@virtuozzo.com,
	den@openvz.org
Subject: Re: [PATCH v5 4/5] vhost: synchronize with RCU readers when freeing workers
Date: Wed, 22 Jul 2026 11:43:29 +0200	[thread overview]
Message-ID: <amCPLOrnpYmr5n0p@sgarzare-redhat> (raw)
In-Reply-To: <20260720102241.371610-5-andrey.drobyshev@virtuozzo.com>

On Mon, Jul 20, 2026 at 01:22:40PM +0300, Andrey Drobyshev wrote:
>vhost_vq_work_queue() only holds the RCU read lock while it dereferences
>vq->worker and queues work on it.  vhost_workers_free() however clears
>the vq->worker pointers and immediately frees the workers, without
>waiting for a grace period.  A caller that fetched the worker right
>before the pointer was cleared can therefore still be queueing work on
>it while it is freed.  And even when the queueing itself wins the race,
>the work is never run, so its VHOST_WORK_QUEUED bit stays set and all
>future attempts to queue it are silently skipped.
>
>None of the current callers can actually hit this: net and scsi stop
>their virtqueues before the workers are freed, and vsock unhashes the
>device and does synchronize_rcu() of its own in vhost_vsock_dev_release()
>before the workers go away.  But the upcoming VHOST_RESET_OWNER support
>in vhost-vsock keeps the device hashed while its workers are freed, so
>the lockless send/cancel paths become able to race with the teardown.
>
>Fix this by clearing the vq->worker pointers, waiting for a grace
>period, and then flushing the workers so any work the last readers
>queued runs before the workers are freed.
>
>Fixes: 228a27cf78af ("vhost: Allow worker switching while work is queueing")
>Suggested-by: Stefano Garzarella <sgarzare@redhat.com>
>Signed-off-by: Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
>---
> drivers/vhost/vhost.c | 11 +++++++++++
> 1 file changed, 11 insertions(+)

Sashiko reported some potential issues here:
https://sashiko.dev/#/patchset/20260720102241.371610-1-andrey.drobyshev@virtuozzo.com?part=4

IMO the first one is pre-existing, but not really sure it is a real 
issue since happening when the worker/vmm is going to be killed.

The second one also not sure if it's an issue since the sender is not 
lockless IIUC.

But, please can you double check them?

Thanks,
Stefano

>
>diff --git a/drivers/vhost/vhost.c b/drivers/vhost/vhost.c
>index 4c525b3e16ea..d6e235c25254 100644
>--- a/drivers/vhost/vhost.c
>+++ b/drivers/vhost/vhost.c
>@@ -729,6 +729,17 @@ static void vhost_workers_free(struct vhost_dev *dev)
>
> 	for (i = 0; i < dev->nvqs; i++)
> 		rcu_assign_pointer(dev->vqs[i]->worker, NULL);
>+
>+	/*
>+	 * vhost_vq_work_queue() reads vq->worker under rcu_read_lock(), so a
>+	 * reader that fetched a worker before we cleared the pointers above
>+	 * may still be queueing work on it.  Wait for those readers to
>+	 * finish, then flush so any work they queued runs (clearing
>+	 * VHOST_WORK_QUEUED) before the workers are freed.
>+	 */
>+	synchronize_rcu();
>+	vhost_dev_flush(dev);
>+
> 	/*
> 	 * Free the default worker we created and cleanup workers userspace
> 	 * created but couldn't clean up (it forgot or crashed).
>-- 
>2.47.1
>


  reply	other threads:[~2026-07-22  9:43 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-20 10:22 [PATCH v5 0/5] vhost/vsock: add support for VHOST_RESET_OWNER and CPR migration Andrey Drobyshev
2026-07-20 10:22 ` [PATCH v5 1/5] vhost/vsock: split out vhost_vsock_drop_backends helper Andrey Drobyshev
2026-07-20 10:22 ` [PATCH v5 2/5] vhost/vsock: suppress EHOSTUNREACH fast-fail during CPR pause Andrey Drobyshev
2026-07-22  9:14   ` Stefano Garzarella
2026-07-20 10:22 ` [PATCH v5 3/5] vhost/vsock: re-scan TX virtqueue on device start Andrey Drobyshev
2026-07-22  9:14   ` Stefano Garzarella
2026-07-20 10:22 ` [PATCH v5 4/5] vhost: synchronize with RCU readers when freeing workers Andrey Drobyshev
2026-07-22  9:43   ` Stefano Garzarella [this message]
2026-07-20 10:22 ` [PATCH v5 5/5] vhost/vsock: add VHOST_RESET_OWNER ioctl Andrey Drobyshev
2026-07-22  9:43   ` Stefano Garzarella

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amCPLOrnpYmr5n0p@sgarzare-redhat \
    --to=sgarzare@redhat.com \
    --cc=andrey.drobyshev@virtuozzo.com \
    --cc=bchaney@akamai.com \
    --cc=den@openvz.org \
    --cc=dongli.zhang@oracle.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=maciej.szmigiero@oracle.com \
    --cc=mark.kanda@oracle.com \
    --cc=mst@redhat.com \
    --cc=netdev@vger.kernel.org \
    --cc=ptikhomirov@virtuozzo.com \
    --cc=stefanha@redhat.com \
    --cc=virtualization@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox