All of lore.kernel.org
 help / color / mirror / Atom feed
From: Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
To: Stefano Garzarella <sgarzare@redhat.com>
Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
	virtualization@lists.linux.dev, netdev@vger.kernel.org,
	mst@redhat.com, stefanha@redhat.com, dongli.zhang@oracle.com,
	maciej.szmigiero@oracle.com, bchaney@akamai.com,
	mark.kanda@oracle.com, ptikhomirov@virtuozzo.com, den@openvz.org
Subject: Re: [PATCH v5 4/5] vhost: synchronize with RCU readers when freeing workers
Date: Thu, 23 Jul 2026 17:49:54 +0300	[thread overview]
Message-ID: <c631dabe-ebde-4b4d-b930-d78db4cbabc4@virtuozzo.com> (raw)
In-Reply-To: <amInL2hPEG9ejUd0@sgarzare-redhat>

On 7/23/26 5:39 PM, Stefano Garzarella wrote:
> On Thu, Jul 23, 2026 at 05:29:12PM +0300, Andrey Drobyshev wrote:
>> On 7/23/26 5:03 PM, Stefano Garzarella wrote:
>>> On Thu, Jul 23, 2026 at 04:57:47PM +0300, Andrey Drobyshev wrote:
>>>> On 7/22/26 12:43 PM, Stefano Garzarella wrote:
>>>>> On Mon, Jul 20, 2026 at 01:22:40PM +0300, Andrey Drobyshev wrote:
>>>>>> vhost_vq_work_queue() only holds the RCU read lock while it dereferences
>>>>>> vq->worker and queues work on it.  vhost_workers_free() however clears
>>>>>> the vq->worker pointers and immediately frees the workers, without
>>>>>> waiting for a grace period.  A caller that fetched the worker right
>>>>>> before the pointer was cleared can therefore still be queueing work on
>>>>>> it while it is freed.  And even when the queueing itself wins the race,
>>>>>> the work is never run, so its VHOST_WORK_QUEUED bit stays set and all
>>>>>> future attempts to queue it are silently skipped.
>>>>>>
>>>>>> None of the current callers can actually hit this: net and scsi stop
>>>>>> their virtqueues before the workers are freed, and vsock unhashes the
>>>>>> device and does synchronize_rcu() of its own in vhost_vsock_dev_release()
>>>>>> before the workers go away.  But the upcoming VHOST_RESET_OWNER support
>>>>>> in vhost-vsock keeps the device hashed while its workers are freed, so
>>>>>> the lockless send/cancel paths become able to race with the teardown.
>>>>>>
>>>>>> Fix this by clearing the vq->worker pointers, waiting for a grace
>>>>>> period, and then flushing the workers so any work the last readers
>>>>>> queued runs before the workers are freed.
>>>>>>
>>>>>> Fixes: 228a27cf78af ("vhost: Allow worker switching while work is queueing")
>>>>>> Suggested-by: Stefano Garzarella <sgarzare@redhat.com>
>>>>>> Signed-off-by: Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
>>>>>> ---
>>>>>> drivers/vhost/vhost.c | 11 +++++++++++
>>>>>> 1 file changed, 11 insertions(+)
>>>>>
>>>>> Sashiko reported some potential issues here:
>>>>> https://sashiko.dev/#/patchset/20260720102241.371610-1-andrey.drobyshev@virtuozzo.com?part=4
>>>>>
>>>>> IMO the first one is pre-existing, but not really sure it is a real
>>>>> issue since happening when the worker/vmm is going to be killed.
>>>>>
>>>>
>>>> Sashiko claims:
>>>>
>>>>> Will this leave the queued work unexecuted and permanently break the
>>>> virtqueue by leaving VHOST_WORK_QUEUED set?
>>>>
>>>> I agree this issue is pre-existing and doesn't have much to do with our
>>>> series here.  It looks real, but in reality should be harmless since we
>>>> may only hit it while the device already dying.  Means there's no VQ
>>>> state to be saved.  The only potentially observable artifact I guess is
>>>> a warning here:
>>>>
>>>> vhost_workers_free()
>>>>  vhost_worker_destroy()
>>>>    WARN_ON(!llist_empty())
>>>>
>>>> So more of a cosmetic noise on a dying device.  Again, not relevant to
>>>> this series.  But one optional way to make it go away would be to clear
>>>> vq->worker under vq->mutex in vhost_workers_free() (mirroring
>>>> vhost_worker_killed()).
>>>>> The second one also not sure if it's an issue since the sender is not
>>>>> lockless IIUC.
>>>>>
>>>>
>>>> The term 'sender' is confusing here:
>>>>
>>>> * vhost_transport_send_pkt() is the .send_pkt() method of struct
>>>> virtio_transport.  It only stores skbs into the queue, doesn't process
>>>> them. It indeed is lockless as it doesn't take vq->mutex.
>>>>
>>>> * vhost_transport_send_pkt_work() is the .fn() method of send_pkt_work.
>>>> It's called by the worker thread to process skbs in the queue, calls
>>>> vhost_transport_do_send_pkt() which does in turn take vq->mutex.
>>>>
>>>> I think sashiko points out to the former.  Still, I don't think it's an
>>>> actual bug.  Look:
>>>>
>>>> 1) By invoking flush, we wake the worker thread:
>>>> vhost_dev_flush()
>>>>  __vhost_worker_flush()
>>>>    vhost_worker_queue(flush.work)
>>>>      worker->ops->wakeup()
>>>>
>>>> 2) Then woken worker does:
>>>> vhost_run_work_list()
>>>>  llist_for_each_entry_safe(work) {
>>>>    clear_bit(VHOST_WORK_QUEUED, &work->flags)
>>>>    work->fn(work) // for send_pkt_work = vhost_transport_send_pkt_work
>>>>  }
>>>>
>>>> 3) And then in work->fn() (vhost_transport_send_pkt_work):
>>>> vhost_transport_send_pkt_work()
>>>>  vhost_transport_do_send_pkt()
>>>>    if (!vhost_vq_get_backend(vq))
>>>>      goto out;
>>>>
>>>> As you can see, we return early in case backend was unset.
>>>>
>>>> So after doing this flush, we have: 1) QUEUED bit is unset; that makes
>>>> send_pkt_work re-queueable. 2) But the queue doesn't actually get
>>>> processed, and VQ state isn't actually touched here - so I guess
>>>> Sashiko's conclusion is incorrect and there's no actual bug.
>>>>
>>>>> But, please can you double check them?
>>>>>
>>>>
>>>> In general I'd leave this patch as-is as Sashiko's complaints aren't
>>>> very convincing so far.  If you want I can add another patch which wraps
>>>> NULLifying workers in vhost_workers_free() in vq->mutex.
>>>>
>>>> WDYT?
>>>
>>> Yeah, I agree, about the other patch, up to you, but I'll eventually
>>> send it separately.
>>>
>>> Thanks,
>>> Stefano
>>>
>>
>> Alright, then let me resend it once more along with this 6th patch, so
>> that we don't have it uncovered.
> 
> But why it should be part of this series?
> 
> If there is no strong reason, I'd send it as a separate patch.
> 
> E.g. even this patch in theory may be a separate one, but this is 
> related to this series, so makes sense to have this included.
> 

No particular reason, just while we're at it.  Alright, no respin, will
send it separately.  Then I'll consider this series reviewed - unless,
maybe, Michael will want to take a look as well?

Thanks,
Andrey

  reply	other threads:[~2026-07-23 14:49 UTC|newest]

Thread overview: 23+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-20 10:22 [PATCH v5 0/5] vhost/vsock: add support for VHOST_RESET_OWNER and CPR migration Andrey Drobyshev
2026-07-20 10:22 ` [PATCH v5 1/5] vhost/vsock: split out vhost_vsock_drop_backends helper Andrey Drobyshev
2026-07-20 10:22 ` [PATCH v5 2/5] vhost/vsock: suppress EHOSTUNREACH fast-fail during CPR pause Andrey Drobyshev
2026-07-22  9:14   ` Stefano Garzarella
2026-07-20 10:22 ` [PATCH v5 3/5] vhost/vsock: re-scan TX virtqueue on device start Andrey Drobyshev
2026-07-22  9:14   ` Stefano Garzarella
2026-07-20 10:22 ` [PATCH v5 4/5] vhost: synchronize with RCU readers when freeing workers Andrey Drobyshev
2026-07-21 10:23   ` sashiko-bot
2026-07-22  9:43   ` Stefano Garzarella
2026-07-23 13:57     ` Andrey Drobyshev
2026-07-23 14:03       ` Stefano Garzarella
2026-07-23 14:29         ` Andrey Drobyshev
2026-07-23 14:31           ` Denis V. Lunev
2026-07-23 14:39           ` Stefano Garzarella
2026-07-23 14:49             ` Andrey Drobyshev [this message]
2026-07-23 14:55               ` Stefano Garzarella
2026-07-23 14:03   ` Stefano Garzarella
2026-07-23 15:18   ` Michael S. Tsirkin
2026-07-23 16:46     ` Andrey Drobyshev
2026-07-23 16:50       ` Michael S. Tsirkin
2026-07-23 17:17         ` Andrey Drobyshev
2026-07-20 10:22 ` [PATCH v5 5/5] vhost/vsock: add VHOST_RESET_OWNER ioctl Andrey Drobyshev
2026-07-22  9:43   ` Stefano Garzarella

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c631dabe-ebde-4b4d-b930-d78db4cbabc4@virtuozzo.com \
    --to=andrey.drobyshev@virtuozzo.com \
    --cc=bchaney@akamai.com \
    --cc=den@openvz.org \
    --cc=dongli.zhang@oracle.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=maciej.szmigiero@oracle.com \
    --cc=mark.kanda@oracle.com \
    --cc=mst@redhat.com \
    --cc=netdev@vger.kernel.org \
    --cc=ptikhomirov@virtuozzo.com \
    --cc=sgarzare@redhat.com \
    --cc=stefanha@redhat.com \
    --cc=virtualization@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.