* [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback
@ 2026-09-12 9:04 Jia Jia
2026-09-13 9:04 ` sashiko-bot
2026-09-14 20:58 ` Mike Christie
0 siblings, 2 replies; 8+ messages in thread
From: Jia Jia @ 2026-09-12 9:04 UTC (permalink / raw)
To: mst, jasowangio, michael.christie, pbonzini, stefanha, eperezma
Cc: virtualization, kvm, netdev, linux-kernel
vhost_scsi_send_evt() is called with the event virtqueue mutex held.
If the worker is gone, the fallback currently calls
vhost_scsi_complete_events(), which tries to acquire the same mutex again
and deadlocks the caller.
Split event completion into a helper for callers that already hold the
mutex and a locking wrapper for the event worker. Use the helper on the
fallback path.
Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/
Fixes: b1b2ce58ed23 ("vhost-scsi: Handle vhost_vq_work_queue failures for events")
Signed-off-by: Jia Jia <physicalmtea@gmail.com>
---
Sashiko AI flagged this while reviewing
the vhost-scsi event queue fix.
This is a pre-existing self-deadlock. It was reproduced in a follow-up test.
Trigger: vq->worker == NULL. vhost_vq_work_queue() then returns false,
and a subsequent vhost_scsi_do_plug() call deadlocks. I do not know what
normal condition gets us here; the normal vhost-scsi worker detach/reset
paths do not reach this code. The only reproduction I could come up with
was killing the vhost-scsi worker. This still looks like a low-probability
condition.
drivers/vhost/scsi.c | 15 +++++++++++----
1 file changed, 11 insertions(+), 4 deletions(-)
diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c
index 7a1f39a32..7c905e9ad 100644
--- a/drivers/vhost/scsi.c
+++ b/drivers/vhost/scsi.c
@@ -630,19 +630,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt)
vhost_scsi_log_write(vq, vq_log, log_num);
}
-static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop)
+/* Caller must hold the event virtqueue mutex. */
+static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop)
{
- struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq;
struct vhost_scsi_evt *evt, *t;
struct llist_node *llnode;
- mutex_lock(&vq->mutex);
llnode = llist_del_all(&vs->vs_event_list);
llist_for_each_entry_safe(evt, t, llnode, list) {
if (!drop)
vhost_scsi_do_evt_work(vs, evt);
vhost_scsi_free_evt(vs, evt);
}
+}
+
+static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop)
+{
+ struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq;
+
+ mutex_lock(&vq->mutex);
+ __vhost_scsi_complete_events(vs, drop);
mutex_unlock(&vq->mutex);
}
@@ -1829,7 +1836,7 @@ vhost_scsi_send_evt(struct vhost_scsi *vs, struct vhost_virtqueue *vq,
llist_add(&evt->list, &vs->vs_event_list);
if (!vhost_vq_work_queue(vq, &vs->vs_event_work))
- vhost_scsi_complete_events(vs, true);
+ __vhost_scsi_complete_events(vs, true);
}
static void vhost_scsi_evt_handle_kick(struct vhost_work *work)
^ permalink raw reply related [flat|nested] 8+ messages in thread* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-12 9:04 [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback Jia Jia @ 2026-09-13 9:04 ` sashiko-bot 2026-09-14 8:38 ` Jia Jia 2026-09-14 20:58 ` Mike Christie 1 sibling, 1 reply; 8+ messages in thread From: sashiko-bot @ 2026-09-13 9:04 UTC (permalink / raw) To: Jia Jia; +Cc: virtualization, Michael S. Tsirkin, Eugenio Perez, kvm Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: Pre-existing issues: - [Critical] vhost_scsi_do_evt_work() accesses vq->iov[out] without verifying that the guest actually provided an input buffer, leading to arbitrary user-space memory corruption. - [High] Events are processed in LIFO order instead of FIFO, which can cause state corruption for hotplug/hotunplug sequences. -- commit 05830c30fc2e4dc682c270a9b487ae5abb66218b Author: Jia Jia <physicalmtea@gmail.com> vhost-scsi: do not relock event vq mutex on send_evt fallback This commit resolves a deadlock during fallback event handling when the vhost worker has disconnected. It splits event completion into an internal unlocked helper and a locked wrapper to safely reuse the completion path. Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c > index 4f8c0260bc9e7..42abf18291b75 100644 > --- a/drivers/vhost/scsi.c > +++ b/drivers/vhost/scsi.c > @@ -660,19 +660,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt) > vhost_scsi_log_write(vq, vq_log, log_num); > } > > -static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > +/* Caller must hold the event virtqueue mutex. */ > +static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > { > - struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq; > struct vhost_scsi_evt *evt, *t; > struct llist_node *llnode; > > - mutex_lock(&vq->mutex); > llnode = llist_del_all(&vs->vs_event_list); > llist_for_each_entry_safe(evt, t, llnode, list) { [Severity: High] This is a pre-existing issue, but does this loop deliver events to the guest in reverse order? The events are added to vs_event_list using llist_add(), which constructs a last-in, first-out list. Because lnode is retrieved locklessly with llist_del_all(), and the iteration happens without first calling llist_reverse_order(), the events appear to be processed from newest to oldest. Could this cause state corruption for hotplug and hotunplug sequences? For example, if a device is plugged and then rapidly unplugged, the guest might receive the unplug event first and the plug event second, leaving it erroneously believing the device is still present. > if (!drop) > vhost_scsi_do_evt_work(vs, evt); [Severity: Critical] This is a pre-existing issue, but can this function access out-of-bounds or stale memory if the guest provides an empty input buffer? Looking at vhost_scsi_do_evt_work() in drivers/vhost/scsi.c, the code uses vhost_get_vq_desc() which populates vq->iov: head = vhost_get_vq_desc(vq, vq->iov, ARRAY_SIZE(vq->iov), &out, &in, vq_log, &log_num); ... if ((vq->iov[out].iov_len != sizeof(struct virtio_scsi_event))) { vq_err(vq, "Expecting virtio_scsi_event, got %zu bytes\n", vq->iov[out].iov_len); If a guest submits a descriptor chain with only read-only descriptors, vhost_get_vq_desc() sets in = 0. The out variable holds the count of read-only descriptors, meaning the first writable descriptor expected by the host would be at index out. Does the length check above access vq->iov[out].iov_len without verifying that in > 0? If in is 0, vq->iov[out] was not populated by vhost_get_vq_desc() and contains stale data from a previous request. Could a malicious guest use this to force the kernel to write the event structure to a previously legitimate but now hot-unplugged host virtual address via the __copy_to_user() call later in the function? > vhost_scsi_free_evt(vs, evt); > } > +} [ ... ] -- Sashiko AI review · https://sashiko.dev/#/patchset/20260912090420.47439-1-physicalmtea@gmail.com?part=1 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-13 9:04 ` sashiko-bot @ 2026-09-14 8:38 ` Jia Jia 2026-09-15 11:01 ` Michael S. Tsirkin 0 siblings, 1 reply; 8+ messages in thread From: Jia Jia @ 2026-09-14 8:38 UTC (permalink / raw) To: sashiko-reviews; +Cc: virtualization, Michael S. Tsirkin, Eugenio Perez > > Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: > > Pre-existing issues: > - [Critical] vhost_scsi_do_evt_work() accesses vq->iov[out] without verifying that the guest actually provided an input buffer, leading to arbitrary user-space memory corruption. > - [High] Events are processed in LIFO order instead of FIFO, which can cause state corruption for hotplug/hotunplug sequences. > -- > > commit 05830c30fc2e4dc682c270a9b487ae5abb66218b > Author: Jia Jia <physicalmtea@gmail.com> > > vhost-scsi: do not relock event vq mutex on send_evt fallback > > This commit resolves a deadlock during fallback event handling when the vhost > worker has disconnected. It splits event completion into an internal unlocked > helper and a locked wrapper to safely reuse the completion path. > > Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > > > diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c > > index 4f8c0260bc9e7..42abf18291b75 100644 > > --- a/drivers/vhost/scsi.c > > +++ b/drivers/vhost/scsi.c > > @@ -660,19 +660,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt) > > vhost_scsi_log_write(vq, vq_log, log_num); > > } > > > > -static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > +/* Caller must hold the event virtqueue mutex. */ > > +static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > { > > - struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq; > > struct vhost_scsi_evt *evt, *t; > > struct llist_node *llnode; > > > > - mutex_lock(&vq->mutex); > > llnode = llist_del_all(&vs->vs_event_list); > > llist_for_each_entry_safe(evt, t, llnode, list) { > > [Severity: High] > This is a pre-existing issue, but does this loop deliver events to the guest > in reverse order? > > The events are added to vs_event_list using llist_add(), which constructs a > last-in, first-out list. Because lnode is retrieved locklessly with > llist_del_all(), and the iteration happens without first calling > llist_reverse_order(), the events appear to be processed from newest to > oldest. > > Could this cause state corruption for hotplug and hotunplug sequences? For > example, if a device is plugged and then rapidly unplugged, the guest might > receive the unplug event first and the plug event second, leaving it > erroneously believing the device is still present. > Looks like this could get the acquire order wrong? Need to verify. > > if (!drop) > > vhost_scsi_do_evt_work(vs, evt); > > [Severity: Critical] > This is a pre-existing issue, but can this function access out-of-bounds or > stale memory if the guest provides an empty input buffer? > > Looking at vhost_scsi_do_evt_work() in drivers/vhost/scsi.c, the code uses > vhost_get_vq_desc() which populates vq->iov: > > head = vhost_get_vq_desc(vq, vq->iov, > ARRAY_SIZE(vq->iov), &out, &in, > vq_log, &log_num); > ... > if ((vq->iov[out].iov_len != sizeof(struct virtio_scsi_event))) { > vq_err(vq, "Expecting virtio_scsi_event, got %zu bytes\n", > vq->iov[out].iov_len); > > If a guest submits a descriptor chain with only read-only descriptors, > vhost_get_vq_desc() sets in = 0. The out variable holds the count of > read-only descriptors, meaning the first writable descriptor expected by the > host would be at index out. > This is a dup -- an earlier patch already fixed it: https://lore.kernel.org/all/20260904005240.502787-1-physicalmtea@gmail.com/ It was while fixing that earlier issue that Sashiko spotted this adjacent problem. > Does the length check above access vq->iov[out].iov_len without verifying > that in > 0? > > If in is 0, vq->iov[out] was not populated by vhost_get_vq_desc() and > contains stale data from a previous request. Could a malicious guest use > this to force the kernel to write the event structure to a previously > legitimate but now hot-unplugged host virtual address via the > __copy_to_user() call later in the function? > > > vhost_scsi_free_evt(vs, evt); > > } > > +} > > [ ... ] > > -- > Sashiko AI review · https://sashiko.dev/#/patchset/20260912090420.47439-1-physicalmtea@gmail.com?part=1 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-14 8:38 ` Jia Jia @ 2026-09-15 11:01 ` Michael S. Tsirkin 2026-09-15 13:30 ` Jia Jia 0 siblings, 1 reply; 8+ messages in thread From: Michael S. Tsirkin @ 2026-09-15 11:01 UTC (permalink / raw) To: Jia Jia; +Cc: sashiko-reviews, virtualization, Eugenio Perez On Mon, Sep 14, 2026 at 04:38:57PM +0800, Jia Jia wrote: > > > > Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: > > > > Pre-existing issues: > > - [Critical] vhost_scsi_do_evt_work() accesses vq->iov[out] without verifying that the guest actually provided an input buffer, leading to arbitrary user-space memory corruption. > > - [High] Events are processed in LIFO order instead of FIFO, which can cause state corruption for hotplug/hotunplug sequences. > > -- > > > > commit 05830c30fc2e4dc682c270a9b487ae5abb66218b > > Author: Jia Jia <physicalmtea@gmail.com> > > > > vhost-scsi: do not relock event vq mutex on send_evt fallback > > > > This commit resolves a deadlock during fallback event handling when the vhost > > worker has disconnected. It splits event completion into an internal unlocked > > helper and a locked wrapper to safely reuse the completion path. > > > > Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > > > > > diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c > > > index 4f8c0260bc9e7..42abf18291b75 100644 > > > --- a/drivers/vhost/scsi.c > > > +++ b/drivers/vhost/scsi.c > > > @@ -660,19 +660,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt) > > > vhost_scsi_log_write(vq, vq_log, log_num); > > > } > > > > > > -static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > +/* Caller must hold the event virtqueue mutex. */ > > > +static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > { > > > - struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq; > > > struct vhost_scsi_evt *evt, *t; > > > struct llist_node *llnode; > > > > > > - mutex_lock(&vq->mutex); > > > llnode = llist_del_all(&vs->vs_event_list); > > > llist_for_each_entry_safe(evt, t, llnode, list) { > > > > [Severity: High] > > This is a pre-existing issue, but does this loop deliver events to the guest > > in reverse order? > > > > The events are added to vs_event_list using llist_add(), which constructs a > > last-in, first-out list. Because lnode is retrieved locklessly with > > llist_del_all(), and the iteration happens without first calling > > llist_reverse_order(), the events appear to be processed from newest to > > oldest. > > > > Could this cause state corruption for hotplug and hotunplug sequences? For > > example, if a device is plugged and then rapidly unplugged, the guest might > > receive the unplug event first and the plug event second, leaving it > > erroneously believing the device is still present. > > > > Looks like this could get the acquire order wrong? Need to verify. didn't look into this yet - did u reach any conclusion? > > > if (!drop) > > > vhost_scsi_do_evt_work(vs, evt); > > > > [Severity: Critical] > > This is a pre-existing issue, but can this function access out-of-bounds or > > stale memory if the guest provides an empty input buffer? > > > > Looking at vhost_scsi_do_evt_work() in drivers/vhost/scsi.c, the code uses > > vhost_get_vq_desc() which populates vq->iov: > > > > head = vhost_get_vq_desc(vq, vq->iov, > > ARRAY_SIZE(vq->iov), &out, &in, > > vq_log, &log_num); > > ... > > if ((vq->iov[out].iov_len != sizeof(struct virtio_scsi_event))) { > > vq_err(vq, "Expecting virtio_scsi_event, got %zu bytes\n", > > vq->iov[out].iov_len); > > > > If a guest submits a descriptor chain with only read-only descriptors, > > vhost_get_vq_desc() sets in = 0. The out variable holds the count of > > read-only descriptors, meaning the first writable descriptor expected by the > > host would be at index out. > > > > This is a dup -- an earlier patch already fixed it: > https://lore.kernel.org/all/20260904005240.502787-1-physicalmtea@gmail.com/ > It was while fixing that earlier issue that Sashiko spotted this > adjacent problem. > > > Does the length check above access vq->iov[out].iov_len without verifying > > that in > 0? > > > > If in is 0, vq->iov[out] was not populated by vhost_get_vq_desc() and > > contains stale data from a previous request. Could a malicious guest use > > this to force the kernel to write the event structure to a previously > > legitimate but now hot-unplugged host virtual address via the > > __copy_to_user() call later in the function? > > > > > vhost_scsi_free_evt(vs, evt); > > > } > > > +} > > > > [ ... ] > > > > -- > > Sashiko AI review · https://sashiko.dev/#/patchset/20260912090420.47439-1-physicalmtea@gmail.com?part=1 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-15 11:01 ` Michael S. Tsirkin @ 2026-09-15 13:30 ` Jia Jia 2026-09-15 14:46 ` Michael S. Tsirkin 0 siblings, 1 reply; 8+ messages in thread From: Jia Jia @ 2026-09-15 13:30 UTC (permalink / raw) To: Michael S. Tsirkin; +Cc: sashiko-reviews, virtualization, Eugenio Perez > > On Mon, Sep 14, 2026 at 04:38:57PM +0800, Jia Jia wrote: > > > > > > Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: > > > > > > Pre-existing issues: > > > - [Critical] vhost_scsi_do_evt_work() accesses vq->iov[out] without verifying that the guest actually provided an input buffer, leading to arbitrary user-space memory corruption. > > > - [High] Events are processed in LIFO order instead of FIFO, which can cause state corruption for hotplug/hotunplug sequences. > > > -- > > > > > > commit 05830c30fc2e4dc682c270a9b487ae5abb66218b > > > Author: Jia Jia <physicalmtea@gmail.com> > > > > > > vhost-scsi: do not relock event vq mutex on send_evt fallback > > > > > > This commit resolves a deadlock during fallback event handling when the vhost > > > worker has disconnected. It splits event completion into an internal unlocked > > > helper and a locked wrapper to safely reuse the completion path. > > > > > > Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > > > > > > > diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c > > > > index 4f8c0260bc9e7..42abf18291b75 100644 > > > > --- a/drivers/vhost/scsi.c > > > > +++ b/drivers/vhost/scsi.c > > > > @@ -660,19 +660,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt) > > > > vhost_scsi_log_write(vq, vq_log, log_num); > > > > } > > > > > > > > -static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > > +/* Caller must hold the event virtqueue mutex. */ > > > > +static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > > { > > > > - struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq; > > > > struct vhost_scsi_evt *evt, *t; > > > > struct llist_node *llnode; > > > > > > > > - mutex_lock(&vq->mutex); > > > > llnode = llist_del_all(&vs->vs_event_list); > > > > llist_for_each_entry_safe(evt, t, llnode, list) { > > > > > > [Severity: High] > > > This is a pre-existing issue, but does this loop deliver events to the guest > > > in reverse order? > > > > > > The events are added to vs_event_list using llist_add(), which constructs a > > > last-in, first-out list. Because lnode is retrieved locklessly with > > > llist_del_all(), and the iteration happens without first calling > > > llist_reverse_order(), the events appear to be processed from newest to > > > oldest. > > > > > > Could this cause state corruption for hotplug and hotunplug sequences? For > > > example, if a device is plugged and then rapidly unplugged, the guest might > > > receive the unplug event first and the plug event second, leaving it > > > erroneously believing the device is still present. > > > > > > > Looks like this could get the acquire order wrong? Need to verify. > > > didn't look into this yet - did u reach any conclusion? > Already reproduced this using a test program. To hit what Sashiko AI described, need multiple vhost_scsi_send_evt->llist_add (at least 2) before the worker schedules vhost_scsi_complete_events(false). I personally think this is not very easy to hit, first this kind of hotplug/unplug event is not very frequent, even if there are multiple calls, as long as the worker is idle it will run vhost_scsi_complete_events(false) quickly, so pending is often only 1. So to reproduce, I created multiple LUNs, and concurrently called 8 ln -s (only triggers RESCAN, no unplug, extra LUNs are just so we can see the lun numbers better). This increases the chance that vhost_scsi_send_evt races in before vhost_scsi_complete_events(false) runs, so there are at least 2+ pending llist. Finally when vhost_scsi_complete_events(false) runs, the guest sees the later-linked LUN first. (before testing I thought this needed a lot of SCSI I/O to reproduce, actually it doesn't) I only tested plug, the guest did see add LUN order reversed. > > > > if (!drop) > > > > vhost_scsi_do_evt_work(vs, evt); > > > > > > [Severity: Critical] > > > This is a pre-existing issue, but can this function access out-of-bounds or > > > stale memory if the guest provides an empty input buffer? > > > > > > Looking at vhost_scsi_do_evt_work() in drivers/vhost/scsi.c, the code uses > > > vhost_get_vq_desc() which populates vq->iov: > > > > > > head = vhost_get_vq_desc(vq, vq->iov, > > > ARRAY_SIZE(vq->iov), &out, &in, > > > vq_log, &log_num); > > > ... > > > if ((vq->iov[out].iov_len != sizeof(struct virtio_scsi_event))) { > > > vq_err(vq, "Expecting virtio_scsi_event, got %zu bytes\n", > > > vq->iov[out].iov_len); > > > > > > If a guest submits a descriptor chain with only read-only descriptors, > > > vhost_get_vq_desc() sets in = 0. The out variable holds the count of > > > read-only descriptors, meaning the first writable descriptor expected by the > > > host would be at index out. > > > > > > > This is a dup -- an earlier patch already fixed it: > > https://lore.kernel.org/all/20260904005240.502787-1-physicalmtea@gmail.com/ > > It was while fixing that earlier issue that Sashiko spotted this > > adjacent problem. > > > > > Does the length check above access vq->iov[out].iov_len without verifying > > > that in > 0? > > > > > > If in is 0, vq->iov[out] was not populated by vhost_get_vq_desc() and > > > contains stale data from a previous request. Could a malicious guest use > > > this to force the kernel to write the event structure to a previously > > > legitimate but now hot-unplugged host virtual address via the > > > __copy_to_user() call later in the function? > > > > > > > vhost_scsi_free_evt(vs, evt); > > > > } > > > > +} > > > > > > [ ... ] > > > > > > -- > > > Sashiko AI review · https://sashiko.dev/#/patchset/20260912090420.47439-1-physicalmtea@gmail.com?part=1 > ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-15 13:30 ` Jia Jia @ 2026-09-15 14:46 ` Michael S. Tsirkin 2026-09-16 0:24 ` Jia Jia 0 siblings, 1 reply; 8+ messages in thread From: Michael S. Tsirkin @ 2026-09-15 14:46 UTC (permalink / raw) To: Jia Jia; +Cc: sashiko-reviews, virtualization, Eugenio Perez On Tue, Sep 15, 2026 at 09:30:27PM +0800, Jia Jia wrote: > > > > On Mon, Sep 14, 2026 at 04:38:57PM +0800, Jia Jia wrote: > > > > > > > > Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: > > > > > > > > Pre-existing issues: > > > > - [Critical] vhost_scsi_do_evt_work() accesses vq->iov[out] without verifying that the guest actually provided an input buffer, leading to arbitrary user-space memory corruption. > > > > - [High] Events are processed in LIFO order instead of FIFO, which can cause state corruption for hotplug/hotunplug sequences. > > > > -- > > > > > > > > commit 05830c30fc2e4dc682c270a9b487ae5abb66218b > > > > Author: Jia Jia <physicalmtea@gmail.com> > > > > > > > > vhost-scsi: do not relock event vq mutex on send_evt fallback > > > > > > > > This commit resolves a deadlock during fallback event handling when the vhost > > > > worker has disconnected. It splits event completion into an internal unlocked > > > > helper and a locked wrapper to safely reuse the completion path. > > > > > > > > Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > > > > > > > > > diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c > > > > > index 4f8c0260bc9e7..42abf18291b75 100644 > > > > > --- a/drivers/vhost/scsi.c > > > > > +++ b/drivers/vhost/scsi.c > > > > > @@ -660,19 +660,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt) > > > > > vhost_scsi_log_write(vq, vq_log, log_num); > > > > > } > > > > > > > > > > -static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > > > +/* Caller must hold the event virtqueue mutex. */ > > > > > +static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > > > { > > > > > - struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq; > > > > > struct vhost_scsi_evt *evt, *t; > > > > > struct llist_node *llnode; > > > > > > > > > > - mutex_lock(&vq->mutex); > > > > > llnode = llist_del_all(&vs->vs_event_list); > > > > > llist_for_each_entry_safe(evt, t, llnode, list) { > > > > > > > > [Severity: High] > > > > This is a pre-existing issue, but does this loop deliver events to the guest > > > > in reverse order? > > > > > > > > The events are added to vs_event_list using llist_add(), which constructs a > > > > last-in, first-out list. Because lnode is retrieved locklessly with > > > > llist_del_all(), and the iteration happens without first calling > > > > llist_reverse_order(), the events appear to be processed from newest to > > > > oldest. > > > > > > > > Could this cause state corruption for hotplug and hotunplug sequences? For > > > > example, if a device is plugged and then rapidly unplugged, the guest might > > > > receive the unplug event first and the plug event second, leaving it > > > > erroneously believing the device is still present. > > > > > > > > > > Looks like this could get the acquire order wrong? Need to verify. > > > > > > didn't look into this yet - did u reach any conclusion? > > > > Already reproduced this using a test program. To hit what Sashiko AI described, > need multiple vhost_scsi_send_evt->llist_add (at least 2) before the > worker schedules > vhost_scsi_complete_events(false). > > I personally think this is not very easy to hit, first this kind of > hotplug/unplug event is not very frequent, > even if there are multiple calls, as long as the worker is idle it > will run vhost_scsi_complete_events(false) > quickly, so pending is often only 1. > > So to reproduce, I created multiple LUNs, and concurrently called 8 ln > -s (only triggers RESCAN, > no unplug, extra LUNs are just so we can see the lun numbers better). > This increases the chance that > vhost_scsi_send_evt races in before vhost_scsi_complete_events(false) > runs, so there are at least 2+ pending llist. > Finally when vhost_scsi_complete_events(false) runs, the guest sees > the later-linked LUN first. > (before testing I thought this needed a lot of SCSI I/O to reproduce, > actually it doesn't) > I only tested plug, the guest did see add LUN order reversed. okay... so llist_reverse_order ? Will that fix it? > > > > > > if (!drop) > > > > > vhost_scsi_do_evt_work(vs, evt); > > > > > > > > [Severity: Critical] > > > > This is a pre-existing issue, but can this function access out-of-bounds or > > > > stale memory if the guest provides an empty input buffer? > > > > > > > > Looking at vhost_scsi_do_evt_work() in drivers/vhost/scsi.c, the code uses > > > > vhost_get_vq_desc() which populates vq->iov: > > > > > > > > head = vhost_get_vq_desc(vq, vq->iov, > > > > ARRAY_SIZE(vq->iov), &out, &in, > > > > vq_log, &log_num); > > > > ... > > > > if ((vq->iov[out].iov_len != sizeof(struct virtio_scsi_event))) { > > > > vq_err(vq, "Expecting virtio_scsi_event, got %zu bytes\n", > > > > vq->iov[out].iov_len); > > > > > > > > If a guest submits a descriptor chain with only read-only descriptors, > > > > vhost_get_vq_desc() sets in = 0. The out variable holds the count of > > > > read-only descriptors, meaning the first writable descriptor expected by the > > > > host would be at index out. > > > > > > > > > > This is a dup -- an earlier patch already fixed it: > > > https://lore.kernel.org/all/20260904005240.502787-1-physicalmtea@gmail.com/ > > > It was while fixing that earlier issue that Sashiko spotted this > > > adjacent problem. > > > > > > > Does the length check above access vq->iov[out].iov_len without verifying > > > > that in > 0? > > > > > > > > If in is 0, vq->iov[out] was not populated by vhost_get_vq_desc() and > > > > contains stale data from a previous request. Could a malicious guest use > > > > this to force the kernel to write the event structure to a previously > > > > legitimate but now hot-unplugged host virtual address via the > > > > __copy_to_user() call later in the function? > > > > > > > > > vhost_scsi_free_evt(vs, evt); > > > > > } > > > > > +} > > > > > > > > [ ... ] > > > > > > > > -- > > > > Sashiko AI review · https://sashiko.dev/#/patchset/20260912090420.47439-1-physicalmtea@gmail.com?part=1 > > ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-15 14:46 ` Michael S. Tsirkin @ 2026-09-16 0:24 ` Jia Jia 0 siblings, 0 replies; 8+ messages in thread From: Jia Jia @ 2026-09-16 0:24 UTC (permalink / raw) To: Michael S. Tsirkin; +Cc: sashiko-reviews, virtualization, Eugenio Perez > > On Tue, Sep 15, 2026 at 09:30:27PM +0800, Jia Jia wrote: > > > > > > On Mon, Sep 14, 2026 at 04:38:57PM +0800, Jia Jia wrote: > > > > > > > > > > Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: > > > > > > > > > > Pre-existing issues: > > > > > - [Critical] vhost_scsi_do_evt_work() accesses vq->iov[out] without verifying that the guest actually provided an input buffer, leading to arbitrary user-space memory corruption. > > > > > - [High] Events are processed in LIFO order instead of FIFO, which can cause state corruption for hotplug/hotunplug sequences. > > > > > -- > > > > > > > > > > commit 05830c30fc2e4dc682c270a9b487ae5abb66218b > > > > > Author: Jia Jia <physicalmtea@gmail.com> > > > > > > > > > > vhost-scsi: do not relock event vq mutex on send_evt fallback > > > > > > > > > > This commit resolves a deadlock during fallback event handling when the vhost > > > > > worker has disconnected. It splits event completion into an internal unlocked > > > > > helper and a locked wrapper to safely reuse the completion path. > > > > > > > > > > Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > > > > > > > > > > > diff --git a/drivers/vhost/scsi.c b/drivers/vhost/scsi.c > > > > > > index 4f8c0260bc9e7..42abf18291b75 100644 > > > > > > --- a/drivers/vhost/scsi.c > > > > > > +++ b/drivers/vhost/scsi.c > > > > > > @@ -660,19 +660,26 @@ vhost_scsi_do_evt_work(struct vhost_scsi *vs, struct vhost_scsi_evt *evt) > > > > > > vhost_scsi_log_write(vq, vq_log, log_num); > > > > > > } > > > > > > > > > > > > -static void vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > > > > +/* Caller must hold the event virtqueue mutex. */ > > > > > > +static void __vhost_scsi_complete_events(struct vhost_scsi *vs, bool drop) > > > > > > { > > > > > > - struct vhost_virtqueue *vq = &vs->vqs[VHOST_SCSI_VQ_EVT].vq; > > > > > > struct vhost_scsi_evt *evt, *t; > > > > > > struct llist_node *llnode; > > > > > > > > > > > > - mutex_lock(&vq->mutex); > > > > > > llnode = llist_del_all(&vs->vs_event_list); > > > > > > llist_for_each_entry_safe(evt, t, llnode, list) { > > > > > > > > > > [Severity: High] > > > > > This is a pre-existing issue, but does this loop deliver events to the guest > > > > > in reverse order? > > > > > > > > > > The events are added to vs_event_list using llist_add(), which constructs a > > > > > last-in, first-out list. Because lnode is retrieved locklessly with > > > > > llist_del_all(), and the iteration happens without first calling > > > > > llist_reverse_order(), the events appear to be processed from newest to > > > > > oldest. > > > > > > > > > > Could this cause state corruption for hotplug and hotunplug sequences? For > > > > > example, if a device is plugged and then rapidly unplugged, the guest might > > > > > receive the unplug event first and the plug event second, leaving it > > > > > erroneously believing the device is still present. > > > > > > > > > > > > > Looks like this could get the acquire order wrong? Need to verify. > > > > > > > > > didn't look into this yet - did u reach any conclusion? > > > > > > > Already reproduced this using a test program. To hit what Sashiko AI described, > > need multiple vhost_scsi_send_evt->llist_add (at least 2) before the > > worker schedules > > vhost_scsi_complete_events(false). > > > > I personally think this is not very easy to hit, first this kind of > > hotplug/unplug event is not very frequent, > > even if there are multiple calls, as long as the worker is idle it > > will run vhost_scsi_complete_events(false) > > quickly, so pending is often only 1. > > > > So to reproduce, I created multiple LUNs, and concurrently called 8 ln > > -s (only triggers RESCAN, > > no unplug, extra LUNs are just so we can see the lun numbers better). > > This increases the chance that > > vhost_scsi_send_evt races in before vhost_scsi_complete_events(false) > > runs, so there are at least 2+ pending llist. > > Finally when vhost_scsi_complete_events(false) runs, the guest sees > > the later-linked LUN first. > > (before testing I thought this needed a lot of SCSI I/O to reproduce, > > actually it doesn't) > > I only tested plug, the guest did see add LUN order reversed. > > okay... so llist_reverse_order ? Will that fix it? > Yes. After reproducing I already sent another patch: <20260915094244.7900-1-physicalmtea@gmail.com> > > > > > > > > if (!drop) > > > > > > vhost_scsi_do_evt_work(vs, evt); > > > > > > > > > > [Severity: Critical] > > > > > This is a pre-existing issue, but can this function access out-of-bounds or > > > > > stale memory if the guest provides an empty input buffer? > > > > > > > > > > Looking at vhost_scsi_do_evt_work() in drivers/vhost/scsi.c, the code uses > > > > > vhost_get_vq_desc() which populates vq->iov: > > > > > > > > > > head = vhost_get_vq_desc(vq, vq->iov, > > > > > ARRAY_SIZE(vq->iov), &out, &in, > > > > > vq_log, &log_num); > > > > > ... > > > > > if ((vq->iov[out].iov_len != sizeof(struct virtio_scsi_event))) { > > > > > vq_err(vq, "Expecting virtio_scsi_event, got %zu bytes\n", > > > > > vq->iov[out].iov_len); > > > > > > > > > > If a guest submits a descriptor chain with only read-only descriptors, > > > > > vhost_get_vq_desc() sets in = 0. The out variable holds the count of > > > > > read-only descriptors, meaning the first writable descriptor expected by the > > > > > host would be at index out. > > > > > > > > > > > > > This is a dup -- an earlier patch already fixed it: > > > > https://lore.kernel.org/all/20260904005240.502787-1-physicalmtea@gmail.com/ > > > > It was while fixing that earlier issue that Sashiko spotted this > > > > adjacent problem. > > > > > > > > > Does the length check above access vq->iov[out].iov_len without verifying > > > > > that in > 0? > > > > > > > > > > If in is 0, vq->iov[out] was not populated by vhost_get_vq_desc() and > > > > > contains stale data from a previous request. Could a malicious guest use > > > > > this to force the kernel to write the event structure to a previously > > > > > legitimate but now hot-unplugged host virtual address via the > > > > > __copy_to_user() call later in the function? > > > > > > > > > > > vhost_scsi_free_evt(vs, evt); > > > > > > } > > > > > > +} > > > > > > > > > > [ ... ] > > > > > > > > > > -- > > > > > Sashiko AI review · https://sashiko.dev/#/patchset/20260912090420.47439-1-physicalmtea@gmail.com?part=1 > > > > ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback 2026-09-12 9:04 [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback Jia Jia 2026-09-13 9:04 ` sashiko-bot @ 2026-09-14 20:58 ` Mike Christie 1 sibling, 0 replies; 8+ messages in thread From: Mike Christie @ 2026-09-14 20:58 UTC (permalink / raw) To: Jia Jia, mst, jasowangio, pbonzini, stefanha, eperezma Cc: virtualization, kvm, netdev, linux-kernel On 9/12/26 4:04 AM, Jia Jia wrote: > vhost_scsi_send_evt() is called with the event virtqueue mutex held. > If the worker is gone, the fallback currently calls > vhost_scsi_complete_events(), which tries to acquire the same mutex again > and deadlocks the caller. > > Split event completion into a helper for callers that already hold the > mutex and a locking wrapper for the event worker. Use the helper on the > fallback path. > > Link: https://lore.kernel.org/all/20260905005352.1E5B01F00A3D@smtp.kernel.org/ > Fixes: b1b2ce58ed23 ("vhost-scsi: Handle vhost_vq_work_queue failures for events") > Signed-off-by: Jia Jia <physicalmtea@gmail.com> > --- > Sashiko AI flagged this while reviewing > the vhost-scsi event queue fix. > This is a pre-existing self-deadlock. It was reproduced in a follow-up test. > > Trigger: vq->worker == NULL. vhost_vq_work_queue() then returns false, > and a subsequent vhost_scsi_do_plug() call deadlocks. I do not know what > normal condition gets us here; the normal vhost-scsi worker detach/reset > paths do not reach this code. The only reproduction I could come up with > was killing the vhost-scsi worker. This still looks like a low-probability > condition. Reviewed-by: Mike Christie <michael.christie@oracle.com> ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-09-16 0:25 UTC | newest] Thread overview: 8+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-12 9:04 [PATCH] vhost-scsi: do not relock event vq mutex on send_evt fallback Jia Jia 2026-09-13 9:04 ` sashiko-bot 2026-09-14 8:38 ` Jia Jia 2026-09-15 11:01 ` Michael S. Tsirkin 2026-09-15 13:30 ` Jia Jia 2026-09-15 14:46 ` Michael S. Tsirkin 2026-09-16 0:24 ` Jia Jia 2026-09-14 20:58 ` Mike Christie
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.