From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C845D36F90D for ; Thu, 23 Jul 2026 16:50:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784825449; cv=none; b=btrkR9PxogeVki8hhbYeKa1ott3rWJf9lBL9/4h4f1mjnDeickd8I+0e2j9GI/XkVNt/L5XpqswdyvU5i/5ZiezPo4RzK2+3CbftinwtSsEwbXwKZ3aqrEMXbfYxGZ13YwXdpRzspi+N5WKt+Oo0hMCNuPo2z2H9mGBljWOjUVE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784825449; c=relaxed/simple; bh=HVzEoDIUOtKuu1wfn87SZ/HEqRdQyUU7D8rY8b4wr/U=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=UeLK1nAmcnu7evik5PLzVpJtTJOfJ34cY9ZU6/Yc1A1Vdwyx/hKFvncizCE2doqG+HTGrpvcFS+GBeXAIEm64rDSMTl02hFLXiyhEdTB2HDrWdHbo4RxxZBeOtcx2dwQhO/SYm8V9L68M/oboAZPFUyX/uQkhQ3ZeucPAhheDfs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=TjSfeZ6p; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=PmlQH47i; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="TjSfeZ6p"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="PmlQH47i" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1784825446; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=skv50H0aRdZe1+yzOYqLWkWcMMVGnxgMC07zr3GCyxM=; b=TjSfeZ6pu5R8DPjQuuTY4xDpXprvKssADW8j7hARr9DqlmkCzOHyBZ55ey0Axw2CicRtGP XW9+Z8D6ofDT6Wv+2wuUjdHo1gad7BUUtqI85ngb6eGT+nghtrP7JwPjaoLTZ0L4EhFz11 OPf31VtPVILbv9cAt40a615RnshLfh0= Received: from mail-wr1-f71.google.com (mail-wr1-f71.google.com [209.85.221.71]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-468-5kzhJgzcMV6RabdRW01ivw-1; Thu, 23 Jul 2026 12:50:45 -0400 X-MC-Unique: 5kzhJgzcMV6RabdRW01ivw-1 X-Mimecast-MFC-AGG-ID: 5kzhJgzcMV6RabdRW01ivw_1784825444 Received: by mail-wr1-f71.google.com with SMTP id ffacd0b85a97d-472a798fc7cso582414f8f.1 for ; Thu, 23 Jul 2026 09:50:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1784825444; x=1785430244; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=skv50H0aRdZe1+yzOYqLWkWcMMVGnxgMC07zr3GCyxM=; b=PmlQH47iXzx8Dp+yvaKQPhyPR03ooN4skoNw+INVlu+jl7dP3mSWuzh+ZUyLMDqYJh 3WP2xFjDptuTn0Z0WctOYFtaZZvpw8iJwV8VUDDdRDMc+8W6hl0zYgGqEejQNzmsQrld xw+ru70kgD9/iuzhCWGBfrmh4QxRv5icMomCc3rQh6kg1XmzLBpGUe8HtziHOXp1ozuW sNU2bKf7VWN+CtEb8NkFCgxHTVI9ti2jb6LGYpZTOEe1o25+cl/TNE2RSBMfxQyijasf vb96pNuBH0RfYeOrs5b2e2KfIZHr/2GHvgIZrbC7J6wHa39Cw2uW/yICfGlBVgvY2Wqa vYbw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784825444; x=1785430244; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=skv50H0aRdZe1+yzOYqLWkWcMMVGnxgMC07zr3GCyxM=; b=qbbPkWy1ThYF+6IA0PPc5brjxSVeRtoXgy0khaUPiD5kBQuwMxMO6WRQmxAotpGIzT kx6BeTV5F+a+P9Llrvu2GT/tQvL0ghCSTTNuKJ0O4ITlDQZEZurUplxXyphk6Vsc4WKE jjXafFAvv+xtzKdLqQpK0FY0BwOcve82aGtWV0Jxv3NELC6ckXi9qX4NVfSvCeAy453S 6v+a8k47ZkzwfSEXWjQQ7bldHsiKxc52y8KyOmkZLub26GJOXZsB6dfHcT8StLUFK5Os 4DzaAQ0267hV1Uws+E87XNxSFySVzg3MNOttgLSoKfn9raRqYG6XL9k4WxHwJwadvzYa ahaw== X-Gm-Message-State: AOJu0YwZ0lfSFMHrzHDa5h5mdVaUN199eVKVsyWOx17S7JAkGGog35Pd F4+Njpo8fZ10Xh1Ukrtg/JVqYUfCNeI4Rd7JiLVS0fvMVju6LFzd0BAPJ3EbTmFMHSx3HM1cR42 rrsXQLOigT1J8nXf4AnQcBJmJ3B5JysLFlxqJt/kcShHUOy5R+uGx15BYL0QrpVSWYw== X-Gm-Gg: AR+sD13srBpF24mTNsU5tNdT7Ksynji53aCTOxYnCGjrtPUXdh/AGisbfX+/Uob6Pwv djzSnSBey1z6ygBakvWHEo4EucxxuReoDJCsYhZH2OPYfzfTbV/SZ7OpbSxQvBMYMtzzQ7zbKUA UlexasDE9X1Qj2X+orPclfwBtdxzVoX+y9FzRHn44iUbQc2ssw0EzpGia7lRjSlqtXyAvz8zRfY elAf3wD/QF/Y2JcjrlmM5GJneTCVDlMgKaq91P6QahZKwvs7JgY91glqWVBE7HunNSUCIo0TqTd kHX2SNs93VwNMHku5bQtmdZxf6J7e6+WeIC+ha/eRi7f21JAbZtkZHXST3hkzWs5KSwSgx4Olpr /T5TuYOVtha6twTM9azLzyg== X-Received: by 2002:a05:600c:350a:b0:495:441a:39d2 with SMTP id 5b1f17b1804b1-49573cd8932mr43271925e9.21.1784825443873; Thu, 23 Jul 2026 09:50:43 -0700 (PDT) X-Received: by 2002:a05:600c:350a:b0:495:441a:39d2 with SMTP id 5b1f17b1804b1-49573cd8932mr43270805e9.21.1784825442214; Thu, 23 Jul 2026 09:50:42 -0700 (PDT) Received: from redhat.com (IGLD-80-230-37-66.inter.net.il. [80.230.37.66]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4957af5d910sm6009915e9.3.2026.07.23.09.50.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 23 Jul 2026 09:50:41 -0700 (PDT) Date: Thu, 23 Jul 2026 12:50:38 -0400 From: "Michael S. Tsirkin" To: Andrey Drobyshev Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org, virtualization@lists.linux.dev, netdev@vger.kernel.org, sgarzare@redhat.com, stefanha@redhat.com, dongli.zhang@oracle.com, maciej.szmigiero@oracle.com, bchaney@akamai.com, mark.kanda@oracle.com, ptikhomirov@virtuozzo.com, den@openvz.org Subject: Re: [PATCH v5 4/5] vhost: synchronize with RCU readers when freeing workers Message-ID: <20260723124908-mutt-send-email-mst@kernel.org> References: <20260720102241.371610-1-andrey.drobyshev@virtuozzo.com> <20260720102241.371610-5-andrey.drobyshev@virtuozzo.com> <20260723111557-mutt-send-email-mst@kernel.org> <0c7172c4-0dc4-4d7e-9310-4f045e27efb9@virtuozzo.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <0c7172c4-0dc4-4d7e-9310-4f045e27efb9@virtuozzo.com> On Thu, Jul 23, 2026 at 07:46:58PM +0300, Andrey Drobyshev wrote: > On 7/23/26 6:18 PM, Michael S. Tsirkin wrote: > > On Mon, Jul 20, 2026 at 01:22:40PM +0300, Andrey Drobyshev wrote: > >> vhost_vq_work_queue() only holds the RCU read lock while it dereferences > >> vq->worker and queues work on it. vhost_workers_free() however clears > >> the vq->worker pointers and immediately frees the workers, without > >> waiting for a grace period. A caller that fetched the worker right > >> before the pointer was cleared can therefore still be queueing work on > >> it while it is freed. And even when the queueing itself wins the race, > >> the work is never run, so its VHOST_WORK_QUEUED bit stays set and all > >> future attempts to queue it are silently skipped. > >> > >> None of the current callers can actually hit this: net and scsi stop > >> their virtqueues before the workers are freed, and vsock unhashes the > >> device and does synchronize_rcu() of its own in vhost_vsock_dev_release() > >> before the workers go away. But the upcoming VHOST_RESET_OWNER support > >> in vhost-vsock keeps the device hashed while its workers are freed, so > >> the lockless send/cancel paths become able to race with the teardown. > >> > >> Fix this by clearing the vq->worker pointers, waiting for a grace > >> period, and then flushing the workers so any work the last readers > >> queued runs before the workers are freed. > > > > > > > > > >> Fixes: 228a27cf78af ("vhost: Allow worker switching while work is queueing") > >> Suggested-by: Stefano Garzarella > >> Signed-off-by: Andrey Drobyshev > >> --- > >> drivers/vhost/vhost.c | 11 +++++++++++ > >> 1 file changed, 11 insertions(+) > >> > >> diff --git a/drivers/vhost/vhost.c b/drivers/vhost/vhost.c > >> index 4c525b3e16ea..d6e235c25254 100644 > >> --- a/drivers/vhost/vhost.c > >> +++ b/drivers/vhost/vhost.c > >> @@ -729,6 +729,17 @@ static void vhost_workers_free(struct vhost_dev *dev) > >> > >> for (i = 0; i < dev->nvqs; i++) > >> rcu_assign_pointer(dev->vqs[i]->worker, NULL); > >> + > >> + /* > >> + * vhost_vq_work_queue() reads vq->worker under rcu_read_lock(), so a > >> + * reader that fetched a worker before we cleared the pointers above > >> + * may still be queueing work on it. Wait for those readers to > >> + * finish, then flush so any work they queued runs (clearing > >> + * VHOST_WORK_QUEUED) before the workers are freed. > >> + */ > >> + synchronize_rcu(); > > > > > > > > Any way not to add this for all devices that don't need it? > > Or preferably, even for vsock in absense of the new ioctl? > > > > This code was initially local to vsock, and was moved here in v3->v4 > after we discussed with Stefano that the issue looks more generic and > should probably be fixed in vhost.c (see > https://lore.kernel.org/virtualization/akO6tps94iFxCAWv@sgarzare-redhat). > > As a compromise, we can keep this code here, but only call it > conditionally. Namely, create a bool flag on 'struct vhost_dev' which > is always false, only set it to true on RESET_OWNER, and only call this > code once it's set. Clumsy, but this way no other code path would have > to wait the full grace period. > > What are your thoughts on that? > > Thanks, > Andrey Or just thread a bool parameter to it? > > > >> + vhost_dev_flush(dev); > >> + > >> /* > >> * Free the default worker we created and cleanup workers userspace > >> * created but couldn't clean up (it forgot or crashed). > >> -- > >> 2.47.1 > >