From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1505733710F for ; Sun, 2 Aug 2026 19:54:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700475; cv=none; b=PftqU8My0QOdfcJczlwoAL9NCYjGWDmqW63GrkiBQuEnwe9aRNjKYhZsviljOGjn4U1cHAY42JuEtsIMxIZkOcTLQVya1h1xVPCydlsM2RNFg9DW6oFUZ/Uq4uz2cvFjdubTn0/MXZj/btE9Z3wh2Bbtc6yVzJCjaBuuWA/5OKQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700475; c=relaxed/simple; bh=ehK5JhoQna7+Uy0Gif3hJvuLiOiFhkY/+8zP8o4ZtQU=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=QzLZc5otPURQgT0iT5c8bcT9AFUJ6jQcvzaZ3TYzUq+hN7ISJj7q51BZybR0TFKMQfcJX+gZvdO1yl0VMq487WPEvwRvLaRB6TgQyzYPdFa1ABvJgCCI3c3l5GhuDsMaUgSFkmqXN/kZx2jnVA35kjVW5ATwPmkuVttYAw98VS0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=NP1+iIql; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=n3v77MB9; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="NP1+iIql"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="n3v77MB9" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785700473; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=xjP8Nj5ayArKITQglbdxDINIj5XZHeMrEUA74eJ1q04=; b=NP1+iIqlDuMMM0mwDbrzQq9fbB2Ot19kvy6cGVPRKIA5SggUCJDlW41ACMw45XQ1ahK6/9 vI06HBHlKzOkZkinjCsKpIPfS8FO7YHLSHdEknyec5rSSrw5mj4h4PitF896WjavCbTDq+ eCAi1tlXAGX5vNGLGIxqv0WC5ubhWAE= Received: from mail-wm1-f72.google.com (mail-wm1-f72.google.com [209.85.128.72]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-341-JHUzhCPaOBGQPKrX49-ecg-1; Sun, 02 Aug 2026 15:54:29 -0400 X-MC-Unique: JHUzhCPaOBGQPKrX49-ecg-1 X-Mimecast-MFC-AGG-ID: JHUzhCPaOBGQPKrX49-ecg_1785700469 Received: by mail-wm1-f72.google.com with SMTP id 5b1f17b1804b1-495517d39fbso7651375e9.2 for ; Sun, 02 Aug 2026 12:54:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785700468; x=1786305268; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=xjP8Nj5ayArKITQglbdxDINIj5XZHeMrEUA74eJ1q04=; b=n3v77MB949fhXTyKhcatUQOXICMnb2JMgTmGXXuIB7puFXs1EqORGZEqxOYDcmOS6J YaS3I5OyGAg9tkQmUsZVAVMDo2mXkkN1Z/5OZicdvkD0jD61Q8AIyCiFFl6nY5S6ZnTM MFVHp9dSR97IihgnUuM77Zw7fE3JV73t21gTQLqbFd/VLQITyin+LuVV4QQBwK4aC7eJ QdZAv7JTw+QOLs7eDdtBHbzGkJ5sEloL0QvkyOVe/6LUMrGRpbm0tG5CG4a9Y6trndUr ZE628iexjMQJQuDZiQPXveBWFgjRBvJuKSf42w9URxlsSe5mYqWwrBoUg1ZPy7pMQX6m 082g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785700468; x=1786305268; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=xjP8Nj5ayArKITQglbdxDINIj5XZHeMrEUA74eJ1q04=; b=XW+uId6uxgmXAZsy/bEe4XBazmmFBwDpfl+EUXI01Vr+sXUVgsjL4IBjLSkAKlF3k4 OX2KwT0rlDwqkdz/ksxnn+kPbDtN2PSGByi3nsVden3aonDYE92j8sUMs+1wGUsIDtut XXLln+32KyedrFZRjqkSJhzAKN7nnjx19nXEXThLdp5h/HAgHJjXzGq0tLRdxF+LaNDk njCfB9zNHYea8wAqMoe3br7l7NFpGnB/sU1ww2TajxXR7zunSflvzB8DAUYRON14Ovhv fqAOifKdv2r7lDOR9elVs2TbPsUp29Tto7a+RLeShqR8xg4tn+8RZrTCb/xcfmm4sgPa voLA== X-Forwarded-Encrypted: i=1; AHgh+RqhL8tu5GXVnRPon1VLYH9nOmwCA7t+nlRWCCWsNtjWEMq5hp016uKjpQF7p0drdxSE9OE8uSwZubJHw+E=@vger.kernel.org X-Gm-Message-State: AOJu0YxUhak/DFBhHEfk8XScfckn7eeYFiOqlZ4qaZFAC15yZmGD9KTo XX9sZCoJgYrffjCDtffFaRScandowOWxQRU0Pq626ChVOtpUJgLHa1N4AxhusCotT+3XzBIXfcF 3L4WHuIWvzIMzwnhpReNUO//a1pF5xdCwkA4JcnHACDyUzgb667Qx1BZ6XbtIO+gZWg== X-Gm-Gg: AR+sD10Id9np9soJ92qjb+T0RxFt8Xpsa2cxJtRRkBGaZwVEvIj/6XzjjjpyioB1cQa Q0CZqDX+TfL5nt/z36joa7P97NOveGIYDU9KtrfVFseebD+SRZlZKC58oxwMKHbZ4Ky6u+5A48+ 6ENbcVqqgPGMVJ6v13lB5TU+l/tntW5aiCQzuhQ/A0YWZaJi7xOBRRzWaV6pT9WTSa7sGNWch6E 6BZn1StjGF0SUzSVigojYLS2wJSM9Ten8O3Yz/nu+99py3N3Q/LIcOo8Lo1HIUaBSqEtghsYWJX YX3W6KYex6qhHsR1NtAZxR7mHNzjY2JXkXE7otSsG6RbBtUrKSQK9T4Qun8BTw7kUF64BM1LsN+ nPlzwovIuqfPRcXJVZIRuSg== X-Received: by 2002:a05:600c:8b33:b0:495:779a:ed33 with SMTP id 5b1f17b1804b1-4980c66c798mr174723745e9.7.1785700468578; Sun, 02 Aug 2026 12:54:28 -0700 (PDT) X-Received: by 2002:a05:600c:8b33:b0:495:779a:ed33 with SMTP id 5b1f17b1804b1-4980c66c798mr174723515e9.7.1785700468061; Sun, 02 Aug 2026 12:54:28 -0700 (PDT) Received: from redhat.com (IGLD-80-230-28-14.inter.net.il. [80.230.28.14]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49807b21e66sm144240835e9.0.2026.08.02.12.54.26 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 02 Aug 2026 12:54:27 -0700 (PDT) Date: Sun, 2 Aug 2026 15:54:24 -0400 From: "Michael S. Tsirkin" To: Abhin Parekadan Jose Cc: jasowangio@gmail.com, xuanzhuo@linux.alibaba.com, eperezma@redhat.com, virtualization@lists.linux.dev, linux-kernel@vger.kernel.org Subject: Re: [PATCH 0/2] virtio_pci_modern: fix vp_reset() hang on unresponsive device Message-ID: <20260802155133-mutt-send-email-mst@kernel.org> References: <20260802174059.4082-1-abhinjoses@gmail.com> <20260802134443-mutt-send-email-mst@kernel.org> <20260802145944-mutt-send-email-mst@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Sun, Aug 02, 2026 at 09:48:26PM +0200, Abhin Parekadan Jose wrote: > On Sun, Aug 02, 2026 at 03:08:12PM -0400, Michael S. Tsirkin wrote: > > On Sun, Aug 02, 2026 at 06:28:03PM +0000, Abhin Parekadan Jose wrote: > > > On Sun, Aug 02, 2026 at 01:47:01PM -0400, Michael S. Tsirkin wrote: > > > > On Sun, Aug 02, 2026 at 05:40:57PM +0000, Abhin Parekadan Jose wrote: > > > > > While investigating a syzbot report of a WARN_ON_ONCE firing in > > > > > virtio_dev_remove() [1], > > > > > > > > > > > > And I responded to that syzbot report, and I quote: > > > > > > > > So it writes 0 into pci command, effectively killing the device, > > > > and then is unhappy that the driver prints warnings? > > > > Who thought it's a good idea? Why? > > > > > > I was learning how to reproduce syzbot bugs when I found this > > > issue by writing 0 to PCI_COMMAND to simulate an unresponsive > > > device. > > > > Yea I have no idea where does this syzbot "bug report" > > come from. Poking at random at device registers is ... not > > a very good idea. > > > > > While doing that I noticed that echo 1 > /sys/../remove > > > hung completely rather than just printing the warning. Since the > > > device_status register lives in the virtio common config MMIO > > > space and has defined values(based on the bits set) in the spec. > > > I thought it made sense for virtio to detect this and handle it > > > gracefully rather than spin forever, so I wrote up a small fix > > > for that. > > > > > > > > I found a related but more serious issue: > > > > > vp_reset() in the modern virtio-pci transport can hang indefinitely > > > > > if PCI_COMMAND memory-space decode is disabled while the device is > > > > > bound (e.g. surprise removal, hardware fault, or -- as reproduced > > > > > here -- a direct write to the PCI_COMMAND register). The status > > > > > register poll loop has no way to distinguish "device still resetting" > > > > > from "device unreachable," so it never terminates. > > > > > > > > > > Patch 1 adds a VIRTIO_STATUS_ERROR() check that recognizes an > > > > > all-ones status read as invalid (per spec, bits 4-5 are reserved and > > > > > can never legitimately be set) and warns once at the point the bad > > > > > read actually happens. > > > > > > > > > > Patch 2 uses that check to break out of vp_reset()'s poll loop > > > > > instead of spinning forever. > > > > > > > > Was all this including the cover letter written with ai assistance? > > > > if yes pls disclose this. > > > > > > Yes, I used AI assistance (Claude). The commit messages were written > > > by me and then refined with AI for spelling and grammar; the cover > > > letter was generated by Claude and reviewed by me. > > > > I suggest limiting it to fixing spelling and grammar exclusively. It > > tends to do things like dramatize, e.g. "more serious issue", like it > > did here. > > > > > The code, testing, > > > and debugging were done by me -- I reproduced the hang in QEMU, > > > debugged to reach the hanging loop, and wrote the actual fix. > > > > > > I should have disclosed this upfront. I'll do so in future > > > submissions. > > > > > > Do I need to add Assisted-by: Claude to the > > > commit messages? > > > > Assisted-by: Claude:claude-sonnet-4-6 > > > > > > > > > > P.S. This is my first kernel patch set. > > > > > > Thanks, keep at it. Bonus points if you find a real fix for > > issues raised in thread about surprise removal, see e.g. here > > cover.1752094439.git.mst@redhat.com > > but don't expect it to be easy. > > > > That looks interesting (haven't gone through in detail but got a gist of it). > I'll try it out and make suggestions if I find a good solution. > > As for the current patch set, does it make sense to drop macro > `VIRTIO_STATUS_ERROR` and use `PCI_POSSIBLE_ERROR` to break out of the loop, > I could test it by actually doing a surprise removal on qemu via the monitor > or just drop this patch set I'd drop this, I'm not interested in working around one source of hangs if others in the same exact path remain unfixable. > and try on cover.1752094439.git.mst@redhat.com patch set? it's not a question of "trying it on" it's a question of the fact that the pci core serializes probe/removal events so a driver inside probe/remove never sees the removal event. in this instance it is polling so it can check (at the cost of adding cpu overhead, mostly for nothing) but in most places it can't, we need the event to reach it. > > -- > > MST > >