Linux block layer
 help / color / mirror / Atom feed
From: "syzbot" <syzbot@kernel.org>
To: syzkaller-bugs@googlegroups.com,
	Kusaram Devineni <kusaram@devineni.in>,
	"Jens Axboe" <axboe@kernel.dk>,
	"Josef Bacik" <josef@toxicpanda.com>,
	<linux-block@vger.kernel.org>, <nbd@other.debian.org>
Cc: linux-kernel@vger.kernel.org, syzbot@lists.linux.dev
Subject: [PATCH] nbd: fix I/O hang on dead socket and rate-limit console error
Date: Fri, 17 Jul 2026 08:58:18 +0000 (UTC)	[thread overview]
Message-ID: <c4c2aa69-a3bc-4d1b-96ed-64ae891f921e@mail.kernel.org> (raw)

From: Kusaram Devineni <kusaram@devineni.in>

When an NBD device is configured with a timeout of 0, a closed socket can
lead to a permanent I/O hang. The sequence is as follows: a request is sent
and marked in-flight, but then the socket becomes dead (e.g., due to a
connection failure). Since the socket is dead, no reply will ever arrive.
In nbd_xmit_timeout(), if the configured timeout is 0, the code currently
only checks if the socket has been replaced by comparing cookies. If the
cookie still matches, the request timer is unconditionally reset and the
request stays in-flight forever. This causes tasks to hang indefinitely in
TASK_UNINTERRUPTIBLE, triggering the hung task detector:

INFO: task udevd:5915 blocked in I/O wait for more than 143 seconds.
Call Trace:
 <TASK>
 schedule+0x164/0x2b0
 io_schedule+0x7f/0xd0
 folio_wait_bit_common+0x836/0xbc0
 do_read_cache_folio+0x1ac/0x590
 read_part_sector+0xb6/0x2b0
 adfspart_check_POWERTEC+0x9a/0x7a0
 bdev_disk_changed+0x851/0x17a0
 blkdev_get_whole+0x372/0x510
 bdev_open+0x324/0xd70
 ...
 </TASK>

Fix this by checking nsock->dead in addition to the cookie check in
nbd_xmit_timeout(). If the socket is dead, the command is requeued.
Requeuing returns the request to the existing fallback, reconnect, or
failure policy. For the reported configuration where no reconnect is
possible, this results in the request being properly failed, which
terminates the hung operation.

A secondary issue was observed where repeated attempts to connect to an
already-in-use NBD device cause console spam because the "nbd%d already in
use" error message in nbd_genl_connect() is not rate-limited. This can
delay console_unlock() and trigger NMI backtraces. This is fixed by
changing the pr_err() to pr_err_ratelimited().

Fixes: 2c272542baee ("nbd: requeue command if the soecket is changed")
Assisted-by: Gemini:gemini-3.1-pro-preview Gemini:gemini-3-flash-preview syzbot
Reported-by: syzbot+82de77d3f217960f087d@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=82de77d3f217960f087d
Link: https://syzkaller.appspot.com/ai_job?id=40e56b38-627e-4b49-b7ac-d3a21e77dba2
Signed-off-by: Kusaram Devineni <kusaram@devineni.in>

---
diff --git a/drivers/block/nbd.c b/drivers/block/nbd.c
index 8f10762e9..eedb1c870 100644
--- a/drivers/block/nbd.c
+++ b/drivers/block/nbd.c
@@ -523,7 +523,7 @@ static enum blk_eh_timer_return nbd_xmit_timeout(struct request *req)
 			blk_rq_bytes(req), (req->timeout / HZ) * cmd->retries);
 
 		mutex_lock(&nsock->tx_lock);
-		if (cmd->cookie != nsock->cookie) {
+		if (cmd->cookie != nsock->cookie || nsock->dead) {
 			nbd_requeue_cmd(cmd);
 			mutex_unlock(&nsock->tx_lock);
 			mutex_unlock(&cmd->lock);
@@ -2172,7 +2172,7 @@ static int nbd_genl_connect(struct sk_buff *skb, struct genl_info *info)
 		nbd_put(nbd);
 		if (index == -1)
 			goto again;
-		pr_err("nbd%d already in use\n", index);
+		pr_err_ratelimited("nbd%d already in use\n", index);
 		return -EBUSY;
 	}
 


base-commit: 8cdeaa50eae8dad34885515f62559ee83e7e8dda
-- 
See https://goo.gle/syzbot-ai-patches for information about AI-generated patches.
You can comment on the patch as usual, syzbot will try to address
the comments and send a new version of the patch if necessary.
syzbot engineers can be reached at syzkaller@googlegroups.com.

                 reply	other threads:[~2026-07-17  8:58 UTC|newest]

Thread overview: [no followups] expand[flat|nested]  mbox.gz  Atom feed

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c4c2aa69-a3bc-4d1b-96ed-64ae891f921e@mail.kernel.org \
    --to=syzbot@kernel.org \
    --cc=axboe@kernel.dk \
    --cc=josef@toxicpanda.com \
    --cc=kusaram@devineni.in \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=nbd@other.debian.org \
    --cc=syzbot@lists.linux.dev \
    --cc=syzkaller-bugs@googlegroups.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox