From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id C458FC88E41 for ; Thu, 10 Sep 2026 22:32:37 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:In-Reply-To:Content-Type: MIME-Version:References:Message-ID:Subject:Cc:To:From:Date:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=AvaXwx2LcNRb7aleMdEJjUSF+MigurOCyRurKBuzJ8I=; b=oVmdBZrATFuG85ulLdTo+ebuEr qH3xPzet1LJ9hfEawgqgZFk1o/vUhF0ntpe1rqdH6NydPvnrsfNLprLDIDghinNrEuw7lMxv7w5yd 29IFv65FEoTwORlyxTfiqk6qQbiEdLsgZl3HrZGc41gR0jRnNK3CUuTyFIVPhZLcP10dj8DY8eQqU NRVsKHL/0zTVO5bIC89ZbvEHa4fZFQzCw4go4dREspzZ/kzyHax/MaizqDArXxDppmIldtB7Z3Jj3 RR5q54L46REHT5f+Lz/O1Tn+l/YfioZD+0tM3Eb6VjMHSJdw7wbOjukUQJJND3RFM0A9pDA5wctnP Yda8LS9w==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1x4nJr-0000000FUNW-2z58; Thu, 10 Sep 2026 22:32:35 +0000 Received: from sea.source.kernel.org ([172.234.252.31]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1x4nJp-0000000FUN9-3fmC for linux-nvme@lists.infradead.org; Thu, 10 Sep 2026 22:32:33 +0000 Received: from smtp.kernel.org (quasi.space.kernel.org [100.103.45.18]) by sea.source.kernel.org (Postfix) with ESMTP id 3A16C436A2; Thu, 10 Sep 2026 22:32:33 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id C425C1F000FF; Thu, 10 Sep 2026 22:32:32 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1789079553; bh=AvaXwx2LcNRb7aleMdEJjUSF+MigurOCyRurKBuzJ8I=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=gW3tTTD7xAgWUf3poPvahx39VxXOsCnnGGymP4z3rdxXzvuHYr1cd/ysJqDKY0Ify +o8/qUIVV8jS2UCrfSlgH+afQbeZESOlvsfFghCqyN4bkflGgKnYBYHud3oV5UL5FN EN5m5fbSzbDWj2Pok4JIErgF4Rq79ZLXZ7ypSpvAVjORdUFxlrvJYhzvanVG9kjK+r IzFD//vk/zxWs/CkWESmhXx7CFOI8D9gmd4yk586iIBk85SjBZJe5u+xJ/K8fobfIT zp1aVZT7PqMwuno+sf6IX6zCcdFvzAYbKDYWJNXLBa3ZYHIMfOBhsDFGWoKZcYbjnp uHxrFEW7vgYhA== Date: Thu, 10 Sep 2026 16:32:31 -0600 From: Keith Busch To: Shivam Kumar Cc: linux-nvme@lists.infradead.org, Greg KH , security@kernel.org, Christoph Hellwig , Sagi Grimberg , Chaitanya Kulkarni , stable@vger.kernel.org Subject: Re: [PATCH] nvmet-tcp: fix a hang on queue teardown with data digest Message-ID: References: <20260901014718.2835558-1-kumar.shivam43666@gmail.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260901014718.2835558-1-kumar.shivam43666@gmail.com> X-BeenThere: linux-nvme@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "Linux-nvme" Errors-To: linux-nvme-bounces+linux-nvme=archiver.kernel.org@lists.infradead.org On Mon, Aug 31, 2026 at 09:47:18PM -0400, Shivam Kumar wrote: > With data digest on, a command that has received all its data waits for > the digest in NVMET_TCP_RECV_DDGST. has_data_in() is already false there, > so need_data_in() is too, and nvmet_tcp_uninit_data_in_cmds() skips it on > teardown even though it still holds the nvmet_req_init() reference. The SQ > percpu_ref never drains, nvmet_sq_destroy() blocks forever in > wait_for_completion(), and the nvmet-wq release worker is stuck. > > An unauthenticated host on an allow_any_host subsystem hits this by > negotiating data digest, sending a write's data but not the trailing > digest, and closing the connection. Each leaked command wedges a release > worker; a few stall queue teardown entirely, and with hung_task_panic the > box goes down. > > Drop the reference for a command left in RECV_DDGST from > nvmet_tcp_release_queue_work(), before rcv_state is cleared. A command > that failed nvmet_req_init() never took one, so skip it. "I have made this letter longer than usual because I lack the time to make it shorter." - Blaise Pascal Take some time to actually write your own (and hopefully more concise) message to demonstrate you understand what you're changing. AI messages are overly verbose. > + /* > + * A command that has received all of its data and is only waiting > + * for the data digest sits in RECV_DDGST: need_data_in() is already > + * false, so nvmet_tcp_uninit_data_in_cmds() below skips it, yet it > + * still holds the reference from nvmet_req_init(). Drop it here, > + * while rcv_state still reflects it, so the SQ percpu_ref can drain > + * and nvmet_sq_destroy() can complete. INIT_FAILED took no ref. > + */ Same here. > + if (queue->rcv_state == NVMET_TCP_RECV_DDGST && queue->cmd && > + !(queue->cmd->flags & NVMET_TCP_F_INIT_FAILED)) > + nvmet_req_uninit(&queue->cmd->req); I don't think this criteria is even correct. RECV_DDGST doesn't mean the command has all its data since that state is per-PDU. nvmet_tcp_try_recv_data() enters that state at every H2C boundary, not just the last one.