Linux-mediatek Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: <alice.chao@mediatek.com>
To: <stable@vger.kernel.org>
Cc: <linux-scsi@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<martin.petersen@oracle.com>,
	<James.Bottomley@HansenPartnership.com>, <bvanassche@acm.org>,
	<avri.altman@sandisk.com>, <alim.akhtar@samsung.com>,
	<wsd_upstream@mediatek.com>, <linux-mediatek@lists.infradead.org>,
	<peter.wang@mediatek.com>, <chun-hung.wu@mediatek.com>,
	<alice.chao@mediatek.com>, <cc.chou@mediatek.com>,
	<chaotian.jing@mediatek.com>, <tun-yu.yu@mediatek.com>,
	<naomi.chu@mediatek.com>, <ed.tsai@mediatek.com>
Subject: [PATCH 6.18.y] scsi: ufs: core: Re-arm the device command completion before submitting
Date: Tue, 15 Sep 2026 13:26:37 +0800	[thread overview]
Message-ID: <20260915052638.459390-1-alice.chao@mediatek.com> (raw)

From: Alice Chao <alice.chao@mediatek.com>

Commit 20b97acc4caf ("scsi: ufs: core: Fix a race condition related to
device commands") moved the device management command completion into
struct ufs_hba, initialized once by ufshcd_init(), and dropped the

	hba->dev_cmd.complete = NULL;

assignments that used to make ufshcd_compl_one_cqe() discard completions
the submitter had already given up on. Nothing replaced them, so the
completion is never reset between two device commands:

  Task A (device command submitter) IRQ (tag == hba->reserved_slot)
  --------------------------------- ------------------------------
  ufshcd_read_desc_param()
   ufshcd_query_descriptor_retry()
    __ufshcd_query_descriptor()
     ufshcd_exec_dev_cmd()
      ufshcd_issue_dev_cmd()
       ufshcd_send_command()
       ufshcd_wait_for_dev_cmd()
        wait_for_completion_timeout()
        /* times out, done == 0 */
        ufshcd_clear_cmd()
        /* returns 0, no effect */
       return -EAGAIN
                                    ufs_mtk_mcq_intr()
                                     ufshcd_mcq_poll_cqe_lock()
                                      ufshcd_mcq_process_cqe()
                                       ufshcd_compl_one_cqe()
                                        /* lrbp->cmd == NULL */
                                        complete(&hba->dev_cmd.complete)
                                        /* done: 0 -> 1 */
  ufshcd_query_attr_retry()
   ufshcd_query_attr()
    ufshcd_exec_dev_cmd()
     ufshcd_issue_dev_cmd()
      ufshcd_send_command()
      ufshcd_wait_for_dev_cmd()
       wait_for_completion_timeout()
       /* returns at once, done: 1 -> 0 */
       ufshcd_dev_cmd_completion()
       /* response UPIU not written yet */
       return -EINVAL

Task A then rejects what it reads out of the response UPIU:

  ufshcd_dev_cmd_completion: Invalid device management cmd response: 0
  ufshcd_dev_cmd_completion: unexpected response in Query RSP: ff

The skew does not self-correct. On a UFS 4.0 controller in MCQ mode it
persisted across more than a thousand consecutive device commands,
failing every descriptor and attribute read until the link was reset.

The controller can still complete the timed-out command because in MCQ
mode ufshcd_clear_cmd() only issues an SQ cleanup (SQRTC.ICU), which
shows the command left the submission queue but not that a CQE is not
already posted. The MCQ path also skips the hba->outstanding_reqs
re-check that the SDB path does, so it returns -EAGAIN with the
completion still armed.

Re-arm the completion in ufshcd_issue_dev_cmd(), immediately before
submitting. All submitters - ufshcd_exec_dev_cmd(),
ufshcd_issue_devman_upiu_cmd() and ufshcd_advanced_rpmb_op() - reach it
holding hba->dev_cmd.lock, so no extra serialization is needed.

This narrows the window rather than closing it: the CQE only carries the
tag and every device command uses hba->reserved_slot, so a late
completion is still indistinguishable from the expected one. It no
longer spans the idle time between two commands.

No mainline commit: commit 08b12cda6c44 ("scsi: ufs: core: Switch to
scsi_get_internal_cmd()") moved this path onto the block layer and
removed struct ufs_dev_cmd::complete and ufshcd_wait_for_dev_cmd(). Each
device command now waits on its own request via blk_execute_rq(), so
mainline has no shared completion to skew. That refactor is not a
reasonable stable backport; this is the minimal alternative.

Affected versions: v6.15 through v6.18, i.e. the kernels that carry the
commit named in the Fixes: tag but not the mainline rewrite above.

Fixes: 20b97acc4caf ("scsi: ufs: core: Fix a race condition related to device commands")
Cc: stable@vger.kernel.org
Signed-off-by: Alice Chao <alice.chao@mediatek.com>
---
 drivers/ufs/core/ufshcd.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/drivers/ufs/core/ufshcd.c b/drivers/ufs/core/ufshcd.c
index 87578e8824d2..003fa8af4f4d 100644
--- a/drivers/ufs/core/ufshcd.c
+++ b/drivers/ufs/core/ufshcd.c
@@ -3312,6 +3312,15 @@ static int ufshcd_issue_dev_cmd(struct ufs_hba *hba, struct ufshcd_lrb *lrbp,
 {
 	int err;
 
+	/*
+	 * A device command that timed out may still be completed by the
+	 * controller later on. hba->dev_cmd.complete is shared by all device
+	 * commands, so re-arm it here, immediately before submitting, to keep
+	 * such a late completion from being mistaken for the completion of
+	 * this command.
+	 */
+	reinit_completion(&hba->dev_cmd.complete);
+
 	ufshcd_add_query_upiu_trace(hba, UFS_QUERY_SEND, lrbp->ucd_req_ptr);
 	ufshcd_send_command(hba, tag, hba->dev_cmd_queue);
 	err = ufshcd_wait_for_dev_cmd(hba, lrbp, timeout);
-- 
2.45.2



             reply	other threads:[~2026-09-15  5:27 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-15  5:26 alice.chao [this message]
2026-09-16  1:52 ` [PATCH 6.18.y] scsi: ufs: core: Re-arm the device command completion before submitting Sasha Levin
2026-10-05  8:34   ` Alice Chao
2026-10-05 11:54 ` Bean Huo
2026-10-06  9:55   ` Alice Chao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260915052638.459390-1-alice.chao@mediatek.com \
    --to=alice.chao@mediatek.com \
    --cc=James.Bottomley@HansenPartnership.com \
    --cc=alim.akhtar@samsung.com \
    --cc=avri.altman@sandisk.com \
    --cc=bvanassche@acm.org \
    --cc=cc.chou@mediatek.com \
    --cc=chaotian.jing@mediatek.com \
    --cc=chun-hung.wu@mediatek.com \
    --cc=ed.tsai@mediatek.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mediatek@lists.infradead.org \
    --cc=linux-scsi@vger.kernel.org \
    --cc=martin.petersen@oracle.com \
    --cc=naomi.chu@mediatek.com \
    --cc=peter.wang@mediatek.com \
    --cc=stable@vger.kernel.org \
    --cc=tun-yu.yu@mediatek.com \
    --cc=wsd_upstream@mediatek.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox