Linux SCSI subsystem development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Nilesh Javali" <njavali@marvell.com>
Cc: linux-scsi@vger.kernel.org
Subject: Re: [PATCH v6 51/56] scsi: qla2xxx: Initialize NVMe abort_work once at submission
Date: Thu, 23 Jul 2026 19:26:19 +0000	[thread overview]
Message-ID: <20260723192619.EF8971F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260723050413.3897522-52-njavali@marvell.com>

Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:

New issues:
- [Critical] The patch moves INIT_WORK to the command submission paths but fails to order it safely against the publication of priv->sp, causing a NULL pointer dereference.
- [High] Calling qla_nvme_fcp_abort or qla_nvme_ls_abort multiple times for the same command leaks its kref and permanently hangs the request.

Pre-existing issues:
- [Critical] The command state (priv->sp) is exposed to concurrent abort handlers before the spinlock (priv->cmd_lock) is initialized, leading to spinlock corruption.
- [Critical] In the error paths of command submission, priv->sp is cleared without holding cmd_lock, and the command structure is freed directly, causing a Use-After-Free.
- [High] A race condition allows qla_nvme_fcp_abort to spuriously abort a completely different, newly submitted command due to tag reuse (ABA problem).
--

--- Patch [51]: [PATCH v6 51/56] scsi: qla2xxx: Initialize NVMe abort_work once at submission ---
commit 9a9f1f14cd9974710e9176cf206ade9b1d04e0e9
Author: Nilesh Javali <njavali@marvell.com>

scsi: qla2xxx: Initialize NVMe abort_work once at submission

This patch moves the initialization of priv->abort_work to the command
submission path to avoid corrupting the workqueue list if an abort is
issued more than once for the same command.

> diff --git a/drivers/scsi/qla2xxx/qla_nvme.c b/drivers/scsi/qla2xxx/qla_nvme.c
> index 0038b6274d444..3b2f255a5d7d4 100644
> --- a/drivers/scsi/qla2xxx/qla_nvme.c
> +++ b/drivers/scsi/qla2xxx/qla_nvme.c
> @@ -463,7 +463,6 @@ static void qla_nvme_ls_abort(struct nvme_fc_local_port *lport,
>  	}
>  	spin_unlock_irqrestore(&priv->cmd_lock, flags);
>  
> -	INIT_WORK(&priv->abort_work, qla_nvme_abort_work);
>  	schedule_work(&priv->abort_work);
>  }

[Severity: High]
If the transport layer calls the abort functions multiple times for the same
command, is there a kref leak here?

Earlier in this function, we unconditionally increment the reference count via
kref_get_unless_zero(&priv->sp->cmd_kref).

Since the work item is no longer re-initialized, schedule_work() will return
false and do nothing if the work is already pending. However, the return
value is ignored and the extra kref taken is never dropped.

Since the work function qla_nvme_abort_work() only executes once and drops a
single kref, will this permanently hang the NVMe core's I/O queue by
preventing the kref from ever reaching zero?

[ ... ]

> @@ -501,6 +500,7 @@ static int qla_nvme_ls_req(struct nvme_fc_local_port *lport,
>  	priv->sp = sp;
>  	kref_init(&sp->cmd_kref);
>  	spin_lock_init(&priv->cmd_lock);
> +	INIT_WORK(&priv->abort_work, qla_nvme_abort_work);
>  	nvme = &sp->u.iocb_cmd;
>  	priv->fd = fd;

[Severity: Critical]
Could this sequence cause a panic due to an uninitialized work item?

We publish the command state with priv->sp = sp explicitly before INIT_WORK()
executes. Because the transport allocates the private data with kzalloc(), the
work struct's function pointer is initially NULL.

If a concurrent error recovery thread calls qla_nvme_ls_abort() while
qla_nvme_ls_req() is executing, it can observe priv->sp != NULL and call
schedule_work(&priv->abort_work).

Will this queue a work item with a NULL function pointer, causing a kernel
panic when the worker thread executes it?

[Severity: Critical]
This is a pre-existing issue, but does initializing the spinlock after
publishing priv->sp lead to lock corruption?

Similar to the work item, priv->sp is exposed to concurrent abort handlers
before spin_lock_init(&priv->cmd_lock) is called. A concurrent abort from
nvme_fc_abort_lsops can see priv->sp != NULL and successfully acquire the
0-initialized lock.

If the submission thread then unconditionally resets the lock internals back
to 0 while the abort handler is inside its critical section, will it cause
memory corruption or lockdep panics when the abort handler tries to unlock it?

[Severity: Critical]
This is a pre-existing issue, but can the error path further down in this
function cause a use-after-free if an abort runs concurrently?

If hardware submission fails, the error path in qla_nvme_ls_req() does this:

    rval = qla2x00_start_sp(sp);
    if (rval != QLA_SUCCESS) {
        ...
        sp->priv = NULL;
        priv->sp = NULL;
        qla2x00_rel_sp(sp);
        return rval;
    }

It sets priv->sp = NULL without acquiring priv->cmd_lock, bypasses kref_put(),
and immediately frees the command structure via qla2x00_rel_sp(). 

If a concurrent abort handler has already checked priv->sp != NULL and
incremented the kref inside qla_nvme_ls_abort(), it is left holding a
reference to freed memory. When the queued abort_work runs, won't it invoke
abort logic on the freed command structure?

[ ... ]

> @@ -545,7 +545,6 @@ static void qla_nvme_fcp_abort(struct nvme_fc_local_port *lport,
>  	}
>  	spin_unlock_irqrestore(&priv->cmd_lock, flags);
>  
> -	INIT_WORK(&priv->abort_work, qla_nvme_abort_work);
>  	schedule_work(&priv->abort_work);
>  }

[Severity: High]
This is a pre-existing issue, but is qla_nvme_fcp_abort() vulnerable to a tag
reuse race condition?

In the NVMe FC core, __nvme_fc_abort_op() initiates an abort without
synchronizing with the I/O completion path. If a hardware completion interrupt
fires and completes the request, the block layer can free and immediately reuse
the same tag (and thus the same fd->private memory) for a completely new
command.

When qla_nvme_fcp_abort() runs and acquires priv->cmd_lock:

    static void qla_nvme_fcp_abort(...)
    {
        ...
        spin_lock_irqsave(&priv->cmd_lock, flags);

it will see the new priv->sp. Does this cause a spurious abort of the newly
submitted, innocent command?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260723050413.3897522-1-njavali@marvell.com?part=51

  reply	other threads:[~2026-07-23 19:26 UTC|newest]

Thread overview: 96+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-23  5:03 [PATCH v6 00/56] scsi: qla2xxx: Add QLA29xx series adapter support Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 01/56] scsi: qla2xxx: Add 29xx series PCI device ID support Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 02/56] scsi: qla2xxx: Add flash read/write interface for 29xx Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 03/56] scsi: qla2xxx: Add NVRAM config support for 29xx adapters Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 04/56] scsi: qla2xxx: Add 29xx support in queue initialisation path Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 05/56] scsi: qla2xxx: Add FC operational firmware load for 29xx Nilesh Javali
2026-07-23  6:44   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 06/56] scsi: qla2xxx: Remove redundant VPD flash read in sysfs read path Nilesh Javali
2026-07-23  6:53   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 07/56] scsi: qla2xxx: Add flash block read/write BSG support for 29xx Nilesh Javali
2026-07-23  7:11   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 08/56] scsi: qla2xxx: Add BSG MPI firmware load/dump " Nilesh Javali
2026-07-23  7:24   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 09/56] scsi: qla2xxx: Add 128-byte IOCB definitions " Nilesh Javali
2026-07-23  7:35   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 10/56] scsi: qla2xxx: Add extended status continuation and marker IOCBs Nilesh Javali
2026-07-23  7:43   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 11/56] scsi: qla2xxx: Update IO path to use 128-byte IOCBs for 29xx Nilesh Javali
2026-07-23  8:25   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 12/56] scsi: qla2xxx: Skip image-set-valid attribute " Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 13/56] scsi: qla2xxx: Skip unsupported sysfs attributes " Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 14/56] scsi: qla2xxx: Enable get_fw_version mailbox " Nilesh Javali
2026-07-23  9:15   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 15/56] scsi: qla2xxx: Extend execute_fw mailbox to include 29xx Nilesh Javali
2026-07-23  9:26   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 16/56] scsi: qla2xxx: Enable get_adapter_id mailbox for 29xx Nilesh Javali
2026-07-23  9:36   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 17/56] scsi: qla2xxx: Enable init_firmware " Nilesh Javali
2026-07-23  9:45   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 18/56] scsi: qla2xxx: Enable get_firmware_state " Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 19/56] scsi: qla2xxx: Enable serdes, resource count and FCE trace " Nilesh Javali
2026-07-23 10:10   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 20/56] scsi: qla2xxx: Enable set_els_cmds and echo_test " Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 21/56] scsi: qla2xxx: Add support for QLA29XX in data rate functions Nilesh Javali
2026-07-23 10:25   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 22/56] scsi: qla2xxx: Enable qla2x00_shutdown for 29xx Nilesh Javali
2026-07-23 10:34   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 23/56] scsi: qla2xxx: Use ring-slot helpers in __qla2x00_alloc_iocbs Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 24/56] scsi: qla2xxx: Add support for QLA29XX in memory allocation Nilesh Javali
2026-07-23 10:56   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 25/56] scsi: qla2xxx: Handle sts_cont_entry_ext_t for 29xx adapters Nilesh Javali
2026-07-23 11:12   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 26/56] scsi: qla2xxx: Update handling of status entries for 29xx series Nilesh Javali
2026-07-23 13:25   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 27/56] scsi: qla2xxx: Enhance ct_entry_24xx_ext iocb handling " Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 28/56] scsi: qla2xxx: Enhance purex_entry " Nilesh Javali
2026-07-23 14:09   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 29/56] scsi: qla2xxx: Update handling of ELS IOCBs " Nilesh Javali
2026-07-23 14:25   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 30/56] scsi: qla2xxx: Add size check for ELS status entry layout on 29xx Nilesh Javali
2026-07-23 14:39   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 31/56] scsi: qla2xxx: Add 29xx extended logio IOCB support Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 32/56] scsi: qla2xxx: Enhance task management IOCB handling for 29xx series Nilesh Javali
2026-07-23 15:08   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 33/56] scsi: qla2xxx: Add abort command " Nilesh Javali
2026-07-23 15:28   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 34/56] scsi: qla2xxx: Enhance ABTS processing " Nilesh Javali
2026-07-23 15:53   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 35/56] scsi: qla2xxx: Update VP control IOCB handling " Nilesh Javali
2026-07-23 16:12   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 36/56] scsi: qla2xxx: Add build-time size check for VP config IOCB layout Nilesh Javali
2026-07-23  5:03 ` [PATCH v6 37/56] scsi: qla2xxx: Add size check for extended VP report ID entry Nilesh Javali
2026-07-23 16:31   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 38/56] scsi: qla2xxx: Add LS4 pass-through IOCB handling for 29xx series Nilesh Javali
2026-07-23 16:47   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 39/56] scsi: qla2xxx: Adjust feature gating in BSG paths for 29xx support Nilesh Javali
2026-07-23 17:05   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 40/56] scsi: qla2xxx: Fix queue teardown NULL dma_free and bitmap locking Nilesh Javali
2026-07-23 17:16   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 41/56] scsi: qla2xxx: Replace __le16 bitfields with scalar and accessors Nilesh Javali
2026-07-23 17:29   ` sashiko-bot
2026-07-23  5:03 ` [PATCH v6 42/56] scsi: qla2xxx: Fix endianness annotations in vp_rpt_id_entry structures Nilesh Javali
2026-07-23  5:04 ` [PATCH v6 43/56] scsi: qla2xxx: Use 64-bit FPM word counters for 29xx host stats Nilesh Javali
2026-07-23  5:04 ` [PATCH v6 44/56] scsi: qla2xxx: Add 64G/128G port speed setting support Nilesh Javali
2026-07-23 18:07   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 45/56] scsi: qla2xxx: Fix 64G link speed reporting in get_data_rate Nilesh Javali
2026-07-23 18:17   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 46/56] scsi: qla2xxx: edif: Fix NULL pointer deref in RX SA delete check Nilesh Javali
2026-07-23  5:04 ` [PATCH v6 47/56] scsi: qla2xxx: Fix Name Server logout detection on FWI2 adapters Nilesh Javali
2026-07-23  5:04 ` [PATCH v6 48/56] scsi: qla2xxx: Bound VP index against VP_CTRL IOCB bitmap size Nilesh Javali
2026-07-23  5:04 ` [PATCH v6 49/56] scsi: qla2xxx: Check entry_status in qla24xx_modify_vp_config() Nilesh Javali
2026-07-23 18:56   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 50/56] scsi: qla2xxx: Hold vport reference in qla24xx_report_id_acquisition() Nilesh Javali
2026-07-23 19:12   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 51/56] scsi: qla2xxx: Initialize NVMe abort_work once at submission Nilesh Javali
2026-07-23 19:26   ` sashiko-bot [this message]
2026-07-23  5:04 ` [PATCH v6 52/56] scsi: qla2xxx: Hold qpair lock when sending NVMe LS reject Nilesh Javali
2026-07-23 19:36   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 53/56] scsi: qla2xxx: Zero dport diagnostics buffer to avoid info leak Nilesh Javali
2026-07-23 19:49   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 54/56] scsi: qla2xxx: Fix BSG job leak on validate flash image error path Nilesh Javali
2026-07-23 20:05   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 55/56] scsi: qla2xxx: Bound image count in qla2x00_update_fru_versions() Nilesh Javali
2026-07-23 20:16   ` sashiko-bot
2026-07-23  5:04 ` [PATCH v6 56/56] scsi: qla2xxx: Update version to 12.00.00.2607b1 Nilesh Javali
2026-07-23 20:27   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260723192619.EF8971F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-scsi@vger.kernel.org \
    --cc=njavali@marvell.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox