Linux SCSI subsystem development
 help / color / mirror / Atom feed
* Re: [PATCH v5 06/10] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
  2026-09-28 18:17 ` [PATCH v5 06/10] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
@ 2026-09-28 18:09   ` sashiko-bot
  0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-09-28 18:09 UTC (permalink / raw)
  To: Nigel Kirkland; +Cc: linux-scsi

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] lpfc: UAF and rport leak in lpfc_register_remote_port() due to race with devloss
- [High] lpfc: Duplicate NVMe transport registrations in lpfc_nlp_reg_node() due to missing synchronization
--

commit 01b3084183c4a09b5084e2e3651c384b815d645a
Author: Nigel Kirkland <nkirkland2304@gmail.com>

lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence

The commit refactors transport registration flags to separate SCSI and
NVME states, preventing premature release of node objects during transport
callbacks. It also introduces symmetric reference-count tracking for NVMET
target nodes.

> diff --git a/drivers/scsi/lpfc/lpfc_hbadisc.c b/drivers/scsi/lpfc/lpfc_hbadisc.c
> --- a/drivers/scsi/lpfc/lpfc_hbadisc.c
> +++ b/drivers/scsi/lpfc/lpfc_hbadisc.c

[ ... ]

> @@ -4481,26 +4481,36 @@ lpfc_register_remote_port(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
>  	struct lpfc_hba  *phba = vport->phba;
>  	unsigned long flags;
>  
> -	if (vport->cfg_enable_fc4_type == LPFC_ENABLE_NVME)
> +	/* Don't add the remote port if unloading. */
> +	if (test_bit(FC_UNLOADING, &vport->load_flag))
>  		return;
>  
> +	spin_lock_irqsave(&ndlp->lock, flags);
> +	if ((ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | SCSI_XPT_UNREG_WAIT)) ==
> +	    SCSI_XPT_REGD) {
> +		/* Prevent racing registrations */
> +		spin_unlock_irqrestore(&ndlp->lock, flags);
> +		return;
> +	}
> +	ndlp->fc4_xpt_flags &= ~SCSI_XPT_UNREG_WAIT;
> +	ndlp->fc4_xpt_flags |= SCSI_XPT_REGD;
> +	spin_unlock_irqrestore(&ndlp->lock, flags);

[Severity: High]
Could this flag manipulation open a race window with devloss? By setting 
SCSI_XPT_REGD and dropping ndlp->lock before calling fc_remote_port_add(), 
ndlp->rport still points to the old rport.

If a concurrent dev_loss_work executes lpfc_dev_loss_tmo_callbk() on the
old rport before the new one is assigned, it will observe SCSI_XPT_REGD as
set. 

If the driver is unloading, lpfc_nlp_put() is called prematurely here,
causing a use-after-free when lpfc_register_remote_port() resumes.
Otherwise, the callback clears SCSI_XPT_REGD, leaving the new rport
permanently unregistered and leaked since subsequent unregister calls will
bypass removal.

>  	/* Remote port has reappeared. Re-register w/ FC transport */
>  	rport_ids.node_name = wwn_to_u64(ndlp->nlp_nodename.u.wwn);
>  	rport_ids.port_name = wwn_to_u64(ndlp->nlp_portname.u.wwn);

[ ... ]

> @@ -4623,7 +4640,10 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
>  	lpfc_check_nlp_post_devloss(vport, ndlp);
>  
>  	spin_lock_irqsave(&ndlp->lock, iflags);
> -	if (ndlp->fc4_xpt_flags & NLP_XPT_REGD) {
> +	if (((ndlp->fc4_xpt_flags & SCSI_XPT_REGD) &&
> +	     !(ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT)) ||
> +	    ((ndlp->fc4_xpt_flags & NVME_XPT_REGD) &&
> +	     !(ndlp->fc4_xpt_flags & NVME_XPT_UNREG_WAIT))) {
>  		/* Already registered with backend, trigger rescan */
>  		spin_unlock_irqrestore(&ndlp->lock, iflags);

[ ... ]

> -	ndlp->fc4_xpt_flags |= NLP_XPT_REGD;
>  	spin_unlock_irqrestore(&ndlp->lock, iflags);

[ ... ]

> -	if (vport->phba->sli_rev >= LPFC_SLI_REV4 &&
> -			ndlp->nlp_fc4_type & NLP_FC4_NVME) {
> -		if (vport->phba->nvmet_support == 0) {
> -			/* Register this rport with the transport.
> -			 * Only NVME Target Rports are registered with
> -			 * the transport.
> -			 */
> -			if (ndlp->nlp_type & NLP_NVME_TARGET) {
> -				vport->phba->nport_event_cnt++;
> -				lpfc_nvme_register_port(vport, ndlp);
> +	if (vport->phba->nvmet_support == 0) {
> +		/* Register this rport with the transport.
> +		 * Only NVME Target Rports are registered with
> +		 * the transport.
> +		 */
> +		if (ndlp->nlp_type & NLP_NVME_TARGET) {
> +			vport->phba->nport_event_cnt++;
> +			lpfc_nvme_register_port(vport, ndlp);

[Severity: High]
Does removing the unified NLP_XPT_REGD flag setting before the unlock
allow duplicate NVMe transport registrations?

By replacing the single check with separate SCSI and NVME flag checks and
dropping ndlp->lock immediately, concurrent executions of lpfc_nlp_reg_node()
(for instance, from RSCN spam) can both observe NVME_XPT_REGD as 0 and
proceed to call lpfc_nvme_register_port().

Since lpfc_nvme_register_port() does not internally synchronize and check
NVME_XPT_REGD before calling nvme_fc_register_remoteport(), both threads
will register a new remoteport and increment the ndlp reference count. The
first remoteport will be overwritten in ndlp->nrport, permanently leaking
the transport object and its reference.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260928181757.21959-1-nkirkland2304@gmail.com?part=6

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v5 09/10] lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI
  2026-09-28 18:17 ` [PATCH v5 09/10] lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI Nigel Kirkland
@ 2026-09-28 18:14   ` sashiko-bot
  0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-09-28 18:14 UTC (permalink / raw)
  To: Nigel Kirkland; +Cc: linux-scsi

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] scsi: lpfc: memory leak of ndlp nodes on unsolicited PLOGI rejection
--

commit dd6a700ba9afcbae128094473fe756a0c9d5e1df
Author: Nigel Kirkland <nkirkland2304@gmail.com>

lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI

This commit updates the reference counting logic when rejecting an
unsolicited PLOGI to account for dynamically changing domain controller
FC_IDs. It modifies the cleanup path in lpfc_els_unsol_buffer to remove
references from an explicitly saved alloc_ndlp instead of ndlp.

> diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
> index 1793d9c6d8b80..0ca1659d5d102 100644
> --- a/drivers/scsi/lpfc/lpfc_els.c
> +++ b/drivers/scsi/lpfc/lpfc_els.c

[ ... ]

> @@ -10580,7 +10582,9 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
>  			did, vport->port_state, ndlp->nlp_flag);
>  
>  		phba->fc_stat.elsRcvPLOGI++;
> +		alloc_ndlp = ndlp;
>  		ndlp = lpfc_plogi_confirm_nport(phba, payload, ndlp);
> +
>  		if (phba->sli_rev == LPFC_SLI_REV4 &&

[ ... ]

> @@ -10896,13 +10900,28 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
>  		stat.un.b.lsRjtRsnCodeExp = rjt_exp;
>  		lpfc_els_rsp_reject(vport, stat.un.lsRjtError, elsiocb, ndlp,
>  				    NULL);
> -		/* Remove the reference from above for new nodes. */
> -		if (newnode)
> -			lpfc_disc_state_machine(vport, ndlp, NULL,
> -					NLP_EVT_DEVICE_RM);
> +
> +		/* Remove the safety reference from routine start. */
> +		if (newnode) {
> +			/* alloc_ndlp is assigned for an unsolicited PLOGI from
> +			 * the ndlp allocated by this routine. The routine
> +			 * lpfc_plogi_confirm_nport may have swapped the initial
> +			 * ndlp to an existing ndlp. In that case, alloc_ndlp
> +			 * is the pre-swap ndlp that needs a DEVICE_RM because
> +			 * it is the newnode. For all other paths to lsrjt
> +			 * (non-PLOGI ELS) that were rejected by the early
> +			 * port_state check, alloc_ndlp is NULL and ndlp is
> +			 * still the originally allocated node.
> +			 */
> +			lpfc_disc_state_machine(vport,
> +						alloc_ndlp ? alloc_ndlp : ndlp,
> +						NULL, NLP_EVT_DEVICE_RM);

[Severity: High]
Does this code leak both struct lpfc_nodelist objects when rejecting an
unsolicited PLOGI from an unrecognized WWPN?

When the PLOGI is received, lpfc_els_unsol_buffer() allocates an initial
node (alloc_ndlp). Then lpfc_plogi_confirm_nport() allocates a new_ndlp:

drivers/scsi/lpfc/lpfc_els.c:lpfc_plogi_confirm_nport() {
	...
	new_ndlp = lpfc_nlp_init(vport, ndlp->nlp_DID);
	...
}

It swaps the states, places alloc_ndlp into NLP_STE_UNUSED_NODE, and returns
new_ndlp as ndlp.

If the PLOGI is then rejected (e.g., due to port state), this modified
cleanup path calls NLP_EVT_DEVICE_RM on alloc_ndlp instead of ndlp. This
seems to abandon new_ndlp (which is now assigned to ndlp) without dropping
its reference, causing it to leak.

Additionally, because alloc_ndlp is already in the NLP_STE_UNUSED_NODE state
after the swap, its cleanup handler (lpfc_device_rm_unused_node) might simply
return NLP_STE_FREED_NODE without dropping the reference, meaning alloc_ndlp
also leaks.

> +		}
>  	}
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260928181757.21959-1-nkirkland2304@gmail.com?part=9

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v5 05/10] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure
  2026-09-28 18:17 ` [PATCH v5 05/10] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
@ 2026-09-28 18:17   ` sashiko-bot
  0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-09-28 18:17 UTC (permalink / raw)
  To: Nigel Kirkland; +Cc: linux-scsi

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] lpfc: clear_bit(NLP_PLOGI_SND) is a no-op since flag is never set
--

commit 0b98cf6de331fab63e77fe03e2cd19be5c4624b1
Author: Nigel Kirkland <nkirkland2304@gmail.com>

lpfc: Add handling for when PLOGI or PRLI is dropped during link failure

This commit adds handling for dropped PLOGI or PRLI during link failures by
clearing their respective send flags and decrementing counters in
lpfc_cmpl_els_link_down(). This is intended to prevent logistical mixups and
allow the login session to recover correctly.

> diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
> index 45aad4cd2dc80..cd431c7bd9f08 100644
> --- a/drivers/scsi/lpfc/lpfc_els.c
> +++ b/drivers/scsi/lpfc/lpfc_els.c
[ ... ]
> @@ -1237,17 +1239,40 @@ lpfc_cmpl_els_link_down(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
[ ... ]
> +	if (cmd == ELS_CMD_PLOGI) {
> +		/* A PLOGI ELS IO needs to clear the PLOGI_SND flag to
> +		 * acknowledge the ELS completion and allow recovery. Otherwise
> +		 * a subsequent PLOGI gets rejected as a duplicate.
> +		 */
> +		clear_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);

[Severity: Medium]
Is the NLP_PLOGI_SND flag still being actively used by this state machine?

It appears that this flag is never set anywhere in the lpfc driver codebase.
If NLP_PLOGI_SND is never set, this clear_bit() operation is a silent no-op.

Could this mean the intended PLOGI recovery mechanism is non-functional, and
subsequent PLOGIs might still get rejected?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260928181757.21959-1-nkirkland2304@gmail.com?part=5

^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1
@ 2026-09-28 18:17 Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 01/10] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
                   ` (10 more replies)
  0 siblings, 11 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

Update lpfc to revision 15.0.0.1

This patch set contains bug fixes related to cleanup handling in both
normal and error paths, and refactoring of duplicate code.

The patches were cut against Martin's 7.4/scsi-queue tree.

Nigel Kirkland (10):
  lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
  lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not
    set
  lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted
    cmd
  lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
  lpfc: Add handling for when PLOGI or PRLI is dropped during link
    failure
  lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery
    sequence
  lpfc: Rework I/O flush ordering when unloading driver
  lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
  lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI
  lpfc: Update lpfc version to 15.0.0.1

 drivers/scsi/lpfc/lpfc_ct.c        |   2 -
 drivers/scsi/lpfc/lpfc_disc.h      |   5 +-
 drivers/scsi/lpfc/lpfc_els.c       | 109 +++++++++++++-----
 drivers/scsi/lpfc/lpfc_hbadisc.c   | 170 +++++++++++++++++------------
 drivers/scsi/lpfc/lpfc_init.c      |  16 ++-
 drivers/scsi/lpfc/lpfc_nportdisc.c |  30 ++++-
 drivers/scsi/lpfc/lpfc_nvme.c      |   9 ++
 drivers/scsi/lpfc/lpfc_sli.c       |   8 +-
 drivers/scsi/lpfc/lpfc_version.h   |   2 +-
 9 files changed, 242 insertions(+), 109 deletions(-)

-- 
2.38.0


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v5 01/10] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 02/10] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
                   ` (9 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

In lpfc_cmpl_ct_cmd_vmid, there is an early call to lpfc_ct_free_iocb
when cmd is SLI_CTAS_DALLAPP_ID.  Within lpfc_ct_free_iocb the
cmdiocb->rsp_dmabuf will be freed.  This means any ctrsp ptr
dereference for SLI_CT_RESPONSE_FS_RJT or even ctrsp->ReasonCode and
ctrsp->Explanation when handling a CT LS_RJT response is a
use-after-free.

Remove the early lpfc_ct_free_iocb call for SLI_CTAS_DALLAPP_ID.
There already is a free_res label that calls lpfc_ct_free_iocb so
there doesn't need to be an early lpfc_ct_free_iocb at the start of
lpfc_cmpl_ct_cmd_vmid.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_ct.c | 2 --
 1 file changed, 2 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_ct.c b/drivers/scsi/lpfc/lpfc_ct.c
index 0734ab3be3e3..f709e5087577 100644
--- a/drivers/scsi/lpfc/lpfc_ct.c
+++ b/drivers/scsi/lpfc/lpfc_ct.c
@@ -3576,8 +3576,6 @@ lpfc_cmpl_ct_cmd_vmid(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
 	int i;
 
 	cmd = be16_to_cpu(ctcmd->CommandResponse.bits.CmdRsp);
-	if (cmd == SLI_CTAS_DALLAPP_ID)
-		lpfc_ct_free_iocb(phba, cmdiocb);
 
 	if (lpfc_els_chk_latt(vport) || get_job_ulpstatus(phba, rspiocb)) {
 		if (cmd != SLI_CTAS_DALLAPP_ID)
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 02/10] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 01/10] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 03/10] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
                   ` (8 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

It is possible that a dev_loss_tmo callback fires during an hba reset.
The ELS pring structure is cleared by the hba reset path and the
dev_loss_tmo callback executing lpfc_els_abort could be using a stale ELS
pring pointer. To significantly reduce exposure, check if HBA_SETUP flag is set
before proceeding to use the ELS pring pointer in lpfc_els_abort. There is
no point to issue aborts when the sli port is not setup anyways.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_nportdisc.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index 9c449055a55e..2c8d995a45bf 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -227,6 +227,11 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
 	struct lpfc_iocbq *iocb, *next_iocb;
 	int retval = 0;
 
+	/* Exit early to prevent race with queue teardown. */
+	if (unlikely(phba->sli_rev == LPFC_SLI_REV4 &&
+		     !test_bit(HBA_SETUP, &phba->hba_flag)))
+		return;
+
 	pring = lpfc_phba_elsring(phba);
 
 	/* In case of error recovery path, we might have a NULL pring here */
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 03/10] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 01/10] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 02/10] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 04/10] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
                   ` (7 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

A kernel oops at dma_unmap_sg_attrs may occur due to a race between an
aborted scsi I/O completion and a scsi error handler issued TUR using the
same repurposed scsi_cmnd structure.

The LPFC_DRIVER_ABORTED cmd_flag is set when inflight I/Os are aborted
via lpfc_sli_abort_taskmgmt, and this flag is not cleared until after
scsi_done is called.  The inflight I/O test is changed to check scsi I/O
for either LFPC_IO_ON_TXCMPLQ or LPFC_DRIVER_ABORTED.  If either cmd_flag
is set, then the I/O should still be counted.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_sli.c | 8 ++++++--
 1 file changed, 6 insertions(+), 2 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
index cd285e87c278..b370a64e6b67 100644
--- a/drivers/scsi/lpfc/lpfc_sli.c
+++ b/drivers/scsi/lpfc/lpfc_sli.c
@@ -12730,8 +12730,12 @@ lpfc_sli_sum_iocb(struct lpfc_vport *vport, uint16_t tgt_id, uint64_t lun_id,
 
 		if (!iocbq || iocbq->vport != vport)
 			continue;
-		if (!(iocbq->cmd_flag & LPFC_IO_FCP) ||
-		    !(iocbq->cmd_flag & LPFC_IO_ON_TXCMPLQ))
+		/* Only count FCP i/o */
+		if (!(iocbq->cmd_flag & LPFC_IO_FCP))
+			continue;
+		/* Count i/o whilst LLDD retains an interest in the scsi_cmnd */
+		if (!(iocbq->cmd_flag &
+				(LPFC_IO_ON_TXCMPLQ | LPFC_DRIVER_ABORTED)))
 			continue;
 
 		/* Include counting outstanding aborts */
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 04/10] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (2 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 03/10] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:19   ` sashiko-bot
  2026-09-28 18:17 ` [PATCH v5 05/10] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
                   ` (6 subsequent siblings)
  10 siblings, 1 reply; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

The current initial kref count drop logic for an ndlp that fails FDISC
assumes that the ndlp has never registered with transport layer and thus
the lpfc_dev_loss_tmo_callbk never called.  However, a failed FDISC can
occur after a successful transport layer registration too.  So,
lpfc_dev_loss_tmo_callbk can occur and there is a potential use-after-free
on the ndlp.

Check ndlp->fc4_xpt_flags if previously registered with an upper layer
transport and check ndlp->nlp_flags if there is a LPFC_EVT_DEV_LOSS work
pending.  If not previously registered nor LPFC_EVT_DEV_LOSS work pending,
then set the NLP_DROPPED flag as before and decrement the initial kref on
FDISC error.  However, if ndlp has been previously registered, then let the
pre-existing logic for each transport's respective dev_loss_tmo_callbk
perform the initial kref decrement.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_els.c | 20 +++++++++++++++-----
 1 file changed, 15 insertions(+), 5 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 6f6394a0047c..45aad4cd2dc8 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -11416,7 +11416,6 @@ lpfc_cmpl_els_fdisc(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
 		ulp_status, ulp_word4, vport->fc_prevDID);
 
 	if (ulp_status) {
-
 		if (lpfc_fabric_login_reqd(phba, cmdiocb, rspiocb)) {
 			lpfc_retry_pport_discovery(phba);
 			goto out;
@@ -11427,11 +11426,22 @@ lpfc_cmpl_els_fdisc(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
 			goto out;
 		/* Warn FDISC status */
 		lpfc_vlog_msg(vport, KERN_WARNING, LOG_ELS,
-			      "0126 FDISC cmpl status: x%x/x%x)\n",
-			      ulp_status, ulp_word4);
+			      "0126 FDISC cmpl status: (x%x/x%x) ndlp x%px "
+			      "Data: x%lx x%x x%x x%x x%x x%x x%x x%x x%x\n",
+			      ulp_status, ulp_word4, ndlp, ndlp->nlp_flag,
+			      ndlp->nlp_DID, ndlp->nlp_last_elscmd,
+			      ndlp->nlp_type, ndlp->nlp_rpi, ndlp->nlp_state,
+			      ndlp->nlp_prev_state, ndlp->fc4_xpt_flags,
+			      kref_read(&ndlp->kref));
 
-		/* drop initial reference */
-		if (!test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
+		/* If have not previously registered with transport layer and no
+		 * LPFC_EVT_DEV_LOSS work pending, then drop initial reference.
+		 * Otherwise, let the dev_loss_tmo_callbk drop the initial
+		 * reference.
+		 */
+		if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD)) &&
+		    !test_bit(NLP_IN_DEV_LOSS, &ndlp->nlp_flag) &&
+		    !test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
 			lpfc_nlp_put(ndlp);
 
 		goto fdisc_failed;
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 05/10] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (3 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 04/10] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:17   ` sashiko-bot
  2026-09-28 18:17 ` [PATCH v5 06/10] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
                   ` (5 subsequent siblings)
  10 siblings, 1 reply; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

PLOGI and PRLI typically complete via the lpfc_cmpl_els_plogi and
lpfc_cmpl_els_prli handler respectively, but when the link drops they
complete via the lpfc_cmpl_els_link_down handler.  When this occurs, normal
cleanup completion actions are missed and we may fail to recover the login
session due to ndlp logistical mixups.  Fix by clearing the NLP_PLOGI_SND
or NLP_PRLI_SND flag and decrement the outstanding prli_sent counters in
the lpfc_cmpl_els_link_down handler.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_els.c | 35 ++++++++++++++++++++++++++++++-----
 1 file changed, 30 insertions(+), 5 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 45aad4cd2dc8..cd431c7bd9f0 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -1230,6 +1230,8 @@ lpfc_cmpl_els_link_down(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
 	uint32_t *pcmd;
 	uint32_t cmd;
 	u32 ulp_status, ulp_word4;
+	struct lpfc_vport *vport = cmdiocb->vport;
+	struct lpfc_nodelist *ndlp = cmdiocb->ndlp;
 
 	pcmd = (uint32_t *)cmdiocb->cmd_dmabuf->virt;
 	cmd = *pcmd;
@@ -1237,17 +1239,40 @@ lpfc_cmpl_els_link_down(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
 	ulp_status = get_job_ulpstatus(phba, rspiocb);
 	ulp_word4 = get_job_word4(phba, rspiocb);
 
-	lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
-			"6445 ELS completes after LINK_DOWN: "
-			" Status %x/%x cmd x%x flg x%x iotag x%x\n",
-			ulp_status, ulp_word4, cmd,
-			cmdiocb->cmd_flag, cmdiocb->iotag);
+	lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+			 "6445 ELS completes after LINK_DOWN: "
+			 "Status %x/%x cmd x%x data: x%x x%x x%x x%px x%px\n",
+			 ulp_status, ulp_word4, cmd,
+			 cmdiocb->cmd_flag, cmdiocb->iotag,
+			 ndlp->nlp_state, vport, ndlp);
+
+	if (cmd == ELS_CMD_PLOGI) {
+		/* A PLOGI ELS IO needs to clear the PLOGI_SND flag to
+		 * acknowledge the ELS completion and allow recovery. Otherwise
+		 * a subsequent PLOGI gets rejected as a duplicate.
+		 */
+		clear_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);
+	} else if (cmd == ELS_CMD_PRLI || cmd == ELS_CMD_NVMEPRLI) {
+		/* A PRLI ELS IO needs to decrement the fc4_prli_sent count
+		 * added by the lpfc_issue_els_prli function.  A nonzero count
+		 * stops transport registrations.
+		 */
+		clear_bit(NLP_PRLI_SND, &ndlp->nlp_flag);
+		spin_lock_irq(&ndlp->lock);
+		vport->fc_prli_sent--;
+		ndlp->fc4_prli_sent--;
+		spin_unlock_irq(&ndlp->lock);
+	}
 
 	if (cmdiocb->cmd_flag & LPFC_IO_FABRIC) {
 		cmdiocb->cmd_flag &= ~LPFC_IO_FABRIC;
 		atomic_dec(&phba->fabric_iocb_count);
 	}
+
 	lpfc_els_free_iocb(phba, cmdiocb);
+
+	/* lpfc took a reference in the issue.  Release it now. */
+	lpfc_nlp_put(ndlp);
 }
 
 /**
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 06/10] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (4 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 05/10] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:09   ` sashiko-bot
  2026-09-28 18:17 ` [PATCH v5 07/10] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
                   ` (4 subsequent siblings)
  10 siblings, 1 reply; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

In large SAN configurations when a target port fails over, RSCNs may
be spammed triggering a repeat of restarting discovery events for an
ndlp object.

In the case when discovery reaches PRLI state, but the PRLI operation
is interrupted, this leaves the nlp_fc4_type and nlp_type flags
cleared. On the next cycle through lpfc_nlp_reg_node, the NLP_XPT_REGD
flag is set but registraton with the fc transport is bypassed because
lpfc_valid_xpt_node returns false.

This sets up a condition whereby the next call to lpfc_nlp_unreg_node
results in a premature release of the ndlp, and a callback from the
transport results in a use-after-free condition.

To address this issue, refactor lpfc_fc4_xpt_flags such that both
SCSI and NVME have separate flags indicating registration with their
respective transport.  The flags also indicate a request to unregister
had been made. In dev-loss or transport callback processing, the
SCSI_XPT_UNREG_WAIT and NVME_XPT_UNREG_WAIT flags indicate whether
the ndlp reference has already been released.

Introduce explicit, symmetric reference-count tracking for NVMET
target nodes via the new NVMET_XPT_TGT flag and close a race in
lpfc_unregister_remote_port() by marking UNREG_WAIT before
triggering fc_remote_port_delete() rather than after.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_disc.h    |   5 +-
 drivers/scsi/lpfc/lpfc_hbadisc.c | 163 ++++++++++++++++++-------------
 drivers/scsi/lpfc/lpfc_nvme.c    |   9 ++
 3 files changed, 107 insertions(+), 70 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_disc.h b/drivers/scsi/lpfc/lpfc_disc.h
index a377e97cbe65..0ca0a785514c 100644
--- a/drivers/scsi/lpfc/lpfc_disc.h
+++ b/drivers/scsi/lpfc/lpfc_disc.h
@@ -83,11 +83,12 @@ struct lpfc_enc_info {
 };
 
 enum lpfc_fc4_xpt_flags {
-	NLP_XPT_REGD		= 0x1,
+	SCSI_XPT_UNREG_WAIT	= 0x1,
 	SCSI_XPT_REGD		= 0x2,
 	NVME_XPT_REGD		= 0x4,
 	NVME_XPT_UNREG_WAIT	= 0x8,
-	NLP_XPT_HAS_HH		= 0x10
+	NLP_XPT_HAS_HH		= 0x10,
+	NVMET_XPT_TGT		= 0x20
 };
 
 enum lpfc_nlp_save_flags { /* mask bits */
diff --git a/drivers/scsi/lpfc/lpfc_hbadisc.c b/drivers/scsi/lpfc/lpfc_hbadisc.c
index 4c673dffa671..83e29eed14fa 100644
--- a/drivers/scsi/lpfc/lpfc_hbadisc.c
+++ b/drivers/scsi/lpfc/lpfc_hbadisc.c
@@ -200,26 +200,18 @@ lpfc_dev_loss_tmo_callbk(struct fc_rport *rport)
 		/* The scsi_transport is done with the rport so lpfc cannot
 		 * call to unregister.
 		 */
-		if (ndlp->fc4_xpt_flags & SCSI_XPT_REGD) {
+		if ((ndlp->fc4_xpt_flags & SCSI_XPT_REGD) &&
+		    !(ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT)) {
+			/* Reference held since no unreg call made */
 			ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;
+			spin_unlock_irqrestore(&ndlp->lock, iflags);
 
-			/* If NLP_XPT_REGD was cleared in lpfc_nlp_unreg_node,
-			 * unregister calls were made to the scsi and nvme
-			 * transports and refcnt was already decremented. Clear
-			 * the NLP_XPT_REGD flag only if the NVME nrport is
-			 * confirmed unregistered.
-			 */
-			if (ndlp->fc4_xpt_flags & NLP_XPT_REGD) {
-				if (!(ndlp->fc4_xpt_flags & NVME_XPT_REGD))
-					ndlp->fc4_xpt_flags &= ~NLP_XPT_REGD;
-				spin_unlock_irqrestore(&ndlp->lock, iflags);
-
-				/* Release scsi transport reference */
-				lpfc_nlp_put(ndlp);
-			} else {
-				spin_unlock_irqrestore(&ndlp->lock, iflags);
-			}
+			/* Release scsi transport reference */
+			lpfc_nlp_put(ndlp);
 		} else {
+			/* Clear scsi xpt flags */
+			ndlp->fc4_xpt_flags &= ~(SCSI_XPT_REGD |
+						 SCSI_XPT_UNREG_WAIT);
 			spin_unlock_irqrestore(&ndlp->lock, iflags);
 		}
 
@@ -270,7 +262,7 @@ lpfc_dev_loss_tmo_callbk(struct fc_rport *rport)
 	 * The backend does not expect any more calls associated with this
 	 * rport. Remove the association between rport and ndlp.
 	 */
-	ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;
+	ndlp->fc4_xpt_flags &= ~(SCSI_XPT_REGD | SCSI_XPT_UNREG_WAIT);
 	((struct lpfc_rport_data *)rport->dd_data)->pnode = NULL;
 	ndlp->rport = NULL;
 	spin_unlock_irqrestore(&ndlp->lock, iflags);
@@ -606,7 +598,7 @@ lpfc_dev_loss_tmo_handler(struct lpfc_nodelist *ndlp)
 		return fcf_inuse;
 	}
 
-	if (!(ndlp->fc4_xpt_flags & NVME_XPT_REGD))
+	if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD)))
 		lpfc_disc_state_machine(vport, ndlp, NULL, NLP_EVT_DEVICE_RM);
 
 	return fcf_inuse;
@@ -4346,7 +4338,8 @@ lpfc_mbx_cmpl_ns_reg_login(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
 		 */
 		if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD))) {
 			clear_bit(NLP_NPR_2B_DISC, &ndlp->nlp_flag);
-			lpfc_nlp_put(ndlp);
+			if (!test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
+				lpfc_nlp_put(ndlp);
 		}
 
 		if (phba->fc_topology == LPFC_TOPOLOGY_LOOP) {
@@ -4488,26 +4481,36 @@ lpfc_register_remote_port(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 	struct lpfc_hba  *phba = vport->phba;
 	unsigned long flags;
 
-	if (vport->cfg_enable_fc4_type == LPFC_ENABLE_NVME)
+	/* Don't add the remote port if unloading. */
+	if (test_bit(FC_UNLOADING, &vport->load_flag))
 		return;
 
+	spin_lock_irqsave(&ndlp->lock, flags);
+	if ((ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | SCSI_XPT_UNREG_WAIT)) ==
+	    SCSI_XPT_REGD) {
+		/* Prevent racing registrations */
+		spin_unlock_irqrestore(&ndlp->lock, flags);
+		return;
+	}
+	ndlp->fc4_xpt_flags &= ~SCSI_XPT_UNREG_WAIT;
+	ndlp->fc4_xpt_flags |= SCSI_XPT_REGD;
+	spin_unlock_irqrestore(&ndlp->lock, flags);
+
 	/* Remote port has reappeared. Re-register w/ FC transport */
 	rport_ids.node_name = wwn_to_u64(ndlp->nlp_nodename.u.wwn);
 	rport_ids.port_name = wwn_to_u64(ndlp->nlp_portname.u.wwn);
 	rport_ids.port_id = ndlp->nlp_DID;
 	rport_ids.roles = FC_RPORT_ROLE_UNKNOWN;
 
-
 	lpfc_debugfs_disc_trc(vport, LPFC_DISC_TRC_RPORT,
 			      "rport add:       did:x%x flg:x%lx type x%x",
 			      ndlp->nlp_DID, ndlp->nlp_flag, ndlp->nlp_type);
 
-	/* Don't add the remote port if unloading. */
-	if (test_bit(FC_UNLOADING, &vport->load_flag))
-		return;
-
 	ndlp->rport = rport = fc_remote_port_add(shost, 0, &rport_ids);
 	if (!rport) {
+		spin_lock_irqsave(&ndlp->lock, flags);
+		ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;
+		spin_unlock_irqrestore(&ndlp->lock, flags);
 		dev_printk(KERN_WARNING, &phba->pcidev->dev,
 			   "Warning: fc_remote_port_add failed\n");
 		return;
@@ -4519,6 +4522,9 @@ lpfc_register_remote_port(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 	rdata = rport->dd_data;
 	rdata->pnode = lpfc_nlp_get(ndlp);
 	if (!rdata->pnode) {
+		spin_lock_irqsave(&ndlp->lock, flags);
+		ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;
+		spin_unlock_irqrestore(&ndlp->lock, flags);
 		dev_warn(&phba->pcidev->dev,
 			 "Warning - node ref failed. Unreg rport\n");
 		fc_remote_port_delete(rport);
@@ -4526,10 +4532,6 @@ lpfc_register_remote_port(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 		return;
 	}
 
-	spin_lock_irqsave(&ndlp->lock, flags);
-	ndlp->fc4_xpt_flags |= SCSI_XPT_REGD;
-	spin_unlock_irqrestore(&ndlp->lock, flags);
-
 	if (ndlp->nlp_type & NLP_FCP_TARGET)
 		rport_ids.roles |= FC_PORT_ROLE_FCP_TARGET;
 	if (ndlp->nlp_type & NLP_FCP_INITIATOR)
@@ -4562,6 +4564,7 @@ lpfc_unregister_remote_port(struct lpfc_nodelist *ndlp)
 {
 	struct fc_rport *rport = ndlp->rport;
 	struct lpfc_vport *vport = ndlp->vport;
+	unsigned long flags;
 
 	if (vport->cfg_enable_fc4_type == LPFC_ENABLE_NVME)
 		return;
@@ -4576,7 +4579,21 @@ lpfc_unregister_remote_port(struct lpfc_nodelist *ndlp)
 			 ndlp->nlp_DID, rport, ndlp->fc4_xpt_flags,
 			 kref_read(&ndlp->kref));
 
+	/* There are certain cases where the following call could result in an
+	 * almost immediate dev-loss callback. Set unreg pending flag before
+	 * making the call.
+	 */
+	spin_lock_irqsave(&ndlp->lock, flags);
+	if (ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT) {
+		spin_unlock_irqrestore(&ndlp->lock, flags);
+		return;
+	}
+	ndlp->fc4_xpt_flags |= SCSI_XPT_UNREG_WAIT;
+	spin_unlock_irqrestore(&ndlp->lock, flags);
+
 	fc_remote_port_delete(rport);
+
+	/* Release reference */
 	lpfc_nlp_put(ndlp);
 }
 
@@ -4623,7 +4640,10 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 	lpfc_check_nlp_post_devloss(vport, ndlp);
 
 	spin_lock_irqsave(&ndlp->lock, iflags);
-	if (ndlp->fc4_xpt_flags & NLP_XPT_REGD) {
+	if (((ndlp->fc4_xpt_flags & SCSI_XPT_REGD) &&
+	     !(ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT)) ||
+	    ((ndlp->fc4_xpt_flags & NVME_XPT_REGD) &&
+	     !(ndlp->fc4_xpt_flags & NVME_XPT_UNREG_WAIT))) {
 		/* Already registered with backend, trigger rescan */
 		spin_unlock_irqrestore(&ndlp->lock, iflags);
 
@@ -4633,41 +4653,46 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 		}
 		return;
 	}
-
-	ndlp->fc4_xpt_flags |= NLP_XPT_REGD;
 	spin_unlock_irqrestore(&ndlp->lock, iflags);
 
-	if (lpfc_valid_xpt_node(ndlp)) {
-		vport->phba->nport_event_cnt++;
-		/*
-		 * Tell the fc transport about the port, if we haven't
-		 * already. If we have, and it's a scsi entity, be
-		 */
-		lpfc_register_remote_port(vport, ndlp);
+	if (vport->cfg_enable_fc4_type & LPFC_ENABLE_FCP) {
+		if (lpfc_valid_xpt_node(ndlp)) {
+			vport->phba->nport_event_cnt++;
+			/* Tell the fc transport about the port */
+			lpfc_register_remote_port(vport, ndlp);
+		}
 	}
 
 	/* We are done if we do not have any NVME remote node */
 	if (!(ndlp->nlp_fc4_type & NLP_FC4_NVME))
 		return;
 
+	if (vport->phba->sli_rev < LPFC_SLI_REV4)
+		return;
+
 	/* Notify the NVME transport of this new rport. */
-	if (vport->phba->sli_rev >= LPFC_SLI_REV4 &&
-			ndlp->nlp_fc4_type & NLP_FC4_NVME) {
-		if (vport->phba->nvmet_support == 0) {
-			/* Register this rport with the transport.
-			 * Only NVME Target Rports are registered with
-			 * the transport.
-			 */
-			if (ndlp->nlp_type & NLP_NVME_TARGET) {
-				vport->phba->nport_event_cnt++;
-				lpfc_nvme_register_port(vport, ndlp);
-			}
-		} else {
-			/* Just take an NDLP ref count since the
-			 * target does not register rports.
-			 */
-			lpfc_nlp_get(ndlp);
+	if (vport->phba->nvmet_support == 0) {
+		/* Register this rport with the transport.
+		 * Only NVME Target Rports are registered with
+		 * the transport.
+		 */
+		if (ndlp->nlp_type & NLP_NVME_TARGET) {
+			vport->phba->nport_event_cnt++;
+			lpfc_nvme_register_port(vport, ndlp);
+		}
+	} else {
+		/* Just take an NDLP ref count since the
+		 * target does not register rports.
+		 */
+		spin_lock_irqsave(&ndlp->lock, iflags);
+		if (ndlp->fc4_xpt_flags & NVMET_XPT_TGT) {
+			spin_unlock_irqrestore(&ndlp->lock, iflags);
+			return;
 		}
+		ndlp->fc4_xpt_flags |= NVMET_XPT_TGT;
+		spin_unlock_irqrestore(&ndlp->lock, iflags);
+
+		lpfc_nlp_get(ndlp);
 	}
 }
 
@@ -4678,7 +4703,15 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 	unsigned long iflags;
 
 	spin_lock_irqsave(&ndlp->lock, iflags);
-	if (!(ndlp->fc4_xpt_flags & NLP_XPT_REGD)) {
+	if (vport->phba->nvmet_support != 0) {
+		if (ndlp->fc4_xpt_flags & NVMET_XPT_TGT) {
+			ndlp->fc4_xpt_flags &= ~NVMET_XPT_TGT;
+			spin_unlock_irqrestore(&ndlp->lock, iflags);
+			lpfc_nlp_put(ndlp);
+			return;
+		}
+	}
+	if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD))) {
 		spin_unlock_irqrestore(&ndlp->lock, iflags);
 		lpfc_printf_vlog(vport, KERN_INFO,
 				 LOG_ELS | LOG_NODE | LOG_DISCOVERY,
@@ -4688,12 +4721,11 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 				  ndlp->nlp_flag, ndlp->fc4_xpt_flags);
 		return;
 	}
-
-	ndlp->fc4_xpt_flags &= ~NLP_XPT_REGD;
 	spin_unlock_irqrestore(&ndlp->lock, iflags);
 
 	if (ndlp->rport &&
-	    ndlp->fc4_xpt_flags & SCSI_XPT_REGD) {
+	    ((ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | SCSI_XPT_UNREG_WAIT)) ==
+	     SCSI_XPT_REGD)) {
 		vport->phba->nport_event_cnt++;
 		lpfc_unregister_remote_port(ndlp);
 	} else if (!ndlp->rport) {
@@ -4706,16 +4738,11 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 				 kref_read(&ndlp->kref));
 	}
 
-	if (ndlp->fc4_xpt_flags & NVME_XPT_REGD) {
+	if ((ndlp->fc4_xpt_flags & (NVME_XPT_REGD | NVME_XPT_UNREG_WAIT)) ==
+	    NVME_XPT_REGD) {
 		vport->phba->nport_event_cnt++;
-		if (vport->phba->nvmet_support == 0) {
-			lpfc_nvme_unregister_port(vport, ndlp);
-		} else {
-			/* NVMET has no upcall. */
-			lpfc_nlp_put(ndlp);
-		}
+		lpfc_nvme_unregister_port(vport, ndlp);
 	}
-
 }
 
 /*
diff --git a/drivers/scsi/lpfc/lpfc_nvme.c b/drivers/scsi/lpfc/lpfc_nvme.c
index 71714ea390d9..1b854306d577 100644
--- a/drivers/scsi/lpfc/lpfc_nvme.c
+++ b/drivers/scsi/lpfc/lpfc_nvme.c
@@ -2601,6 +2601,15 @@ lpfc_nvme_unregister_port(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 		 * The transport will update it.
 		 */
 		spin_lock_irq(&ndlp->lock);
+		/* Prevent racing unregister requests. */
+		if (ndlp->fc4_xpt_flags & NVME_XPT_UNREG_WAIT) {
+			spin_unlock_irq(&ndlp->lock);
+			lpfc_vlog_msg(vport, KERN_NOTICE, LOG_NVME_DISC,
+				      "6034 NVME unregister in progress, "
+				      "remote port x%px, refcnt %d\n",
+				      remoteport, kref_read(&ndlp->kref));
+			return;
+		}
 		ndlp->fc4_xpt_flags |= NVME_XPT_UNREG_WAIT;
 		spin_unlock_irq(&ndlp->lock);
 
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 07/10] lpfc: Rework I/O flush ordering when unloading driver
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (5 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 06/10] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:18   ` sashiko-bot
  2026-09-28 18:17 ` [PATCH v5 08/10] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
                   ` (3 subsequent siblings)
  10 siblings, 1 reply; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

The lpfc_els_abort routine has a code path that cancels outstanding
I/Os on the ELS ring when attempted aborts fail.  The failed aborts are
queued to a drv_cmpl_list and then cancelled after the ELS pring->txcmplq
is fully traversed.  However if the abort failure returns IOCB_ABORTING,
then the driver should not have cancelled it.  Doing so starts two threads
working on the same iocb and ndlp, leading to unintended race conditions.

Fix by capturing the IOCB_ABORTING return value in lpfc_els_abort and not
adding it to the list of iocbs for cancelling.  We should allow the iocb
scheduled for abort to complete naturally.  This avoids simultaneous
threads acting on the same iocb and ndlp objects.

The lpfc_free_iocb_list is moved to execute after lpfc_sli4_hba_unset
allowing the routine to flush I/O before freeing it.  And, in
lpfc_pci_remove_one_s4 a call to flush the phba->wq is added.  This makes
the unload logic consistent with offline handling logic.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_init.c      | 16 ++++++++++++++--
 drivers/scsi/lpfc/lpfc_nportdisc.c | 11 +++++++++--
 2 files changed, 23 insertions(+), 4 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_init.c b/drivers/scsi/lpfc/lpfc_init.c
index 6460127bcc7b..7352cb6e584b 100644
--- a/drivers/scsi/lpfc/lpfc_init.c
+++ b/drivers/scsi/lpfc/lpfc_init.c
@@ -13514,6 +13514,9 @@ lpfc_sli4_hba_unset(struct lpfc_hba *phba)
 	/* Stop the SLI4 device port */
 	if (phba->pport)
 		phba->pport->work_port_events = 0;
+
+	/* All IO completed and queues released. Free the IOCBs. */
+	lpfc_free_iocb_list(phba);
 }
 
 /*
@@ -14948,11 +14951,20 @@ lpfc_pci_remove_one_s4(struct pci_dev *pdev)
 
 	/* Perform scsi free before driver resource_unset since scsi
 	 * buffers are released to their corresponding pools here.
+	 * lpfc_sli4_hba_unset() issues aborts via lpfc_sli_hba_iocb_abort(),
+	 * which allocates abort IOCBs from phba->lpfc_iocb_list; the pool
+	 * must still exist, so lpfc_free_iocb_list() runs only after unset.
 	 */
 	lpfc_io_free(phba);
-	lpfc_free_iocb_list(phba);
-	lpfc_sli4_hba_unset(phba);
 
+	/* Flush the PHBA WQ - there could be a race with ELS IOs while lpfc
+	 * is unloading.  This stops a race between completions, aborts and
+	 * resource recovery.
+	 */
+	if (phba->wq)
+		flush_workqueue(phba->wq);
+
+	lpfc_sli4_hba_unset(phba);
 	lpfc_unset_driver_resource_phase2(phba);
 	lpfc_sli4_driver_resource_unset(phba);
 
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index 2c8d995a45bf..f917a5bcfd02 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -255,8 +255,9 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
 	spin_lock_irq(&phba->hbalock);
 	if (phba->sli_rev == LPFC_SLI_REV4)
 		spin_lock(&pring->ring_lock);
+
 	list_for_each_entry_safe(iocb, next_iocb, &pring->txcmplq, list) {
-	/* Add to abort_list on on NDLP match. */
+		/* Add to abort_list on NDLP match. */
 		if (lpfc_check_sli_ndlp(phba, pring, iocb, ndlp))
 			list_add_tail(&iocb->dlist, &abort_list);
 	}
@@ -271,7 +272,13 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
 		retval = lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
 		spin_unlock_irq(&phba->hbalock);
 
-		if (retval && test_bit(FC_UNLOADING, &phba->pport->load_flag)) {
+		/* An abort that fails here is just cancelled when the driver is
+		 * going offline.  However, if the abort failure is because the
+		 * IOCB is already getting aborted, don't cancel.  Just let it
+		 * complete.
+		 */
+		if (test_bit(FC_UNLOADING, &phba->pport->load_flag) &&
+		    retval && retval != IOCB_ABORTING) {
 			list_del_init(&iocb->list);
 			list_add_tail(&iocb->list, &drv_cmpl_list);
 		}
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 08/10] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (6 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 07/10] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:17 ` [PATCH v5 09/10] lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI Nigel Kirkland
                   ` (2 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

The lpfc_set_disctmo routine is not used for all cases when the driver
needs to restart discovery on the fc_disctmo timer.  Not doing so, makes
discovery timer actions invisible in some cases as they do not get logged.
This patch substitutes calls on fc_disctmo to use lpfc_set_disctmo in
lpfc_els_rcv_rscn.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_els.c | 17 ++++++-----------
 1 file changed, 6 insertions(+), 11 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index cd431c7bd9f0..1793d9c6d8b8 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -8171,7 +8171,7 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
 	uint32_t payload_len, length, nportid, *cmd;
 	int rscn_cnt;
 	int rscn_id = 0, hba_id = 0;
-	int i, tmo;
+	int i;
 
 	pcmd = cmdiocb->cmd_dmabuf;
 	lp = (uint32_t *) pcmd->virt;
@@ -8250,11 +8250,8 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
 			lpfc_els_rsp_acc(vport, ELS_CMD_ACC, cmdiocb,
 				ndlp, NULL);
 			/* Restart disctmo if its already running */
-			if (test_bit(FC_DISC_TMO, &vport->fc_flag)) {
-				tmo = ((phba->fc_ratov * 3) + 3);
-				mod_timer(&vport->fc_disctmo,
-					  jiffies + secs_to_jiffies(tmo));
-			}
+			if (test_bit(FC_DISC_TMO, &vport->fc_flag))
+				lpfc_set_disctmo(vport);
 			return 0;
 		}
 	}
@@ -8285,11 +8282,9 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
 		set_bit(FC_RSCN_DEFERRED, &vport->fc_flag);
 
 		/* Restart disctmo if its already running */
-		if (test_bit(FC_DISC_TMO, &vport->fc_flag)) {
-			tmo = ((phba->fc_ratov * 3) + 3);
-			mod_timer(&vport->fc_disctmo,
-				  jiffies + secs_to_jiffies(tmo));
-		}
+		if (test_bit(FC_DISC_TMO, &vport->fc_flag))
+			lpfc_set_disctmo(vport);
+
 		if ((rscn_cnt < FC_MAX_HOLD_RSCN) &&
 		    !test_bit(FC_RSCN_DISCOVERY, &vport->fc_flag)) {
 			set_bit(FC_RSCN_MODE, &vport->fc_flag);
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 09/10] lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (7 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 08/10] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-09-28 18:14   ` sashiko-bot
  2026-09-28 18:17 ` [PATCH v5 10/10] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
  2026-10-01 18:12 ` [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
  10 siblings, 1 reply; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

The domain controller FC_ID can change whenever a fabric switch with
an ISL link is replaced or the admin reconfigures the fabric switch
with active links. In such situations, reference count updates should
be applied to the correct nodelist structure.

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_els.c       | 37 ++++++++++++++++++++++--------
 drivers/scsi/lpfc/lpfc_hbadisc.c   |  7 ++++++
 drivers/scsi/lpfc/lpfc_nportdisc.c | 14 +++++++++++
 3 files changed, 49 insertions(+), 9 deletions(-)

diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 1793d9c6d8b8..0ca1659d5d10 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -5977,11 +5977,12 @@ lpfc_els_rsp_reject(struct lpfc_vport *vport, uint32_t rejectError,
 	/* Xmit ELS RJT <err> response tag <ulpIoTag> */
 	lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
 			 "0129 Xmit ELS RJT x%x response tag x%x "
-			 "xri x%x, did x%x, nlp_flag x%lx, nlp_state x%x, "
-			 "rpi x%x\n",
+			 "xri x%x, ndlp x%px, did x%x, nlp_flag x%lx, "
+			 "nlp_state x%x, rpi x%x ref_cnt %d\n",
 			 rejectError, elsiocb->iotag,
-			 get_job_ulpcontext(phba, elsiocb), ndlp->nlp_DID,
-			 ndlp->nlp_flag, ndlp->nlp_state, ndlp->nlp_rpi);
+			 get_job_ulpcontext(phba, elsiocb), ndlp,
+			 ndlp->nlp_DID, ndlp->nlp_flag, ndlp->nlp_state,
+			 ndlp->nlp_rpi, kref_read(&ndlp->kref));
 	lpfc_debugfs_disc_trc(vport, LPFC_DISC_TRC_ELS_RSP,
 		"Issue LS_RJT:    did:x%x flg:x%lx err:x%x",
 		ndlp->nlp_DID, ndlp->nlp_flag, rejectError);
@@ -10482,6 +10483,7 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
 		      struct lpfc_vport *vport, struct lpfc_iocbq *elsiocb)
 {
 	struct lpfc_nodelist *ndlp;
+	struct lpfc_nodelist *alloc_ndlp = NULL;
 	struct ls_rjt stat;
 	u32 *payload, payload_len;
 	u32 cmd = 0, did = 0, newnode, status = 0;
@@ -10580,7 +10582,9 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
 			did, vport->port_state, ndlp->nlp_flag);
 
 		phba->fc_stat.elsRcvPLOGI++;
+		alloc_ndlp = ndlp;
 		ndlp = lpfc_plogi_confirm_nport(phba, payload, ndlp);
+
 		if (phba->sli_rev == LPFC_SLI_REV4 &&
 		    test_bit(FC_PT2PT, &phba->pport->fc_flag)) {
 			vport->fc_prevDID = vport->fc_myDID;
@@ -10896,13 +10900,28 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
 		stat.un.b.lsRjtRsnCodeExp = rjt_exp;
 		lpfc_els_rsp_reject(vport, stat.un.lsRjtError, elsiocb, ndlp,
 				    NULL);
-		/* Remove the reference from above for new nodes. */
-		if (newnode)
-			lpfc_disc_state_machine(vport, ndlp, NULL,
-					NLP_EVT_DEVICE_RM);
+
+		/* Remove the safety reference from routine start. */
+		if (newnode) {
+			/* alloc_ndlp is assigned for an unsolicited PLOGI from
+			 * the ndlp allocated by this routine. The routine
+			 * lpfc_plogi_confirm_nport may have swapped the initial
+			 * ndlp to an existing ndlp. In that case, alloc_ndlp
+			 * is the pre-swap ndlp that needs a DEVICE_RM because
+			 * it is the newnode. For all other paths to lsrjt
+			 * (non-PLOGI ELS) that were rejected by the early
+			 * port_state check, alloc_ndlp is NULL and ndlp is
+			 * still the originally allocated node.
+			 */
+			lpfc_disc_state_machine(vport,
+						alloc_ndlp ? alloc_ndlp : ndlp,
+						NULL, NLP_EVT_DEVICE_RM);
+		}
 	}
 
-	/* Release the reference on this elsiocb, not the ndlp. */
+	/* This elsiocb is not the IOCB issued for a response. Remove
+	 * the safety reference allocated earlier.
+	 */
 	lpfc_nlp_put(elsiocb->ndlp);
 	elsiocb->ndlp = NULL;
 
diff --git a/drivers/scsi/lpfc/lpfc_hbadisc.c b/drivers/scsi/lpfc/lpfc_hbadisc.c
index 83e29eed14fa..6665c3e5b62d 100644
--- a/drivers/scsi/lpfc/lpfc_hbadisc.c
+++ b/drivers/scsi/lpfc/lpfc_hbadisc.c
@@ -4978,6 +4978,13 @@ lpfc_drop_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
 	 */
 	if (ndlp->nlp_state == NLP_STE_UNUSED_NODE)
 		return;
+
+	lpfc_printf_vlog(vport, KERN_INFO, LOG_NODE | LOG_ELS | LOG_DISCOVERY,
+			"3421 Mark ndlp x%px Dropped. DID x%06x, nflags x%lx "
+			"xflags x%x ref_cnt %d\n",
+			 ndlp, ndlp->nlp_DID, ndlp->nlp_flag,
+			 ndlp->fc4_xpt_flags, kref_read(&ndlp->kref));
+
 	lpfc_nlp_set_state(vport, ndlp, NLP_STE_UNUSED_NODE);
 	if (vport->phba->sli_rev == LPFC_SLI_REV4) {
 		lpfc_cleanup_vports_rrqs(vport, ndlp);
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index f917a5bcfd02..ffec352c8ab2 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -2878,6 +2878,20 @@ lpfc_device_rm_npr_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
 		set_bit(NLP_NODEV_REMOVE, &ndlp->nlp_flag);
 		return ndlp->nlp_state;
 	}
+
+	/* It is an error to drop a node while either of its transport
+	 * registrations have not completed.
+	 */
+	if (ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD)) {
+		lpfc_printf_vlog(vport, KERN_WARNING,
+				 LOG_ELS | LOG_NODE | LOG_DISCOVERY,
+				 "3423 Not dropping ndlp x%px, DID x%x xflags x%x "
+				 "with dev_loss pending. ref_cnt %d\n",
+				 ndlp, ndlp->nlp_DID, ndlp->fc4_xpt_flags,
+				 kref_read(&ndlp->kref));
+		return ndlp->nlp_state;
+	}
+
 	lpfc_drop_node(vport, ndlp);
 	return NLP_STE_FREED_NODE;
 }
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH v5 10/10] lpfc: Update lpfc version to 15.0.0.1
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (8 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 09/10] lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI Nigel Kirkland
@ 2026-09-28 18:17 ` Nigel Kirkland
  2026-10-01 18:12 ` [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
  10 siblings, 0 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-09-28 18:17 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304

Update lpfc version to 15.0.0.1

Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
 drivers/scsi/lpfc/lpfc_version.h | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/scsi/lpfc/lpfc_version.h b/drivers/scsi/lpfc/lpfc_version.h
index d6e6e436fbfc..7df63d118234 100644
--- a/drivers/scsi/lpfc/lpfc_version.h
+++ b/drivers/scsi/lpfc/lpfc_version.h
@@ -20,7 +20,7 @@
  * included with this package.                                     *
  *******************************************************************/
 
-#define LPFC_DRIVER_VERSION "15.0.0.0"
+#define LPFC_DRIVER_VERSION "15.0.0.1"
 #define LPFC_DRIVER_NAME		"lpfc"
 
 /* Used for SLI 2/3 */
-- 
2.38.0


^ permalink raw reply related	[flat|nested] 17+ messages in thread

* Re: [PATCH v5 07/10] lpfc: Rework I/O flush ordering when unloading driver
  2026-09-28 18:17 ` [PATCH v5 07/10] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
@ 2026-09-28 18:18   ` sashiko-bot
  0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-09-28 18:18 UTC (permalink / raw)
  To: Nigel Kirkland; +Cc: linux-scsi

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] scsi: lpfc: memory leak of iocbq left on txcmplq during unload
--

commit 0147a0023ef276c5ea7aa246240c55db9bc5ed52
Author: Nigel Kirkland <nkirkland2304@gmail.com>

lpfc: Rework I/O flush ordering when unloading driver

This commit reworks I/O flush ordering during driver unload by avoiding
cancelling IOCB_ABORTING returns in lpfc_els_abort() and flushing the
PHBA WQ prior to resource unset.

> diff --git a/drivers/scsi/lpfc/lpfc_init.c b/drivers/scsi/lpfc/lpfc_init.c
> index 6460127bcc7b0..7352cb6e584bf 100644
> --- a/drivers/scsi/lpfc/lpfc_init.c
> +++ b/drivers/scsi/lpfc/lpfc_init.c
> @@ -13514,6 +13514,9 @@ lpfc_sli4_hba_unset(struct lpfc_hba *phba)
>  	/* Stop the SLI4 device port */
>  	if (phba->pport)
>  		phba->pport->work_port_events = 0;
> +
> +	/* All IO completed and queues released. Free the IOCBs. */
> +	lpfc_free_iocb_list(phba);
>  }

[ ... ]

> diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
> index 2c8d995a45bf8..f917a5bcfd025 100644
> --- a/drivers/scsi/lpfc/lpfc_nportdisc.c
> +++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
[ ... ]
> @@ -271,7 +272,13 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
>  		retval = lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
>  		spin_unlock_irq(&phba->hbalock);
>  
> -		if (retval && test_bit(FC_UNLOADING, &phba->pport->load_flag)) {
> +		/* An abort that fails here is just cancelled when the driver is
> +		 * going offline.  However, if the abort failure is because the
> +		 * IOCB is already getting aborted, don't cancel.  Just let it
> +		 * complete.
> +		 */
> +		if (test_bit(FC_UNLOADING, &phba->pport->load_flag) &&
> +		    retval && retval != IOCB_ABORTING) {

[Severity: Medium]
If an ELS abort returns IOCB_ABORTING during driver unload, this explicitly
leaves the IOCB on the ring's txcmplq to complete naturally.

However, does the unload path guarantee this will complete? In lpfc_init.c,
lpfc_sli4_hba_unset() calls lpfc_sli4_xri_exchange_busy_wait() which checks
abts_* lists, but does not wait for txcmplq to empty before calling
lpfc_sli4_disable_intr() to disable interrupts.

When lpfc_sli4_hba_unset() finally calls lpfc_free_iocb_list(), it only frees
the available IOCBs in phba->lpfc_iocb_list:

lpfc_init.c:lpfc_free_iocb_list() {
    ...
    list_for_each_entry_safe(iocbq_entry, iocbq_next,
                             &phba->lpfc_iocb_list, list) {
        list_del(&iocbq_entry->list);
        kfree(iocbq_entry);
    ...
}

Because the pending IOCB is still linked in txcmplq and hasn't been returned
to the free list, does this result in a memory leak of the lpfc_iocbq
structure?

>  			list_del_init(&iocb->list);
>  			list_add_tail(&iocb->list, &drv_cmpl_list);
>  		}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260928181757.21959-1-nkirkland2304@gmail.com?part=7

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v5 04/10] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
  2026-09-28 18:17 ` [PATCH v5 04/10] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
@ 2026-09-28 18:19   ` sashiko-bot
  0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-09-28 18:19 UTC (permalink / raw)
  To: Nigel Kirkland; +Cc: linux-scsi

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] lpfc: TOCTOU race in lpfc_cmpl_els_fdisc() leads to Use-After-Free
--

commit 46ec54ed5cbeb514e9e9362594a41ca2f9b3aa3d
Author: Nigel Kirkland <nkirkland2304@gmail.com>

lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error

The current initial kref count drop logic for an ndlp that fails FDISC
assumes that the ndlp has never registered with transport layer and thus
the lpfc_dev_loss_tmo_callbk never called.  However, a failed FDISC can
occur after a successful transport layer registration too.  So,
lpfc_dev_loss_tmo_callbk can occur and there is a potential use-after-free
on the ndlp.

> diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
> index 6f6394a0047c2..45aad4cd2dc80 100644
> --- a/drivers/scsi/lpfc/lpfc_els.c
> +++ b/drivers/scsi/lpfc/lpfc_els.c
[ ... ]
> @@ -11427,11 +11426,22 @@ lpfc_cmpl_els_fdisc(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
[ ... ]
> -		/* drop initial reference */
> -		if (!test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
> +		/* If have not previously registered with transport layer and no
> +		 * LPFC_EVT_DEV_LOSS work pending, then drop initial reference.
> +		 * Otherwise, let the dev_loss_tmo_callbk drop the initial
> +		 * reference.
> +		 */
> +		if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD)) &&
> +		    !test_bit(NLP_IN_DEV_LOSS, &ndlp->nlp_flag) &&
> +		    !test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
>  			lpfc_nlp_put(ndlp);

[Severity: High]
Does this lockless check introduce a time-of-check to time-of-use race
condition that could lead to a use-after-free?

If an FDISC error occurs concurrently with a devloss timeout for the same
ndlp, lpfc_dev_loss_tmo_callbk() could execute on another CPU:

CPU1 in lpfc_dev_loss_tmo_callbk():
    set_bit(NLP_IN_DEV_LOSS, &ndlp->nlp_flag);

    spin_lock_irqsave(&ndlp->lock, iflags);
    /* ... */
    ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;

Because set_bit() lacks a full memory barrier and spin_lock() provides only
an ACQUIRE barrier, could the clearing of SCSI_XPT_REGD become visible to
CPU2 before NLP_IN_DEV_LOSS is seen?

If CPU2 is executing this new check in lpfc_cmpl_els_fdisc() locklessly and
without read barriers, it might see both SCSI_XPT_REGD == 0 and
NLP_IN_DEV_LOSS == 0.

Could this cause CPU2 to erroneously evaluate the condition as true, execute
lpfc_nlp_put(ndlp), and prematurely drop the initial reference, leading to a
use-after-free when CPU1 later tries to use the ndlp pointer?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260928181757.21959-1-nkirkland2304@gmail.com?part=4

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1
  2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
                   ` (9 preceding siblings ...)
  2026-09-28 18:17 ` [PATCH v5 10/10] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
@ 2026-10-01 18:12 ` Nigel Kirkland
  10 siblings, 0 replies; 17+ messages in thread
From: Nigel Kirkland @ 2026-10-01 18:12 UTC (permalink / raw)
  To: linux-scsi, nigel.kirkland; +Cc: paul.ely

On Mon, Sep 28, 2026 at 10:54 AM Nigel Kirkland <nkirkland2304@gmail.com> wrote:
>
> Update lpfc to revision 15.0.0.1
>
> This patch set contains bug fixes related to cleanup handling in both
> normal and error paths, and refactoring of duplicate code.
>
> The patches were cut against Martin's 7.4/scsi-queue tree.
>
> Nigel Kirkland (10):
>   lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
>   lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not
>     set
>   lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted
>     cmd
>   lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error

Responses to sashiko-bot findings inline

patch 4 - new finding (was not reported against v4) - will fix

>   lpfc: Add handling for when PLOGI or PRLI is dropped during link
>     failure

patch 5 - companion patch was erroneously omitted, will fix

>   lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery
>     sequence

patch 6 - some findings are not a regression - however, will be improved
in next submission

>   lpfc: Rework I/O flush ordering when unloading driver

patch 7 - we believe the fix is valid. There is longstanding scope for
future hardening. We plan to address separately in a future patch
set.

>   lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
>   lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI

patch 9 - we believe the fix is valid and it is well tested. There is scope
for additional hardening. We plan to address separately in a future
patch set.

>   lpfc: Update lpfc version to 15.0.0.1
>
>  drivers/scsi/lpfc/lpfc_ct.c        |   2 -
>  drivers/scsi/lpfc/lpfc_disc.h      |   5 +-
>  drivers/scsi/lpfc/lpfc_els.c       | 109 +++++++++++++-----
>  drivers/scsi/lpfc/lpfc_hbadisc.c   | 170 +++++++++++++++++------------
>  drivers/scsi/lpfc/lpfc_init.c      |  16 ++-
>  drivers/scsi/lpfc/lpfc_nportdisc.c |  30 ++++-
>  drivers/scsi/lpfc/lpfc_nvme.c      |   9 ++
>  drivers/scsi/lpfc/lpfc_sli.c       |   8 +-
>  drivers/scsi/lpfc/lpfc_version.h   |   2 +-
>  9 files changed, 242 insertions(+), 109 deletions(-)
>
> --
> 2.38.0
>

^ permalink raw reply	[flat|nested] 17+ messages in thread

end of thread, other threads:[~2026-10-01 18:12 UTC | newest]

Thread overview: 17+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 18:17 [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-28 18:17 ` [PATCH v5 01/10] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
2026-09-28 18:17 ` [PATCH v5 02/10] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
2026-09-28 18:17 ` [PATCH v5 03/10] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
2026-09-28 18:17 ` [PATCH v5 04/10] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
2026-09-28 18:19   ` sashiko-bot
2026-09-28 18:17 ` [PATCH v5 05/10] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
2026-09-28 18:17   ` sashiko-bot
2026-09-28 18:17 ` [PATCH v5 06/10] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
2026-09-28 18:09   ` sashiko-bot
2026-09-28 18:17 ` [PATCH v5 07/10] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
2026-09-28 18:18   ` sashiko-bot
2026-09-28 18:17 ` [PATCH v5 08/10] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
2026-09-28 18:17 ` [PATCH v5 09/10] lpfc: Update correct ndlp refcnt when rejecting an unsolicited PLOGI Nigel Kirkland
2026-09-28 18:14   ` sashiko-bot
2026-09-28 18:17 ` [PATCH v5 10/10] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
2026-10-01 18:12 ` [PATCH v5 00/10] lpfc: Update lpfc to revision 15.0.0.1 Nigel Kirkland

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox