* Re: [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
2026-09-17 22:20 ` [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
@ 2026-09-17 22:10 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-09-17 22:10 UTC (permalink / raw)
To: Nigel Kirkland; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] scsi: lpfc: Use-after-free in NVMe unregistration due to missing serialization
- [High] scsi: lpfc: ndlp memory leak due to concurrent registration race
--
commit 6da5a3b652f8d5c08807e25e9db6e426dc7022ff
Author: Nigel Kirkland <nkirkland2304@gmail.com>
lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
This commit refactors how registration state is tracked by separating the SCSI
and NVME transport registration flags. It introduces new wait flags to track
unregistration and a new flag for NVMET target nodes to address use-after-free
issues during target port failovers and RSCN storms.
> diff --git a/drivers/scsi/lpfc/lpfc_hbadisc.c b/drivers/scsi/lpfc/lpfc_hbadisc.c
> index 4c673dffa671b..b4a5c7d5c2a0b 100644
> --- a/drivers/scsi/lpfc/lpfc_hbadisc.c
> +++ b/drivers/scsi/lpfc/lpfc_hbadisc.c
[ ... ]
> @@ -4623,7 +4632,10 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
> lpfc_check_nlp_post_devloss(vport, ndlp);
>
> spin_lock_irqsave(&ndlp->lock, iflags);
> - if (ndlp->fc4_xpt_flags & NLP_XPT_REGD) {
> + if (((ndlp->fc4_xpt_flags & SCSI_XPT_REGD) &&
> + !(ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT)) ||
> + ((ndlp->fc4_xpt_flags & NVME_XPT_REGD) &&
> + !(ndlp->fc4_xpt_flags & NVME_XPT_UNREG_WAIT))) {
> /* Already registered with backend, trigger rescan */
> spin_unlock_irqrestore(&ndlp->lock, iflags);
[ ... ]
> - ndlp->fc4_xpt_flags |= NLP_XPT_REGD;
> spin_unlock_irqrestore(&ndlp->lock, iflags);
[Severity: High]
Does removing the atomic test-and-set of the registration flag here introduce
an ndlp memory leak?
If multiple threads execute lpfc_nlp_reg_node() concurrently (e.g., from
repeated RSCNs), they can all pass the registration check above while the flags
are still cleared. All threads will drop the lock and proceed to call
lpfc_register_remote_port(), which will unconditionally increment the ndlp
reference count via lpfc_nlp_get().
During unregistration, the SCSI path correctly prevents double unregistration
by checking SCSI_XPT_UNREG_WAIT, so it will only drop one reference. The extra
references taken during the concurrent registrations will permanently leak.
>
> if (lpfc_valid_xpt_node(ndlp)) {
[ ... ]
> @@ -4706,16 +4728,11 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
> kref_read(&ndlp->kref));
> }
>
> - if (ndlp->fc4_xpt_flags & NVME_XPT_REGD) {
> + if ((ndlp->fc4_xpt_flags & (NVME_XPT_REGD | NVME_XPT_UNREG_WAIT)) ==
> + NVME_XPT_REGD) {
[Severity: High]
Can this lockless flag check cause a use-after-free on the ndlp object?
Because the flag check happens outside of ndlp->lock in lpfc_nlp_unreg_node(),
multiple threads can concurrently pass this check and call
lpfc_nvme_unregister_port().
Unlike the SCSI unregistration path, lpfc_nvme_unregister_port()
unconditionally sets the NVME_XPT_UNREG_WAIT flag under the lock without
verifying if a concurrent thread has already set it. Both threads will then
call nvme_fc_unregister_remoteport() and lpfc_nlp_put(). Since the ndlp
reference was only incremented once during NVMe registration, the double
lpfc_nlp_put() could result in a use-after-free.
> vport->phba->nport_event_cnt++;
> - if (vport->phba->nvmet_support == 0) {
> - lpfc_nvme_unregister_port(vport, ndlp);
> - } else {
> - /* NVMET has no upcall. */
> - lpfc_nlp_put(ndlp);
> - }
> + lpfc_nvme_unregister_port(vport, ndlp);
> }
> -
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917222015.61053-1-nkirkland2304@gmail.com?part=6
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations
2026-09-17 22:20 ` [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations Nigel Kirkland
@ 2026-09-17 22:12 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-09-17 22:12 UTC (permalink / raw)
To: Nigel Kirkland; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] lpfc: Deferred PLOGIs are erroneously rejected in lpfc_nlp_logo_unreg due to state check
- [Medium] lpfc: TOCTOU race in NLP_PLOGI_SND flag tracking allows multiple PLOGIs in flight
--
commit 3d16993b649c134340e20dec0bd637799a114aff
Author: Nigel Kirkland <nkirkland2304@gmail.com>
lpfc: Improve PLOGI retry handling for large SAN configurations
In large SAN configurations with link perturbations, rediscovery of target
ports is problematic due to PLOGI retry race conditions. This patch improves
target rediscovery by ensuring PLOGI retries are serialized in unregistration
and retry handler paths.
> diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
> index cd431c7bd9f08..3e26654a8286a 100644
> --- a/drivers/scsi/lpfc/lpfc_els.c
> +++ b/drivers/scsi/lpfc/lpfc_els.c
[ ... ]
> @@ -2329,17 +2331,43 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
[ ... ]
> + if (test_bit(NLP_PLOGI_SND, &ndlp->nlp_flag)) {
> + lpfc_printf_vlog(vport, KERN_INFO,
> + LOG_ELS | LOG_NODE | LOG_DISCOVERY,
> + "4113 Reject PLOGI issue, PLOGI in-flight "
> + "x%px, DID x%x nflag x%lx\n",
> + ndlp, ndlp->nlp_DID, ndlp->nlp_flag);
> + return 1;
> + }
[Severity: Medium]
Can concurrent worker threads race on this lockless flag check?
If two threads execute lpfc_issue_els_plogi() simultaneously (for example,
from an unregistration completion and a timer), both could evaluate
test_bit(NLP_PLOGI_SND, &ndlp->nlp_flag) as false.
Both threads would then continue executing, allocate their IOCBs, and reach
the set_bit(NLP_PLOGI_SND, &ndlp->nlp_flag) call at the end of the function.
This sequence would issue duplicate PLOGI commands to the same NPort.
> +
> + if (ndlp->nlp_state > NLP_STE_PLOGI_ISSUE &&
> + ndlp->nlp_state <= NLP_STE_MAPPED_NODE) {
> + lpfc_printf_vlog(vport, KERN_INFO,
> + LOG_ELS | LOG_NODE | LOG_DISCOVERY,
> + "4114 Reject PLOGI issue, Node in "
> + "unexpected state x%px, DID x%x nflag x%lx "
> + "in State x%x\n",
> + ndlp, ndlp->nlp_DID,
> + ndlp->nlp_flag, ndlp->nlp_state);
> + return 1;
> }
[Severity: High]
Will this strict state check inadvertently drop deferred PLOGIs during LOGO
unregistration?
During target recovery after a LOGO storm or link toggle, the unregistration
handler lpfc_nlp_logo_unreg() uses lpfc_issue_els_logo() to transition the
node state to NLP_STE_LOGO_ISSUE.
When the deferred PLOGI is subsequently issued via lpfc_issue_els_plogi(),
this new check (ndlp->nlp_state > NLP_STE_PLOGI_ISSUE &&
ndlp->nlp_state <= NLP_STE_MAPPED_NODE) evaluates to true because
NLP_STE_LOGO_ISSUE (0x5) falls between 0x1 and 0x7. This causes the deferred
PLOGI to be erroneously rejected, potentially stalling the target discovery
process.
[ ... ]
> @@ -2410,11 +2438,26 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
[ ... ]
> lpfc_els_free_iocb(phba, elsiocb);
> lpfc_nlp_put(ndlp);
> return 1;
> }
>
> + set_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);
> return 0;
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917222015.61053-1-nkirkland2304@gmail.com?part=8
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received
2026-09-17 22:20 ` [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received Nigel Kirkland
@ 2026-09-17 22:15 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-09-17 22:15 UTC (permalink / raw)
To: Nigel Kirkland; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] scsi: lpfc: txcmplq list corruption in lpfc_sli_abort_iocb_ring()
- [High] scsi: lpfc: unintended abort of reused XRI/IoTag in lpfc_els_retry()
--
commit d8f6ae5839dc18257873d48c2e2262ffa79b7bec
Author: Nigel Kirkland <nkirkland2304@gmail.com>
Subject: lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received
This commit modifies the LPFC driver to explicitly set the Inhibit Abort (IA)
bit when sending an ABORT_WQE for a timed-out PLOGI CQE. It introduces a
boolean argument to lpfc_sli_issue_abort_iotag to control this behavior.
> diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
> --- a/drivers/scsi/lpfc/lpfc_els.c
> +++ b/drivers/scsi/lpfc/lpfc_els.c
[ ... ]
> @@ -4921,7 +4925,38 @@ lpfc_els_retry(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
> /* Reset the Link */
> link_reset = 1;
> break;
> + } else if (cmd == ELS_CMD_PLOGI) {
[ ... ]
> + pring = lpfc_phba_elsring(phba);
> + spin_lock_irqsave(&phba->hbalock, iflags);
> + lpfc_sli_issue_abort_iotag(phba, pring, cmdiocb,
> + true, NULL);
> + spin_unlock_irqrestore(&phba->hbalock, iflags);
[Severity: High]
Could this inadvertently abort an unrelated command?
Because lpfc_els_retry() schedules a retry, the completion handler
lpfc_cmpl_els_plogi() will immediately free cmdiocb locally, returning the
XRI/IoTag to the free pool.
Since the abort is issued here without validating if the hardware actually
kept the exchange busy (LPFC_EXCHANGE_BUSY) or checking if this is SLI-4
hardware, another CPU could concurrently allocate the freed tag, issue a new
command, and have it erroneously aborted.
> }
> +
> retry = 1;
> delay = 100;
> break;
> diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
> --- a/drivers/scsi/lpfc/lpfc_sli.c
> +++ b/drivers/scsi/lpfc/lpfc_sli.c
[ ... ]
> @@ -4620,12 +4621,19 @@ lpfc_sli_abort_iocb_ring(struct lpfc_hba *phba, struct lpfc_sli_ring *pring)
> iocb->cmd_flag &= ~LPFC_IO_ON_TXCMPLQ;
> list_splice_init(&pring->txcmplq, &tx_completions);
> pring->txcmplq_cnt = 0;
> + spin_unlock_irqrestore(plock, iflag);
> } else {
> + /* lpfc_sli_issue_abort_iotag expects the hba_lock held, but not
> + * the ring_lock.
> + */
> + spin_unlock_irqrestore(plock, iflag);
> + spin_lock_irqsave(&phba->hbalock, iflag);
> /* Issue ABTS for everything on the txcmplq */
> list_for_each_entry_safe(iocb, next_iocb, &pring->txcmplq, list)
> - lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
> + lpfc_sli_issue_abort_iotag(phba, pring, iocb, false,
> + NULL);
> + spin_unlock_irqrestore(&phba->hbalock, iflag);
[Severity: High]
Does this code risk list corruption on the txcmplq list?
In lpfc_sli_abort_iocb_ring(), plock (which aliases ring_lock for SLI-4
hardware) is dropped, leaving txcmplq unprotected against concurrent removals
by completion handlers.
While list_for_each_entry_safe() protects against the current element being
removed, it is not thread-safe if a concurrent interrupt handler removes
next_iocb.
Could this lead to poisoned or corrupted list pointers during host reset or
link down events?
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917222015.61053-1-nkirkland2304@gmail.com?part=9
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v4 00/14] Update lpfc to revision 15.0.0.1
@ 2026-09-17 22:20 Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 01/14] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
` (14 more replies)
0 siblings, 15 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
Update lpfc to revision 15.0.0.1
This patch set contains bug fixes related to cleanup handling in both
normal and error paths, discovery rework for large SAN configurations, and
refactoring of duplicate code.
The patches were cut against Martin's 7.4/scsi-queue tree.
Nigel Kirkland (14):
lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not
set
lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted
cmd
lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
lpfc: Add handling for when PLOGI or PRLI is dropped during link
failure
lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery
sequence
lpfc: Rework I/O flush ordering when unloading driver
lpfc: Improve PLOGI retry handling for large SAN configurations
lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received
lpfc: Remove slowpath cqe process limiter in slow ring event handler
lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL
unavailable
lpfc: Update ELS ACC logging for diagnostic troubleshooting
lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
lpfc: Update lpfc version to 15.0.0.1
drivers/scsi/lpfc/lpfc_bsg.c | 7 +-
drivers/scsi/lpfc/lpfc_crtn.h | 12 +-
drivers/scsi/lpfc/lpfc_ct.c | 23 +-
drivers/scsi/lpfc/lpfc_disc.h | 5 +-
drivers/scsi/lpfc/lpfc_els.c | 503 ++++++++++++++++++++++-------
drivers/scsi/lpfc/lpfc_hbadisc.c | 130 ++++----
drivers/scsi/lpfc/lpfc_init.c | 21 +-
drivers/scsi/lpfc/lpfc_nportdisc.c | 100 +++++-
drivers/scsi/lpfc/lpfc_nvme.c | 2 +-
drivers/scsi/lpfc/lpfc_scsi.c | 2 +-
drivers/scsi/lpfc/lpfc_sli.c | 227 ++++++++++---
drivers/scsi/lpfc/lpfc_sli.h | 18 ++
drivers/scsi/lpfc/lpfc_version.h | 2 +-
13 files changed, 812 insertions(+), 240 deletions(-)
--
2.38.0
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v4 01/14] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 02/14] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
` (13 subsequent siblings)
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
In lpfc_cmpl_ct_cmd_vmid, there is an early call to lpfc_ct_free_iocb when
cmd is SLI_CTAS_DALLAPP_ID. Within lpfc_ct_free_iocb the
cmdiocb->rsp_dmabuf will be freed. This means any ctrsp ptr dereference
for SLI_CT_RESPONSE_FS_RJT or even ctrsp->ReasonCode and ctrsp->Explanation
when handling a CT LS_RJT response is a use-after-free.
Remove the early lpfc_ct_free_iocb call for SLI_CTAS_DALLAPP_ID. There
already is a free_res label that calls lpfc_ct_free_iocb so there doesn't
need to be an early lpfc_ct_free_iocb at the start of
lpfc_cmpl_ct_cmd_vmid.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_ct.c | 2 --
1 file changed, 2 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_ct.c b/drivers/scsi/lpfc/lpfc_ct.c
index 0734ab3be3e3..f709e5087577 100644
--- a/drivers/scsi/lpfc/lpfc_ct.c
+++ b/drivers/scsi/lpfc/lpfc_ct.c
@@ -3576,8 +3576,6 @@ lpfc_cmpl_ct_cmd_vmid(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
int i;
cmd = be16_to_cpu(ctcmd->CommandResponse.bits.CmdRsp);
- if (cmd == SLI_CTAS_DALLAPP_ID)
- lpfc_ct_free_iocb(phba, cmdiocb);
if (lpfc_els_chk_latt(vport) || get_job_ulpstatus(phba, rspiocb)) {
if (cmd != SLI_CTAS_DALLAPP_ID)
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 02/14] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 01/14] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 03/14] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
` (12 subsequent siblings)
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
It is possible that a dev_loss_tmo callback fires during an hba reset.
The ELS pring structure is cleared by the hba reset path and the
dev_loss_tmo callback executing lpfc_els_abort could be using a stale ELS
pring pointer. To significantly reduce exposure, check if HBA_SETUP flag is set
before proceeding to use the ELS pring pointer in lpfc_els_abort. There is
no point to issue aborts when the sli port is not setup anyways.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_nportdisc.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index 9c449055a55e..2c8d995a45bf 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -227,6 +227,11 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
struct lpfc_iocbq *iocb, *next_iocb;
int retval = 0;
+ /* Exit early to prevent race with queue teardown. */
+ if (unlikely(phba->sli_rev == LPFC_SLI_REV4 &&
+ !test_bit(HBA_SETUP, &phba->hba_flag)))
+ return;
+
pring = lpfc_phba_elsring(phba);
/* In case of error recovery path, we might have a NULL pring here */
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 03/14] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 01/14] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 02/14] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 04/14] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
` (11 subsequent siblings)
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
A kernel oops at dma_unmap_sg_attrs may occur due to a race between an
aborted scsi I/O completion and a scsi error handler issued TUR using the
same repurposed scsi_cmnd structure.
The LPFC_DRIVER_ABORTED cmd_flag is set when inflight I/Os are aborted
via lpfc_sli_abort_taskmgmt, and this flag is not cleared until after
scsi_done is called. The inflight I/O test is changed to check scsi I/O
for either LFPC_IO_ON_TXCMPLQ or LPFC_DRIVER_ABORTED. If either cmd_flag
is set, then the I/O should still be counted.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_sli.c | 8 ++++++--
1 file changed, 6 insertions(+), 2 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
index cd285e87c278..b370a64e6b67 100644
--- a/drivers/scsi/lpfc/lpfc_sli.c
+++ b/drivers/scsi/lpfc/lpfc_sli.c
@@ -12730,8 +12730,12 @@ lpfc_sli_sum_iocb(struct lpfc_vport *vport, uint16_t tgt_id, uint64_t lun_id,
if (!iocbq || iocbq->vport != vport)
continue;
- if (!(iocbq->cmd_flag & LPFC_IO_FCP) ||
- !(iocbq->cmd_flag & LPFC_IO_ON_TXCMPLQ))
+ /* Only count FCP i/o */
+ if (!(iocbq->cmd_flag & LPFC_IO_FCP))
+ continue;
+ /* Count i/o whilst LLDD retains an interest in the scsi_cmnd */
+ if (!(iocbq->cmd_flag &
+ (LPFC_IO_ON_TXCMPLQ | LPFC_DRIVER_ABORTED)))
continue;
/* Include counting outstanding aborts */
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 04/14] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (2 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 03/14] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 05/14] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
` (10 subsequent siblings)
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
The current initial kref count drop logic for an ndlp that fails FDISC
assumes that the ndlp has never registered with transport layer and thus
the lpfc_dev_loss_tmo_callbk never called. However, a failed FDISC can
occur after a successful transport layer registration too. So,
lpfc_dev_loss_tmo_callbk can occur and there is a potential use-after-free
on the ndlp.
Check ndlp->fc4_xpt_flags if previously registered with an upper layer
transport and check ndlp->nlp_flags if there is a LPFC_EVT_DEV_LOSS work
pending. If not previously registered nor LPFC_EVT_DEV_LOSS work pending,
then set the NLP_DROPPED flag as before and decrement the initial kref on
FDISC error. However, if ndlp has been previously registered, then let the
pre-existing logic for each transport's respective dev_loss_tmo_callbk
perform the initial kref decrement.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_els.c | 20 +++++++++++++++-----
1 file changed, 15 insertions(+), 5 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 6f6394a0047c..45aad4cd2dc8 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -11416,7 +11416,6 @@ lpfc_cmpl_els_fdisc(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
ulp_status, ulp_word4, vport->fc_prevDID);
if (ulp_status) {
-
if (lpfc_fabric_login_reqd(phba, cmdiocb, rspiocb)) {
lpfc_retry_pport_discovery(phba);
goto out;
@@ -11427,11 +11426,22 @@ lpfc_cmpl_els_fdisc(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
goto out;
/* Warn FDISC status */
lpfc_vlog_msg(vport, KERN_WARNING, LOG_ELS,
- "0126 FDISC cmpl status: x%x/x%x)\n",
- ulp_status, ulp_word4);
+ "0126 FDISC cmpl status: (x%x/x%x) ndlp x%px "
+ "Data: x%lx x%x x%x x%x x%x x%x x%x x%x x%x\n",
+ ulp_status, ulp_word4, ndlp, ndlp->nlp_flag,
+ ndlp->nlp_DID, ndlp->nlp_last_elscmd,
+ ndlp->nlp_type, ndlp->nlp_rpi, ndlp->nlp_state,
+ ndlp->nlp_prev_state, ndlp->fc4_xpt_flags,
+ kref_read(&ndlp->kref));
- /* drop initial reference */
- if (!test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
+ /* If have not previously registered with transport layer and no
+ * LPFC_EVT_DEV_LOSS work pending, then drop initial reference.
+ * Otherwise, let the dev_loss_tmo_callbk drop the initial
+ * reference.
+ */
+ if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD)) &&
+ !test_bit(NLP_IN_DEV_LOSS, &ndlp->nlp_flag) &&
+ !test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
lpfc_nlp_put(ndlp);
goto fdisc_failed;
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 05/14] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (3 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 04/14] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
` (9 subsequent siblings)
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
PLOGI and PRLI typically complete via the lpfc_cmpl_els_plogi and
lpfc_cmpl_els_prli handler respectively, but when the link drops they
complete via the lpfc_cmpl_els_link_down handler. When this occurs, normal
cleanup completion actions are missed and we may fail to recover the login
session due to ndlp logistical mixups. Fix by clearing the NLP_PLOGI_SND
or NLP_PRLI_SND flag and decrement the outstanding prli_sent counters in
the lpfc_cmpl_els_link_down handler.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_els.c | 35 ++++++++++++++++++++++++++++++-----
1 file changed, 30 insertions(+), 5 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 45aad4cd2dc8..cd431c7bd9f0 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -1230,6 +1230,8 @@ lpfc_cmpl_els_link_down(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
uint32_t *pcmd;
uint32_t cmd;
u32 ulp_status, ulp_word4;
+ struct lpfc_vport *vport = cmdiocb->vport;
+ struct lpfc_nodelist *ndlp = cmdiocb->ndlp;
pcmd = (uint32_t *)cmdiocb->cmd_dmabuf->virt;
cmd = *pcmd;
@@ -1237,17 +1239,40 @@ lpfc_cmpl_els_link_down(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
ulp_status = get_job_ulpstatus(phba, rspiocb);
ulp_word4 = get_job_word4(phba, rspiocb);
- lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
- "6445 ELS completes after LINK_DOWN: "
- " Status %x/%x cmd x%x flg x%x iotag x%x\n",
- ulp_status, ulp_word4, cmd,
- cmdiocb->cmd_flag, cmdiocb->iotag);
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "6445 ELS completes after LINK_DOWN: "
+ "Status %x/%x cmd x%x data: x%x x%x x%x x%px x%px\n",
+ ulp_status, ulp_word4, cmd,
+ cmdiocb->cmd_flag, cmdiocb->iotag,
+ ndlp->nlp_state, vport, ndlp);
+
+ if (cmd == ELS_CMD_PLOGI) {
+ /* A PLOGI ELS IO needs to clear the PLOGI_SND flag to
+ * acknowledge the ELS completion and allow recovery. Otherwise
+ * a subsequent PLOGI gets rejected as a duplicate.
+ */
+ clear_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);
+ } else if (cmd == ELS_CMD_PRLI || cmd == ELS_CMD_NVMEPRLI) {
+ /* A PRLI ELS IO needs to decrement the fc4_prli_sent count
+ * added by the lpfc_issue_els_prli function. A nonzero count
+ * stops transport registrations.
+ */
+ clear_bit(NLP_PRLI_SND, &ndlp->nlp_flag);
+ spin_lock_irq(&ndlp->lock);
+ vport->fc_prli_sent--;
+ ndlp->fc4_prli_sent--;
+ spin_unlock_irq(&ndlp->lock);
+ }
if (cmdiocb->cmd_flag & LPFC_IO_FABRIC) {
cmdiocb->cmd_flag &= ~LPFC_IO_FABRIC;
atomic_dec(&phba->fabric_iocb_count);
}
+
lpfc_els_free_iocb(phba, cmdiocb);
+
+ /* lpfc took a reference in the issue. Release it now. */
+ lpfc_nlp_put(ndlp);
}
/**
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (4 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 05/14] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:10 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 07/14] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
` (8 subsequent siblings)
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
In large SAN configurations when a target port fails over, RSCNs may be
spammed triggering a repeat of restarting discovery events for an ndlp
object.
In the case when discovery reaches PRLI state, but the PRLI operation is
interrupted, this leaves the nlp_fc4_type and nlp_type flags cleared. And,
on the next cycle through lpfc_nlp_reg_node, the NLP_XPT_REGD flag is set
but registraton with the fc transport is bypassed because
lpfc_valid_xpt_node returns false.
This sets up a condition whereby the next call to lpfc_nlp_unreg_node
results in a premature release of the ndlp, and a callback from the
transport results in a use-after-free condition.
To address this issue, refactor lpfc_fc4_xpt_flags such that both SCSI and
NVME have separate flags indicating registration with their respective
transport. The flags also indicate a request to unregister had been made.
In dev-loss or transport callback processing, the SCSI_XPT_UNREG_WAIT and
NVME_XPT_UNREG_WAIT flags indicate whether the ndlp reference has already
been released.
Introduce explicit, symmetric reference-count tracking for NVMET target
nodes via the new NVMET_XPT_TGT flag and close a race in
lpfc_unregister_remote_port() by marking UNREG_WAIT before triggering
fc_remote_port_delete() rather than after.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_disc.h | 5 +-
drivers/scsi/lpfc/lpfc_hbadisc.c | 127 ++++++++++++++++++-------------
2 files changed, 75 insertions(+), 57 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_disc.h b/drivers/scsi/lpfc/lpfc_disc.h
index a377e97cbe65..0ca0a785514c 100644
--- a/drivers/scsi/lpfc/lpfc_disc.h
+++ b/drivers/scsi/lpfc/lpfc_disc.h
@@ -83,11 +83,12 @@ struct lpfc_enc_info {
};
enum lpfc_fc4_xpt_flags {
- NLP_XPT_REGD = 0x1,
+ SCSI_XPT_UNREG_WAIT = 0x1,
SCSI_XPT_REGD = 0x2,
NVME_XPT_REGD = 0x4,
NVME_XPT_UNREG_WAIT = 0x8,
- NLP_XPT_HAS_HH = 0x10
+ NLP_XPT_HAS_HH = 0x10,
+ NVMET_XPT_TGT = 0x20
};
enum lpfc_nlp_save_flags { /* mask bits */
diff --git a/drivers/scsi/lpfc/lpfc_hbadisc.c b/drivers/scsi/lpfc/lpfc_hbadisc.c
index 4c673dffa671..b4a5c7d5c2a0 100644
--- a/drivers/scsi/lpfc/lpfc_hbadisc.c
+++ b/drivers/scsi/lpfc/lpfc_hbadisc.c
@@ -200,26 +200,18 @@ lpfc_dev_loss_tmo_callbk(struct fc_rport *rport)
/* The scsi_transport is done with the rport so lpfc cannot
* call to unregister.
*/
- if (ndlp->fc4_xpt_flags & SCSI_XPT_REGD) {
+ if ((ndlp->fc4_xpt_flags & SCSI_XPT_REGD) &&
+ !(ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT)) {
+ /* Reference held since no unreg call made */
ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;
+ spin_unlock_irqrestore(&ndlp->lock, iflags);
- /* If NLP_XPT_REGD was cleared in lpfc_nlp_unreg_node,
- * unregister calls were made to the scsi and nvme
- * transports and refcnt was already decremented. Clear
- * the NLP_XPT_REGD flag only if the NVME nrport is
- * confirmed unregistered.
- */
- if (ndlp->fc4_xpt_flags & NLP_XPT_REGD) {
- if (!(ndlp->fc4_xpt_flags & NVME_XPT_REGD))
- ndlp->fc4_xpt_flags &= ~NLP_XPT_REGD;
- spin_unlock_irqrestore(&ndlp->lock, iflags);
-
- /* Release scsi transport reference */
- lpfc_nlp_put(ndlp);
- } else {
- spin_unlock_irqrestore(&ndlp->lock, iflags);
- }
+ /* Release scsi transport reference */
+ lpfc_nlp_put(ndlp);
} else {
+ /* Clear scsi xpt flags */
+ ndlp->fc4_xpt_flags &= ~(SCSI_XPT_REGD |
+ SCSI_XPT_UNREG_WAIT);
spin_unlock_irqrestore(&ndlp->lock, iflags);
}
@@ -270,7 +262,7 @@ lpfc_dev_loss_tmo_callbk(struct fc_rport *rport)
* The backend does not expect any more calls associated with this
* rport. Remove the association between rport and ndlp.
*/
- ndlp->fc4_xpt_flags &= ~SCSI_XPT_REGD;
+ ndlp->fc4_xpt_flags &= ~(SCSI_XPT_REGD | SCSI_XPT_UNREG_WAIT);
((struct lpfc_rport_data *)rport->dd_data)->pnode = NULL;
ndlp->rport = NULL;
spin_unlock_irqrestore(&ndlp->lock, iflags);
@@ -606,7 +598,7 @@ lpfc_dev_loss_tmo_handler(struct lpfc_nodelist *ndlp)
return fcf_inuse;
}
- if (!(ndlp->fc4_xpt_flags & NVME_XPT_REGD))
+ if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD)))
lpfc_disc_state_machine(vport, ndlp, NULL, NLP_EVT_DEVICE_RM);
return fcf_inuse;
@@ -4346,7 +4338,8 @@ lpfc_mbx_cmpl_ns_reg_login(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
*/
if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD))) {
clear_bit(NLP_NPR_2B_DISC, &ndlp->nlp_flag);
- lpfc_nlp_put(ndlp);
+ if (!test_and_set_bit(NLP_DROPPED, &ndlp->nlp_flag))
+ lpfc_nlp_put(ndlp);
}
if (phba->fc_topology == LPFC_TOPOLOGY_LOOP) {
@@ -4527,6 +4520,7 @@ lpfc_register_remote_port(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
}
spin_lock_irqsave(&ndlp->lock, flags);
+ ndlp->fc4_xpt_flags &= ~SCSI_XPT_UNREG_WAIT;
ndlp->fc4_xpt_flags |= SCSI_XPT_REGD;
spin_unlock_irqrestore(&ndlp->lock, flags);
@@ -4562,6 +4556,7 @@ lpfc_unregister_remote_port(struct lpfc_nodelist *ndlp)
{
struct fc_rport *rport = ndlp->rport;
struct lpfc_vport *vport = ndlp->vport;
+ unsigned long flags;
if (vport->cfg_enable_fc4_type == LPFC_ENABLE_NVME)
return;
@@ -4576,7 +4571,21 @@ lpfc_unregister_remote_port(struct lpfc_nodelist *ndlp)
ndlp->nlp_DID, rport, ndlp->fc4_xpt_flags,
kref_read(&ndlp->kref));
+ /* There are certain cases where the following call could result in an
+ * almost immediate dev-loss callback. Set unreg pending flag before
+ * making the call.
+ */
+ spin_lock_irqsave(&ndlp->lock, flags);
+ if (ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT) {
+ spin_unlock_irqrestore(&ndlp->lock, flags);
+ return;
+ }
+ ndlp->fc4_xpt_flags |= SCSI_XPT_UNREG_WAIT;
+ spin_unlock_irqrestore(&ndlp->lock, flags);
+
fc_remote_port_delete(rport);
+
+ /* Release reference */
lpfc_nlp_put(ndlp);
}
@@ -4623,7 +4632,10 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
lpfc_check_nlp_post_devloss(vport, ndlp);
spin_lock_irqsave(&ndlp->lock, iflags);
- if (ndlp->fc4_xpt_flags & NLP_XPT_REGD) {
+ if (((ndlp->fc4_xpt_flags & SCSI_XPT_REGD) &&
+ !(ndlp->fc4_xpt_flags & SCSI_XPT_UNREG_WAIT)) ||
+ ((ndlp->fc4_xpt_flags & NVME_XPT_REGD) &&
+ !(ndlp->fc4_xpt_flags & NVME_XPT_UNREG_WAIT))) {
/* Already registered with backend, trigger rescan */
spin_unlock_irqrestore(&ndlp->lock, iflags);
@@ -4633,16 +4645,11 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
}
return;
}
-
- ndlp->fc4_xpt_flags |= NLP_XPT_REGD;
spin_unlock_irqrestore(&ndlp->lock, iflags);
if (lpfc_valid_xpt_node(ndlp)) {
vport->phba->nport_event_cnt++;
- /*
- * Tell the fc transport about the port, if we haven't
- * already. If we have, and it's a scsi entity, be
- */
+ /* Tell the fc transport about the port */
lpfc_register_remote_port(vport, ndlp);
}
@@ -4650,24 +4657,32 @@ lpfc_nlp_reg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
if (!(ndlp->nlp_fc4_type & NLP_FC4_NVME))
return;
+ if (vport->phba->sli_rev < LPFC_SLI_REV4)
+ return;
+
/* Notify the NVME transport of this new rport. */
- if (vport->phba->sli_rev >= LPFC_SLI_REV4 &&
- ndlp->nlp_fc4_type & NLP_FC4_NVME) {
- if (vport->phba->nvmet_support == 0) {
- /* Register this rport with the transport.
- * Only NVME Target Rports are registered with
- * the transport.
- */
- if (ndlp->nlp_type & NLP_NVME_TARGET) {
- vport->phba->nport_event_cnt++;
- lpfc_nvme_register_port(vport, ndlp);
- }
- } else {
- /* Just take an NDLP ref count since the
- * target does not register rports.
- */
- lpfc_nlp_get(ndlp);
+ if (vport->phba->nvmet_support == 0) {
+ /* Register this rport with the transport.
+ * Only NVME Target Rports are registered with
+ * the transport.
+ */
+ if (ndlp->nlp_type & NLP_NVME_TARGET) {
+ vport->phba->nport_event_cnt++;
+ lpfc_nvme_register_port(vport, ndlp);
+ }
+ } else {
+ /* Just take an NDLP ref count since the
+ * target does not register rports.
+ */
+ spin_lock_irqsave(&ndlp->lock, iflags);
+ if (ndlp->fc4_xpt_flags & NVMET_XPT_TGT) {
+ spin_unlock_irqrestore(&ndlp->lock, iflags);
+ return;
}
+ ndlp->fc4_xpt_flags |= NVMET_XPT_TGT;
+ spin_unlock_irqrestore(&ndlp->lock, iflags);
+
+ lpfc_nlp_get(ndlp);
}
}
@@ -4678,7 +4693,15 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
unsigned long iflags;
spin_lock_irqsave(&ndlp->lock, iflags);
- if (!(ndlp->fc4_xpt_flags & NLP_XPT_REGD)) {
+ if (vport->phba->nvmet_support != 0) {
+ if (ndlp->fc4_xpt_flags & NVMET_XPT_TGT) {
+ ndlp->fc4_xpt_flags &= ~NVMET_XPT_TGT;
+ spin_unlock_irqrestore(&ndlp->lock, iflags);
+ lpfc_nlp_put(ndlp);
+ return;
+ }
+ }
+ if (!(ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | NVME_XPT_REGD))) {
spin_unlock_irqrestore(&ndlp->lock, iflags);
lpfc_printf_vlog(vport, KERN_INFO,
LOG_ELS | LOG_NODE | LOG_DISCOVERY,
@@ -4688,12 +4711,11 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
ndlp->nlp_flag, ndlp->fc4_xpt_flags);
return;
}
-
- ndlp->fc4_xpt_flags &= ~NLP_XPT_REGD;
spin_unlock_irqrestore(&ndlp->lock, iflags);
if (ndlp->rport &&
- ndlp->fc4_xpt_flags & SCSI_XPT_REGD) {
+ ((ndlp->fc4_xpt_flags & (SCSI_XPT_REGD | SCSI_XPT_UNREG_WAIT)) ==
+ SCSI_XPT_REGD)) {
vport->phba->nport_event_cnt++;
lpfc_unregister_remote_port(ndlp);
} else if (!ndlp->rport) {
@@ -4706,16 +4728,11 @@ lpfc_nlp_unreg_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
kref_read(&ndlp->kref));
}
- if (ndlp->fc4_xpt_flags & NVME_XPT_REGD) {
+ if ((ndlp->fc4_xpt_flags & (NVME_XPT_REGD | NVME_XPT_UNREG_WAIT)) ==
+ NVME_XPT_REGD) {
vport->phba->nport_event_cnt++;
- if (vport->phba->nvmet_support == 0) {
- lpfc_nvme_unregister_port(vport, ndlp);
- } else {
- /* NVMET has no upcall. */
- lpfc_nlp_put(ndlp);
- }
+ lpfc_nvme_unregister_port(vport, ndlp);
}
-
}
/*
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 07/14] lpfc: Rework I/O flush ordering when unloading driver
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (5 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations Nigel Kirkland
` (7 subsequent siblings)
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
The lpfc_els_abort routine has a code path that cancels outstanding
I/Os on the ELS ring when attempted aborts fail. The failed aborts are
queued to a drv_cmpl_list and then cancelled after the ELS pring->txcmplq
is fully traversed. However if the abort failure returns IOCB_ABORTING,
then the driver should not have cancelled it. Doing so starts two threads
working on the same iocb and ndlp, leading to unintended race conditions.
Fix by capturing the IOCB_ABORTING return value in lpfc_els_abort and not
adding it to the list of iocbs for cancelling. We should allow the iocb
scheduled for abort to complete naturally. This avoids simultaneous
threads acting on the same iocb and ndlp objects.
The lpfc_free_iocb_list is moved to execute after lpfc_sli4_hba_unset
allowing the routine to flush I/O before freeing it. And, in
lpfc_pci_remove_one_s4 a call to flush the phba->wq is added. This makes
the unload logic consistent with offline handling logic.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_init.c | 16 ++++++++++++++--
drivers/scsi/lpfc/lpfc_nportdisc.c | 11 +++++++++--
2 files changed, 23 insertions(+), 4 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_init.c b/drivers/scsi/lpfc/lpfc_init.c
index 6460127bcc7b..7352cb6e584b 100644
--- a/drivers/scsi/lpfc/lpfc_init.c
+++ b/drivers/scsi/lpfc/lpfc_init.c
@@ -13514,6 +13514,9 @@ lpfc_sli4_hba_unset(struct lpfc_hba *phba)
/* Stop the SLI4 device port */
if (phba->pport)
phba->pport->work_port_events = 0;
+
+ /* All IO completed and queues released. Free the IOCBs. */
+ lpfc_free_iocb_list(phba);
}
/*
@@ -14948,11 +14951,20 @@ lpfc_pci_remove_one_s4(struct pci_dev *pdev)
/* Perform scsi free before driver resource_unset since scsi
* buffers are released to their corresponding pools here.
+ * lpfc_sli4_hba_unset() issues aborts via lpfc_sli_hba_iocb_abort(),
+ * which allocates abort IOCBs from phba->lpfc_iocb_list; the pool
+ * must still exist, so lpfc_free_iocb_list() runs only after unset.
*/
lpfc_io_free(phba);
- lpfc_free_iocb_list(phba);
- lpfc_sli4_hba_unset(phba);
+ /* Flush the PHBA WQ - there could be a race with ELS IOs while lpfc
+ * is unloading. This stops a race between completions, aborts and
+ * resource recovery.
+ */
+ if (phba->wq)
+ flush_workqueue(phba->wq);
+
+ lpfc_sli4_hba_unset(phba);
lpfc_unset_driver_resource_phase2(phba);
lpfc_sli4_driver_resource_unset(phba);
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index 2c8d995a45bf..f917a5bcfd02 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -255,8 +255,9 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
spin_lock_irq(&phba->hbalock);
if (phba->sli_rev == LPFC_SLI_REV4)
spin_lock(&pring->ring_lock);
+
list_for_each_entry_safe(iocb, next_iocb, &pring->txcmplq, list) {
- /* Add to abort_list on on NDLP match. */
+ /* Add to abort_list on NDLP match. */
if (lpfc_check_sli_ndlp(phba, pring, iocb, ndlp))
list_add_tail(&iocb->dlist, &abort_list);
}
@@ -271,7 +272,13 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
retval = lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
spin_unlock_irq(&phba->hbalock);
- if (retval && test_bit(FC_UNLOADING, &phba->pport->load_flag)) {
+ /* An abort that fails here is just cancelled when the driver is
+ * going offline. However, if the abort failure is because the
+ * IOCB is already getting aborted, don't cancel. Just let it
+ * complete.
+ */
+ if (test_bit(FC_UNLOADING, &phba->pport->load_flag) &&
+ retval && retval != IOCB_ABORTING) {
list_del_init(&iocb->list);
list_add_tail(&iocb->list, &drv_cmpl_list);
}
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (6 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 07/14] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:12 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received Nigel Kirkland
` (6 subsequent siblings)
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
In large SAN configurations with link perturbations, rediscovery of target
ports is problematic due to PLOGI retry race conditions.
This patch improves target rediscovery by ensuring PLOGI retries are
serialized in unregistration and retry handler paths.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_els.c | 92 ++++++++++++++++++++++++++++--
drivers/scsi/lpfc/lpfc_nportdisc.c | 72 ++++++++++++++++++++++-
drivers/scsi/lpfc/lpfc_sli.c | 41 +++++++++++--
3 files changed, 191 insertions(+), 14 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index cd431c7bd9f0..3e26654a8286 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -2159,6 +2159,8 @@ lpfc_cmpl_els_plogi(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
goto out_freeiocb;
}
+ clear_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);
+
/* Since ndlp can be freed in the disc state machine, note if this node
* is being used during discovery.
*/
@@ -2329,17 +2331,43 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
test_bit(NLP_UNREG_INP, &ndlp->nlp_flag)) &&
((ndlp->nlp_DID & Fabric_DID_MASK) != Fabric_DID_MASK) &&
!test_bit(FC_OFFLINE_MODE, &vport->fc_flag)) {
- lpfc_printf_vlog(vport, KERN_INFO, LOG_DISCOVERY,
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_ELS | LOG_NODE | LOG_DISCOVERY,
"4110 Issue PLOGI x%x deferred "
"on NPort x%x rpi x%x flg x%lx Data:"
" x%px\n",
ndlp->nlp_defer_did, ndlp->nlp_DID,
ndlp->nlp_rpi, ndlp->nlp_flag, ndlp);
- /* We can only defer 1st PLOGI */
- if (ndlp->nlp_defer_did == NLP_EVT_NOTHING_PENDING)
+ /* Don't defer a PLOGI that is already in that condition.
+ * Also set the nlp_last_elscmd to PLOGI to get the retry.
+ */
+ if (ndlp->nlp_defer_did == NLP_EVT_NOTHING_PENDING) {
ndlp->nlp_defer_did = did;
- return 0;
+ ndlp->nlp_last_elscmd = ELS_CMD_PLOGI;
+ }
+ return 1;
+ }
+
+ if (test_bit(NLP_PLOGI_SND, &ndlp->nlp_flag)) {
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_ELS | LOG_NODE | LOG_DISCOVERY,
+ "4113 Reject PLOGI issue, PLOGI in-flight "
+ "x%px, DID x%x nflag x%lx\n",
+ ndlp, ndlp->nlp_DID, ndlp->nlp_flag);
+ return 1;
+ }
+
+ if (ndlp->nlp_state > NLP_STE_PLOGI_ISSUE &&
+ ndlp->nlp_state <= NLP_STE_MAPPED_NODE) {
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_ELS | LOG_NODE | LOG_DISCOVERY,
+ "4114 Reject PLOGI issue, Node in "
+ "unexpected state x%px, DID x%x nflag x%lx "
+ "in State x%x\n",
+ ndlp, ndlp->nlp_DID,
+ ndlp->nlp_flag, ndlp->nlp_state);
+ return 1;
}
cmdsize = (sizeof(uint32_t) + sizeof(struct serv_parm));
@@ -2410,11 +2438,26 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
ret = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
if (ret) {
+ lpfc_vlog_msg(vport, KERN_NOTICE,
+ LOG_ELS | LOG_DISCOVERY | LOG_NODE,
+ "0157 PLOGI WQE Put returned %d\n",
+ ret);
+
+ /* Under heavy vpi counts, the driver's host_index can catch up
+ * to the hba_index causing a put error. Catch this case and
+ * put the IO on phba->txq.
+ */
+ if (ret == IOCB_FAILED_PUT && phba->sli_rev == LPFC_SLI_REV4) {
+ lpfc_sli4_queue_io_for_retry(phba, elsiocb, false);
+ return 0;
+ }
+
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
}
+ set_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);
return 0;
}
@@ -2735,7 +2778,21 @@ lpfc_issue_els_prli(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (rc) {
+ lpfc_vlog_msg(vport, KERN_NOTICE,
+ LOG_ELS | LOG_DISCOVERY | LOG_NODE,
+ "0155 PRLI WQE Put returned %d\n",
+ rc);
+
+ /* Under heavy vpi counts, the driver's host_index can catch up
+ * to the hba_index causing a put error. Catch this case and
+ * put the IO on phba->txq.
+ */
+ if (rc == IOCB_FAILED_PUT && phba->sli_rev == LPFC_SLI_REV4) {
+ lpfc_sli4_queue_io_for_retry(phba, elsiocb, false);
+ return 0;
+ }
+
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -4614,6 +4671,31 @@ lpfc_els_retry_delay_handler(struct lpfc_nodelist *ndlp)
lpfc_issue_els_flogi(vport, ndlp, retry);
break;
case ELS_CMD_PLOGI:
+ /* The driver delayed a PLOGI via the nlp_delayfunc, but
+ * it's possible the PLOGI is already on a deferred retry.
+ * Catch this case and skip this delayed PLOGI. This prevents
+ * multiple PLOGIs in flight. The defer code flow cleans
+ * up.
+ */
+ if ((test_bit(NLP_IGNR_REG_CMPL, &ndlp->nlp_flag) ||
+ test_bit(NLP_UNREG_INP, &ndlp->nlp_flag)) &&
+ ndlp->nlp_defer_did != NLP_EVT_NOTHING_PENDING &&
+ ((ndlp->nlp_DID & Fabric_DID_MASK) != Fabric_DID_MASK) &&
+ !test_bit(FC_OFFLINE_MODE, &vport->fc_flag)) {
+ /* When UNREG_RPI completes we need to have the
+ * nlp_last_elscmd set.
+ */
+ ndlp->nlp_last_elscmd = ELS_CMD_PLOGI;
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_ELS | LOG_NODE | LOG_DISCOVERY,
+ "4112 Skip delayed PLOGI x%x deferred "
+ "on NPort x%x rpi x%x flg x%lx Data:"
+ " x%px\n",
+ ndlp->nlp_defer_did, ndlp->nlp_DID,
+ ndlp->nlp_rpi, ndlp->nlp_flag, ndlp);
+ break;
+ }
+
if (!lpfc_issue_els_plogi(vport, ndlp->nlp_DID, retry)) {
ndlp->nlp_prev_state = ndlp->nlp_state;
lpfc_nlp_set_state(vport, ndlp, NLP_STE_PLOGI_ISSUE);
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index f917a5bcfd02..b0af279480fb 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -920,6 +920,64 @@ lpfc_rcv_logo(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
ndlp->nlp_DID, ndlp->nlp_state,
ndlp->nlp_type, vport->fc_flag);
+ /* The driver wants to schedule a delayed PLOGI to recover
+ * the remote Nport. However, there are two cases that
+ * stop this so that multiple PLOGI are not inflight to
+ * the same NPortID
+ *
+ * Do not schedule a delayed PLOGI if the deferred PLOGI
+ * code already set up a PLOGI retry after an UNREG_RPI
+ * mailbox completes.
+ */
+ if (test_bit(NLP_UNREG_INP, &ndlp->nlp_flag) &&
+ ndlp->nlp_defer_did == ndlp->nlp_DID &&
+ ndlp->nlp_last_elscmd == ELS_CMD_PLOGI) {
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_NODE | LOG_ELS | LOG_DISCOVERY,
+ "3206 No PLOGI delay, defer PLOGI "
+ "waiting on DID x%06x UNREG_RPI "
+ "nflag x%lx state x%x lastels x%x "
+ "defer_did x%x\n",
+ ndlp->nlp_DID, ndlp->nlp_flag,
+ ndlp->nlp_state, ndlp->nlp_last_elscmd,
+ ndlp->nlp_defer_did);
+ goto out;
+ }
+
+ /* A delayed PLOGI retry is not required if the ndlp's delay
+ * timer is running and the last command was PLOGI.
+ */
+ if (test_bit(NLP_DELAY_TMO, &ndlp->nlp_flag) &&
+ ndlp->nlp_last_elscmd == ELS_CMD_PLOGI) {
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_NODE | LOG_ELS | LOG_DISCOVERY,
+ "3207 No PLOGI delay, PLOGI_DELAY_TMO "
+ "active on DID x%06x "
+ "nflag x%lx state x%x lastels x%x "
+ "defer_did x%x\n",
+ ndlp->nlp_DID, ndlp->nlp_flag,
+ ndlp->nlp_state, ndlp->nlp_last_elscmd,
+ ndlp->nlp_defer_did);
+ goto out;
+ }
+
+ /* Do not schedule a PLOGI retry if the ndlp state is NPR
+ * and vport has received an RSCN
+ */
+ if (ndlp->nlp_state == NLP_STE_NPR_NODE &&
+ test_bit(FC_RSCN_MODE, &vport->fc_flag)) {
+ lpfc_printf_vlog(vport, KERN_INFO,
+ LOG_NODE | LOG_ELS | LOG_DISCOVERY,
+ "3939 No PLOGI delay, RSCN in "
+ "progress for NPR DID x%06x "
+ "nflag x%lx state x%x last_els x%x "
+ "defer_did x%06x\n",
+ ndlp->nlp_DID, ndlp->nlp_flag,
+ ndlp->nlp_state, ndlp->nlp_last_elscmd,
+ ndlp->nlp_defer_did);
+ goto out;
+ }
+
/* Special cases for rports that recover post LOGO. */
if ((!(ndlp->nlp_type == NLP_FABRIC) &&
(ndlp->nlp_type & (NLP_FCP_TARGET | NLP_NVME_TARGET) ||
@@ -1651,9 +1709,14 @@ lpfc_rcv_plogi_adisc_issue(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
}
return ndlp->nlp_state;
}
- ndlp->nlp_prev_state = NLP_STE_ADISC_ISSUE;
- lpfc_issue_els_plogi(vport, ndlp->nlp_DID, 0);
+
+ /* lpfc_issue_els_plogi rejects a request if the ndlp state
+ * has advanced beyond PLOGI_ISSUE. Since the ADISC was aborted
+ * and the lpfc_rcv_plogi failed, ensure the PLOGI is sent.
+ */
+ ndlp->nlp_prev_state = ndlp->nlp_state;
lpfc_nlp_set_state(vport, ndlp, NLP_STE_PLOGI_ISSUE);
+ lpfc_issue_els_plogi(vport, ndlp->nlp_DID, 0);
return ndlp->nlp_state;
}
@@ -2675,7 +2738,10 @@ lpfc_rcv_plogi_npr_node(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
clear_bit(NLP_NPR_ADISC, &ndlp->nlp_flag);
clear_bit(NLP_NPR_2B_DISC, &ndlp->nlp_flag);
} else if (!test_bit(NLP_NPR_2B_DISC, &ndlp->nlp_flag)) {
- /* send PLOGI immediately, move to PLOGI issue state */
+ /* Provided a delay timer isn't running, send a PLOGI.
+ * Otherwise the nlp_delay_func will expire and send
+ * a PLOGI to the remote node.
+ */
if (!test_bit(NLP_DELAY_TMO, &ndlp->nlp_flag)) {
ndlp->nlp_prev_state = NLP_STE_NPR_NODE;
lpfc_nlp_set_state(vport, ndlp, NLP_STE_PLOGI_ISSUE);
diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
index b370a64e6b67..dab5411f876d 100644
--- a/drivers/scsi/lpfc/lpfc_sli.c
+++ b/drivers/scsi/lpfc/lpfc_sli.c
@@ -2915,7 +2915,19 @@ lpfc_sli_def_mbox_cmpl(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
ndlp->nlp_defer_did != NLP_EVT_NOTHING_PENDING) {
clear_bit(NLP_UNREG_INP, &ndlp->nlp_flag);
ndlp->nlp_defer_did = NLP_EVT_NOTHING_PENDING;
- lpfc_issue_els_plogi(vport, ndlp->nlp_DID, 0);
+
+ if (!test_bit(NLP_DELAY_TMO, &ndlp->nlp_flag) &&
+ ndlp->nlp_last_elscmd == ELS_CMD_PLOGI) {
+ rc = lpfc_issue_els_plogi(vport,
+ ndlp->nlp_DID,
+ 0);
+ if (!rc) {
+ ndlp->nlp_prev_state =
+ ndlp->nlp_state;
+ lpfc_nlp_set_state(vport, ndlp,
+ NLP_STE_PLOGI_ISSUE);
+ }
+ }
} else {
clear_bit(NLP_UNREG_INP, &ndlp->nlp_flag);
}
@@ -2966,6 +2978,7 @@ lpfc_sli4_unreg_rpi_cmpl_clr(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
struct lpfc_vport *vport = pmb->vport;
struct lpfc_nodelist *ndlp;
bool unreg_inp;
+ int rc = 0;
ndlp = pmb->ctx_ndlp;
if (pmb->u.mb.mbxCommand == MBX_UNREG_LOGIN) {
@@ -3003,15 +3016,31 @@ lpfc_sli4_unreg_rpi_cmpl_clr(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
LOG_MBOX | LOG_SLI | LOG_NODE,
"4111 UNREG cmpl deferred "
"clr x%x on "
- "NPort x%x Data: x%x x%px\n",
+ "NPort x%x Data: x%x x%x x%px\n",
ndlp->nlp_rpi, ndlp->nlp_DID,
- ndlp->nlp_defer_did, ndlp);
+ ndlp->nlp_defer_did,
+ ndlp->nlp_last_elscmd,
+ ndlp);
ndlp->nlp_defer_did =
NLP_EVT_NOTHING_PENDING;
- lpfc_issue_els_plogi(
- vport, ndlp->nlp_DID, 0);
- }
+ if (!test_bit(NLP_DELAY_TMO,
+ &ndlp->nlp_flag) &&
+ ndlp->nlp_last_elscmd ==
+ ELS_CMD_PLOGI) {
+ rc = lpfc_issue_els_plogi(vport,
+ ndlp->nlp_DID, 0);
+ if (rc)
+ goto out;
+
+ ndlp->nlp_prev_state =
+ ndlp->nlp_state;
+ lpfc_nlp_set_state(vport,
+ ndlp,
+ NLP_STE_PLOGI_ISSUE);
+ }
+ }
+out:
lpfc_nlp_put(ndlp);
}
}
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (7 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:15 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler Nigel Kirkland
` (5 subsequent siblings)
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
It is unlikely that an BA_ACC will be received for a sent ABTS knowing
that a previously sent PLOGI E_D_TOV timed out.
By sending an ABORT_WQE with IA=1, the XRI_ABORTED CQE for the PLOGI CQE
SEQUENCE_TIMEOUT with XB=1 will return immediately compared to waiting
E_D_TOV for a BA_ACC.
Add a new bool ia argument variable to lpfc_sli_issue_abort_iotag, to
explicitly set the IA bit when filling out ABORT_WQE. When the ia argument
is false, we fall back to the old logic of implicitly setting the IA bit
under previous conditions. Setting the ia argument to true, is currently
only used for PLOGI CQE LOCAL_REJECT/SEQUENCE_TIMEOUT.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_bsg.c | 5 ++--
drivers/scsi/lpfc/lpfc_crtn.h | 5 ++--
drivers/scsi/lpfc/lpfc_els.c | 48 ++++++++++++++++++++++++++----
drivers/scsi/lpfc/lpfc_hbadisc.c | 3 +-
drivers/scsi/lpfc/lpfc_nportdisc.c | 3 +-
drivers/scsi/lpfc/lpfc_nvme.c | 2 +-
drivers/scsi/lpfc/lpfc_scsi.c | 2 +-
drivers/scsi/lpfc/lpfc_sli.c | 28 ++++++++++-------
8 files changed, 72 insertions(+), 24 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_bsg.c b/drivers/scsi/lpfc/lpfc_bsg.c
index 7354ae9ba8e5..63b6839230df 100644
--- a/drivers/scsi/lpfc/lpfc_bsg.c
+++ b/drivers/scsi/lpfc/lpfc_bsg.c
@@ -1,7 +1,7 @@
/*******************************************************************
* This file is part of the Emulex Linux Device Driver for *
* Fibre Channel Host Bus Adapters. *
- * Copyright (C) 2017-2024 Broadcom. All Rights Reserved. The term *
+ * Copyright (C) 2017-2026 Broadcom. All Rights Reserved. The term *
* “Broadcom” refers to Broadcom Inc. and/or its subsidiaries. *
* Copyright (C) 2009-2015 Emulex. All rights reserved. *
* EMULEX and SLI are trademarks of Emulex. *
@@ -5855,7 +5855,8 @@ lpfc_bsg_timeout(struct bsg_job *job)
}
}
if (list_empty(&completions))
- lpfc_sli_issue_abort_iotag(phba, pring, cmdiocb, NULL);
+ lpfc_sli_issue_abort_iotag(phba, pring, cmdiocb, false,
+ NULL);
spin_unlock_irqrestore(&phba->hbalock, flags);
if (!list_empty(&completions)) {
lpfc_sli_cancel_iocbs(phba, &completions,
diff --git a/drivers/scsi/lpfc/lpfc_crtn.h b/drivers/scsi/lpfc/lpfc_crtn.h
index 8a5b76bdea06..2ca5f0229ca1 100644
--- a/drivers/scsi/lpfc/lpfc_crtn.h
+++ b/drivers/scsi/lpfc/lpfc_crtn.h
@@ -402,8 +402,9 @@ int lpfc_sli_hbq_count(void);
int lpfc_sli_hbqbuf_add_hbqs(struct lpfc_hba *, uint32_t);
void lpfc_sli_hbqbuf_free_all(struct lpfc_hba *);
int lpfc_sli_hbq_size(void);
-int lpfc_sli_issue_abort_iotag(struct lpfc_hba *, struct lpfc_sli_ring *,
- struct lpfc_iocbq *, void *);
+int lpfc_sli_issue_abort_iotag(struct lpfc_hba *phba,
+ struct lpfc_sli_ring *pring,
+ struct lpfc_iocbq *cmdiocb, bool ia, void *cmpl);
int lpfc_sli_sum_iocb(struct lpfc_vport *, uint16_t, uint64_t, lpfc_ctx_cmd);
int lpfc_sli_abort_iocb(struct lpfc_vport *vport, u16 tgt_id, u64 lun_id,
lpfc_ctx_cmd abort_cmd);
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 3e26654a8286..7309fb948bea 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -1556,7 +1556,7 @@ lpfc_els_abort_flogi(struct lpfc_hba *phba)
iocb->fabric_cmd_cmpl =
lpfc_ignore_els_cmpl;
lpfc_sli_issue_abort_iotag(phba, pring, iocb,
- NULL);
+ false, NULL);
}
}
}
@@ -2129,7 +2129,7 @@ lpfc_cmpl_els_plogi(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
struct lpfc_dmabuf *prsp;
bool disc;
struct serv_parm *sp = NULL;
- u32 ulp_status, ulp_word4, did, iotag;
+ u32 ulp_status, ulp_word4, did, iotag, word3;
bool release_node = false;
/* we pass cmdiocb to state machine which needs rspiocb as well */
@@ -2141,9 +2141,11 @@ lpfc_cmpl_els_plogi(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
if (phba->sli_rev == LPFC_SLI_REV4) {
iotag = get_wqe_reqtag(cmdiocb);
+ word3 = rspiocb->wcqe_cmpl.word3;
} else {
irsp = &rspiocb->iocb;
iotag = irsp->ulpIoTag;
+ word3 = 0;
}
lpfc_debugfs_disc_trc(vport, LPFC_DISC_TRC_ELS_CMD,
@@ -2169,10 +2171,10 @@ lpfc_cmpl_els_plogi(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
/* PLOGI completes to NPort <nlp_DID> */
lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
"0102 PLOGI completes to NPort x%06x "
- "IoTag x%x Data: x%x x%x x%x x%x x%x\n",
+ "IoTag x%x Data: x%x x%x x%x x%x x%x x%x\n",
ndlp->nlp_DID, iotag,
ndlp->nlp_fc4_type,
- ulp_status, ulp_word4,
+ ulp_status, ulp_word4, word3,
disc, vport->num_disc_nodes);
/* Check to see if link went down during discovery */
@@ -4813,6 +4815,7 @@ lpfc_els_retry(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
union lpfc_wqe128 *irsp = &rspiocb->wqe;
struct lpfc_nodelist *ndlp = cmdiocb->ndlp;
struct lpfc_dmabuf *pcmd = cmdiocb->cmd_dmabuf;
+ struct lpfc_sli_ring *pring;
uint32_t *elscmd;
struct ls_rjt stat;
int retry = 0, maxretry = lpfc_max_els_tries, delay = 0;
@@ -4820,6 +4823,7 @@ lpfc_els_retry(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
uint32_t cmd = 0;
uint32_t did;
int link_reset = 0, rc;
+ unsigned long iflags;
u32 ulp_status = get_job_ulpstatus(phba, rspiocb);
u32 ulp_word4 = get_job_word4(phba, rspiocb);
u8 rsn_code_exp = 0;
@@ -4921,7 +4925,38 @@ lpfc_els_retry(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
/* Reset the Link */
link_reset = 1;
break;
+ } else if (cmd == ELS_CMD_PLOGI) {
+
+ /* if invalid ndlp, do not retry */
+ if (unlikely(!ndlp)) {
+ retry = 0;
+ break;
+ }
+
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "0159 PLOGI Sequence TMO for "
+ "ndlp x%px x%lx x%x x%x x%x "
+ "x%x x%x %u x%x\n",
+ ndlp, ndlp->nlp_flag,
+ ndlp->nlp_DID, ndlp->nlp_type,
+ ndlp->nlp_fc4_type,
+ ndlp->nlp_rpi, ndlp->nlp_state,
+ kref_read(&ndlp->kref),
+ ndlp->fc4_xpt_flags);
+
+ /* Abort all outstanding ELS and auto-ABTS. If
+ * no response for E_D_TOV, then it is unlikely
+ * auto-ABTS would receive a response too. So,
+ * inhibit the abort for faster XRI release.
+ * However, still proceed with delayed retry.
+ */
+ pring = lpfc_phba_elsring(phba);
+ spin_lock_irqsave(&phba->hbalock, iflags);
+ lpfc_sli_issue_abort_iotag(phba, pring, cmdiocb,
+ true, NULL);
+ spin_unlock_irqrestore(&phba->hbalock, iflags);
}
+
retry = 1;
delay = 100;
break;
@@ -9791,7 +9826,7 @@ lpfc_els_timeout_handler(struct lpfc_vport *vport)
spin_lock_irq(&phba->hbalock);
list_del_init(&piocb->dlist);
- lpfc_sli_issue_abort_iotag(phba, pring, piocb, NULL);
+ lpfc_sli_issue_abort_iotag(phba, pring, piocb, false, NULL);
spin_unlock_irq(&phba->hbalock);
}
@@ -9909,7 +9944,8 @@ lpfc_els_flush_cmd(struct lpfc_vport *vport)
if (mbx_tmo_err || !(phba->sli.sli_flag & LPFC_SLI_ACTIVE))
list_move_tail(&piocb->list, &cancel_list);
else
- lpfc_sli_issue_abort_iotag(phba, pring, piocb, NULL);
+ lpfc_sli_issue_abort_iotag(phba, pring, piocb, false,
+ NULL);
spin_unlock_irqrestore(&phba->hbalock, iflags);
}
diff --git a/drivers/scsi/lpfc/lpfc_hbadisc.c b/drivers/scsi/lpfc/lpfc_hbadisc.c
index b4a5c7d5c2a0..c100c8c50682 100644
--- a/drivers/scsi/lpfc/lpfc_hbadisc.c
+++ b/drivers/scsi/lpfc/lpfc_hbadisc.c
@@ -6009,7 +6009,8 @@ lpfc_free_tx(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
if (ulp_command == CMD_ELS_REQUEST64_CR ||
ulp_command == CMD_XMIT_ELS_RSP64_CX) {
- lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
+ lpfc_sli_issue_abort_iotag(phba, pring, iocb, false,
+ NULL);
}
}
spin_unlock_irq(&phba->hbalock);
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index b0af279480fb..27a891f7fcf4 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -269,7 +269,8 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
list_for_each_entry_safe(iocb, next_iocb, &abort_list, dlist) {
spin_lock_irq(&phba->hbalock);
list_del_init(&iocb->dlist);
- retval = lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
+ retval = lpfc_sli_issue_abort_iotag(phba, pring, iocb, false,
+ NULL);
spin_unlock_irq(&phba->hbalock);
/* An abort that fails here is just cancelled when the driver is
diff --git a/drivers/scsi/lpfc/lpfc_nvme.c b/drivers/scsi/lpfc/lpfc_nvme.c
index 71714ea390d9..45b5966d9e6b 100644
--- a/drivers/scsi/lpfc/lpfc_nvme.c
+++ b/drivers/scsi/lpfc/lpfc_nvme.c
@@ -743,7 +743,7 @@ __lpfc_nvme_ls_abort(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
spin_unlock(&pring->ring_lock);
if (foundit)
- lpfc_sli_issue_abort_iotag(phba, pring, wqe, NULL);
+ lpfc_sli_issue_abort_iotag(phba, pring, wqe, false, NULL);
spin_unlock_irq(&phba->hbalock);
if (foundit)
diff --git a/drivers/scsi/lpfc/lpfc_scsi.c b/drivers/scsi/lpfc/lpfc_scsi.c
index 50616b05488a..58d08476e586 100644
--- a/drivers/scsi/lpfc/lpfc_scsi.c
+++ b/drivers/scsi/lpfc/lpfc_scsi.c
@@ -5603,7 +5603,7 @@ lpfc_abort_handler(struct scsi_cmnd *cmnd)
lpfc_sli_abort_fcp_cmpl);
} else {
pring = &phba->sli.sli3_ring[LPFC_FCP_RING];
- ret_val = lpfc_sli_issue_abort_iotag(phba, pring, iocb,
+ ret_val = lpfc_sli_issue_abort_iotag(phba, pring, iocb, false,
lpfc_sli_abort_fcp_cmpl);
}
diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
index dab5411f876d..10d9030a1e88 100644
--- a/drivers/scsi/lpfc/lpfc_sli.c
+++ b/drivers/scsi/lpfc/lpfc_sli.c
@@ -4598,6 +4598,7 @@ lpfc_sli_abort_iocb_ring(struct lpfc_hba *phba, struct lpfc_sli_ring *pring)
spinlock_t *plock; /* for transmit queue access */
struct lpfc_iocbq *iocb, *next_iocb;
int offline;
+ unsigned long iflag;
if (phba->sli_rev >= LPFC_SLI_REV4)
plock = &pring->ring_lock;
@@ -4610,7 +4611,7 @@ lpfc_sli_abort_iocb_ring(struct lpfc_hba *phba, struct lpfc_sli_ring *pring)
offline = pci_channel_offline(phba->pcidev);
/* Cancel everything on txq */
- spin_lock_irq(plock);
+ spin_lock_irqsave(plock, iflag);
list_splice_init(&pring->txq, &tx_completions);
pring->txq_cnt = 0;
@@ -4620,12 +4621,19 @@ lpfc_sli_abort_iocb_ring(struct lpfc_hba *phba, struct lpfc_sli_ring *pring)
iocb->cmd_flag &= ~LPFC_IO_ON_TXCMPLQ;
list_splice_init(&pring->txcmplq, &tx_completions);
pring->txcmplq_cnt = 0;
+ spin_unlock_irqrestore(plock, iflag);
} else {
+ /* lpfc_sli_issue_abort_iotag expects the hba_lock held, but not
+ * the ring_lock.
+ */
+ spin_unlock_irqrestore(plock, iflag);
+ spin_lock_irqsave(&phba->hbalock, iflag);
/* Issue ABTS for everything on the txcmplq */
list_for_each_entry_safe(iocb, next_iocb, &pring->txcmplq, list)
- lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
+ lpfc_sli_issue_abort_iotag(phba, pring, iocb, false,
+ NULL);
+ spin_unlock_irqrestore(&phba->hbalock, iflag);
}
- spin_unlock_irq(plock);
if (!offline)
lpfc_issue_hb_tmo(phba);
@@ -12017,7 +12025,7 @@ lpfc_sli_host_down(struct lpfc_vport *vport)
if (iocb->vport != vport)
continue;
lpfc_sli_issue_abort_iotag(phba, pring, iocb,
- NULL);
+ false, NULL);
}
pring->flag = prev_pring_flag;
}
@@ -12045,7 +12053,7 @@ lpfc_sli_host_down(struct lpfc_vport *vport)
if (iocb->vport != vport)
continue;
lpfc_sli_issue_abort_iotag(phba, pring, iocb,
- NULL);
+ false, NULL);
}
pring->flag = prev_pring_flag;
}
@@ -12462,6 +12470,7 @@ lpfc_ignore_els_cmpl(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
* @phba: Pointer to HBA context object.
* @pring: Pointer to driver SLI ring object.
* @cmdiocb: Pointer to driver command iocb object.
+ * @ia: Flag to explicitly or implicitly inhibit abort.
* @cmpl: completion function.
*
* This function issues an abort iocb for the provided command iocb. In case
@@ -12474,7 +12483,7 @@ lpfc_ignore_els_cmpl(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
**/
int
lpfc_sli_issue_abort_iotag(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
- struct lpfc_iocbq *cmdiocb, void *cmpl)
+ struct lpfc_iocbq *cmdiocb, bool ia, void *cmpl)
{
struct lpfc_vport *vport = cmdiocb->vport;
struct lpfc_iocbq *abtsiocbp;
@@ -12483,7 +12492,6 @@ lpfc_sli_issue_abort_iotag(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
struct lpfc_nodelist *ndlp = NULL;
u32 ulp_command = get_job_cmnd(phba, cmdiocb);
u16 ulp_context, iotag;
- bool ia;
/*
* There are certain command types we don't want to abort. And we
@@ -12533,7 +12541,7 @@ lpfc_sli_issue_abort_iotag(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
}
/* Just close the exchange under certain conditions. */
- if (test_bit(FC_UNLOADING, &vport->load_flag) ||
+ if (ia || test_bit(FC_UNLOADING, &vport->load_flag) ||
phba->link_state < LPFC_LINK_UP ||
(phba->sli_rev == LPFC_SLI_REV4 &&
phba->sli4_hba.link_state.status == LPFC_FC_LA_TYPE_LINK_DOWN) ||
@@ -12576,7 +12584,7 @@ lpfc_sli_issue_abort_iotag(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
abort_iotag_exit:
- lpfc_printf_vlog(vport, KERN_INFO, LOG_SLI,
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS | LOG_SLI,
"0339 Abort IO XRI x%x, Original iotag x%x, "
"abort tag x%x Cmdjob : x%px Abortjob : x%px "
"retval x%x : IA %d cmd_cmpl %ps\n",
@@ -12870,7 +12878,7 @@ lpfc_sli_abort_iocb(struct lpfc_vport *vport, u16 tgt_id, u64 lun_id,
} else if (phba->sli_rev == LPFC_SLI_REV4) {
pring = lpfc_sli4_calc_ring(phba, iocbq);
}
- ret_val = lpfc_sli_issue_abort_iotag(phba, pring, iocbq,
+ ret_val = lpfc_sli_issue_abort_iotag(phba, pring, iocbq, false,
lpfc_sli_abort_fcp_cmpl);
spin_unlock_irqrestore(&phba->hbalock, iflags);
if (ret_val != IOCB_SUCCESS)
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (8 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable Nigel Kirkland
` (4 subsequent siblings)
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
There is a cqe process limit of 64 cqes in
lpfc_sli_handle_slow_ring_event_s4. The HBA_SP_QUEUE_EVT flag is not set
nor the worker thread rescheduled when reaching this 64 limit. This means
a burst of over 64 cqes can incur a delayed processing penalty, which can
be problematic in large SAN configurations waiting for rediscovery after a
link perturbation.
Remove the slowpath cqe process limiter in
lpfc_sli_handle_slow_ring_event_s4 to ensure the slow path CQ is drained.
Add log messages to notify when LPFC_WQE_DEF_COUNT cqe count is reached.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_sli.c | 29 +++++++++++++++++++++++------
1 file changed, 23 insertions(+), 6 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
index 10d9030a1e88..06dc01c9539e 100644
--- a/drivers/scsi/lpfc/lpfc_sli.c
+++ b/drivers/scsi/lpfc/lpfc_sli.c
@@ -4542,7 +4542,7 @@ lpfc_sli_handle_slow_ring_event_s4(struct lpfc_hba *phba,
struct hbq_dmabuf *dmabuf;
struct lpfc_cq_event *cq_event;
unsigned long iflag;
- int count = 0;
+ u32 count = 0;
clear_bit(HBA_SP_QUEUE_EVT, &phba->hba_flag);
while (!list_empty(&phba->sli4_hba.sp_queue_event)) {
@@ -4562,22 +4562,39 @@ lpfc_sli_handle_slow_ring_event_s4(struct lpfc_hba *phba,
if (irspiocbq)
lpfc_sli_sp_handle_rspiocb(phba, pring,
irspiocbq);
- count++;
break;
case CQE_CODE_RECEIVE:
case CQE_CODE_RECEIVE_V1:
dmabuf = container_of(cq_event, struct hbq_dmabuf,
cq_event);
lpfc_sli4_handle_received_buffer(phba, dmabuf);
- count++;
break;
default:
+ lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
+ "7771 Unknown WCQE completion code "
+ "x%x, ignoring.\n",
+ bf_get(lpfc_wcqe_c_code,
+ &cq_event->cqe.wcqe_cmpl));
break;
}
- /* Limit the number of events to 64 to avoid soft lockups */
- if (count == 64)
- break;
+ /* This loop runs until the ELS/CT CQ is empty. Post a one
+ * time message for debug support when ELS WQ ecount
+ * completions are processed - this represent 1 full ELS WQ
+ * wrap.
+ */
+ if (++count == LPFC_WQE_DEF_COUNT) {
+ lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
+ "7772 %s SP CQE count %d\n",
+ __func__, count);
+ }
+ }
+
+ /* Log a final message to note how many CQEs were processed. */
+ if (count > LPFC_WQE_DEF_COUNT) {
+ lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
+ "7773 %s SP CQEs complete, count %d\n",
+ __func__, count);
}
}
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (9 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting Nigel Kirkland
` (3 subsequent siblings)
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
When ELS/CT commands can't be sent due to ELS WQ full, queue the iocbq on
the phba->txq tail for retrying submission later in lpfc_drain_txq.
lpfc_drain_txq is flagged to be called by the worker thread when an ELS CQE
completes lpfc_sli4_sp_handle_els_wcqe through the HBA_SP_QUEUE_EVT flag.
lpfc_drain_txq is also updated to queue an iocbq back on to the head of
phba->txq when lpfc_drain_txq itself still observes ELS WQ full or SGL
unavailable events.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_bsg.c | 2 +-
drivers/scsi/lpfc/lpfc_crtn.h | 7 +-
drivers/scsi/lpfc/lpfc_ct.c | 21 ++++-
drivers/scsi/lpfc/lpfc_els.c | 110 +++++++++++++++++---------
drivers/scsi/lpfc/lpfc_init.c | 5 +-
drivers/scsi/lpfc/lpfc_nportdisc.c | 9 ++-
drivers/scsi/lpfc/lpfc_sli.c | 121 ++++++++++++++++++++++++-----
drivers/scsi/lpfc/lpfc_sli.h | 18 +++++
8 files changed, 227 insertions(+), 66 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_bsg.c b/drivers/scsi/lpfc/lpfc_bsg.c
index 63b6839230df..7919d3bf2cce 100644
--- a/drivers/scsi/lpfc/lpfc_bsg.c
+++ b/drivers/scsi/lpfc/lpfc_bsg.c
@@ -2968,7 +2968,7 @@ static int lpfcdiag_sli3_loop_post_rxbufs(struct lpfc_hba *phba, uint16_t rxxri,
iocb_stat = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, cmdiocbq,
0);
- if (iocb_stat == IOCB_ERROR) {
+ if (lpfc_iocb_failed(iocb_stat)) {
diag_cmd_data_free(phba,
(struct lpfc_dmabufext *)mp[0]);
if (mp[1])
diff --git a/drivers/scsi/lpfc/lpfc_crtn.h b/drivers/scsi/lpfc/lpfc_crtn.h
index 2ca5f0229ca1..47a593fe5e69 100644
--- a/drivers/scsi/lpfc/lpfc_crtn.h
+++ b/drivers/scsi/lpfc/lpfc_crtn.h
@@ -536,8 +536,11 @@ int lpfc_bsg_timeout(struct bsg_job *);
int lpfc_bsg_ct_unsol_event(struct lpfc_hba *, struct lpfc_sli_ring *,
struct lpfc_iocbq *);
int lpfc_bsg_ct_unsol_abort(struct lpfc_hba *, struct hbq_dmabuf *);
-void __lpfc_sli_ringtx_put(struct lpfc_hba *, struct lpfc_sli_ring *,
- struct lpfc_iocbq *);
+void lpfc_sli4_queue_io_for_retry(struct lpfc_hba *phba,
+ struct lpfc_iocbq *iocb,
+ bool head);
+void __lpfc_sli_ringtx_put(struct lpfc_hba *phba, struct lpfc_sli_ring *ring,
+ struct lpfc_iocbq *iocb, bool head);
struct lpfc_iocbq *lpfc_sli_ringtx_get(struct lpfc_hba *,
struct lpfc_sli_ring *);
int __lpfc_sli_issue_iocb(struct lpfc_hba *, uint32_t,
diff --git a/drivers/scsi/lpfc/lpfc_ct.c b/drivers/scsi/lpfc/lpfc_ct.c
index f709e5087577..372fa5718f3a 100644
--- a/drivers/scsi/lpfc/lpfc_ct.c
+++ b/drivers/scsi/lpfc/lpfc_ct.c
@@ -246,7 +246,7 @@ lpfc_ct_reject_event(struct lpfc_nodelist *ndlp,
goto ct_no_ndlp;
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, cmdiocbq, 0);
- if (rc) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_nlp_put(ndlp);
goto ct_no_ndlp;
}
@@ -639,11 +639,24 @@ lpfc_gen_req(struct lpfc_vport *vport, struct lpfc_dmabuf *bmp,
goto out;
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, geniocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
+ lpfc_vlog_msg(vport, KERN_NOTICE,
+ LOG_ELS | LOG_DISCOVERY | LOG_NODE,
+ "0156 %ps WQE Put returned %d\n",
+ cmpl, rc);
+
+ /* Under heavy vpi counts, the driver's host_index can catch up
+ * to the hba_index causing a put error. Catch this case and
+ * put the IO on phba->txq.
+ */
+ if (rc == IOCB_FAILED_PUT && phba->sli_rev == LPFC_SLI_REV4) {
+ lpfc_sli4_queue_io_for_retry(phba, geniocb, false);
+ return 0;
+ }
+
lpfc_nlp_put(ndlp);
goto out;
}
-
return 0;
out:
lpfc_sli_release_iocbq(phba, geniocb);
@@ -2258,7 +2271,7 @@ lpfc_cmpl_ct_disc_fdmi(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
/* Retry the same FDMI command */
err = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING,
cmdiocb, 0);
- if (err == IOCB_ERROR)
+ if (lpfc_iocb_failed(err))
break;
return;
default:
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index 7309fb948bea..bf71b5a3e55e 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -1431,7 +1431,7 @@ lpfc_issue_els_flogi(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
set_bit(HBA_FLOGI_OUTSTANDING, &phba->hba_flag);
rc = lpfc_issue_fabric_iocb(phba, elsiocb);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
clear_bit(HBA_FLOGI_ISSUED, &phba->hba_flag);
clear_bit(HBA_FLOGI_OUTSTANDING, &phba->hba_flag);
lpfc_nlp_put(ndlp);
@@ -2439,7 +2439,7 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
}
ret = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (ret) {
+ if (lpfc_iocb_failed(ret)) {
lpfc_vlog_msg(vport, KERN_NOTICE,
LOG_ELS | LOG_DISCOVERY | LOG_NODE,
"0157 PLOGI WQE Put returned %d\n",
@@ -2450,8 +2450,15 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
* put the IO on phba->txq.
*/
if (ret == IOCB_FAILED_PUT && phba->sli_rev == LPFC_SLI_REV4) {
+ /* The iocb will always successfully get queued
+ * for a retry
+ */
lpfc_sli4_queue_io_for_retry(phba, elsiocb, false);
- return 0;
+
+ /* Goto the out label to mark the nlp_flag correctly.
+ * The IO will eventually go out.
+ */
+ goto out;
}
lpfc_els_free_iocb(phba, elsiocb);
@@ -2459,6 +2466,7 @@ lpfc_issue_els_plogi(struct lpfc_vport *vport, uint32_t did, uint8_t retry)
return 1;
}
+out:
set_bit(NLP_PLOGI_SND, &ndlp->nlp_flag);
return 0;
}
@@ -2780,7 +2788,7 @@ lpfc_issue_els_prli(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_vlog_msg(vport, KERN_NOTICE,
LOG_ELS | LOG_DISCOVERY | LOG_NODE,
"0155 PRLI WQE Put returned %d\n",
@@ -2791,8 +2799,15 @@ lpfc_issue_els_prli(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
* put the IO on phba->txq.
*/
if (rc == IOCB_FAILED_PUT && phba->sli_rev == LPFC_SLI_REV4) {
+ /* The iocb will always successfully get queued for a
+ * retry
+ */
lpfc_sli4_queue_io_for_retry(phba, elsiocb, false);
- return 0;
+
+ /* Jump to out_next because the prli counters and an NVME
+ * PRLI need to be exercised.
+ */
+ goto out_next;
}
lpfc_els_free_iocb(phba, elsiocb);
@@ -2800,6 +2815,7 @@ lpfc_issue_els_prli(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
return 1;
}
+out_next:
/* The vport counters are used for lpfc_scan_finished, but
* the ndlp is used to track outstanding PRLIs for different
* FC4 types.
@@ -3121,7 +3137,7 @@ lpfc_issue_els_adisc(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
ndlp->nlp_DID, kref_read(&ndlp->kref), 0);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
goto err;
@@ -3331,7 +3347,7 @@ lpfc_issue_els_logo(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
ndlp->nlp_DID, kref_read(&ndlp->kref), 0);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
goto err;
@@ -3701,7 +3717,7 @@ lpfc_issue_els_scr(struct lpfc_vport *vport, uint8_t retry)
ndlp->nlp_DID, kref_read(&ndlp->kref), 0);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR)
+ if (lpfc_iocb_failed(rc))
goto out_iocb_error;
return 0;
@@ -3804,7 +3820,7 @@ lpfc_issue_els_rscn(struct lpfc_vport *vport, uint8_t retry)
ndlp->nlp_DID, kref_read(&ndlp->kref), 0);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -3899,7 +3915,7 @@ lpfc_issue_els_farpr(struct lpfc_vport *vport, uint32_t nportid, uint8_t retry)
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
/* The additional lpfc_nlp_put will cause the following
* lpfc_els_free_iocb routine to trigger the release of
* the node.
@@ -4000,7 +4016,7 @@ lpfc_issue_els_rdf(struct lpfc_vport *vport, uint8_t retry)
ndlp->nlp_DID, kref_read(&ndlp->kref), 0);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
err = -EIO;
goto out_iocb_error;
}
@@ -4529,7 +4545,7 @@ lpfc_issue_els_edc(struct lpfc_vport *vport, uint8_t retry)
"Issue EDC: did:x%x refcnt %d",
ndlp->nlp_DID, kref_read(&ndlp->kref), 0);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
/* The additional lpfc_nlp_put will cause the following
* lpfc_els_free_iocb routine to trigger the release of
* the node.
@@ -6010,7 +6026,21 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
+ lpfc_vlog_msg(vport, KERN_NOTICE,
+ LOG_ELS | LOG_DISCOVERY | LOG_NODE,
+ "0154 LS_ACC WQE Put returned %d\n",
+ rc);
+
+ /* Under heavy vpi counts, the driver's host_index can catch up
+ * to the hba_index causing a put error. Catch this case and
+ * put the IO on phba->txq.
+ */
+ if (rc == IOCB_FAILED_PUT && phba->sli_rev == LPFC_SLI_REV4) {
+ lpfc_sli4_queue_io_for_retry(phba, elsiocb, false);
+ return 0;
+ }
+
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -6112,7 +6142,7 @@ lpfc_els_rsp_reject(struct lpfc_vport *vport, uint32_t rejectError,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -6204,7 +6234,7 @@ lpfc_issue_els_edc_rsp(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -6310,7 +6340,7 @@ lpfc_els_rsp_adisc_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -6502,7 +6532,7 @@ lpfc_els_rsp_prli_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -6616,7 +6646,7 @@ lpfc_els_rsp_rnid_acc(struct lpfc_vport *vport, uint8_t format,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -6750,7 +6780,7 @@ lpfc_els_rsp_echo_acc(struct lpfc_vport *vport, uint8_t *data,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -7418,7 +7448,7 @@ lpfc_els_rdp_cmpl(struct lpfc_hba *phba, struct lpfc_rdp_context *rdp_context,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
}
@@ -7461,7 +7491,7 @@ lpfc_els_rdp_cmpl(struct lpfc_hba *phba, struct lpfc_rdp_context *rdp_context,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
}
@@ -7824,7 +7854,7 @@ lpfc_els_lcb_rsp(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
}
@@ -7870,7 +7900,7 @@ lpfc_els_lcb_rsp(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
}
@@ -8972,7 +9002,7 @@ lpfc_els_rsp_rls_acc(struct lpfc_hba *phba, LPFC_MBOXQ_t *pmb)
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
}
@@ -9136,7 +9166,7 @@ lpfc_els_rcv_rtv(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
}
@@ -9211,7 +9241,7 @@ lpfc_issue_els_rrq(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
goto io_err;
ret = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (ret == IOCB_ERROR) {
+ if (lpfc_iocb_failed(ret)) {
lpfc_nlp_put(ndlp);
goto io_err;
}
@@ -9332,7 +9362,7 @@ lpfc_els_rsp_rpl_acc(struct lpfc_vport *vport, uint16_t cmdsize,
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return 1;
@@ -10611,6 +10641,7 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
uint8_t rjt_exp, rjt_err = 0, init_link = 0;
struct lpfc_wcqe_complete *wcqe_cmpl = NULL;
LPFC_MBOXQ_t *mbox;
+ u16 rcv_oxid;
if (!vport || !elsiocb->cmd_dmabuf)
goto dropit;
@@ -10680,11 +10711,18 @@ lpfc_els_unsol_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
if ((cmd & ELS_CMD_MASK) == ELS_CMD_RSCN) {
cmd &= ELS_CMD_MASK;
}
+
+ /* Fetch the incoming ox_id for SLI4 only - SLI3 can't do this. */
+ rcv_oxid = 0xffff;
+ if (phba->sli_rev == LPFC_SLI_REV4)
+ rcv_oxid = get_job_rcvoxid(phba, elsiocb);
+
/* ELS command <elsCmd> received from NPORT <did> */
lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
"0112 ELS command x%x received from NPORT x%x "
- "refcnt %d Data: x%x x%lx x%x x%x\n",
- cmd, did, kref_read(&ndlp->kref), vport->port_state,
+ "refcnt %d rcvoxid x%x Data: x%x x%lx x%x x%x\n",
+ cmd, did, kref_read(&ndlp->kref),
+ rcv_oxid, vport->port_state,
vport->fc_flag, vport->fc_myDID, vport->fc_prevDID);
/* reject till our FLOGI completes or PLOGI assigned DID via PT2PT */
@@ -11773,7 +11811,7 @@ lpfc_issue_els_fdisc(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp,
goto err_out;
rc = lpfc_issue_fabric_iocb(phba, elsiocb);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_nlp_put(ndlp);
goto err_out;
}
@@ -11911,7 +11949,7 @@ lpfc_issue_els_npiv_logo(struct lpfc_vport *vport, struct lpfc_nodelist *ndlp)
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
goto err;
@@ -11992,8 +12030,7 @@ lpfc_resume_fabric_iocbs(struct lpfc_hba *phba)
iocb->vport->port_state, 0, 0);
ret = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, iocb, 0);
-
- if (ret == IOCB_ERROR) {
+ if (lpfc_iocb_failed(ret)) {
iocb->cmd_cmpl = iocb->fabric_cmd_cmpl;
iocb->fabric_cmd_cmpl = NULL;
iocb->cmd_flag &= ~LPFC_IO_FABRIC;
@@ -12157,8 +12194,7 @@ lpfc_issue_fabric_iocb(struct lpfc_hba *phba, struct lpfc_iocbq *iocb)
iocb->vport->port_state, 0, 0);
ret = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, iocb, 0);
-
- if (ret == IOCB_ERROR) {
+ if (lpfc_iocb_failed(ret)) {
iocb->cmd_cmpl = iocb->fabric_cmd_cmpl;
iocb->fabric_cmd_cmpl = NULL;
iocb->cmd_flag &= ~LPFC_IO_FABRIC;
@@ -12580,7 +12616,7 @@ int lpfc_issue_els_qfpa(struct lpfc_vport *vport)
}
ret = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 2);
- if (ret != IOCB_SUCCESS) {
+ if (lpfc_iocb_failed(ret)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
return -EIO;
@@ -12664,7 +12700,7 @@ lpfc_vmid_uvem(struct lpfc_vport *vport,
}
ret = lpfc_sli_issue_iocb(vport->phba, LPFC_ELS_RING, elsiocb, 0);
- if (ret != IOCB_SUCCESS) {
+ if (lpfc_iocb_failed(ret)) {
lpfc_els_free_iocb(vport->phba, elsiocb);
lpfc_nlp_put(ndlp);
goto out;
diff --git a/drivers/scsi/lpfc/lpfc_init.c b/drivers/scsi/lpfc/lpfc_init.c
index 7352cb6e584b..a423bde283c9 100644
--- a/drivers/scsi/lpfc/lpfc_init.c
+++ b/drivers/scsi/lpfc/lpfc_init.c
@@ -2813,6 +2813,7 @@ lpfc_sli3_post_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring, int cn
IOCB_t *icmd;
struct lpfc_iocbq *iocb;
struct lpfc_dmabuf *mp1, *mp2;
+ int rc;
cnt += pring->missbufcnt;
@@ -2875,8 +2876,8 @@ lpfc_sli3_post_buffer(struct lpfc_hba *phba, struct lpfc_sli_ring *pring, int cn
icmd->ulpCommand = CMD_QUE_RING_BUF64_CN;
icmd->ulpLe = 1;
- if (lpfc_sli_issue_iocb(phba, pring->ringno, iocb, 0) ==
- IOCB_ERROR) {
+ rc = lpfc_sli_issue_iocb(phba, pring->ringno, iocb, 0);
+ if (lpfc_iocb_failed(rc)) {
lpfc_mbuf_free(phba, mp1->virt, mp1->phys);
kfree(mp1);
cnt++;
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index 27a891f7fcf4..49653ff35c56 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -1998,6 +1998,7 @@ lpfc_cmpl_reglogin_reglogin_issue(struct lpfc_vport *vport,
LPFC_MBOXQ_t *pmb = (LPFC_MBOXQ_t *) arg;
MAILBOX_t *mb = &pmb->u.mb;
uint32_t did = mb->un.varWords[1];
+ int rc;
if (mb->mbxStatus) {
/* RegLogin failed */
@@ -2076,7 +2077,13 @@ lpfc_cmpl_reglogin_reglogin_issue(struct lpfc_vport *vport,
ndlp->nlp_prev_state = NLP_STE_REG_LOGIN_ISSUE;
lpfc_nlp_set_state(vport, ndlp, NLP_STE_PRLI_ISSUE);
- if (lpfc_issue_els_prli(vport, ndlp, 0)) {
+ rc = lpfc_issue_els_prli(vport, ndlp, 0);
+ if (rc) {
+ lpfc_vlog_msg(vport, KERN_NOTICE,
+ LOG_ELS | LOG_DISCOVERY | LOG_NODE,
+ "3015 PRLI Issue returning %d to DID "
+ "x%06x, Send LOGO\n",
+ rc, ndlp->nlp_DID);
lpfc_issue_els_logo(vport, ndlp, 0);
ndlp->nlp_prev_state = NLP_STE_REG_LOGIN_ISSUE;
lpfc_nlp_set_state(vport, ndlp, NLP_STE_NPR_NODE);
diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
index 06dc01c9539e..bc46b14eca41 100644
--- a/drivers/scsi/lpfc/lpfc_sli.c
+++ b/drivers/scsi/lpfc/lpfc_sli.c
@@ -244,6 +244,38 @@ lpfc_sli4_pcimem_bcopy(void *srcp, void *destp, uint32_t cnt)
#define lpfc_sli4_pcimem_bcopy(a, b, c) lpfc_sli_pcimem_bcopy(a, b, c)
#endif
+/**
+ * lpfc_sli4_queue_io_for_retry - Put an ELS/CT IO request on the txq.
+ * @phba - pointer to the hba instance.
+ * @iocb - the IO request to put on the txq.
+ * @head - A boolean to request a put to the head (true) or tail
+ * (false) of the txq.
+ *
+ * This routine is for SLI4 interfaces only.
+ * When ELS/CT commands can't be sent because of ELS WQ full put status,
+ * queue the iocbq on the phba->txq, at the head or tail as requested by
+ * the caller, for a retry later in lpfc_drain_txq. The lpfc_drain_txq
+ * routine is called per ELS CQE completion.
+ */
+void lpfc_sli4_queue_io_for_retry(struct lpfc_hba *phba,
+ struct lpfc_iocbq *iocb,
+ bool head)
+{
+ unsigned long iflags;
+ struct lpfc_sli_ring *pring;
+
+ /* Mark the IO as in RETRY to stop a full reprep of the SGL/XRI. */
+ iocb->cmd_flag |= LPFC_IO_IN_RETRY;
+ pring = phba->sli4_hba.els_wq->pring;
+
+ /* Insert to the head or tail of the txq ring as indicated
+ * by the caller's head argument.
+ */
+ spin_lock_irqsave(&pring->ring_lock, iflags);
+ __lpfc_sli_ringtx_put(phba, pring, iocb, head);
+ spin_unlock_irqrestore(&pring->ring_lock, iflags);
+}
+
/**
* lpfc_sli4_wq_put - Put a Work Queue Entry on an Work Queue
* @q: The Work Queue to operate on.
@@ -276,6 +308,10 @@ lpfc_sli4_wq_put(struct lpfc_queue *q, union lpfc_wqe128 *wqe)
/* If the host has not yet processed the next entry then we are done */
idx = ((q->host_index + 1) % q->entry_count);
if (idx == q->hba_index) {
+ lpfc_printf_log(q->phba, KERN_WARNING, LOG_SLI,
+ "9998 No available WQ Slots on "
+ "q_id x%x, host x%x, hba x%x\n",
+ q->queue_id, idx, q->hba_index);
q->WQ_overflow++;
return -EBUSY;
}
@@ -10445,6 +10481,7 @@ lpfc_mbox_api_table_setup(struct lpfc_hba *phba, uint8_t dev_grp)
* @phba: Pointer to HBA context object.
* @pring: Pointer to driver SLI ring object.
* @piocb: Pointer to address of newly added command iocb.
+ * @head: put at head (true) or tail (false)
*
* This function is called with hbalock held for SLI3 ports or
* the ring lock held for SLI4 ports to add a command
@@ -10453,14 +10490,20 @@ lpfc_mbox_api_table_setup(struct lpfc_hba *phba, uint8_t dev_grp)
**/
void
__lpfc_sli_ringtx_put(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
- struct lpfc_iocbq *piocb)
+ struct lpfc_iocbq *piocb, bool head)
{
if (phba->sli_rev == LPFC_SLI_REV4)
lockdep_assert_held(&pring->ring_lock);
else
lockdep_assert_held(&phba->hbalock);
- /* Insert the caller's iocb in the txq tail for later processing. */
- list_add_tail(&piocb->list, &pring->txq);
+
+ /* Insert the caller's iocb in the txq head or tail for later
+ * processing.
+ */
+ if (head)
+ list_add(&piocb->list, &pring->txq);
+ else
+ list_add_tail(&piocb->list, &pring->txq);
}
/**
@@ -10468,6 +10511,7 @@ __lpfc_sli_ringtx_put(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
* @phba: Pointer to HBA context object.
* @pring: Pointer to driver SLI ring object.
* @piocb: Pointer to address of newly added command iocb.
+ * @head: put at head (true) or tail (false)
*
* This function is called with hbalock held before a new
* iocb is submitted to the firmware. This function checks
@@ -10613,7 +10657,7 @@ __lpfc_sli_issue_iocb_s3(struct lpfc_hba *phba, uint32_t ring_number,
out_busy:
if (!(flag & SLI_IOCB_RET_IOCB)) {
- __lpfc_sli_ringtx_put(phba, pring, piocb);
+ __lpfc_sli_ringtx_put(phba, pring, piocb, false);
return IOCB_SUCCESS;
}
@@ -10751,6 +10795,7 @@ __lpfc_sli_issue_iocb_s4(struct lpfc_hba *phba, uint32_t ring_number,
struct lpfc_queue *wq;
struct lpfc_sli_ring *pring;
u32 ulp_command = get_job_cmnd(phba, piocb);
+ int rc = IOCB_SUCCESS;
/* Get the WQ */
if ((piocb->cmd_flag & LPFC_IO_FCP) ||
@@ -10766,9 +10811,11 @@ __lpfc_sli_issue_iocb_s4(struct lpfc_hba *phba, uint32_t ring_number,
/*
* The WQE can be either 64 or 128 bytes,
*/
-
lockdep_assert_held(&pring->ring_lock);
wqe = &piocb->wqe;
+ if (piocb->cmd_flag & LPFC_IO_IN_RETRY)
+ goto retry_io;
+
if (piocb->sli4_xritag == NO_XRI) {
if (ulp_command == CMD_ABORT_XRI_CX)
sglq = NULL;
@@ -10778,7 +10825,7 @@ __lpfc_sli_issue_iocb_s4(struct lpfc_hba *phba, uint32_t ring_number,
if (!(flag & SLI_IOCB_RET_IOCB)) {
__lpfc_sli_ringtx_put(phba,
pring,
- piocb);
+ piocb, false);
return IOCB_SUCCESS;
} else {
return IOCB_BUSY;
@@ -10819,12 +10866,19 @@ __lpfc_sli_issue_iocb_s4(struct lpfc_hba *phba, uint32_t ring_number,
return IOCB_ERROR;
}
- if (lpfc_sli4_wq_put(wq, wqe))
- return IOCB_ERROR;
-
- lpfc_sli_ringtxcmpl_put(phba, pring, piocb);
+ retry_io:
+ piocb->cmd_flag &= ~LPFC_IO_IN_RETRY;
- return 0;
+ /* Push the wqe to the wq and if successful, push to the txcmplq.
+ * A return of -EBUSY the WQ is currently full and a retry possible.
+ * Otherwise the put failed for another, nonretryable reason.
+ */
+ rc = lpfc_sli4_wq_put(wq, wqe);
+ if (!rc)
+ lpfc_sli_ringtxcmpl_put(phba, pring, piocb);
+ else if (rc == -EBUSY)
+ return IOCB_FAILED_PUT;
+ return rc;
}
/*
@@ -12608,7 +12662,7 @@ lpfc_sli_issue_abort_iotag(struct lpfc_hba *phba, struct lpfc_sli_ring *pring,
ulp_context, (phba->sli_rev == LPFC_SLI_REV4) ?
cmdiocb->iotag : iotag, iotag, cmdiocb, abtsiocbp,
retval, ia, abtsiocbp->cmd_cmpl);
- if (retval) {
+ if (lpfc_iocb_failed(retval)) {
cmdiocb->cmd_flag &= ~LPFC_DRIVER_ABORTED;
__lpfc_sli_release_iocbq(phba, abtsiocbp);
}
@@ -13060,7 +13114,7 @@ lpfc_sli_abort_taskmgmt(struct lpfc_vport *vport, struct lpfc_sli_ring *pring,
spin_unlock(&lpfc_cmd->buf_lock);
- if (ret_val == IOCB_ERROR)
+ if (lpfc_iocb_failed(ret_val))
__lpfc_sli_release_iocbq(phba, abtsiocbq);
else
sum++;
@@ -19200,7 +19254,7 @@ lpfc_sli4_seq_abort_rsp(struct lpfc_vport *vport,
ctiocb->abort_rctl, oxid, phba->link_state);
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, ctiocb, 0);
- if (rc == IOCB_ERROR) {
+ if (lpfc_iocb_failed(rc)) {
lpfc_printf_vlog(vport, KERN_ERR, LOG_TRACE_EVENT,
"2925 Failed to issue CT ABTS RSP x%x on "
"xri x%x, Data x%x\n",
@@ -19581,7 +19635,7 @@ lpfc_sli4_handle_mds_loopback(struct lpfc_vport *vport,
iocbq->cmd_cmpl = lpfc_sli4_mds_loopback_cmpl;
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, iocbq, 0);
- if (rc == IOCB_ERROR)
+ if (lpfc_iocb_failed(rc))
goto exit;
lpfc_in_buf_free(phba, &dmabuf->dbuf);
@@ -21301,12 +21355,15 @@ lpfc_drain_txq(struct lpfc_hba *phba)
}
txq_cnt--;
+ /* Capture the return values that indicate an error during IO
+ * submit and cannot be retried. Prefix message 2822.
+ */
ret = __lpfc_sli_issue_iocb(phba, pring->ringno, piocbq, 0);
-
- if (ret && ret != IOCB_BUSY) {
+ if (ret && ret != IOCB_BUSY && ret != IOCB_FAILED_PUT) {
fail_msg = " - Cannot send IO ";
piocbq->cmd_flag &= ~LPFC_DRIVER_ABORTED;
}
+
if (fail_msg) {
piocbq->cmd_flag |= LPFC_DRIVER_ABORTED;
/* Failed means we can't issue and need to cancel */
@@ -21319,9 +21376,35 @@ lpfc_drain_txq(struct lpfc_hba *phba)
list_add_tail(&piocbq->list, &completions);
fail_msg = NULL;
}
- spin_unlock_irqrestore(&pring->ring_lock, iflags);
- if (txq_cnt == 0 || ret == IOCB_BUSY)
+
+ if (txq_cnt == 0 || ret == IOCB_BUSY ||
+ ret == IOCB_FAILED_PUT) {
+ /* IOCB_FAILED_PUT is unique to SLI4 and means SGL/XRI
+ * resources are allocated. For this case, set the
+ * in retry flag. For SLI3 and 4, push the IO to the
+ * txq for retry.
+ */
+ lpfc_printf_log(phba, KERN_INFO, LOG_SLI,
+ "2820 IOCB iotag x%x xri x%x ret %d "
+ "cmd_flg x%x txq_cnt x%x\n",
+ piocbq->iotag, piocbq->sli4_xritag, ret,
+ piocbq->cmd_flag, txq_cnt);
+ switch (ret) {
+ case IOCB_FAILED_PUT:
+ piocbq->cmd_flag |= LPFC_IO_IN_RETRY;
+ __lpfc_sli_ringtx_put(phba, pring, piocbq,
+ true);
+ break;
+ case IOCB_BUSY:
+ __lpfc_sli_ringtx_put(phba, pring, piocbq,
+ true);
+ break;
+ }
+ spin_unlock_irqrestore(&pring->ring_lock, iflags);
break;
+ }
+
+ spin_unlock_irqrestore(&pring->ring_lock, iflags);
}
/* Cancel all the IOCBs that cannot be issued */
lpfc_sli_cancel_iocbs(phba, &completions, IOSTAT_LOCAL_REJECT,
diff --git a/drivers/scsi/lpfc/lpfc_sli.h b/drivers/scsi/lpfc/lpfc_sli.h
index cf7c42ec0306..37559efd0032 100644
--- a/drivers/scsi/lpfc/lpfc_sli.h
+++ b/drivers/scsi/lpfc/lpfc_sli.h
@@ -123,6 +123,7 @@ struct lpfc_iocbq {
#define LPFC_IO_NVMET 0x800000 /* NVMET command */
#define LPFC_IO_VMID 0x1000000 /* VMID tagged IO */
#define LPFC_IO_CMF 0x4000000 /* CMF command */
+#define LPFC_IO_IN_RETRY 0x8000000 /* Caller is retrying IO. */
uint32_t drvrTimeout; /* driver timeout in seconds */
struct lpfc_vport *vport;/* virtual port pointer */
@@ -160,6 +161,23 @@ struct lpfc_iocbq {
#define IOCB_ABORTED 4
#define IOCB_ABORTING 5
#define IOCB_NORESOURCE 6
+#define IOCB_FAILED_PUT 7
+
+/**
+ * lpfc_iocb_failed - test an IOCB issue return status for failure
+ * @rc: return status from an IOCB issue routine (e.g. lpfc_sli_issue_iocb).
+ *
+ * IOCB_SUCCESS is the only non-failure return value. Callers must test
+ * rc != IOCB_SUCCESS rather than rc == IOCB_ERROR, since that misses
+ * IOCB_FAILED_PUT (SLI-4 WQ full) and treats it as success.
+ *
+ * Return: true if the IOCB issue did not succeed.
+ */
+static inline bool
+lpfc_iocb_failed(int rc)
+{
+ return rc != IOCB_SUCCESS;
+}
#define SLI_WQE_RET_WQE 1 /* Return WQE if cmd ring full */
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (10 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:22 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 13/14] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
` (2 subsequent siblings)
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
Currently, there are ELS ACC routines that lack debug log messages to
indicate when ACC frame transmission run into issues. The generic
lpfc_els_rsp_acc and more specific ACC routines are updated to log when
there is an issue with transmitting the frame. The routines are also
updated to return different return codes when encountering various
transmission issues and their function comment header is updated.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_els.c | 187 ++++++++++++++++++++++++++---------
1 file changed, 141 insertions(+), 46 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index bf71b5a3e55e..cb06dbc9edfa 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -4053,13 +4053,8 @@ lpfc_els_rcv_rdf(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
int rc;
rc = lpfc_els_rsp_acc(vport, ELS_CMD_RDF, cmdiocb, ndlp, NULL);
- /* Send LS_ACC */
- if (rc) {
- lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS | LOG_CGN_MGMT,
- "1623 Failed to RDF_ACC from x%x for x%x Data: %d\n",
- ndlp->nlp_DID, vport->fc_myDID, rc);
+ if (rc)
return -EIO;
- }
rc = lpfc_issue_els_rdf(vport, 0);
/* Issue new RDF for reregistering */
@@ -5793,7 +5788,10 @@ lpfc_cmpl_els_rsp(struct lpfc_hba *phba, struct lpfc_iocbq *cmdiocb,
*
* Return code
* 0 - Successfully issued acc response
- * 1 - Failed to issue acc response
+ * -ENOMEM - IOCB not prepped successfully
+ * -EIO - The IOCB failed to issue successfully
+ * -ENODEV - No associated node for IOCB
+ * -EACCES - ACC unhandled for this command
**/
int
lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
@@ -5812,6 +5810,8 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
int rc;
ELS_PKT *els_pkt_ptr;
struct fc_els_rdf_resp *rdf_resp;
+ int err;
+ uint32_t old_opcode = 0;
switch (flag) {
case ELS_CMD_ACC:
@@ -5820,7 +5820,8 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
ndlp, ndlp->nlp_DID, ELS_CMD_ACC);
if (!elsiocb) {
clear_bit(NLP_LOGO_ACC, &ndlp->nlp_flag);
- return 1;
+ err = -ENOMEM;
+ goto err_out;
}
if (phba->sli_rev == LPFC_SLI_REV4) {
@@ -5855,8 +5856,10 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
cmdsize = (sizeof(struct serv_parm) + sizeof(uint32_t));
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry,
ndlp, ndlp->nlp_DID, ELS_CMD_ACC);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -5933,8 +5936,10 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
cmdsize = sizeof(uint32_t) + sizeof(PRLO);
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry,
ndlp, ndlp->nlp_DID, ELS_CMD_PRLO);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -5971,8 +5976,10 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
cmdsize = sizeof(*rdf_resp);
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry,
ndlp, ndlp->nlp_DID, ELS_CMD_ACC);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -6007,7 +6014,8 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
rdf_resp->lsri.rqst_w0.cmd = ELS_RDF;
break;
default:
- return 1;
+ err = -EACCES;
+ goto err_out;
}
if (test_bit(NLP_LOGO_ACC, &ndlp->nlp_flag)) {
if (!test_bit(NLP_RPI_REGISTERED, &ndlp->nlp_flag) &&
@@ -6022,7 +6030,8 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
elsiocb->ndlp = lpfc_nlp_get(ndlp);
if (!elsiocb->ndlp) {
lpfc_els_free_iocb(phba, elsiocb);
- return 1;
+ err = -ENODEV;
+ goto err_out;
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
@@ -6043,7 +6052,8 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
- return 1;
+ err = -EIO;
+ goto err_out;
}
/* Xmit ELS ACC response tag <ulpIoTag> */
@@ -6055,6 +6065,17 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
ndlp->nlp_DID, ndlp->nlp_flag, ndlp->nlp_state,
ndlp->nlp_rpi, vport->fc_flag, kref_read(&ndlp->kref));
return 0;
+
+err_out:
+ if (oldiocb->cmd_dmabuf && oldiocb->cmd_dmabuf->virt)
+ old_opcode = *(uint32_t *)oldiocb->cmd_dmabuf->virt;
+
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "1027 Xmit ELS ACC Unsuccessful: "
+ "cmd: x%x, error_code: %d "
+ "S_ID: x%x\n", old_opcode, err,
+ vport->fc_myDID);
+ return err;
}
/**
@@ -6269,7 +6290,9 @@ lpfc_issue_els_edc_rsp(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
*
* Return code
* 0 - Successfully issued acc adisc response
- * 1 - Failed to issue adisc acc response
+ * -ENOMEM - IOCB not prepped successfully
+ * -EIO - The IOCB failed to issue successfully
+ * -ENODEV - No associated node for IOCB
**/
int
lpfc_els_rsp_adisc_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
@@ -6284,12 +6307,15 @@ lpfc_els_rsp_adisc_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
uint16_t cmdsize;
int rc;
u32 ulp_context;
+ int err;
cmdsize = sizeof(uint32_t) + sizeof(ADISC);
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry, ndlp,
ndlp->nlp_DID, ELS_CMD_ACC);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -6336,17 +6362,26 @@ lpfc_els_rsp_adisc_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
elsiocb->ndlp = lpfc_nlp_get(ndlp);
if (!elsiocb->ndlp) {
lpfc_els_free_iocb(phba, elsiocb);
- return 1;
+ err = -ENODEV;
+ goto err_out;
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
- return 1;
+ err = -EIO;
+ goto err_out;
}
return 0;
+
+err_out:
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "1025 Xmit ADISC ACC Unsuccessful: "
+ "error_code: %d S_ID: x%x\n",
+ err, vport->fc_myDID);
+ return err;
}
/**
@@ -6366,7 +6401,10 @@ lpfc_els_rsp_adisc_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
*
* Return code
* 0 - Successfully issued acc prli response
- * 1 - Failed to issue acc prli response
+ * -ENOMEM - IOCB not prepped successfully
+ * -EIO - The IOCB failed to issue successfully
+ * -ENODEV - No associated node for IOCB
+ * -EACCES - Acc not needed for this command
**/
int
lpfc_els_rsp_prli_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
@@ -6386,6 +6424,7 @@ lpfc_els_rsp_prli_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
struct lpfc_dmabuf *req_buf;
int rc;
u32 elsrspcmd, ulp_context;
+ int err;
/* Need the incoming PRLI payload to determine if the ACC is for an
* FC4 or NVME PRLI type. The PRLI type is at word 1.
@@ -6407,13 +6446,16 @@ lpfc_els_rsp_prli_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
cmdsize = sizeof(uint32_t) + sizeof(struct lpfc_nvme_prli);
elsrspcmd = (ELS_CMD_ACC | (ELS_CMD_NVMEPRLI & ~ELS_RSP_MASK));
} else {
- return 1;
+ err = -EACCES;
+ goto err_out;
}
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry, ndlp,
ndlp->nlp_DID, elsrspcmd);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -6528,17 +6570,28 @@ lpfc_els_rsp_prli_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
elsiocb->ndlp = lpfc_nlp_get(ndlp);
if (!elsiocb->ndlp) {
lpfc_els_free_iocb(phba, elsiocb);
- return 1;
+ err = -ENODEV;
+ goto err_out;
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
- return 1;
+ err = -EIO;
+ goto err_out;
}
return 0;
+
+err_out:
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "1026 Xmit PRLI ACC Unsuccessful: "
+ "cmd: x%x, error_code: %d "
+ "S_ID: x%x\n",
+ *(uint32_t *)oldiocb->cmd_dmabuf->virt, err,
+ vport->fc_myDID);
+ return err;
}
/**
@@ -6559,7 +6612,9 @@ lpfc_els_rsp_prli_acc(struct lpfc_vport *vport, struct lpfc_iocbq *oldiocb,
*
* Return code
* 0 - Successfully issued acc rnid response
- * 1 - Failed to issue acc rnid response
+ * -ENOMEM - IOCB not prepped successfully
+ * -EIO - The IOCB failed to issue successfully
+ * -ENODEV - No associated node for IOCB
**/
static int
lpfc_els_rsp_rnid_acc(struct lpfc_vport *vport, uint8_t format,
@@ -6574,6 +6629,7 @@ lpfc_els_rsp_rnid_acc(struct lpfc_vport *vport, uint8_t format,
uint16_t cmdsize;
int rc;
u32 ulp_context;
+ int err;
cmdsize = sizeof(uint32_t) + sizeof(uint32_t)
+ (2 * sizeof(struct lpfc_name));
@@ -6582,8 +6638,10 @@ lpfc_els_rsp_rnid_acc(struct lpfc_vport *vport, uint8_t format,
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry, ndlp,
ndlp->nlp_DID, ELS_CMD_ACC);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -6642,17 +6700,26 @@ lpfc_els_rsp_rnid_acc(struct lpfc_vport *vport, uint8_t format,
elsiocb->ndlp = lpfc_nlp_get(ndlp);
if (!elsiocb->ndlp) {
lpfc_els_free_iocb(phba, elsiocb);
- return 1;
+ err = -ENODEV;
+ goto err_out;
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
- return 1;
+ err = -EIO;
+ goto err_out;
}
return 0;
+
+err_out:
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "1028 Xmit RNID ACC Unsuccessful: "
+ "error_code: %d S_ID: x%x\n",
+ err, vport->fc_myDID);
+ return err;
}
/**
@@ -6712,7 +6779,9 @@ lpfc_els_clear_rrq(struct lpfc_vport *vport,
*
* Return code
* 0 - Successfully issued acc echo response
- * 1 - Failed to issue acc echo response
+ * -ENOMEM - IOCB not prepped successfully
+ * -EIO - The IOCB failed to issue successfully
+ * -ENODEV - No associated node for IOCB
**/
static int
lpfc_els_rsp_echo_acc(struct lpfc_vport *vport, uint8_t *data,
@@ -6726,6 +6795,7 @@ lpfc_els_rsp_echo_acc(struct lpfc_vport *vport, uint8_t *data,
uint16_t cmdsize;
int rc;
u32 ulp_context;
+ int err;
if (phba->sli_rev == LPFC_SLI_REV4)
cmdsize = oldiocb->wcqe_cmpl.total_data_placed;
@@ -6739,8 +6809,10 @@ lpfc_els_rsp_echo_acc(struct lpfc_vport *vport, uint8_t *data,
cmdsize = LPFC_BPL_SIZE;
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry, ndlp,
ndlp->nlp_DID, ELS_CMD_ACC);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
if (phba->sli_rev == LPFC_SLI_REV4) {
wqe = &elsiocb->wqe;
@@ -6776,17 +6848,26 @@ lpfc_els_rsp_echo_acc(struct lpfc_vport *vport, uint8_t *data,
elsiocb->ndlp = lpfc_nlp_get(ndlp);
if (!elsiocb->ndlp) {
lpfc_els_free_iocb(phba, elsiocb);
- return 1;
+ err = -ENODEV;
+ goto err_out;
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
- return 1;
+ err = -EIO;
+ goto err_out;
}
return 0;
+
+err_out:
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "1029 Xmit ECHO ACC Unsuccessful: "
+ "error_code: %d S_ID: x%x\n",
+ err, vport->fc_myDID);
+ return err;
}
/**
@@ -8489,14 +8570,14 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
vport->fc_rscn_id_list[vport->fc_rscn_id_cnt++] = pcmd;
/* Indicate we are done walking fc_rscn_id_list on this vport */
vport->fc_rscn_flush = 0;
+ /* Send back ACC */
+ lpfc_els_rsp_acc(vport, ELS_CMD_ACC, cmdiocb, ndlp, NULL);
/*
* If we zero, cmdiocb->cmd_dmabuf, the calling routine will
* not try to free it.
*/
cmdiocb->cmd_dmabuf = NULL;
lpfc_set_disctmo(vport);
- /* Send back ACC */
- lpfc_els_rsp_acc(vport, ELS_CMD_ACC, cmdiocb, ndlp, NULL);
/* send RECOVERY event for ALL nodes that match RSCN payload */
lpfc_rscn_recovery_check(vport);
return lpfc_els_handle_rscn(vport);
@@ -9296,7 +9377,9 @@ lpfc_send_rrq(struct lpfc_hba *phba, struct lpfc_node_rrq *rrq)
*
* Return code
* 0 - Successfully issued ACC RPL ELS command
- * 1 - Failed to issue ACC RPL ELS command
+ * -ENOMEM - IOCB not prepped successfully
+ * -EIO - The IOCB failed to issue successfully
+ * -ENODEV - No associated node for IOCB
**/
static int
lpfc_els_rsp_rpl_acc(struct lpfc_vport *vport, uint16_t cmdsize,
@@ -9310,12 +9393,15 @@ lpfc_els_rsp_rpl_acc(struct lpfc_vport *vport, uint16_t cmdsize,
struct lpfc_iocbq *elsiocb;
uint8_t *pcmd;
u32 ulp_context;
+ int err;
elsiocb = lpfc_prep_els_iocb(vport, 0, cmdsize, oldiocb->retry, ndlp,
ndlp->nlp_DID, ELS_CMD_ACC);
- if (!elsiocb)
- return 1;
+ if (!elsiocb) {
+ err = -ENOMEM;
+ goto err_out;
+ }
ulp_context = get_job_ulpcontext(phba, elsiocb);
if (phba->sli_rev == LPFC_SLI_REV4) {
@@ -9358,17 +9444,26 @@ lpfc_els_rsp_rpl_acc(struct lpfc_vport *vport, uint16_t cmdsize,
elsiocb->ndlp = lpfc_nlp_get(ndlp);
if (!elsiocb->ndlp) {
lpfc_els_free_iocb(phba, elsiocb);
- return 1;
+ err = -ENODEV;
+ goto err_out;
}
rc = lpfc_sli_issue_iocb(phba, LPFC_ELS_RING, elsiocb, 0);
if (lpfc_iocb_failed(rc)) {
lpfc_els_free_iocb(phba, elsiocb);
lpfc_nlp_put(ndlp);
- return 1;
+ err = -EIO;
+ goto err_out;
}
return 0;
+
+err_out:
+ lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
+ "1030 Xmit ELS RPL ACC Unsuccessful: "
+ "error_code: %d S_ID: x%x\n",
+ err, vport->fc_myDID);
+ return err;
}
/**
@@ -9603,7 +9698,7 @@ lpfc_els_rcv_fan(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
* @ndlp: pointer to a node-list data structure.
*
* Return code
- * 0 - Successfully processed echo iocb (currently always return 0)
+ * 0 - Successfully processed edc iocb (currently always return 0)
**/
static int
lpfc_els_rcv_edc(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 13/14] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (11 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 14/14] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
2026-09-19 7:45 ` [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
The lpfc_set_disctmo routine is not used for all cases when the driver
needs to restart discovery on the fc_disctmo timer. Not doing so, makes
discovery timer actions invisible in some cases as they do not get logged.
This patch substitutes calls on fc_disctmo to use lpfc_set_disctmo in
lpfc_els_rcv_rscn.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_els.c | 17 ++++++-----------
1 file changed, 6 insertions(+), 11 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
index cb06dbc9edfa..0f2a8c71d8db 100644
--- a/drivers/scsi/lpfc/lpfc_els.c
+++ b/drivers/scsi/lpfc/lpfc_els.c
@@ -8399,7 +8399,7 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
uint32_t payload_len, length, nportid, *cmd;
int rscn_cnt;
int rscn_id = 0, hba_id = 0;
- int i, tmo;
+ int i;
pcmd = cmdiocb->cmd_dmabuf;
lp = (uint32_t *) pcmd->virt;
@@ -8478,11 +8478,8 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
lpfc_els_rsp_acc(vport, ELS_CMD_ACC, cmdiocb,
ndlp, NULL);
/* Restart disctmo if its already running */
- if (test_bit(FC_DISC_TMO, &vport->fc_flag)) {
- tmo = ((phba->fc_ratov * 3) + 3);
- mod_timer(&vport->fc_disctmo,
- jiffies + secs_to_jiffies(tmo));
- }
+ if (test_bit(FC_DISC_TMO, &vport->fc_flag))
+ lpfc_set_disctmo(vport);
return 0;
}
}
@@ -8513,11 +8510,9 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
set_bit(FC_RSCN_DEFERRED, &vport->fc_flag);
/* Restart disctmo if its already running */
- if (test_bit(FC_DISC_TMO, &vport->fc_flag)) {
- tmo = ((phba->fc_ratov * 3) + 3);
- mod_timer(&vport->fc_disctmo,
- jiffies + secs_to_jiffies(tmo));
- }
+ if (test_bit(FC_DISC_TMO, &vport->fc_flag))
+ lpfc_set_disctmo(vport);
+
if ((rscn_cnt < FC_MAX_HOLD_RSCN) &&
!test_bit(FC_RSCN_DISCOVERY, &vport->fc_flag)) {
set_bit(FC_RSCN_MODE, &vport->fc_flag);
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [PATCH v4 14/14] lpfc: Update lpfc version to 15.0.0.1
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (12 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 13/14] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
@ 2026-09-17 22:20 ` Nigel Kirkland
2026-09-19 7:45 ` [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
14 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-17 22:20 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, nkirkland2304
Update lpfc version to 15.0.0.1
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_version.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/scsi/lpfc/lpfc_version.h b/drivers/scsi/lpfc/lpfc_version.h
index d6e6e436fbfc..7df63d118234 100644
--- a/drivers/scsi/lpfc/lpfc_version.h
+++ b/drivers/scsi/lpfc/lpfc_version.h
@@ -20,7 +20,7 @@
* included with this package. *
*******************************************************************/
-#define LPFC_DRIVER_VERSION "15.0.0.0"
+#define LPFC_DRIVER_VERSION "15.0.0.1"
#define LPFC_DRIVER_NAME "lpfc"
/* Used for SLI 2/3 */
--
2.38.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* Re: [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable
2026-09-17 22:20 ` [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable Nigel Kirkland
@ 2026-09-17 22:20 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-09-17 22:20 UTC (permalink / raw)
To: Nigel Kirkland; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] scsi: lpfc: txq list churning in lpfc_drain_txq due to missing SLI_IOCB_RET_IOCB flag
--
commit 8b96ef178c664b569b564d0486b1112a8fe88e10
Author: Nigel Kirkland <nkirkland2304@gmail.com>
lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable
This commit adds logic to queue ELS/CT commands to the head or tail of the
txq for retry when the ELS WQ is full or SGLs are unavailable. It updates
lpfc_drain_txq to handle these retry scenarios.
> diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
> --- a/drivers/scsi/lpfc/lpfc_sli.c
> +++ b/drivers/scsi/lpfc/lpfc_sli.c
[ ... ]
> @@ -10819,12 +10866,19 @@ __lpfc_sli_issue_iocb_s4(struct lpfc_hba *phba, uint32_t ring_number,
> if (!sglq) {
> if (!(flag & SLI_IOCB_RET_IOCB)) {
> __lpfc_sli_ringtx_put(phba,
> pring,
> - piocb);
> + piocb, false);
> return IOCB_SUCCESS;
> } else {
> return IOCB_BUSY;
> }
[Severity: Medium]
If lpfc_drain_txq() passes 0 for the flag, won't it trigger this block?
When SLI_IOCB_RET_IOCB is omitted, this path places the IOCB at the tail of
the txq and returns IOCB_SUCCESS. This defeats the newly added back-off
mechanism in lpfc_drain_txq() which relies on receiving an IOCB_BUSY status
to execute its retry logic and push the failing IOCB to the head.
[ ... ]
> @@ -21301,12 +21355,15 @@ lpfc_drain_txq(struct lpfc_hba *phba)
> }
> txq_cnt--;
>
> + /* Capture the return values that indicate an error during IO
> + * submit and cannot be retried. Prefix message 2822.
> + */
> ret = __lpfc_sli_issue_iocb(phba, pring->ringno, piocbq, 0);
[Severity: Medium]
Does this missing flag cause txq list churning during resource exhaustion?
Because the flag is passed as 0 here (instead of SLI_IOCB_RET_IOCB), the
underlying __lpfc_sli_issue_iocb() call will not return IOCB_BUSY when SGL
resources are exhausted. Since IOCB_SUCCESS is returned instead, the loop
below won't break early and will continuously pop and push items back to the
tail of the txq.
> -
> - if (ret && ret != IOCB_BUSY) {
> + if (ret && ret != IOCB_BUSY && ret != IOCB_FAILED_PUT) {
> fail_msg = " - Cannot send IO ";
> piocbq->cmd_flag &= ~LPFC_DRIVER_ABORTED;
> }
> +
> if (fail_msg) {
> piocbq->cmd_flag |= LPFC_DRIVER_ABORTED;
> /* Failed means we can't issue and need to cancel */
> @@ -21319,9 +21376,35 @@ lpfc_drain_txq(struct lpfc_hba *phba)
> list_add_tail(&piocbq->list, &completions);
> fail_msg = NULL;
> }
> - spin_unlock_irqrestore(&pring->ring_lock, iflags);
> - if (txq_cnt == 0 || ret == IOCB_BUSY)
> +
> + if (txq_cnt == 0 || ret == IOCB_BUSY ||
> + ret == IOCB_FAILED_PUT) {
> + /* IOCB_FAILED_PUT is unique to SLI4 and means SGL/XRI
> + * resources are allocated. For this case, set the
> + * in retry flag. For SLI3 and 4, push the IO to the
> + * txq for retry.
> + */
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917222015.61053-1-nkirkland2304@gmail.com?part=11
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler
2026-09-17 22:20 ` [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler Nigel Kirkland
@ 2026-09-17 22:20 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-09-17 22:20 UTC (permalink / raw)
To: Nigel Kirkland; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] scsi: lpfc: CPU soft lockup in lpfc_sli_handle_slow_ring_event_s4()
--
commit 2c325e87e41c5052e2704f826b9f808f153d5397
Author: Nigel Kirkland <nkirkland2304@gmail.com>
lpfc: Remove slowpath cqe process limiter in slow ring event handler
This commit removes the 64 CQE limit in the slow path event handler
to ensure the slow path CQ is drained completely. It also adds logging
when the CQE count reaches a full ELS WQ wrap.
> diff --git a/drivers/scsi/lpfc/lpfc_sli.c b/drivers/scsi/lpfc/lpfc_sli.c
> index 10d9030a1e883..06dc01c9539ed 100644
> --- a/drivers/scsi/lpfc/lpfc_sli.c
> +++ b/drivers/scsi/lpfc/lpfc_sli.c
[ ... ]
> @@ -4562,22 +4562,39 @@ lpfc_sli_handle_slow_ring_event_s4(struct lpfc_hba *phba,
> if (irspiocbq)
> lpfc_sli_sp_handle_rspiocb(phba, pring,
> irspiocbq);
> - count++;
> break;
> case CQE_CODE_RECEIVE:
> case CQE_CODE_RECEIVE_V1:
> dmabuf = container_of(cq_event, struct hbq_dmabuf,
> cq_event);
> lpfc_sli4_handle_received_buffer(phba, dmabuf);
> - count++;
> break;
> default:
> + lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
> + "7771 Unknown WCQE completion code "
> + "x%x, ignoring.\n",
> + bf_get(lpfc_wcqe_c_code,
> + &cq_event->cqe.wcqe_cmpl));
> break;
> }
>
> - /* Limit the number of events to 64 to avoid soft lockups */
> - if (count == 64)
> - break;
[Severity: High]
Could removing this limit cause a CPU soft lockup under certain failure
conditions?
When an MDS loopback or unsolicited data frame is received (CQE_CODE_RECEIVE),
and the driver's IOCBQ pool is momentarily empty,
lpfc_sli4_handle_mds_loopback() fails to allocate an iocbq via
lpfc_sli_get_iocbq() and adds the event back to the sp_queue_event list:
lpfc_sli4_handle_mds_loopback() {
...
list_add_tail(&dmabuf->cq_event.list,
&phba->sli4_hba.sp_queue_event);
...
}
Since this patch removes the 64-iteration limit, this while loop no
longer yields. It will immediately dequeue the same event and retry the
allocation in a tight spin loop. Without the loop break, the worker
thread can no longer reach wait_event_interruptible() to yield the CPU,
resulting in a 100% CPU lockup.
Is there another mechanism to prevent the loop from spinning infinitely
when events are requeued?
> + /* This loop runs until the ELS/CT CQ is empty. Post a one
> + * time message for debug support when ELS WQ ecount
> + * completions are processed - this represent 1 full ELS WQ
> + * wrap.
> + */
> + if (++count == LPFC_WQE_DEF_COUNT) {
> + lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
> + "7772 %s SP CQE count %d\n",
> + __func__, count);
> + }
> + }
> +
> + /* Log a final message to note how many CQEs were processed. */
> + if (count > LPFC_WQE_DEF_COUNT) {
> + lpfc_printf_log(phba, KERN_INFO, LOG_ELS,
> + "7773 %s SP CQEs complete, count %d\n",
> + __func__, count);
> }
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917222015.61053-1-nkirkland2304@gmail.com?part=10
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting
2026-09-17 22:20 ` [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting Nigel Kirkland
@ 2026-09-17 22:22 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-09-17 22:22 UTC (permalink / raw)
To: Nigel Kirkland; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] scsi: lpfc: UAF in lpfc_els_rsp_acc() error path due to concurrent RSCN flush
--
commit 47dddef465d8869f241a192661f8107a7fb86c19
Author: Nigel Kirkland <nkirkland2304@gmail.com>
lpfc: Update ELS ACC logging for diagnostic troubleshooting
Currently, there are ELS ACC routines that lack debug log messages to
indicate when ACC frame transmission run into issues. The generic
lpfc_els_rsp_acc and more specific ACC routines are updated to log when
there is an issue with transmitting the frame. The routines are also
updated to return different return codes when encountering various
transmission issues and their function comment header is updated.
> diff --git a/drivers/scsi/lpfc/lpfc_els.c b/drivers/scsi/lpfc/lpfc_els.c
> index bf71b5a3e55e5..cb06dbc9edfab 100644
> --- a/drivers/scsi/lpfc/lpfc_els.c
> +++ b/drivers/scsi/lpfc/lpfc_els.c
[ ... ]
> @@ -6055,6 +6065,17 @@ lpfc_els_rsp_acc(struct lpfc_vport *vport, uint32_t flag,
> ndlp->nlp_DID, ndlp->nlp_flag, ndlp->nlp_state,
> ndlp->nlp_rpi, vport->fc_flag, kref_read(&ndlp->kref));
> return 0;
> +
> +err_out:
> + if (oldiocb->cmd_dmabuf && oldiocb->cmd_dmabuf->virt)
> + old_opcode = *(uint32_t *)oldiocb->cmd_dmabuf->virt;
[Severity: High]
Can this dereference a freed pointer?
In lpfc_els_rcv_rscn(), setting vport->fc_rscn_flush = 0 publishes the buffer
for concurrent processing. If a parallel thread processes and frees this buffer,
and lpfc_els_rsp_acc() takes this new error path, oldiocb->cmd_dmabuf->virt
would point to freed memory.
> +
> + lpfc_printf_vlog(vport, KERN_INFO, LOG_ELS,
> + "1027 Xmit ELS ACC Unsuccessful: "
> + "cmd: x%x, error_code: %d "
> + "S_ID: x%x\n", old_opcode, err,
> + vport->fc_myDID);
> + return err;
> }
[ ... ]
> @@ -8489,14 +8570,14 @@ lpfc_els_rcv_rscn(struct lpfc_vport *vport, struct lpfc_iocbq *cmdiocb,
> vport->fc_rscn_id_list[vport->fc_rscn_id_cnt++] = pcmd;
> /* Indicate we are done walking fc_rscn_id_list on this vport */
> vport->fc_rscn_flush = 0;
> + /* Send back ACC */
> + lpfc_els_rsp_acc(vport, ELS_CMD_ACC, cmdiocb, ndlp, NULL);
[Severity: High]
Does moving this lpfc_els_rsp_acc() call introduce a use-after-free risk?
By calling lpfc_els_rsp_acc() here, cmdiocb->cmd_dmabuf has been published by
clearing vport->fc_rscn_flush, but is not yet set to NULL. A concurrent thread
could free cmdiocb->cmd_dmabuf before the error path in lpfc_els_rsp_acc()
attempts to read from it.
> /*
> * If we zero, cmdiocb->cmd_dmabuf, the calling routine will
> * not try to free it.
> */
> cmdiocb->cmd_dmabuf = NULL;
> lpfc_set_disctmo(vport);
> - /* Send back ACC */
> - lpfc_els_rsp_acc(vport, ELS_CMD_ACC, cmdiocb, ndlp, NULL);
> /* send RECOVERY event for ALL nodes that match RSCN payload */
> lpfc_rscn_recovery_check(vport);
> return lpfc_els_handle_rscn(vport);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917222015.61053-1-nkirkland2304@gmail.com?part=12
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v4 00/14] Update lpfc to revision 15.0.0.1
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
` (13 preceding siblings ...)
2026-09-17 22:20 ` [PATCH v4 14/14] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
@ 2026-09-19 7:45 ` Nigel Kirkland
2026-09-28 17:13 ` Nigel Kirkland
14 siblings, 1 reply; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-19 7:45 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely, Nigel Kirkland
On Thu, Sep 17, 2026 at 2:57 PM Nigel Kirkland <nkirkland2304@gmail.com> wrote:
>
> Update lpfc to revision 15.0.0.1
>
> This patch set contains bug fixes related to cleanup handling in both
> normal and error paths, discovery rework for large SAN configurations, and
> refactoring of duplicate code.
>
> The patches were cut against Martin's 7.4/scsi-queue tree.
>
> Nigel Kirkland (14):
> lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
> lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set
> lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd
> lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
> lpfc: Add handling for when PLOGI or PRLI is dropped during link failure
> lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
> lpfc: Rework I/O flush ordering when unloading driver
> lpfc: Improve PLOGI retry handling for large SAN configurations
> lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received
> lpfc: Remove slowpath cqe process limiter in slow ring event handler
> lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable
> lpfc: Update ELS ACC logging for diagnostic troubleshooting
> lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
> lpfc: Update lpfc version to 15.0.0.1
>
Preliminary response regarding Sashiko findings.
Patch #6: Sashiko pointed out 2 [high] issues that appear valid in
theory, but appear
highly improbably in the real world. We are reviewing these findings
and whether to
address in a v5 drop or a future patch set.
Patch #8: Sashiko pointed out 1 [medium] and 1 [high] issue. We are
evaluating if
these are a real concern.
Patch #9: Sashiko pointed out 2 [high] issues. We are evaluating if
these are a real
concern.
Patch #10: Sashiko pointed out 1 [high] issue. Whilst valid in theory,
the patch eliminates
a real world "silent-drop/delayed-reschedule" bug. Removal of the
maximum processing
count does introduce a theoretical concern in the event of a sustained
flood of slow
path CQEs. However, the original issue that prompted addition of the
processing limit was
related to loopback testing. Loopback processing was subsequently
moved away from the
slow path, eliminating that as a producer. Removal of the processing
limit was overlooked at
that time.
Patch #11: Sashiko pointed out 1 [medium] issue. We are reviewing
these findings to
determine whether this should be addressed in a v5 patch set.
Patch #12: Sashiko pointed out 1 [high] theoretical concern, but in
our analysis the race is
unlikely. If seen, the worst case outcome appears to be reporting of
an incorrect cmd code
value in log message 1027. Closing this theoretical window is
something we plan to address
in a future patch set.
In summary: Analysis of the findings is ongoing with a complete
response anticipated
in a few days.
> drivers/scsi/lpfc/lpfc_bsg.c | 7 +-
> drivers/scsi/lpfc/lpfc_crtn.h | 12 +-
> drivers/scsi/lpfc/lpfc_ct.c | 23 +-
> drivers/scsi/lpfc/lpfc_disc.h | 5 +-
> drivers/scsi/lpfc/lpfc_els.c | 503 ++++++++++++++++++++++-------
> drivers/scsi/lpfc/lpfc_hbadisc.c | 130 ++++----
> drivers/scsi/lpfc/lpfc_init.c | 21 +-
> drivers/scsi/lpfc/lpfc_nportdisc.c | 100 +++++-
> drivers/scsi/lpfc/lpfc_nvme.c | 2 +-
> drivers/scsi/lpfc/lpfc_scsi.c | 2 +-
> drivers/scsi/lpfc/lpfc_sli.c | 227 ++++++++++---
> drivers/scsi/lpfc/lpfc_sli.h | 18 ++
> drivers/scsi/lpfc/lpfc_version.h | 2 +-
> 13 files changed, 812 insertions(+), 240 deletions(-)
>
> --
> 2.38.0
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v4 00/14] Update lpfc to revision 15.0.0.1
2026-09-19 7:45 ` [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
@ 2026-09-28 17:13 ` Nigel Kirkland
0 siblings, 0 replies; 23+ messages in thread
From: Nigel Kirkland @ 2026-09-28 17:13 UTC (permalink / raw)
To: linux-scsi, nigel.kirkland; +Cc: paul.ely
On Sat, Sep 19, 2026 at 12:45 AM Nigel Kirkland <nkirkland2304@gmail.com> wrote:
>
> On Thu, Sep 17, 2026 at 2:57 PM Nigel Kirkland <nkirkland2304@gmail.com> wrote:
> >
> > Update lpfc to revision 15.0.0.1
> >
> > This patch set contains bug fixes related to cleanup handling in both
> > normal and error paths, discovery rework for large SAN configurations, and
> > refactoring of duplicate code.
> >
> > The patches were cut against Martin's 7.4/scsi-queue tree.
> >
> > Nigel Kirkland (14):
> > lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid
> > lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set
> > lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd
> > lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error
> > lpfc: Add handling for when PLOGI or PRLI is dropped during link failure
> > lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence
> > lpfc: Rework I/O flush ordering when unloading driver
> > lpfc: Improve PLOGI retry handling for large SAN configurations
> > lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received
> > lpfc: Remove slowpath cqe process limiter in slow ring event handler
> > lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable
> > lpfc: Update ELS ACC logging for diagnostic troubleshooting
> > lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler
> > lpfc: Update lpfc version to 15.0.0.1
> >
>
> Preliminary response regarding Sashiko findings.
>
> Patch #6: Sashiko pointed out 2 [high] issues that appear valid in
> theory, but appear
> highly improbably in the real world. We are reviewing these findings
> and whether to
> address in a v5 drop or a future patch set.
>
> Patch #8: Sashiko pointed out 1 [medium] and 1 [high] issue. We are
> evaluating if
> these are a real concern.
>
> Patch #9: Sashiko pointed out 2 [high] issues. We are evaluating if
> these are a real
> concern.
>
> Patch #10: Sashiko pointed out 1 [high] issue. Whilst valid in theory,
> the patch eliminates
> a real world "silent-drop/delayed-reschedule" bug. Removal of the
> maximum processing
> count does introduce a theoretical concern in the event of a sustained
> flood of slow
> path CQEs. However, the original issue that prompted addition of the
> processing limit was
> related to loopback testing. Loopback processing was subsequently
> moved away from the
> slow path, eliminating that as a producer. Removal of the processing
> limit was overlooked at
> that time.
>
> Patch #11: Sashiko pointed out 1 [medium] issue. We are reviewing
> these findings to
> determine whether this should be addressed in a v5 patch set.
>
> Patch #12: Sashiko pointed out 1 [high] theoretical concern, but in
> our analysis the race is
> unlikely. If seen, the worst case outcome appears to be reporting of
> an incorrect cmd code
> value in log message 1027. Closing this theoretical window is
> something we plan to address
> in a future patch set.
>
> In summary: Analysis of the findings is ongoing with a complete
> response anticipated
> in a few days.
>
Out of an abundance of caution, we will be submitting a v5 patch set.
>
> > drivers/scsi/lpfc/lpfc_bsg.c | 7 +-
> > drivers/scsi/lpfc/lpfc_crtn.h | 12 +-
> > drivers/scsi/lpfc/lpfc_ct.c | 23 +-
> > drivers/scsi/lpfc/lpfc_disc.h | 5 +-
> > drivers/scsi/lpfc/lpfc_els.c | 503 ++++++++++++++++++++++-------
> > drivers/scsi/lpfc/lpfc_hbadisc.c | 130 ++++----
> > drivers/scsi/lpfc/lpfc_init.c | 21 +-
> > drivers/scsi/lpfc/lpfc_nportdisc.c | 100 +++++-
> > drivers/scsi/lpfc/lpfc_nvme.c | 2 +-
> > drivers/scsi/lpfc/lpfc_scsi.c | 2 +-
> > drivers/scsi/lpfc/lpfc_sli.c | 227 ++++++++++---
> > drivers/scsi/lpfc/lpfc_sli.h | 18 ++
> > drivers/scsi/lpfc/lpfc_version.h | 2 +-
> > 13 files changed, 812 insertions(+), 240 deletions(-)
> >
> > --
> > 2.38.0
> >
^ permalink raw reply [flat|nested] 23+ messages in thread
end of thread, other threads:[~2026-09-28 17:13 UTC | newest]
Thread overview: 23+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 01/14] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 02/14] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 03/14] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 04/14] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 05/14] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
2026-09-17 22:10 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 07/14] lpfc: Rework I/O flush ordering when unloading driver Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations Nigel Kirkland
2026-09-17 22:12 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received Nigel Kirkland
2026-09-17 22:15 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler Nigel Kirkland
2026-09-17 22:20 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable Nigel Kirkland
2026-09-17 22:20 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting Nigel Kirkland
2026-09-17 22:22 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 13/14] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 14/14] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
2026-09-19 7:45 ` [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-28 17:13 ` Nigel Kirkland
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox