From: Nigel Kirkland <nkirkland2304@gmail.com>
To: linux-scsi@vger.kernel.org, nigel.kirkland@broadcom.com
Cc: paul.ely@broadcom.com, nkirkland2304@gmail.com
Subject: [PATCH v4 07/14] lpfc: Rework I/O flush ordering when unloading driver
Date: Thu, 17 Sep 2026 15:20:08 -0700 [thread overview]
Message-ID: <20260917222015.61053-8-nkirkland2304@gmail.com> (raw)
In-Reply-To: <20260917222015.61053-1-nkirkland2304@gmail.com>
The lpfc_els_abort routine has a code path that cancels outstanding
I/Os on the ELS ring when attempted aborts fail. The failed aborts are
queued to a drv_cmpl_list and then cancelled after the ELS pring->txcmplq
is fully traversed. However if the abort failure returns IOCB_ABORTING,
then the driver should not have cancelled it. Doing so starts two threads
working on the same iocb and ndlp, leading to unintended race conditions.
Fix by capturing the IOCB_ABORTING return value in lpfc_els_abort and not
adding it to the list of iocbs for cancelling. We should allow the iocb
scheduled for abort to complete naturally. This avoids simultaneous
threads acting on the same iocb and ndlp objects.
The lpfc_free_iocb_list is moved to execute after lpfc_sli4_hba_unset
allowing the routine to flush I/O before freeing it. And, in
lpfc_pci_remove_one_s4 a call to flush the phba->wq is added. This makes
the unload logic consistent with offline handling logic.
Signed-off-by: Nigel Kirkland <nkirkland2304@gmail.com>
---
drivers/scsi/lpfc/lpfc_init.c | 16 ++++++++++++++--
drivers/scsi/lpfc/lpfc_nportdisc.c | 11 +++++++++--
2 files changed, 23 insertions(+), 4 deletions(-)
diff --git a/drivers/scsi/lpfc/lpfc_init.c b/drivers/scsi/lpfc/lpfc_init.c
index 6460127bcc7b..7352cb6e584b 100644
--- a/drivers/scsi/lpfc/lpfc_init.c
+++ b/drivers/scsi/lpfc/lpfc_init.c
@@ -13514,6 +13514,9 @@ lpfc_sli4_hba_unset(struct lpfc_hba *phba)
/* Stop the SLI4 device port */
if (phba->pport)
phba->pport->work_port_events = 0;
+
+ /* All IO completed and queues released. Free the IOCBs. */
+ lpfc_free_iocb_list(phba);
}
/*
@@ -14948,11 +14951,20 @@ lpfc_pci_remove_one_s4(struct pci_dev *pdev)
/* Perform scsi free before driver resource_unset since scsi
* buffers are released to their corresponding pools here.
+ * lpfc_sli4_hba_unset() issues aborts via lpfc_sli_hba_iocb_abort(),
+ * which allocates abort IOCBs from phba->lpfc_iocb_list; the pool
+ * must still exist, so lpfc_free_iocb_list() runs only after unset.
*/
lpfc_io_free(phba);
- lpfc_free_iocb_list(phba);
- lpfc_sli4_hba_unset(phba);
+ /* Flush the PHBA WQ - there could be a race with ELS IOs while lpfc
+ * is unloading. This stops a race between completions, aborts and
+ * resource recovery.
+ */
+ if (phba->wq)
+ flush_workqueue(phba->wq);
+
+ lpfc_sli4_hba_unset(phba);
lpfc_unset_driver_resource_phase2(phba);
lpfc_sli4_driver_resource_unset(phba);
diff --git a/drivers/scsi/lpfc/lpfc_nportdisc.c b/drivers/scsi/lpfc/lpfc_nportdisc.c
index 2c8d995a45bf..f917a5bcfd02 100644
--- a/drivers/scsi/lpfc/lpfc_nportdisc.c
+++ b/drivers/scsi/lpfc/lpfc_nportdisc.c
@@ -255,8 +255,9 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
spin_lock_irq(&phba->hbalock);
if (phba->sli_rev == LPFC_SLI_REV4)
spin_lock(&pring->ring_lock);
+
list_for_each_entry_safe(iocb, next_iocb, &pring->txcmplq, list) {
- /* Add to abort_list on on NDLP match. */
+ /* Add to abort_list on NDLP match. */
if (lpfc_check_sli_ndlp(phba, pring, iocb, ndlp))
list_add_tail(&iocb->dlist, &abort_list);
}
@@ -271,7 +272,13 @@ lpfc_els_abort(struct lpfc_hba *phba, struct lpfc_nodelist *ndlp)
retval = lpfc_sli_issue_abort_iotag(phba, pring, iocb, NULL);
spin_unlock_irq(&phba->hbalock);
- if (retval && test_bit(FC_UNLOADING, &phba->pport->load_flag)) {
+ /* An abort that fails here is just cancelled when the driver is
+ * going offline. However, if the abort failure is because the
+ * IOCB is already getting aborted, don't cancel. Just let it
+ * complete.
+ */
+ if (test_bit(FC_UNLOADING, &phba->pport->load_flag) &&
+ retval && retval != IOCB_ABORTING) {
list_del_init(&iocb->list);
list_add_tail(&iocb->list, &drv_cmpl_list);
}
--
2.38.0
next prev parent reply other threads:[~2026-09-17 21:57 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 22:20 [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 01/14] lpfc: Fix use-after-free in lpfc_cmpl_ct_cmd_vmid Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 02/14] lpfc: Early return out of lpfc_els_abort when HBA_SETUP flag is not set Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 03/14] lpfc: Fix kernel oops when unmapping scsi dma buffers for an aborted cmd Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 04/14] lpfc: Check fc4_xpt_flags before decrementing ndlp kref on FDISC error Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 05/14] lpfc: Add handling for when PLOGI or PRLI is dropped during link failure Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 06/14] lpfc: Fix ndlp use-after-free during repeated RSCN and rediscovery sequence Nigel Kirkland
2026-09-17 22:10 ` sashiko-bot
2026-09-17 22:20 ` Nigel Kirkland [this message]
2026-09-17 22:20 ` [PATCH v4 08/14] lpfc: Improve PLOGI retry handling for large SAN configurations Nigel Kirkland
2026-09-17 22:12 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 09/14] lpfc: Send inhibited ABORT_WQE when PLOGI CQE SEQUENCE_TMO is received Nigel Kirkland
2026-09-17 22:15 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 10/14] lpfc: Remove slowpath cqe process limiter in slow ring event handler Nigel Kirkland
2026-09-17 22:20 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 11/14] lpfc: Put iocbq on phba->txq when ELS WQ is full or ELS SGL unavailable Nigel Kirkland
2026-09-17 22:20 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 12/14] lpfc: Update ELS ACC logging for diagnostic troubleshooting Nigel Kirkland
2026-09-17 22:22 ` sashiko-bot
2026-09-17 22:20 ` [PATCH v4 13/14] lpfc: Refactor calls on fc_disctmo to lpfc_set_disctmo in RSCN handler Nigel Kirkland
2026-09-17 22:20 ` [PATCH v4 14/14] lpfc: Update lpfc version to 15.0.0.1 Nigel Kirkland
2026-09-19 7:45 ` [PATCH v4 00/14] Update lpfc to revision 15.0.0.1 Nigel Kirkland
2026-09-28 17:13 ` Nigel Kirkland
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260917222015.61053-8-nkirkland2304@gmail.com \
--to=nkirkland2304@gmail.com \
--cc=linux-scsi@vger.kernel.org \
--cc=nigel.kirkland@broadcom.com \
--cc=paul.ely@broadcom.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox