* [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition
@ 2026-08-06 2:12 Pei Xiao
2026-08-06 2:37 ` sashiko-bot
2026-08-06 15:32 ` Logan Gunthorpe
0 siblings, 2 replies; 3+ messages in thread
From: Pei Xiao @ 2026-08-06 2:12 UTC (permalink / raw)
To: kurt.schwemmer, logang, bhelgaas, linux-pci, linux-kernel
Cc: Pei Xiao, stable
In stdev_create, &stdev->mrpc_work is bound with mrpc_event_work, and
&stdev->link_event_work is bound with link_event_work. The IRQ handlers
switchtec_event_isr and switchtec_dma_mrpc_isr can schedule these works
on system_wq (via schedule_work() in the ISRs and via
check_link_state_events()).
If we remove the device, switchtec_pci_remove makes cleanup and the
memory allocated for stdev is released by put_device() ->
stdev_release() -> kfree(stdev), while the works mentioned above may
still be pending or running. The sequence of operations that may lead
to a UAF bug is as follows:
CPU0 CPU1
| switchtec_event_isr
| schedule_work(&stdev->mrpc_work)
switchtec_pci_remove |
cdev_device_del(&stdev->cdev, |
&stdev->dev) |
stdev_kill(stdev) |
switchtec_exit_pci(stdev) |
pci_dev_put(stdev->pdev) |
put_device(&stdev->dev) |
// stdev_release -> kfree(stdev) |
| mrpc_event_work
| // use stdev (use-after-free)
Fix it by quiescing the interrupt sources before canceling the works:
stdev_kill() first clears PCI bus mastering, then explicitly frees both
IRQs, so no handler can be running and scheduling new work while the works
are canceled.
Fixes: 080b47def5e5 ("MicroSemi Switchtec management interface driver")
Fixes: 48c302dc8f3a ("NTB: switchtec: Add link event notifier callback")
Cc: stable@vger.kernel.org
Assisted-by: Codex:deepseek-v4-flash
Signed-off-by: Pei Xiao <xiaopei01@kylinos.cn>
---
changes in v2:
1.Add explicitly devm_free_irq
2.cacel mrpc_work and link_event_work move to before mrpc_timeout
3.Add event_irq and dma_mrpc_irq in struct stdev
4.Add Cc: stable@vger.kernel.org
---
drivers/pci/switch/switchtec.c | 16 +++++++++++++++-
include/linux/switchtec.h | 2 ++
2 files changed, 17 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/switch/switchtec.c b/drivers/pci/switch/switchtec.c
index 5711aaa5df11..235ca1877b6c 100644
--- a/drivers/pci/switch/switchtec.c
+++ b/drivers/pci/switch/switchtec.c
@@ -1318,6 +1318,13 @@ static void stdev_kill(struct switchtec_dev *stdev)
pci_clear_master(stdev->pdev);
+ if (stdev->event_irq >= 0)
+ devm_free_irq(&stdev->pdev->dev, stdev->event_irq, stdev);
+ if (stdev->dma_mrpc_irq >= 0)
+ devm_free_irq(&stdev->pdev->dev, stdev->dma_mrpc_irq, stdev);
+
+ cancel_work_sync(&stdev->mrpc_work);
+ cancel_work_sync(&stdev->link_event_work);
cancel_delayed_work_sync(&stdev->mrpc_timeout);
/* Mark the hardware as unavailable and complete all completions */
@@ -1356,6 +1363,8 @@ static struct switchtec_dev *stdev_create(struct pci_dev *pdev)
INIT_LIST_HEAD(&stdev->mrpc_queue);
mutex_init(&stdev->mrpc_mutex);
stdev->mrpc_busy = 0;
+ stdev->event_irq = -1;
+ stdev->dma_mrpc_irq = -1;
INIT_WORK(&stdev->mrpc_work, mrpc_event_work);
INIT_DELAYED_WORK(&stdev->mrpc_timeout, mrpc_timeout_work);
INIT_WORK(&stdev->link_event_work, link_event_work);
@@ -1513,6 +1522,7 @@ static int switchtec_init_isr(struct switchtec_dev *stdev)
if (rc)
return rc;
+ stdev->event_irq = event_irq;
if (!stdev->dma_mrpc)
return rc;
@@ -1529,7 +1539,11 @@ static int switchtec_init_isr(struct switchtec_dev *stdev)
switchtec_dma_mrpc_isr, 0,
KBUILD_MODNAME, stdev);
- return rc;
+ if (rc)
+ return rc;
+ stdev->dma_mrpc_irq = dma_mrpc_irq;
+
+ return 0;
}
static void init_pff(struct switchtec_dev *stdev)
diff --git a/include/linux/switchtec.h b/include/linux/switchtec.h
index 724da6c08bf7..fd38d3e7f3b6 100644
--- a/include/linux/switchtec.h
+++ b/include/linux/switchtec.h
@@ -500,6 +500,8 @@ struct switchtec_dev {
struct mutex mrpc_mutex;
struct list_head mrpc_queue;
int mrpc_busy;
+ int event_irq;
+ int dma_mrpc_irq;
struct work_struct mrpc_work;
struct delayed_work mrpc_timeout;
bool alive;
--
2.25.1
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition
2026-08-06 2:12 [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition Pei Xiao
@ 2026-08-06 2:37 ` sashiko-bot
2026-08-06 15:32 ` Logan Gunthorpe
1 sibling, 0 replies; 3+ messages in thread
From: sashiko-bot @ 2026-08-06 2:37 UTC (permalink / raw)
To: Pei Xiao; +Cc: linux-pci
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
Pre-existing issues:
- [High] The delayed work `mrpc_timeout` can be concurrently re-armed by `switchtec_dev_write()` after being cancelled in `stdev_kill()`, leading to UAF of devm-freed MMIO or `stdev`.
- [High] Probe error path bypasses `stdev_kill()`, leading to UAF of `stdev` by active devm-managed IRQs and uncancelled workqueues.
- [High] Missing `dma_rmb()` between reading DMA status and DMA data.
--
--- Patch [1]: [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition ---
commit 1cdcc605287dd43ead581c21636460c8793d36f2
Author: Pei Xiao <xiaopei01@kylinos.cn>
Subject: PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition
This commit attempts to fix a use-after-free bug during device removal by
explicitly clearing bus master and freeing IRQs before cancelling workqueues.
> diff --git a/drivers/pci/switch/switchtec.c b/drivers/pci/switch/switchtec.c
> index 41fc4b5127081..4b0988013c69d 100644
> --- a/drivers/pci/switch/switchtec.c
> +++ b/drivers/pci/switch/switchtec.c
> @@ -1318,6 +1318,13 @@ static void stdev_kill(struct switchtec_dev *stdev)
>
> pci_clear_master(stdev->pdev);
>
> + if (stdev->event_irq >= 0)
> + devm_free_irq(&stdev->pdev->dev, stdev->event_irq, stdev);
> + if (stdev->dma_mrpc_irq >= 0)
> + devm_free_irq(&stdev->pdev->dev, stdev->dma_mrpc_irq, stdev);
> +
> + cancel_work_sync(&stdev->mrpc_work);
> + cancel_work_sync(&stdev->link_event_work);
> cancel_delayed_work_sync(&stdev->mrpc_timeout);
>
> /* Mark the hardware as unavailable and complete all completions */
[Severity: High]
This is a pre-existing issue, but can the delayed work mrpc_timeout be
concurrently re-armed after it is cancelled here?
During device removal, stdev_kill() cancels mrpc_timeout before holding
mrpc_mutex and setting alive = false:
drivers/pci/switch/switchtec.c:stdev_kill() {
...
cancel_delayed_work_sync(&stdev->mrpc_timeout);
scoped_guard (mutex, &stdev->mrpc_mutex) {
stdev->alive = false;
...
}
Could userspace concurrently call switchtec_dev_write(), pass the alive
check, and re-arm the delayed work?
drivers/pci/switch/switchtec.c:switchtec_dev_write() {
...
rc = lock_mutex_and_test_alive(stdev);
if (rc)
return rc;
...
schedule_delayed_work(&stdev->mrpc_timeout, ...);
}
If the timeout fires later, could this result in a use-after-free when
stdev has already been freed by devres?
[ ... ]
> @@ -1529,7 +1539,11 @@ static int switchtec_init_isr(struct switchtec_dev *stdev)
> switchtec_dma_mrpc_isr, 0,
> KBUILD_MODNAME, stdev);
>
> - return rc;
> + if (rc)
> + return rc;
[Severity: High]
This isn't a bug introduced by this patch, but if switchtec_init_isr() fails
here, does the probe error path leak resources or cause a use-after-free?
If we return rc here, the caller switchtec_pci_probe() jumps to err_exit_pci:
drivers/pci/switch/switchtec.c:switchtec_pci_probe() {
...
rc = switchtec_init_isr(stdev);
if (rc) {
dev_err(&stdev->dev, "failed to init isr.\n");
goto err_exit_pci;
}
...
err_devadd:
stdev_kill(stdev);
err_exit_pci:
switchtec_exit_pci(stdev);
err_put:
ida_free(&switchtec_minor_ida, MINOR(stdev->dev.devt));
put_device(&stdev->dev);
}
Since err_exit_pci bypasses stdev_kill(), active devm-managed IRQs and
uncancelled workqueues (like mrpc_work) are not cleaned up before
put_device() frees stdev.
Could an interrupt fire during this window and execute switchtec_event_isr(),
leading to a use-after-free of stdev?
> + stdev->dma_mrpc_irq = dma_mrpc_irq;
> +
> + return 0;
> }
Additionally, while reviewing the code, I noticed another pre-existing issue
in mrpc_complete_cmd():
[Severity: High]
This is a pre-existing issue, but is there a missing dma_rmb() between
reading the DMA status and the DMA data?
drivers/pci/switch/switchtec.c:mrpc_complete_cmd() {
if (stdev->dma_mrpc)
stuser->status = stdev->dma_mrpc->status;
else
stuser->status = ioread32(&stdev->mmio_mrpc->status);
if (stuser->status == SWITCHTEC_MRPC_STATUS_INPROGRESS)
return;
...
if (stdev->dma_mrpc)
memcpy(stuser->data, &stdev->dma_mrpc->data,
stuser->read_len);
}
Since the DMA buffer is allocated with dma_alloc_coherent(), memory accesses
are coherent but not strictly ordered on weakly-ordered architectures like
ARM64.
Can the CPU speculatively read the payload data before the status
read completes, resulting in stale data being processed?
--
Sashiko AI review · https://sashiko.dev/#/patchset/1f52d931ec7a7f5b36837cabf07b36e0463bf182.1785982054.git.xiaopei01@kylinos.cn?part=1
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition
2026-08-06 2:12 [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition Pei Xiao
2026-08-06 2:37 ` sashiko-bot
@ 2026-08-06 15:32 ` Logan Gunthorpe
1 sibling, 0 replies; 3+ messages in thread
From: Logan Gunthorpe @ 2026-08-06 15:32 UTC (permalink / raw)
To: Pei Xiao, kurt.schwemmer, bhelgaas, linux-pci, linux-kernel; +Cc: stable
On 2026-08-05 8:12 p.m., Pei Xiao wrote:
> In stdev_create, &stdev->mrpc_work is bound with mrpc_event_work, and
> &stdev->link_event_work is bound with link_event_work. The IRQ handlers
> switchtec_event_isr and switchtec_dma_mrpc_isr can schedule these works
> on system_wq (via schedule_work() in the ISRs and via
> check_link_state_events()).
>
> If we remove the device, switchtec_pci_remove makes cleanup and the
> memory allocated for stdev is released by put_device() ->
> stdev_release() -> kfree(stdev), while the works mentioned above may
> still be pending or running. The sequence of operations that may lead
> to a UAF bug is as follows:
>
> CPU0 CPU1
>
> | switchtec_event_isr
> | schedule_work(&stdev->mrpc_work)
> switchtec_pci_remove |
> cdev_device_del(&stdev->cdev, |
> &stdev->dev) |
> stdev_kill(stdev) |
> switchtec_exit_pci(stdev) |
> pci_dev_put(stdev->pdev) |
> put_device(&stdev->dev) |
> // stdev_release -> kfree(stdev) |
> | mrpc_event_work
> | // use stdev (use-after-free)
>
> Fix it by quiescing the interrupt sources before canceling the works:
> stdev_kill() first clears PCI bus mastering, then explicitly frees both
> IRQs, so no handler can be running and scheduling new work while the works
> are canceled.
>
> Fixes: 080b47def5e5 ("MicroSemi Switchtec management interface driver")
> Fixes: 48c302dc8f3a ("NTB: switchtec: Add link event notifier callback")
> Cc: stable@vger.kernel.org
> Assisted-by: Codex:deepseek-v4-flash
> Signed-off-by: Pei Xiao <xiaopei01@kylinos.cn>
> ---
> changes in v2:
> 1.Add explicitly devm_free_irq
> 2.cacel mrpc_work and link_event_work move to before mrpc_timeout
> 3.Add event_irq and dma_mrpc_irq in struct stdev
> 4.Add Cc: stable@vger.kernel.org
> ---
> drivers/pci/switch/switchtec.c | 16 +++++++++++++++-
> include/linux/switchtec.h | 2 ++
> 2 files changed, 17 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/pci/switch/switchtec.c b/drivers/pci/switch/switchtec.c
> index 5711aaa5df11..235ca1877b6c 100644
> --- a/drivers/pci/switch/switchtec.c
> +++ b/drivers/pci/switch/switchtec.c
> @@ -1318,6 +1318,13 @@ static void stdev_kill(struct switchtec_dev *stdev)
>
> pci_clear_master(stdev->pdev);
>
> + if (stdev->event_irq >= 0)
> + devm_free_irq(&stdev->pdev->dev, stdev->event_irq, stdev);
> + if (stdev->dma_mrpc_irq >= 0)
> + devm_free_irq(&stdev->pdev->dev, stdev->dma_mrpc_irq, stdev);
> +
> + cancel_work_sync(&stdev->mrpc_work);
> + cancel_work_sync(&stdev->link_event_work);
> cancel_delayed_work_sync(&stdev->mrpc_timeout);
>
> /* Mark the hardware as unavailable and complete all completions */
> @@ -1356,6 +1363,8 @@ static struct switchtec_dev *stdev_create(struct pci_dev *pdev)
> INIT_LIST_HEAD(&stdev->mrpc_queue);
> mutex_init(&stdev->mrpc_mutex);
> stdev->mrpc_busy = 0;
> + stdev->event_irq = -1;
> + stdev->dma_mrpc_irq = -1;
I think the patch as a whole is correct. But one minor nit that could
maybe clean it up slightly. pci_irq_vector() cannot return 0. So I think
it would be fine to just leave these explictily initialized to zero and
check for non-zero in the tests above.
Other than that minor point:
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Thanks!
Logan
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-08-06 15:33 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-06 2:12 [PATCH v2] PCI: switchtec: Fix use-after-free in switchtec_pci_remove due to race condition Pei Xiao
2026-08-06 2:37 ` sashiko-bot
2026-08-06 15:32 ` Logan Gunthorpe
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox