From: Vidya Sagar <vidyas@nvidia.com>
To: <bhelgaas@google.com>, <tglx@kernel.org>,
<wangruikang@iscas.ac.cn>, <Frank.Li@nxp.com>,
<lihaoxiang@isrc.iscas.ac.cn>, <18255117159@163.com>,
<shawn.lin@rock-chips.com>, <xiangzao@linux.alibaba.com>
Cc: <vsethi@nvidia.com>, <sdonthineni@nvidia.com>,
<kthota@nvidia.com>, <mmaddireddy@nvidia.com>,
<kumarahul@nvidia.com>, <sagar.tv@gmail.com>,
<linux-pci@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
Vidya Sagar <vidyas@nvidia.com>
Subject: [PATCH V3] PCI/MSI: Don't touch the MSI-X table while the Link is contained
Date: Tue, 25 Aug 2026 22:57:19 +0530 [thread overview]
Message-ID: <20260825172719.4153402-1-vidyas@nvidia.com> (raw)
In-Reply-To: <20260825140952.4066140-1-vidyas@nvidia.com>
The MSI-X Table lives in device MMIO space behind a BAR, so it is only
reachable while the Link is up. While a Downstream Port has the Link
contained by DPC it completes accesses to the Table with Unsupported
Request, and reads return all ones.
If the upstream Root Port implements the RP Extensions for DPC, it
additionally reports that UR completion as an RP PIO error and answers
with a DPC of its own, which contains every other device below it. So a
containment event on a single Downstream Port can escalate into one at
the Root Port and take down unrelated devices.
pci_free_irq_vectors() is called from driver error_detected() and
prepare-for-reset callbacks, i.e. while the Link is contained, and it
masks every descriptor. Each mask is an MMIO write followed by a
non-posted flush read, so this is reached on every contained device whose
driver tears down its interrupts before the reset.
Skip the hardware access when the device is not in pci_channel_io_normal,
in addition to the existing surprise removal check. The msix_ctrl cache
is still updated, so __pci_restore_msix_state() replays the intended mask
state once the Link is back up. report_slot_reset() returns the device to
pci_channel_io_normal before invoking the driver callback, so re-enabling
and restoring MSI-X during recovery is unaffected.
error_state is set once containment has already occurred, so these checks
cover the case where the kernel knows the Link is down. They are not
mutual exclusion against a containment event that begins concurrently.
pci_msix_write_tph_tag() flushes its Vector Control update with an
unconditional read, which would otherwise be issued for a write that was
skipped, so return -EIO there instead. pcie_tph_set_st_entry() responds
by disabling TPH, which is preferable to reporting a Steering Tag update
that never reached the device.
Signed-off-by: Vidya Sagar <vidyas@nvidia.com>
---
Changes in v3:
- Move the pci_msi_dev_inaccessible() check in pci_msix_write_tph_tag()
under irq_desc::lock, next to the accesses it guards, instead of before
msi_descs_lock which can sleep (reported by Sashiko AI review).
Changes in v2:
- Return -EIO from pci_msix_write_tph_tag() so its unconditional flush
read is not issued for a skipped write (reported by Sashiko AI review).
- Rename pci_msix_mmio_unsafe() to pci_msi_dev_inaccessible(), since in
__pci_write_msi_msg() it also gates the Configuration Space MSI path.
- Note in the log why MSI-X restore during recovery is unaffected.
drivers/pci/msi/msi.c | 11 ++++++++++-
drivers/pci/msi/msi.h | 21 +++++++++++++++++++++
2 files changed, 31 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/msi/msi.c b/drivers/pci/msi/msi.c
index 80a9db417dc8..975948502f8e 100644
--- a/drivers/pci/msi/msi.c
+++ b/drivers/pci/msi/msi.c
@@ -249,7 +249,7 @@ void __pci_write_msi_msg(struct msi_desc *entry, struct msi_msg *msg)
{
struct pci_dev *dev = msi_desc_to_pci_dev(entry);
- if (dev->current_state != PCI_D0 || pci_dev_is_disconnected(dev)) {
+ if (dev->current_state != PCI_D0 || pci_msi_dev_inaccessible(dev)) {
/* Don't touch the hardware now */
} else if (entry->pci.msi_attrib.is_msix) {
pci_write_msg_msix(entry, msg);
@@ -976,6 +976,15 @@ int pci_msix_write_tph_tag(struct pci_dev *pdev, unsigned int index, u16 tag)
if (!msi_desc || msi_desc->pci.msi_attrib.is_virtual)
return -ENXIO;
+ /*
+ * The tag update below is a write to the MSI-X Table followed by a
+ * flush read, neither of which can be completed while the Link is
+ * contained. Check as late as possible, i.e. under irq_desc::lock, as
+ * containment can begin at any point. Let the caller disable TPH.
+ */
+ if (pci_msi_dev_inaccessible(pdev))
+ return -EIO;
+
FIELD_MODIFY(PCI_MSIX_ENTRY_CTRL_ST, &msi_desc->pci.msix_ctrl, tag);
pci_msix_write_vector_ctrl(msi_desc, msi_desc->pci.msix_ctrl);
/* Flush the write */
diff --git a/drivers/pci/msi/msi.h b/drivers/pci/msi/msi.h
index 0b420b319f50..c3194d8425c8 100644
--- a/drivers/pci/msi/msi.h
+++ b/drivers/pci/msi/msi.h
@@ -26,6 +26,20 @@ static inline void __iomem *pci_msix_desc_addr(struct msi_desc *desc)
return desc->pci.mask_base + desc->msi_index * PCI_MSIX_ENTRY_SIZE;
}
+/*
+ * The MSI-X Table lives in device MMIO space and the MSI Capability in
+ * Configuration Space, so both are only reachable while the Link is usable.
+ * While a Downstream Port has the Link contained by DPC it completes these
+ * accesses with Unsupported Request. If the upstream Root Port implements the
+ * RP Extensions for DPC, it reports that completion as an RP PIO error and
+ * answers with a DPC of its own, taking down every other device below it.
+ */
+static inline bool pci_msi_dev_inaccessible(struct pci_dev *pdev)
+{
+ return pdev->error_state != pci_channel_io_normal ||
+ pci_dev_is_disconnected(pdev);
+}
+
/*
* This internal function does not flush PCI writes to the device. All
* users must ensure that they read from the device before either assuming
@@ -36,6 +50,9 @@ static inline void pci_msix_write_vector_ctrl(struct msi_desc *desc, u32 ctrl)
{
void __iomem *desc_addr = pci_msix_desc_addr(desc);
+ if (pci_msi_dev_inaccessible(msi_desc_to_pci_dev(desc)))
+ return;
+
if (desc->pci.msi_attrib.can_mask)
writel(ctrl, desc_addr + PCI_MSIX_ENTRY_VECTOR_CTRL);
}
@@ -43,6 +60,10 @@ static inline void pci_msix_write_vector_ctrl(struct msi_desc *desc, u32 ctrl)
static inline void pci_msix_mask(struct msi_desc *desc)
{
desc->pci.msix_ctrl |= PCI_MSIX_ENTRY_CTRL_MASKBIT;
+
+ if (pci_msi_dev_inaccessible(msi_desc_to_pci_dev(desc)))
+ return;
+
pci_msix_write_vector_ctrl(desc, desc->pci.msix_ctrl);
/* Flush write to device */
readl(desc->pci.mask_base);
--
2.43.0
next prev parent reply other threads:[~2026-08-25 17:28 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-17 19:56 [PATCH V1] PCI/MSI: Don't touch the MSI-X table while the Link is contained Vidya Sagar
2026-08-17 20:12 ` sashiko-bot
2026-08-25 12:45 ` Vidya Sagar
2026-08-25 14:09 ` [PATCH V2] " Vidya Sagar
2026-08-25 14:25 ` sashiko-bot
2026-08-25 14:56 ` Vidya Sagar
2026-08-25 17:27 ` Vidya Sagar [this message]
2026-08-25 17:45 ` [PATCH V3] " sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825172719.4153402-1-vidyas@nvidia.com \
--to=vidyas@nvidia.com \
--cc=18255117159@163.com \
--cc=Frank.Li@nxp.com \
--cc=bhelgaas@google.com \
--cc=kthota@nvidia.com \
--cc=kumarahul@nvidia.com \
--cc=lihaoxiang@isrc.iscas.ac.cn \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=mmaddireddy@nvidia.com \
--cc=sagar.tv@gmail.com \
--cc=sdonthineni@nvidia.com \
--cc=shawn.lin@rock-chips.com \
--cc=tglx@kernel.org \
--cc=vsethi@nvidia.com \
--cc=wangruikang@iscas.ac.cn \
--cc=xiangzao@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox