From: Leon Romanovsky <leon@kernel.org>
To: Bjorn Helgaas <bhelgaas@google.com>,
Logan Gunthorpe <logang@deltatee.com>,
Chaitanya Kulkarni <kch@nvidia.com>,
Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
Jens Axboe <axboe@kernel.dk>, Alex Williamson <alex@shazbot.org>,
Leon Romanovsky <leon@kernel.org>,
Ankit Agrawal <ankita@nvidia.com>, Jason Gunthorpe <jgg@ziepe.ca>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
"Joerg Roedel (AMD)" <joro@8bytes.org>,
Will Deacon <will@kernel.org>,
Robin Murphy <robin.murphy@arm.com>
Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-doc@vger.kernel.org, iommu@lists.linux.dev
Subject: [PATCH v3 12/17] PCI/P2PDMA: Honor ACS egress control vectors
Date: Tue, 11 Aug 2026 12:30:54 +0300 [thread overview]
Message-ID: <20260811-fix-p2p-acs-v3-12-efc488ee7c03@nvidia.com> (raw)
In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com>
From: Leon Romanovsky <leonro@nvidia.com>
An enabled Egress Control bit does not itself send a peer request
upstream. PCIe r7.0, sec 6.12.3, table 6-11 makes the outcome depend on
the Egress Control Vector bit for the target port: a clear bit routes the
request directly regardless of P2P Request Redirect.
Read the vector where the paths diverge below their common upstream port.
Keep a clear vector bit on the direct path, subject to P2P Completion
Redirect.
A set bit with Request Redirect clear is an ACS Violation. ACS acts only
on peer-to-peer Requests, so route it, and an indeterminate vector,
through the host bridge.
Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory")
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 114 +++++++++++++++++++++++++++++++++------------------
1 file changed, 75 insertions(+), 39 deletions(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index 69cef8ca9557..879c92d66f5b 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -541,12 +541,13 @@ static struct pci_dev *find_parent_pci_dev(struct device *dev)
enum pci_acs_p2pdma_state {
PCI_ACS_P2PDMA_DIRECT,
PCI_ACS_P2PDMA_REDIRECT,
+ PCI_ACS_P2PDMA_NOT_SUPPORTED,
};
static enum pci_acs_p2pdma_state
pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci_dev *target)
{
- int pos;
+ int pos, ret;
u16 ctrl;
pos = pdev->acs_cap;
@@ -554,26 +555,26 @@ pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci_dev *target)
return PCI_ACS_P2PDMA_DIRECT;
if (pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl))
- return PCI_ACS_P2PDMA_REDIRECT;
+ return PCI_ACS_P2PDMA_NOT_SUPPORTED;
- if (!(ctrl & PCI_ACS_EC))
+ /* EC applies only at the path divergence where the target is known. */
+ if (!target || !(ctrl & PCI_ACS_EC))
return ctrl & (PCI_ACS_RR | PCI_ACS_CR) ?
PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT;
/*
- * The vector cannot be read without the peer target, so redirect
- * upstream until the paths diverge.
+ * PCIe r7.0, sec 6.12.3, table 6-11: a set Egress Control Vector
+ * bit redirects the request only when Request Redirect is set. With
+ * Request Redirect clear, the request is handled as an ACS Violation.
+ * A clear vector bit permits direct routing, subject to Completion
+ * Redirect.
*/
- if (!target)
- return PCI_ACS_P2PDMA_REDIRECT;
-
- /*
- * PCIe r7.0, sec 6.12.3, table 6-11: a set or indeterminate egress
- * control vector bit keeps the request off the direct path; a clear
- * bit permits it, subject only to completion redirect.
- */
- if (pci_acs_egress_ctrl_is_set(pdev, target))
- return PCI_ACS_P2PDMA_REDIRECT;
+ ret = pci_acs_egress_ctrl_is_set(pdev, target);
+ if (ret < 0)
+ return PCI_ACS_P2PDMA_NOT_SUPPORTED;
+ if (ret)
+ return ctrl & PCI_ACS_RR ? PCI_ACS_P2PDMA_REDIRECT :
+ PCI_ACS_P2PDMA_NOT_SUPPORTED;
return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT :
PCI_ACS_P2PDMA_DIRECT;
@@ -754,9 +755,9 @@ static unsigned long map_types_idx(struct pci_dev *client)
* then to Device B. The mapping type returned depends on the ACS
* redirection setting of the ports along the path.
*
- * If ACS redirect is set on any port in the path, traffic between the
- * devices will go through the host bridge, so return
- * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return
+ * If ACS redirects traffic on any port in the path, or blocks the direct
+ * path or leaves its routing indeterminate, return
+ * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE. Otherwise, return
* PCI_P2PDMA_MAP_BUS_ADDR.
*
* Any two devices that have a data path that goes through the host bridge
@@ -770,10 +771,13 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
int *dist, bool verbose)
{
enum pci_p2pdma_map_type map_type = PCI_P2PDMA_MAP_THRU_HOST_BRIDGE;
- struct pci_dev *a = provider, *b = client, *bb;
+ struct pci_dev *a = provider, *b = client, *bb, *target;
+ struct pci_dev *a_child = NULL, *b_child = NULL;
+ struct pci_dev *acs_unsupported = NULL;
+ enum pci_acs_p2pdma_state state;
struct pci_p2pdma *p2pdma;
struct seq_buf acs_list;
- int acs_cnt = 0;
+ int acs_redirect_cnt = 0;
int dist_a = 0;
int dist_b = 0;
char buf[128];
@@ -787,60 +791,92 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
*/
while (a) {
dist_b = 0;
-
- if (pci_acs_p2pdma_state(a, NULL) ==
- PCI_ACS_P2PDMA_REDIRECT) {
- seq_buf_print_bus_devfn(&acs_list, a);
- acs_cnt++;
- }
-
+ b_child = NULL;
bb = b;
while (bb) {
if (a == bb)
- goto check_b_path_acs;
+ goto check_paths_acs;
+ b_child = bb;
bb = pci_upstream_bridge(bb);
dist_b++;
}
+ a_child = a;
a = pci_upstream_bridge(a);
dist_a++;
}
+ /*
+ * The paths share no upstream bridge, so there is no direct path for
+ * ACS to gate: PCI_P2PDMA_MAP_BUS_ADDR is not reachable here and the
+ * request can only get to the peer through the host bridge.
+ */
*dist = dist_a + dist_b;
goto map_through_host_bridge;
-check_b_path_acs:
- bb = b;
+check_paths_acs:
+ *dist = dist_a + dist_b;
+ bb = provider;
while (bb) {
+ target = bb == a_child ? b_child : NULL;
+ state = pci_acs_p2pdma_state(bb, target);
+ if (state != PCI_ACS_P2PDMA_DIRECT) {
+ seq_buf_print_bus_devfn(&acs_list, bb);
+ if (state == PCI_ACS_P2PDMA_REDIRECT)
+ acs_redirect_cnt++;
+ else if (!acs_unsupported)
+ acs_unsupported = bb;
+ }
+
if (a == bb)
break;
- if (pci_acs_p2pdma_state(bb, NULL) ==
- PCI_ACS_P2PDMA_REDIRECT) {
+ bb = pci_upstream_bridge(bb);
+ }
+
+ bb = client;
+
+ while (bb && a != bb) {
+ target = bb == b_child ? a_child : NULL;
+ state = pci_acs_p2pdma_state(bb, target);
+ if (state != PCI_ACS_P2PDMA_DIRECT) {
seq_buf_print_bus_devfn(&acs_list, bb);
- acs_cnt++;
+ if (state == PCI_ACS_P2PDMA_REDIRECT)
+ acs_redirect_cnt++;
+ else if (!acs_unsupported)
+ acs_unsupported = bb;
}
bb = pci_upstream_bridge(bb);
}
- *dist = dist_a + dist_b;
-
- if (!acs_cnt) {
+ /*
+ * Below a shared upstream bridge, a path that no port redirects or
+ * blocks routes the request directly.
+ */
+ if (!acs_unsupported && !acs_redirect_cnt) {
map_type = PCI_P2PDMA_MAP_BUS_ADDR;
goto done;
}
+ /*
+ * ACS controls only act on Requests routed peer-to-peer, so a blocked
+ * or indeterminate direct path still leaves the host-bridge route.
+ */
if (verbose) {
/* Drop the final semicolon; the list is not empty here. */
if (!seq_buf_has_overflowed(&acs_list))
acs_list.buffer[acs_list.len - 1] = '\0';
- pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
- pci_name(provider));
- pci_warn(client, "to disable ACS redirect for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
+ if (acs_unsupported)
+ pci_warn(client, "ACS leaves no usable direct P2P path to provider %s at %s\n",
+ pci_name(provider), pci_name(acs_unsupported));
+ else
+ pci_warn(client, "ACS redirect is set between the client and provider (%s)\n",
+ pci_name(provider));
+ pci_warn(client, "to disable ACS controls for this path, add the kernel parameter: pci=disable_acs_redir=%s\n",
seq_buf_str(&acs_list));
}
--
2.55.0
next prev parent reply other threads:[~2026-08-11 9:32 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-11 9:30 [PATCH v3 00/17] PCI/P2PDMA: Fix ACS egress control handling Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 01/17] PCI/P2PDMA: Do not tear down the allocate attribute on registration failure Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 02/17] PCI/P2PDMA: Wait for RCU readers before freeing state Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 03/17] PCI/P2PDMA: Restrict the p2pmem search to pool backed providers Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 04/17] PCI/P2PDMA: Safely terminate ACS redirect lists Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 05/17] PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 06/17] PCI/P2PDMA: Gate the host bridge whitelist warning on verbose Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 07/17] PCI/P2PDMA: Document the Address Type assumption Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 08/17] PCI: Account for Direct Translated P2P in ACS isolation checks Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 09/17] PCI: Add ACS egress control vector accessor Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 10/17] PCI: Account for ACS egress control in isolation checks Leon Romanovsky
2026-08-11 9:30 ` Leon Romanovsky [this message]
2026-08-11 9:30 ` [PATCH v3 13/17] PCI/P2PDMA: Document ACS egress control handling Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 14/17] PCI/P2PDMA: Extract pure ACS routing decision helpers Leon Romanovsky
2026-08-11 9:30 ` [PATCH v3 15/17] PCI/P2PDMA: Add KUnit tests for ACS routing decisions Leon Romanovsky
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260811-fix-p2p-acs-v3-12-efc488ee7c03@nvidia.com \
--to=leon@kernel.org \
--cc=alex@shazbot.org \
--cc=ankita@nvidia.com \
--cc=axboe@kernel.dk \
--cc=bhelgaas@google.com \
--cc=corbet@lwn.net \
--cc=gregkh@linuxfoundation.org \
--cc=iommu@lists.linux.dev \
--cc=jgg@ziepe.ca \
--cc=joro@8bytes.org \
--cc=kch@nvidia.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=logang@deltatee.com \
--cc=robin.murphy@arm.com \
--cc=skhan@linuxfoundation.org \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox