* I
@ 2007-04-08 5:22 Clem P. Latham
0 siblings, 0 replies; 3+ messages in thread
From: Clem P. Latham @ 2007-04-08 5:22 UTC (permalink / raw)
To: nfs
[-- Attachment #1.1.1: Type: text/plain, Size: 3286 bytes --]
The other application is for use by health departments.
Never overload an inverter or leave it unattended. For more information or to register, visit us online at www. Obviously, most of us think that we are doing something to create a viable thermal infrared image. To help ensure accuracy, thermographers should be trained to at least Level II and, when possible, work with an experienced mentor until they have gained sufficient field experience. The ESS systems will be installed on HH-60 and HH-65 helicopters to enhance the Coast Guard's airborne use of force, interdiction and search and rescue missions.
Recent advances in technology have resulted in power inverters that are both small and dependable.
The Guideline applies to mammals, insects, and wood destroying organisms such as mold and fungi.
Overheating of tubes can reduce operational life or lead to catastrophic failure.
(Miami, FL) at the National Pest Control Association Annual Convention (October 2006), was a portable 4-camera system that could be wheeled into a facility and could be set up within minutes.
High temperature environments make contact measurements difficult or impossible. The stored results could then be downloaded from the DVR either onto a removable DVR or onto a flash memory device for remote viewing and report writing at a later time. For the vast majority of us, this information will provide us with some insights into a new technology and may provide us with a solution for a future problem. When connected to a water supply and placed in front of a building wall, a spray rack can be used to deliver a deluge of water to an area of interest. Insulated windows are a common feature found on modern commercial and residential structures.
Consult your local weather forecast before you set out and consider postponing your trip if extreme weather is predicted. Process heaters are similar to steam boilers in their construction except that hydrocarbon is passed through the firebox tubes instead of water. 4 million over ten years, is for the Coast Guard's Electro Optical Sensor System (ESS). The system was designated as a Critter Activity Tracking System, (CATS).
Developed by Infraspection and experts within the pest management industry, this 10 page document outlines procedures for detecting pests and pest related damage using a thermal imager. Under the right circumstances, infrared thermography can be used to provide qualitative and quantitative data for in-service heater tubes.
"We are delighted the Coast Guard has again chosen FLIR to provide these systems for such critical missions," he concluded. More than 290 top science students from the Corona Norco Unified School District, with an enrollment I excess of 50,000 competed for a spot to go on to the county competition.
Coast Guard and demonstrates the unique capabilities and extreme ruggedness of our systems," said Earl R. Process heaters are large, refractory-lined structures used to heat hydrocarbon product during refining. Edwards secured his spot at the District Level after taking first place at his individual school competition in February. When connected to a water supply and placed in front of a building wall, a spray rack can be used to deliver a deluge of water to an area of interest.
[-- Attachment #1.1.2: Type: text/html, Size: 4247 bytes --]
[-- Attachment #1.2: sunburn.gif --]
[-- Type: image/gif, Size: 10096 bytes --]
[-- Attachment #2: Type: text/plain, Size: 345 bytes --]
-------------------------------------------------------------------------
Take Surveys. Earn Cash. Influence the Future of IT
Join SourceForge.net's Techsay panel and you'll get the chance to share your
opinions on IT & business topics through brief surveys-and earn cash
http://www.techsay.com/default.php?page=join.php&p=sourceforge&CID=DEVDEV
[-- Attachment #3: Type: text/plain, Size: 140 bytes --]
_______________________________________________
NFS maillist - NFS@lists.sourceforge.net
https://lists.sourceforge.net/lists/listinfo/nfs
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v4 00/18] PCI/P2PDMA: Fix ACS egress control handling
@ 2026-08-21 19:38 Leon Romanovsky
2026-08-21 19:38 ` [PATCH v4 05/18] PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules Leon Romanovsky
0 siblings, 1 reply; 3+ messages in thread
From: Leon Romanovsky @ 2026-08-21 19:38 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Chaitanya Kulkarni,
Greg Kroah-Hartman, Jens Axboe, Alex Williamson, Leon Romanovsky,
Ankit Agrawal, Jason Gunthorpe, Jonathan Corbet, Shuah Khan,
Joerg Roedel (AMD), Will Deacon, Robin Murphy
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
Matt Evans
PCI P2PDMA treats any enabled ACS P2P Egress Control bit as an upstream
redirect. PCIe r7.0, sec 6.12.3, table 6-11 says the Egress Control
Vector bit for the target port decides instead: a clear bit routes a peer
request directly, regardless of P2P Request Redirect. Firmware can
therefore enable Egress Control with a permissive vector while Linux
incorrectly rejects a valid direct P2P path.
Table 6-11, where E is ACS P2P Egress Control Enable, R is ACS P2P
Request Redirect Enable and V the Egress Control Vector bit for the
target port:
E R V Required Handling for Peer-to-Peer Requests
- - - ------------------------------------------
0 0 x Route directly to peer-to-peer target
0 1 x Redirect Upstream
1 0 1 Handle as an ACS Violation
1 0 0 Route directly to peer-to-peer target
1 1 1 Redirect Upstream
1 1 0 Route directly to peer-to-peer target
P2P Completion Redirect lies outside this table and forces host-bridge
routing when set at the provider-side path divergence.
The same interaction affects target-independent ACS isolation checks.
Request Redirect does not guarantee that peer requests are forwarded
upstream while Egress Control is enabled because a clear vector bit
overrides it. Such checks cannot identify every potential target, so
treat Request Redirect as ineffective while Egress Control is enabled,
which merges the affected devices into one IOMMU group.
ACS Direct Translated P2P routes a Request carrying a Translated address
to the peer regardless of Request Redirect and Egress Control, so it
voids the same guarantee unless Translation Blocking rejects the Request
first.
That last rule holds only for a caller that needs Request Redirect to
isolate peers. pci_enable_pasid() asks for it so that a Request carrying
a PASID reaches the translation agent (sec 2.2.10.4), and a Translated
Request already carries an address the agent produced for that PASID
(sec 10.1.3). pci_acs_enabled() and pci_acs_path_enabled() therefore
take a scope, and Direct Translated P2P applies only to
PCI_ACS_SCOPE_ALL.
A pre-existing gap comes first. The routing analysis covers only Requests
carrying an Untranslated address; ACS Direct Translated P2P overrides
those controls, so that scope is now written down rather than implied.
It is nearly impossible to test all possible combinations due to limited
hardware availability, so I added KUnit coverage for ACS routing
decisions, isolation checks, Egress Control Vector lookups, and
provider-to-client path traversal over a fabricated PCIe fabric.
Disclaimer:
All patches were prepared with AI assistance, with a significant
difference between the code changes and the KUnit tests. The code
changes were thoroughly reviewed and rewritten.
In contrast, the KUnit patches were produced entirely by AI with
minimal human interaction, and multiple AI tools (Claude, Codex,
and Gemini) with frontier models were used to verify that the tests
comply with the PCI specification.
Thanks
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
Changes in v4:
- Added debug prints (we can drop it) patch which is very useful for automatic
root cause analysis. Just feed the output of these prints, together
with topology and kernel boot command line to your favorite LLM and it
will give you reliable RCA why ACS didn't work.
- Reject ACS Violations and unreadable routing state instead of treating
them as host-bridge redirects
- Added Tested-by tags from Tushar Dave
- Added support to asymmetric ACS routing
- Limited redirect checks to the two ports at the path divergence
- Added standalone ACS routing diagnostics for hardware retesting
- Link to v3: https://patch.msgid.link/20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com
Changes in v3:
- Fixed pci_p2pdma_add_resource() error unwinding
- Made pdev->p2pdma teardown wait unconditionally for RCU readers
- Restricted pci_p2pmem_find_many() to pool-backed providers
- Documented the pdev->p2pdma lifetime and RCU rules
- Fixed calc_map_type_and_dist() handling of the verbose argument
- Required the ACS port and target to share a bus before indexing the
Egress Control Vector
- Gave pci_acs_enabled() and pci_acs_path_enabled() a scope, so the ACS
Direct Translated P2P rule no longer stops pci_enable_pasid() from
enabling PASID
- Dropped "Report ACS ports when the paths share no upstream bridge":
the mapping type cannot change without a shared upstream bridge, so
the pci=disable_acs_redir= hint was not actionable there and the ACS
walk only cost config space reads
- Folded the Request Redirect rule into pci_acs_rr_ineffective(), so
pci_acs_flags_enabled() and the Intel SPT PCH quirk share one copy
- Renamed pci_acs_egress_ctrl_set() to pci_acs_egress_ctrl_is_set(), it
reads the bit rather than setting it
- Reworded the blocked-path warning: ACS may also leave the direct route
indeterminate rather than blocked
- Added KUnit coverage for the shared-bus guard, a device with no ACS
capability and an unreadable ACS Control register
- Added the missing Fixes: tags, a second one on the
pci_p2pdma_add_resource() unwinding fix (the dangling devres action
dates to f58ef9d1d135) and one on the Egress Control isolation change
- Link to v2: https://patch.msgid.link/20260806-fix-p2p-acs-v2-0-0cec14812965@nvidia.com
Changes in v2:
- Added Logan's ROB tags
- Added commas in Documentation patch
- Link to v1: https://patch.msgid.link/20260802-fix-p2p-acs-v1-0-a7c5eb64fff6@nvidia.com
---
Leon Romanovsky (18):
PCI/P2PDMA: Do not tear down the allocate attribute on registration failure
PCI/P2PDMA: Wait for RCU readers before freeing state
PCI/P2PDMA: Restrict the p2pmem search to pool backed providers
PCI/P2PDMA: Safely terminate ACS redirect lists
PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules
PCI/P2PDMA: Gate the host bridge whitelist warning on verbose
PCI/P2PDMA: Document the Address Type assumption
PCI: Account for Direct Translated P2P in ACS isolation checks
PCI: Add ACS egress control vector accessor
PCI: Account for ACS egress control in isolation checks
PCI/P2PDMA: Derive peer-to-peer routing from ACS control bits
PCI/P2PDMA: Honor ACS egress control vectors
PCI/P2PDMA: Document ACS egress control handling
PCI/P2PDMA: Extract pure ACS routing decision helpers
PCI/P2PDMA: Add KUnit tests for ACS routing decisions
PCI/P2PDMA: Add KUnit coverage for the ACS P2P routing walk
PCI: Add KUnit coverage for ACS isolation checks
PCI/P2PDMA: Log detailed ACS routing diagnostics
Documentation/admin-guide/kernel-parameters.txt | 9 +-
Documentation/driver-api/pci/p2pdma.rst | 24 +
drivers/iommu/iommu.c | 8 +-
drivers/pci/Kconfig | 15 +
drivers/pci/Makefile | 1 +
drivers/pci/ats.c | 11 +-
drivers/pci/p2pdma.c | 496 +++++++++++--
drivers/pci/pci.c | 115 ++-
drivers/pci/pci.h | 69 +-
drivers/pci/pci_acs_test.c | 920 ++++++++++++++++++++++++
drivers/pci/quirks.c | 62 +-
include/linux/pci-p2pdma.h | 8 +-
include/linux/pci.h | 30 +-
13 files changed, 1661 insertions(+), 107 deletions(-)
---
base-commit: 43598807f71ac1c9164f26004acf2496d4038daf
change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
Best regards,
--
Leon Romanovsky <leonro@nvidia.com>
^ permalink raw reply [flat|nested] 3+ messages in thread* [PATCH v4 05/18] PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules
2026-08-21 19:38 [PATCH v4 00/18] PCI/P2PDMA: Fix ACS egress control handling Leon Romanovsky
@ 2026-08-21 19:38 ` Leon Romanovsky
2026-08-24 20:29 ` I Logan Gunthorpe
0 siblings, 1 reply; 3+ messages in thread
From: Leon Romanovsky @ 2026-08-21 19:38 UTC (permalink / raw)
To: Bjorn Helgaas, Logan Gunthorpe, Chaitanya Kulkarni,
Greg Kroah-Hartman, Jens Axboe, Alex Williamson, Leon Romanovsky,
Ankit Agrawal, Jason Gunthorpe, Jonathan Corbet, Shuah Khan,
Joerg Roedel (AMD), Will Deacon, Robin Murphy
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
Matt Evans
From: Leon Romanovsky <leonro@nvidia.com>
pdev->p2pdma has two lifetime models. Provider-based entry points are
quiesced by their driver before remove completes. pci_p2pmem_find_many()
and the p2pmem sysfs attributes can race with unbind and therefore rely
on the teardown grace period.
Document publication, teardown, and how the grace period protects both
the struct pci_p2pdma object and its optional gen_pool.
Tested-by: Tushar Dave <tdave@nvidia.com>
Cc: Alex Williamson <alex@shazbot.org>
Cc: Matt Evans <matt@ozlabs.org>
Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
---
drivers/pci/p2pdma.c | 54 +++++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 53 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
index a77ef9deb3c6..49bc8cf06240 100644
--- a/drivers/pci/p2pdma.c
+++ b/drivers/pci/p2pdma.c
@@ -21,6 +21,39 @@
#include <linux/seq_buf.h>
#include <linux/xarray.h>
+/*
+ * Lifetime and RCU usage
+ *
+ * Within one driver bind, pdev->p2pdma is published exactly once, by
+ * pcim_p2pdma_init(), and cleared exactly once, by the pci_p2pdma_release()
+ * devres action that the same function installs. It is never re-pointed at a
+ * second struct pci_p2pdma, so a reader that observes a non-NULL pointer
+ * always observes the same, fully initialised object. That object is devres
+ * memory allocated before the action is installed, so devres frees it only
+ * after pci_p2pdma_release() has returned.
+ *
+ * Most exported entry points reach pdev->p2pdma through a struct pci_dev or a
+ * struct p2pdma_provider owned by the provider driver, and
+ * pcim_p2pdma_provider() requires callers to drop those references before the
+ * driver's remove() completes. Those cannot run concurrently with
+ * pci_p2pdma_release(), and their rcu_dereference() calls are simply how an
+ * __rcu pointer is read.
+ *
+ * pci_p2pmem_find_many() and the p2pmem sysfs attributes are the exceptions.
+ * The first walks every PCI device, so it can reach a provider whose driver is
+ * unbinding: pci_get_device() pins the struct pci_dev, not the driver. The
+ * second is reachable from userspace until sysfs_remove_group() runs at the end
+ * of the release. pci_has_p2pmem() must dereference the object to determine
+ * whether it owns a gen_pool, so even a poolless object must remain alive until
+ * that RCU reader exits. The sysfs group is created with the pool.
+ *
+ * The grace period in pci_p2pdma_release() first protects the struct
+ * pci_p2pdma itself from being freed while pci_has_p2pmem() is using it. For a
+ * pool-backed provider it also fences the gen_pool: gen_pool_alloc_owner()
+ * walks pool->chunks under RCU and gen_pool_destroy() frees those chunks
+ * without waiting for a grace period of its own, so pci_alloc_p2pmem() and
+ * p2pmem_alloc_mmap() hold rcu_read_lock() across the allocation.
+ */
struct pci_p2pdma {
struct gen_pool *pool;
bool p2pmem_published;
@@ -235,9 +268,19 @@ static void pci_p2pdma_release(void *data)
if (!p2pdma)
return;
- /* Flush and disable pci_alloc_p2p_mem() */
+ /*
+ * Stop new RCU readers and wait for readers that observed p2pdma before
+ * allowing devres to free it. This is required even without a pool,
+ * because pci_has_p2pmem() dereferences every non-NULL p2pdma it finds.
+ * For a pool-backed provider this also fences gen_pool_destroy().
+ */
RCU_INIT_POINTER(pdev->p2pdma, NULL);
synchronize_rcu();
+
+ /*
+ * The grace period also ensures no RCU reader can still be accessing
+ * map_types here.
+ */
xa_destroy(&p2pdma->map_types);
if (!p2pdma->pool)
@@ -255,6 +298,9 @@ static void pci_p2pdma_release(void *data)
* for a PCI device. It allocates and sets up the necessary data
* structures to support P2PDMA operations, including mapping type
* tracking.
+ *
+ * The state is published once per driver bind and torn down by a devres
+ * action on unbind. Repeated calls for the same device are a no-op.
*/
int pcim_p2pdma_init(struct pci_dev *pdev)
{
@@ -786,6 +832,12 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
}
done:
+ /*
+ * pci_p2pmem_find_many() reaches this with a provider whose driver may
+ * be unbinding, so the store runs under RCU: pci_p2pdma_release()
+ * clears the pointer and waits for readers before destroying
+ * map_types. See "Lifetime and RCU usage" above.
+ */
rcu_read_lock();
p2pdma = rcu_dereference(provider->p2pdma);
if (p2pdma)
--
2.55.0
^ permalink raw reply related [flat|nested] 3+ messages in thread* I
2026-08-21 19:38 ` [PATCH v4 05/18] PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules Leon Romanovsky
@ 2026-08-24 20:29 ` Logan Gunthorpe
2026-08-30 9:14 ` I Leon Romanovsky
0 siblings, 1 reply; 3+ messages in thread
From: Logan Gunthorpe @ 2026-08-24 20:29 UTC (permalink / raw)
To: Leon Romanovsky, Bjorn Helgaas, Chaitanya Kulkarni,
Greg Kroah-Hartman, Jens Axboe, Alex Williamson, Ankit Agrawal,
Jason Gunthorpe, Jonathan Corbet, Shuah Khan, Joerg Roedel (AMD),
Will Deacon, Robin Murphy
Cc: linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
Matt Evans
On 2026-08-21 13:38, Leon Romanovsky wrote:
> From: Leon Romanovsky <leonro@nvidia.com>
>
> pdev->p2pdma has two lifetime models. Provider-based entry points are
> quiesced by their driver before remove completes. pci_p2pmem_find_many()
> and the p2pmem sysfs attributes can race with unbind and therefore rely
> on the teardown grace period.
>
> Document publication, teardown, and how the grace period protects both
> the struct pci_p2pdma object and its optional gen_pool.
>
> Tested-by: Tushar Dave <tdave@nvidia.com>
> Cc: Alex Williamson <alex@shazbot.org>
> Cc: Matt Evans <matt@ozlabs.org>
> Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> ---
> drivers/pci/p2pdma.c | 54 +++++++++++++++++++++++++++++++++++++++++++++++++++-
> 1 file changed, 53 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> index a77ef9deb3c6..49bc8cf06240 100644
> --- a/drivers/pci/p2pdma.c
> +++ b/drivers/pci/p2pdma.c
> @@ -21,6 +21,39 @@
> #include <linux/seq_buf.h>
> #include <linux/xarray.h>
>
> +/*
> + * Lifetime and RCU usage
> + *
> + * Within one driver bind, pdev->p2pdma is published exactly once, by
Is published the right verb here? Seems like we use published for
different purposes in p2pdma and calling pcim_p2pdma_init() publishing
reads strangely.
> + * pcim_p2pdma_init(), and cleared exactly once, by the pci_p2pdma_release()
> + * devres action that the same function installs. It is never re-pointed at a
"It is never re-pointed at" is some strange wording. I had to read it a
few times to understand what it is saying.
> + * second struct pci_p2pdma, so a reader that observes a non-NULL pointer
> + * always observes the same, fully initialised object. That object is devres
> + * memory allocated before the action is installed, so devres frees it only
> + * after pci_p2pdma_release() has returned.
I don't quite follow the point of this paragraph. It's like it's
building to some kind of gotcha, but all it seems to be saying is is the
life cycle of pdev->p2pdma is the same as the lifecycle of pdev.
> + * Most exported entry points reach pdev->p2pdma through a struct pci_dev or a
> + * struct p2pdma_provider owned by the provider driver, and
> + * pcim_p2pdma_provider() requires callers to drop those references before the
> + * driver's remove() completes. Those cannot run concurrently with
> + * pci_p2pdma_release(), and their rcu_dereference() calls are simply how an
> + * __rcu pointer is read.
> + *
> + * pci_p2pmem_find_many() and the p2pmem sysfs attributes are the exceptions.
> + * The first walks every PCI device, so it can reach a provider whose driver is
> + * unbinding: pci_get_device() pins the struct pci_dev, not the driver. The
> + * second is reachable from userspace until sysfs_remove_group() runs at the end
> + * of the release. pci_has_p2pmem() must dereference the object to determine
> + * whether it owns a gen_pool, so even a poolless object must remain alive until
Maybe a hyphen with pool-less or maybe better to use plain language: "so
even a device without a pool must remain alive..."
> + * that RCU reader exits. The sysfs group is created with the pool.
> + *
> + * The grace period in pci_p2pdma_release() first protects the struct
> + * pci_p2pdma itself from being freed while pci_has_p2pmem() is using it. For a
> + * pool-backed provider it also fences the gen_pool: gen_pool_alloc_owner()
> + * walks pool->chunks under RCU and gen_pool_destroy() frees those chunks
> + * without waiting for a grace period of its own, so pci_alloc_p2pmem() and
> + * p2pmem_alloc_mmap() hold rcu_read_lock() across the allocation.
> + */
> struct pci_p2pdma {
> struct gen_pool *pool;
> bool p2pmem_published;
> @@ -235,9 +268,19 @@ static void pci_p2pdma_release(void *data)
> if (!p2pdma)
> return;
>
> - /* Flush and disable pci_alloc_p2p_mem() */
> + /*
> + * Stop new RCU readers and wait for readers that observed p2pdma before
> + * allowing devres to free it. This is required even without a pool,
> + * because pci_has_p2pmem() dereferences every non-NULL p2pdma it finds.
> + * For a pool-backed provider this also fences gen_pool_destroy().
> + */
> RCU_INIT_POINTER(pdev->p2pdma, NULL);
> synchronize_rcu();
> +
> + /*
> + * The grace period also ensures no RCU reader can still be accessing
> + * map_types here.
> + */
> xa_destroy(&p2pdma->map_types);
>
> if (!p2pdma->pool)
> @@ -255,6 +298,9 @@ static void pci_p2pdma_release(void *data)
> * for a PCI device. It allocates and sets up the necessary data
> * structures to support P2PDMA operations, including mapping type
> * tracking.
> + *
> + * The state is published once per driver bind and torn down by a devres
I still find the use of "published" here a bit odd and I'm not sure what
"state" it is referring to.
> + * action on unbind. Repeated calls for the same device are a no-op.
> */
> int pcim_p2pdma_init(struct pci_dev *pdev)
> {
> @@ -786,6 +832,12 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
> }
> done:
> + /*
> + * pci_p2pmem_find_many() reaches this with a provider whose driver may
> + * be unbinding, so the store runs under RCU: pci_p2pdma_release()
> + * clears the pointer and waits for readers before destroying
> + * map_types. See "Lifetime and RCU usage" above.
> + */
I don't know, but this seems like we're just describing basic RCU usage
here. I'm not sure I personally find much value in the comment.
> rcu_read_lock();
> p2pdma = rcu_dereference(provider->p2pdma);
> if (p2pdma)
>
Thanks,
Logan
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: I
2026-08-24 20:29 ` I Logan Gunthorpe
@ 2026-08-30 9:14 ` Leon Romanovsky
0 siblings, 0 replies; 3+ messages in thread
From: Leon Romanovsky @ 2026-08-30 9:14 UTC (permalink / raw)
To: Logan Gunthorpe
Cc: Bjorn Helgaas, Chaitanya Kulkarni, Greg Kroah-Hartman, Jens Axboe,
Alex Williamson, Ankit Agrawal, Jason Gunthorpe, Jonathan Corbet,
Shuah Khan, Joerg Roedel (AMD), Will Deacon, Robin Murphy,
linux-pci, linux-kernel, linux-doc, iommu, Tushar Dave,
Matt Evans
On Mon, Aug 24, 2026 at 02:29:34PM -0600, Logan Gunthorpe wrote:
>
>
> On 2026-08-21 13:38, Leon Romanovsky wrote:
> > From: Leon Romanovsky <leonro@nvidia.com>
> >
> > pdev->p2pdma has two lifetime models. Provider-based entry points are
> > quiesced by their driver before remove completes. pci_p2pmem_find_many()
> > and the p2pmem sysfs attributes can race with unbind and therefore rely
> > on the teardown grace period.
> >
> > Document publication, teardown, and how the grace period protects both
> > the struct pci_p2pdma object and its optional gen_pool.
> >
> > Tested-by: Tushar Dave <tdave@nvidia.com>
> > Cc: Alex Williamson <alex@shazbot.org>
> > Cc: Matt Evans <matt@ozlabs.org>
> > Signed-off-by: Leon Romanovsky <leonro@nvidia.com>
> > ---
> > drivers/pci/p2pdma.c | 54 +++++++++++++++++++++++++++++++++++++++++++++++++++-
> > 1 file changed, 53 insertions(+), 1 deletion(-)
> >
> > diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c
> > index a77ef9deb3c6..49bc8cf06240 100644
> > --- a/drivers/pci/p2pdma.c
> > +++ b/drivers/pci/p2pdma.c
> > @@ -21,6 +21,39 @@
> > #include <linux/seq_buf.h>
> > #include <linux/xarray.h>
> >
> > +/*
> > + * Lifetime and RCU usage
> > + *
> > + * Within one driver bind, pdev->p2pdma is published exactly once, by
>
> Is published the right verb here? Seems like we use published for
> different purposes in p2pdma and calling pcim_p2pdma_init() publishing
> reads strangely.
It should probably say "pdev->p2pdma is set exactly once".
>
> > + * pcim_p2pdma_init(), and cleared exactly once, by the pci_p2pdma_release()
> > + * devres action that the same function installs. It is never re-pointed at a
>
> "It is never re-pointed at" is some strange wording. I had to read it a
> few times to understand what it is saying.
Sorry about that. AI + non-english speaker = "re-pointed".
>
> > + * second struct pci_p2pdma, so a reader that observes a non-NULL pointer
> > + * always observes the same, fully initialised object. That object is devres
> > + * memory allocated before the action is installed, so devres frees it only
> > + * after pci_p2pdma_release() has returned.
>
> I don't quite follow the point of this paragraph. It's like it's
> building to some kind of gotcha, but all it seems to be saying is is the
> life cycle of pdev->p2pdma is the same as the lifecycle of pdev.
Yes
>
> > + * Most exported entry points reach pdev->p2pdma through a struct pci_dev or a
> > + * struct p2pdma_provider owned by the provider driver, and
> > + * pcim_p2pdma_provider() requires callers to drop those references before the
> > + * driver's remove() completes. Those cannot run concurrently with
> > + * pci_p2pdma_release(), and their rcu_dereference() calls are simply how an
> > + * __rcu pointer is read.
> > + *
> > + * pci_p2pmem_find_many() and the p2pmem sysfs attributes are the exceptions.
> > + * The first walks every PCI device, so it can reach a provider whose driver is
> > + * unbinding: pci_get_device() pins the struct pci_dev, not the driver. The
> > + * second is reachable from userspace until sysfs_remove_group() runs at the end
> > + * of the release. pci_has_p2pmem() must dereference the object to determine
> > + * whether it owns a gen_pool, so even a poolless object must remain alive until
>
> Maybe a hyphen with pool-less or maybe better to use plain language: "so
> even a device without a pool must remain alive..."
I will try to simplify the language.
>
> > + * that RCU reader exits. The sysfs group is created with the pool.
> > + *
> > + * The grace period in pci_p2pdma_release() first protects the struct
> > + * pci_p2pdma itself from being freed while pci_has_p2pmem() is using it. For a
> > + * pool-backed provider it also fences the gen_pool: gen_pool_alloc_owner()
> > + * walks pool->chunks under RCU and gen_pool_destroy() frees those chunks
> > + * without waiting for a grace period of its own, so pci_alloc_p2pmem() and
> > + * p2pmem_alloc_mmap() hold rcu_read_lock() across the allocation.
> > + */
> > struct pci_p2pdma {
> > struct gen_pool *pool;
> > bool p2pmem_published;
> > @@ -235,9 +268,19 @@ static void pci_p2pdma_release(void *data)
> > if (!p2pdma)
> > return;
> >
> > - /* Flush and disable pci_alloc_p2p_mem() */
> > + /*
> > + * Stop new RCU readers and wait for readers that observed p2pdma before
> > + * allowing devres to free it. This is required even without a pool,
> > + * because pci_has_p2pmem() dereferences every non-NULL p2pdma it finds.
> > + * For a pool-backed provider this also fences gen_pool_destroy().
> > + */
> > RCU_INIT_POINTER(pdev->p2pdma, NULL);
> > synchronize_rcu();
> > +
> > + /*
> > + * The grace period also ensures no RCU reader can still be accessing
> > + * map_types here.
> > + */
> > xa_destroy(&p2pdma->map_types);
> >
> > if (!p2pdma->pool)
> > @@ -255,6 +298,9 @@ static void pci_p2pdma_release(void *data)
> > * for a PCI device. It allocates and sets up the necessary data
> > * structures to support P2PDMA operations, including mapping type
> > * tracking.
> > + *
> > + * The state is published once per driver bind and torn down by a devres
>
> I still find the use of "published" here a bit odd and I'm not sure what
> "state" it is referring to.
published == assigned.
>
> > + * action on unbind. Repeated calls for the same device are a no-op.
> > */
> > int pcim_p2pdma_init(struct pci_dev *pdev)
> > {
> > @@ -786,6 +832,12 @@ calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client,
> > map_type = PCI_P2PDMA_MAP_NOT_SUPPORTED;
> > }
> > done:
> > + /*
> > + * pci_p2pmem_find_many() reaches this with a provider whose driver may
> > + * be unbinding, so the store runs under RCU: pci_p2pdma_release()
> > + * clears the pointer and waits for readers before destroying
> > + * map_types. See "Lifetime and RCU usage" above.
> > + */
>
> I don't know, but this seems like we're just describing basic RCU usage
> here. I'm not sure I personally find much value in the comment.
It was mainly intended to help AI review P2P patches according to the
lifetime model.
>
> > rcu_read_lock();
> > p2pdma = rcu_dereference(provider->p2pdma);
> > if (p2pdma)
> >
>
> Thanks,
>
> Logan
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-08-30 9:15 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2007-04-08 5:22 I Clem P. Latham
-- strict thread matches above, loose matches on Subject: below --
2026-08-21 19:38 [PATCH v4 00/18] PCI/P2PDMA: Fix ACS egress control handling Leon Romanovsky
2026-08-21 19:38 ` [PATCH v4 05/18] PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules Leon Romanovsky
2026-08-24 20:29 ` I Logan Gunthorpe
2026-08-30 9:14 ` I Leon Romanovsky
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.