Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
@ 2026-05-03 15:02 Yonatan Nachum
  2026-05-04  7:34 ` Jason Gunthorpe
  0 siblings, 1 reply; 14+ messages in thread
From: Yonatan Nachum @ 2026-05-03 15:02 UTC (permalink / raw)
  To: jgg, leon, linux-rdma
  Cc: mrgolin, sleybo, matua, gal.pressman, Yonatan Nachum,
	Yehuda Yitschak

Expose device P2P DMA support using the query device verbs.
If the device support P2P DMA, it can DMA directly to and from a peer
PCIe device

Reviewed-by: Michael Margolin <mrgolin@amazon.com>
Reviewed-by: Yehuda Yitschak <yehuday@amazon.com>
Signed-off-by: Yonatan Nachum <ynachum@amazon.com>
---
 drivers/infiniband/hw/efa/efa_admin_cmds_defs.h | 10 +++++++++-
 drivers/infiniband/hw/efa/efa_verbs.c           |  3 +++
 include/uapi/rdma/efa-abi.h                     |  1 +
 3 files changed, 13 insertions(+), 1 deletion(-)

diff --git a/drivers/infiniband/hw/efa/efa_admin_cmds_defs.h b/drivers/infiniband/hw/efa/efa_admin_cmds_defs.h
index ad34ea5da6b0..097b3303f3e9 100644
--- a/drivers/infiniband/hw/efa/efa_admin_cmds_defs.h
+++ b/drivers/infiniband/hw/efa/efa_admin_cmds_defs.h
@@ -725,7 +725,11 @@ struct efa_admin_feature_device_attr_desc {
 	 *    on TX queues
 	 * 4 : unsolicited_write_recv - If set, unsolicited
 	 *    write with imm. receive is supported
-	 * 31:5 : reserved - MBZ
+	 * 5 : event_counters - If set, event counters are
+	 *    supported
+	 * 6 : p2p_dma - If set the device can DMA directly
+	 *    to and from a peer PCIe device
+	 * 31:7 : reserved - MBZ
 	 */
 	u32 device_caps;
 
@@ -1132,6 +1136,10 @@ struct efa_admin_host_info {
 #define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_DATA_POLLING_128_MASK BIT(2)
 #define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_RDMA_WRITE_MASK  BIT(3)
 #define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_UNSOLICITED_WRITE_RECV_MASK BIT(4)
+#define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_EVENT_COUNTERS_SHIFT 5
+#define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_EVENT_COUNTERS_MASK BIT(5)
+#define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_P2P_DMA_SHIFT    6
+#define EFA_ADMIN_FEATURE_DEVICE_ATTR_DESC_P2P_DMA_MASK     BIT(6)
 
 /* create_eq_cmd */
 #define EFA_ADMIN_CREATE_EQ_CMD_ENTRY_SIZE_WORDS_MASK       GENMASK(4, 0)
diff --git a/drivers/infiniband/hw/efa/efa_verbs.c b/drivers/infiniband/hw/efa/efa_verbs.c
index 7bd0838ebc99..b16f470f7d30 100644
--- a/drivers/infiniband/hw/efa/efa_verbs.c
+++ b/drivers/infiniband/hw/efa/efa_verbs.c
@@ -270,6 +270,9 @@ int efa_query_device(struct ib_device *ibdev,
 		if (EFA_DEV_CAP(dev, UNSOLICITED_WRITE_RECV))
 			resp.device_caps |= EFA_QUERY_DEVICE_CAPS_UNSOLICITED_WRITE_RECV;
 
+		if (EFA_DEV_CAP(dev, P2P_DMA))
+			resp.device_caps |= EFA_QUERY_DEVICE_CAPS_P2P_DMA;
+
 		if (dev->neqs)
 			resp.device_caps |= EFA_QUERY_DEVICE_CAPS_CQ_NOTIFICATIONS;
 
diff --git a/include/uapi/rdma/efa-abi.h b/include/uapi/rdma/efa-abi.h
index d5c18f8de182..d19cb59d822d 100644
--- a/include/uapi/rdma/efa-abi.h
+++ b/include/uapi/rdma/efa-abi.h
@@ -133,6 +133,7 @@ enum {
 	EFA_QUERY_DEVICE_CAPS_RDMA_WRITE = 1 << 5,
 	EFA_QUERY_DEVICE_CAPS_UNSOLICITED_WRITE_RECV = 1 << 6,
 	EFA_QUERY_DEVICE_CAPS_CQ_WITH_EXT_MEM = 1 << 7,
+	EFA_QUERY_DEVICE_CAPS_P2P_DMA = 1 << 8,
 };
 
 struct efa_ibv_ex_query_device_resp {
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-05-03 15:02 [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query Yonatan Nachum
@ 2026-05-04  7:34 ` Jason Gunthorpe
  2026-05-05  8:15   ` Yonatan Nachum
  0 siblings, 1 reply; 14+ messages in thread
From: Jason Gunthorpe @ 2026-05-04  7:34 UTC (permalink / raw)
  To: Yonatan Nachum
  Cc: leon, linux-rdma, mrgolin, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Sun, May 03, 2026 at 03:02:46PM +0000, Yonatan Nachum wrote:
> Expose device P2P DMA support using the query device verbs.
> If the device support P2P DMA, it can DMA directly to and from a peer
> PCIe device

This doesn't seem right, this should be policed by failing to
established p2p mappings and to fail mapping dmabufs not with random
user space bits like this.

There are lots of things in our system that need this feedback to go
down that path.

Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-05-04  7:34 ` Jason Gunthorpe
@ 2026-05-05  8:15   ` Yonatan Nachum
  2026-05-05 15:40     ` Jason Gunthorpe
  0 siblings, 1 reply; 14+ messages in thread
From: Yonatan Nachum @ 2026-05-05  8:15 UTC (permalink / raw)
  To: Jason Gunthorpe
  Cc: leon, linux-rdma, mrgolin, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Mon, May 04, 2026 at 04:34:12AM -0300, Jason Gunthorpe wrote:
> On Sun, May 03, 2026 at 03:02:46PM +0000, Yonatan Nachum wrote:
> > Expose device P2P DMA support using the query device verbs.
> > If the device support P2P DMA, it can DMA directly to and from a peer
> > PCIe device
> 
> This doesn't seem right, this should be policed by failing to
> established p2p mappings and to fail mapping dmabufs not with random
> user space bits like this.

The motivation here is to avoid requiring userspace to speculatively
attempt registering accelerator MRs just to discover whether the device
supports P2P DMA. Beyond the API awkwardness, this also has a real
performance impact — initializing an accelerator context can take
seconds, and by advertising this as a device capability, userspace can
know upfront whether it's worth going down that path.

If I understand your suggestion correctly, you'd prefer to enforce this
in the reg MR path itself rather than exposing a capability bit. I see
two issues with that approach:
1. Userspace would still need to speculatively attempt a reg MR to
discover P2P support.

2. When we get a dmabuf, I can't distinguish what is the backing memory,
for example DRAM or accelerator memory and we can't just blindly fail
all cases.

Can you please clarify how you imagine this capbility to be used ?

> There are lots of things in our system that need this feedback to go
> down that path.

Can you please clarify what you mean here ?

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-05-05  8:15   ` Yonatan Nachum
@ 2026-05-05 15:40     ` Jason Gunthorpe
  2026-09-23 16:38       ` Michael Margolin
  0 siblings, 1 reply; 14+ messages in thread
From: Jason Gunthorpe @ 2026-05-05 15:40 UTC (permalink / raw)
  To: Yonatan Nachum
  Cc: leon, linux-rdma, mrgolin, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Tue, May 05, 2026 at 08:15:14AM +0000, Yonatan Nachum wrote:
> On Mon, May 04, 2026 at 04:34:12AM -0300, Jason Gunthorpe wrote:
> > On Sun, May 03, 2026 at 03:02:46PM +0000, Yonatan Nachum wrote:
> > > Expose device P2P DMA support using the query device verbs.
> > > If the device support P2P DMA, it can DMA directly to and from a peer
> > > PCIe device
> > 
> > This doesn't seem right, this should be policed by failing to
> > established p2p mappings and to fail mapping dmabufs not with random
> > user space bits like this.
> 
> The motivation here is to avoid requiring userspace to speculatively
> attempt registering accelerator MRs just to discover whether the device
> supports P2P DMA.

This isn't how the kernel works there isn't a "this device can do
p2p", the question is always answered with pairs of devices. There is
no way for a single device to know it doesn't do p2p with any other
device in the system.

> Beyond the API awkwardness, this also has a real
> performance impact — initializing an accelerator context can take
> seconds, and by advertising this as a device capability, userspace can
> know upfront whether it's worth going down that path.

This seems like a misconfiguration. Since so much of this topologcal
information is not discoverable at run time usually systems have an
external data file explaining how to best use the system. I wouldn't
expect the runtime to be trying to probe this at run time.

If you *really* want probing then it has to be dmabuf based because
you must probe with pairs of PCI devices.

> 2. When we get a dmabuf, I can't distinguish what is the backing memory,
> for example DRAM or accelerator memory and we can't just blindly fail
> all cases.

You must rely on the p2p subsystem to make this determination and the
dmabuf exporter has to call into it. This will deal with those
problems.

Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-05-05 15:40     ` Jason Gunthorpe
@ 2026-09-23 16:38       ` Michael Margolin
  2026-09-23 17:17         ` Jason Gunthorpe
  0 siblings, 1 reply; 14+ messages in thread
From: Michael Margolin @ 2026-09-23 16:38 UTC (permalink / raw)
  To: Jason Gunthorpe
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Tue, May 05, 2026 at 12:40:10PM -0300, Jason Gunthorpe wrote:
> On Tue, May 05, 2026 at 08:15:14AM +0000, Yonatan Nachum wrote:
> > On Mon, May 04, 2026 at 04:34:12AM -0300, Jason Gunthorpe wrote:
> > > On Sun, May 03, 2026 at 03:02:46PM +0000, Yonatan Nachum wrote:
> > > > Expose device P2P DMA support using the query device verbs.
> > > > If the device support P2P DMA, it can DMA directly to and from a peer
> > > > PCIe device
> > > 
> > > This doesn't seem right, this should be policed by failing to
> > > established p2p mappings and to fail mapping dmabufs not with random
> > > user space bits like this.
> > 
> > The motivation here is to avoid requiring userspace to speculatively
> > attempt registering accelerator MRs just to discover whether the device
> > supports P2P DMA.
> 
> This isn't how the kernel works there isn't a "this device can do
> p2p", the question is always answered with pairs of devices. There is
> no way for a single device to know it doesn't do p2p with any other
> device in the system.
> 
> > Beyond the API awkwardness, this also has a real
> > performance impact — initializing an accelerator context can take
> > seconds, and by advertising this as a device capability, userspace can
> > know upfront whether it's worth going down that path.
> 
> This seems like a misconfiguration. Since so much of this topologcal
> information is not discoverable at run time usually systems have an
> external data file explaining how to best use the system. I wouldn't
> expect the runtime to be trying to probe this at run time.
> 
> If you *really* want probing then it has to be dmabuf based because
> you must probe with pairs of PCI devices.
> 
> > 2. When we get a dmabuf, I can't distinguish what is the backing memory,
> > for example DRAM or accelerator memory and we can't just blindly fail
> > all cases.
> 
> You must rely on the p2p subsystem to make this determination and the
> dmabuf exporter has to call into it. This will deal with those
> problems.
> 
> Jason

Jason,

I'm reviving this conversation since it is becoming an
increasing need to reduce the long startup times. The major
part of this time is related to CUDA device initialization
needed only to allocate GPU memory, export a dmabuf, and
test accessibility via MR registration. Since obtaining a
dmabuf fd requires CUDA initialization, which is exactly
the cost we are trying to avoid, we are looking for a
solution that doesn't require dmabuf.

I have two options in mind:

1. Introduce a new (common or EFA-specific) verb that takes
   a peer device BDF and queries the p2p subsystem,
   returning something like:

   pci_p2pdma_distance(peer, self, false) > 0 &&
   EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)

2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
   directly to userspace. As you already stated, this is a
   bit awkward as a device isn't expected to be aware of
   system topology, but since EFA lives in the cloud it is
   partially aware of the platform configuration and in
   particular the ability to access peer devices over PCIe.

I would like to hear your perspective on this.

Michael


^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-23 16:38       ` Michael Margolin
@ 2026-09-23 17:17         ` Jason Gunthorpe
  2026-09-24 15:29           ` Michael Margolin
  0 siblings, 1 reply; 14+ messages in thread
From: Jason Gunthorpe @ 2026-09-23 17:17 UTC (permalink / raw)
  To: Michael Margolin
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Wed, Sep 23, 2026 at 04:38:30PM +0000, Michael Margolin wrote:

> I'm reviving this conversation since it is becoming an
> increasing need to reduce the long startup times. The major
> part of this time is related to CUDA device initialization
> needed only to allocate GPU memory, export a dmabuf, and
> test accessibility via MR registration. Since obtaining a
> dmabuf fd requires CUDA initialization, which is exactly
> the cost we are trying to avoid, we are looking for a
> solution that doesn't require dmabuf.

You have to use dmabuf. Go complain to the cuda people their stuff is
too slow if that is the actual problem.

Precompute it and put it in a topology file like everyone else does :\

> I have two options in mind:
> 
> 1. Introduce a new (common or EFA-specific) verb that takes
>    a peer device BDF and queries the p2p subsystem,
>    returning something like:

No way. No BDFs in uapis.

> 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
>    directly to userspace. As you already stated, this is a
>    bit awkward as a device isn't expected to be aware of
>    system topology, but since EFA lives in the cloud it is
>    partially aware of the platform configuration and in
>    particular the ability to access peer devices over PCIe.

If you already know your VMs are always safe what is even the issue?
What are you probing for?

Lets please not hack around bad userspace with even worse kernel uAPI.

Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-23 17:17         ` Jason Gunthorpe
@ 2026-09-24 15:29           ` Michael Margolin
  2026-09-24 18:22             ` Jason Gunthorpe
  0 siblings, 1 reply; 14+ messages in thread
From: Michael Margolin @ 2026-09-24 15:29 UTC (permalink / raw)
  To: Jason Gunthorpe
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Wed, Sep 23, 2026 at 02:17:41PM -0300, Jason Gunthorpe wrote:
> On Wed, Sep 23, 2026 at 04:38:30PM +0000, Michael Margolin wrote:
> 
> > I'm reviving this conversation since it is becoming an
> > increasing need to reduce the long startup times. The major
> > part of this time is related to CUDA device initialization
> > needed only to allocate GPU memory, export a dmabuf, and
> > test accessibility via MR registration. Since obtaining a
> > dmabuf fd requires CUDA initialization, which is exactly
> > the cost we are trying to avoid, we are looking for a
> > solution that doesn't require dmabuf.
> 
> You have to use dmabuf. Go complain to the cuda people their stuff is
> too slow if that is the actual problem.
> 
> Precompute it and put it in a topology file like everyone else does :\
> 
> > I have two options in mind:
> > 
> > 1. Introduce a new (common or EFA-specific) verb that takes
> >    a peer device BDF and queries the p2p subsystem,
> >    returning something like:
> 
> No way. No BDFs in uapis.
> 
> > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> >    directly to userspace. As you already stated, this is a
> >    bit awkward as a device isn't expected to be aware of
> >    system topology, but since EFA lives in the cloud it is
> >    partially aware of the platform configuration and in
> >    particular the ability to access peer devices over PCIe.
> 
> If you already know your VMs are always safe what is even the issue?
> What are you probing for?
> 
> Lets please not hack around bad userspace with even worse kernel uAPI.

The EFA device already knows this attribute. The problem is that this
information is not exposed to userspace today.

So if I go back to our original proposal - exposing the
EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
attribute that says "this device does not block PCIe peer access on its
end." It's not a topology claim, it's a device property, similar to how
we already expose RDMA read/write support.

The actual reachability is still a combination of multiple factors.

Why is this a bad uAPI?

Michael

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-24 15:29           ` Michael Margolin
@ 2026-09-24 18:22             ` Jason Gunthorpe
  2026-09-24 19:30               ` Michael Margolin
  0 siblings, 1 reply; 14+ messages in thread
From: Jason Gunthorpe @ 2026-09-24 18:22 UTC (permalink / raw)
  To: Michael Margolin
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:

> > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > >    directly to userspace. As you already stated, this is a
> > >    bit awkward as a device isn't expected to be aware of
> > >    system topology, but since EFA lives in the cloud it is
> > >    partially aware of the platform configuration and in
> > >    particular the ability to access peer devices over PCIe.
> > 
> > If you already know your VMs are always safe what is even the issue?
> > What are you probing for?
> > 
> > Lets please not hack around bad userspace with even worse kernel uAPI.
> 
> The EFA device already knows this attribute. The problem is that this
> information is not exposed to userspace today.
> 
> So if I go back to our original proposal - exposing the
> EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> attribute that says "this device does not block PCIe peer access on its
> end." It's not a topology claim, it's a device property, similar to how
> we already expose RDMA read/write support.

And how does that help you if you actually need reachability? You
haven't said directly what you are even trying to achieve (though I
have a pretty good guess).

If all your VMs can do P2P then what is the issue? If you need
pairwise reachability then you have to use dmabuf.

Precompute it and use a topology file like everyone else?

Inject a topology file into your VMs, like some others?

Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-24 18:22             ` Jason Gunthorpe
@ 2026-09-24 19:30               ` Michael Margolin
  2026-09-25 12:55                 ` Jason Gunthorpe
  0 siblings, 1 reply; 14+ messages in thread
From: Michael Margolin @ 2026-09-24 19:30 UTC (permalink / raw)
  To: Jason Gunthorpe
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Thu, Sep 24, 2026 at 03:22:01PM -0300, Jason Gunthorpe wrote:
> On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:
> 
> > > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > > >    directly to userspace. As you already stated, this is a
> > > >    bit awkward as a device isn't expected to be aware of
> > > >    system topology, but since EFA lives in the cloud it is
> > > >    partially aware of the platform configuration and in
> > > >    particular the ability to access peer devices over PCIe.
> > > 
> > > If you already know your VMs are always safe what is even the issue?
> > > What are you probing for?
> > > 
> > > Lets please not hack around bad userspace with even worse kernel uAPI.
> > 
> > The EFA device already knows this attribute. The problem is that this
> > information is not exposed to userspace today.
> > 
> > So if I go back to our original proposal - exposing the
> > EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> > attribute that says "this device does not block PCIe peer access on its
> > end." It's not a topology claim, it's a device property, similar to how
> > we already expose RDMA read/write support.
> 
> And how does that help you if you actually need reachability? You
> haven't said directly what you are even trying to achieve (though I
> have a pretty good guess).

On some platforms P2P is blocked at the device level despite the fact
that the peer GPU is reachable over PCIe. This mainly happens in
multi-instance setups where there is a need to manage PCIe bandwidth
or prevent noisy neighbor effects. Today we check this attribute on the
device during MR registration, so libfabric's EFA provider allocates
GPU memory and attempts to register an MR to get this information.
This process is costly, thus we are trying to replace it with a simple
capability check.

> 
> If all your VMs can do P2P then what is the issue? If you need
> pairwise reachability then you have to use dmabuf.

As I describe above, this is not the case.

> 
> Precompute it and use a topology file like everyone else?

This is an option but it seems like a big overkill for the simple check
that we need.

> 
> Inject a topology file into your VMs, like some others?
> 
> Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-24 19:30               ` Michael Margolin
@ 2026-09-25 12:55                 ` Jason Gunthorpe
  2026-09-30 13:26                   ` Michael Margolin
  0 siblings, 1 reply; 14+ messages in thread
From: Jason Gunthorpe @ 2026-09-25 12:55 UTC (permalink / raw)
  To: Michael Margolin
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Thu, Sep 24, 2026 at 07:30:26PM +0000, Michael Margolin wrote:
> On Thu, Sep 24, 2026 at 03:22:01PM -0300, Jason Gunthorpe wrote:
> > On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:
> > 
> > > > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > > > >    directly to userspace. As you already stated, this is a
> > > > >    bit awkward as a device isn't expected to be aware of
> > > > >    system topology, but since EFA lives in the cloud it is
> > > > >    partially aware of the platform configuration and in
> > > > >    particular the ability to access peer devices over PCIe.
> > > > 
> > > > If you already know your VMs are always safe what is even the issue?
> > > > What are you probing for?
> > > > 
> > > > Lets please not hack around bad userspace with even worse kernel uAPI.
> > > 
> > > The EFA device already knows this attribute. The problem is that this
> > > information is not exposed to userspace today.
> > > 
> > > So if I go back to our original proposal - exposing the
> > > EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> > > attribute that says "this device does not block PCIe peer access on its
> > > end." It's not a topology claim, it's a device property, similar to how
> > > we already expose RDMA read/write support.
> > 
> > And how does that help you if you actually need reachability? You
> > haven't said directly what you are even trying to achieve (though I
> > have a pretty good guess).
> 
> On some platforms P2P is blocked at the device level despite the fact
> that the peer GPU is reachable over PCIe.

This idea just doesn't exist in Linux.. It is completely inappropriate
for a PCI device to declare it knows something about the
topology. That's not how PCI works, it is the wrong place for your VM
to signal information about the PCI routing through the EFA driver.

That needs to come up to the P2P subsystem.

> This mainly happens in multi-instance setups where there is a need
> to manage PCIe bandwidth or prevent noisy neighbor effects. Today we
> check this attribute on the device during MR registration, so
> libfabric's EFA provider allocates GPU memory and attempts to
> register an MR to get this information.  This process is costly,
> thus we are trying to replace it with a simple capability check.

Which is the right thing to do, again if it is slow go talk to the
people who are making it slow..

> > Precompute it and use a topology file like everyone else?
> 
> This is an option but it seems like a big overkill for the simple
> check that we need.

Maybe there is some other place you can inject into the VM a 'this vm
does not support p2p' flag that userspace can see?

Leon has been working on some ACPI enhancements for this, perhaps that
new table could be distilled down to a single sysfs under an acpi
directory?

Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-25 12:55                 ` Jason Gunthorpe
@ 2026-09-30 13:26                   ` Michael Margolin
  2026-10-01 15:53                     ` Leon Romanovsky
  0 siblings, 1 reply; 14+ messages in thread
From: Michael Margolin @ 2026-09-30 13:26 UTC (permalink / raw)
  To: Jason Gunthorpe
  Cc: Yonatan Nachum, leon, linux-rdma, sleybo, matua, gal.pressman,
	Yehuda Yitschak

On Fri, Sep 25, 2026 at 09:55:12AM -0300, Jason Gunthorpe wrote:
> On Thu, Sep 24, 2026 at 07:30:26PM +0000, Michael Margolin wrote:
> > On Thu, Sep 24, 2026 at 03:22:01PM -0300, Jason Gunthorpe wrote:
> > > On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:
> > > 
> > > > > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > > > > >    directly to userspace. As you already stated, this is a
> > > > > >    bit awkward as a device isn't expected to be aware of
> > > > > >    system topology, but since EFA lives in the cloud it is
> > > > > >    partially aware of the platform configuration and in
> > > > > >    particular the ability to access peer devices over PCIe.
> > > > > 
> > > > > If you already know your VMs are always safe what is even the issue?
> > > > > What are you probing for?
> > > > > 
> > > > > Lets please not hack around bad userspace with even worse kernel uAPI.
> > > > 
> > > > The EFA device already knows this attribute. The problem is that this
> > > > information is not exposed to userspace today.
> > > > 
> > > > So if I go back to our original proposal - exposing the
> > > > EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> > > > attribute that says "this device does not block PCIe peer access on its
> > > > end." It's not a topology claim, it's a device property, similar to how
> > > > we already expose RDMA read/write support.
> > > 
> > > And how does that help you if you actually need reachability? You
> > > haven't said directly what you are even trying to achieve (though I
> > > have a pretty good guess).
> > 
> > On some platforms P2P is blocked at the device level despite the fact
> > that the peer GPU is reachable over PCIe.
> 
> This idea just doesn't exist in Linux.. It is completely inappropriate
> for a PCI device to declare it knows something about the
> topology. That's not how PCI works, it is the wrong place for your VM
> to signal information about the PCI routing through the EFA driver.
> 
> That needs to come up to the P2P subsystem.
> 
> > This mainly happens in multi-instance setups where there is a need
> > to manage PCIe bandwidth or prevent noisy neighbor effects. Today we
> > check this attribute on the device during MR registration, so
> > libfabric's EFA provider allocates GPU memory and attempts to
> > register an MR to get this information.  This process is costly,
> > thus we are trying to replace it with a simple capability check.
> 
> Which is the right thing to do, again if it is slow go talk to the
> people who are making it slow..
> 
> > > Precompute it and use a topology file like everyone else?
> > 
> > This is an option but it seems like a big overkill for the simple
> > check that we need.
> 
> Maybe there is some other place you can inject into the VM a 'this vm
> does not support p2p' flag that userspace can see?
> 
> Leon has been working on some ACPI enhancements for this, perhaps that
> new table could be distilled down to a single sysfs under an acpi
> directory?

Looked at Leon's patches and it does seem like the right place to expose
such peer relations. Although it wouldn't help with our immediate need,
maybe we can add a sysfs interface on top of Leon's patches to expose
the gathered accessiblity (and maybe also performance) information.

Leon, what is your opinion on this?

Michael

> 
> Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-09-30 13:26                   ` Michael Margolin
@ 2026-10-01 15:53                     ` Leon Romanovsky
  2026-10-02 14:52                       ` Jason Gunthorpe
  0 siblings, 1 reply; 14+ messages in thread
From: Leon Romanovsky @ 2026-10-01 15:53 UTC (permalink / raw)
  To: Michael Margolin
  Cc: Jason Gunthorpe, Yonatan Nachum, linux-rdma, sleybo, matua,
	gal.pressman, Yehuda Yitschak

On Wed, Sep 30, 2026 at 01:26:13PM +0000, Michael Margolin wrote:
> On Fri, Sep 25, 2026 at 09:55:12AM -0300, Jason Gunthorpe wrote:
> > On Thu, Sep 24, 2026 at 07:30:26PM +0000, Michael Margolin wrote:
> > > On Thu, Sep 24, 2026 at 03:22:01PM -0300, Jason Gunthorpe wrote:
> > > > On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:
> > > > 
> > > > > > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > > > > > >    directly to userspace. As you already stated, this is a
> > > > > > >    bit awkward as a device isn't expected to be aware of
> > > > > > >    system topology, but since EFA lives in the cloud it is
> > > > > > >    partially aware of the platform configuration and in
> > > > > > >    particular the ability to access peer devices over PCIe.
> > > > > > 
> > > > > > If you already know your VMs are always safe what is even the issue?
> > > > > > What are you probing for?
> > > > > > 
> > > > > > Lets please not hack around bad userspace with even worse kernel uAPI.
> > > > > 
> > > > > The EFA device already knows this attribute. The problem is that this
> > > > > information is not exposed to userspace today.
> > > > > 
> > > > > So if I go back to our original proposal - exposing the
> > > > > EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> > > > > attribute that says "this device does not block PCIe peer access on its
> > > > > end." It's not a topology claim, it's a device property, similar to how
> > > > > we already expose RDMA read/write support.
> > > > 
> > > > And how does that help you if you actually need reachability? You
> > > > haven't said directly what you are even trying to achieve (though I
> > > > have a pretty good guess).
> > > 
> > > On some platforms P2P is blocked at the device level despite the fact
> > > that the peer GPU is reachable over PCIe.
> > 
> > This idea just doesn't exist in Linux.. It is completely inappropriate
> > for a PCI device to declare it knows something about the
> > topology. That's not how PCI works, it is the wrong place for your VM
> > to signal information about the PCI routing through the EFA driver.
> > 
> > That needs to come up to the P2P subsystem.
> > 
> > > This mainly happens in multi-instance setups where there is a need
> > > to manage PCIe bandwidth or prevent noisy neighbor effects. Today we
> > > check this attribute on the device during MR registration, so
> > > libfabric's EFA provider allocates GPU memory and attempts to
> > > register an MR to get this information.  This process is costly,
> > > thus we are trying to replace it with a simple capability check.
> > 
> > Which is the right thing to do, again if it is slow go talk to the
> > people who are making it slow..
> > 
> > > > Precompute it and use a topology file like everyone else?
> > > 
> > > This is an option but it seems like a big overkill for the simple
> > > check that we need.
> > 
> > Maybe there is some other place you can inject into the VM a 'this vm
> > does not support p2p' flag that userspace can see?
> > 
> > Leon has been working on some ACPI enhancements for this, perhaps that
> > new table could be distilled down to a single sysfs under an acpi
> > directory?
> 
> Looked at Leon's patches and it does seem like the right place to expose
> such peer relations. Although it wouldn't help with our immediate need,
> maybe we can add a sysfs interface on top of Leon's patches to expose
> the gathered accessiblity (and maybe also performance) information.
> 
> Leon, what is your opinion on this?

ACPI tables are available here /sys/firmware/acpi/tables/, this will
include our new p2p HMAT extension.

Thanks

> 
> Michael
> 
> > 
> > Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-10-01 15:53                     ` Leon Romanovsky
@ 2026-10-02 14:52                       ` Jason Gunthorpe
  2026-10-07  9:01                         ` Michael Margolin
  0 siblings, 1 reply; 14+ messages in thread
From: Jason Gunthorpe @ 2026-10-02 14:52 UTC (permalink / raw)
  To: Leon Romanovsky
  Cc: Michael Margolin, Yonatan Nachum, linux-rdma, sleybo, matua,
	gal.pressman, Yehuda Yitschak

On Thu, Oct 01, 2026 at 06:53:33PM +0300, Leon Romanovsky wrote:

> > Looked at Leon's patches and it does seem like the right place to expose
> > such peer relations. Although it wouldn't help with our immediate need,
> > maybe we can add a sysfs interface on top of Leon's patches to expose
> > the gathered accessiblity (and maybe also performance) information.
> > 
> > Leon, what is your opinion on this?
> 
> ACPI tables are available here /sys/firmware/acpi/tables/, this will
> include our new p2p HMAT extension.

ACPI is nice and general, but..

Doesn't AWS have something already to convay VM instance type
configuration information? That's pretty normal these days isn't it?

Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
  2026-10-02 14:52                       ` Jason Gunthorpe
@ 2026-10-07  9:01                         ` Michael Margolin
  0 siblings, 0 replies; 14+ messages in thread
From: Michael Margolin @ 2026-10-07  9:01 UTC (permalink / raw)
  To: Jason Gunthorpe
  Cc: Leon Romanovsky, Yonatan Nachum, linux-rdma, sleybo, matua,
	gal.pressman, Yehuda Yitschak

On Fri, Oct 02, 2026 at 11:52:41AM -0300, Jason Gunthorpe wrote:
> On Thu, Oct 01, 2026 at 06:53:33PM +0300, Leon Romanovsky wrote:
> 
> > > Looked at Leon's patches and it does seem like the right place to expose
> > > such peer relations. Although it wouldn't help with our immediate need,
> > > maybe we can add a sysfs interface on top of Leon's patches to expose
> > > the gathered accessiblity (and maybe also performance) information.
> > > 
> > > Leon, what is your opinion on this?
> > 
> > ACPI tables are available here /sys/firmware/acpi/tables/, this will
> > include our new p2p HMAT extension.

I was iterating over Jason's idea and what I imagine is an interface
that contains already parsed data at PCI level, like:

/sys/firmware/acpi/p2p/
├── 0000:55:00.0/          # RDMA NIC
│   ├── 0000:59:00.0/      # → GPU (reachable peer)
│   │   ├── latency_ns     # 200
│   │   └── bandwidth_mbs  # 32000
│   └── 0000:5a:00.0/      # → GPU (reachable peer)
│       ├── latency_ns     # 200
│       └── bandwidth_mbs  # 32000
│
├── 0000:59:00.0/          # GPU
│   ├── 0000:55:00.0/      # → RDMA NIC (reachable peer)
│   │   ├── latency_ns     # 200
│   │   └── bandwidth_mbs  # 32000
│   └── 0000:56:00.0/      # → RDMA NIC (reachable peer)
│       ├── latency_ns     # 200
│       └── bandwidth_mbs  # 32000
│
├── 0000:9a:00.0/          # NVMe
│   └── 0000:55:00.0/      # → RDMA NIC (reachable peer)
│       ├── latency_ns     # 300
│       └── bandwidth_mbs  # 12000

> 
> ACPI is nice and general, but..
> 
> Doesn't AWS have something already to convay VM instance type
> configuration information? That's pretty normal these days isn't it?

There is but both options doesn't help in the short term and long term I
think, as you mentioned, that ACPI is a more general solution and it
doesn't require any additional dependencies.

Michael

> Jason

^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-10-07  9:01 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-05-03 15:02 [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query Yonatan Nachum
2026-05-04  7:34 ` Jason Gunthorpe
2026-05-05  8:15   ` Yonatan Nachum
2026-05-05 15:40     ` Jason Gunthorpe
2026-09-23 16:38       ` Michael Margolin
2026-09-23 17:17         ` Jason Gunthorpe
2026-09-24 15:29           ` Michael Margolin
2026-09-24 18:22             ` Jason Gunthorpe
2026-09-24 19:30               ` Michael Margolin
2026-09-25 12:55                 ` Jason Gunthorpe
2026-09-30 13:26                   ` Michael Margolin
2026-10-01 15:53                     ` Leon Romanovsky
2026-10-02 14:52                       ` Jason Gunthorpe
2026-10-07  9:01                         ` Michael Margolin

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox