From: Leon Romanovsky <leon@kernel.org>
To: Michael Margolin <mrgolin@amazon.com>
Cc: Jason Gunthorpe <jgg@nvidia.com>,
Yonatan Nachum <ynachum@amazon.com>,
linux-rdma@vger.kernel.org, sleybo@amazon.com, matua@amazon.com,
gal.pressman@linux.dev, Yehuda Yitschak <yehuday@amazon.com>
Subject: Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
Date: Thu, 1 Oct 2026 18:53:33 +0300 [thread overview]
Message-ID: <20261001155333.GP3401365@unreal> (raw)
In-Reply-To: <20260930132613.GA24387@dev-dsk-mrgolin-1c-b2091117.eu-west-1.amazon.com>
On Wed, Sep 30, 2026 at 01:26:13PM +0000, Michael Margolin wrote:
> On Fri, Sep 25, 2026 at 09:55:12AM -0300, Jason Gunthorpe wrote:
> > On Thu, Sep 24, 2026 at 07:30:26PM +0000, Michael Margolin wrote:
> > > On Thu, Sep 24, 2026 at 03:22:01PM -0300, Jason Gunthorpe wrote:
> > > > On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:
> > > >
> > > > > > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > > > > > > directly to userspace. As you already stated, this is a
> > > > > > > bit awkward as a device isn't expected to be aware of
> > > > > > > system topology, but since EFA lives in the cloud it is
> > > > > > > partially aware of the platform configuration and in
> > > > > > > particular the ability to access peer devices over PCIe.
> > > > > >
> > > > > > If you already know your VMs are always safe what is even the issue?
> > > > > > What are you probing for?
> > > > > >
> > > > > > Lets please not hack around bad userspace with even worse kernel uAPI.
> > > > >
> > > > > The EFA device already knows this attribute. The problem is that this
> > > > > information is not exposed to userspace today.
> > > > >
> > > > > So if I go back to our original proposal - exposing the
> > > > > EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> > > > > attribute that says "this device does not block PCIe peer access on its
> > > > > end." It's not a topology claim, it's a device property, similar to how
> > > > > we already expose RDMA read/write support.
> > > >
> > > > And how does that help you if you actually need reachability? You
> > > > haven't said directly what you are even trying to achieve (though I
> > > > have a pretty good guess).
> > >
> > > On some platforms P2P is blocked at the device level despite the fact
> > > that the peer GPU is reachable over PCIe.
> >
> > This idea just doesn't exist in Linux.. It is completely inappropriate
> > for a PCI device to declare it knows something about the
> > topology. That's not how PCI works, it is the wrong place for your VM
> > to signal information about the PCI routing through the EFA driver.
> >
> > That needs to come up to the P2P subsystem.
> >
> > > This mainly happens in multi-instance setups where there is a need
> > > to manage PCIe bandwidth or prevent noisy neighbor effects. Today we
> > > check this attribute on the device during MR registration, so
> > > libfabric's EFA provider allocates GPU memory and attempts to
> > > register an MR to get this information. This process is costly,
> > > thus we are trying to replace it with a simple capability check.
> >
> > Which is the right thing to do, again if it is slow go talk to the
> > people who are making it slow..
> >
> > > > Precompute it and use a topology file like everyone else?
> > >
> > > This is an option but it seems like a big overkill for the simple
> > > check that we need.
> >
> > Maybe there is some other place you can inject into the VM a 'this vm
> > does not support p2p' flag that userspace can see?
> >
> > Leon has been working on some ACPI enhancements for this, perhaps that
> > new table could be distilled down to a single sysfs under an acpi
> > directory?
>
> Looked at Leon's patches and it does seem like the right place to expose
> such peer relations. Although it wouldn't help with our immediate need,
> maybe we can add a sysfs interface on top of Leon's patches to expose
> the gathered accessiblity (and maybe also performance) information.
>
> Leon, what is your opinion on this?
ACPI tables are available here /sys/firmware/acpi/tables/, this will
include our new p2p HMAT extension.
Thanks
>
> Michael
>
> >
> > Jason
next prev parent reply other threads:[~2026-10-01 15:53 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-05-03 15:02 [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query Yonatan Nachum
2026-05-04 7:34 ` Jason Gunthorpe
2026-05-05 8:15 ` Yonatan Nachum
2026-05-05 15:40 ` Jason Gunthorpe
2026-09-23 16:38 ` Michael Margolin
2026-09-23 17:17 ` Jason Gunthorpe
2026-09-24 15:29 ` Michael Margolin
2026-09-24 18:22 ` Jason Gunthorpe
2026-09-24 19:30 ` Michael Margolin
2026-09-25 12:55 ` Jason Gunthorpe
2026-09-30 13:26 ` Michael Margolin
2026-10-01 15:53 ` Leon Romanovsky [this message]
2026-10-02 14:52 ` Jason Gunthorpe
2026-10-07 9:01 ` Michael Margolin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001155333.GP3401365@unreal \
--to=leon@kernel.org \
--cc=gal.pressman@linux.dev \
--cc=jgg@nvidia.com \
--cc=linux-rdma@vger.kernel.org \
--cc=matua@amazon.com \
--cc=mrgolin@amazon.com \
--cc=sleybo@amazon.com \
--cc=yehuday@amazon.com \
--cc=ynachum@amazon.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox