Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: Jason Gunthorpe <jgg@nvidia.com>
To: Michael Margolin <mrgolin@amazon.com>
Cc: Yonatan Nachum <ynachum@amazon.com>,
	leon@kernel.org, linux-rdma@vger.kernel.org, sleybo@amazon.com,
	matua@amazon.com, gal.pressman@linux.dev,
	Yehuda Yitschak <yehuday@amazon.com>
Subject: Re: [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query
Date: Fri, 25 Sep 2026 09:55:12 -0300	[thread overview]
Message-ID: <20260925125512.GM9354@nvidia.com> (raw)
In-Reply-To: <20260924193026.GA13498@dev-dsk-mrgolin-1c-b2091117.eu-west-1.amazon.com>

On Thu, Sep 24, 2026 at 07:30:26PM +0000, Michael Margolin wrote:
> On Thu, Sep 24, 2026 at 03:22:01PM -0300, Jason Gunthorpe wrote:
> > On Thu, Sep 24, 2026 at 03:29:13PM +0000, Michael Margolin wrote:
> > 
> > > > > 2. Expose EFA_DEV_CAP(dev, PCIE_PEER_ACCESS_ENABLED)
> > > > >    directly to userspace. As you already stated, this is a
> > > > >    bit awkward as a device isn't expected to be aware of
> > > > >    system topology, but since EFA lives in the cloud it is
> > > > >    partially aware of the platform configuration and in
> > > > >    particular the ability to access peer devices over PCIe.
> > > > 
> > > > If you already know your VMs are always safe what is even the issue?
> > > > What are you probing for?
> > > > 
> > > > Lets please not hack around bad userspace with even worse kernel uAPI.
> > > 
> > > The EFA device already knows this attribute. The problem is that this
> > > information is not exposed to userspace today.
> > > 
> > > So if I go back to our original proposal - exposing the
> > > EFA_DEV_CAP(dev, P2P_DMA) bit through query device - this is a firmware
> > > attribute that says "this device does not block PCIe peer access on its
> > > end." It's not a topology claim, it's a device property, similar to how
> > > we already expose RDMA read/write support.
> > 
> > And how does that help you if you actually need reachability? You
> > haven't said directly what you are even trying to achieve (though I
> > have a pretty good guess).
> 
> On some platforms P2P is blocked at the device level despite the fact
> that the peer GPU is reachable over PCIe.

This idea just doesn't exist in Linux.. It is completely inappropriate
for a PCI device to declare it knows something about the
topology. That's not how PCI works, it is the wrong place for your VM
to signal information about the PCI routing through the EFA driver.

That needs to come up to the P2P subsystem.

> This mainly happens in multi-instance setups where there is a need
> to manage PCIe bandwidth or prevent noisy neighbor effects. Today we
> check this attribute on the device during MR registration, so
> libfabric's EFA provider allocates GPU memory and attempts to
> register an MR to get this information.  This process is costly,
> thus we are trying to replace it with a simple capability check.

Which is the right thing to do, again if it is slow go talk to the
people who are making it slow..

> > Precompute it and use a topology file like everyone else?
> 
> This is an option but it seems like a big overkill for the simple
> check that we need.

Maybe there is some other place you can inject into the VM a 'this vm
does not support p2p' flag that userspace can see?

Leon has been working on some ACPI enhancements for this, perhaps that
new table could be distilled down to a single sysfs under an acpi
directory?

Jason

  reply	other threads:[~2026-09-25 12:55 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-05-03 15:02 [PATCH for-next] RDMA/efa: Expose device P2P DMA support via device query Yonatan Nachum
2026-05-04  7:34 ` Jason Gunthorpe
2026-05-05  8:15   ` Yonatan Nachum
2026-05-05 15:40     ` Jason Gunthorpe
2026-09-23 16:38       ` Michael Margolin
2026-09-23 17:17         ` Jason Gunthorpe
2026-09-24 15:29           ` Michael Margolin
2026-09-24 18:22             ` Jason Gunthorpe
2026-09-24 19:30               ` Michael Margolin
2026-09-25 12:55                 ` Jason Gunthorpe [this message]
2026-09-30 13:26                   ` Michael Margolin
2026-10-01 15:53                     ` Leon Romanovsky
2026-10-02 14:52                       ` Jason Gunthorpe
2026-10-07  9:01                         ` Michael Margolin

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260925125512.GM9354@nvidia.com \
    --to=jgg@nvidia.com \
    --cc=gal.pressman@linux.dev \
    --cc=leon@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=matua@amazon.com \
    --cc=mrgolin@amazon.com \
    --cc=sleybo@amazon.com \
    --cc=yehuday@amazon.com \
    --cc=ynachum@amazon.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox