* [PATCH V2 0/3] Add UALink infrastructure series 2
@ 2026-08-31 18:33 Alex Deucher
2026-08-31 18:33 ` [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 Alex Deucher
` (2 more replies)
0 siblings, 3 replies; 12+ messages in thread
From: Alex Deucher @ 2026-08-31 18:33 UTC (permalink / raw)
To: amd-gfx, dri-devel, Felix.Kuehling; +Cc: Alex Deucher
This adds the core infrastructure for supporting UALink (Ultra Accelerator Link)
connected scale up pods. I've split this into two series, one to add the core
infrastructure, and one to expose the new IOCTL interface and add the
documentation to avoid spamming the larger audience with the implemenation
defails. This is the second series.
The first series can be found here:
V2: https://lists.freedesktop.org/archives/amd-gfx/2026-August/151914.html
V1: https://lists.freedesktop.org/archives/amd-gfx/2026-August/151200.html
Both series can be accessed via this git branch:
https://gitlab.freedesktop.org/agd5f/linux/-/commits/ualink?ref_type=heads
This implements UALoE (UALink over Ethernet). An overview of the complete
solution can be found here:
https://www.amd.com/en/products/rackscale-solutions/helios.html
Overview
Connected GPUs in a pod can directly access the remove memory on another
GPU over UALink. Unlike RMDA, there is no copy involved; it is direct
loads/stores over the fabric. Shared memory can only be accessed by
a remote GPU if the memory was exported and the importer has been authorized.
For the memory to be shared, it must be part of a unified physical
address space shared between nodes. This address space is called NPA (Nework
Physical Address) space. This address space is partitioned between
the GPUs so that each GPU has it's own segment of the address space in which
to export its memory. Each GPU maintains a dedicated set of page tables
for their NPA space similar to GPUVM.
Exported memory is pinned. This is similar to how dma-bufs in VRAM are pinned
for P2P access. Remote TLB shootdowns from the exporter are used when the
exported memory is freed in order to remove access by remote GPUs
To access remote memory, the driver can map NPA addresses into its per
process GPUVM page tables just like local memory. Applications use
opaque handles to represent remote memory. GPUs in a pod communicate
with eachother directly to exchange NPA addresses between importers
and exporters. If a node goes offline or is reset, their peers will
clean up any remaining refrences that are lost when that happens.
User interface
Export Memory
To export memory, a handle must be created for an allocation
that can be shared with another node in the pod. To do this
the exporter calls the GEM UALink IOCTL with the GEM handle
to the buffer it wants to export. The IOCTL returns a
unique handle which can be shared with the remote host.
Import Memory
To import the memory, the handle from the remote node must be converted
from a unique handle to a local GEM object which represents the local reference
to the NPA space on the importer. If the memory has already been
imported, it just returns a new reference to the existing object. If not,
the importer queries the exporter to get the NPA address. Once it has that
the importer can create a dma-buf to represent the NPA space used by the
allocation and that is returned to the application.
Proposed Userspace
https://github.com/ROCm/rocm-systems/blob/35959f8e1260c7cee3e51a740e320be3856ee4ff/projects/rocr-runtime/libhsakmt/src/memory.c#L971
https://github.com/ROCm/rocm-systems/blob/35959f8e1260c7cee3e51a740e320be3856ee4ff/projects/rocr-runtime/libhsakmt/src/memory.c#L1036
v2:
- Update the Documentation based on review feedback
Alex Deucher (1):
Documentation: add initial UALink Documentation
Mukul Joshi (2):
drm/gem: Add callback for when handle count goes to 0
drm/amdgpu: Add ioctl infra for exporting/importing UALink handles
Documentation/gpu/amdgpu/index.rst | 1 +
Documentation/gpu/amdgpu/ualink.rst | 75 ++++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c | 2 +
drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c | 9 +++
drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c | 44 +++++++++++++
drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h | 2 +
drivers/gpu/drm/drm_gem.c | 5 +-
include/drm/drm_gem.h | 11 ++++
include/uapi/drm/amdgpu_drm.h | 30 +++++++++
9 files changed, 178 insertions(+), 1 deletion(-)
create mode 100644 Documentation/gpu/amdgpu/ualink.rst
--
2.55.0
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-08-31 18:33 [PATCH V2 0/3] Add UALink infrastructure series 2 Alex Deucher @ 2026-08-31 18:33 ` Alex Deucher 2026-09-05 5:36 ` Thorsten Leemhuis 2026-08-31 18:33 ` [PATCH 2/3] drm/amdgpu: Add ioctl infra for exporting/importing UALink handles Alex Deucher 2026-08-31 18:33 ` [PATCH 3/3] Documentation: add initial UALink Documentation Alex Deucher 2 siblings, 1 reply; 12+ messages in thread From: Alex Deucher @ 2026-08-31 18:33 UTC (permalink / raw) To: amd-gfx, dri-devel, Felix.Kuehling Cc: Mukul Joshi, Christian König, Felix Kuehling, Alex Deucher From: Mukul Joshi <mukul.joshi@amd.com> Add an optional callback for driver-specific cleanup when the GEM handle of an object is freed. This will be used by AMDGPU to enable freeing of memory exported to other nodes in a UALink pod once all user mode references are gone. The callback is called outside the object_name_lock and before releasing the reference count on the GEM object Suggested-by: Christian König <christian.koenig@amd.com> Signed-off-by: Mukul Joshi <mukul.joshi@amd.com> Reviewed-by: Felix Kuehling <felix.kuehling@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> --- drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c | 9 +++++++++ drivers/gpu/drm/drm_gem.c | 5 ++++- include/drm/drm_gem.h | 11 +++++++++++ 3 files changed, 24 insertions(+), 1 deletion(-) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c index f754a4a3a1c22..0d579517c03ce 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c @@ -386,6 +386,14 @@ static int amdgpu_gem_object_mmap(struct drm_gem_object *obj, struct vm_area_str return drm_gem_ttm_mmap(obj, vma); } +static void amdgpu_gem_object_handle_free(struct drm_gem_object *gobj) +{ + struct amdgpu_bo *aobj = gem_to_amdgpu_bo(gobj); + + amdgpu_ualink_revoke_exported_memory(aobj); + +} + const struct drm_gem_object_funcs amdgpu_gem_object_funcs = { .free = amdgpu_gem_object_free, .open = amdgpu_gem_object_open, @@ -395,6 +403,7 @@ const struct drm_gem_object_funcs amdgpu_gem_object_funcs = { .vunmap = drm_gem_ttm_vunmap, .mmap = amdgpu_gem_object_mmap, .vm_ops = &amdgpu_gem_vm_ops, + .handle_free = amdgpu_gem_object_handle_free }; static bool amdgpu_gem_are_domains_valid(u32 domains) diff --git a/drivers/gpu/drm/drm_gem.c b/drivers/gpu/drm/drm_gem.c index e3ed684ddcf29..6a86bd2a0343e 100644 --- a/drivers/gpu/drm/drm_gem.c +++ b/drivers/gpu/drm/drm_gem.c @@ -354,8 +354,11 @@ void drm_gem_object_handle_put_unlocked(struct drm_gem_object *obj) } mutex_unlock(&dev->object_name_lock); - if (final) + if (final) { + if (obj->funcs->handle_free) + obj->funcs->handle_free(obj); drm_gem_object_put(obj); + } } /* diff --git a/include/drm/drm_gem.h b/include/drm/drm_gem.h index 8a704f6a65c15..95d8ae6f85df7 100644 --- a/include/drm/drm_gem.h +++ b/include/drm/drm_gem.h @@ -227,6 +227,17 @@ struct drm_gem_object_funcs { */ size_t (*rss)(struct drm_gem_object *obj); + /** + * @handle_free: + * + * This callback is called when the GEM handle count goes down to 0. + * It is currently used by AMDGPU driver to release their exported BO + * handles. + * + * This callback is optional. + */ + void (*handle_free)(struct drm_gem_object *obj); + /** * @vm_ops: * -- 2.55.0 ^ permalink raw reply related [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-08-31 18:33 ` [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 Alex Deucher @ 2026-09-05 5:36 ` Thorsten Leemhuis 2026-09-05 9:31 ` Manuel Ebner 0 siblings, 1 reply; 12+ messages in thread From: Thorsten Leemhuis @ 2026-09-05 5:36 UTC (permalink / raw) To: Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling Cc: Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On 8/31/26 20:33, Alex Deucher wrote: > From: Mukul Joshi <mukul.joshi@amd.com> > > Add an optional callback for driver-specific cleanup when the GEM > handle of an object is freed. This will be used by AMDGPU to enable > freeing of memory exported to other nodes in a UALink pod once all > user mode references are gone. > > The callback is called outside the object_name_lock and before > releasing the reference count on the GEM object This showed up in -next yesterday and afaics broke build on arm64 and x86_64 for me with the following error from rust: """ >> error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` >> --> rust/kernel/drm/gem/mod.rs:265:58 >> | >> 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { >> | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` >> >> error: aborting due to 1 previous error >> >> For more information about this error, try `rustc --explain E0063`. >> make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 >> make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 >> make: *** [Makefile:256: __sub-make] Error 2 """ Full log: https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz Reverting this change and 2/3 from this set fixed the problem for me. Ciao, Thorsten > Suggested-by: Christian König <christian.koenig@amd.com> > Signed-off-by: Mukul Joshi <mukul.joshi@amd.com> > Reviewed-by: Felix Kuehling <felix.kuehling@amd.com> > Signed-off-by: Alex Deucher <alexander.deucher@amd.com> > --- > drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c | 9 +++++++++ > drivers/gpu/drm/drm_gem.c | 5 ++++- > include/drm/drm_gem.h | 11 +++++++++++ > 3 files changed, 24 insertions(+), 1 deletion(-) > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c > index f754a4a3a1c22..0d579517c03ce 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c > @@ -386,6 +386,14 @@ static int amdgpu_gem_object_mmap(struct drm_gem_object *obj, struct vm_area_str > return drm_gem_ttm_mmap(obj, vma); > } > > +static void amdgpu_gem_object_handle_free(struct drm_gem_object *gobj) > +{ > + struct amdgpu_bo *aobj = gem_to_amdgpu_bo(gobj); > + > + amdgpu_ualink_revoke_exported_memory(aobj); > + > +} > + > const struct drm_gem_object_funcs amdgpu_gem_object_funcs = { > .free = amdgpu_gem_object_free, > .open = amdgpu_gem_object_open, > @@ -395,6 +403,7 @@ const struct drm_gem_object_funcs amdgpu_gem_object_funcs = { > .vunmap = drm_gem_ttm_vunmap, > .mmap = amdgpu_gem_object_mmap, > .vm_ops = &amdgpu_gem_vm_ops, > + .handle_free = amdgpu_gem_object_handle_free > }; > > static bool amdgpu_gem_are_domains_valid(u32 domains) > diff --git a/drivers/gpu/drm/drm_gem.c b/drivers/gpu/drm/drm_gem.c > index e3ed684ddcf29..6a86bd2a0343e 100644 > --- a/drivers/gpu/drm/drm_gem.c > +++ b/drivers/gpu/drm/drm_gem.c > @@ -354,8 +354,11 @@ void drm_gem_object_handle_put_unlocked(struct drm_gem_object *obj) > } > mutex_unlock(&dev->object_name_lock); > > - if (final) > + if (final) { > + if (obj->funcs->handle_free) > + obj->funcs->handle_free(obj); > drm_gem_object_put(obj); > + } > } > > /* > diff --git a/include/drm/drm_gem.h b/include/drm/drm_gem.h > index 8a704f6a65c15..95d8ae6f85df7 100644 > --- a/include/drm/drm_gem.h > +++ b/include/drm/drm_gem.h > @@ -227,6 +227,17 @@ struct drm_gem_object_funcs { > */ > size_t (*rss)(struct drm_gem_object *obj); > > + /** > + * @handle_free: > + * > + * This callback is called when the GEM handle count goes down to 0. > + * It is currently used by AMDGPU driver to release their exported BO > + * handles. > + * > + * This callback is optional. > + */ > + void (*handle_free)(struct drm_gem_object *obj); > + > /** > * @vm_ops: > * ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-09-05 5:36 ` Thorsten Leemhuis @ 2026-09-05 9:31 ` Manuel Ebner 2026-09-05 22:25 ` Klara Modin 0 siblings, 1 reply; 12+ messages in thread From: Manuel Ebner @ 2026-09-05 9:31 UTC (permalink / raw) To: Thorsten Leemhuis, Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling Cc: Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On Sat, 2026-09-05 at 07:36 +0200, Thorsten Leemhuis wrote: > On 8/31/26 20:33, Alex Deucher wrote: > > From: Mukul Joshi <mukul.joshi@amd.com> > > ... > ... > > > > error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` > > > --> rust/kernel/drm/gem/mod.rs:265:58 > > > | > > > 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { > > > | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` > > > > > > error: aborting due to 1 previous error > > > > > > For more information about this error, try `rustc --explain E0063`. > > > make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 > > > make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 > > > make: *** [Makefile:256: __sub-make] Error 2 > """ > Full log: > https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz > > Reverting this change and 2/3 from this set fixed the problem for me. Verified. Thanks Manuel > > Ciao, Thorsten > > ... ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-09-05 9:31 ` Manuel Ebner @ 2026-09-05 22:25 ` Klara Modin 2026-09-05 23:41 ` Gary Guo ` (2 more replies) 0 siblings, 3 replies; 12+ messages in thread From: Klara Modin @ 2026-09-05 22:25 UTC (permalink / raw) To: Manuel Ebner Cc: Thorsten Leemhuis, Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling, Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On 2026-09-05 11:31:41 +0200, Manuel Ebner wrote: > On Sat, 2026-09-05 at 07:36 +0200, Thorsten Leemhuis wrote: > > On 8/31/26 20:33, Alex Deucher wrote: > > > From: Mukul Joshi <mukul.joshi@amd.com> > > > ... > > ... > > > > > > error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` > > > > --> rust/kernel/drm/gem/mod.rs:265:58 > > > > | > > > > 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { > > > > | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` > > > > > > > > error: aborting due to 1 previous error > > > > > > > > For more information about this error, try `rustc --explain E0063`. > > > > make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 > > > > make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 > > > > make: *** [Makefile:256: __sub-make] Error 2 > > """ > > Full log: > > https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz > > > > Reverting this change and 2/3 from this set fixed the problem for me. That leaves only the documentation without the feature so might as well drop 3/3 too in that case? Otherwise, the following also fixes it for me: diff --git a/rust/kernel/drm/gem/mod.rs b/rust/kernel/drm/gem/mod.rs index e1ebad77ebe2..847e0ee38863 100644 --- a/rust/kernel/drm/gem/mod.rs +++ b/rust/kernel/drm/gem/mod.rs @@ -278,6 +278,7 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> { vm_ops: core::ptr::null_mut(), evict: None, rss: None, + handle_free: None, }; /// Returns the `Device` that owns this GEM object. > > Verified. > > Thanks > Manuel > > > > Ciao, Thorsten > > > ... Regards, Klara Modin ^ permalink raw reply related [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-09-05 22:25 ` Klara Modin @ 2026-09-05 23:41 ` Gary Guo 2026-09-08 4:33 ` Thorsten Leemhuis 2026-09-08 16:03 ` Alex Deucher 2 siblings, 0 replies; 12+ messages in thread From: Gary Guo @ 2026-09-05 23:41 UTC (permalink / raw) To: Klara Modin, Manuel Ebner Cc: Thorsten Leemhuis, Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling, Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On Sat Sep 5, 2026 at 11:25 PM BST, Klara Modin wrote: > On 2026-09-05 11:31:41 +0200, Manuel Ebner wrote: >> On Sat, 2026-09-05 at 07:36 +0200, Thorsten Leemhuis wrote: >> > On 8/31/26 20:33, Alex Deucher wrote: >> > > From: Mukul Joshi <mukul.joshi@amd.com> >> > > ... >> > ... >> > >> > > > error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` >> > > > --> rust/kernel/drm/gem/mod.rs:265:58 >> > > > | >> > > > 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { >> > > > | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` >> > > > >> > > > error: aborting due to 1 previous error >> > > > >> > > > For more information about this error, try `rustc --explain E0063`. >> > > > make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 >> > > > make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 >> > > > make: *** [Makefile:256: __sub-make] Error 2 >> > """ >> > Full log: >> > https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz >> > >> > Reverting this change and 2/3 from this set fixed the problem for me. > > That leaves only the documentation without the feature so might as well > drop 3/3 too in that case? > > Otherwise, the following also fixes it for me: > > diff --git a/rust/kernel/drm/gem/mod.rs b/rust/kernel/drm/gem/mod.rs > index e1ebad77ebe2..847e0ee38863 100644 > --- a/rust/kernel/drm/gem/mod.rs > +++ b/rust/kernel/drm/gem/mod.rs > @@ -278,6 +278,7 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> { > vm_ops: core::ptr::null_mut(), > evict: None, > rss: None, > + handle_free: None, IMO this should use ..pin_init::zeroed() to set all other fields as zero rather than explicitly listing all of them. Best, Gary > }; > > /// Returns the `Device` that owns this GEM object. > >> >> Verified. >> >> Thanks >> Manuel >> > >> > Ciao, Thorsten >> > > ... > > Regards, > Klara Modin ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-09-05 22:25 ` Klara Modin 2026-09-05 23:41 ` Gary Guo @ 2026-09-08 4:33 ` Thorsten Leemhuis 2026-09-08 16:03 ` Alex Deucher 2 siblings, 0 replies; 12+ messages in thread From: Thorsten Leemhuis @ 2026-09-08 4:33 UTC (permalink / raw) To: Klara Modin, Manuel Ebner Cc: Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling, Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On 9/6/26 00:25, Klara Modin wrote: > On 2026-09-05 11:31:41 +0200, Manuel Ebner wrote: >> On Sat, 2026-09-05 at 07:36 +0200, Thorsten Leemhuis wrote: >>> On 8/31/26 20:33, Alex Deucher wrote: >>>> From: Mukul Joshi <mukul.joshi@amd.com> > >>>>> error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` >>>>> --> rust/kernel/drm/gem/mod.rs:265:58 >>>>> | >>>>> 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { >>>>> | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` >>>>> >>>>> error: aborting due to 1 previous error >>>>> >>>>> For more information about this error, try `rustc --explain E0063`. >>>>> make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 >>>>> make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 >>>>> make: *** [Makefile:256: __sub-make] Error 2 >>> Full log: >>> https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz >>> >>> Reverting this change and 2/3 from this set fixed the problem for me. > > That leaves only the documentation without the feature so might as well > drop 3/3 too in that case? My intention was just to confirm that these patches are the problem, not suggest this as a solution, so I didn't bother about the docs. :-D > Otherwise, the following also fixes it for me: Thx, for me, too. Ciao, Thorsten > diff --git a/rust/kernel/drm/gem/mod.rs b/rust/kernel/drm/gem/mod.rs > index e1ebad77ebe2..847e0ee38863 100644 > --- a/rust/kernel/drm/gem/mod.rs > +++ b/rust/kernel/drm/gem/mod.rs > @@ -278,6 +278,7 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> { > vm_ops: core::ptr::null_mut(), > evict: None, > rss: None, > + handle_free: None, > }; > > /// Returns the `Device` that owns this GEM object. ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-09-05 22:25 ` Klara Modin 2026-09-05 23:41 ` Gary Guo 2026-09-08 4:33 ` Thorsten Leemhuis @ 2026-09-08 16:03 ` Alex Deucher 2026-09-08 17:57 ` Klara Modin 2 siblings, 1 reply; 12+ messages in thread From: Alex Deucher @ 2026-09-08 16:03 UTC (permalink / raw) To: Klara Modin Cc: Manuel Ebner, Thorsten Leemhuis, Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling, Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On Mon, Sep 7, 2026 at 3:25 AM Klara Modin <klarasmodin@gmail.com> wrote: > > On 2026-09-05 11:31:41 +0200, Manuel Ebner wrote: > > On Sat, 2026-09-05 at 07:36 +0200, Thorsten Leemhuis wrote: > > > On 8/31/26 20:33, Alex Deucher wrote: > > > > From: Mukul Joshi <mukul.joshi@amd.com> > > > > ... > > > ... > > > > > > > > error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` > > > > > --> rust/kernel/drm/gem/mod.rs:265:58 > > > > > | > > > > > 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { > > > > > | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` > > > > > > > > > > error: aborting due to 1 previous error > > > > > > > > > > For more information about this error, try `rustc --explain E0063`. > > > > > make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 > > > > > make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 > > > > > make: *** [Makefile:256: __sub-make] Error 2 > > > """ > > > Full log: > > > https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz > > > > > > Reverting this change and 2/3 from this set fixed the problem for me. > > That leaves only the documentation without the feature so might as well > drop 3/3 too in that case? > > Otherwise, the following also fixes it for me: > > diff --git a/rust/kernel/drm/gem/mod.rs b/rust/kernel/drm/gem/mod.rs > index e1ebad77ebe2..847e0ee38863 100644 > --- a/rust/kernel/drm/gem/mod.rs > +++ b/rust/kernel/drm/gem/mod.rs > @@ -278,6 +278,7 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> { > vm_ops: core::ptr::null_mut(), > evict: None, > rss: None, > + handle_free: None, > }; > > /// Returns the `Device` that owns this GEM object. > Can you send this as a proper patch? Thanks, Alex > > > > Verified. > > > > Thanks > > Manuel > > > > > > Ciao, Thorsten > > > > ... > > Regards, > Klara Modin ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 2026-09-08 16:03 ` Alex Deucher @ 2026-09-08 17:57 ` Klara Modin 0 siblings, 0 replies; 12+ messages in thread From: Klara Modin @ 2026-09-08 17:57 UTC (permalink / raw) To: Alex Deucher Cc: Manuel Ebner, Thorsten Leemhuis, Alex Deucher, amd-gfx, dri-devel, Felix.Kuehling, Mukul Joshi, Christian König, Mark Brown, Miguel Ojeda, rust-for-linux, Linux Next Mailing List On 2026-09-08 12:03:57 -0400, Alex Deucher wrote: > On Mon, Sep 7, 2026 at 3:25 AM Klara Modin <klarasmodin@gmail.com> wrote: > > > > On 2026-09-05 11:31:41 +0200, Manuel Ebner wrote: > > > On Sat, 2026-09-05 at 07:36 +0200, Thorsten Leemhuis wrote: > > > > On 8/31/26 20:33, Alex Deucher wrote: > > > > > From: Mukul Joshi <mukul.joshi@amd.com> > > > > > ... > > > > ... > > > > > > > > > > error[E0063]: missing field `handle_free` in initializer of `drm_gem_object_funcs` > > > > > > --> rust/kernel/drm/gem/mod.rs:265:58 > > > > > > | > > > > > > 265 | const OBJECT_FUNCS: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs { > > > > > > | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ missing `handle_free` > > > > > > > > > > > > error: aborting due to 1 previous error > > > > > > > > > > > > For more information about this error, try `rustc --explain E0063`. > > > > > > make[2]: *** [rust/Makefile:799: rust/kernel.o] Error 1 > > > > > > make[1]: *** [/builddir/build/BUILD/kernel-7.3.0-build/kernel-next-20260904/linux-7.3.0-0.0.next.20260904.221.vanilla.fc43.aarch64/Makefile:1442: prepare] Error 2 > > > > > > make: *** [Makefile:256: __sub-make] Error 2 > > > > """ > > > > Full log: > > > > https://download.copr.fedorainfracloud.org/results/@kernel-vanilla/next/fedora-43-aarch64/10951622-next-next-all/builder-live.log.gz > > > > > > > > Reverting this change and 2/3 from this set fixed the problem for me. > > > > That leaves only the documentation without the feature so might as well > > drop 3/3 too in that case? > > > > Otherwise, the following also fixes it for me: > > > > diff --git a/rust/kernel/drm/gem/mod.rs b/rust/kernel/drm/gem/mod.rs > > index e1ebad77ebe2..847e0ee38863 100644 > > --- a/rust/kernel/drm/gem/mod.rs > > +++ b/rust/kernel/drm/gem/mod.rs > > @@ -278,6 +278,7 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> { > > vm_ops: core::ptr::null_mut(), > > evict: None, > > rss: None, > > + handle_free: None, > > }; > > > > /// Returns the `Device` that owns this GEM object. > > > > Can you send this as a proper patch? Here: https://lore.kernel.org/lkml/20260908175427.47207-1-klarasmodin@gmail.com I liked Gary's suggestion so went with that instead. > > Thanks, > > Alex Regards, Klara Modin > > > > > > > Verified. > > > > > > Thanks > > > Manuel > > > > > > > > Ciao, Thorsten > > > > > ... > > > > Regards, > > Klara Modin ^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH 2/3] drm/amdgpu: Add ioctl infra for exporting/importing UALink handles 2026-08-31 18:33 [PATCH V2 0/3] Add UALink infrastructure series 2 Alex Deucher 2026-08-31 18:33 ` [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 Alex Deucher @ 2026-08-31 18:33 ` Alex Deucher 2026-08-31 18:33 ` [PATCH 3/3] Documentation: add initial UALink Documentation Alex Deucher 2 siblings, 0 replies; 12+ messages in thread From: Alex Deucher @ 2026-08-31 18:33 UTC (permalink / raw) To: amd-gfx, dri-devel, Felix.Kuehling Cc: Mukul Joshi, Horatio Zhang, Alex Deucher From: Mukul Joshi <mukul.joshi@amd.com> Add the ioctl infrastructure to support exporting and importing BOs to facilitate NPA based memory sharing across GPUs in a rack scale setup. Proposed userspace: https://github.com/ROCm/rocm-systems/blob/35959f8e1260c7cee3e51a740e320be3856ee4ff/projects/rocr-runtime/libhsakmt/src/memory.c#L971 https://github.com/ROCm/rocm-systems/blob/35959f8e1260c7cee3e51a740e320be3856ee4ff/projects/rocr-runtime/libhsakmt/src/memory.c#L1036 v2: Move the ioctl wire-up to the end of the series. Signed-off-by: Mukul Joshi <mukul.joshi@amd.com> Signed-off-by: Horatio Zhang <hongkun.zhang@amd.com> Signed-off-by: Alex Deucher <alexander.deucher@amd.com> --- drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c | 2 + drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c | 44 ++++++++++++++++++++++ drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h | 2 + include/uapi/drm/amdgpu_drm.h | 30 +++++++++++++++ 4 files changed, 78 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c index 287c9455969b2..9c5e93cd3ee64 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c @@ -54,6 +54,7 @@ #include "amdgpu_userq.h" #include "amdgpu_userq_fence.h" #include "../amdxcp/amdgpu_xcp_drv.h" +#include "amdgpu_ualink.h" /* * KMS wrapper. @@ -3120,6 +3121,7 @@ const struct drm_ioctl_desc amdgpu_ioctls_kms[] = { DRM_IOCTL_DEF_DRV(AMDGPU_USERQ_WAIT, amdgpu_userq_wait_ioctl, DRM_AUTH|DRM_RENDER_ALLOW), DRM_IOCTL_DEF_DRV(AMDGPU_GEM_LIST_HANDLES, amdgpu_gem_list_handles_ioctl, DRM_AUTH|DRM_RENDER_ALLOW), DRM_IOCTL_DEF_DRV(AMDGPU_PROC_OPTIONS, amdgpu_proc_options_ioctl, DRM_AUTH|DRM_RENDER_ALLOW), + DRM_IOCTL_DEF_DRV(AMDGPU_UALINK_HANDLE, amdgpu_gem_ualink_handle_ioctl, DRM_AUTH|DRM_RENDER_ALLOW) }; static const struct drm_driver amdgpu_kms_driver = { diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c index 3f96044a3d54f..f82bc9a9d0837 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c @@ -4008,6 +4008,50 @@ int amdgpu_ualink_export_handle(struct drm_device *dev, struct drm_file *filp, return r; } +int amdgpu_gem_ualink_handle_ioctl(struct drm_device *dev, void *data, + struct drm_file *filp) +{ + union drm_amdgpu_ualink_handle *args = data; + struct amdgpu_device *adev = drm_to_adev(dev); + struct amdgpu_ualink_handle handle = {}; + u32 gem_handle; + int r, fd = -1; + + if (adev->ualink.info->accel_state != + AMDGPU_UALINK_ACCEL_STATE_ACTIVE) { + dev_err(adev->dev, + "ualink device is not in active state in vpod\n"); + return -EOPNOTSUPP; + } + + /* The input and output members of the ioctl argument alias each other. + * Latch every input field before invoking the handlers, and only write + * the output fields afterwards. + */ + switch (args->in.op) { + case DRM_AMDGPU_UALINK_HANDLE_OP_EXPORT: + gem_handle = args->in.gem_handle; + r = amdgpu_ualink_export_handle(dev, filp, gem_handle, &handle); + if (!r) { + args->out.export_ualink_handle[0] = handle.handle_lo; + args->out.export_ualink_handle[1] = handle.handle_hi; + } + break; + case DRM_AMDGPU_UALINK_HANDLE_OP_IMPORT: + handle.handle_lo = args->in.import_ualink_handle[0]; + handle.handle_hi = args->in.import_ualink_handle[1]; + r = amdgpu_ualink_import_handle(dev, &handle, &fd); + if (!r) + args->out.import_dmabuf_handle = fd; + break; + default: + r = -EINVAL; + break; + } + + return r; +} + int amdgpu_ualink_manager_start(struct amdgpu_device *adev) { int i, r; diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h index 1bd7181147dd4..85c9a9634a4b0 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h @@ -413,6 +413,8 @@ int amdgpu_ualink_export_handle(struct drm_device *dev, struct drm_file *filp, int amdgpu_ualink_import_handle(struct drm_device *dev, const struct amdgpu_ualink_handle *ualink_handle, int *fd_out); +int amdgpu_gem_ualink_handle_ioctl(struct drm_device *dev, void *data, + struct drm_file *filp); void amdgpu_ualink_revoke_exported_memory(struct amdgpu_bo *bo); int ualink_ip_hw_init(struct amdgpu_ip_block *ip_block); diff --git a/include/uapi/drm/amdgpu_drm.h b/include/uapi/drm/amdgpu_drm.h index d6d5402a1e789..0e8115673d8d2 100644 --- a/include/uapi/drm/amdgpu_drm.h +++ b/include/uapi/drm/amdgpu_drm.h @@ -59,6 +59,7 @@ extern "C" { #define DRM_AMDGPU_USERQ_WAIT 0x18 #define DRM_AMDGPU_GEM_LIST_HANDLES 0x19 #define DRM_AMDGPU_PROC_OPTIONS 0x1A +#define DRM_AMDGPU_UALINK_HANDLE 0x1B #define DRM_IOCTL_AMDGPU_GEM_CREATE DRM_IOWR(DRM_COMMAND_BASE + DRM_AMDGPU_GEM_CREATE, union drm_amdgpu_gem_create) #define DRM_IOCTL_AMDGPU_GEM_MMAP DRM_IOWR(DRM_COMMAND_BASE + DRM_AMDGPU_GEM_MMAP, union drm_amdgpu_gem_mmap) @@ -81,6 +82,7 @@ extern "C" { #define DRM_IOCTL_AMDGPU_USERQ_WAIT DRM_IOWR(DRM_COMMAND_BASE + DRM_AMDGPU_USERQ_WAIT, struct drm_amdgpu_userq_wait) #define DRM_IOCTL_AMDGPU_GEM_LIST_HANDLES DRM_IOWR(DRM_COMMAND_BASE + DRM_AMDGPU_GEM_LIST_HANDLES, struct drm_amdgpu_gem_list_handles) #define DRM_IOCTL_AMDGPU_PROC_OPTIONS DRM_IOWR(DRM_COMMAND_BASE + DRM_AMDGPU_PROC_OPTIONS, struct drm_amdgpu_proc_options) +#define DRM_IOCTL_AMDGPU_UALINK_HANDLE DRM_IOWR(DRM_COMMAND_BASE + DRM_AMDGPU_UALINK_HANDLE, union drm_amdgpu_ualink_handle) /** * DOC: memory domains @@ -1705,6 +1707,34 @@ struct drm_amdgpu_proc_options { } kfd_sigbus_delay; }; +#define DRM_AMDGPU_UALINK_HANDLE_OP_EXPORT 0 +#define DRM_AMDGPU_UALINK_HANDLE_OP_IMPORT 1 + +struct drm_amdgpu_ualink_handle_in { + /* Export or import */ + __u32 op; + /* For future use, no flags defined so far */ + __u32 flags; + union { + /* GEM handle of the BO to export */ + __u32 gem_handle; + /* UALink handle to import */ + __u64 import_ualink_handle[2]; + }; +}; + +union drm_amdgpu_ualink_handle_out { + /* Exported UALink handle */ + __u64 export_ualink_handle[2]; + /** DMABuf representing the imported UALink handle */ + __u32 import_dmabuf_handle; +}; + +union drm_amdgpu_ualink_handle { + struct drm_amdgpu_ualink_handle_in in; + union drm_amdgpu_ualink_handle_out out; +}; + #if defined(__cplusplus) } #endif -- 2.55.0 ^ permalink raw reply related [flat|nested] 12+ messages in thread
* [PATCH 3/3] Documentation: add initial UALink Documentation 2026-08-31 18:33 [PATCH V2 0/3] Add UALink infrastructure series 2 Alex Deucher 2026-08-31 18:33 ` [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 Alex Deucher 2026-08-31 18:33 ` [PATCH 2/3] drm/amdgpu: Add ioctl infra for exporting/importing UALink handles Alex Deucher @ 2026-08-31 18:33 ` Alex Deucher 2 siblings, 0 replies; 12+ messages in thread From: Alex Deucher @ 2026-08-31 18:33 UTC (permalink / raw) To: amd-gfx, dri-devel, Felix.Kuehling; +Cc: Alex Deucher, Joseph.Greathouse Document the details of UALink on the GPU. v2: add updates from Felix Cc: Felix.Kuehling@amd.com Cc: Joseph.Greathouse@amd.com Signed-off-by: Alex Deucher <alexander.deucher@amd.com> --- Documentation/gpu/amdgpu/index.rst | 1 + Documentation/gpu/amdgpu/ualink.rst | 75 +++++++++++++++++++++++++++++ 2 files changed, 76 insertions(+) create mode 100644 Documentation/gpu/amdgpu/ualink.rst diff --git a/Documentation/gpu/amdgpu/index.rst b/Documentation/gpu/amdgpu/index.rst index b2ab182236efb..ba2ee73278672 100644 --- a/Documentation/gpu/amdgpu/index.rst +++ b/Documentation/gpu/amdgpu/index.rst @@ -23,4 +23,5 @@ Next (GCN), Radeon DNA (RDNA), and Compute DNA (CDNA) architectures. debugfs process-isolation amdgpu-glossary + ualink ptl diff --git a/Documentation/gpu/amdgpu/ualink.rst b/Documentation/gpu/amdgpu/ualink.rst new file mode 100644 index 0000000000000..e15c8621e5a1a --- /dev/null +++ b/Documentation/gpu/amdgpu/ualink.rst @@ -0,0 +1,75 @@ +============== +UALink Support +============== + +Overview +======== + +Connected GPUs in a pod can directly access the remove memory on another GPU +over UALink. Unlike RMDA, there is no copy involved; it is direct loads/stores +over the fabric. Shared memory can only be accessed by a remote GPU if the +memory was exported and the importer has been authorized. For the memory to be +shared, it must be part of a unified physical address space shared between +nodes. This address space is called NPA (Nework Physical Address) space. This +address space is partitioned between the GPUs so that each GPU has its own +segment of the address space in which to export its memory. Each GPU maintains +a dedicated set of page tables for their NPA space similar to GPUVM. Note that +this mechanism only allows for GPU access to remote memory. The remote memory +is not CPU accessible. + +Exported memory is pinned. This is similar to how dma-bufs in VRAM are pinned +for P2P access. Remote TLB shootdowns from the exporter are used when the +exported memory is freed in order to remove access by remote GPUs. + +To access remote memory, the driver can map NPA addresses into its per process +GPUVM page tables just like local memory. Applications use opaque handles to +represent remote memory. GPUs in a pod communicate with eachother directly to +exchange NPA addresses between importers and exporters. If a node goes offline +or is reset, their peers will clean up any remaining refrences that are lost +when that happens. NPA addresses are exchanged directly between the kernel +drivers using the scale up fabric. NPA addresses are never exposed to user +space. + +On the importer, the NPA space is like another physical address space. NPA +addresses can be used as physical addresses for GPUVM to provide GPU virtual +addresses to the memory for processes using the GPU. + +On the exporter, the NPA space provides a way to expose discontiguous local +memory as a contiguous address range for remote GPUs. This allows the exporter +to locally manage the pages mapped into the NPA space. + +Remote NPAs are managed like another device specific TTM pool similar to +doorbells or VRAM, however they cannot be CPU mapped. + + +User Interface +============== +Two IOCTLs are provided to export and import remote memory. + +Export Memory +------------- +To export memory, a UALINK handle must be created for an allocation that can be +shared with another node in the pod. To do this the exporter calls the GEM +UALink IOCTL with the GEM handle to the buffer it wants to export. The IOCTL +returns a unique 128 bit handle which can be shared with the remote host. +Calling export on the same GEM handle always returns the same UALink handle. +The UALink handle is destroyed when the GEM object reference count reaches 0. + +Import Memory +------------- +To import remote memory, the UALink handle from the remote node must be +converted from a UALink handle to a local GEM object which represents the local +reference to the NPA space on the importer. If the memory has already been +imported, it just returns a new reference to the existing GEM object. If not, +the importer queries the exporter to get the NPA address. Once it has that, the +importer can create the GEM to represent the NPA space used by the allocation. +The GEM object is then exported to the caller as a dma-buf. The dma-buf is +leveraged for dynamic attachment which provides the ability to revoke access +when necessary. + + +Device to Device Communications +=============================== + +Devices communicate via a protocol implemented in firmware. Mesages sent to a +remote node generate an interrupt on that node for servicing. -- 2.55.0 ^ permalink raw reply related [flat|nested] 12+ messages in thread
* [PATCH V2 0/3] Add UALink infrastructure series 2 @ 2026-08-24 17:01 Alex Deucher 0 siblings, 0 replies; 12+ messages in thread From: Alex Deucher @ 2026-08-24 17:01 UTC (permalink / raw) To: amd-gfx, dri-devel, Felix.Kuehling; +Cc: Alex Deucher This adds the core infrastructure for supporting UALink (Ultra Accelerator Link) connected scale up pods. I've split this into two series, one to add the core infrastructure, and one to expose the new IOCTL interface and add the documentation to avoid spamming the larger audience with the implemenation defails. This is the second series. The first series can be found here: https://lists.freedesktop.org/archives/amd-gfx/2026-August/151200.html Both series can be accessed via this git branch: https://gitlab.freedesktop.org/agd5f/linux/-/commits/ualink?ref_type=heads This implements UALoE (UALink over Ethernet). An overview of the complete solution can be found here: https://www.amd.com/en/products/rackscale-solutions/helios.html Overview Connected GPUs in a pod can directly access the remove memory on another GPU over UALink. Unlike RMDA, there is no copy involved; it is direct loads/stores over the fabric. Shared memory can only be accessed by a remote GPU if the memory was exported and the importer has been authorized. For the memory to be shared, it must be part of a unified physical address space shared between nodes. This address space is called NPA (Nework Physical Address) space. This address space is partitioned between the GPUs so that each GPU has it's own segment of the address space in which to export its memory. Each GPU maintains a dedicated set of page tables for their NPA space similar to GPUVM. Exported memory is pinned. This is similar to how dma-bufs in VRAM are pinned for P2P access. Remote TLB shootdowns from the exporter are used when the exported memory is freed in order to remove access by remote GPUs. To access remote memory, the driver can map NPA addresses into its per process GPUVM page tables just like local memory. Applications use opaque handles to represent remote memory. GPUs in a pod communicate with eachother directly to exchange NPA addresses between importers and exporters. If a node goes offline or is reset, their peers will clean up any remaining refrences that are lost when that happens. User interface Export Memory To export memory, a handle must be created for an allocation that can be shared with another node in the pod. To do this the exporter calls the GEM UALink IOCTL with the GEM handle to the buffer it wants to export. The IOCTL returns a unique handle which can be shared with the remote host. Import Memory To import the memory, the handle from the remote node must be converted from a unique handle to a local GEM object which represents the local reference to the NPA space on the importer. If the memory has already been imported, it just returns a new reference to the existing object. If not, the importer queries the exporter to get the NPA address. Once it has that the importer can create a dma-buf to represent the NPA space used by the allocation and that is returned to the application. Proposed Userspace https://github.com/ROCm/rocm-systems/blob/35959f8e1260c7cee3e51a740e320be3856ee4ff/projects/rocr-runtime/libhsakmt/src/memory.c#L971 https://github.com/ROCm/rocm-systems/blob/35959f8e1260c7cee3e51a740e320be3856ee4ff/projects/rocr-runtime/libhsakmt/src/memory.c#L1036 V2: Update the documentation per Felix' feedback. Clarify that exported memory is pinned. Alex Deucher (1): Documentation: add initial UALink Documentation Mukul Joshi (2): drm/gem: Add callback for when handle count goes to 0 drm/amdgpu: Add ioctl infra for exporting/importing UALink handles Documentation/gpu/amdgpu/index.rst | 1 + Documentation/gpu/amdgpu/ualink.rst | 75 ++++++++++++++++++++++ drivers/gpu/drm/amd/amdgpu/amdgpu_drv.c | 2 + drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c | 9 +++ drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.c | 44 +++++++++++++ drivers/gpu/drm/amd/amdgpu/amdgpu_ualink.h | 2 + drivers/gpu/drm/drm_gem.c | 5 +- include/drm/drm_gem.h | 11 ++++ include/uapi/drm/amdgpu_drm.h | 30 +++++++++ 9 files changed, 178 insertions(+), 1 deletion(-) create mode 100644 Documentation/gpu/amdgpu/ualink.rst -- 2.55.0 ^ permalink raw reply [flat|nested] 12+ messages in thread
end of thread, other threads:[~2026-09-09 6:55 UTC | newest] Thread overview: 12+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-31 18:33 [PATCH V2 0/3] Add UALink infrastructure series 2 Alex Deucher 2026-08-31 18:33 ` [PATCH 1/3] drm/gem: Add callback for when handle count goes to 0 Alex Deucher 2026-09-05 5:36 ` Thorsten Leemhuis 2026-09-05 9:31 ` Manuel Ebner 2026-09-05 22:25 ` Klara Modin 2026-09-05 23:41 ` Gary Guo 2026-09-08 4:33 ` Thorsten Leemhuis 2026-09-08 16:03 ` Alex Deucher 2026-09-08 17:57 ` Klara Modin 2026-08-31 18:33 ` [PATCH 2/3] drm/amdgpu: Add ioctl infra for exporting/importing UALink handles Alex Deucher 2026-08-31 18:33 ` [PATCH 3/3] Documentation: add initial UALink Documentation Alex Deucher -- strict thread matches above, loose matches on Subject: below -- 2026-08-24 17:01 [PATCH V2 0/3] Add UALink infrastructure series 2 Alex Deucher
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.