From: Dhruv Bhogaonkar <dhruv.b@linux.ibm.com>
To: Felix Kuehling <felix.kuehling@amd.com>,
Donet Tom <donettom@linux.ibm.com>,
amd-gfx@lists.freedesktop.org,
Alex Deucher <alexander.deucher@amd.com>,
Alex Deucher <alexdeucher@gmail.com>,
christian.koenig@amd.com, Philip Yang <yangp@amd.com>
Cc: David.YatSin@amd.com, Kent.Russell@amd.com,
Ritesh Harjani <ritesh.list@gmail.com>,
Vaidyanathan Srinivasan <svaidy@linux.ibm.com>,
David Airlie <airlied@gmail.com>, Simona Vetter <simona@ffwll.ch>
Subject: Re: [PATCH 0/7] drm/amdgpu: Fix topology device creation and proximity domain mappings for sparse and CPU-less NUMA systems
Date: Wed, 23 Sep 2026 11:28:28 +0530 [thread overview]
Message-ID: <cafaaba4-7d45-46ef-b0bf-67fe8cfa70ac@linux.ibm.com> (raw)
In-Reply-To: <00a10239-1e09-470b-8970-db00814af486@amd.com>
On 23/09/26 04:19, Felix Kuehling wrote:
> On 2026-09-22 12:59, Dhruv Bhogaonkar wrote:
>>
>> On 21/09/26 22:35, Kuehling, Felix wrote:
>>> On 2026-09-21 00:56, Dhruv Bhogaonkar wrote:
>>>> On 05/08/26 21:46, Kuehling, Felix wrote:
>>>>>
>>>>> On 2026-08-05 03:09, Donet Tom wrote:
>>>>>>
>>>>>> On 8/5/26 3:57 AM, Felix Kuehling wrote:
>>>>>>>
>>>>>>> On 2026-08-04 05:52, Donet Tom wrote:
>>>>>>>> This series fixes topology device creation and proximity domain
>>>>>>>> mappings in
>>>>>>>> AMDKFD when a system does not provide a CRAT table and the driver
>>>>>>>> generates a
>>>>>>>> Virtual CRAT (VCRAT).
>>>>>>>>
>>>>>>>> The current implementation assumes that CPU NUMA node IDs are
>>>>>>>> contiguous and
>>>>>>>> that every NUMA node contains CPUs. During VCRAT generation, CPU
>>>>>>>> topology
>>>>>>>> entries and proximity domains are created only for NUMA nodes that
>>>>>>>> have CPUs.
>>>>>>>> GPU proximity domains are then allocated immediately after the CPU
>>>>>>>> proximity
>>>>>>>> domains, and the GPU I/O link (proximity_domain_to) is initialized
>>>>>>>> using the
>>>>>>>> NUMA node ID to which the GPU is attached, implicitly assuming
>>>>>>>> that
>>>>>>>> NUMA node
>>>>>>>> IDs and proximity domains have a one-to-one mapping.
>>>>>>>>
>>>>>>>> These assumptions break on systems with:
>>>>>>>>
>>>>>>>> Sparse (non-contiguous) NUMA node IDs
>>>>>>>> CPU-less NUMA nodes
>>>>>>>> CPU-less and memory-less NUMA nodes
>>>>>>>>
>>>>>>>> For example:
>>>>>>>>
>>>>>>>> available: 3 nodes (0,2-3)
>>>>>>>>
>>>>>>>> node 0: CPUs present
>>>>>>>> node 2: CPU-less
>>>>>>>> node 3: CPU-less
>>>>>>>>
>>>>>>>> In this case, the driver creates a CPU topology device only for
>>>>>>>> node 0 and
>>>>>>>> assigns a single CPU proximity domain (0). However, the GPU VCRAT
>>>>>>>> still
>>>>>>>> references the NUMA node ID to which the GPU is attached (for
>>>>>>>> example, node 3)
>>>>>>>> in the proximity_domain_to field. Since no corresponding CPU
>>>>>>>> proximity domain
>>>>>>>> exists for node 3, the parser cannot find a matching proximity
>>>>>>>> domain during
>>>>>>>> VCRAT parsing, causing topology initialization to fail.
>>>>>>>
>>>>>>> Hi Tom,
>>>>>>
>>>>>>
>>>>>> Hi Felix,
>>>>>>
>>>>>>
>>>>>>>
>>>>>>> Thank you for the explanation and the patch series. I may have some
>>>>>>> gaps in my understanding that I would like to clarify. In my
>>>>>>> mind, I
>>>>>>> was using "proximity domain" and "NUMA node" interchangeably.
>>>>>>> You're
>>>>>>> demonstrating that they are not the same thing. Is that just a
>>>>>>> different way of labeling the same thing, or are NUMA nodes and
>>>>>>> proximity domains fundamentally different concepts.
>>>>>>>
>>>>>>> Your code in patch 5 (kfd_proximity_domain_to_numa_node and
>>>>>>> kfd_numa_node_to_proximity_domain) seems to imply that there is, in
>>>>>>> fact, a 1:1 mapping, as it assumes that there is a unique
>>>>>>> translation in both directions. Am I missing something?
>>>>>>
>>>>>>
>>>>>>
>>>>>> Thanks for the comment.
>>>>>>
>>>>>> IIUC, NUMA node IDs and proximity domains are different numbering
>>>>>> schemes. We have proximity domains for both CPU and GPU devices. CPU
>>>>>> proximity domains start from 0, and GPU proximity domains start
>>>>>> after
>>>>>> the last CPU proximity domain.
>>>>>>
>>>>>> If the NUMA node IDs are contiguous, the CPU proximity domains
>>>>>> happen
>>>>>> to match the NUMA node IDs. However, if the NUMA node IDs are
>>>>>> discontiguous, the CPU proximity domains and NUMA node IDs no longer
>>>>>> have a one-to-one mapping.
>>>>>>
>>>>>>>
>>>>>>> If they are just different numbering systems, do we really need to
>>>>>>> keep track of both proximity domains and NUMA nodes in the KFD
>>>>>>> topology? Or would it be sufficient to only track NUMA nodes, if
>>>>>>> that's what we really care about in the uAPI (KFD sysfs)?
>>>>>>
>>>>>>
>>>>>>
>>>>>> Thanks for the suggestion. I also think we don't need to keep track
>>>>>> of both the proximity domains and the NUMA node IDs in KFD. Do you
>>>>>> think the approach below would be reasonable?
>>>>>>
>>>>>> Just to make sure I understand correctly, for CPU devices the
>>>>>> proximity domain will be the same as the NUMA node ID, and the GPU
>>>>>> proximity domains will start after the last CPU proximity domain.
>>>>>>
>>>>>> In that case, the CPU proximity domains and NUMA node IDs will
>>>>>> always
>>>>>> have a one-to-one mapping, and the GPU proximity domains will follow
>>>>>> after them.
>>>>>>
>>>>>> For example, if a system has three NUMA nodes and two GPUs:
>>>>>>
>>>>>> NUMA node IDs: 0 2 3
>>>>>> CPU proximity domains: 0 2 3
>>>>>> GPU proximity domains: 4 5
>>>>>>
>>>>>> Would it be okay to proceed with this approach?
>>>>>
>>>>> Yes, this looks good to me.
>>>>
>>>>
>>>> Hi Felix,
>>>>
>>>> I have been looking into this issue along with Donet.
>>>>
>>>> From my understanding, there are three related but distinct concepts
>>>> involved here:
>>>>
>>>> * NUMA node: A system can have sparse NUMA node IDs. For example,
>>>> online nodes can be 0, 2, and 3, with node 1 absent.
>>>>
>>>> * Proximity domain: The proximity domain is the identifier
>>>> associated with a KFD topology device. In the current
>>>> implementation,
>>>> proximity domains for CPU topology devices are assigned using a
>>>> sequential counter, and GPU proximity domains are allocated after
>>>> the
>>>> CPU proximity domains.
>>>>
>>>> * KFD sysfs topology node: KFD exposes each topology device through
>>>> a node under its sysfs topology interface. Userspace reads these
>>>> sysfs
>>>> nodes to discover the CPU/GPU topology and their relationships.
>>>> In the
>>>> current implementation, the sysfs node numbering corresponds to the
>>>> proximity domain numbering.
>>>>
>>>> IIUC, the proximity-domain numbering and KFD sysfs topology-node
>>>> numbering are expected to correspond, while the Linux NUMA node ID can
>>>> be a separate identifier that is mapped to the corresponding topology
>>>> or proximity domain.
>>>>
>>>> As discussed in the thread, we implemented the
>>>> 'NUMA node ==Proximity Domain' approach. However, after implementing
>>>> this for sparse nodes, we found that the userspace libraries using the
>>>> KFD sysfs nodes have an assumption that the topology devices are
>>>> contiguous.
>>>>
>>>> With this approach, discontiguous NUMA nodes result in discontiguous
>>>> proximity-domain and sysfs node numbers. This would break userspace
>>>> libraries that assume the topology devices are contiguous.
>>>>
>>>> To support this approach, we would need to make corresponding changes
>>>> in the userspace libraries, which would break applications using the
>>>> existing userspace and introduce backward-compatibility concerns.
>>>>
>>>> With this in mind, we think the current RFC v1 approach would be the
>>>> better way to proceed, keeping the NUMA node IDs separate from the
>>>> proximity-domain/sysfs numbering.
>>>
>>> I'm not sure about this. We have code in user mode that uses the KFD
>>> proximity domains as CPU NUMA nodes, e.g. in calls to mbind. E.g.
>>> https://github.com/ROCm/rocm-systems/blob/f437a7135fa67b96b48da4baeb72d5f75c406052/projects/rocr-runtime/libhsakmt/src/fmm.c#L2058
>>>
>>>
>>> So if you change the KFD CPU proximity domain numbers to something
>>> contiguous, you're probably breaking user mode either way.
>>
>>
>> Hi Felix,
>>
>> Thanks for pointing this out.
>>
>> I think the approach you originally suggested seems to be the better
>> one:
>> keeping the CPU proximity domain numbers aligned with the NUMA node IDs,
>> including sparse NUMA node IDs.
>>
>> This will require corresponding userspace changes to support
>> non-contiguous NUMA node IDs, but I don't expect this to break any
>> existing applications or introduce backward compatibility issues. For
>> systems with contiguous NUMA node IDs, the existing behavior will work
>> as is, while systems with non-contiguous NUMA node IDs will work after
>> the userspace changes.
>>
>> For example, one of the areas that will need to be updated is here:
>>
>> https://github.com/ROCm/rocm-systems/blob/2d743306cb341b5973cb0d6af7b549569d127724/
>>
>> projects/rocr-runtime/libhsakmt/src/topology.c#L831
>
> The code here is already designed to handle sysfs nodes that are not
> accessible because GPUs are not available in the process' cgroup.
> Maybe we need additional fixes here to handle the case where a
> directory is completely missing in sysfs. The biggest problem is
> probably, that user mode tries to remap the node-IDs into a contiguous
> range here:
> https://github.com/ROCm/rocm-systems/blob/2d743306cb341b5973cb0d6af7b549569d127724/projects/rocr-runtime/libhsakmt/src/topology.c#L836.
> We'd need to avoid that for CPU nodes so that we preserve the identity
> mapping from node IDs to NUMA node IDs. Maybe we can create dummy
> nodes with no CPU cores and no memory as place-holders.
>
>
>>
>>
>> If you have any pointers or suggestions regarding this approach,
>> especially for the userspace changes, I would really appreciate your
>> feedback.
>>
>>
>>
>>>
>>>
>>>
>>>
>>> Can you point out what specific problems you run into with
>>> non-contiguous NUMA nodes topologies?
>>
>>
>> The issue we see on a system with discontiguous NUMA nodes is that the
>> GPU VCRAT parsing fails when we load the GPU driver, resulting in the
>> following errors in dmesg:
>>
>> amdgpu: Virtual CRAT table created for GPU
>> amdgpu: Error parsing VCRAT
>> kfd: amdgpu: Error adding device to topology
>> kfd: amdgpu: Error initializing KFD node
>>
>> On our system, we have 3 NUMA nodes: 0, 2, and 3, with 2 GPUs attached
>> to NUMA node 3.
>
> OK, that's all stuff that would need to be addressed with kernel mode
> driver patches. I'm hoping it would be a simplified version of the
> patches that Donet already proposed.
>
> Regards,
> Felix
Hi Felix,
Thanks for your suggestions.
I'll start working in this direction and post a patch once
it's ready.
Thanks,
Dhruv B
>
>
>>
>> With the current upstream implementation, CPU proximity domains are
>> assigned sequentially. Thus, NUMA nodes 0, 2, and 3 get CPU proximity
>> domains 0, 1, and 2 respectively. The two GPUs then get proximity
>> domains 3 and 4.
>>
>> However, when generating the GPU VCRAT, `proximity_domain_to` is
>> populated directly with the NUMA node ID. Since both GPUs are attached
>> to NUMA node 3, the VCRAT contains:
>>
>> GPU0 -> proximity_domain_to = 3
>> GPU1 -> proximity_domain_to = 3
>>
>> https://elixir.bootlin.com/linux/v7.1/source/drivers/gpu/drm/amd/amdkfd/kfd_crat.c#L2175
>>
>>
>> Here, the CPU proximity domain corresponding to NUMA node 3 is actually
>> 2, while proximity domain 3 belongs to GPU0.
>>
>> During VCRAT parsing, the I/O link parser looks up this `id_to`
>> proximity domain, which is 3, and returns `-ENODEV` because there is no
>> CPU topology device with proximity domain 3:
>>
>> https://elixir.bootlin.com/linux/v7.1/source/drivers/gpu/drm/amd/amdkfd/kfd_crat.c#L1284
>>
>>
>> Thus, the GPU I/O link is referencing the NUMA node ID as if it were
>> the CPU proximity domain, which causes the VCRAT parsing to fail.
>>
>> Since the parsing fails, the driver does not get loaded.
>>
>> Thanks,
>> Dhruv B
>>>
>>>
>>>
>>>
>>>
>>>
>>> Regards,
>>> Felix
>>>
>>>
>>>>
>>>> Would you agree with this approach? If so, I can rebase the RFC v1
>>>> series onto the latest kernel and post a new version for review.
>>>>
>>>> Thanks,
>>>> Dhruv B
>>>>>
>>>>> Regards,
>>>>> Felix
>>>>>
>>>>>
>>>>>>
>>>>>> I think this approach should also resolve the driver loading issue.
>>>>>>
>>>>>>
>>>>>> Thanks
>>>>>> Donet Tom
>>>>>>
>>>>>>
>>>>>>>
>>>>>>> Thanks,
>>>>>>> Felix
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>> The failure is observed as:
>>>>>>>>
>>>>>>>> amdgpu: Virtual CRAT table created for GPU
>>>>>>>> amdgpu: Error parsing VCRAT
>>>>>>>> kfd: amdgpu: Error adding device to topology
>>>>>>>> kfd: amdgpu: Error initializing KFD node
>>>>>>>>
>>>>>>>>
>>>>>>>> Since every online NUMA node is a valid topology object and can
>>>>>>>> contain
>>>>>>>> CPUs, memory, I/O links, or any combination of these, topology
>>>>>>>> devices
>>>>>>>> and proximity domains should be created for every online NUMA node
>>>>>>>> rather than only for NUMA nodes that contain CPUs.
>>>>>>>>
>>>>>>>> To address this, this series introduces a new VCRAT subtype that
>>>>>>>> records the mapping between the NUMA node ID and the generated
>>>>>>>> VCRAT
>>>>>>>> proximity domain. When topology devices are created, this
>>>>>>>> information
>>>>>>>> is stored in the corresponding topology device, allowing the
>>>>>>>> driver to
>>>>>>>> translate a NUMA node ID into its associated proximity domain
>>>>>>>> whenever
>>>>>>>> required.
>>>>>>>>
>>>>>>>> Returning to the previous example, the system contains three
>>>>>>>> online
>>>>>>>> NUMA nodes, so three CPU topology devices and three proximity
>>>>>>>> domains
>>>>>>>> are created, even though only one NUMA node contains CPUs. The
>>>>>>>> NUMA
>>>>>>>> node ID is stored in each topology device together with its
>>>>>>>> generated
>>>>>>>> proximity domain.
>>>>>>>>
>>>>>>>> Later, when the GPU VCRAT is generated, the driver only knows the
>>>>>>>> NUMA
>>>>>>>> node ID to which the GPU is attached (for example, node 3).
>>>>>>>> Instead of
>>>>>>>> assuming that the NUMA node ID is equal to the proximity
>>>>>>>> domain, the
>>>>>>>> driver walks the existing topology devices to locate the
>>>>>>>> corresponding
>>>>>>>> NUMA node and retrieves its generated proximity domain. In this
>>>>>>>> example, NUMA node 3 maps to proximity domain 2, so
>>>>>>>> proximity_domain_to is populated with the correct value.
>>>>>>>>
>>>>>>>> Since the GPU I/O link now references a valid proximity domain,
>>>>>>>> VCRAT
>>>>>>>> parsing completes successfully and topology initialization
>>>>>>>> proceeds
>>>>>>>> without errors on systems with sparse NUMA node IDs, CPU-less NUMA
>>>>>>>> nodes, and CPU-less/memory-less NUMA nodes.
>>>>>>>>
>>>>>>>> This series consists of the following patches:
>>>>>>>>
>>>>>>>> Patch 1 removes an unused argument from
>>>>>>>> kfd_create_crat_image_virtual() as a preparatory cleanup.
>>>>>>>>
>>>>>>>> Patch 2 introduces a new VCRAT NUMA affinity subtype that
>>>>>>>> stores the
>>>>>>>> NUMA node ID and its corresponding proximity domain.
>>>>>>>>
>>>>>>>> Patch 3 populates the NUMA affinity entries during VCRAT
>>>>>>>> generation
>>>>>>>> for every online NUMA node.
>>>>>>>>
>>>>>>>> Patch 4 parses the NUMA affinity entries from the VCRAT and
>>>>>>>> stores the
>>>>>>>> NUMA node ID and proximity domain in the corresponding topology
>>>>>>>> device.
>>>>>>>>
>>>>>>>> Patch 5 fixes GPU VCRAT proximity domain mappings by
>>>>>>>> translating the
>>>>>>>> GPU's NUMA node ID to the corresponding CPU proximity domain
>>>>>>>> before
>>>>>>>> programming proximity_domain_to.
>>>>>>>>
>>>>>>>> Patch 6 creates proximity domains and VCRAT entries for all online
>>>>>>>> NUMA nodes, including CPU-less and memory-less nodes.
>>>>>>>>
>>>>>>>> Patch 7 adds a numa_node sysfs attribute for CPU and GPU topology
>>>>>>>> devices, exposing the associated NUMA node through the topology
>>>>>>>> sysfs
>>>>>>>> interface.
>>>>>>>>
>>>>>>>> Please note that the changes in this series are on a best effort
>>>>>>>> basis from our
>>>>>>>> end. Therefore, requesting the amd-gfx community (who have deeper
>>>>>>>> knowledge of the
>>>>>>>> HW & SW stack) to kindly help with the review and provide feedback
>>>>>>>> / comments on
>>>>>>>> these patches
>>>>>>>>
>>>>>>>> Donet Tom (7):
>>>>>>>> drm/amdgpu: Remove unused argument from
>>>>>>>> kfd_create_crat_image_virtual
>>>>>>>> drm/amdgpu: Add VCRAT NUMA affinity entry
>>>>>>>> drm/amdgpu: Populate NUMA affinity entries in VCRAT
>>>>>>>> drm/amdgpu: Parse NUMA affinity entries from VCRAT
>>>>>>>> drm/amdgpu: Fix VCRAT proximity domain mappings for GPU nodes
>>>>>>>> drm/amdgpu: Create proximity domains and VCRAT entries for
>>>>>>>> CPU-less
>>>>>>>> and memory-less NUMA nodes
>>>>>>>> drm/amdgpu: Add numa_node in cpu/gpu topology device sysfs
>>>>>>>> entry
>>>>>>>>
>>>>>>>> drivers/gpu/drm/amd/amdkfd/kfd_crat.c | 152
>>>>>>>> +++++++++++++++-------
>>>>>>>> drivers/gpu/drm/amd/amdkfd/kfd_crat.h | 19 ++-
>>>>>>>> drivers/gpu/drm/amd/amdkfd/kfd_topology.c | 49 +++++--
>>>>>>>> drivers/gpu/drm/amd/amdkfd/kfd_topology.h | 4 +
>>>>>>>> 4 files changed, 168 insertions(+), 56 deletions(-)
>>>>>>>>
next prev parent reply other threads:[~2026-09-23 7:56 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-04 9:52 [PATCH 0/7] drm/amdgpu: Fix topology device creation and proximity domain mappings for sparse and CPU-less NUMA systems Donet Tom
2026-08-04 9:52 ` [PATCH 1/7] drm/amdgpu: Remove unused argument from kfd_create_crat_image_virtual Donet Tom
2026-08-04 9:52 ` [PATCH 2/7] drm/amdgpu: Add VCRAT NUMA affinity entry Donet Tom
2026-08-04 9:52 ` [PATCH 3/7] drm/amdgpu: Populate NUMA affinity entries in VCRAT Donet Tom
2026-08-04 9:52 ` [PATCH 4/7] drm/amdgpu: Parse NUMA affinity entries from VCRAT Donet Tom
2026-08-04 9:52 ` [PATCH 5/7] drm/amdgpu: Fix VCRAT proximity domain mappings for GPU nodes Donet Tom
2026-08-04 9:52 ` [PATCH 6/7] drm/amdgpu: Create proximity domains and VCRAT entries for CPU-less and memory-less NUMA nodes Donet Tom
2026-08-04 9:52 ` [PATCH 7/7] drm/amdgpu: Add numa_node in cpu/gpu topology device sysfs entry Donet Tom
2026-08-04 22:27 ` [PATCH 0/7] drm/amdgpu: Fix topology device creation and proximity domain mappings for sparse and CPU-less NUMA systems Felix Kuehling
2026-08-05 7:09 ` Donet Tom
2026-08-05 16:16 ` Kuehling, Felix
2026-08-05 17:12 ` Donet Tom
2026-09-21 4:56 ` Dhruv Bhogaonkar
2026-09-21 17:05 ` Kuehling, Felix
2026-09-22 16:59 ` Dhruv Bhogaonkar
2026-09-22 22:49 ` Felix Kuehling
2026-09-23 5:58 ` Dhruv Bhogaonkar [this message]
2026-09-25 14:13 ` Dhruv Bhogaonkar
2026-10-02 21:14 ` Felix Kuehling
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cafaaba4-7d45-46ef-b0bf-67fe8cfa70ac@linux.ibm.com \
--to=dhruv.b@linux.ibm.com \
--cc=David.YatSin@amd.com \
--cc=Kent.Russell@amd.com \
--cc=airlied@gmail.com \
--cc=alexander.deucher@amd.com \
--cc=alexdeucher@gmail.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
--cc=donettom@linux.ibm.com \
--cc=felix.kuehling@amd.com \
--cc=ritesh.list@gmail.com \
--cc=simona@ffwll.ch \
--cc=svaidy@linux.ibm.com \
--cc=yangp@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox