All of lore.kernel.org
 help / color / mirror / Atom feed
From: Felix Kuehling <felix.kuehling@amd.com>
To: "Xiaogang.Chen" <xiaogang.chen@amd.com>, amd-gfx@lists.freedesktop.org
Subject: Re: [PATCH v2 2/4] drm/amdkfd: Change migration size in CPU/GPU page fault handler to THP size
Date: Tue, 6 Oct 2026 17:43:54 -0400	[thread overview]
Message-ID: <528f1f07-cece-4bec-96e2-7c43e335bbd3@amd.com> (raw)
In-Reply-To: <20260904195421.42919-3-xiaogang.chen@amd.com>

On 2026-09-04 15:54, Xiaogang.Chen wrote:
> From: Xiaogang Chen <xiaogang.chen@amd.com>
>
> When use HPAGE_PMD_SIZE based device private pages during migration core HMM
> treats device private memory in HPAGE_PMD_SIZE compound folio if possible.
> Current kfd driver uses prange->granularity that can be changed by user. Need
> have migration size in CPU and GPU page fault handler in HPAGE_PMD_SIZE based.
>
> For AMD GPU that exposes private device memory choose HPAGE_PMD_SIZE as
> minimums migration size in CPU and GPU page fault handler. For x86 it is
> same as default prange->granularity.

As we discussed before, this breaks the API semantics. The default 
setting is fine for allowing THP. But we want applications to be able to 
use smaller granularity for use cases that access or distribute data 
between multiple devices at finer granularity. In those cases, the 
additional memory management and TLB overhead is offset by reduced 
migrations or thrashing.

Please just drop this patch.

Regards,
   Felix


>
> Signed-off-by: Xiaogang Chen <xiaogang.chen@amd.com>
> ---
>   drivers/gpu/drm/amd/amdkfd/kfd_migrate.c |  6 ++++--
>   drivers/gpu/drm/amd/amdkfd/kfd_svm.c     | 13 +++++++++++--
>   drivers/gpu/drm/amd/amdkfd/kfd_svm.h     | 12 ++++++++++++
>   3 files changed, 27 insertions(+), 4 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_migrate.c b/drivers/gpu/drm/amd/amdkfd/kfd_migrate.c
> index 813f3c1d29dc..bbf0fefd5722 100644
> --- a/drivers/gpu/drm/amd/amdkfd/kfd_migrate.c
> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_migrate.c
> @@ -1026,8 +1026,10 @@ static vm_fault_t svm_migrate_to_ram(struct vm_fault *vmf)
>   	if (!prange->actual_loc)
>   		goto out_unlock_prange;
>   
> -	/* Align migration range start and size to granularity size */
> -	size = 1UL << prange->granularity;
> +	/* Align migration range start and size to max of
> +	 * THP with HPAGE_PMD_ORDER and granularity size
> +	 */
> +	size = 1UL << max(prange->granularity, HPAGE_PMD_ORDER);
>   	start = max(ALIGN_DOWN(addr, size), prange->start);
>   	last = min(ALIGN(addr + 1, size) - 1, prange->last);
>   
> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_svm.c b/drivers/gpu/drm/amd/amdkfd/kfd_svm.c
> index 6b783d12bce4..482cd4e7eee5 100644
> --- a/drivers/gpu/drm/amd/amdkfd/kfd_svm.c
> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_svm.c
> @@ -3071,6 +3071,7 @@ svm_range_restore_pages(struct amdgpu_device *adev, unsigned int pasid,
>   	struct kfd_node *node;
>   	int32_t best_loc;
>   	int32_t gpuid, gpuidx = MAX_GPU_INSTANCE;
> +	bool is_private_device = false;
>   	bool write_locked = false;
>   	struct vm_area_struct *vma;
>   	bool migration = false;
> @@ -3087,6 +3088,7 @@ svm_range_restore_pages(struct amdgpu_device *adev, unsigned int pasid,
>   		return 0;
>   	}
>   	svms = &p->svms;
> +	is_private_device = svm_is_private_zone(adev);
>   
>   	pr_debug("restoring svms 0x%p fault address 0x%llx\n", svms, addr);
>   
> @@ -3224,8 +3226,15 @@ svm_range_restore_pages(struct amdgpu_device *adev, unsigned int pasid,
>   	kfd_smi_event_page_fault_start(node, p->lead_thread, addr,
>   				       write_fault, timestamp);
>   
> -	/* Align migration range start and size to granularity size */
> -	size = 1UL << prange->granularity;
> +	if (is_private_device)
> +		/* Align migration range start and size to max of
> +		 * THP and granularity size
> +		 */
> +		size = 1UL << max(prange->granularity, HPAGE_PMD_ORDER);
> +	else
> +		/* Align migration range start and size to granularity size */
> +		size = 1UL << prange->granularity;
> +
>   	start = max_t(unsigned long, ALIGN_DOWN(addr, size), prange->start);
>   	last = min_t(unsigned long, ALIGN(addr + 1, size) - 1, prange->last);
>   	if (prange->actual_loc != 0 || best_loc != 0) {
> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_svm.h b/drivers/gpu/drm/amd/amdkfd/kfd_svm.h
> index f2b3a05cd8cf..e78ee94ba33c 100644
> --- a/drivers/gpu/drm/amd/amdkfd/kfd_svm.h
> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_svm.h
> @@ -215,6 +215,13 @@ void svm_range_bo_unref_async(struct svm_range_bo *svm_bo);
>   void svm_range_set_max_pages(struct amdgpu_device *adev);
>   int svm_range_switch_xnack_reserve_mem(struct kfd_process *p, bool xnack_enabled);
>   
> +/* check adev has device private zone memory */
> +static inline bool svm_is_private_zone(struct amdgpu_device *adev)
> +{
> +	struct amdgpu_kfd_dev *kfddev = &adev->kfd;
> +	return (kfddev->pgmap.type == MEMORY_DEVICE_PRIVATE);
> +}
> +
>   #else
>   
>   struct kfd_process;
> @@ -277,6 +284,11 @@ static inline void svm_range_set_max_pages(struct amdgpu_device *adev)
>   {
>   }
>   
> +static inline bool svm_is_private_zone(struct amdgpu_device *adev)
> +{
> +	return false;
> +}
> +
>   #define KFD_IS_SVM_API_SUPPORTED(dev) false
>   
>   #endif /* IS_ENABLED(CONFIG_HSA_AMD_SVM) */

  reply	other threads:[~2026-10-06 21:44 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 19:54 [PATCH v2 0/4] drm/amdkfd: Enable device private memory THP support in kfd svm driver Xiaogang.Chen
2026-09-04 19:54 ` [PATCH v2 1/4] drm/amdkfd: Add awareness of THP of device and system RAM " Xiaogang.Chen
2026-10-06 21:40   ` Felix Kuehling
2026-10-06 22:49     ` Chen, Xiaogang
2026-10-06 23:01       ` Felix Kuehling
2026-09-04 19:54 ` [PATCH v2 2/4] drm/amdkfd: Change migration size in CPU/GPU page fault handler to THP size Xiaogang.Chen
2026-10-06 21:43   ` Felix Kuehling [this message]
2026-09-04 19:54 ` [PATCH v2 3/4] drm/amdkfd: Apply HMM THP zone device-private memory migration in kfd driver Xiaogang.Chen
2026-10-06 22:44   ` Felix Kuehling
2026-10-07 18:57     ` Chen, Xiaogang
2026-09-04 19:54 ` [PATCH v2 4/4] drm/amdkfd: Apply AMDGPU_PTE_FRAG to pte of gart page table for THP mapping Xiaogang.Chen

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=528f1f07-cece-4bec-96e2-7c43e335bbd3@amd.com \
    --to=felix.kuehling@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=xiaogang.chen@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.