All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Himal Prasad Ghimiray" <himal.prasad.ghimiray@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [RFC v2 11/21] drm/xe/svm: Handle svm vma for acc_ctr trigger
Date: Wed, 09 Sep 2026 12:52:22 +0000	[thread overview]
Message-ID: <20260909125223.56A941F00A3A@smtp.kernel.org> (raw)
In-Reply-To: <20260909124457.1610890-34-himal.prasad.ghimiray@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] The early graceful bailout for non-migratable ranges on access counter triggers is unreachable on discrete GPUs, causing an incorrect `-EACCES` error return.
- [Medium] Valid page fault statistics are incorrectly incremented for access counter triggers, skewing hardware telemetry.
--

commit 67501dc98da68ff714f7eec618d994bd502ef87c
Author: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Subject: drm/xe/svm: Handle svm vma for acc_ctr trigger

This commit generalizes page fault handling into a range setup function to
support hardware access-counter triggers. It aims to skip page-fault-specific
bookkeeping and forcefully migrate triggered ranges to device memory if
possible.

> diff --git a/drivers/gpu/drm/xe/xe_svm.c b/drivers/gpu/drm/xe/xe_svm.c
> --- a/drivers/gpu/drm/xe/xe_svm.c
> +++ b/drivers/gpu/drm/xe/xe_svm.c
[ ... ]
> @@ -1314,7 +1316,8 @@ static int __xe_svm_handle_pagefault(struct xe_vm *vm, struct xe_vma *vma,
>  	if (IS_ERR(range))
>  		return PTR_ERR(range);
>  
> -	xe_svm_range_fault_count_stats_incr(gt, range);
> +	if (!acc_ctr_trigger)
> +		xe_svm_range_fault_count_stats_incr(gt, range);

[Severity: Medium]
While this successfully skips the primary page fault stat increment, does it
leave the valid page fault statistics skewed for access counter triggers?

If the range is already valid, __xe_svm_range_setup() evaluates
xe_svm_range_is_valid() and unconditionally increments the valid fault stat:

        if (xe_svm_range_is_valid(range, tile, ctx.devmem_only, dpagemap)) {
                xe_svm_range_valid_fault_count_stats_incr(gt, range);
                range_debug(range, acc_ctr_trigger ? ...
                goto out;
        }

Should xe_svm_range_valid_fault_count_stats_incr() also be guarded by
!acc_ctr_trigger to prevent falsely inflating hardware telemetry during normal
operation?

>  
>  	mutex_lock(&range->lock);
>  
> @@ -1336,6 +1339,10 @@ static int __xe_svm_handle_pagefault(struct xe_vm *vm, struct xe_vma *vma,
>  
>  	range_debug(range, "PAGE FAULT");
>  
> +	if (acc_ctr_trigger && !range_flags.migrate_devmem) {
> +		goto out;
> +	}
> +

[Severity: Medium]
Will this newly added graceful bailout ever be reached on discrete GPUs?

Because ctx.devmem_only is initialized to (need_vram || acc_ctr_trigger) &&
devmem_possible, it evaluates to true during an access counter trigger on a
discrete GPU.

Earlier in __xe_svm_range_setup(), there is a pre-existing check before
reaching this code block:

        if (ctx.devmem_only && !range_flags.migrate_devmem) {
                err = -EACCES;
                goto err_out;
        }

Since both ctx.devmem_only and !range_flags.migrate_devmem are true for
non-migratable ranges during an access counter trigger, the function appears
to immediately return -EACCES. Does this bypass the intended graceful exit
added here, creating dead code and returning an incorrect error instead?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260909124457.1610890-23-himal.prasad.ghimiray@intel.com?part=11

  reply	other threads:[~2026-09-09 12:52 UTC|newest]

Thread overview: 32+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-09 12:44 [RFC v2 00/21] drm/xe: Access counter support for migration hints Himal Prasad Ghimiray
2026-09-09 12:44 ` [RFC v2 01/21] drm/xe: Add xe_usm_queue generic USM circular buffer Himal Prasad Ghimiray
2026-09-09 12:51   ` sashiko-bot
2026-09-09 12:44 ` [RFC v2 02/21] drm/xe: Stub out new access_counter layer Himal Prasad Ghimiray
2026-09-09 12:44 ` [RFC v2 03/21] drm/xe: Implement xe_access_counter_init Himal Prasad Ghimiray
2026-09-09 12:55   ` sashiko-bot
2026-09-09 12:44 ` [RFC v2 04/21] drm/xe: Implement xe_access_counter_handler Himal Prasad Ghimiray
2026-09-09 12:58   ` sashiko-bot
2026-09-09 12:44 ` [RFC v2 05/21] drm/xe: Extract xe_vma_lock_and_validate helper Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 06/21] drm/xe: Move ASID to FAULT VM lookup to xe_device Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 07/21] drm/xe/pf: Use xe_device_asid_to_vm in xe_pagefault_save_to_vm Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 08/21] drm/xe: Implement xe_access_counter_queue_work Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 09/21] drm/xe: Implement xe_access_counter_service Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 10/21] drm/xe/trace: Add xe_vma_acc trace event for access counter notifications Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 11/21] drm/xe/svm: Handle svm vma for acc_ctr trigger Himal Prasad Ghimiray
2026-09-09 12:52   ` sashiko-bot [this message]
2026-09-09 12:45 ` [RFC v2 12/21] drm/xe: Service all VMAs in an access counter granularity window Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 13/21] drm/xe: Add xe_guc_access_counter layer Himal Prasad Ghimiray
2026-09-09 12:54   ` sashiko-bot
2026-09-09 12:45 ` [RFC v2 14/21] drm/xe/uapi: Add access counter parameter extension for exec queue Himal Prasad Ghimiray
2026-09-09 12:52   ` sashiko-bot
2026-09-09 12:45 ` [RFC v2 15/21] drm/xe/lrc: Pass exec_queue to xe_lrc_create for access counter params Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 16/21] drm/xe/vm: Add xe_vma_supports_access_ctr() helper Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 17/21] drm/xe/pt: Set NC PTE bit for VMAs ineligible for access counting Himal Prasad Ghimiray
2026-09-09 12:56   ` sashiko-bot
2026-09-09 12:45 ` [RFC v2 18/21] drm/xe/svm: Define access counter migration policy Himal Prasad Ghimiray
2026-09-09 13:01   ` sashiko-bot
2026-09-09 12:45 ` [RFC v2 19/21] drm/xe/svm: Add MIGRATE_ON_ACCESS_COUNTER bind flag Himal Prasad Ghimiray
2026-09-09 12:57   ` sashiko-bot
2026-09-09 12:45 ` [RFC v2 20/21] drm/xe/svm: Move EVICTED PAGES debug log to callers Himal Prasad Ghimiray
2026-09-09 12:45 ` [RFC v2 21/21] drm/xe/svm: Distinguish access-counter-triggered range setup in logs Himal Prasad Ghimiray
2026-09-09 12:58   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260909125223.56A941F00A3A@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=himal.prasad.ghimiray@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.