From: sashiko-bot@kernel.org
To: "Mika Kuoppala" <mika.kuoppala@linux.intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v10 27/27] drm/xe/eudebug: Enable EU pagefault handling
Date: Thu, 03 Sep 2026 15:46:45 +0000 [thread overview]
Message-ID: <20260903154646.687761F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260903145952.848051-28-mika.kuoppala@linux.intel.com>
Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] Temporary NULL VMAs for subsequent eudebug pagefaults are destroyed immediately after ACKing the fault, creating a race condition that leads to livelock.
- [High] Successful SVM pagefaults bypass private tagging, which permanently stalls the finalization of the active eudebug session.
- [High] Pagefault cache hits bypass eudebug tagging, permanently hanging the debugger session.
- [High] `xe_vm_svm_vma_subtract()` is called with unaligned address boundaries, corrupting the VM layout.
--
commit bd171058301961cced31067feb0527d9f0b38c87
Author: Gwan-gyeong Mun <gwan-gyeong.mun@intel.com>
drm/xe/eudebug: Enable EU pagefault handling
This commit enables EU pagefault handling for XE2/PVC hardware, allowing
the driver to unhalt pagefaulted EU threads via SIP mode and a temporary
null page.
> diff --git a/drivers/gpu/drm/xe/xe_guc_pagefault.c b/drivers/gpu/drm/xe/xe_guc_pagefault.c
> --- a/drivers/gpu/drm/xe/xe_guc_pagefault.c
> +++ b/drivers/gpu/drm/xe/xe_guc_pagefault.c
> @@ -59,13 +60,14 @@ static void guc_ack_fault(struct xe_pagefault *pf, int err)
> write_only);
> }
>
> -static void guc_ack_fault_end(void *private)
> +static void guc_ack_fault_end(void *private, int err)
> {
> - struct xe_guc *guc = private;
> + struct xe_guc *guc = xe_eudebug_pagefault_get_private(private);
>
> if ((guc->pagefault_ack_counter & (XE_GUC_PAGEFAULT_FLUSH_PERIOD - 1)) != 1)
> xe_guc_ct_send_flush(&guc->ct);
> xe_guc_ct_unlock(&guc->ct);
> + xe_eudebug_pagefault_end(private, err);
[Severity: High]
Does this create a race condition that leads to a livelock?
When this calls xe_eudebug_pagefault_end(), temporary null VMAs for
subsequent pagefaults are destroyed immediately (destroy_pf = 1) after
ACKing the fault and signaling the hardware to resume.
Since there is no synchronization to guarantee the EU thread has re-executed
the faulting instruction before the KMD unbinds the VMA, if the VMA is
unbound before the EU thread accesses the address, won't the thread
pagefault again on the same instruction and loop infinitely?
> }
>
> diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
> --- a/drivers/gpu/drm/xe/xe_pagefault.c
> +++ b/drivers/gpu/drm/xe/xe_pagefault.c
[ ... ]
> @@ -227,8 +228,56 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
>
> vma = xe_vm_find_vma_by_addr(vm, pf->consumer.page_addr);
> if (!vma) {
[ ... ]
> + } else {
> + /*
> + * For non-SVM case:
> + * Eudebug with active pagefault always needs to be attached
> + * to pagefault since it waits for all pagefaults with matching
> + * asid to be resolved.
> + */
> + if (!xe_vma_is_cpu_addr_mirror(vma))
> + xe_eudebug_pagefault_set_private(pf, vm);
[Severity: High]
Does skipping the private tagging for SVM (cpu_addr_mirror) VMAs cause the
eudebug session to hang permanently?
If an active eudebug session exists for the same ASID and its finalization
was postponed because this SVM fault was still in the queue,
xe_eudebug_pagefault_end() will see no eudebug tag and return immediately
without triggering eudebug_pagefault_finalize(). This seems like it would
leave the EU threads permanently stuck in SIP mode.
> + }
>
> if (xe_vma_read_only(vma) &&
> @@ -239,11 +288,52 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
>
> atomic = xe_pagefault_access_is_atomic(pf->consumer.access_type);
>
> - if (xe_vma_is_cpu_addr_mirror(vma))
> + if (xe_vma_is_cpu_addr_mirror(vma)) {
> err = xe_svm_handle_pagefault(vm, vma, pf, gt,
> pf->consumer.page_addr, atomic);
> - else
> +
> +#if IS_ENABLED(CONFIG_DRM_XE_EUDEBUG)
> + /*
> + * If err is -ENOENT, it means that the cpu-address-space-mirrored
> + * xe vma exists, but there is no mm vma allocated in
> + * the CPU address space. This indicates that no memory has been
> + * allocated in the CPU address space.
> + */
> + if (err == -ENOENT &&
> + !xe_vm_is_closed_or_banned(vm) &&
> + xe_eudebug_pagefault_creatable(gt, vm)) {
> + u32 page_size = vm->flags & XE_VM_FLAG_64K ? SZ_64K : SZ_4K;
[ ... ]
> + vma = xe_vm_find_vma_by_addr(vm, pf->consumer.page_addr);
> + if (vma && xe_vma_is_cpu_addr_mirror(vma))
> + vma = xe_vm_svm_vma_subtract(vm, vma,
> + pf->consumer.page_addr,
> + pf->consumer.page_addr + page_size);
[Severity: High]
Is it safe to pass pf->consumer.page_addr directly as the subtraction
boundaries?
The hardware fault address is not guaranteed to be aligned to page_size.
Passing unaligned boundaries to xe_vm_svm_vma_subtract() will create
unaligned head/tail VMAs via xe_vma_create().
Inserting unaligned VMAs into the interval tree and page tables violates
core memory management invariants and could lead to page table corruption or
kernel panics.
> + else
> + vma = ERR_PTR(-EINVAL);
[ ... ]
> @@ -598,6 +688,12 @@ static void xe_pagefault_queue_work(struct work_struct *w)
> }
>
> ack_fault:
> + /*
> + * set private after xe_pagefault_service() since eudebug could swap
> + * the pf->producer.private field. Also needed when cache was hit.
> + */
> + private = pf->producer.private;
[Severity: High]
Does a pagefault cache hit permanently hang the eudebug session?
When a pagefault matches the cache, the code jumps directly to ack_fault,
skipping xe_pagefault_service(). Since the service routine is bypassed,
xe_eudebug_pagefault_set_private() is never called, leaving the fault
untagged.
Similar to the SVM bypass above, if the session postponed finalization for
this queued fault, it will never trigger eudebug_pagefault_finalize() upon
completion.
> +
> xe_assert(xe, pf->consumer.alloc_state ==
> XE_PAGEFAULT_ALLOC_STATE_ACTIVE);
> xe_assert(xe, pf == pf_work->cache.pf);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260903145952.848051-1-mika.kuoppala@linux.intel.com?part=27
next prev parent reply other threads:[~2026-09-03 15:46 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 14:59 [PATCH v10 00/27] Intel Xe GPU Debug Support (eudebug) v10 Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 01/27] drm/xe/eudebug: Introduce eudebug interface Mika Kuoppala
2026-09-03 15:16 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 02/27] drm/xe/eudebug: Add documentation Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 03/27] drm/xe/eudebug: Add connection establishment documentation Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 04/27] drm/xe/eudebug: Introduce discovery for resources Mika Kuoppala
2026-09-03 15:22 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 05/27] drm/xe: Add EUDEBUG_ENABLE exec queue property Mika Kuoppala
2026-09-03 15:14 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 06/27] drm/xe/eudebug: Introduce exec_queue events Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 07/27] drm/xe/eudebug: Mark guc contexts as debuggable Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 08/27] drm/xe: Remove ifdef in DRM_GPUVA_OP_DRIVER svm subop checking Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 09/27] drm/xe: Introduce ADD_DEBUG_DATA and REMOVE_DEBUG_DATA vm bind ops Mika Kuoppala
2026-09-03 15:22 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 10/27] drm/xe/eudebug: Introduce vm bind and vm bind debug data events Mika Kuoppala
2026-09-03 15:26 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 11/27] drm/xe/eudebug: Add ufence events with acks Mika Kuoppala
2026-09-03 15:20 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 12/27] drm/xe/eudebug: Add vm open/pread/pwrite Mika Kuoppala
2026-09-03 15:27 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 13/27] drm/xe/eudebug: Add userptr vm pread/pwrite Mika Kuoppala
2026-09-03 15:24 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 14/27] drm/xe/eudebug: Add hw enablement Mika Kuoppala
2026-09-03 15:15 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 15/27] drm/xe/eudebug: Introduce EU control interface Mika Kuoppala
2026-09-03 15:34 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 16/27] drm/xe/eudebug: Introduce per device attention scan worker Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 17/27] drm/xe/eudebug_test: Introduce eudebug live tests Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 18/27] drm/xe: Implement SR-IOV and eudebug exclusivity Mika Kuoppala
2026-09-03 15:32 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 19/27] drm/xe: Add xe_client_debugfs and introduce debug_data file Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 20/27] drm/xe/pagefault: export pagefault queue properties Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 21/27] drm/xe/eudebug: Add read/count/compare helper for eu attention Mika Kuoppala
2026-09-03 15:31 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 22/27] drm/xe/vm: Support for adding null page VMA to VM on request Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 23/27] drm/xe/vm: Add xe_vm_svm_vma_subtract() to carve out a sub-range from an SVM VMA Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 24/27] drm/xe: Support for xe_vma_unbind() Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 25/27] drm/xe: export prep_vma_destroy as xe_vm_prep_vma_destroy Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 26/27] drm/xe/eudebug: Introduce EU pagefault handling interface Mika Kuoppala
2026-09-03 15:43 ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 27/27] drm/xe/eudebug: Enable EU pagefault handling Mika Kuoppala
2026-09-03 15:46 ` sashiko-bot [this message]
2026-09-03 15:35 ` ✗ CI.checkpatch: warning for Intel Xe GPU Debug Support (eudebug) v10 Patchwork
2026-09-03 15:37 ` ✓ CI.KUnit: success " Patchwork
2026-09-03 15:53 ` ✗ CI.checksparse: warning " Patchwork
2026-09-03 16:17 ` ✓ Xe.CI.BAT: success " Patchwork
2026-09-03 16:30 ` [PATCH v10 00/27] " Rodrigo Vivi
2026-09-04 3:21 ` ✗ Xe.CI.FULL: failure for " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260903154646.687761F000E9@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=mika.kuoppala@linux.intel.com \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox