* [PATCH 0/1] drm/xe: fixing endless PF range response
@ 2026-10-06 21:34 fei.yang
2026-10-06 21:34 ` [PATCH 1/1] " fei.yang
2026-10-06 21:36 ` ✗ CI.KUnit: failure for " Patchwork
0 siblings, 2 replies; 4+ messages in thread
From: fei.yang @ 2026-10-06 21:34 UTC (permalink / raw)
To: intel-xe; +Cc: matthew.brost, Fei Yang
From: Fei Yang <fei.yang@intel.com>
xe4_guc_ack_fault_range() reads vma_info, which is not updated by the
SVM handler. This causes the ack fault to loop through the whole address
space with 1G pages, and that is wrong and the loop takes a long while
to complete.
Fei Yang (1):
drm/xe: fixing endless PF range response
drivers/gpu/drm/xe/xe_pagefault.c | 7 ++++++-
drivers/gpu/drm/xe/xe_pagefault.h | 2 ++
2 files changed, 8 insertions(+), 1 deletion(-)
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 1/1] drm/xe: fixing endless PF range response
2026-10-06 21:34 [PATCH 0/1] drm/xe: fixing endless PF range response fei.yang
@ 2026-10-06 21:34 ` fei.yang
2026-10-06 21:36 ` sashiko-bot
2026-10-06 21:36 ` ✗ CI.KUnit: failure for " Patchwork
1 sibling, 1 reply; 4+ messages in thread
From: fei.yang @ 2026-10-06 21:34 UTC (permalink / raw)
To: intel-xe; +Cc: matthew.brost, Fei Yang
From: Fei Yang <fei.yang@intel.com>
[334.635636] PF range response: start=0x80200000 end=0xc0200000 pages=262144 (more ranges pending)
...
[519.065737] PF range response: start=0x5ca140200000 ... (still going, 94888 messages logged)
SVM handler narrows the serviced range through functions
xe_pagefault_set_[start|end]_addr, which updates only the
consumer.[page|end]_addr. However a different field is used
for "immediate ack" cache, xe4_guc_ack_fault_range() reads
vma_info. This causes the ack fault to loop through the
whole address space with 1G pages, and that is wrong and
the loop takes a long while to complete.
VLK-99796
Assisted-by: GitHub_Copilot:claude-opus-5
Signed-off-by: Fei Yang <fei.yang@intel.com>
---
drivers/gpu/drm/xe/xe_pagefault.c | 7 ++++++-
drivers/gpu/drm/xe/xe_pagefault.h | 2 ++
2 files changed, 8 insertions(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
index 5247321c805ec..72dfe27025736 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_pagefault.c
@@ -694,6 +694,8 @@ static void xe_pagefault_queue_work(struct work_struct *w)
/* Last fault same address, ack immediately */
if (xe_pagefault_match(pf, cache_start, cache_end, cache_asid)) {
xe_gt_stats_incr(gt, XE_GT_STATS_ID_LAST_PAGEFAULT_COUNT, 1);
+ xe_pagefault_set_start_addr(pf, cache_start);
+ xe_pagefault_set_end_addr(pf, cache_end);
goto ack_fault;
}
@@ -751,9 +753,12 @@ static void xe_pagefault_queue_work(struct work_struct *w)
pf = xe_pagefault_queue_unchain_requeue(pf_queue, pf, gt);
/* Ensure resets are safe */
- if (pf)
+ if (pf) {
pf->consumer.alloc_state =
XE_PAGEFAULT_ALLOC_STATE_ACTIVE;
+ xe_pagefault_set_start_addr(pf, cache_start);
+ xe_pagefault_set_end_addr(pf, cache_end);
+ }
spin_unlock_irq(&pf_queue->lock);
}
diff --git a/drivers/gpu/drm/xe/xe_pagefault.h b/drivers/gpu/drm/xe/xe_pagefault.h
index 799c984dee84f..5ce2f6ed10c56 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.h
+++ b/drivers/gpu/drm/xe/xe_pagefault.h
@@ -98,6 +98,7 @@ xe_pagefault_set_end_addr(struct xe_pagefault *pf, u64 end_addr)
{
pf->consumer.end_addr &= ~XE_PAGEFAULT_END_ADDR_MASK;
pf->consumer.end_addr |= end_addr;
+ pf->consumer.vma_info.end_addr = end_addr;
}
/**
@@ -137,6 +138,7 @@ xe_pagefault_set_start_addr(struct xe_pagefault *pf, u64 start_addr)
{
pf->consumer.page_addr &= XE_PAGEFAULT_ERROR_MASK;
pf->consumer.page_addr |= (start_addr & ~XE_PAGEFAULT_ERROR_MASK);
+ pf->consumer.vma_info.start_addr = start_addr;
}
/**
--
2.43.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* ✗ CI.KUnit: failure for drm/xe: fixing endless PF range response
2026-10-06 21:34 [PATCH 0/1] drm/xe: fixing endless PF range response fei.yang
2026-10-06 21:34 ` [PATCH 1/1] " fei.yang
@ 2026-10-06 21:36 ` Patchwork
1 sibling, 0 replies; 4+ messages in thread
From: Patchwork @ 2026-10-06 21:36 UTC (permalink / raw)
To: fei.yang; +Cc: intel-xe
== Series Details ==
Series: drm/xe: fixing endless PF range response
URL : https://patchwork.freedesktop.org/series/175636/
State : failure
== Summary ==
+ trap cleanup EXIT
+ kunitconfigs=('/kernel/drivers/gpu/tests/.kunitconfig' '/kernel/drivers/gpu/drm/xe/.kunitconfig' '/kernel/drivers/gpu/drm/tests/.kunitconfig' '/kernel/drivers/gpu/drm/ttm/tests/.kunitconfig' '/kernel/drivers/dma-buf/.kunitconfig')
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/tests/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/tests/.kunitconfig
[21:35:52] Configuring KUnit Kernel ...
Generating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[21:35:57] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[21:36:17] Starting KUnit Kernel (1/1)...
[21:36:17] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[21:36:17] ================= gpu_buddy (14 subtests) ==================
[21:36:17] [PASSED] gpu_test_buddy_alloc_limit
[21:36:17] [PASSED] gpu_test_buddy_alloc_optimistic
[21:36:17] [PASSED] gpu_test_buddy_alloc_pessimistic
[21:36:17] [PASSED] gpu_test_buddy_alloc_pathological
[21:36:17] [PASSED] gpu_test_buddy_alloc_contiguous
[21:36:17] [PASSED] gpu_test_buddy_alloc_clear
[21:36:17] [PASSED] gpu_test_buddy_alloc_range
[21:36:17] [PASSED] gpu_test_buddy_alloc_range_bias
[21:36:18] [PASSED] gpu_test_buddy_fragmentation_performance
[21:36:19] [PASSED] gpu_test_buddy_dirty_tracker_performance
[21:36:19] [PASSED] gpu_test_buddy_alloc_exceeds_max_order
[21:36:19] [PASSED] gpu_test_buddy_offset_aligned_allocation
[21:36:19] [PASSED] gpu_test_buddy_subtree_offset_alignment_stress
[21:36:19] [PASSED] gpu_test_buddy_addr_to_block
[21:36:19] ==================== [PASSED] gpu_buddy ====================
[21:36:19] ============================================================
[21:36:19] Testing complete. Ran 14 tests: passed: 14
[21:36:19] Elapsed time: 26.641s total, 4.466s configuring, 20.307s building, 1.822s running
+ for kcfg in "${kunitconfigs[@]}"
+ [[ ! -f /kernel/drivers/gpu/drm/xe/.kunitconfig ]]
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/xe/.kunitconfig
ERROR:root:In file included from ../drivers/gpu/drm/xe/xe_device.c:59:
../drivers/gpu/drm/xe/xe_pagefault.h: In function ‘xe_pagefault_set_end_addr’:
../drivers/gpu/drm/xe/xe_pagefault.h:101:21: error: ‘struct <anonymous>’ has no member named ‘vma_info’
101 | pf->consumer.vma_info.end_addr = end_addr;
| ^
../drivers/gpu/drm/xe/xe_pagefault.h: In function ‘xe_pagefault_set_start_addr’:
../drivers/gpu/drm/xe/xe_pagefault.h:141:21: error: ‘struct <anonymous>’ has no member named ‘vma_info’
141 | pf->consumer.vma_info.start_addr = start_addr;
| ^
make[7]: *** [../scripts/Makefile.build:290: drivers/gpu/drm/xe/xe_device.o] Error 1
make[7]: *** Waiting for unfinished jobs....
make[6]: *** [../scripts/Makefile.build:551: drivers/gpu/drm/xe] Error 2
make[6]: *** Waiting for unfinished jobs....
make[5]: *** [../scripts/Makefile.build:551: drivers/gpu/drm] Error 2
make[4]: *** [../scripts/Makefile.build:551: drivers/gpu] Error 2
make[4]: *** Waiting for unfinished jobs....
make[3]: *** [../scripts/Makefile.build:551: drivers] Error 2
make[3]: *** Waiting for unfinished jobs....
make[2]: *** [/kernel/Makefile:2229: .] Error 2
make[1]: *** [/kernel/Makefile:248: __sub-make] Error 2
make: *** [Makefile:248: __sub-make] Error 2
[21:36:19] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[21:36:21] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
+ cleanup
++ stat -c %u:%g /kernel
+ chown -R 1003:1003 /kernel
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH 1/1] drm/xe: fixing endless PF range response
2026-10-06 21:34 ` [PATCH 1/1] " fei.yang
@ 2026-10-06 21:36 ` sashiko-bot
0 siblings, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-10-06 21:36 UTC (permalink / raw)
To: fei.yang; +Cc: intel-xe
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] drm/xe: build failure due to hallucinated vma_info field
--
commit 634f5ddbb1931a170d6586933045107086b28191
Author: Fei Yang <fei.yang@intel.com>
Subject: drm/xe: fixing endless PF range response
This commit modifies the page fault queue worker to immediately update the
start and end addresses of a page fault. It also adds assignments to a new
vma_info structure inside the page fault consumer.
> diff --git a/drivers/gpu/drm/xe/xe_pagefault.h b/drivers/gpu/drm/xe/xe_pagefault.h
> index 799c984dee84f..5ce2f6ed10c56 100644
> --- a/drivers/gpu/drm/xe/xe_pagefault.h
> +++ b/drivers/gpu/drm/xe/xe_pagefault.h
> @@ -98,6 +98,7 @@ xe_pagefault_set_end_addr(struct xe_pagefault *pf, u64 end_addr)
> {
> pf->consumer.end_addr &= ~XE_PAGEFAULT_END_ADDR_MASK;
> pf->consumer.end_addr |= end_addr;
> + pf->consumer.vma_info.end_addr = end_addr;
[Severity: Low]
Does the consumer structure actually have a vma_info member?
Looking at the definition in drivers/gpu/drm/xe/xe_pagefault_types.h:
struct xe_pagefault {
...
struct {
u64 page_addr;
union {
struct {
u8 alloc_state;
...
u32 id;
};
u64 end_addr;
};
struct xe_pagefault *next;
} consumer;
...
};
It appears consumer lacks a vma_info field. Will this code compile, or will
it result in a build error complaining about no member named vma_info?
[ ... ]
> @@ -137,6 +138,7 @@ xe_pagefault_set_start_addr(struct xe_pagefault *pf, u64 start_addr)
> {
> pf->consumer.page_addr &= XE_PAGEFAULT_ERROR_MASK;
> pf->consumer.page_addr |= (start_addr & ~XE_PAGEFAULT_ERROR_MASK);
> + pf->consumer.vma_info.start_addr = start_addr;
[Severity: Low]
Similarly, since vma_info doesn't exist in consumer, will this assignment also
cause a compiler error?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006213449.4027749-1-fei.yang@intel.com?part=1
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-06 21:36 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06 21:34 [PATCH 0/1] drm/xe: fixing endless PF range response fei.yang
2026-10-06 21:34 ` [PATCH 1/1] " fei.yang
2026-10-06 21:36 ` sashiko-bot
2026-10-06 21:36 ` ✗ CI.KUnit: failure for " Patchwork
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox