* drm/msm: unrecoverable a8xx hang-recovery loop on Glymur ("cx gdsc didn't collapse"); killing the offending client does not stop the loop
@ 2026-08-20 13:44 Greg Ociepka
2026-08-21 13:35 ` Rob Clark
0 siblings, 1 reply; 2+ messages in thread
From: Greg Ociepka @ 2026-08-20 13:44 UTC (permalink / raw)
To: linux-arm-msm, dri-devel, freedreno
Cc: Rob Clark, Sean Paul, Konrad Dybcio, Akhil P Oommen
Hi,
On an ASUS Zenbook A16 (UX3607OA, Snapdragon X2 Elite Extreme "Glymur", Adreno GPU, chip id 0x44070041) we are hitting an unrecoverable GPU recovery loop that, once triggered, can only be cleared by a full reboot.
Environment:
- kernel: linux-next next-20260730, plus a handful of local patches for this laptop (USB wakeup-source/dma-coherent DTS, CPU thermal trips, an eDP link-training retry tweak and a DP debug-print ratelimit); none of them touch the GPU, GMU, GDSC or GPU clock paths
- Mesa 26.1.5 (freedreno, a8xx)
- GMU firmware v5.2.38 (qcom/gen80100_gmu.bin) and gen80100_sqe.fw from linux-firmware
- niri (smithay-based Wayland compositor), ordinary desktop clients
Symptom: a regular client (rotating across sessions: chromium, firefox, Xwayland, thunderbird) eventually triggers a GPU fault. From that point recovery loops at roughly 1 Hz; every iteration logs:
adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* gpu fault ring 2 fence 28f9c status 00880005 gfx_status 00880004
adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* BR: status 00880004 rb 0002/0078 ib1 0000000000000000/0000 ib2 0000000000000000/0000 ib3 0000000000000000/0000
adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* BV: status 00880004 rb 0000/0078 ib1 0000000000000000/0000 ib2 0000000000000000/0000 ib3 0000000000000000/0000
msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 44070041: hangcheck recover!
msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 44070041: offending task: <client>
adreno 3d00000.gpu: [drm:a8xx_recover [msm]] *ERROR* cx gdsc didn't collapse
Observations across four separate incidents (2026-08-09 to 2026-08-20; longest 60+ minutes; worst measured rate 161 recover cycles in 5 minutes):
1. The reported offending client differs between incidents and can change within a single incident.
2. Killing the offending client does not stop the loop. We verified this explicitly: after SIGTERM-ing the reported victim the loop continued at an unchanged rate within seconds, and the next client that submitted work became the new "offending task". The corrupted state evidently survives in hardware/firmware rather than in the client.
3. The loop never converges on its own; only a full reboot restores the GPU. While it runs, the victim's surface stays black and the whole desktop degrades, since every recovery stalls the GPU.
4. Onset correlates with many clients creating GPU contexts in parallel, e.g. right after a session restore spawns half a dozen applications at once.
Since every iteration reports "cx gdsc didn't collapse", the recovery path apparently never succeeds in power-cycling the GPU, which would explain why the fault state persists across recovery attempts. Two questions this raises:
- Is the CX GDSC on Glymur shared with, or kept up by, other consumers (the display path, for instance) so that it cannot collapse while the panel is active - which would make this recovery flow structurally ineffective on this SoC?
- Or does gen8/GMU require a different stop/reset sequence before the collapse is attempted?
We are happy to test patches, run instrumented builds, and provide full journals or devcoredumps - the machine reproduces this every couple of days under normal desktop use.
Thanks,
Greg Ociepka
^ permalink raw reply [flat|nested] 2+ messages in thread* Re: drm/msm: unrecoverable a8xx hang-recovery loop on Glymur ("cx gdsc didn't collapse"); killing the offending client does not stop the loop
2026-08-20 13:44 drm/msm: unrecoverable a8xx hang-recovery loop on Glymur ("cx gdsc didn't collapse"); killing the offending client does not stop the loop Greg Ociepka
@ 2026-08-21 13:35 ` Rob Clark
0 siblings, 0 replies; 2+ messages in thread
From: Rob Clark @ 2026-08-21 13:35 UTC (permalink / raw)
To: Greg Ociepka
Cc: linux-arm-msm, dri-devel, freedreno, Sean Paul, Konrad Dybcio,
Akhil P Oommen
On Thu, Aug 20, 2026 at 6:44 AM Greg Ociepka <greg@ferrisoft.com> wrote:
>
> Hi,
>
> On an ASUS Zenbook A16 (UX3607OA, Snapdragon X2 Elite Extreme "Glymur", Adreno GPU, chip id 0x44070041) we are hitting an unrecoverable GPU recovery loop that, once triggered, can only be cleared by a full reboot.
>
> Environment:
> - kernel: linux-next next-20260730, plus a handful of local patches for this laptop (USB wakeup-source/dma-coherent DTS, CPU thermal trips, an eDP link-training retry tweak and a DP debug-print ratelimit); none of them touch the GPU, GMU, GDSC or GPU clock paths
> - Mesa 26.1.5 (freedreno, a8xx)
> - GMU firmware v5.2.38 (qcom/gen80100_gmu.bin) and gen80100_sqe.fw from linux-firmware
> - niri (smithay-based Wayland compositor), ordinary desktop clients
>
> Symptom: a regular client (rotating across sessions: chromium, firefox, Xwayland, thunderbird) eventually triggers a GPU fault. From that point recovery loops at roughly 1 Hz; every iteration logs:
>
> adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* gpu fault ring 2 fence 28f9c status 00880005 gfx_status 00880004
> adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* BR: status 00880004 rb 0002/0078 ib1 0000000000000000/0000 ib2 0000000000000000/0000 ib3 0000000000000000/0000
> adreno 3d00000.gpu: [drm:a8xx_irq [msm]] *ERROR* BV: status 00880004 rb 0000/0078 ib1 0000000000000000/0000 ib2 0000000000000000/0000 ib3 0000000000000000/0000
> msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 44070041: hangcheck recover!
> msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 44070041: offending task: <client>
> adreno 3d00000.gpu: [drm:a8xx_recover [msm]] *ERROR* cx gdsc didn't collapse
>
> Observations across four separate incidents (2026-08-09 to 2026-08-20; longest 60+ minutes; worst measured rate 161 recover cycles in 5 minutes):
>
> 1. The reported offending client differs between incidents and can change within a single incident.
> 2. Killing the offending client does not stop the loop. We verified this explicitly: after SIGTERM-ing the reported victim the loop continued at an unchanged rate within seconds, and the next client that submitted work became the new "offending task". The corrupted state evidently survives in hardware/firmware rather than in the client.
> 3. The loop never converges on its own; only a full reboot restores the GPU. While it runs, the victim's surface stays black and the whole desktop degrades, since every recovery stalls the GPU.
> 4. Onset correlates with many clients creating GPU contexts in parallel, e.g. right after a session restore spawns half a dozen applications at once.
>
> Since every iteration reports "cx gdsc didn't collapse", the recovery path apparently never succeeds in power-cycling the GPU, which would explain why the fault state persists across recovery attempts. Two questions this raises:
>
> - Is the CX GDSC on Glymur shared with, or kept up by, other consumers (the display path, for instance) so that it cannot collapse while the panel is active - which would make this recovery flow structurally ineffective on this SoC?
> - Or does gen8/GMU require a different stop/reset sequence before the collapse is attempted?
Hmm, I've not seen this on glymur slim7x, and I've crashed the GPU
plenty while burning down the remaining clcts/glcts/vkcts issues. And
certainly not any GPU hangs/crashes in normal desktop usage. (Vanilla
gnome-shell / fedora f44 here.)
System mesa is 26.1.6, but for testing/debug it has been newer ToT
mesa. I kinda suspect that the mesa version does not matter, this
sounds more like a power issue.
Fwiw, the kernel I've been using: [1].. but that doesn't have the a16
dts. I need to move forward to something newer, hopefully I'll get to
that today. We'll see if I start seeing this issue. Or maybe
something x2e vs x2ee is different regarding pmic/supplies?
My device doesn't have the highest OPP (1.85ghz), which I think your
device does. We might be still missing some power related features
that are (more?) needed at higher freq's, or maybe I have a more lucky
piece of silicon? Akhil would know more about this. You might try
removing the highest OPP in gpu_opp_table (in glymur.dtsi) to see if
that improves stability? (Repeat until it is stable if you have the
same issues w/ 1.7ghz.)
BR,
-R
[1] https://gitlab.freedesktop.org/robclark/msm/-/commits/glymur-next?ref_type=heads
>
> We are happy to test patches, run instrumented builds, and provide full journals or devcoredumps - the machine reproduces this every couple of days under normal desktop use.
>
> Thanks,
> Greg Ociepka
>
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-08-21 13:35 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-20 13:44 drm/msm: unrecoverable a8xx hang-recovery loop on Glymur ("cx gdsc didn't collapse"); killing the offending client does not stop the loop Greg Ociepka
2026-08-21 13:35 ` Rob Clark
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.