From: Akhil P Oommen <akhilpo@oss.qualcomm.com>
To: rob.clark@oss.qualcomm.com
Cc: "Sean Paul" <sean@poorly.run>,
"Konrad Dybcio" <konradybcio@kernel.org>,
"Dmitry Baryshkov" <lumag@kernel.org>,
"Abhinav Kumar" <abhinav.kumar@linux.dev>,
"Jessica Zhang" <jesszhan0024@gmail.com>,
"Marijn Suijten" <marijn.suijten@somainline.org>,
"David Airlie" <airlied@gmail.com>,
"Simona Vetter" <simona@ffwll.ch>,
"Puranam V G Tejaswi" <quic_pvgtejas@quicinc.com>,
"Jie Zhang" <quic_jiezh@quicinc.com>,
"Maíra Canal" <mcanal@igalia.com>,
linux-arm-msm@vger.kernel.org, dri-devel@lists.freedesktop.org,
freedreno@lists.freedesktop.org, linux-kernel@vger.kernel.org,
"Jie Zhang" <jie.zhang@oss.qualcomm.com>
Subject: Re: [PATCH 5/6] drm/msm/a6xx: Fix IRQ storm during msm_recovery test
Date: Tue, 9 Jun 2026 03:24:59 +0530 [thread overview]
Message-ID: <49b8530f-24d3-4201-b22c-0f8eaea9f4e0@oss.qualcomm.com> (raw)
In-Reply-To: <CACSVV01dbQcjE+nTic+9R4VfCtNGvpwODH8BMZi8B7LFtcCCfQ@mail.gmail.com>
On 6/5/2026 12:20 PM, Rob Clark wrote:
> On Thu, Jun 4, 2026 at 1:10 PM Akhil P Oommen <akhilpo@oss.qualcomm.com> wrote:
>>
>> From: Jie Zhang <jie.zhang@oss.qualcomm.com>
>>
>> Once a hang is triggered by the msm_recovery test, the gpu error irq
>> remains asserted and triggers an interrupt storm. In the worst case,
>> this IRQ storm lands on the CPU core where the hangcheck timer is
>> scheduled, blocking it from running. This eventually leads to CPU
>> watchdog timeouts.
>>
>> To fix this, mask the gpu error irqs during msm_recovery test and
>> enable them back during the recovery.
>>
>> Fixes: 5edf2750d998 ("drm/msm: Add debugfs to disable hw err handling")
>> Signed-off-by: Jie Zhang <jie.zhang@oss.qualcomm.com>
>> Signed-off-by: Akhil P Oommen <akhilpo@oss.qualcomm.com>
>> ---
>> drivers/gpu/drm/msm/adreno/a5xx_gpu.c | 5 +++++
>> drivers/gpu/drm/msm/adreno/a6xx_gpu.c | 5 ++++-
>> drivers/gpu/drm/msm/adreno/a8xx_gpu.c | 5 ++++-
>> drivers/gpu/drm/msm/msm_gpu.c | 2 ++
>> 4 files changed, 15 insertions(+), 2 deletions(-)
>>
>> diff --git a/drivers/gpu/drm/msm/adreno/a5xx_gpu.c b/drivers/gpu/drm/msm/adreno/a5xx_gpu.c
>> index 2c0bbac43c52..f1df2514c613 100644
>> --- a/drivers/gpu/drm/msm/adreno/a5xx_gpu.c
>> +++ b/drivers/gpu/drm/msm/adreno/a5xx_gpu.c
>> @@ -1275,6 +1275,11 @@ static irqreturn_t a5xx_irq(struct msm_gpu *gpu)
>> status & ~A5XX_RBBM_INT_0_MASK_RBBM_AHB_ERROR);
>>
>> if (priv->disable_err_irq) {
>> + /* Turn off interrupts to avoid interrupt storm */
>> + gpu_write(gpu, REG_A5XX_RBBM_INT_0_MASK,
>> + A5XX_RBBM_INT_0_MASK_CP_CACHE_FLUSH_TS |
>> + A5XX_RBBM_INT_0_MASK_CP_SW);
>> +
>> status &= A5XX_RBBM_INT_0_MASK_CP_CACHE_FLUSH_TS |
>> A5XX_RBBM_INT_0_MASK_CP_SW;
>> }
>> diff --git a/drivers/gpu/drm/msm/adreno/a6xx_gpu.c b/drivers/gpu/drm/msm/adreno/a6xx_gpu.c
>> index 8b3bb2fd433b..9a4f9d0e1780 100644
>> --- a/drivers/gpu/drm/msm/adreno/a6xx_gpu.c
>> +++ b/drivers/gpu/drm/msm/adreno/a6xx_gpu.c
>> @@ -1911,8 +1911,11 @@ static irqreturn_t a6xx_irq(struct msm_gpu *gpu)
>>
>> gpu_write(gpu, REG_A6XX_RBBM_INT_CLEAR_CMD, status);
>>
>> - if (priv->disable_err_irq)
>> + if (priv->disable_err_irq) {
>> + /* Turn off interrupts to avoid interrupt storm */
>> + gpu_write(gpu, REG_A6XX_RBBM_INT_0_MASK, A6XX_RBBM_INT_0_MASK_CP_CACHE_FLUSH_TS);
>> status &= A6XX_RBBM_INT_0_MASK_CP_CACHE_FLUSH_TS;
>> + }
>>
>> if (status & A6XX_RBBM_INT_0_MASK_RBBM_HANG_DETECT)
>> a6xx_fault_detect_irq(gpu);
>> diff --git a/drivers/gpu/drm/msm/adreno/a8xx_gpu.c b/drivers/gpu/drm/msm/adreno/a8xx_gpu.c
>> index 9e44fd1ae634..0f6fd35bd587 100644
>> --- a/drivers/gpu/drm/msm/adreno/a8xx_gpu.c
>> +++ b/drivers/gpu/drm/msm/adreno/a8xx_gpu.c
>> @@ -1211,8 +1211,11 @@ irqreturn_t a8xx_irq(struct msm_gpu *gpu)
>>
>> gpu_write(gpu, REG_A8XX_RBBM_INT_CLEAR_CMD, status);
>>
>> - if (priv->disable_err_irq)
>> + if (priv->disable_err_irq) {
>> + /* Turn off interrupts to avoid interrupt storm */
>> + gpu_write(gpu, REG_A8XX_RBBM_INT_0_MASK, A6XX_RBBM_INT_0_MASK_CP_CACHE_FLUSH_TS);
>> status &= A6XX_RBBM_INT_0_MASK_CP_CACHE_FLUSH_TS;
>> + }
>>
>> if (status & A6XX_RBBM_INT_0_MASK_RBBM_HANG_DETECT)
>> a8xx_fault_detect_irq(gpu);
>> diff --git a/drivers/gpu/drm/msm/msm_gpu.c b/drivers/gpu/drm/msm/msm_gpu.c
>> index 9ac7740a87f0..48ac51f4119b 100644
>> --- a/drivers/gpu/drm/msm/msm_gpu.c
>> +++ b/drivers/gpu/drm/msm/msm_gpu.c
>> @@ -552,6 +552,8 @@ static void recover_worker(struct kthread_work *work)
>> msm_update_fence(ring->fctx, fence);
>> }
>>
>> + priv->disable_err_irq = false;
>
> Ok, so we rely on recovery to re-enable the error irqs.. that is
> probably ok, given the intended purpose of the debugfs file. And,
> well, it is debugfs. But why do we clear disable_err_irq here?
Now that we are updating the IRQ mask register which won't reset until
there is a gpu suspend, its side effect will be felt even after
userspace deasserts the debugfs knob, potentially into the next
testcase. This is different from the older behavior. So, I felt it would
be better to reset this flag during the recovery, considering
msm_recovery is the only user of this knob, afaiu.
I should have explicitly called out this new behavior of disable_err_irq
in the commit text, but I forgot.
-Akhil.
>
> BR,
> -R
>
>> +
>> gpu->funcs->recover(gpu);
>>
>> /* retire completed submits, plus the one that hung: */
>>
>> --
>> 2.51.0
>>
next prev parent reply other threads:[~2026-06-08 21:55 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-04 20:08 [PATCH 0/6] drm/msm: Assorted fixes - June/26 Akhil P Oommen
2026-06-04 20:08 ` [PATCH 1/6] drm/msm/a6xx: Fix stale rpmh votes after suspend Akhil P Oommen
2026-06-05 13:09 ` Neil Armstrong
2026-06-06 12:04 ` Dmitry Baryshkov
2026-06-08 8:10 ` Konrad Dybcio
2026-06-04 20:08 ` [PATCH 2/6] drm/msm: Recover HW before retire hung submit Akhil P Oommen
2026-06-08 8:11 ` Konrad Dybcio
2026-06-04 20:08 ` [PATCH 3/6] drm/msm/a6xx: Fix A663 GPUCC register list for state capture Akhil P Oommen
2026-06-06 12:04 ` Dmitry Baryshkov
2026-06-04 20:08 ` [PATCH 4/6] drm/msm/a6xx: Fix A621 " Akhil P Oommen
2026-06-06 12:05 ` Dmitry Baryshkov
2026-06-04 20:08 ` [PATCH 5/6] drm/msm/a6xx: Fix IRQ storm during msm_recovery test Akhil P Oommen
2026-06-05 6:50 ` Rob Clark
2026-06-08 21:54 ` Akhil P Oommen [this message]
2026-06-09 0:20 ` Rob Clark
2026-06-09 13:08 ` Akhil P Oommen
2026-06-09 14:25 ` Rob Clark
2026-06-09 20:03 ` Akhil P Oommen
2026-06-04 20:08 ` [PATCH 6/6] drm/msm: Fix task_struct reference leak in recover_worker Akhil P Oommen
2026-06-08 8:16 ` Konrad Dybcio
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=49b8530f-24d3-4201-b22c-0f8eaea9f4e0@oss.qualcomm.com \
--to=akhilpo@oss.qualcomm.com \
--cc=abhinav.kumar@linux.dev \
--cc=airlied@gmail.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=freedreno@lists.freedesktop.org \
--cc=jesszhan0024@gmail.com \
--cc=jie.zhang@oss.qualcomm.com \
--cc=konradybcio@kernel.org \
--cc=linux-arm-msm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lumag@kernel.org \
--cc=marijn.suijten@somainline.org \
--cc=mcanal@igalia.com \
--cc=quic_jiezh@quicinc.com \
--cc=quic_pvgtejas@quicinc.com \
--cc=rob.clark@oss.qualcomm.com \
--cc=sean@poorly.run \
--cc=simona@ffwll.ch \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox