All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Mallesh, Koujalagi" <mallesh.koujalagi@intel.com>
To: "Tauro, Riana" <riana.tauro@intel.com>,
	Michal Wajdeczko <michal.wajdeczko@intel.com>,
	<intel-xe@lists.freedesktop.org>, <rodrigo.vivi@intel.com>,
	<matthew.brost@intel.com>, <aravind.iddamsetty@linux.intel.com>
Cc: <anshuman.gupta@intel.com>, <badal.nilawar@intel.com>,
	<vinay.belgaumkar@intel.com>, <karthik.poosa@intel.com>,
	<sk.anirban@intel.com>, <raag.jadav@intel.com>,
	<umesh.nerlige.ramappa@intel.com>,
	<dnyaneshwar.bhadane@intel.com>, <anoop.c.vijay@intel.com>
Subject: Re: [PATCH v4 1/7] drm/xe/sysctrl: Return error codes from sysctrl_wait_bit_clear()
Date: Fri, 4 Sep 2026 13:55:39 +0530	[thread overview]
Message-ID: <87216310-a723-43ce-ae62-6b85d2525bbd@intel.com> (raw)
In-Reply-To: <292aaec1-662d-466a-83a6-dc820dace637@intel.com>


On 25-08-2026 01:05 pm, Tauro, Riana wrote:
>
> On 24-08-2026 16:34, Michal Wajdeczko wrote:
>>
>> On 8/24/2026 7:44 AM, Tauro, Riana wrote:
>>> Hi Mallesh/Michal
>>>
>>> On 20-08-2026 15:46, Mallesh Koujalagi wrote:
>>>> Make sysctrl_wait_bit_clear() return an error code rather than a bool.
>>>> and update callers to use xe_log_err() with the propagated error code.
>>> Bit confused here on the usage of SIGID. My assumption based on 
>>> documentation and discussion was
>>> that SIG ID is only needed for error logs based on severity and 
>>> requires a resolution which can be documented.
>>> Am i missing something?
>> from "How to pick a SIGID (the uniqueness rule)"
>>
>>   * A single underlying failure therefore legitimately produces a
>>   * *chain* of reports from different layers, each with its own SIGID 
>> -- e.g. a
>>   * GuC communication failure is reported as %XE_SIGID_RUNTIME_FW by 
>> the firmware
>>   * path, the failed recovery as %XE_SIGID_GT_TDR by the reset path, 
>> and an
>>   * aborted bind as %XE_SIGID_PROBE by the probe path. That chain 
>> lets triage
>>   * follow a fault from origin to final effect; it is not a duplicate.
>>
>> so in this case, some sysctrl timeout may eventually lead to a reset 
>> and/or
>> wedge and/or runtime-survivability-mode, and this intermediate FW 
>> SIGID may
>> help in any postmortem triage, and if we don't reach that final state 
>> then
>> we can still provide some generic resolution based on site and errno 
>> (COLLECT)
>
> What if the paths are also used in non-critical scenarios for ex: 
> sysctrl commands are used for
> general counter telemetry, gpu health querying and firmware status.
> Wouldn't this be lead to a lot of sigids in dmesg in case of some 
> unrelated failure ..
> Why not only add sigid for critical paths when we are sure it will 
> cause a chain of failure
> instead of these prints on every trivial command failure.

Ops I miss this.

I agree that we should not add SIGID for every minor or noisy failure. 
The intent is not to log every non-critical cases,

but to capture important intermediate failures that can help explain the 
real root cause. If we log only the final critical

failure, we may miss the earlier cases that actually caused the issue. 
That makes triage and postmortem harder, since

we lose visibility into how the failure occurred. The idea is to use 
SIGID only for meaningful failure in key execution paths.

Even when these failure are not fatal, they can provide valuable 
context, make debugging easier and reduce

mean-time-to triage (MTTT).


Thanks,

-/Mallesh

>
> Thanks
> Riana
>
>>
>>> How do these errors indicate the need for a resolution.
>>> Do we need it for all kmd logs with components or for errors that 
>>> need a recovery or can be recoverable?
>>>
>>> Thanks
>>> Riana
>>>
>>>> Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
>>>> ---
>>>>    drivers/gpu/drm/xe/xe_sysctrl_mailbox.c | 26 
>>>> ++++++++++++-------------
>>>>    1 file changed, 13 insertions(+), 13 deletions(-)
>>>>
>>>> diff --git a/drivers/gpu/drm/xe/xe_sysctrl_mailbox.c 
>>>> b/drivers/gpu/drm/xe/xe_sysctrl_mailbox.c
>>>> index e13eebaac1d0..ef847f0a8f2c 100644
>>>> --- a/drivers/gpu/drm/xe/xe_sysctrl_mailbox.c
>>>> +++ b/drivers/gpu/drm/xe/xe_sysctrl_mailbox.c
>>>> @@ -11,6 +11,7 @@
>>>>      #include "regs/xe_sysctrl_regs.h"
>>>>    #include "xe_device.h"
>>>> +#include "xe_log.h"
>>>>    #include "xe_mmio.h"
>>>>    #include "xe_pm.h"
>>>>    #include "xe_printk.h"
>>>> @@ -34,15 +35,11 @@ struct xe_sysctrl_mailbox_msg_hdr {
>>>>    #define XE_SYSCTRL_HDR_RESULT(hdr) \
>>>>        FIELD_GET(SYSCTRL_HDR_RESULT_MASK, le32_to_cpu((hdr)->data))
>>>>    -static bool sysctrl_wait_bit_clear(struct xe_sysctrl *sc, u32 
>>>> bit_mask,
>>>> -                   unsigned int timeout_ms)
>>>> +static int sysctrl_wait_bit_clear(struct xe_sysctrl *sc, u32 
>>>> bit_mask,
>>>> +                  unsigned int timeout_ms)
>>>>    {
>>>> -    int ret;
>>>> -
>>>> -    ret = xe_mmio_wait32_not(sc->mmio, SYSCTRL_MB_CTRL, bit_mask, 
>>>> bit_mask,
>>>> +    return xe_mmio_wait32_not(sc->mmio, SYSCTRL_MB_CTRL, bit_mask, 
>>>> bit_mask,
>>>>                     timeout_ms * 1000, NULL, false);
>>>> -
>>>> -    return ret == 0;
>>>>    }
>>>>      static bool sysctrl_wait_bit_set(struct xe_sysctrl *sc, u32 
>>>> bit_mask,
>>>> @@ -145,12 +142,14 @@ static int sysctrl_send_frames(struct 
>>>> xe_sysctrl *sc,
>>>>        struct xe_device *xe = sc_to_xe(sc);
>>>>        u32 ctrl_reg, total_frames, frame;
>>>>        size_t bytes_sent, frame_size;
>>>> +    int ret;
>>>>          total_frames = DIV_ROUND_UP(cmd_size, 
>>>> XE_SYSCTRL_MB_FRAME_SIZE);
>>>>    -    if (!sysctrl_wait_bit_clear(sc, SYSCTRL_MB_CTRL_RUN_BUSY, 
>>>> timeout_ms)) {
>>>> -        xe_err(xe, "sysctrl: Mailbox busy\n");
>>>> -        return -EBUSY;
>>>> +    ret = sysctrl_wait_bit_clear(sc, SYSCTRL_MB_CTRL_RUN_BUSY, 
>>>> timeout_ms);
>>>> +    if (ret) {
>>>> +        xe_log_err(xe, SYSCTRL, ret, "Mailbox busy\n");
>>>> +        return ret;
>>>>        }
>>>>          sc->phase_bit ^= 1;
>>>> @@ -173,10 +172,11 @@ static int sysctrl_send_frames(struct 
>>>> xe_sysctrl *sc,
>>>>              xe_mmio_write32(sc->mmio, SYSCTRL_MB_CTRL, ctrl_reg);
>>>>    -        if (!sysctrl_wait_bit_clear(sc, 
>>>> SYSCTRL_MB_CTRL_RUN_BUSY, timeout_ms)) {
>>>> -            xe_err(xe, "sysctrl: Frame %u acknowledgment 
>>>> timeout\n", frame);
>>>> +        ret = sysctrl_wait_bit_clear(sc, SYSCTRL_MB_CTRL_RUN_BUSY, 
>>>> timeout_ms);
>>>> +        if (ret) {
>>>> +            xe_log_err(xe, SYSCTRL, ret, "Frame %u acknowledgment 
>>>> timeout\n", frame);
>>>>                sc->phase_bit = 0;
>>>> -            return -ETIMEDOUT;
>>>> +            return ret;
>>>>            }
>>>>              bytes_sent += frame_size;

  reply	other threads:[~2026-09-04  8:26 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-20 10:16 [PATCH v4 0/7] drm/xe/sysctrl: Clean up error handling in Mallesh Koujalagi
2026-08-20 10:16 ` [PATCH v4 1/7] drm/xe/sysctrl: Return error codes from sysctrl_wait_bit_clear() Mallesh Koujalagi
2026-08-20 22:23   ` Michal Wajdeczko
2026-08-24  5:44   ` Tauro, Riana
2026-08-24 11:04     ` Michal Wajdeczko
2026-08-25  7:35       ` Tauro, Riana
2026-09-04  8:25         ` Mallesh, Koujalagi [this message]
2026-09-04  9:06           ` Tauro, Riana
2026-09-04 10:18             ` Mallesh, Koujalagi
2026-09-04 11:11               ` Tauro, Riana
2026-09-04 11:57                 ` Michal Wajdeczko
2026-08-20 10:16 ` [PATCH v4 2/7] drm/xe/sysctrl: Return error codes from sysctrl_wait_bit_set() Mallesh Koujalagi
2026-08-20 22:26   ` Michal Wajdeczko
2026-08-20 10:16 ` [PATCH v4 3/7] drm/xe/sysctrl: Make sysctrl_write_frame() void Mallesh Koujalagi
2026-08-20 22:28   ` Michal Wajdeczko
2026-08-20 10:16 ` [PATCH v4 4/7] drm/xe/sysctrl: Use xe_assert() for payload size validation Mallesh Koujalagi
2026-08-20 22:30   ` Michal Wajdeczko
2026-08-20 10:16 ` [PATCH v4 5/7] drm/xe/sysctrl: Improve firmware response error logging Mallesh Koujalagi
2026-08-20 10:28   ` sashiko-bot
2026-08-20 22:40   ` Michal Wajdeczko
2026-08-20 10:16 ` [PATCH v4 6/7] drm/xe/sysctrl: Log group and command ID on mailbox failure Mallesh Koujalagi
2026-08-20 22:44   ` Michal Wajdeczko
2026-08-20 10:16 ` [PATCH v4 7/7] drm/xe/sysctrl: Report 'System Controller event' error using SIGID Mallesh Koujalagi
2026-08-20 22:58   ` Michal Wajdeczko
2026-08-20 10:25 ` ✓ CI.KUnit: success for drm/xe/sysctrl: Clean up error handling in Patchwork
2026-08-20 11:01 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-20 12:14 ` ✓ Xe.CI.FULL: " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=87216310-a723-43ce-ae62-6b85d2525bbd@intel.com \
    --to=mallesh.koujalagi@intel.com \
    --cc=anoop.c.vijay@intel.com \
    --cc=anshuman.gupta@intel.com \
    --cc=aravind.iddamsetty@linux.intel.com \
    --cc=badal.nilawar@intel.com \
    --cc=dnyaneshwar.bhadane@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=karthik.poosa@intel.com \
    --cc=matthew.brost@intel.com \
    --cc=michal.wajdeczko@intel.com \
    --cc=raag.jadav@intel.com \
    --cc=riana.tauro@intel.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=sk.anirban@intel.com \
    --cc=umesh.nerlige.ramappa@intel.com \
    --cc=vinay.belgaumkar@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.