From: Michal Wajdeczko <michal.wajdeczko@intel.com>
To: Mallesh Koujalagi <mallesh.koujalagi@intel.com>,
<intel-xe@lists.freedesktop.org>, <rodrigo.vivi@intel.com>,
<matthew.brost@intel.com>
Cc: <anshuman.gupta@intel.com>, <badal.nilawar@intel.com>,
<vinay.belgaumkar@intel.com>, <riana.tauro@intel.com>,
<karthik.poosa@intel.com>, <sk.anirban@intel.com>,
<raag.jadav@intel.com>, <aravind.iddamsetty@linux.intel.com>,
Dnyaneshwar Bhadane <dnyaneshwar.bhadane@intel.com>
Subject: Re: [PATCH 2/4] drm/xe: Use xe_log SIGID API for probe-path error reporting
Date: Thu, 6 Aug 2026 19:52:32 +0200 [thread overview]
Message-ID: <4592d2a2-f241-4ff2-b922-997a3ef81cdb@intel.com> (raw)
In-Reply-To: <20260806110039.1245282-8-mallesh.koujalagi@intel.com>
On 8/6/2026 1:00 PM, Mallesh Koujalagi wrote:
> From: Dnyaneshwar Bhadane <dnyaneshwar.bhadane@intel.com>
>
> Replace ad-hoc drm_err()/xe_err() calls in the probe path with
> structured SIGID logging helpers. This gives fleet tooling a stable,
> machine-parseable identifier (SIGID) for each recognised fault
> situation rather than relying on fragile string matching.
>
> Signed-off-by: Dnyaneshwar Bhadane <dnyaneshwar.bhadane@intel.com>
> Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
> ---
> drivers/gpu/drm/xe/xe_device.c | 6 ++++--
> drivers/gpu/drm/xe/xe_hwmon.c | 4 +++-
> drivers/gpu/drm/xe/xe_irq.c | 9 +++++----
> drivers/gpu/drm/xe/xe_mmio.c | 5 +++--
> drivers/gpu/drm/xe/xe_pat.c | 6 ++++--
> drivers/gpu/drm/xe/xe_pci.c | 14 +++++++++-----
> drivers/gpu/drm/xe/xe_pcode.c | 5 +++--
> drivers/gpu/drm/xe/xe_vram.c | 5 +++--
> 8 files changed, 34 insertions(+), 20 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> index d25d02b24898..52bafe708d74 100644
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
> @@ -48,6 +48,7 @@
> #include "xe_i2c.h"
> #include "xe_irq.h"
> #include "xe_late_bind_fw.h"
> +#include "xe_log.h"
> #include "xe_mmio.h"
> #include "xe_module.h"
> #include "xe_nvm.h"
> @@ -586,7 +587,8 @@ int xe_device_init_early(struct xe_device *xe)
> * Cleanup done in xe_device_destroy via
> * drmm_add_action_or_reset register above
> */
> - drm_err(&xe->drm, "Failed to allocate xe workqueues\n");
> + xe_log_err_fatal(xe, PROBE, -ENOMEM, "Failed to allocate xe workqueues\n");
OOM errors will be already printed by the MM, we likely shouldn't have this drm_err here...
can't we just leave it as-is and add xe_log_err somewhere in xe_pci_probe()?
also, maybe we should wait with submitting more patches that use xe_log() until we actually get some agreement and merge that first?
> +
> return -ENOMEM;
> }
>
> @@ -717,7 +719,7 @@ static int xe_set_dma_info(struct xe_device *xe)
> return 0;
>
> mask_err:
> - drm_err(&xe->drm, "Can't set DMA mask/consistent mask (%d)\n", err);
> + xe_log_err(xe, PROBE, err, "Can't set DMA mask/consistent mask\n");
> return err;
> }
>
> diff --git a/drivers/gpu/drm/xe/xe_hwmon.c b/drivers/gpu/drm/xe/xe_hwmon.c
> index de3f2aeffc3f..e4c319598dd0 100644
> --- a/drivers/gpu/drm/xe/xe_hwmon.c
> +++ b/drivers/gpu/drm/xe/xe_hwmon.c
> @@ -15,6 +15,7 @@
> #include "regs/xe_pcode_regs.h"
> #include "xe_device.h"
> #include "xe_hwmon.h"
> +#include "xe_log.h"
> #include "xe_mmio.h"
> #include "xe_pcode.h"
> #include "xe_pcode_api.h"
> @@ -1575,7 +1576,8 @@ int xe_hwmon_register(struct xe_device *xe)
> &hwmon_chip_info,
> hwmon_groups);
> if (IS_ERR(hwmon->hwmon_dev)) {
> - drm_err(&xe->drm, "Failed to register xe hwmon (%pe)\n", hwmon->hwmon_dev);
> + xe_log_err(xe, PROBE, PTR_ERR(hwmon->hwmon_dev),
> + "Failed to register xe hwmon\n");
while this is during the probe, shouldn't we use HWMON as a component?
we allow chained SIGID reports from different layers, and this error
will eventually lead to final xe_log_err(PROBE) anyway
also, there are at least 2 other cases where hwmon initialization can fail
why only this one is so special that requires use of xe_log ?
and IMO it would be good to add each new SIGID usage (across
different files) in separate patches
> xe->hwmon = NULL;
> return PTR_ERR(hwmon->hwmon_dev);
> }
> diff --git a/drivers/gpu/drm/xe/xe_irq.c b/drivers/gpu/drm/xe/xe_irq.c
> index 9e49e2241da4..74d2b63cd505 100644
> --- a/drivers/gpu/drm/xe/xe_irq.c
> +++ b/drivers/gpu/drm/xe/xe_irq.c
> @@ -19,6 +19,7 @@
> #include "xe_hw_engine.h"
> #include "xe_hw_error.h"
> #include "xe_i2c.h"
> +#include "xe_log.h"
> #include "xe_memirq.h"
> #include "xe_mert.h"
> #include "xe_mmio.h"
> @@ -759,14 +760,14 @@ static int xe_irq_msi_request_irqs(struct xe_device *xe)
>
> irq_handler = xe_irq_handler(xe);
> if (!irq_handler) {
> - drm_err(&xe->drm, "No supported interrupt handler");
> + xe_log_err(xe, PROBE, -EINVAL, "No supported interrupt handler\n");
this is already a dead code, as xe_irq_handler() never returns NULL
we can use xe_assert instead
> return -EINVAL;
> }
>
> irq = pci_irq_vector(pdev, 0);
> err = request_irq(irq, irq_handler, IRQF_SHARED, DRIVER_NAME, xe);
> if (err < 0) {
> - drm_err(&xe->drm, "Failed to request MSI IRQ %d\n", err);
> + xe_log_err(xe, PROBE, err, "Failed to request MSI IRQ\n");
> return err;
> }
>
> @@ -822,7 +823,7 @@ int xe_irq_install(struct xe_device *xe)
>
> err = pci_alloc_irq_vectors(pdev, nvec, nvec, irq_flags);
> if (err < 0) {
> - drm_err(&xe->drm, "Failed to allocate IRQ vectors: %d\n", err);
> + xe_log_err(xe, PROBE, err, "Failed to allocate IRQ vectors\n");
> return err;
> }
>
> @@ -891,7 +892,7 @@ static int xe_irq_msix_init(struct xe_device *xe)
> return 0; /* MSI */
>
> if (nvec < 0) {
> - drm_err(&xe->drm, "Failed getting MSI-X vectors count: %d\n", nvec);
> + xe_log_err(xe, PROBE, nvec, "Failed getting MSI-X vectors count\n");
maybe we should have IRQ component with SIGID_SW and use it here?
> return nvec;
> }
>
> diff --git a/drivers/gpu/drm/xe/xe_mmio.c b/drivers/gpu/drm/xe/xe_mmio.c
> index 7fa18dfcb5a2..bcf36a381424 100644
> --- a/drivers/gpu/drm/xe/xe_mmio.c
> +++ b/drivers/gpu/drm/xe/xe_mmio.c
> @@ -15,6 +15,7 @@
> #include "regs/xe_bars.h"
> #include "xe_device.h"
> #include "xe_gt_sriov_vf.h"
> +#include "xe_log.h"
> #include "xe_printk.h"
> #include "xe_sriov.h"
> #include "xe_tile_printk.h"
> @@ -105,13 +106,13 @@ int xe_mmio_probe_early(struct xe_device *xe)
>
> xe->mmio.regs = pcim_iomap(pdev, GTTMMADR_BAR, 0);
> if (!xe->mmio.regs) {
> - xe_err(xe, "Failed to map GTTMMADR_BAR\n");
> + xe_log_err(xe, PROBE, -EIO, "Failed to map GTTMMADR_BAR\n");
shouldn't we use REGS component (SIGID_IO_BUS) here?
and this is a FATAL error, no?
> return -EIO;
> }
>
> xe->mmio.size = pci_resource_len(pdev, GTTMMADR_BAR);
> if (xe->mmio.size < SZ_16M) {
> - xe_err(xe, "GTTMMADR_BAR is too small: %zu\n", xe->mmio.size);
> + xe_log_err(xe, PROBE, -EIO, "GTTMMADR_BAR is too small: %zu\n", xe->mmio.size);
ditto
> return -EIO;
> }
>
> diff --git a/drivers/gpu/drm/xe/xe_pat.c b/drivers/gpu/drm/xe/xe_pat.c
> index a5fe1beec652..aad15a4b0ac3 100644
> --- a/drivers/gpu/drm/xe/xe_pat.c
> +++ b/drivers/gpu/drm/xe/xe_pat.c
> @@ -16,6 +16,7 @@
> #include "xe_force_wake.h"
> #include "xe_gt.h"
> #include "xe_gt_mcr.h"
> +#include "xe_log.h"
> #include "xe_mmio.h"
> #include "xe_sriov.h"
> #include "xe_wa.h"
> @@ -686,8 +687,9 @@ void xe_pat_init_early(struct xe_device *xe)
> * raise an error rather than trying to silently inherit the
> * most recent platform's behavior.
> */
> - drm_err(&xe->drm, "Missing PAT table for platform with graphics version %d.%02d!\n",
> - GRAPHICS_VER(xe), GRAPHICS_VERx100(xe) % 100);
> + xe_log_err(xe, PROBE, -ENODEV,
> + "Missing PAT table for platform with graphics version %d.%02d!\n",
> + GRAPHICS_VER(xe), GRAPHICS_VERx100(xe) % 100);
this here is pure SW gap and it is broken as right after
logging this error we might crash in the xe_assert() below
or later in xe_pat_init() as xe->pat.ops is NULL
maybe instead of blindly adding xe_log_err we can just
fix that first and return some error to abort the probe?
also maybe we should add PAT component (with SIGID_IO_BUS?)
> }
>
> xe_assert(xe, xe->pat.ops->dump);
> diff --git a/drivers/gpu/drm/xe/xe_pci.c b/drivers/gpu/drm/xe/xe_pci.c
> index 36d62927b969..3c452f399e6d 100644
> --- a/drivers/gpu/drm/xe/xe_pci.c
> +++ b/drivers/gpu/drm/xe/xe_pci.c
> @@ -25,6 +25,7 @@
> #include "xe_gt_printk.h"
> #include "xe_gt_sriov_vf.h"
> #include "xe_guc.h"
> +#include "xe_log.h"
> #include "xe_mmio.h"
> #include "xe_module.h"
> #include "xe_pci_error.h"
> @@ -718,8 +719,9 @@ static int handle_gmdid(struct xe_device *xe,
>
> *graphics_ip = find_graphics_ip(ver);
> if (!*graphics_ip) {
> - drm_err(&xe->drm, "Hardware reports unknown graphics version %u.%02u\n",
> - ver / 100, ver % 100);
> + xe_log_err(xe, PROBE, -ENODEV,
> + "Hardware reports unknown graphics version %u.%02u\n",
> + ver / 100, ver % 100);
maybe use REGS component here (SIGID_IO_BUS)
also, isn't that actually broken?
with graphics_ip = NULL set here we will crash later in xe_info_init()
> }
>
> ret = read_gmdid(xe, GMDID_MEDIA, &ver, media_revid);
> @@ -732,8 +734,9 @@ static int handle_gmdid(struct xe_device *xe,
>
> *media_ip = find_media_ip(ver);
> if (!*media_ip) {
> - drm_err(&xe->drm, "Hardware reports unknown media version %u.%02u\n",
> - ver / 100, ver % 100);
> + xe_log_err(xe, PROBE, -ENODEV,
> + "Hardware reports unknown media version %u.%02u\n",
> + ver / 100, ver % 100);
> }
>
> return 0;
> @@ -1088,7 +1091,8 @@ static int xe_info_init(struct xe_device *xe,
> * required for VRAM management).
> */
> if (!tile->primary_gt) {
> - drm_err(&xe->drm, "Cannot probe device with without a primary GT\n");
> + xe_log_err_fatal(xe, PROBE, -ENODEV,
> + "Cannot probe device without a primary GT\n");
> return -ENODEV;
> }
>
> diff --git a/drivers/gpu/drm/xe/xe_pcode.c b/drivers/gpu/drm/xe/xe_pcode.c
> index ccc3bdeed6bb..c492737242d3 100644
> --- a/drivers/gpu/drm/xe/xe_pcode.c
> +++ b/drivers/gpu/drm/xe/xe_pcode.c
> @@ -14,6 +14,7 @@
> #include "regs/xe_pmt.h"
> #include "xe_assert.h"
> #include "xe_device.h"
> +#include "xe_log.h"
> #include "xe_mmio.h"
> #include "xe_pcode_api.h"
> #include "xe_pm.h"
> @@ -320,8 +321,8 @@ int xe_pcode_ready(struct xe_device *xe, bool locked)
> mutex_unlock(&tile->pcode.lock);
>
> if (ret)
> - drm_err(&xe->drm,
> - "PCODE initialization timedout after: 3 min\n");
> + xe_log_err(xe, PCODE, -ret,
> + "PCODE initialization timedout after: 3 min\n");
drop "PCODE" prefix
and move to separate patch
>
> return ret;
> }
> diff --git a/drivers/gpu/drm/xe/xe_vram.c b/drivers/gpu/drm/xe/xe_vram.c
> index 23eb7edbdd57..dada430d7ed5 100644
> --- a/drivers/gpu/drm/xe/xe_vram.c
> +++ b/drivers/gpu/drm/xe/xe_vram.c
> @@ -17,6 +17,7 @@
> #include "xe_device.h"
> #include "xe_force_wake.h"
> #include "xe_gt_mcr.h"
> +#include "xe_log.h"
> #include "xe_mmio.h"
> #include "xe_sriov.h"
> #include "xe_tile_sriov_vf.h"
> @@ -43,7 +44,7 @@ static int determine_lmem_bar_size(struct xe_device *xe, struct xe_vram_region *
> struct pci_dev *pdev = to_pci_dev(xe->drm.dev);
>
> if (!resource_is_valid(pdev, LMEM_BAR)) {
> - drm_err(&xe->drm, "pci resource is not valid\n");
> + xe_log_err(xe, PROBE, -ENXIO, "pci resource is not valid\n");
shouldn't we use REGS here ? (SIGID_IO_BUS)
or add VRAM component (also SIGID_IO_BUS?)
> return -ENXIO;
> }
>
> @@ -237,7 +238,7 @@ static int vram_region_init(struct xe_device *xe, struct xe_vram_region *vram,
> vram->io_size = min_t(u64, usable_size, remain_io_size);
>
> if (!vram->io_size) {
> - drm_err(&xe->drm, "Tile without any CPU visible VRAM. Aborting.\n");
> + xe_log_err(xe, PROBE, -ENODEV, "Tile without any CPU visible VRAM. Aborting.\n");
we should first fix the vram_region_init() as now it is used
in two scenarios: to init each tile region, and to init device
level region info
and this error is meaningful for tile-based usage, where
we should use tile as a location
> return -ENODEV;
> }
>
next prev parent reply other threads:[~2026-08-06 17:52 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-06 11:00 [PATCH 0/4] drm/xe: Adopt xe_log SIGID API for structured error reporting Mallesh Koujalagi
2026-08-06 11:00 ` [PATCH 1/4] drm/xe/log: DO NOT REVIEW Mallesh Koujalagi
2026-08-06 11:00 ` [PATCH 2/4] drm/xe: Use xe_log SIGID API for probe-path error reporting Mallesh Koujalagi
2026-08-06 17:52 ` Michal Wajdeczko [this message]
2026-08-06 11:00 ` [PATCH 3/4] drm/xe/pcode: Improve pcode timeout logging Mallesh Koujalagi
2026-08-06 14:42 ` Michal Wajdeczko
2026-08-06 11:00 ` [PATCH 4/4] drm/xe/sysctrl: Add better sysctrl error reporting Mallesh Koujalagi
2026-08-06 14:41 ` Michal Wajdeczko
2026-08-07 8:48 ` Mallesh, Koujalagi
2026-08-06 19:14 ` Umesh Nerlige Ramappa
2026-08-07 8:59 ` Mallesh, Koujalagi
2026-08-06 11:07 ` ✗ CI.checkpatch: warning for drm/xe: Adopt xe_log SIGID API for structured " Patchwork
2026-08-06 11:08 ` ✓ CI.KUnit: success " Patchwork
2026-08-06 12:13 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-06 20:44 ` ✓ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4592d2a2-f241-4ff2-b922-997a3ef81cdb@intel.com \
--to=michal.wajdeczko@intel.com \
--cc=anshuman.gupta@intel.com \
--cc=aravind.iddamsetty@linux.intel.com \
--cc=badal.nilawar@intel.com \
--cc=dnyaneshwar.bhadane@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=karthik.poosa@intel.com \
--cc=mallesh.koujalagi@intel.com \
--cc=matthew.brost@intel.com \
--cc=raag.jadav@intel.com \
--cc=riana.tauro@intel.com \
--cc=rodrigo.vivi@intel.com \
--cc=sk.anirban@intel.com \
--cc=vinay.belgaumkar@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox