Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Tales A. Mendonça" <talesam@gmail.com>
Cc: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v4 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout
Date: Thu, 17 Sep 2026 16:50:42 +0000	[thread overview]
Message-ID: <20260917165042.7466A1F00893@smtp.kernel.org> (raw)
In-Reply-To: <20260917163553.1742580-2-talesam@gmail.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] drm/xe: data race in xe_tlb_inval_fence_timeout() reading seqno_recv
--

commit ef794f6bbaafa679257b9fec9405245af4942e68
Author: Tales A. Mendonça <talesam@gmail.com>

drm/xe: Capture devcoredump on TLB invalidation timeout

This commit adds xe_devcoredump_gt(), a variant of xe_devcoredump() for hangs
that are not tied to an exec queue or job. It hooks this new capture mechanism
into the TLB invalidation timeout path to capture the GuC log and CT state.

Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678

> diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
> index 7c3bf32d9258b..7a0c04fac60aa 100644
> --- a/drivers/gpu/drm/xe/xe_tlb_inval.c
> +++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
[ ... ]
> @@ -97,6 +107,16 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
>  		queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->fence_tdr,
>  				   timeout_delay);
>  	spin_unlock_irq(&tlb_inval->pending_lock);
> +
> +	/*
> +	 * Capture the GuC log and CT state so the firmware side of the hang
> +	 * can be inspected; there is no queue or job to blame here. Must be
> +	 * outside pending_lock as the capture takes sleeping locks.
> +	 */
> +	if (timedout_seqno)
> +		xe_devcoredump_gt(tlb_inval_to_gt(tlb_inval),
> +				  "TLB invalidation fence timeout, seqno=%d recv=%d",
> +				  timedout_seqno, tlb_inval->seqno_recv);

[Severity: Low]
Does this code introduce a data race when reading tlb_inval->seqno_recv?

In xe_tlb_inval_fence_timeout() in drivers/gpu/drm/xe/xe_tlb_inval.c, the
pending_lock spinlock is dropped just before this block. Because seqno_recv
is read locklessly here, and concurrent updates (such as those in
xe_tlb_inval_done_handler()) use WRITE_ONCE(), reading it directly could
result in a data race.

Could this lead to KCSAN warnings or logging a corrupted value in the
devcoredump reason string? Would it be safer to read this using READ_ONCE()
or cache the value while the lock is held?

>  }
>  
>  /**

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260917163553.1742580-1-talesam@gmail.com?part=1

  reply	other threads:[~2026-09-17 16:50 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17 16:35 [PATCH v4 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-17 16:35 ` [PATCH v4 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-17 16:50   ` sashiko-bot [this message]
2026-09-17 23:18     ` Tales A. Mendonça
2026-09-18 22:01   ` Matthew Brost
2026-09-21 18:06     ` Tales A. Mendonça
2026-09-17 16:35 ` [PATCH v4 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-09-18 22:06   ` Matthew Brost
2026-09-21  5:30     ` Matthew Brost
2026-09-21 18:10       ` Tales A. Mendonça
2026-09-21 18:19         ` Matthew Brost
2026-09-17 16:35 ` [PATCH v4 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-17 16:42 ` ✗ CI.checkpatch: warning for drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) (rev3) Patchwork
2026-09-17 16:44 ` ✓ CI.KUnit: success " Patchwork
2026-09-17 17:54 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-17 18:45 ` ✓ Xe.CI.FULL: " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260917165042.7466A1F00893@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=talesam@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox