Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Mika Kuoppala" <mika.kuoppala@linux.intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v10 04/27] drm/xe/eudebug: Introduce discovery for resources
Date: Thu, 03 Sep 2026 15:22:23 +0000	[thread overview]
Message-ID: <20260903152224.AD48C1F00A3F@smtp.kernel.org> (raw)
In-Reply-To: <20260903145952.848051-5-mika.kuoppala@linux.intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] Global ordered workqueue stall via copy_from_user block while holding ioctl_lock.
- [Medium] Event sequence numbers (seqno) can be queued out-of-order in the event stream due to a race between seqno generation and event queuing.
--

commit e2061e72d8e954a17bde35e9c970ee41bc2db0ad
Author: Mika Kuoppala <mika.kuoppala@linux.intel.com>

    drm/xe/eudebug: Introduce discovery for resources

This commit introduces a resource discovery mechanism for debugger connections.
When a debugger connects, all existing resources for the target client are
relayed as events via an ordered workqueue, and DRM ioctls are blocked using
an rw_semaphore until discovery is complete.

> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> index 032a7ec5fbc8d..f10959b898297 100644
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
[ ... ]
> @@ -238,8 +239,12 @@ static long xe_drm_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
>  
>  	ACQUIRE(xe_pm_runtime_ioctl, pm)(xe);
>  	ret = ACQUIRE_ERR(xe_pm_runtime_ioctl, &pm);
> -	if (ret >= 0)
> +	if (ret >= 0) {
> +		bool lock = xe_eudebug_discovery_lock(file, cmd);
>  		ret = drm_ioctl(file, cmd, arg);

[Severity: High]
Does holding the discovery lock across drm_ioctl() create a vector for a
workqueue stall? 

If an unprivileged user maps the ioctl argument pointer to a blocking FUSE
or userfaultfd region, drm_ioctl() can be stalled indefinitely during
copy_from_user() while holding xef->eudebug.ioctl_lock for read.

> +		if (lock)
> +			xe_eudebug_discovery_unlock(file, cmd);
> +	}
>  
>  	return ret;
>  }
[ ... ]
> diff --git a/drivers/gpu/drm/xe/xe_eudebug.c b/drivers/gpu/drm/xe/xe_eudebug.c
> index 835dbd7309af5..7a5b14346e059 100644
> --- a/drivers/gpu/drm/xe/xe_eudebug.c
> +++ b/drivers/gpu/drm/xe/xe_eudebug.c
[ ... ]
> @@ -659,6 +705,87 @@ void xe_eudebug_vm_destroy(struct xe_file *xef, struct xe_vm *vm)
>  	xe_eudebug_put(d);
>  }
>  
> +static struct xe_file *xe_eudebug_target_get(struct xe_eudebug *d)
> +{
[ ... ]
> +static void discover_client(struct xe_eudebug *d)
> +{
> +	struct xe_file *xef;
> +	struct xe_vm *vm;
> +	unsigned long i;
> +	unsigned int vm_count = 0;
> +	int err = 0;
> +
> +	xef = xe_eudebug_target_get(d);
> +	if (!xef) {
> +		complete_all(&d->discovery);
> +		return;
> +	}
> +
> +	/*
> +	 * An xe_eudebug ref is taken for the discovery worker, which holds a
> +	 * ref to the target xe_file; the xe_file in turn holds the vm and
> +	 * exec_queue refs.
> +	 *
> +	 * All the relevant ioctls that manipulate the xe_file resources eudebug
> +	 * cares about are gated behind down_read(&xef->eudebug.ioctl_lock).
> +	 * Thus taking the write lock excludes all those ioctls, which lets us
> +	 * peek into the resource lists without taking their individual locks.
> +	 */
> +	down_write(&xef->eudebug.ioctl_lock);

[Severity: High]
If the target process is blocked in copy_from_user() inside drm_ioctl()
while holding the read lock, will this down_write() block indefinitely?

Because discovery runs on a global ordered workqueue (configured below in
xe_eudebug_init), this single blocked work item could stall all eudebug
discovery operations system-wide, causing a Denial of Service for all
debugger connections.

> +
> +	if (xe_eudebug_detached(d))
> +		goto out;
[ ... ]
> @@ -1126,12 +1266,20 @@ void xe_eudebug_init_early(struct xe_device *xe)
>  void xe_eudebug_init(struct xe_device *xe)
>  {
>  	struct drm_device *dev = &xe->drm;
> +	struct workqueue_struct *wq;
>  	int err;
>  
>  	/* early init failed */
>  	if (xe->eudebug.cap_state == XE_EUDEBUG_CAP_NOT_SUPPORTED)
>  		return;
>  
> +	wq = drmm_alloc_ordered_workqueue(dev, "xe-eudebug-discovery", 0);

[Severity: High]
This allocates the workqueue as ordered (max_active=1), which means a single
stalled work item blocks all other discovery work.

Can the workqueue be allocated differently, or can the ioctl_lock be managed
to avoid wrapping the copy_from_user() operations in drm_ioctl()?

> +	if (IS_ERR(wq)) {
> +		err = PTR_ERR(wq);
> +		goto out_err;
> +	}
> +	xe->eudebug.discovery_wq = wq;
> +
>  	err = sysfs_create_file(&dev->dev->kobj,
>  				&dev_attr_enable_eudebug.attr);
>  	if (err)
[ ... ]

[Severity: Medium]
Can sequence numbers be queued out-of-order in the event stream?

Looking at send_vm_bind_event() in xe_eudebug.c:

	*seqno = atomic_long_inc_return(&d->events.seqno);

	spin_lock(&d->events.lock);
	event = xe_eudebug_prepare_event(d, DRM_XE_EUDEBUG_EVENT_VM_BIND,
					 *seqno, base_flags, sz);

Since the sequence number is incremented atomically outside of the FIFO
spinlock, could two threads acquire sequence numbers N and N+1, but acquire
the FIFO lock in reverse order?

This same pattern appears in _xe_eudebug_add_handle(). Pushing N+1 before
N would break the expected monotonicity of the event stream and could cause
the debugger to drop events.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260903145952.848051-1-mika.kuoppala@linux.intel.com?part=4

  reply	other threads:[~2026-09-03 15:22 UTC|newest]

Thread overview: 48+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 14:59 [PATCH v10 00/27] Intel Xe GPU Debug Support (eudebug) v10 Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 01/27] drm/xe/eudebug: Introduce eudebug interface Mika Kuoppala
2026-09-03 15:16   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 02/27] drm/xe/eudebug: Add documentation Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 03/27] drm/xe/eudebug: Add connection establishment documentation Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 04/27] drm/xe/eudebug: Introduce discovery for resources Mika Kuoppala
2026-09-03 15:22   ` sashiko-bot [this message]
2026-09-03 14:59 ` [PATCH v10 05/27] drm/xe: Add EUDEBUG_ENABLE exec queue property Mika Kuoppala
2026-09-03 15:14   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 06/27] drm/xe/eudebug: Introduce exec_queue events Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 07/27] drm/xe/eudebug: Mark guc contexts as debuggable Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 08/27] drm/xe: Remove ifdef in DRM_GPUVA_OP_DRIVER svm subop checking Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 09/27] drm/xe: Introduce ADD_DEBUG_DATA and REMOVE_DEBUG_DATA vm bind ops Mika Kuoppala
2026-09-03 15:22   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 10/27] drm/xe/eudebug: Introduce vm bind and vm bind debug data events Mika Kuoppala
2026-09-03 15:26   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 11/27] drm/xe/eudebug: Add ufence events with acks Mika Kuoppala
2026-09-03 15:20   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 12/27] drm/xe/eudebug: Add vm open/pread/pwrite Mika Kuoppala
2026-09-03 15:27   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 13/27] drm/xe/eudebug: Add userptr vm pread/pwrite Mika Kuoppala
2026-09-03 15:24   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 14/27] drm/xe/eudebug: Add hw enablement Mika Kuoppala
2026-09-03 15:15   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 15/27] drm/xe/eudebug: Introduce EU control interface Mika Kuoppala
2026-09-03 15:34   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 16/27] drm/xe/eudebug: Introduce per device attention scan worker Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 17/27] drm/xe/eudebug_test: Introduce eudebug live tests Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 18/27] drm/xe: Implement SR-IOV and eudebug exclusivity Mika Kuoppala
2026-09-03 15:32   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 19/27] drm/xe: Add xe_client_debugfs and introduce debug_data file Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 20/27] drm/xe/pagefault: export pagefault queue properties Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 21/27] drm/xe/eudebug: Add read/count/compare helper for eu attention Mika Kuoppala
2026-09-03 15:31   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 22/27] drm/xe/vm: Support for adding null page VMA to VM on request Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 23/27] drm/xe/vm: Add xe_vm_svm_vma_subtract() to carve out a sub-range from an SVM VMA Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 24/27] drm/xe: Support for xe_vma_unbind() Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 25/27] drm/xe: export prep_vma_destroy as xe_vm_prep_vma_destroy Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 26/27] drm/xe/eudebug: Introduce EU pagefault handling interface Mika Kuoppala
2026-09-03 15:43   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 27/27] drm/xe/eudebug: Enable EU pagefault handling Mika Kuoppala
2026-09-03 15:46   ` sashiko-bot
2026-09-03 15:35 ` ✗ CI.checkpatch: warning for Intel Xe GPU Debug Support (eudebug) v10 Patchwork
2026-09-03 15:37 ` ✓ CI.KUnit: success " Patchwork
2026-09-03 15:53 ` ✗ CI.checksparse: warning " Patchwork
2026-09-03 16:17 ` ✓ Xe.CI.BAT: success " Patchwork
2026-09-03 16:30 ` [PATCH v10 00/27] " Rodrigo Vivi
2026-09-04  3:21 ` ✗ Xe.CI.FULL: failure for " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260903152224.AD48C1F00A3F@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=mika.kuoppala@linux.intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox