All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Mika Kuoppala" <mika.kuoppala@linux.intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v10 04/27] drm/xe/eudebug: Introduce discovery for resources
Date: Thu, 03 Sep 2026 15:22:23 +0000	[thread overview]
Message-ID: <20260903152224.AD48C1F00A3F@smtp.kernel.org> (raw)
In-Reply-To: <20260903145952.848051-5-mika.kuoppala@linux.intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] Global ordered workqueue stall via copy_from_user block while holding ioctl_lock.
- [Medium] Event sequence numbers (seqno) can be queued out-of-order in the event stream due to a race between seqno generation and event queuing.
--

commit e2061e72d8e954a17bde35e9c970ee41bc2db0ad
Author: Mika Kuoppala <mika.kuoppala@linux.intel.com>

    drm/xe/eudebug: Introduce discovery for resources

This commit introduces a resource discovery mechanism for debugger connections.
When a debugger connects, all existing resources for the target client are
relayed as events via an ordered workqueue, and DRM ioctls are blocked using
an rw_semaphore until discovery is complete.

> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> index 032a7ec5fbc8d..f10959b898297 100644
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
[ ... ]
> @@ -238,8 +239,12 @@ static long xe_drm_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
>  
>  	ACQUIRE(xe_pm_runtime_ioctl, pm)(xe);
>  	ret = ACQUIRE_ERR(xe_pm_runtime_ioctl, &pm);
> -	if (ret >= 0)
> +	if (ret >= 0) {
> +		bool lock = xe_eudebug_discovery_lock(file, cmd);
>  		ret = drm_ioctl(file, cmd, arg);

[Severity: High]
Does holding the discovery lock across drm_ioctl() create a vector for a
workqueue stall? 

If an unprivileged user maps the ioctl argument pointer to a blocking FUSE
or userfaultfd region, drm_ioctl() can be stalled indefinitely during
copy_from_user() while holding xef->eudebug.ioctl_lock for read.

> +		if (lock)
> +			xe_eudebug_discovery_unlock(file, cmd);
> +	}
>  
>  	return ret;
>  }
[ ... ]
> diff --git a/drivers/gpu/drm/xe/xe_eudebug.c b/drivers/gpu/drm/xe/xe_eudebug.c
> index 835dbd7309af5..7a5b14346e059 100644
> --- a/drivers/gpu/drm/xe/xe_eudebug.c
> +++ b/drivers/gpu/drm/xe/xe_eudebug.c
[ ... ]
> @@ -659,6 +705,87 @@ void xe_eudebug_vm_destroy(struct xe_file *xef, struct xe_vm *vm)
>  	xe_eudebug_put(d);
>  }
>  
> +static struct xe_file *xe_eudebug_target_get(struct xe_eudebug *d)
> +{
[ ... ]
> +static void discover_client(struct xe_eudebug *d)
> +{
> +	struct xe_file *xef;
> +	struct xe_vm *vm;
> +	unsigned long i;
> +	unsigned int vm_count = 0;
> +	int err = 0;
> +
> +	xef = xe_eudebug_target_get(d);
> +	if (!xef) {
> +		complete_all(&d->discovery);
> +		return;
> +	}
> +
> +	/*
> +	 * An xe_eudebug ref is taken for the discovery worker, which holds a
> +	 * ref to the target xe_file; the xe_file in turn holds the vm and
> +	 * exec_queue refs.
> +	 *
> +	 * All the relevant ioctls that manipulate the xe_file resources eudebug
> +	 * cares about are gated behind down_read(&xef->eudebug.ioctl_lock).
> +	 * Thus taking the write lock excludes all those ioctls, which lets us
> +	 * peek into the resource lists without taking their individual locks.
> +	 */
> +	down_write(&xef->eudebug.ioctl_lock);

[Severity: High]
If the target process is blocked in copy_from_user() inside drm_ioctl()
while holding the read lock, will this down_write() block indefinitely?

Because discovery runs on a global ordered workqueue (configured below in
xe_eudebug_init), this single blocked work item could stall all eudebug
discovery operations system-wide, causing a Denial of Service for all
debugger connections.

> +
> +	if (xe_eudebug_detached(d))
> +		goto out;
[ ... ]
> @@ -1126,12 +1266,20 @@ void xe_eudebug_init_early(struct xe_device *xe)
>  void xe_eudebug_init(struct xe_device *xe)
>  {
>  	struct drm_device *dev = &xe->drm;
> +	struct workqueue_struct *wq;
>  	int err;
>  
>  	/* early init failed */
>  	if (xe->eudebug.cap_state == XE_EUDEBUG_CAP_NOT_SUPPORTED)
>  		return;
>  
> +	wq = drmm_alloc_ordered_workqueue(dev, "xe-eudebug-discovery", 0);

[Severity: High]
This allocates the workqueue as ordered (max_active=1), which means a single
stalled work item blocks all other discovery work.

Can the workqueue be allocated differently, or can the ioctl_lock be managed
to avoid wrapping the copy_from_user() operations in drm_ioctl()?

> +	if (IS_ERR(wq)) {
> +		err = PTR_ERR(wq);
> +		goto out_err;
> +	}
> +	xe->eudebug.discovery_wq = wq;
> +
>  	err = sysfs_create_file(&dev->dev->kobj,
>  				&dev_attr_enable_eudebug.attr);
>  	if (err)
[ ... ]

[Severity: Medium]
Can sequence numbers be queued out-of-order in the event stream?

Looking at send_vm_bind_event() in xe_eudebug.c:

	*seqno = atomic_long_inc_return(&d->events.seqno);

	spin_lock(&d->events.lock);
	event = xe_eudebug_prepare_event(d, DRM_XE_EUDEBUG_EVENT_VM_BIND,
					 *seqno, base_flags, sz);

Since the sequence number is incremented atomically outside of the FIFO
spinlock, could two threads acquire sequence numbers N and N+1, but acquire
the FIFO lock in reverse order?

This same pattern appears in _xe_eudebug_add_handle(). Pushing N+1 before
N would break the expected monotonicity of the event stream and could cause
the debugger to drop events.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260903145952.848051-1-mika.kuoppala@linux.intel.com?part=4

  reply	other threads:[~2026-09-03 15:22 UTC|newest]

Thread overview: 61+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 14:59 [PATCH v10 00/27] Intel Xe GPU Debug Support (eudebug) v10 Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 01/27] drm/xe/eudebug: Introduce eudebug interface Mika Kuoppala
2026-09-03 15:16   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 02/27] drm/xe/eudebug: Add documentation Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 03/27] drm/xe/eudebug: Add connection establishment documentation Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 04/27] drm/xe/eudebug: Introduce discovery for resources Mika Kuoppala
2026-09-03 15:22   ` sashiko-bot [this message]
2026-09-07 13:24     ` Joonas Lahtinen
2026-09-09 10:00       ` FUSE deadlocks vs. copy_from_user() and locks (Was: Re: [PATCH v10 04/27] drm/xe/eudebug: Introduce discovery for resources) Joonas Lahtinen
2026-09-09 10:21         ` Simona Vetter
2026-09-09 11:02           ` Christian König
2026-09-09 11:24             ` Joonas Lahtinen
2026-09-09 11:44               ` Miklos Szeredi
2026-09-09 13:03               ` Christian König
2026-09-09 14:51                 ` Joonas Lahtinen
2026-09-09 10:30         ` Miklos Szeredi
2026-09-03 14:59 ` [PATCH v10 05/27] drm/xe: Add EUDEBUG_ENABLE exec queue property Mika Kuoppala
2026-09-03 15:14   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 06/27] drm/xe/eudebug: Introduce exec_queue events Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 07/27] drm/xe/eudebug: Mark guc contexts as debuggable Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 08/27] drm/xe: Remove ifdef in DRM_GPUVA_OP_DRIVER svm subop checking Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 09/27] drm/xe: Introduce ADD_DEBUG_DATA and REMOVE_DEBUG_DATA vm bind ops Mika Kuoppala
2026-09-03 15:22   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 10/27] drm/xe/eudebug: Introduce vm bind and vm bind debug data events Mika Kuoppala
2026-09-03 15:26   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 11/27] drm/xe/eudebug: Add ufence events with acks Mika Kuoppala
2026-09-03 15:20   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 12/27] drm/xe/eudebug: Add vm open/pread/pwrite Mika Kuoppala
2026-09-03 15:27   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 13/27] drm/xe/eudebug: Add userptr vm pread/pwrite Mika Kuoppala
2026-09-03 15:24   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 14/27] drm/xe/eudebug: Add hw enablement Mika Kuoppala
2026-09-03 15:15   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 15/27] drm/xe/eudebug: Introduce EU control interface Mika Kuoppala
2026-09-03 15:34   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 16/27] drm/xe/eudebug: Introduce per device attention scan worker Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 17/27] drm/xe/eudebug_test: Introduce eudebug live tests Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 18/27] drm/xe: Implement SR-IOV and eudebug exclusivity Mika Kuoppala
2026-09-03 15:32   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 19/27] drm/xe: Add xe_client_debugfs and introduce debug_data file Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 20/27] drm/xe/pagefault: export pagefault queue properties Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 21/27] drm/xe/eudebug: Add read/count/compare helper for eu attention Mika Kuoppala
2026-09-03 15:31   ` sashiko-bot
2026-09-03 14:59 ` [PATCH v10 22/27] drm/xe/vm: Support for adding null page VMA to VM on request Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 23/27] drm/xe/vm: Add xe_vm_svm_vma_subtract() to carve out a sub-range from an SVM VMA Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 24/27] drm/xe: Support for xe_vma_unbind() Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 25/27] drm/xe: export prep_vma_destroy as xe_vm_prep_vma_destroy Mika Kuoppala
2026-09-03 14:59 ` [PATCH v10 26/27] drm/xe/eudebug: Introduce EU pagefault handling interface Mika Kuoppala
2026-09-03 15:43   ` sashiko-bot
2026-09-08 15:12     ` Maciej Patelczyk
2026-09-03 14:59 ` [PATCH v10 27/27] drm/xe/eudebug: Enable EU pagefault handling Mika Kuoppala
2026-09-03 15:46   ` sashiko-bot
2026-09-08  9:35     ` Joonas Lahtinen
2026-09-08 15:28     ` Maciej Patelczyk
2026-09-03 15:35 ` ✗ CI.checkpatch: warning for Intel Xe GPU Debug Support (eudebug) v10 Patchwork
2026-09-03 15:37 ` ✓ CI.KUnit: success " Patchwork
2026-09-03 15:53 ` ✗ CI.checksparse: warning " Patchwork
2026-09-03 16:17 ` ✓ Xe.CI.BAT: success " Patchwork
2026-09-03 16:30 ` [PATCH v10 00/27] " Rodrigo Vivi
2026-09-07  5:32   ` Joonas Lahtinen
2026-09-04  3:21 ` ✗ Xe.CI.FULL: failure for " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260903152224.AD48C1F00A3F@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=mika.kuoppala@linux.intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.