* [PATCH] drm/pagemap: Prevent CPU stalls during unbounded pagemap teardown
@ 2026-09-17 17:03 Varun Gupta
2026-09-18 7:05 ` Thomas Hellström
0 siblings, 1 reply; 3+ messages in thread
From: Varun Gupta @ 2026-09-17 17:03 UTC (permalink / raw)
To: dri-devel; +Cc: matthew.brost, thomas.hellstrom
drm_pagemap_dev_unhold_work() processes an unbatched llist of pagemaps
on the system workqueue. This creates potential latency traps during
heavy teardown cycles due to two compounding factors:
1. The loop contains a per-item drm_dbg() trace. If dynamic debug is
enabled on a slow serial console, synchronous print latency scales
linearly with the batch size and can block worker execution.
2. The llist can accumulate large batches during intensive unmap
operations, monopolizing worker resources without yielding.
This causes below log warning:
"workqueue: drm_pagemap_dev_unhold_work [drm_gpusvm_helper] hogged CPU
for >10000us 19 times, consider switching to WQ_UNBOUND"
Fix this by implementing a latency mitigation strategy:
- Move the teardown work to system_unbound_wq to avoid tying up
per-CPU bound worker pools.
- Add cond_resched() inside the teardown loop to yield the CPU during
large batch teardowns, ensuring system responsiveness.
- Convert the drm_dbg() trace to drm_dbg_ratelimited() to mitigate
serial console bottlenecks while preserving debuggability.
Fixes: a26084328ac4 ("drm/pagemap, drm/xe: Manage drm_pagemap provider lifetimes")
Signed-off-by: Varun Gupta <varun.gupta@intel.com>
---
drivers/gpu/drm/drm_pagemap.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/drm_pagemap.c b/drivers/gpu/drm/drm_pagemap.c
index 892b325fa99b..f8d5428f1750 100644
--- a/drivers/gpu/drm/drm_pagemap.c
+++ b/drivers/gpu/drm/drm_pagemap.c
@@ -986,7 +986,7 @@ static void drm_pagemap_release(struct kref *ref)
dpagemap->dev_hold = NULL;
drm_pagemap_shrinker_add(dpagemap);
llist_add(&dev_hold->link, &drm_pagemap_unhold_list);
- schedule_work(&drm_pagemap_work);
+ queue_work(system_unbound_wq, &drm_pagemap_work);
/*
* Here, either the provider device is still alive, since if called from
* page_free(), the caller is holding a reference on the dev_pagemap,
@@ -1009,10 +1009,11 @@ static void drm_pagemap_dev_unhold_work(struct work_struct *work)
struct drm_device *drm = dev_hold->drm;
struct module *module = drm->driver->fops->owner;
- drm_dbg(drm, "Releasing reference on provider device and module.\n");
+ drm_dbg_ratelimited(drm, "Releasing reference on provider device and module.\n");
drm_dev_put(drm);
module_put(module);
kfree(dev_hold);
+ cond_resched();
}
}
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
* Re: [PATCH] drm/pagemap: Prevent CPU stalls during unbounded pagemap teardown
2026-09-17 17:03 [PATCH] drm/pagemap: Prevent CPU stalls during unbounded pagemap teardown Varun Gupta
@ 2026-09-18 7:05 ` Thomas Hellström
2026-09-18 8:06 ` Yadav, Arvind
0 siblings, 1 reply; 3+ messages in thread
From: Thomas Hellström @ 2026-09-18 7:05 UTC (permalink / raw)
To: Varun Gupta, dri-devel; +Cc: matthew.brost
On Thu, 2026-09-17 at 22:33 +0530, Varun Gupta wrote:
> drm_pagemap_dev_unhold_work() processes an unbatched llist of
> pagemaps
> on the system workqueue. This creates potential latency traps during
> heavy teardown cycles due to two compounding factors:
>
> 1. The loop contains a per-item drm_dbg() trace. If dynamic debug is
> enabled on a slow serial console, synchronous print latency scales
> linearly with the batch size and can block worker execution.
> 2. The llist can accumulate large batches during intensive unmap
> operations, monopolizing worker resources without yielding.
>
> This causes below log warning:
> "workqueue: drm_pagemap_dev_unhold_work [drm_gpusvm_helper] hogged
> CPU
> for >10000us 19 times, consider switching to WQ_UNBOUND"
>
> Fix this by implementing a latency mitigation strategy:
> - Move the teardown work to system_unbound_wq to avoid tying up
> per-CPU bound worker pools.
> - Add cond_resched() inside the teardown loop to yield the CPU during
> large batch teardowns, ensuring system responsiveness.
> - Convert the drm_dbg() trace to drm_dbg_ratelimited() to mitigate
> serial console bottlenecks while preserving debuggability.
>
> Fixes: a26084328ac4 ("drm/pagemap, drm/xe: Manage drm_pagemap
> provider lifetimes")
> Signed-off-by: Varun Gupta <varun.gupta@intel.com>
LGTM.
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> ---
> drivers/gpu/drm/drm_pagemap.c | 5 +++--
> 1 file changed, 3 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/drm_pagemap.c
> b/drivers/gpu/drm/drm_pagemap.c
> index 892b325fa99b..f8d5428f1750 100644
> --- a/drivers/gpu/drm/drm_pagemap.c
> +++ b/drivers/gpu/drm/drm_pagemap.c
> @@ -986,7 +986,7 @@ static void drm_pagemap_release(struct kref *ref)
> dpagemap->dev_hold = NULL;
> drm_pagemap_shrinker_add(dpagemap);
> llist_add(&dev_hold->link, &drm_pagemap_unhold_list);
> - schedule_work(&drm_pagemap_work);
> + queue_work(system_unbound_wq, &drm_pagemap_work);
> /*
> * Here, either the provider device is still alive, since if
> called from
> * page_free(), the caller is holding a reference on the
> dev_pagemap,
> @@ -1009,10 +1009,11 @@ static void
> drm_pagemap_dev_unhold_work(struct work_struct *work)
> struct drm_device *drm = dev_hold->drm;
> struct module *module = drm->driver->fops->owner;
>
> - drm_dbg(drm, "Releasing reference on provider device
> and module.\n");
> + drm_dbg_ratelimited(drm, "Releasing reference on
> provider device and module.\n");
> drm_dev_put(drm);
> module_put(module);
> kfree(dev_hold);
> + cond_resched();
> }
> }
>
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] drm/pagemap: Prevent CPU stalls during unbounded pagemap teardown
2026-09-18 7:05 ` Thomas Hellström
@ 2026-09-18 8:06 ` Yadav, Arvind
0 siblings, 0 replies; 3+ messages in thread
From: Yadav, Arvind @ 2026-09-18 8:06 UTC (permalink / raw)
To: Thomas Hellström, Varun Gupta, dri-devel; +Cc: matthew.brost
[-- Attachment #1: Type: text/plain, Size: 2948 bytes --]
On 18-09-2026 12:35, Thomas Hellström wrote:
> On Thu, 2026-09-17 at 22:33 +0530, Varun Gupta wrote:
>> drm_pagemap_dev_unhold_work() processes an unbatched llist of
>> pagemaps
>> on the system workqueue. This creates potential latency traps during
>> heavy teardown cycles due to two compounding factors:
>>
>> 1. The loop contains a per-item drm_dbg() trace. If dynamic debug is
>> enabled on a slow serial console, synchronous print latency scales
>> linearly with the batch size and can block worker execution.
>> 2. The llist can accumulate large batches during intensive unmap
>> operations, monopolizing worker resources without yielding.
>>
>> This causes below log warning:
>> "workqueue: drm_pagemap_dev_unhold_work [drm_gpusvm_helper] hogged
>> CPU
>> for >10000us 19 times, consider switching to WQ_UNBOUND"
>>
>> Fix this by implementing a latency mitigation strategy:
>> - Move the teardown work to system_unbound_wq to avoid tying up
>> per-CPU bound worker pools.
>> - Add cond_resched() inside the teardown loop to yield the CPU during
>> large batch teardowns, ensuring system responsiveness.
>> - Convert the drm_dbg() trace to drm_dbg_ratelimited() to mitigate
>> serial console bottlenecks while preserving debuggability.
>>
>> Fixes: a26084328ac4 ("drm/pagemap, drm/xe: Manage drm_pagemap
>> provider lifetimes")
>> Signed-off-by: Varun Gupta<varun.gupta@intel.com>
> LGTM.
> Reviewed-by: Thomas Hellström<thomas.hellstrom@linux.intel.com>
>
>> ---
>> drivers/gpu/drm/drm_pagemap.c | 5 +++--
>> 1 file changed, 3 insertions(+), 2 deletions(-)
>>
>> diff --git a/drivers/gpu/drm/drm_pagemap.c
>> b/drivers/gpu/drm/drm_pagemap.c
>> index 892b325fa99b..f8d5428f1750 100644
>> --- a/drivers/gpu/drm/drm_pagemap.c
>> +++ b/drivers/gpu/drm/drm_pagemap.c
>> @@ -986,7 +986,7 @@ static void drm_pagemap_release(struct kref *ref)
>> dpagemap->dev_hold = NULL;
>> drm_pagemap_shrinker_add(dpagemap);
>> llist_add(&dev_hold->link, &drm_pagemap_unhold_list);
>> - schedule_work(&drm_pagemap_work);
>> + queue_work(system_unbound_wq, &drm_pagemap_work);
Instead of system_unbound_wq we should use system_dfl_wq because
system_unbound_wq is deprecated.
Apart from this LGTM:
Reviewed-by: Arvind Yadav <arvind.yadav@intel.com>
>> /*
>> * Here, either the provider device is still alive, since if
>> called from
>> * page_free(), the caller is holding a reference on the
>> dev_pagemap,
>> @@ -1009,10 +1009,11 @@ static void
>> drm_pagemap_dev_unhold_work(struct work_struct *work)
>> struct drm_device *drm = dev_hold->drm;
>> struct module *module = drm->driver->fops->owner;
>>
>> - drm_dbg(drm, "Releasing reference on provider device
>> and module.\n");
>> + drm_dbg_ratelimited(drm, "Releasing reference on
>> provider device and module.\n");
>> drm_dev_put(drm);
>> module_put(module);
>> kfree(dev_hold);
>> + cond_resched();
>> }
>> }
>>
[-- Attachment #2: Type: text/html, Size: 4121 bytes --]
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-18 8:07 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 17:03 [PATCH] drm/pagemap: Prevent CPU stalls during unbounded pagemap teardown Varun Gupta
2026-09-18 7:05 ` Thomas Hellström
2026-09-18 8:06 ` Yadav, Arvind
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox