On 19-Sep-26 3:45 AM, Matthew Brost wrote:
On Fri, Sep 18, 2026 at 01:43:21PM +0530, Varun Gupta wrote:
drm_pagemap_dev_unhold_work() processes an unbatched llist of pagemaps
on the system workqueue. This creates potential latency traps during
heavy teardown cycles due to two compounding factors:

1. The loop contains a per-item drm_dbg() trace. If dynamic debug is
enabled on a slow serial console, synchronous print latency scales
linearly with the batch size and can block worker execution.
2. The llist can accumulate large batches during intensive unmap
operations, monopolizing worker resources without yielding.

This causes below log warning:
"workqueue: drm_pagemap_dev_unhold_work [drm_gpusvm_helper] hogged CPU
for >10000us 19 times, consider switching to WQ_UNBOUND"

Fix this by implementing a latency mitigation strategy:
- Move the teardown work to system_unbound_wq to avoid tying up
   per-CPU bound worker pools.
- Add cond_resched() inside the teardown loop to yield the CPU during
   large batch teardowns, ensuring system responsiveness.
- Convert the drm_dbg() trace to drm_dbg_ratelimited() to mitigate
   serial console bottlenecks while preserving debuggability.

v2:
   - Changed workqueue from deprecated system_unbound_wq to system_dfl_wq.
     (Arvind Yadav)

Fixes: a26084328ac4 ("drm/pagemap, drm/xe: Manage drm_pagemap provider 
lifetimes")
Reviewed-by: Arvind Yadav <[email protected]>
Reviewed-by: Thomas Hellström <[email protected]>
Signed-off-by: Varun Gupta <[email protected]>
Did this patch break our CI [1] [2]? Varun can you take a look at a
follow up or work with CI so this isn't flagged?

Thanks,
Matt

[1] 
https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-174512v1/bat-bmg-2/igt@[email protected]
[2] 
https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-174512v1/bat-bmg-vm/igt@[email protected]
Thanks for pointing out Matt, will talk to someone from CI, if they cannot handle at their end,
I'll send a follow-up changing it back to drm_dbg.

Varun

---
  drivers/gpu/drm/drm_pagemap.c | 5 +++--
  1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/drm_pagemap.c b/drivers/gpu/drm/drm_pagemap.c
index 892b325fa99b..0867bfbde5c5 100644
--- a/drivers/gpu/drm/drm_pagemap.c
+++ b/drivers/gpu/drm/drm_pagemap.c
@@ -986,7 +986,7 @@ static void drm_pagemap_release(struct kref *ref)
        dpagemap->dev_hold = NULL;
        drm_pagemap_shrinker_add(dpagemap);
        llist_add(&dev_hold->link, &drm_pagemap_unhold_list);
-       schedule_work(&drm_pagemap_work);
+       queue_work(system_dfl_wq, &drm_pagemap_work);
        /*
         * Here, either the provider device is still alive, since if called from
         * page_free(), the caller is holding a reference on the dev_pagemap,
@@ -1009,10 +1009,11 @@ static void drm_pagemap_dev_unhold_work(struct 
work_struct *work)
                struct drm_device *drm = dev_hold->drm;
                struct module *module = drm->driver->fops->owner;
- drm_dbg(drm, "Releasing reference on provider device and module.\n");
+               drm_dbg_ratelimited(drm, "Releasing reference on provider device and 
module.\n");
                drm_dev_put(drm);
                module_put(module);
                kfree(dev_hold);
+               cond_resched();
        }
  }
--
2.43.0

Reply via email to