drm_pagemap_dev_unhold_work() processes an unbatched llist of pagemaps
on the system workqueue. This creates potential latency traps during
heavy teardown cycles due to two compounding factors:

1. The loop contains a per-item drm_dbg() trace. If dynamic debug is
enabled on a slow serial console, synchronous print latency scales
linearly with the batch size and can block worker execution.
2. The llist can accumulate large batches during intensive unmap
operations, monopolizing worker resources without yielding.

This causes below log warning:
"workqueue: drm_pagemap_dev_unhold_work [drm_gpusvm_helper] hogged CPU
for >10000us 19 times, consider switching to WQ_UNBOUND"

Fix this by implementing a latency mitigation strategy:
- Move the teardown work to system_unbound_wq to avoid tying up
  per-CPU bound worker pools.
- Add cond_resched() inside the teardown loop to yield the CPU during
  large batch teardowns, ensuring system responsiveness.
- Convert the drm_dbg() trace to drm_dbg_ratelimited() to mitigate
  serial console bottlenecks while preserving debuggability.

v2:
  - Changed workqueue from deprecated system_unbound_wq to system_dfl_wq.
    (Arvind Yadav)

Fixes: a26084328ac4 ("drm/pagemap, drm/xe: Manage drm_pagemap provider 
lifetimes")
Reviewed-by: Arvind Yadav <[email protected]>
Reviewed-by: Thomas Hellström <[email protected]>
Signed-off-by: Varun Gupta <[email protected]>
---
 drivers/gpu/drm/drm_pagemap.c | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/drm_pagemap.c b/drivers/gpu/drm/drm_pagemap.c
index 892b325fa99b..0867bfbde5c5 100644
--- a/drivers/gpu/drm/drm_pagemap.c
+++ b/drivers/gpu/drm/drm_pagemap.c
@@ -986,7 +986,7 @@ static void drm_pagemap_release(struct kref *ref)
        dpagemap->dev_hold = NULL;
        drm_pagemap_shrinker_add(dpagemap);
        llist_add(&dev_hold->link, &drm_pagemap_unhold_list);
-       schedule_work(&drm_pagemap_work);
+       queue_work(system_dfl_wq, &drm_pagemap_work);
        /*
         * Here, either the provider device is still alive, since if called from
         * page_free(), the caller is holding a reference on the dev_pagemap,
@@ -1009,10 +1009,11 @@ static void drm_pagemap_dev_unhold_work(struct 
work_struct *work)
                struct drm_device *drm = dev_hold->drm;
                struct module *module = drm->driver->fops->owner;
 
-               drm_dbg(drm, "Releasing reference on provider device and 
module.\n");
+               drm_dbg_ratelimited(drm, "Releasing reference on provider 
device and module.\n");
                drm_dev_put(drm);
                module_put(module);
                kfree(dev_hold);
+               cond_resched();
        }
 }
 
-- 
2.43.0

Reply via email to