On 10-09-2026 16:15, Thomas Hellström wrote:
On Thu, 2026-08-27 at 15:47 +0530, Arvind Yadav wrote:
A permanently wedged device can still access hardware through pending
work,
existing VRAM mappings and background driver paths. This can cause
unwanted
DMA or MMIO access while userspace is preparing device recovery.
This series separates temporary PCI error recovery from permanent
wedging and
introduces a common device-I/O-blocked state.
After a permanent wedge, the driver:
- stops queued VM, SVM and GuC work.
- shuts down the display.
- invalidates existing CPU mappings to VRAM.
- maps later CPU faults to a per-BO dummy page.
- stops interrupts and clears PCI bus mastering.
- rejects new VRAM allocations.
- notifies userspace only after isolation completes.
- drops queued page faults and stops VM, SVM and GuC work.
System suspend waits for wedge isolation to complete before skipping
the
normal suspend path. The worker cannot reach synchronize_srcu() until
all
GTs have been declared wedged and their pending fences have been
signalled.
The device remains isolated until userspace performs the advertised
recovery,
such as driver rebind or bus reset.
The series also fixes IRQ uninstall so requested handlers are freed
after
interrupts have already been suspended.
Arvind Yadav (13):
drm/xe/irq: Always free requested IRQs on uninstall
drm/xe: Separate AER reset state from device wedging
drm/xe: Drop queued page faults when device I/O is blocked
drm/xe: Stop VM work when device I/O is blocked
drm/xe: Send wedged notification from a worker
drm/xe: Reuse one dummy page per BO after wedge
drm/xe: Invalidate existing VRAM mappings on wedge
drm/xe/irq: Serialize IRQ suspend and resume
drm/xe: Isolate a wedged device before notifying userspace
drm/xe/ttm: Reject VRAM allocations on wedged devices
drm/xe/guc: Skip timeout recovery on a wedged device
drm/xe: Skip PM notifier preparation for wedged devices
drm/xe: Block BO VM access when device I/O is unavailable
drivers/gpu/drm/xe/display/xe_display.c | 18 +++-
drivers/gpu/drm/xe/xe_bo.c | 101 +++++++++++++++++++-
drivers/gpu/drm/xe/xe_bo.h | 1 +
drivers/gpu/drm/xe/xe_bo_types.h | 4 +
drivers/gpu/drm/xe/xe_device.c | 118 ++++++++++++++++++++--
--
drivers/gpu/drm/xe/xe_device.h | 12 +++
drivers/gpu/drm/xe/xe_device_types.h | 26 ++++++
drivers/gpu/drm/xe/xe_guc_ct.c | 4 +-
drivers/gpu/drm/xe/xe_guc_pc.c | 10 +-
drivers/gpu/drm/xe/xe_guc_rc.c | 4 +-
drivers/gpu/drm/xe/xe_guc_submit.c | 13 ++-
drivers/gpu/drm/xe/xe_guc_tlb_inval.c | 8 +-
drivers/gpu/drm/xe/xe_irq.c | 32 +++++--
drivers/gpu/drm/xe/xe_pagefault.c | 40 ++++++++
drivers/gpu/drm/xe/xe_pci.c | 23 ++++-
drivers/gpu/drm/xe/xe_pci_error.c | 22 ++---
drivers/gpu/drm/xe/xe_pm.c | 27 ++++++
drivers/gpu/drm/xe/xe_sriov_pf.c | 2 +-
drivers/gpu/drm/xe/xe_svm.c | 21 ++++-
drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 4 +
drivers/gpu/drm/xe/xe_vm.c | 16 ++++
21 files changed, 446 insertions(+), 60 deletions(-)
Hi, Arvind.
I think this series lacks context to some extent. Much of the code is
also inheritly racy [1]. because current hardware accesses aren't
drained before reset or wedging. What I think is needed here is a
design detailing
1) What type of user-space accesses and workqueues are completely
blocked (for example sharing the newly introduced vram_userfault.srcu)
This can perhaps also be combined with the drm SRCU since I figure
there might be multiple places where we ensure hardware is present
using drm_dev_enter / drm_dev_exit and also want to check whether it's
wedged or not.
See an example in
https://gitlab.freedesktop.org/thomash/kernel/-/commits/wedge?ref_type=heads
(Note that this is just an example, not fully adapted to the intentions
in this series).
Thanks Thomas. I will document all userspace and worker paths that need
protection and rework them around one common synchronization domain.
I will also check whether it can share the DRM SRCU.
2) How is user-space supposed to respond to a recoverable PCIe reset?
Currently? If a change is planned, In the future? Is it the case that
currently we flag "In reset" and then force the user to close its drm
connection and re-discover a recovered device? Is this about to change
and hence the distinction between PCIe reset and permanent wedging?
It looks like AMD might have an rwsem that just blocks callers over a
reset and perhaps notifies them that, for example VRAM is lost. In any
case, in addition to locking out new callers we need to drain old
callers:
[1]
Thread A Thread B
is_io_allowed()
disable_io()
access_hardware() reset()
Bang!
Thanks Thomas. Agreed that the current xe_device_io_blocked checks only
block new work. They do not drain a caller that already passed the check.
As per your suggestion I will rework on this series.
Thanks,
Arvind
Thanks,
Thomas