On 3/18/2026 12:10 PM, Mallesh Koujalagi wrote:
Add documentation for the DRM_WEDGE_RECOVERY_COLD_RESET recovery
method introduced for handling power management unit errors. This method is
designated for severe errors that compromise core device functionality
and are unrecoverable via recovery mechanisms such as driver reload or PCIe
bus reset. The documentation clarifies when this recovery method should be
used and its implications for userspace applications.

v2:
- Add several instead of number to avoid update. (Jani)

Signed-off-by: Mallesh Koujalagi <[email protected]>
---
  Documentation/gpu/drm-uapi.rst | 73 +++++++++++++++++++++++++++++++++-
  1 file changed, 72 insertions(+), 1 deletion(-)

diff --git a/Documentation/gpu/drm-uapi.rst b/Documentation/gpu/drm-uapi.rst
index d98428a592f1..5b63f1c17b9b 100644
--- a/Documentation/gpu/drm-uapi.rst
+++ b/Documentation/gpu/drm-uapi.rst
@@ -418,7 +418,7 @@ needed.
  Recovery
  --------
-Current implementation defines four recovery methods, out of which, drivers
+Current implementation defines several recovery methods, out of which, drivers
  can use any one, multiple or none. Method(s) of choice will be sent in the
  uevent environment as ``WEDGED=<method1>[,..,<methodN>]`` in order of less to
  more side-effects. See the section `Vendor Specific Recovery`_
@@ -435,6 +435,7 @@ following expectations.
      rebind          unbind + bind driver
      bus-reset       unbind + bus reset/re-enumeration + bind
      vendor-specific vendor specific recovery method
+    cold-reset      full device cold reset required
      unknown         consumer policy
      =============== ========================================
@@ -446,6 +447,27 @@ telemetry information (devcoredump, syslog). This is useful because the first
  hang is usually the most critical one which can result in consequential hangs 
or
  complete wedging.
+Cold Reset Recovery
+-------------------
+
+The ``WEDGED=cold-reset`` event indicates that the device has encountered
+power management unit errors that affect core functionality that cannot be

Power management errors may be only xe usecase.  Keep the documentation vendor-agnostic.
Some vendors may want to use it for a different usecase

+resolved through recovery mechanisms.
+
+This recovery method is reserved for power management unit error conditions 
where the
Same as above

Thanks
Riana


+device state cannot be restored via:
+
+- Driver unbind/rebind operations
+- PCIe bus reset and re-enumeration
+- Device Function Level Reset (FLR)
+- Warm device resets
+
+Such power management unit error state typically persists across all 
software-based
+recovery attempts. Only a complete device power cycle can restore
+normal operation.
+
+Upon receiving a ``WEDGED=cold-reset`` event, userspace should initiate
+a full cold reset of the affected device to restore functionality.
Vendor Specific Recovery
  ------------------------
@@ -524,6 +546,55 @@ Recovery script::
      echo -n $DEVICE > $DRIVER/unbind
      echo -n $DEVICE > $DRIVER/bind
+Example - cold-reset
+--------------------
+
+Udev rule::
+
+    SUBSYSTEM=="drm", ENV{WEDGED}=="cold-reset", DEVPATH=="*/drm/card[0-9]",
+    RUN+="/path/to/cold-reset.sh $env{DEVPATH}"
+
+Recovery script::
+
+    #!/bin/sh
+
+    [ -z "$1" ] && echo "Usage: $0 <device-path>" && exit 1
+
+    # Get device
+    DEVPATH=$(readlink -f /sys/$1/device 2>/dev/null || readlink -f /sys/$1)
+    DEVICE=$(basename $DEVPATH)
+
+    echo "Cold reset: $DEVICE"
+
+    # Try slot power reset first
+    SLOT=$(find /sys/bus/pci/slots/ -type l 2>/dev/null | while read slot; do
+           ADDR=$(cat "$slot" 2>/dev/null)
+           [ -n "$ADDR" ] && echo "$DEVICE" | grep -q "^$ADDR" && basename $(dirname 
"$slot") && break
+    done)
+
+    if [ -n "$SLOT" ]; then
+       echo "Using slot $SLOT"
+
+       # Unbind driver
+       [ -e "/sys/bus/pci/devices/$DEVICE/driver" ] && \
+       echo "$DEVICE" > /sys/bus/pci/devices/$DEVICE/driver/unbind 2>/dev/null
+
+       # Remove device
+       echo 1 > /sys/bus/pci/devices/$DEVICE/remove
+
+       # Power cycle slot
+       echo 0 > /sys/bus/pci/slots/$SLOT/power
+       sleep 2
+       echo 1 > /sys/bus/pci/slots/$SLOT/power
+       sleep 1
+
+       # Rescan
+       echo 1 > /sys/bus/pci/rescan
+       echo "Done!"
+    else
+       echo "No slot found"
+    fi
+
  Customization
  -------------

Reply via email to