Describe the crash memaction registry, including its kernel and userspace interfaces, allocation lifetime rules, and current limitations.
Signed-off-by: Jan Sebastian Götte <[email protected]> --- Documentation/mm/crash_memaction.rst | 59 ++++++++++++++++++++++++++++++++++++ Documentation/mm/index.rst | 1 + MAINTAINERS | 1 + 3 files changed, 61 insertions(+) diff --git a/Documentation/mm/crash_memaction.rst b/Documentation/mm/crash_memaction.rst new file mode 100644 index 000000000000..3658bff7693c --- /dev/null +++ b/Documentation/mm/crash_memaction.rst @@ -0,0 +1,59 @@ +.. SPDX-License-Identifier: GPL-2.0 + +==================== +Crash memory actions +==================== + +`CONFIG_CRASH_MEMACTION` provides a mechanism for the running kernel to +communicate a single bit attribute per memory page to a kdump kernel. This +feature can be used by the kdump kernel to exclude certain pages (e.g. holding +crypto secrets, or cache) from the system memory dump, or to wipe their contents +after a crash. + +The meaning of the registry bitmap is set by the ``crash_memaction=`` kernel +cmdline parameter. The running kernel only sets the bits for the tracked pages, +and it's up to the kdump kernel to do something with them. + + ``crash_memaction={secret|cache}`` + +Right now, ``secret`` tracks pages containing kernel crypto keys as well as +pages explicitly marked using ``madvise(2)``. ``cache`` currently only tracks +pages marked from userspace using ``madvise(2)``. + +Kernel users +============ + +Kernel code can mark a virtual range in the linear map or in vmalloc space:: + + void crash_memaction_mark(void *addr, size_t size, int types); + void crash_memaction_unmark(void *addr, size_t size); + +Marking is safe from any context and cannot fail. Markings last until the page +frame is handed out to another user to cover stale data left over after the page +is free'd. + +For objects smaller than a page, lib/secret_pool provides allocations from a +"secret" marked kmem_buckets. Objects can be allocated through secret_pool +without the need for any additional lifetime/marking tracking. + +Userspace interface +=================== + +For userspace code, ``MADV_CRASH_SECRET``, ``MADV_CRASH_CACHE`` and +``MADV_CRASH_RESET`` are provided for use with ``madvise(2)``. Marks show up in +``/proc/pid/smaps`` as ``cm`` VMA flag. Only anonymous, shmem and hugetlb ranges +can be marked. Marking is transparent to swapping and lazy allocation. + +Limitations +=========== + +* This mechanism is best-effort. During a crash, nothing can be guaranteed. At + page level granularity, some over- or under-marking is to be expected. +* Marks are tracked only for *mapped* folios, so a shmem file only ever written + with ``write(2)`` cannot be covered. +* A single ``madvise(2)`` call on a folio mapped into several processes globally + (un)marks it for all of them. +* Marks are tracked at page level granularity. There is no refcounting. When + trying to mark smaller objects, use lib/secret_pool or a similar mechanism. +* Memory hotplug is not supported at this time. When memory is hotplugged, the + newly hotplugged memory will not be included in the bitmap. diff --git a/Documentation/mm/index.rst b/Documentation/mm/index.rst index 13a79f5d092c..4cdfc5030a56 100644 --- a/Documentation/mm/index.rst +++ b/Documentation/mm/index.rst @@ -54,6 +54,7 @@ documentation, or deleted if it has served its purpose. allocation-profiling arch_pgtable_helpers balance + crash_memaction damon/index free_page_reporting hmm diff --git a/MAINTAINERS b/MAINTAINERS index bcb11c2138bc..269b8be6c222 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -14256,6 +14256,7 @@ L: [email protected] S: Maintained T: git git://git.kernel.org/pub/scm/linux/kernel/git/liveupdate/linux.git F: Documentation/admin-guide/kdump/ +F: Documentation/mm/crash_memaction.rst F: fs/proc/vmcore.c F: include/linux/crash_core.h F: include/linux/crash_dump.h -- 2.55.0

