On 9/29/26 10:43, Baolin Wang wrote: > > > On 9/29/26 4:38 PM, David Hildenbrand (Arm) wrote: >> On 9/23/26 17:29, Yeoreum Yun wrote: >>> There are intermittent failures in collapse_max_ptes_swap() and >>> collapse_max_ptes_shared() when using the khugepaged_context: >>> >>> // while running ./khugepaged -s 2 >>> >>> # Run test: collapse_max_ptes_shared (khugepaged:anon) >>> # Allocate huge page... OK >>> # Share huge page over fork()... OK >>> # Trigger CoW on page 1023 of 2048... OK >>> # Maybe collapse with max_ptes_shared exceeded.... OK >>> # Trigger CoW on page 1024 of 2048... Fail >>> Bail out! Unexpected huge page >>> # Planned tests != run tests (26 != 23) >>> # Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0 >>> >>> # Run test: collapse_max_ptes_swap (khugepaged:anon) >>> # Swapout 257 of 2048 pages... OK >>> # Maybe collapse with max_ptes_swap exceeded.... OK >>> # Swapout 256 of 2048 pages... OK >>> Bail out! Unexpected huge page >>> # Planned tests != run tests (26 != 17) >>> # Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0 >>> >>> This happens because khugepaged may collapse the pages before >>> wait_for_scan() >>> is called, causing a sanity check that expects uncollapsed pages to fail. >>> >>> For example, in collapse_max_ptes_swap(), after faulting the pages back in >>> and paging out up to max_ptes_swap pages, khugepaged may collapse them again >>> before c->collapse() is called. >>> >>> To prevent this, mark the VMA with MADV_NOHUGEPAGE after it has been >>> collapsed by wait_for_scan() for anon. This prevents khugepaged from >>> collapsing it again before c->collapse() is called. >>> >>> This failure was observed on NVIDIA Spark with 16KB page. >>> >>> Reviewed-by: Baolin Wang <[email protected]> >>> Tested-by: Baolin Wang <[email protected]> >>> Signed-off-by: Yeoreum Yun <[email protected]> >>> --- >>> tools/testing/selftests/mm/khugepaged.c | 3 +++ >>> 1 file changed, 3 insertions(+) >>> >>> diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/ >>> selftests/mm/khugepaged.c >>> index 2aa7c9197158..b0cb02bf1a73 100644 >>> --- a/tools/testing/selftests/mm/khugepaged.c >>> +++ b/tools/testing/selftests/mm/khugepaged.c >>> @@ -618,6 +618,9 @@ static bool wait_for_scan(const char *msg, char *p, >>> size_t len, >>> usleep(TICK); >>> } >>> + if (is_anon(ops)) >>> + madvise(p, len, MADV_NOHUGEPAGE); >>> + >> >> Any reason we just do that unconditionally? > > Although it's a bit messy, as I mentioned before [1], unconditionally setting > MADV_NOHUGEPAGE will break shmem testing. Maybe add some comments.
Ah, thanks for clarifying. The problem really is that we cannot undo a MADV_HUGEPAGE (give me hugepages) cleanly. We can only go to the other extreme (no huge pages). Yes, let's please add a comment describing why we limit it to anon. -- Cheers, David

