On Tue, Sep 29, 2026 at 10:46:50AM +0200, David Hildenbrand (Arm) wrote:
> On 9/29/26 10:43, Baolin Wang wrote:
> > 
> > 
> > On 9/29/26 4:38 PM, David Hildenbrand (Arm) wrote:
> >> On 9/23/26 17:29, Yeoreum Yun wrote:
> >>> There are intermittent failures in collapse_max_ptes_swap() and
> >>> collapse_max_ptes_shared() when using the khugepaged_context:
> >>>
> >>>    // while running ./khugepaged -s 2
> >>>
> >>>    # Run test: collapse_max_ptes_shared (khugepaged:anon)
> >>>    # Allocate huge page... OK
> >>>    # Share huge page over fork()... OK
> >>>    # Trigger CoW on page 1023 of 2048... OK
> >>>    # Maybe collapse with max_ptes_shared exceeded.... OK
> >>>    # Trigger CoW on page 1024 of 2048... Fail
> >>>    Bail out! Unexpected huge page
> >>>    # Planned tests != run tests (26 != 23)
> >>>    # Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0
> >>>
> >>>    # Run test: collapse_max_ptes_swap (khugepaged:anon)
> >>>    # Swapout 257 of 2048 pages... OK
> >>>    # Maybe collapse with max_ptes_swap exceeded.... OK
> >>>    # Swapout 256 of 2048 pages... OK
> >>>    Bail out! Unexpected huge page
> >>>    # Planned tests != run tests (26 != 17)
> >>>    # Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0
> >>>
> >>> This happens because khugepaged may collapse the pages before 
> >>> wait_for_scan()
> >>> is called, causing a sanity check that expects uncollapsed pages to fail.
> >>>
> >>> For example, in collapse_max_ptes_swap(), after faulting the pages back in
> >>> and paging out up to max_ptes_swap pages, khugepaged may collapse them 
> >>> again
> >>> before c->collapse() is called.
> >>>
> >>> To prevent this, mark the VMA with MADV_NOHUGEPAGE after it has been
> >>> collapsed by wait_for_scan() for anon. This prevents khugepaged from
> >>> collapsing it again before c->collapse() is called.
> >>>
> >>> This failure was observed on NVIDIA Spark with 16KB page.
> >>>
> >>> Reviewed-by: Baolin Wang <[email protected]>
> >>> Tested-by: Baolin Wang <[email protected]>
> >>> Signed-off-by: Yeoreum Yun <[email protected]>
> >>> ---
> >>>   tools/testing/selftests/mm/khugepaged.c | 3 +++
> >>>   1 file changed, 3 insertions(+)
> >>>
> >>> diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/
> >>> selftests/mm/khugepaged.c
> >>> index 2aa7c9197158..b0cb02bf1a73 100644
> >>> --- a/tools/testing/selftests/mm/khugepaged.c
> >>> +++ b/tools/testing/selftests/mm/khugepaged.c
> >>> @@ -618,6 +618,9 @@ static bool wait_for_scan(const char *msg, char *p,
> >>> size_t len,
> >>>           usleep(TICK);
> >>>       }
> >>>   +    if (is_anon(ops))
> >>> +        madvise(p, len, MADV_NOHUGEPAGE);
> >>> +
> >>
> >> Any reason we just do that unconditionally?
> > 
> > Although it's a bit messy, as I mentioned before [1], unconditionally 
> > setting
> > MADV_NOHUGEPAGE will break shmem testing. Maybe add some comments.
> 
> Ah, thanks for clarifying. The problem really is that we cannot undo a
> MADV_HUGEPAGE (give me hugepages) cleanly. We can only go to the other extreme
> (no huge pages).
> 
> Yes, let's please add a comment describing why we limit it to anon.

Okay.

-- 
Sincerely,
Yeoreum Yun

Reply via email to