On 9/29/26 10:43, Baolin Wang wrote:
> 
> 
> On 9/29/26 4:38 PM, David Hildenbrand (Arm) wrote:
>> On 9/23/26 17:29, Yeoreum Yun wrote:
>>> There are intermittent failures in collapse_max_ptes_swap() and
>>> collapse_max_ptes_shared() when using the khugepaged_context:
>>>
>>>    // while running ./khugepaged -s 2
>>>
>>>    # Run test: collapse_max_ptes_shared (khugepaged:anon)
>>>    # Allocate huge page... OK
>>>    # Share huge page over fork()... OK
>>>    # Trigger CoW on page 1023 of 2048... OK
>>>    # Maybe collapse with max_ptes_shared exceeded.... OK
>>>    # Trigger CoW on page 1024 of 2048... Fail
>>>    Bail out! Unexpected huge page
>>>    # Planned tests != run tests (26 != 23)
>>>    # Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0
>>>
>>>    # Run test: collapse_max_ptes_swap (khugepaged:anon)
>>>    # Swapout 257 of 2048 pages... OK
>>>    # Maybe collapse with max_ptes_swap exceeded.... OK
>>>    # Swapout 256 of 2048 pages... OK
>>>    Bail out! Unexpected huge page
>>>    # Planned tests != run tests (26 != 17)
>>>    # Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0
>>>
>>> This happens because khugepaged may collapse the pages before 
>>> wait_for_scan()
>>> is called, causing a sanity check that expects uncollapsed pages to fail.
>>>
>>> For example, in collapse_max_ptes_swap(), after faulting the pages back in
>>> and paging out up to max_ptes_swap pages, khugepaged may collapse them again
>>> before c->collapse() is called.
>>>
>>> To prevent this, mark the VMA with MADV_NOHUGEPAGE after it has been
>>> collapsed by wait_for_scan() for anon. This prevents khugepaged from
>>> collapsing it again before c->collapse() is called.
>>>
>>> This failure was observed on NVIDIA Spark with 16KB page.
>>>
>>> Reviewed-by: Baolin Wang <[email protected]>
>>> Tested-by: Baolin Wang <[email protected]>
>>> Signed-off-by: Yeoreum Yun <[email protected]>
>>> ---
>>>   tools/testing/selftests/mm/khugepaged.c | 3 +++
>>>   1 file changed, 3 insertions(+)
>>>
>>> diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/
>>> selftests/mm/khugepaged.c
>>> index 2aa7c9197158..b0cb02bf1a73 100644
>>> --- a/tools/testing/selftests/mm/khugepaged.c
>>> +++ b/tools/testing/selftests/mm/khugepaged.c
>>> @@ -618,6 +618,9 @@ static bool wait_for_scan(const char *msg, char *p,
>>> size_t len,
>>>           usleep(TICK);
>>>       }
>>>   +    if (is_anon(ops))
>>> +        madvise(p, len, MADV_NOHUGEPAGE);
>>> +
>>
>> Any reason we just do that unconditionally?
> 
> Although it's a bit messy, as I mentioned before [1], unconditionally setting
> MADV_NOHUGEPAGE will break shmem testing. Maybe add some comments.

Ah, thanks for clarifying. The problem really is that we cannot undo a
MADV_HUGEPAGE (give me hugepages) cleanly. We can only go to the other extreme
(no huge pages).

Yes, let's please add a comment describing why we limit it to anon.

-- 
Cheers,

David

Reply via email to