Hi Kalesh and David! On 9/29/26 1:17 PM, David Hildenbrand (Arm) wrote: > On 9/29/26 09:32, Kalesh Singh wrote: >> On Mon, Sep 28, 2026 at 10:54 PM Sarthak Sharma <[email protected]> >> wrote: >>> >>> >>> >>> On 9/24/26 10:30 AM, Sarthak Sharma wrote: >>>> This series fixes several correctness issues in mremap_test and >>>> simplifies and strengthens its data validation. >>>> >>>> Patch 1 converts the mremap_test to use kselftest helpers, removes >>>> manual tracking of failed tests and corrects some spelling errors. >>>> >>>> Patches 2 to 5 fix userfaultfd skipping, unexpected mremap success >>>> handling, multi-VMA data validation and failure reporting when data >>>> corruption is detected. >>>> >>>> Patch 6 removes the randomization and uses a simple pattern based approach. >>>> >>>> Patch 7 removes perf tests and timing infrastructure. >>>> >>>> Patch 8 removes validation threshold and always validates complete >>>> mappings. >>>> >>>> Patch 9 strengthens the multi VMA validation by also checking the >>>> mapping state of holes after remapping. >>> >>> Hello everyone! Just wanted to check if someone has had a chance to look >>> at the series. >>> >>> Also, Sashiko has a concern [1], and I had the same while posting the >>> series. I hope I can get some opinion from the community. >>> >>> Currently, for PUD remap tests, we allocate a source mapping of 2GB. We >>> only fault in the first threshold_mb amount of memory, remap the whole >>> region, and validate the already faulted threshold_mb amount of memory, >>> which is, by default, equal to 4MB and can be changed by the user by >>> supplying a command line option. >>> >>> Since we plan to remove all command line options from teh selftests, I >>> removed threshold_mb altogether, following some discussion on the list >>> [2]. This would now cause a 2GB source mapping, and its contents copied >>> from another 2GB buffer. This causes the process to have a 4GB RSS. > > Yeah, that's a lot for small CI systems indeed. > >>> Sashiko says that this can cause OOM killing in small CI machines. >>> >>> Would it be okay to keep a 4GB RSS in this case, or should we find some >>> other way of validating a part of the whole range instead? >> >> Hi Sarthak, >> >> IIRC when I initially introduced the test, John was concerned that >> validating the whole range would significantly increase the duration >> of the mm selftests; this is why the threshold was introduced. Please >> check how much it increases if we validate the full range (with >> David's suggestions) and if it's no longer a concern from other folks. >> I am fine with removing the threshold.
I tested on an Orion O6. Here's the difference before and after removing the threshold right now. Before: 0.02user 0.07system 0:00.10elapsed 98%CPU (0avgtext+0avgdata 87332maxresident)k 0inputs+0outputs (0major+21133minor)pagefaults 0swaps After: 1.67user 4.91system 0:06.63elapsed 99%CPU (0avgtext+0avgdata 4195564maxresident)k 0inputs+0outputs (0major+2225058minor)pagefaults 0swaps I think instead of the time taken, the more concerning thing here is RSS (85MB vs 4GB). > > As discussed off-list, I guess it makes more sense to validate a couple of > pages > at the beginning, the middle and the end? Yup, makes sense. I will implement this then.

