linux.git/mm/mlock.c, branch v6.9

mm: make folios_put() the basis of release_pages()

2024-03-05T01:01:22+00:00

Patch series "Rearrange batched folio freeing", v3.

Other than the obvious "remove calls to compound_head" changes, the
fundamental belief here is that iterating a linked list is much slower
than iterating an array (5-15x slower in my testing).  There's also an
associated belief that since we iterate the batch of folios three times,
we do better when the array is small (ie 15 entries) than we do with a
batch that is hundreds of entries long, which only gives us the
opportunity for the first pages to fall out of cache by the time we get to
the end.

It is possible we should increase the size of folio_batch.  Hopefully the
bots let us know if this introduces any performance regressions.


This patch (of 3):

By making release_pages() call folios_put(), we can get rid of the calls
to compound_head() for the callers that already know they have folios.  We
can also get rid of the lock_batch tracking as we know the size of the
batch is limited by folio_batch.  This does reduce the maximum number of
pages for which the lruvec lock is held, from SWAP_CLUSTER_MAX (32) to
PAGEVEC_SIZE (15).  I do not expect this to make a significant difference,
but if it does, we can increase PAGEVEC_SIZE to 31.

Link: https://lkml.kernel.org/r/20240227174254.710559-1-willy@infradead.org
Link: https://lkml.kernel.org/r/20240227174254.710559-2-willy@infradead.org
Signed-off-by: Matthew Wilcox (Oracle) 
Cc: David Hildenbrand 
Cc: Mel Gorman 
Cc: Ryan Roberts 
Signed-off-by: Andrew Morton

mm: mlock: avoid folio_within_range() on KSM pages

2023-10-25T23:47:14+00:00

Since commit dc68badcede4 ("mm: mlock: update mlock_pte_range to handle
large folio") I've just occasionally seen VM_WARN_ON_FOLIO(folio_test_ksm)
warnings from folio_within_range(), in a splurge after testing with KSM
hyperactive.

folio_referenced_one()'s use of folio_within_vma() is safe because it
checks folio_test_large() first; but allow_mlock_munlock() needs to do the
same to avoid those warnings (or check !folio_test_ksm() itself?  Or move
either check into folio_within_range()?  Hard to tell without more
examples of its use).

Link: https://lkml.kernel.org/r/23852f6a-5bfa-1ffd-30db-30c5560ad426@google.com
Fixes: dc68badcede4 ("mm: mlock: update mlock_pte_range to handle large folio")
Signed-off-by: Hugh Dickins 
Reviewed-by: Yin Fengwei 
Cc: Lorenzo Stoakes 
Cc: Matthew Wilcox (Oracle) 
Cc: Stefan Roesch 
Signed-off-by: Andrew Morton

mm: abstract the vma_merge()/split_vma() pattern for mprotect() et al.

2023-10-18T21:34:18+00:00

mprotect() and other functions which change VMA parameters over a range
each employ a pattern of:-

1. Attempt to merge the range with adjacent VMAs.
2. If this fails, and the range spans a subset of the VMA, split it
   accordingly.

This is open-coded and duplicated in each case. Also in each case most of
the parameters passed to vma_merge() remain the same.

Create a new function, vma_modify(), which abstracts this operation,
accepting only those parameters which can be changed.

To avoid the mess of invoking each function call with unnecessary
parameters, create inline wrapper functions for each of the modify
operations, parameterised only by what is required to perform the action.

We can also significantly simplify the logic - by returning the VMA if we
split (or merged VMA if we do not) we no longer need specific handling for
merge/split cases in any of the call sites.

Note that the userfaultfd_release() case works even though it does not
split VMAs - since start is set to vma->vm_start and end is set to
vma->vm_end, the split logic does not trigger.

In addition, since we calculate pgoff to be equal to vma->vm_pgoff + (start
- vma->vm_start) >> PAGE_SHIFT, and start - vma->vm_start will be 0 in this
instance, this invocation will remain unchanged.

We eliminate a VM_WARN_ON() in mprotect_fixup() as this simply asserts that
vma_merge() correctly ensures that flags remain the same, something that is
already checked in is_mergeable_vma() and elsewhere, and in any case is not
specific to mprotect().

Link: https://lkml.kernel.org/r/0dfa9368f37199a423674bf0ee312e8ea0619044.1697043508.git.lstoakes@gmail.com
Signed-off-by: Lorenzo Stoakes 
Reviewed-by: Vlastimil Babka 
Cc: Alexander Viro 
Cc: Christian Brauner 
Cc: Liam R. Howlett 
Signed-off-by: Andrew Morton

mm: mlock: update mlock_pte_range to handle large folio

2023-10-04T17:32:32+00:00

Current kernel only lock base size folio during mlock syscall.
Add large folio support with following rules:
  - Only mlock large folio when it's in VM_LOCKED VMA range
    and fully mapped to page table.

    fully mapped folio is required as if folio is not fully
    mapped to a VM_LOCKED VMA, if system is in memory pressure,
    page reclaim is allowed to pick up this folio, split it
    and reclaim the pages which are not in VM_LOCKED VMA.

  - munlock will apply to the large folio which is in VMA range
    or cross the VMA boundary.

    This is required to handle the case that the large folio is
    mlocked, later the VMA is split in the middle of large folio.

Link: https://lkml.kernel.org/r/20230918073318.1181104-4-fengwei.yin@intel.com
Signed-off-by: Yin Fengwei 
Cc: David Hildenbrand 
Cc: Hugh Dickins 
Cc: Matthew Wilcox (Oracle) 
Cc: Ryan Roberts 
Cc: Yang Shi 
Cc: Yosry Ahmed 
Cc: Yu Zhao 
Signed-off-by: Andrew Morton

merge mm-hotfixes-stable into mm-stable to pick up depended-upon changes

2023-08-21T21:26:20+00:00

mm: lock vma explicitly before doing vm_flags_reset and vm_flags_reset_once

2023-08-21T20:37:46+00:00

Implicit vma locking inside vm_flags_reset() and vm_flags_reset_once() is
not obvious and makes it hard to understand where vma locking is happening.
Also in some cases (like in dup_userfaultfd()) vma should be locked earlier
than vma_flags modification. To make locking more visible, change these
functions to assert that the vma write lock is taken and explicitly lock
the vma beforehand. Fix userfaultfd functions which should lock the vma
earlier.

Link: https://lkml.kernel.org/r/20230804152724.3090321-5-surenb@google.com
Suggested-by: Linus Torvalds 
Signed-off-by: Suren Baghdasaryan 
Cc: Jann Horn 
Cc: Liam R. Howlett 
Signed-off-by: Andrew Morton

mm: enable page walking API to lock vmas during the walk

2023-08-21T20:07:20+00:00

walk_page_range() and friends often operate under write-locked mmap_lock. 
With introduction of vma locks, the vmas have to be locked as well during
such walks to prevent concurrent page faults in these areas.  Add an
additional member to mm_walk_ops to indicate locking requirements for the
walk.

The change ensures that page walks which prevent concurrent page faults
by write-locking mmap_lock, operate correctly after introduction of
per-vma locks.  With per-vma locks page faults can be handled under vma
lock without taking mmap_lock at all, so write locking mmap_lock would
not stop them.  The change ensures vmas are properly locked during such
walks.

A sample issue this solves is do_mbind() performing queue_pages_range()
to queue pages for migration.  Without this change a concurrent page
can be faulted into the area and be left out of migration.

Link: https://lkml.kernel.org/r/20230804152724.3090321-2-surenb@google.com
Signed-off-by: Suren Baghdasaryan 
Suggested-by: Linus Torvalds 
Suggested-by: Jann Horn 
Cc: David Hildenbrand 
Cc: Davidlohr Bueso 
Cc: Hugh Dickins 
Cc: Johannes Weiner 
Cc: Laurent Dufour 
Cc: Liam Howlett 
Cc: Matthew Wilcox (Oracle) 
Cc: Michal Hocko 
Cc: Michel Lespinasse 
Cc: Peter Xu 
Cc: Vlastimil Babka 
Cc: 
Signed-off-by: Andrew Morton

mm/mlock: fix vma iterator conversion of apply_vma_lock_flags()

2023-07-17T19:53:21+00:00

apply_vma_lock_flags() calls mlock_fixup(), which could merge the VMA
after where the vma iterator is located.  Although this is not an issue,
the next iteration of the loop will check the start of the vma to be equal
to the locally saved 'tmp' variable and cause an incorrect failure
scenario.  Fix the error by setting tmp to the end of the vma iterator
value before restarting the loop.

There is also a potential of the error code being overwritten when the
loop terminates early.  Fix the return issue by directly returning when an
error is encountered since there is nothing to undo after the loop.

Link: https://lkml.kernel.org/r/20230711175020.4091336-1-Liam.Howlett@oracle.com
Fixes: 37598f5a9d8b ("mlock: convert mlock to vma iterator")
Signed-off-by: Liam R. Howlett 
Reported-by: Ryan Roberts 
  Link: https://lore.kernel.org/linux-mm/50341ca1-d582-b33a-e3d0-acb08a65166f@arm.com/
Tested-by: Ryan Roberts 
Cc: 
Signed-off-by: Andrew Morton

mm: ptep_get() conversion

2023-06-19T23:19:25+00:00

Convert all instances of direct pte_t* dereferencing to instead use
ptep_get() helper.  This means that by default, the accesses change from a
C dereference to a READ_ONCE().  This is technically the correct thing to
do since where pgtables are modified by HW (for access/dirty) they are
volatile and therefore we should always ensure READ_ONCE() semantics.

But more importantly, by always using the helper, it can be overridden by
the architecture to fully encapsulate the contents of the pte.  Arch code
is deliberately not converted, as the arch code knows best.  It is
intended that arch code (arm64) will override the default with its own
implementation that can (e.g.) hide certain bits from the core code, or
determine young/dirty status by mixing in state from another source.

Conversion was done using Coccinelle:

----

// $ make coccicheck \
//          COCCI=ptepget.cocci \
//          SPFLAGS="--include-headers" \
//          MODE=patch

virtual patch

@ depends on patch @
pte_t *v;
@@

- *v
+ ptep_get(v)

----

Then reviewed and hand-edited to avoid multiple unnecessary calls to
ptep_get(), instead opting to store the result of a single call in a
variable, where it is correct to do so.  This aims to negate any cost of
READ_ONCE() and will benefit arch-overrides that may be more complex.

Included is a fix for an issue in an earlier version of this patch that
was pointed out by kernel test robot.  The issue arose because config
MMU=n elides definition of the ptep helper functions, including
ptep_get().  HUGETLB_PAGE=n configs still define a simple
huge_ptep_clear_flush() for linking purposes, which dereferences the ptep.
So when both configs are disabled, this caused a build error because
ptep_get() is not defined.  Fix by continuing to do a direct dereference
when MMU=n.  This is safe because for this config the arch code cannot be
trying to virtualize the ptes because none of the ptep helpers are
defined.

Link: https://lkml.kernel.org/r/20230612151545.3317766-4-ryan.roberts@arm.com
Reported-by: kernel test robot 
Link: https://lore.kernel.org/oe-kbuild-all/202305120142.yXsNEo6H-lkp@intel.com/
Signed-off-by: Ryan Roberts 
Cc: Adrian Hunter 
Cc: Alexander Potapenko 
Cc: Alexander Shishkin 
Cc: Alex Williamson 
Cc: Al Viro 
Cc: Andrey Konovalov 
Cc: Andrey Ryabinin 
Cc: Christian Brauner 
Cc: Christoph Hellwig 
Cc: Daniel Vetter 
Cc: Dave Airlie 
Cc: Dimitri Sivanich 
Cc: Dmitry Vyukov 
Cc: Ian Rogers 
Cc: Jason Gunthorpe 
Cc: Jérôme Glisse 
Cc: Jiri Olsa 
Cc: Johannes Weiner 
Cc: Kirill A. Shutemov 
Cc: Lorenzo Stoakes 
Cc: Mark Rutland 
Cc: Matthew Wilcox 
Cc: Miaohe Lin 
Cc: Michal Hocko 
Cc: Mike Kravetz 
Cc: Mike Rapoport (IBM) 
Cc: Muchun Song 
Cc: Namhyung Kim 
Cc: Naoya Horiguchi 
Cc: Oleksandr Tyshchenko 
Cc: Pavel Tatashin 
Cc: Roman Gushchin 
Cc: SeongJae Park 
Cc: Shakeel Butt 
Cc: Uladzislau Rezki (Sony) 
Cc: Vincenzo Frascino 
Cc: Yu Zhao 
Signed-off-by: Andrew Morton

mm/pagewalkers: ACTION_AGAIN if pte_offset_map_lock() fails

2023-06-19T23:19:13+00:00

Simple walk_page_range() users should set ACTION_AGAIN to retry when
pte_offset_map_lock() fails.

No need to check pmd_trans_unstable(): that was precisely to avoid the
possiblity of calling pte_offset_map() on a racily removed or inserted THP
entry, but such cases are now safely handled inside it.  Likewise there is
no need to check pmd_none() or pmd_bad() before calling it.

Link: https://lkml.kernel.org/r/c77d9d10-3aad-e3ce-4896-99e91c7947f3@google.com
Signed-off-by: Hugh Dickins 
Reviewed-by: SeongJae Park  for mm/damon part
Cc: Alistair Popple 
Cc: Anshuman Khandual 
Cc: Axel Rasmussen 
Cc: Christophe Leroy 
Cc: Christoph Hellwig 
Cc: David Hildenbrand 
Cc: "Huang, Ying" 
Cc: Ira Weiny 
Cc: Jason Gunthorpe 
Cc: Kirill A. Shutemov 
Cc: Lorenzo Stoakes 
Cc: Matthew Wilcox 
Cc: Mel Gorman 
Cc: Miaohe Lin 
Cc: Mike Kravetz 
Cc: Mike Rapoport (IBM) 
Cc: Minchan Kim 
Cc: Naoya Horiguchi 
Cc: Pavel Tatashin 
Cc: Peter Xu 
Cc: Peter Zijlstra 
Cc: Qi Zheng 
Cc: Ralph Campbell 
Cc: Ryan Roberts 
Cc: Song Liu 
Cc: Steven Price 
Cc: Suren Baghdasaryan 
Cc: Thomas Hellström 
Cc: Will Deacon 
Cc: Yang Shi 
Cc: Yu Zhao 
Cc: Zack Rusin 
Signed-off-by: Andrew Morton