linux.git/include/linux/swap.h, branch v7.1-rc4

mm: workingset: use lruvec_lru_size() to get the number of lru pages

2026-04-18T07:10:47+00:00

For cgroup v2, count_shadow_nodes() is the only place to read
non-hierarchical stats (lruvec_stats->state_local).  To avoid the need to
consider cgroup v2 during subsequent non-hierarchical stats reparenting,
use lruvec_lru_size() instead of lruvec_page_state_local() to get the
number of lru pages.

For NR_SLAB_RECLAIMABLE_B and NR_SLAB_UNRECLAIMABLE_B cases, it appears
that the statistics here have already been problematic for a while since
slab pages have been reparented.  So just ignore it for now.

Link: https://lore.kernel.org/b1d448c667a8fb377c3390d9aba43bdb7e4d5739.1772711148.git.zhengqi.arch@bytedance.com
Signed-off-by: Qi Zheng 
Acked-by: Shakeel Butt 
Acked-by: Muchun Song 
Cc: Allen Pais 
Cc: Axel Rasmussen 
Cc: Baoquan He 
Cc: Chengming Zhou 
Cc: Chen Ridong 
Cc: David Hildenbrand 
Cc: Hamza Mahfooz 
Cc: Harry Yoo 
Cc: Hugh Dickins 
Cc: Imran Khan 
Cc: Johannes Weiner 
Cc: Kamalesh Babulal 
Cc: Lance Yang 
Cc: Liam Howlett 
Cc: Lorenzo Stoakes (Oracle) 
Cc: Michal Hocko 
Cc: Michal Koutný 
Cc: Mike Rapoport 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Suren Baghdasaryan 
Cc: Usama Arif 
Cc: Vlastimil Babka 
Cc: Wei Xu 
Cc: Yosry Ahmed 
Cc: Yuanchu Xie 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: vmscan: prepare for reparenting traditional LRU folios

2026-04-18T07:10:46+00:00

To resolve the dying memcg issue, we need to reparent LRU folios of child
memcg to its parent memcg.  For traditional LRU list, each lruvec of every
memcg comprises four LRU lists.  Due to the symmetry of the LRU lists, it
is feasible to transfer the LRU lists from a memcg to its parent memcg
during the reparenting process.

This commit implements the specific function, which will be used during
the reparenting process.

Link: https://lore.kernel.org/a92d217a9fc82bd0c401210204a095caaf615b1c.1772711148.git.zhengqi.arch@bytedance.com
Signed-off-by: Qi Zheng 
Reviewed-by: Harry Yoo 
Acked-by: Johannes Weiner 
Acked-by: Muchun Song 
Acked-by: Shakeel Butt 
Cc: Allen Pais 
Cc: Axel Rasmussen 
Cc: Baoquan He 
Cc: Chengming Zhou 
Cc: Chen Ridong 
Cc: David Hildenbrand 
Cc: Hamza Mahfooz 
Cc: Hugh Dickins 
Cc: Imran Khan 
Cc: Kamalesh Babulal 
Cc: Lance Yang 
Cc: Liam Howlett 
Cc: Lorenzo Stoakes (Oracle) 
Cc: Michal Hocko 
Cc: Michal Koutný 
Cc: Mike Rapoport 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Suren Baghdasaryan 
Cc: Usama Arif 
Cc: Vlastimil Babka 
Cc: Wei Xu 
Cc: Yosry Ahmed 
Cc: Yuanchu Xie 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: memcontrol: prepare for reparenting LRU pages for lruvec lock

2026-04-18T07:10:46+00:00

The following diagram illustrates how to ensure the safety of the folio
lruvec lock when LRU folios undergo reparenting.

In the folio_lruvec_lock(folio) function:

    rcu_read_lock();
retry:
    lruvec = folio_lruvec(folio);
    /* There is a possibility of folio reparenting at this point. */
    spin_lock(&lruvec->lru_lock);
    if (unlikely(lruvec_memcg(lruvec) != folio_memcg(folio))) {
        /*
         * The wrong lruvec lock was acquired, and a retry is required.
         * This is because the folio resides on the parent memcg lruvec
         * list.
         */
        spin_unlock(&lruvec->lru_lock);
        goto retry;
    }

    /* Reaching here indicates that folio_memcg() is stable. */


In the memcg_reparent_objcgs(memcg) function:

    spin_lock(&lruvec->lru_lock);
    spin_lock(&lruvec_parent->lru_lock);
    /* Transfer folios from the lruvec list to the parent's. */
    spin_unlock(&lruvec_parent->lru_lock);
    spin_unlock(&lruvec->lru_lock);

After acquiring the lruvec lock, it is necessary to verify whether the
folio has been reparented.  If reparenting has occurred, the new lruvec
lock must be reacquired.  During the LRU folio reparenting process, the
lruvec lock will also be acquired (this will be implemented in a
subsequent patch).  Therefore, folio_memcg() remains unchanged while the
lruvec lock is held.

Given that lruvec_memcg(lruvec) is always equal to folio_memcg(folio)
after the lruvec lock is acquired, the lruvec_memcg_debug() check is
redundant.  Hence, it is removed.

This patch serves as a preparation for the reparenting of LRU folios.

Link: https://lore.kernel.org/23f22cbb1419f277a3483018b32158ae2b86c666.1772711148.git.zhengqi.arch@bytedance.com
Signed-off-by: Muchun Song 
Signed-off-by: Qi Zheng 
Acked-by: Johannes Weiner 
Acked-by: Shakeel Butt 
Cc: Allen Pais 
Cc: Axel Rasmussen 
Cc: Baoquan He 
Cc: Chengming Zhou 
Cc: Chen Ridong 
Cc: David Hildenbrand 
Cc: Hamza Mahfooz 
Cc: Harry Yoo 
Cc: Hugh Dickins 
Cc: Imran Khan 
Cc: Kamalesh Babulal 
Cc: Lance Yang 
Cc: Liam Howlett 
Cc: Lorenzo Stoakes (Oracle) 
Cc: Michal Hocko 
Cc: Michal Koutný 
Cc: Mike Rapoport 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Suren Baghdasaryan 
Cc: Usama Arif 
Cc: Vlastimil Babka 
Cc: Wei Xu 
Cc: Yosry Ahmed 
Cc: Yuanchu Xie 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: remove stray references to struct pagevec

2026-04-05T20:53:06+00:00

Patch series "mm: Remove stray references to pagevec", v2.

struct pagevec was removed in commit 1e0877d58b1e ("mm: remove struct
pagevec").  Remove any stray references to it and rename relevant files
and macros accordingly.

While at it, remove unnecessary #includes of pagevec.h (now folio_batch.h)
in .c files.  There are probably more of these that could be removed in .h
files, but those are more complex to verify.


This patch (of 4):

struct pagevec was removed in commit 1e0877d58b1e ("mm: remove struct
pagevec").  Remove remaining forward declarations and change
__folio_batch_release()'s declaration to match its definition.

Link: https://lkml.kernel.org/r/20260225-pagevec_cleanup-v2-0-716868cc2d11@columbia.edu
Link: https://lkml.kernel.org/r/20260225-pagevec_cleanup-v2-1-716868cc2d11@columbia.edu
Signed-off-by: Tal Zussman 
Reviewed-by: Matthew Wilcox (Oracle) 
Acked-by: David Hildenbrand (Arm) 
Acked-by: Chris Li 
Acked-by: Zi Yan 
Reviewed-by: Lorenzo Stoakes (Oracle) 
Cc: Christian Brauner 
Cc: Jan Kara 
Signed-off-by: Andrew Morton

mm, swap: use the swap table to track the swap count

2026-04-05T20:52:59+00:00

Now all the infrastructures are ready, switch to using the swap table
only.  This is unfortunately a large patch because the whole old counting
mechanism, especially SWP_CONTINUED, has to be gone and switch to the new
mechanism together, with no intermediate steps available.

The swap table is capable of holding up to SWP_TB_COUNT_MAX - 1 counts in
the higher bits of each table entry, so using that, the swap_map can be
completely dropped.

swap_map also had a limit of SWAP_CONT_MAX.  Any value beyond that limit
will require a COUNT_CONTINUED page.  COUNT_CONTINUED is a bit complex to
maintain, so for the swap table, a simpler approach is used: when the
count goes beyond SWP_TB_COUNT_MAX - 1, the cluster will have an
extend_table allocated, which is a swap cluster-sized array of unsigned
int.  The counting is basically offloaded there until the count drops
below SWP_TB_COUNT_MAX again.

Both the swap table and the extend table are cluster-based, so they
exhibit good performance and sparsity.

To make the switch from swap_map to swap table clean, this commit cleans
up and introduces a new set of functions based on the swap table design,
for manipulating swap counts:

- __swap_cluster_dup_entry, __swap_cluster_put_entry,
  __swap_cluster_alloc_entry, __swap_cluster_free_entry:

  Increase/decrease the count of a swap slot, or alloc / free a swap
  slot. This is the internal routine that does the counting work based
  on the swap table and handles all the complexities. The caller will
  need to lock the cluster before calling them.

  All swap count-related update operations are wrapped by these four
  helpers.

- swap_dup_entries_cluster, swap_put_entries_cluster:

  Increase/decrease the swap count of one or a set of swap slots in the
  same cluster range. These two helpers serve as the common routines for
  folio_dup_swap & swap_dup_entry_direct, or
  folio_put_swap & swap_put_entries_direct.

And use these helpers to replace all existing callers. This helps to
simplify the count tracking by a lot, and the swap_map is gone.

[ryncsn@gmail.com: fix build]
  Link: https://lkml.kernel.org/r/aZWuLZi-vYi3vAWe@KASONG-MC4
Link: https://lkml.kernel.org/r/20260218-swap-table-p3-v3-9-f4e34be021a7@tencent.com
Signed-off-by: Kairui Song 
Suggested-by: Chris Li 
Acked-by: Chris Li 
Cc: Baoquan He 
Cc: Barry Song 
Cc: David Hildenbrand 
Cc: Johannes Weiner 
Cc: Kairui Song 
Cc: Kemeng Shi 
Cc: kernel test robot 
Cc: Lorenzo Stoakes 
Cc: Nhat Pham 
Signed-off-by: Andrew Morton

mm, swap: drop the SWAP_HAS_CACHE flag

2026-01-31T22:22:57+00:00

Now, the swap cache is managed by the swap table.  All swap cache users
are checking the swap table directly to check the swap cache state. 
SWAP_HAS_CACHE is now just a temporary pin before the first increase from
0 to 1 of a slot's swap count (swap_dup_entries) after swap allocation
(folio_alloc_swap), or before the final free of slots pinned by folio in
swap cache (put_swap_folio).

Drop these two usages.  For the first dup, SWAP_HAS_CACHE pinning was hard
to kill because it used to have multiple meanings, more than just "a slot
is cached".  We have just simplified that and defined that the first dup
is always done with folio locked in swap cache (folio_dup_swap), so stop
checking the SWAP_HAS_CACHE bit and just check the swap cache (swap table)
directly, and add a WARN if a swap entry's count is being increased for
the first time while the folio is not in swap cache.

As for freeing, just let the swap cache free all swap entries of a folio
that have a swap count of zero directly upon folio removal.  We have also
just cleaned up batch freeing to check the swap cache usage using the swap
table: a slot with swap cache in the swap table will not be freed until
its cache is gone, and no SWAP_HAS_CACHE bit is involved anymore.  And
besides, the removal of a folio and freeing of the slots are being done in
the same critical section now, which should improve the performance.

After these two changes, SWAP_HAS_CACHE no longer has any users.  Swap
cache synchronization is also done by the swap table directly, so using
SWAP_HAS_CACHE to pin a slot before adding the cache is also no longer
needed.  Remove all related logic and helpers.  swap_map is now only used
for tracking the count, so all swap_map users can just read it directly,
ignoring the swap_count helper, which was previously used to filter out
the SWAP_HAS_CACHE bit.

The idea of dropping SWAP_HAS_CACHE and using the swap table directly was
initially from Chris's idea of merging all the metadata usage of all swaps
into one place.

Link: https://lkml.kernel.org/r/20251220-swap-table-p2-v5-18-8862a265a033@tencent.com
Signed-off-by: Kairui Song 
Suggested-by: Chris Li 
Reviewed-by: Baoquan He 
Cc: Baolin Wang 
Cc: Barry Song 
Cc: Nhat Pham 
Cc: Rafael J. Wysocki (Intel) 
Cc: Yosry Ahmed 
Cc: Deepanshu Kartikey 
Cc: Johannes Weiner 
Cc: Kairui Song 
Signed-off-by: Andrew Morton

mm, swap: add folio to swap cache directly on allocation

2026-01-31T22:22:57+00:00

The allocator uses SWAP_HAS_CACHE to pin a swap slot upon allocation. 
SWAP_HAS_CACHE is being deprecated as it caused a lot of confusion.  This
pinning usage here can be dropped by adding the folio to swap cache
directly on allocation.

All swap allocations are folio-based now (except for hibernation), so the
swap allocator can always take the folio as the parameter.  And now both
swap cache (swap table) and swap map are protected by the cluster lock,
scanning the map and inserting the folio can be done in the same critical
section.  This eliminates the time window that a slot is pinned by
SWAP_HAS_CACHE, but it has no cache, and avoids touching the lock multiple
times.

This is both a cleanup and an optimization.

Link: https://lkml.kernel.org/r/20251220-swap-table-p2-v5-15-8862a265a033@tencent.com
Signed-off-by: Kairui Song 
Reviewed-by: Baoquan He 
Cc: Baolin Wang 
Cc: Barry Song 
Cc: Chris Li 
Cc: Nhat Pham 
Cc: Rafael J. Wysocki (Intel) 
Cc: Yosry Ahmed 
Cc: Deepanshu Kartikey 
Cc: Johannes Weiner 
Cc: Kairui Song 
Signed-off-by: Andrew Morton

mm, swap: cleanup swap entry management workflow

2026-01-31T22:22:56+00:00

The current swap entry allocation/freeing workflow has never had a clear
definition.  This makes it hard to debug or add new optimizations.

This commit introduces a proper definition of how swap entries would be
allocated and freed.  Now, most operations are folio based, so they will
never exceed one swap cluster, and we now have a cleaner border between
swap and the rest of mm, making it much easier to follow and debug,
especially with new added sanity checks.  Also making more optimization
possible.

Swap entry will be mostly freed and free with a folio bound.  The folio
lock will be useful for resolving many swap related races.

Now swap allocation (except hibernation) always starts with a folio in the
swap cache, and gets duped/freed protected by the folio lock:

- folio_alloc_swap() - The only allocation entry point now.
  Context: The folio must be locked.
  This allocates one or a set of continuous swap slots for a folio and
  binds them to the folio by adding the folio to the swap cache. The
  swap slots' swap count start with zero value.

- folio_dup_swap() - Increase the swap count of one or more entries.
  Context: The folio must be locked and in the swap cache. For now, the
  caller still has to lock the new swap entry owner (e.g., PTL).
  This increases the ref count of swap entries allocated to a folio.
  Newly allocated swap slots' count has to be increased by this helper
  as the folio got unmapped (and swap entries got installed).

- folio_put_swap() - Decrease the swap count of one or more entries.
  Context: The folio must be locked and in the swap cache. For now, the
  caller still has to lock the new swap entry owner (e.g., PTL).
  This decreases the ref count of swap entries allocated to a folio.
  Typically, swapin will decrease the swap count as the folio got
  installed back and the swap entry got uninstalled

  This won't remove the folio from the swap cache and free the
  slot. Lazy freeing of swap cache is helpful for reducing IO.
  There is already a folio_free_swap() for immediate cache reclaim.
  This part could be further optimized later.

The above locking constraints could be further relaxed when the swap table
is fully implemented.  Currently dup still needs the caller to lock the
swap entry container (e.g.  PTL), or a concurrent zap may underflow the
swap count.

Some swap users need to interact with swap count without involving folio
(e.g.  forking/zapping the page table or mapping truncate without swapin).
In such cases, the caller has to ensure there is no race condition on
whatever owns the swap count and use the below helpers:

- swap_put_entries_direct() - Decrease the swap count directly.
  Context: The caller must lock whatever is referencing the slots to
  avoid a race.

  Typically the page table zapping or shmem mapping truncate will need
  to free swap slots directly. If a slot is cached (has a folio bound),
  this will also try to release the swap cache.

- swap_dup_entry_direct() - Increase the swap count directly.
  Context: The caller must lock whatever is referencing the entries to
  avoid race, and the entries must already have a swap count > 1.

  Typically, forking will need to copy the page table and hence needs to
  increase the swap count of the entries in the table. The page table is
  locked while referencing the swap entries, so the entries all have a
  swap count > 1 and can't be freed.

Hibernation subsystem is a bit different, so two special wrappers are here:

- swap_alloc_hibernation_slot() - Allocate one entry from one device.
- swap_free_hibernation_slot() - Free one entry allocated by the above
  helper.

All hibernation entries are exclusive to the hibernation subsystem and
should not interact with ordinary swap routines.

By separating the workflows, it will be possible to bind folio more
tightly with swap cache and get rid of the SWAP_HAS_CACHE as a temporary
pin.

This commit should not introduce any behavior change

[kasong@tencent.com: fix leak, per Chris Mason.  Remove WARN_ON, per Lai Yi]
  Link: https://lkml.kernel.org/r/CAMgjq7AUz10uETVm8ozDWcB3XohkOqf0i33KGrAquvEVvfp5cg@mail.gmail.com
[ryncsn@gmail.com: fix KSM copy pages for swapoff, per Chris]
  Link: https://lkml.kernel.org/r/aXxkANcET3l2Xu6J@KASONG-MC4
Link: https://lkml.kernel.org/r/20251220-swap-table-p2-v5-14-8862a265a033@tencent.com
Signed-off-by: Kairui Song 
Signed-off-by: Kairui Song 
Acked-by: Rafael J. Wysocki (Intel) 
Reviewed-by: Baoquan He 
Cc: Baolin Wang 
Cc: Barry Song 
Cc: Chris Li 
Cc: Nhat Pham 
Cc: Yosry Ahmed 
Cc: Deepanshu Kartikey 
Cc: Johannes Weiner 
Cc: Kairui Song 
Cc: Chris Mason 
Cc: Chris Mason 
Cc: Lai Yi 
Signed-off-by: Andrew Morton

mm, swap: use swap cache as the swap in synchronize layer

2026-01-31T22:22:56+00:00

Current swap in synchronization mostly uses the swap_map's SWAP_HAS_CACHE
bit.  Whoever sets the bit first does the actual work to swap in a folio.

This has been causing many issues as it's just a poor implementation of a
bit lock.  Raced users have no idea what is pinning a slot, so it has to
loop with a schedule_timeout_uninterruptible(1), which is ugly and causes
long-tailing or other performance issues.  Besides, the abuse of
SWAP_HAS_CACHE has been causing many other troubles for synchronization or
maintenance.

This is the first step to remove this bit completely.

Now all swap in paths are using the swap cache, and both the swap cache
and swap map are protected by the cluster lock.  So we can just resolve
the swap synchronization with the swap cache layer directly using the
cluster lock and folio lock.  Whoever inserts a folio in the swap cache
first does the swap in work.  And because folios are locked during swap
operations, other raced swap operations will just wait on the folio lock.

The SWAP_HAS_CACHE will be removed in later commit.  For now, we still set
it for some remaining users.  But now we do the bit setting and swap cache
folio adding in the same critical section, after swap cache is ready.  No
one will have to spin on the SWAP_HAS_CACHE bit anymore.

This both simplifies the logic and should improve the performance,
eliminating issues like the one solved in commit 01626a1823024 ("mm: avoid
unconditional one-tick sleep when swapcache_prepare fails"), or the
"skip_if_exists" from commit a65b0e7607ccb ("zswap: make shrinking
memcg-aware"), which will be removed very soon.

[kasong@tencent.com: fix cgroup v1 accounting issue]
 Link: https://lkml.kernel.org/r/CAMgjq7CGUnzOVG7uSaYjzw9wD7w2dSKOHprJfaEp4CcGLgE3iw@mail.gmail.com
Link: https://lkml.kernel.org/r/20251220-swap-table-p2-v5-12-8862a265a033@tencent.com
Signed-off-by: Kairui Song 
Reviewed-by: Baoquan He 
Cc: Baolin Wang 
Cc: Barry Song 
Cc: Chris Li 
Cc: Nhat Pham 
Cc: Rafael J. Wysocki (Intel) 
Cc: Yosry Ahmed 
Cc: Deepanshu Kartikey 
Cc: Johannes Weiner 
Cc: Kairui Song 
Signed-off-by: Andrew Morton

mm/shmem, swap: remove SWAP_MAP_SHMEM

2026-01-31T22:22:55+00:00

The SWAP_MAP_SHMEM state was introduced in the commit aaa468653b4a
("swap_info: note SWAP_MAP_SHMEM"), to quickly determine if a swap entry
belongs to shmem during swapoff.

However, swapoff has since been rewritten in the commit b56a2d8af914 ("mm:
rid swapoff of quadratic complexity").  Now having swap count ==
SWAP_MAP_SHMEM value is basically the same as having swap count == 1, and
swap_shmem_alloc() behaves analogously to swap_duplicate().  The only
difference of note is that swap_shmem_alloc() does not check for -ENOMEM
returned from __swap_duplicate(), but it is OK because shmem never
re-duplicates any swap entry it owns.  This will stil be safe if we use
(batched) swap_duplicate() instead.

This commit adds swap_duplicate_nr(), the batched variant of
swap_duplicate(), and removes the SWAP_MAP_SHMEM state and the associated
swap_shmem_alloc() helper to simplify the state machine (both mentally and
in terms of actual code).  We will also have an extra state/special value
that can be repurposed (for swap entries that never gets re-duplicated).

Link: https://lkml.kernel.org/r/20251220-swap-table-p2-v5-8-8862a265a033@tencent.com
Signed-off-by: Kairui Song 
Signed-off-by: Nhat Pham 
Reviewed-by: Baolin Wang 
Tested-by: Baolin Wang 
Cc: Baoquan He 
Cc: Barry Song 
Cc: Chris Li 
Cc: Rafael J. Wysocki (Intel) 
Cc: Yosry Ahmed 
Cc: Deepanshu Kartikey 
Cc: Johannes Weiner 
Cc: Kairui Song 
Signed-off-by: Andrew Morton