linux.git/include/linux/swap.h, branch v7.2-rc1

mm/swap, PM: hibernate: fix swapoff race in uswsusp by pinning swap device

2026-06-09T01:21:31+00:00

Patch series "mm/swap, PM: hibernate: fix swapoff race in uswsusp by
pinning swap device", v8.

Currently, in the uswsusp path, only the swap type value is retrieved at
lookup time without holding a reference. If swapoff races after the type
is acquired, subsequent slot allocations operate on a stale swap device.

Additionally, grabbing and releasing the swap device reference on every
slot allocation is inefficient across the entire hibernation swap path.

This patch series addresses these issues:
- Patch 1: Fixes the swapoff race in uswsusp by pinning the swap device
  from the point it is looked up until the session completes.
- Patch 2: Removes the overhead of per-slot reference counting in alloc/free
  paths and cleans up the redundant SWP_WRITEOK check.


This patch (of 2):

Hibernation via uswsusp (/dev/snapshot ioctls) has a race window: after
selecting the resume swap area but before user space is frozen, swapoff
may run and invalidate the selected swap device.

Fix this by pinning the swap device with SWP_HIBERNATION while it is in
use.  The pin is exclusive, which is sufficient since hibernate_acquire()
already prevents concurrent hibernation sessions.

The kernel swsusp path (sysfs-based hibernate/resume) uses
find_hibernation_swap_type() which is not affected by the pin.  It freezes
user space before touching swap, so swapoff cannot race.

Introduce dedicated helpers:
- pin_hibernation_swap_type(): Look up and pin the swap device.
  Used by the uswsusp path.
- find_hibernation_swap_type(): Lookup without pinning.
  Used by the kernel swsusp path.
- unpin_hibernation_swap_type(): Clear the hibernation pin.

While a swap device is pinned, swapoff is prevented from proceeding.

Link: https://lore.kernel.org/20260323160822.1409904-1-youngjun.park@lge.com
Link: https://lore.kernel.org/20260323160822.1409904-2-youngjun.park@lge.com
Signed-off-by: Youngjun Park 
Reviewed-by: Kairui Song 
Cc: Baoquan He 
Cc: Barry Song 
Cc: Chris Li 
Cc: Kemeng Shi 
Cc: Nhat Pham 
Cc: "Rafael J . Wysocki" 
Signed-off-by: Andrew Morton

mm, swap: merge zeromap into swap table

2026-06-02T22:22:23+00:00

By allocating one additional bit in the swap table entry's flags field
alongside the count, we can store the zeromap inline

For 64 bit systems, zeromap will store in the swap table, avoiding zeromap
allocation.  It reduces the allocated memory.  That is the happy path.

For certain 32-bit archs, there might not be enough bits in the swap table
to contain both PFN and flags.  Therefore, conditionally let each cluster
have a zeromap field at build time, and use that instead.  If the swapfile
cluster is not fully used, it will still save memory for zeromap.  The
empty cluster does not allocate a zeromap.  In the worst case, all cluster
are fully populated.  We will use memory similar to the previous zeromap
implementation.

A few macros were moved to different headers for build time struct
definition.

[akpm@linux-foundation.org: swap_cluster_alloc_table(): remove unused local `ret]
[akpm@linux-foundation.org: fix unused label `err_free']
Link: https://lore.kernel.org/20260517-swap-table-p4-v5-12-88ae43e064c7@tencent.com
Signed-off-by: Kairui Song 
Acked-by: Chris Li 
Reviewed-by: Youngjun Park 
Cc: Baolin Wang 
Cc: Baoquan He 
Cc: Barry Song 
Cc: Chengming Zhou 
Cc: David Hildenbrand 
Cc: Hugh Dickins 
Cc: Johannes Weiner 
Cc: Kemeng Shi 
Cc: Lorenzo Stoakes 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm/memcg, swap: store cgroup id in cluster table directly

2026-06-02T22:22:23+00:00

Drop the usage of the swap_cgroup_ctrl, and use the dynamic cluster table
instead.

The per-cluster memcg table is 1024 / 512 bytes on most archs, and does
not need RCU protection: the cgroup data is only read and written under
the cluster lock.  That keeps things simple, lets the allocation use plain
kmalloc with immediate kfree (no deferred free), and keeps fragmentation
acceptable.

[akpm@linux-foundation.org: memcgv1: don't compile swap functions when CONFIG_SWAP=n]
  Link: https://lore.kernel.org/202605281711.bSeZlErK-lkp@intel.com
[akpm@linux-foundation.org: fix CONFIG_SWAP=n build]
Link: https://lore.kernel.org/20260517-swap-table-p4-v5-10-88ae43e064c7@tencent.com
Signed-off-by: Kairui Song 
Acked-by: Chris Li 
Cc: Baolin Wang 
Cc: Baoquan He 
Cc: Barry Song 
Cc: Chengming Zhou 
Cc: David Hildenbrand 
Cc: Hugh Dickins 
Cc: Johannes Weiner 
Cc: Kemeng Shi 
Cc: Lorenzo Stoakes 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Youngjun Park 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm/memcg, swap: tidy up cgroup v1 memsw swap helpers

2026-06-02T22:22:22+00:00

The cgroup v1 swap helpers always operate on swap cache folios whose swap
entry is stable: the folio is locked and in the swap cache.  There is no
need to pass the swap entry or page count as separate parameters when they
can be derived from the folio itself.

Simplify the redundant parameters and add sanity checks to document the
required preconditions.

Also rename memcg1_swapout to __memcg1_swapout to indicate it requires
special calling context: the folio must be isolated and dying, and the
call must be made with interrupts disabled.

No functional change.

Link: https://lore.kernel.org/20260517-swap-table-p4-v5-6-88ae43e064c7@tencent.com
Signed-off-by: Kairui Song 
Acked-by: Chris Li 
Cc: Baolin Wang 
Cc: Baoquan He 
Cc: Barry Song 
Cc: Chengming Zhou 
Cc: David Hildenbrand 
Cc: Hugh Dickins 
Cc: Johannes Weiner 
Cc: Kemeng Shi 
Cc: Lorenzo Stoakes 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Shakeel Butt 
Cc: Youngjun Park 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: workingset: use lruvec_lru_size() to get the number of lru pages

2026-04-18T07:10:47+00:00

For cgroup v2, count_shadow_nodes() is the only place to read
non-hierarchical stats (lruvec_stats->state_local).  To avoid the need to
consider cgroup v2 during subsequent non-hierarchical stats reparenting,
use lruvec_lru_size() instead of lruvec_page_state_local() to get the
number of lru pages.

For NR_SLAB_RECLAIMABLE_B and NR_SLAB_UNRECLAIMABLE_B cases, it appears
that the statistics here have already been problematic for a while since
slab pages have been reparented.  So just ignore it for now.

Link: https://lore.kernel.org/b1d448c667a8fb377c3390d9aba43bdb7e4d5739.1772711148.git.zhengqi.arch@bytedance.com
Signed-off-by: Qi Zheng 
Acked-by: Shakeel Butt 
Acked-by: Muchun Song 
Cc: Allen Pais 
Cc: Axel Rasmussen 
Cc: Baoquan He 
Cc: Chengming Zhou 
Cc: Chen Ridong 
Cc: David Hildenbrand 
Cc: Hamza Mahfooz 
Cc: Harry Yoo 
Cc: Hugh Dickins 
Cc: Imran Khan 
Cc: Johannes Weiner 
Cc: Kamalesh Babulal 
Cc: Lance Yang 
Cc: Liam Howlett 
Cc: Lorenzo Stoakes (Oracle) 
Cc: Michal Hocko 
Cc: Michal Koutný 
Cc: Mike Rapoport 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Suren Baghdasaryan 
Cc: Usama Arif 
Cc: Vlastimil Babka 
Cc: Wei Xu 
Cc: Yosry Ahmed 
Cc: Yuanchu Xie 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: vmscan: prepare for reparenting traditional LRU folios

2026-04-18T07:10:46+00:00

To resolve the dying memcg issue, we need to reparent LRU folios of child
memcg to its parent memcg.  For traditional LRU list, each lruvec of every
memcg comprises four LRU lists.  Due to the symmetry of the LRU lists, it
is feasible to transfer the LRU lists from a memcg to its parent memcg
during the reparenting process.

This commit implements the specific function, which will be used during
the reparenting process.

Link: https://lore.kernel.org/a92d217a9fc82bd0c401210204a095caaf615b1c.1772711148.git.zhengqi.arch@bytedance.com
Signed-off-by: Qi Zheng 
Reviewed-by: Harry Yoo 
Acked-by: Johannes Weiner 
Acked-by: Muchun Song 
Acked-by: Shakeel Butt 
Cc: Allen Pais 
Cc: Axel Rasmussen 
Cc: Baoquan He 
Cc: Chengming Zhou 
Cc: Chen Ridong 
Cc: David Hildenbrand 
Cc: Hamza Mahfooz 
Cc: Hugh Dickins 
Cc: Imran Khan 
Cc: Kamalesh Babulal 
Cc: Lance Yang 
Cc: Liam Howlett 
Cc: Lorenzo Stoakes (Oracle) 
Cc: Michal Hocko 
Cc: Michal Koutný 
Cc: Mike Rapoport 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Suren Baghdasaryan 
Cc: Usama Arif 
Cc: Vlastimil Babka 
Cc: Wei Xu 
Cc: Yosry Ahmed 
Cc: Yuanchu Xie 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: memcontrol: prepare for reparenting LRU pages for lruvec lock

2026-04-18T07:10:46+00:00

The following diagram illustrates how to ensure the safety of the folio
lruvec lock when LRU folios undergo reparenting.

In the folio_lruvec_lock(folio) function:

    rcu_read_lock();
retry:
    lruvec = folio_lruvec(folio);
    /* There is a possibility of folio reparenting at this point. */
    spin_lock(&lruvec->lru_lock);
    if (unlikely(lruvec_memcg(lruvec) != folio_memcg(folio))) {
        /*
         * The wrong lruvec lock was acquired, and a retry is required.
         * This is because the folio resides on the parent memcg lruvec
         * list.
         */
        spin_unlock(&lruvec->lru_lock);
        goto retry;
    }

    /* Reaching here indicates that folio_memcg() is stable. */


In the memcg_reparent_objcgs(memcg) function:

    spin_lock(&lruvec->lru_lock);
    spin_lock(&lruvec_parent->lru_lock);
    /* Transfer folios from the lruvec list to the parent's. */
    spin_unlock(&lruvec_parent->lru_lock);
    spin_unlock(&lruvec->lru_lock);

After acquiring the lruvec lock, it is necessary to verify whether the
folio has been reparented.  If reparenting has occurred, the new lruvec
lock must be reacquired.  During the LRU folio reparenting process, the
lruvec lock will also be acquired (this will be implemented in a
subsequent patch).  Therefore, folio_memcg() remains unchanged while the
lruvec lock is held.

Given that lruvec_memcg(lruvec) is always equal to folio_memcg(folio)
after the lruvec lock is acquired, the lruvec_memcg_debug() check is
redundant.  Hence, it is removed.

This patch serves as a preparation for the reparenting of LRU folios.

Link: https://lore.kernel.org/23f22cbb1419f277a3483018b32158ae2b86c666.1772711148.git.zhengqi.arch@bytedance.com
Signed-off-by: Muchun Song 
Signed-off-by: Qi Zheng 
Acked-by: Johannes Weiner 
Acked-by: Shakeel Butt 
Cc: Allen Pais 
Cc: Axel Rasmussen 
Cc: Baoquan He 
Cc: Chengming Zhou 
Cc: Chen Ridong 
Cc: David Hildenbrand 
Cc: Hamza Mahfooz 
Cc: Harry Yoo 
Cc: Hugh Dickins 
Cc: Imran Khan 
Cc: Kamalesh Babulal 
Cc: Lance Yang 
Cc: Liam Howlett 
Cc: Lorenzo Stoakes (Oracle) 
Cc: Michal Hocko 
Cc: Michal Koutný 
Cc: Mike Rapoport 
Cc: Muchun Song 
Cc: Nhat Pham 
Cc: Roman Gushchin 
Cc: Suren Baghdasaryan 
Cc: Usama Arif 
Cc: Vlastimil Babka 
Cc: Wei Xu 
Cc: Yosry Ahmed 
Cc: Yuanchu Xie 
Cc: Zi Yan 
Signed-off-by: Andrew Morton

mm: remove stray references to struct pagevec

2026-04-05T20:53:06+00:00

Patch series "mm: Remove stray references to pagevec", v2.

struct pagevec was removed in commit 1e0877d58b1e ("mm: remove struct
pagevec").  Remove any stray references to it and rename relevant files
and macros accordingly.

While at it, remove unnecessary #includes of pagevec.h (now folio_batch.h)
in .c files.  There are probably more of these that could be removed in .h
files, but those are more complex to verify.


This patch (of 4):

struct pagevec was removed in commit 1e0877d58b1e ("mm: remove struct
pagevec").  Remove remaining forward declarations and change
__folio_batch_release()'s declaration to match its definition.

Link: https://lkml.kernel.org/r/20260225-pagevec_cleanup-v2-0-716868cc2d11@columbia.edu
Link: https://lkml.kernel.org/r/20260225-pagevec_cleanup-v2-1-716868cc2d11@columbia.edu
Signed-off-by: Tal Zussman 
Reviewed-by: Matthew Wilcox (Oracle) 
Acked-by: David Hildenbrand (Arm) 
Acked-by: Chris Li 
Acked-by: Zi Yan 
Reviewed-by: Lorenzo Stoakes (Oracle) 
Cc: Christian Brauner 
Cc: Jan Kara 
Signed-off-by: Andrew Morton

mm, swap: use the swap table to track the swap count

2026-04-05T20:52:59+00:00

Now all the infrastructures are ready, switch to using the swap table
only.  This is unfortunately a large patch because the whole old counting
mechanism, especially SWP_CONTINUED, has to be gone and switch to the new
mechanism together, with no intermediate steps available.

The swap table is capable of holding up to SWP_TB_COUNT_MAX - 1 counts in
the higher bits of each table entry, so using that, the swap_map can be
completely dropped.

swap_map also had a limit of SWAP_CONT_MAX.  Any value beyond that limit
will require a COUNT_CONTINUED page.  COUNT_CONTINUED is a bit complex to
maintain, so for the swap table, a simpler approach is used: when the
count goes beyond SWP_TB_COUNT_MAX - 1, the cluster will have an
extend_table allocated, which is a swap cluster-sized array of unsigned
int.  The counting is basically offloaded there until the count drops
below SWP_TB_COUNT_MAX again.

Both the swap table and the extend table are cluster-based, so they
exhibit good performance and sparsity.

To make the switch from swap_map to swap table clean, this commit cleans
up and introduces a new set of functions based on the swap table design,
for manipulating swap counts:

- __swap_cluster_dup_entry, __swap_cluster_put_entry,
  __swap_cluster_alloc_entry, __swap_cluster_free_entry:

  Increase/decrease the count of a swap slot, or alloc / free a swap
  slot. This is the internal routine that does the counting work based
  on the swap table and handles all the complexities. The caller will
  need to lock the cluster before calling them.

  All swap count-related update operations are wrapped by these four
  helpers.

- swap_dup_entries_cluster, swap_put_entries_cluster:

  Increase/decrease the swap count of one or a set of swap slots in the
  same cluster range. These two helpers serve as the common routines for
  folio_dup_swap & swap_dup_entry_direct, or
  folio_put_swap & swap_put_entries_direct.

And use these helpers to replace all existing callers. This helps to
simplify the count tracking by a lot, and the swap_map is gone.

[ryncsn@gmail.com: fix build]
  Link: https://lkml.kernel.org/r/aZWuLZi-vYi3vAWe@KASONG-MC4
Link: https://lkml.kernel.org/r/20260218-swap-table-p3-v3-9-f4e34be021a7@tencent.com
Signed-off-by: Kairui Song 
Suggested-by: Chris Li 
Acked-by: Chris Li 
Cc: Baoquan He 
Cc: Barry Song 
Cc: David Hildenbrand 
Cc: Johannes Weiner 
Cc: Kairui Song 
Cc: Kemeng Shi 
Cc: kernel test robot 
Cc: Lorenzo Stoakes 
Cc: Nhat Pham 
Signed-off-by: Andrew Morton

mm, swap: drop the SWAP_HAS_CACHE flag

2026-01-31T22:22:57+00:00

Now, the swap cache is managed by the swap table.  All swap cache users
are checking the swap table directly to check the swap cache state. 
SWAP_HAS_CACHE is now just a temporary pin before the first increase from
0 to 1 of a slot's swap count (swap_dup_entries) after swap allocation
(folio_alloc_swap), or before the final free of slots pinned by folio in
swap cache (put_swap_folio).

Drop these two usages.  For the first dup, SWAP_HAS_CACHE pinning was hard
to kill because it used to have multiple meanings, more than just "a slot
is cached".  We have just simplified that and defined that the first dup
is always done with folio locked in swap cache (folio_dup_swap), so stop
checking the SWAP_HAS_CACHE bit and just check the swap cache (swap table)
directly, and add a WARN if a swap entry's count is being increased for
the first time while the folio is not in swap cache.

As for freeing, just let the swap cache free all swap entries of a folio
that have a swap count of zero directly upon folio removal.  We have also
just cleaned up batch freeing to check the swap cache usage using the swap
table: a slot with swap cache in the swap table will not be freed until
its cache is gone, and no SWAP_HAS_CACHE bit is involved anymore.  And
besides, the removal of a folio and freeing of the slots are being done in
the same critical section now, which should improve the performance.

After these two changes, SWAP_HAS_CACHE no longer has any users.  Swap
cache synchronization is also done by the swap table directly, so using
SWAP_HAS_CACHE to pin a slot before adding the cache is also no longer
needed.  Remove all related logic and helpers.  swap_map is now only used
for tracking the count, so all swap_map users can just read it directly,
ignoring the swap_count helper, which was previously used to filter out
the SWAP_HAS_CACHE bit.

The idea of dropping SWAP_HAS_CACHE and using the swap table directly was
initially from Chris's idea of merging all the metadata usage of all swaps
into one place.

Link: https://lkml.kernel.org/r/20251220-swap-table-p2-v5-18-8862a265a033@tencent.com
Signed-off-by: Kairui Song 
Suggested-by: Chris Li 
Reviewed-by: Baoquan He 
Cc: Baolin Wang 
Cc: Barry Song 
Cc: Nhat Pham 
Cc: Rafael J. Wysocki (Intel) 
Cc: Yosry Ahmed 
Cc: Deepanshu Kartikey 
Cc: Johannes Weiner 
Cc: Kairui Song 
Signed-off-by: Andrew Morton