diff options
| author | Usama Arif <usama.arif@linux.dev> | 2026-06-30 04:23:32 -0700 |
|---|---|---|
| committer | Andrew Morton <akpm@linux-foundation.org> | 2026-07-30 19:40:28 -0700 |
| commit | a33b5c9116554c48adec9e8476c137f49bdb2043 (patch) | |
| tree | 8ce106191306b23fe005863072adda27744b775b /tools/lib/python/kdoc/parse_data_structs.py | |
| parent | bd1e4c4aa469c9e946d9aa2831b8cd4ff3c198f5 (diff) | |
mm/vmpressure: skip tree=true accounting on cgroup v2
Patch series "mm/vmpressure: reduce CPU, memory and code overhead on
cgroup v2", v3.
The vmpressure subsystem has two distinct consumers, gated by the @tree
argument:
tree=false : in-kernel socket pressure, consumed by TCP/SCTP. This
is cgroup v2 only; v1 sockets read memcg->tcpmem_pressure
instead.
tree=true : cgroup v1 userspace eventfd notifications via the
memory.pressure_level / cgroup.event_control interface.
v2 has no equivalent (userspace gets reclaim signals
through memory.pressure / PSI, which doesn't touch
vmpressure).
So of the four (hierarchy, tree) combinations, only two carry data that
anyone reads. The existing early return in vmpressure() covered v1 +
tree=false; the symmetric v2 + tree=true case was falling through and
doing the full lock / accumulate / schedule_work / parent-walk dance, even
though the events list it eventually iterates is empty on cgroup v2
(vmpressure_register_event() is wired up only through the v1 cftype
"memory.pressure_level" and can't be reached from a v2 memcg).
Patch 1 extends the existing early return to also skip v2 + tree=true. On
a v2-only host this eliminates a contended path where reclaimers can
serialize on a single global sr_lock. bpftrace on a 176-core production
host (cgroup v2, 285 memcgs, sustained reclaim) showed ~16,200 such calls
per minute with tree = true.
Patch 2 follows up with a cleanup: it splits the v1 userspace eventfd
interface (struct vmpressure_event, the events list and its mutex, the
work_struct and its handler, the parent walk, vmpressure_register_event /
unregister_event, and vmpressure_prio) into a new mm/memcontrol-v1.c built
only when CONFIG_MEMCG_V1=y, behind small no-op stubs in the header.
mm/vmpressure.c keeps the shared bits and the tree=false socket-pressure
path. The size of vmpressure.c goes down to half and the code is much
more simpler. The only #ifdef CONFIG_MEMCG_V1 remaining in source is
around the v1-only fields inside struct vmpressure itself. Memory savings
on CONFIG_MEMCG_V1=n:
struct vmpressure : 112B -> 24B
struct mem_cgroup : 1664B -> 1536B
This split is the first step toward eventually making vmpressure
CONFIG_MEMCG_V1 only. The v2 in-kernel socket pressure path (tree=false)
cannot be removed today immediately: PSI is not an exact replacement for
vmpressure, and switching networking socket-buffer back-off to PSI may
regress networking performance or increase memory pressure in workloads
that today rely on vmpressure's hysteresis. The medium-term plan is to
introduce a PSI-based socket-pressure path, keep vmpressure available for
v2 behind a defconfig as an opt-out for several releases, and only then
drop the tree=false path entirely, at which point everything that remains
in mm/memcontrol-v1.c is the whole subsystem.
This patch (of 2):
vmpressure() has two outputs gated by the @tree argument:
@tree=false drives in-kernel socket pressure (mem_cgroup_set_
socket_pressure), consumed by TCP/SCTP. This only
applies on cgroup v2; on v1 socket memory is charged
separately via tcpmem and the consumer reads
memcg->tcpmem_pressure instead.
@tree=true drives userspace eventfd notifications via the v1
memory.pressure_level / cgroup.event_control interface.
v2 has no equivalent: userspace gets reclaim signals
through memory.pressure (PSI), which does not touch
vmpressure.
The existing early return covered v1 + @tree=false. The symmetric v2 +
@tree=true case was falling through and doing the full lock / accumulate /
schedule_work / parent-walk dance for an events list that can never be
populated. bpftrace on a 176-core production host (cgroup v2,
CONFIG_MEMCG_V1=n, 285 memcgs, sustained reclaim) showed ~16,200
@tree=true vmpressure() calls per minute. Add an early return that skips
cgroup v2 + tree = true which avoids us doing all this work. On a v2-only
host this also eliminates a lock contention path that can serialise
reclaimers on a single global sr_lock.
[usama.arif@linux.dev: simplify the guard]
Link: https://lore.kernel.org/e8e1a409-48d8-4fa7-ae98-49485a1607f6@linux.dev
Link: https://lore.kernel.org/20260630112617.1198623-1-usama.arif@linux.dev
Link: https://lore.kernel.org/20260630112617.1198623-2-usama.arif@linux.dev
Signed-off-by: Usama Arif <usama.arif@linux.dev>
Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
Acked-by: Johannes Weiner <hannes@cmpxchg.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Michal Koutný <mkoutny@suse.com>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Tejun Heo <tj@kernel.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Diffstat (limited to 'tools/lib/python/kdoc/parse_data_structs.py')
0 files changed, 0 insertions, 0 deletions
