diff options
| author | Leo Martins <loemra.dev@gmail.com> | 2026-07-01 16:47:10 -0700 |
|---|---|---|
| committer | David Sterba <dsterba@suse.com> | 2026-08-07 19:17:17 +0200 |
| commit | 67bd829a14b7d71ffea28147b98dadf9ee684568 (patch) | |
| tree | eef88ac90ac38e1ba81378a493126514cc5c9475 /tools/perf/scripts/python/stackcollapse.py | |
| parent | f6c02cd048bd7db865a2e2d65cae78aea64444fa (diff) | |
btrfs: replace writeback inhibition xarray with a fixed inline buffer
Commit f9a48549a15a ("btrfs: inhibit extent buffer writeback to prevent
COW amplification") tracks the extent buffers a transaction handle has
inhibited in a per-handle xarray. Keying the tracking to the transaction
handle is correct, but using an xarray for it causes two problems in
production.
First, a write_iops regression. Every COW calls
btrfs_inhibit_eb_writeback() from btrfs_force_cow_block() and
should_cow_block(), which does an xa_store() keyed by eb->start. The
kernel test robot reported a 22.6% fio.write_iops regression on a
single-task 4k randwrite workload (ftruncate ioengine, buffered IO) on
btrfs. The cost is the per-COW xarray store done on every COW'd block.
Replacing it with a non-allocating fixed buffer recovers the lost
throughput, and that buffer does more per-COW bookkeeping yet still
recovers, so the cost is the xarray operation itself rather than the
extra tracking work.
Second, an unbounded cleanup walk. btrfs_uninhibit_all_eb_writeback()
iterates every eb the handle inhibited with xa_for_each(). A single
handle that COWs a very large number of blocks (inode eviction, or
truncate of a file with many extents, where btrfs_truncate_inode_items()
loops over many search_again descents under one handle) makes that walk
arbitrarily long. It runs in __btrfs_end_transaction() before
num_writers is dropped, so it blocks the committing thread; this shows up
as multi-second stalls and RCU stall reports.
Replace the xarray with a fixed inline array on btrfs_trans_handle,
managed with a CLOCK (second-chance) eviction policy. Inhibiting a buffer
becomes an array append with no allocation and no tree walk, and the
end-of-handle cleanup is bounded by the array size.
The set that actually needs protection is the working set the handle
revisits across search_again descents, the search path frontier, which is
on the order of the tree height. It is not every block the handle ever
COWs. should_cow_block() re-inhibiting an already tracked buffer marks it
referenced, so revisited buffers survive eviction while write-once buffers
are reclaimed first. A small fixed buffer is therefore enough where a
non-evicting array would either overflow or have to grow without bound.
BTRFS_INHIBITED_EBS_SLOTS is 8 and the reference bits pack into a u32.
The CLOCK eviction is what justifies the extra complexity over a plain
non-evicting array. The test workload stresses amplification: it removes
16 heavily fragmented 64 MiB files in one transaction while background
writeback keeps writing out in-use metadata. A re-COW event is a buffer
already COWed in the running transaction that was written back and then
COWed again; the figure below is the ratio of re-COW events to first-COW
events summed across the eviction (n=5, lower is better):
tracking re-COW per first-COW
no inhibition 6.1
non-evicting array, 32 slots 3.8
CLOCK array, 8 slots (this patch) 1.6
unbounded xarray (reverted) 1.4
The non-evicting array fills with write-once buffers and stops covering
the buffers the handle keeps revisiting, so even at four times the slots
it leaves most of the amplification. CLOCK evicts the cold buffers and
keeps the revisited ones, recovering almost all of the unbounded benefit.
The eviction policy, not the buffer size, is what closes the gap.
eb->writeback_inhibitors and the WB_SYNC_ALL bypass in
lock_extent_buffer_for_io() are unchanged, so fsync and commit behavior
are unaffected. A reference is taken on each tracked buffer so it cannot
be freed while the array points at it; eviction drops that reference and
the inhibitor count.
There's another testing report, showing 20% latency improvement on
reflink and deduplication synthetic benchmark. Full detailed report at
https://github.com/lcf0399/linux-regression-evidence/tree/main/btrfs-remap-writeback-inhibition-v2 .
Link: https://lore.kernel.org/all/CANGjgd=fQkHht2PdDi-+EAdzWH7UtxxWhhJ7b80Rr17PbpgxOw@mail.gmail.com/
Reported-by: kernel test robot <oliver.sang@intel.com>
Fixes: f9a48549a15a ("btrfs: inhibit extent buffer writeback to prevent COW amplification")
Closes: https://lore.kernel.org/oe-lkp/202603112240.f7605968-lkp@intel.com
Tested-by: Chengfeng Lin <lin2530632123@gmail.com>
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: Sun YangKai <sunk67188@gmail.com>
Signed-off-by: Leo Martins <loemra.dev@gmail.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
Diffstat (limited to 'tools/perf/scripts/python/stackcollapse.py')
0 files changed, 0 insertions, 0 deletions
