From: Julian Sun <sunjunchao@bytedance.com>
To: linux-mm@kvack.org, cgroups@vger.kernel.org
Cc: hannes@cmpxchg.org, mhocko@kernel.org, roman.gushchin@linux.dev,
shakeel.butt@linux.dev, muchun.song@linux.dev,
akpm@linux-foundation.org, axboe@kernel.dk, tj@kernel.org,
jack@suse.cz
Subject: [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes
Date: Wed, 9 Sep 2026 11:43:43 +0800 [thread overview]
Message-ID: <20260909034343.340703-1-sunjunchao@bytedance.com> (raw)
The expectation in wbc_detach_inode() that "concurrent write sharing of an
inode is expected to be very rare" does not hold for bdev inodes. On ext4,
metadata from many memcgs shares the same bdev inode, so dirty throttling
in those memcgs can repeatedly trigger foreign flushes of the owner wb.
These flushes can also write ordinary file data, reducing overwrite
coalescing under continuous buffered overwrites. The resulting extra I/O
leaves less device bandwidth for other workloads.
Skip foreign dirty tracking for bdev inodes. Ordinary writeback and
foreign dirty tracking for non-bdev inodes are unchanged.
In a synthetic ext4 test with cgroup v2 (QEMU/KVM, 4 vCPUs, 8 GiB RAM),
device writes are capped at 200 MiB/s. fio repeatedly overwrites a 256 MiB
file in the root wb while four memory- and I/O-limited memcgs generate
metadata and trigger foreign flushes through dirty throttling. A concurrent
fio performs 16 GiB of sequential direct writes on the same filesystem.
Across five pairs with 9 GiB of buffered overwrites per run, mean results
are:
Baseline Skip bdev
Overwritten-file device writes 5.60 GiB 1.70 GiB (-69.6%)
Foreground completion time 104.2 s 89.1 s (-14.4%)
Foreground bandwidth 157.3 MiB/s 183.8 MiB/s
Foreground p99 completion latency 1434 ms 480 ms
The tradeoff is that memcgs with dirty bdev folios lose this way of
requesting writeback to relieve dirty pressure. In a separate uncapped ext4
test, a memcg with memory.high=256 MiB creates 60,000 files paced at 4,000
files/s, each with a distinct 128-byte xattr, and makes occasional buffered
data writes. The xattrs use external 4 KiB blocks. Across two pairs,
completion is 31-33% slower including final sync, while device writes
remain about 524 MiB.
With 64-byte xattrs stored within the inode, the same paced workload
showed no substantial completion-time regression, including at
memory.high=16 or 32 MiB where the baseline triggered dirty throttling
and foreign flushes.
I believe this external-xattr-heavy workload under memory pressure is
relatively uncommon in practice, and expect the benefits to outweigh the
regression. I therefore consider this an acceptable workaround.
Core fio commands for the overwrite test:
fio --name=hotspot --filename="$dir/hot.data" --rw=write \
--bs=64K --ioengine=psync --direct=0 --thread=1 \
--size=256M --io_size=9216M --rate=64M \
--refill_buffers=1 --invalidate=0 --eta=never \
--output-format=json --output="$out/target.json" &
hot_pid=$!
sleep 8
fio --name=foreground --filename="$dir/foreground.data" \
--rw=write --bs=1M --ioengine=libaio --iodepth=32 \
--direct=1 --thread=1 --invalidate=0 --size=4096M \
--io_size=16384M --eta=never --output-format=json \
--output="$out/foreground.json"
wait "$hot_pid"
In the log below, base is the original kernel and skip includes this
patch. hot_GiB is device writes for the repeatedly overwritten file,
including final sync; the other performance columns describe foreground
fio.
Full five-pair log:
nohup: ignoring input
[2026-09-05 17:52:10] RESULTS=/tmp/cgwb-20260905-175145; WARNING: reformats /dev/nvme5n1; 5 pair(s), limit=200 MiB/s
[2026-09-05 17:52:10] round-001-base: booting
[2026-09-05 17:52:20] round-001-base: running workload
[2026-09-05 17:55:36] round-001-skip: booting
[2026-09-05 17:55:45] round-001-skip: running workload
[2026-09-05 17:58:57] ROUND 1: kind seconds MiB/s IOPS mean_ms p99_ms hot_GiB WARN
[2026-09-05 17:58:57] base 104.574 156.67 156.67 203.82 1434.45 5.862 0
[2026-09-05 17:58:57] skip 89.125 183.83 183.83 173.62 480.25 1.656 0
[2026-09-05 17:58:57] skip vs base: bandwidth +17.33%, time -14.77%, hotspot IO -71.75%
[2026-09-05 17:58:57] round-002-skip: booting
[2026-09-05 17:59:06] round-002-skip: running workload
[2026-09-05 18:02:18] round-002-base: booting
[2026-09-05 18:02:27] round-002-base: running workload
[2026-09-05 18:05:42] ROUND 2: kind seconds MiB/s IOPS mean_ms p99_ms hot_GiB WARN
[2026-09-05 18:05:42] base 103.533 158.25 158.25 201.75 1434.45 5.493 0
[2026-09-05 18:05:42] skip 89.124 183.83 183.83 173.65 480.25 1.730 0
[2026-09-05 18:05:42] skip vs base: bandwidth +16.17%, time -13.92%, hotspot IO -68.50%
[2026-09-05 18:05:42] round-003-base: booting
[2026-09-05 18:05:52] round-003-base: running workload
[2026-09-05 18:09:06] round-003-skip: booting
[2026-09-05 18:09:15] round-003-skip: running workload
[2026-09-05 18:12:28] ROUND 3: kind seconds MiB/s IOPS mean_ms p99_ms hot_GiB WARN
[2026-09-05 18:12:28] base 104.648 156.56 156.56 203.96 1434.45 5.334 1
[2026-09-05 18:12:28] skip 89.243 183.59 183.59 173.92 480.25 1.824 0
[2026-09-05 18:12:28] skip vs base: bandwidth +17.26%, time -14.72%, hotspot IO -65.80%
[2026-09-05 18:12:28] WARN: runtime diagnostics present; inspect each data/dmesg before using this pair.
[2026-09-05 18:12:28] round-004-skip: booting
[2026-09-05 18:12:37] round-004-skip: running workload
[2026-09-05 18:15:50] round-004-base: booting
[2026-09-05 18:16:00] round-004-base: running workload
[2026-09-05 18:19:15] ROUND 4: kind seconds MiB/s IOPS mean_ms p99_ms hot_GiB WARN
[2026-09-05 18:19:15] base 103.565 158.20 158.20 201.84 1434.45 5.497 0
[2026-09-05 18:19:15] skip 89.124 183.83 183.83 173.57 480.25 1.648 0
[2026-09-05 18:19:15] skip vs base: bandwidth +16.20%, time -13.94%, hotspot IO -70.01%
[2026-09-05 18:19:15] round-005-base: booting
[2026-09-05 18:19:24] round-005-base: running workload
[2026-09-05 18:22:41] round-005-skip: booting
[2026-09-05 18:22:50] round-005-skip: running workload
[2026-09-05 18:26:03] ROUND 5: kind seconds MiB/s IOPS mean_ms p99_ms hot_GiB WARN
[2026-09-05 18:26:03] base 104.567 156.68 156.68 203.78 1434.45 5.830 1
[2026-09-05 18:26:03] skip 89.035 184.02 184.02 173.47 480.25 1.648 0
[2026-09-05 18:26:03] skip vs base: bandwidth +17.44%, time -14.85%, hotspot IO -71.72%
[2026-09-05 18:26:03] WARN: runtime diagnostics present; inspect each data/dmesg before using this pair.
[2026-09-05 18:26:03] DONE: 5 pair(s); all test VMs stopped
Fixes: 97b27821b485 ("writeback, memcg: Implement foreign dirty flushing")
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
mm/memcontrol.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1271d390b617..6e8f7ff7d44a 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -3898,6 +3898,15 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
u64 oldest_at = now;
int oldest = -1;
int i;
+ struct address_space *mapping = folio_mapping(folio);
+ struct inode *bdev_inode = mapping ? mapping->host : NULL;
+
+ /*
+ * Bdev inodes are usually shared by many memcgs so foreign
+ * tracking leads to frequent flushes which is counterproductive.
+ */
+ if (bdev_inode && sb_is_blkdev_sb(bdev_inode->i_sb))
+ return;
trace_track_foreign_dirty(folio, wb);
--
2.39.5
next reply other threads:[~2026-09-09 3:43 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 3:43 Julian Sun [this message]
2026-09-09 11:03 ` [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes Jan Kara
2026-09-09 13:18 ` [External] " Julian Sun
2026-09-11 11:20 ` Jan Kara
2026-09-11 11:39 ` Julian Sun
2026-09-11 15:39 ` Jan Kara
2026-09-14 2:18 ` Julian Sun
2026-09-09 18:34 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260909034343.340703-1-sunjunchao@bytedance.com \
--to=sunjunchao@bytedance.com \
--cc=akpm@linux-foundation.org \
--cc=axboe@kernel.dk \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jack@suse.cz \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox