From: Tejun Heo <tj@kernel.org>
To: Jan Kara <jack@suse.cz>
Cc: Julian Sun <sunjunchao@bytedance.com>,
linux-mm@kvack.org, cgroups@vger.kernel.org, hannes@cmpxchg.org,
mhocko@kernel.org, roman.gushchin@linux.dev,
shakeel.butt@linux.dev, muchun.song@linux.dev,
akpm@linux-foundation.org, axboe@kernel.dk
Subject: Re: [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes
Date: Wed, 9 Sep 2026 08:34:03 -1000 [thread overview]
Message-ID: <aqGmmw3rv_2f5Jk4@slm.duckdns.org> (raw)
In-Reply-To: <iqvi6yo62ydxzshad2mhi55rxhyskfsoo62ji2curbldrg5cc4@z3wju4r6c3ab>
Hello,
On Wed, Sep 09, 2026 at 01:03:27PM +0200, Jan Kara wrote:
> On Wed 09-09-26 11:43:43, Julian Sun wrote:
> > The expectation in wbc_detach_inode() that "concurrent write sharing of an
> > inode is expected to be very rare" does not hold for bdev inodes. On ext4,
> > metadata from many memcgs shares the same bdev inode, so dirty throttling
> > in those memcgs can repeatedly trigger foreign flushes of the owner wb.
> > These flushes can also write ordinary file data, reducing overwrite
> > coalescing under continuous buffered overwrites. The resulting extra I/O
> > leaves less device bandwidth for other workloads.
>
> Hum. Are you aware of [1]? The guy is from the same company so I was
> assuming you two are working together on the problem :).
>
> [1] https://lore.kernel.org/all/cover.1788836830.git.yinxin.x@bytedance.com
>
> > Skip foreign dirty tracking for bdev inodes. Ordinary writeback and
> > foreign dirty tracking for non-bdev inodes are unchanged.
>
> Frankly, I'm not yet convinced this is worth the risk of regressions. The
> gains are nice but both the winning and the loosing loads are synthetic and
> none of them is like "this is definitely what most people are going to
> hit".
>
> I think what [1] is proposing is more in the direction where we might go
> (although the second patch in its current form isn't acceptable to me
> either). But what I miss from both your approaches is the verification that
> the change actually helps your *production* workloads. Unless it is an
> obvious fix (like was the case of patch 1 in the series [1]) I'm very
> reluctant to approve any additional heuristics unless it is proven they
> actually help production loads. These heuristics are never going to be
> perfect so you can always come up with a synthetic load breaking it in one
> way or another. The art is in tuning them so they work for the practical
> cases :)
The original design's assumption was that a single inode being actively
written into by multiple cgroups would be rare and rather silly and all we
need to do when that happens is being able to limp along. So, the trade-offs
were made in favor of lower overhead for the common cases where there can be
a huge number of active inodes. This mostly worked out for normal files.
Here, the situation is very different. The sharing isn't silly and there are
only a handful of these inodes. If this is an actual problem, I don't think
it can be solved by slewing the existing mechanism one way or another. Maybe
it will make some scenarios better but it will worsen others. The existing
mechanism simply doesn't have enough knoweldge to make sensible decisions.
Given that, if this is a real problem, I think the right solution likely is
tracking more states for these cases so that it can actually make sensible
decisions rather than trying to guess blindly.
Thanks.
--
tejun
prev parent reply other threads:[~2026-09-09 18:34 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 3:43 [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes Julian Sun
2026-09-09 11:03 ` Jan Kara
2026-09-09 13:18 ` [External] " Julian Sun
2026-09-11 11:20 ` Jan Kara
2026-09-11 11:39 ` Julian Sun
2026-09-11 15:39 ` Jan Kara
2026-09-14 2:18 ` Julian Sun
2026-09-09 18:34 ` Tejun Heo [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqGmmw3rv_2f5Jk4@slm.duckdns.org \
--to=tj@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=axboe@kernel.dk \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jack@suse.cz \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=sunjunchao@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox