From: Julian Sun <sunjunchao@bytedance.com>
To: Jan Kara <jack@suse.cz>
Cc: linux-mm@kvack.org, cgroups@vger.kernel.org, hannes@cmpxchg.org,
mhocko@kernel.org, roman.gushchin@linux.dev,
shakeel.butt@linux.dev, muchun.song@linux.dev,
akpm@linux-foundation.org, axboe@kernel.dk, tj@kernel.org
Subject: Re: [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes
Date: Fri, 11 Sep 2026 19:39:42 +0800 [thread overview]
Message-ID: <db6d62f4-34f0-4237-ab58-c21b0968bee7@bytedance.com> (raw)
In-Reply-To: <ekylbhyzseq25nxq2yivmmdt5xl5glum5proueyd5mrqfux2y6@fpvflg4crr2i>
On 9/11/26 7:20 PM, Jan Kara wrote:
> On Wed 09-09-26 21:18:11, Julian Sun wrote:
>> On Wed, Sep 9, 2026 at 7:13 PM Jan Kara <jack@suse.cz> wrote:
>>>
>>> On Wed 09-09-26 11:43:43, Julian Sun wrote:
>>>> The expectation in wbc_detach_inode() that "concurrent write sharing of an
>>>> inode is expected to be very rare" does not hold for bdev inodes. On ext4,
>>>> metadata from many memcgs shares the same bdev inode, so dirty throttling
>>>> in those memcgs can repeatedly trigger foreign flushes of the owner wb.
>>>> These flushes can also write ordinary file data, reducing overwrite
>>>> coalescing under continuous buffered overwrites. The resulting extra I/O
>>>> leaves less device bandwidth for other workloads.
>>>
>>> Hum. Are you aware of [1]? The guy is from the same company so I was
>>> assuming you two are working together on the problem :).
>>
>> Yes, Xin and I are on the same team, and we are looking at different
>> ways to mitigate this issue.
>
> OK :).
>
>> We will continue trying to gather more evidence from production to
>> better understand its actual impact. The number of calls to
>> `balance_dirty_pages()` and the time spent in it might be useful
>> metrics for this purpose.
>>
>> Jan, if we can show that this patch provides a measurable improvement
>> in production, would you consider accepting it? Or do you think we
>> should still look for a better approach? If you think there is a
>> better way to address this, we'd be very interested in your
>> suggestion.
>
> I definitely acknowledge that there's a problem with foreign flushing due
> to bdev inodes that needs to be fixed because as Tejun wrote, bdev inodes
> break the fundamental assumptions of foreign flushes. But so far I haven't
> seen a solution which I'd definitely support.
>
> Having data from production to understand what exactly is the practical
> problem will hopefully help us understand how the solution should look
> like. Is it just that we trigger too many flushes and that slows down
> things? Or is even one foreign flush of the whole wb due to bdev inode
> making performance bad? Is it that we have many foreign pages and foreign
> flushes fail to properly write them? These are questions we need to answer
> to decide how the solution should look like. And the solution could be
> anything from just limiting number of foreign flushes for a wb to 1
> (because there really isn't big point in more of them being queued) to more
> involved foreign page tracking for bdev inodes so that flushes can be more
> targetted.
Hi Jan,
Thanks for your feedback.
All the foreign flushes we've observed were triggered by bdev inodes.
What we've observed is that the bdev inode has only around 10 MB of
dirty pages, but it triggers writeback of around 40 GB of dirty pages
across the entire wb. The actual foreign pages account for less than
10 MB, yet flushing them triggers writeback of the whole wb, and this
happens repeatedly.
This appears to cause significant write amplification.
Limiting foreign flushes to one may not be sufficient in this case,
since even a single flush can trigger a large amount of unnecessary
writeback. I'm currently working on tracking foreign inodes so that
foreign flushes can target those inodes specifically.
The impact on production workloads is not easy to measure directly,
which is another challenge we're facing. We've had reports of slow
dirty page writeback from production workloads. Although we haven't
established the full causal chain yet, we suspect foreign flushes may
be contributing to the problem.
>
> Honza
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
next prev parent reply other threads:[~2026-09-11 11:39 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 3:43 [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes Julian Sun
2026-09-09 11:03 ` Jan Kara
2026-09-09 13:18 ` [External] " Julian Sun
2026-09-11 11:20 ` Jan Kara
2026-09-11 11:39 ` Julian Sun [this message]
2026-09-11 15:39 ` Jan Kara
2026-09-14 2:18 ` Julian Sun
2026-09-09 18:34 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=db6d62f4-34f0-4237-ab58-c21b0968bee7@bytedance.com \
--to=sunjunchao@bytedance.com \
--cc=akpm@linux-foundation.org \
--cc=axboe@kernel.dk \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jack@suse.cz \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox