From: Julian Sun <sunjunchao@bytedance.com>
To: Jan Kara <jack@suse.cz>
Cc: linux-mm@kvack.org, cgroups@vger.kernel.org, hannes@cmpxchg.org,
mhocko@kernel.org, roman.gushchin@linux.dev,
shakeel.butt@linux.dev, muchun.song@linux.dev,
akpm@linux-foundation.org, axboe@kernel.dk, tj@kernel.org
Subject: Re: [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes
Date: Mon, 14 Sep 2026 10:18:29 +0800 [thread overview]
Message-ID: <e55a7422-1871-4a03-80bd-09cedf9630dd@bytedance.com> (raw)
In-Reply-To: <l2awpukcp7lorvujvwdoqnlwlkkjl6e2kef6tnjvdw7qve2fni@vh4lspozoytt>
On 9/11/26 11:39 PM, Jan Kara wrote:
> On Fri 11-09-26 19:39:42, Julian Sun wrote:
>> On 9/11/26 7:20 PM, Jan Kara wrote:
>>> On Wed 09-09-26 21:18:11, Julian Sun wrote:
>>>> On Wed, Sep 9, 2026 at 7:13 PM Jan Kara <jack@suse.cz> wrote:
>>>>>
>>>>> On Wed 09-09-26 11:43:43, Julian Sun wrote:
>>>>>> The expectation in wbc_detach_inode() that "concurrent write sharing of an
>>>>>> inode is expected to be very rare" does not hold for bdev inodes. On ext4,
>>>>>> metadata from many memcgs shares the same bdev inode, so dirty throttling
>>>>>> in those memcgs can repeatedly trigger foreign flushes of the owner wb.
>>>>>> These flushes can also write ordinary file data, reducing overwrite
>>>>>> coalescing under continuous buffered overwrites. The resulting extra I/O
>>>>>> leaves less device bandwidth for other workloads.
>>>>>
>>>>> Hum. Are you aware of [1]? The guy is from the same company so I was
>>>>> assuming you two are working together on the problem :).
>>>>
>>>> Yes, Xin and I are on the same team, and we are looking at different
>>>> ways to mitigate this issue.
>>>
>>> OK :).
>>>
>>>> We will continue trying to gather more evidence from production to
>>>> better understand its actual impact. The number of calls to
>>>> `balance_dirty_pages()` and the time spent in it might be useful
>>>> metrics for this purpose.
>>>>
>>>> Jan, if we can show that this patch provides a measurable improvement
>>>> in production, would you consider accepting it? Or do you think we
>>>> should still look for a better approach? If you think there is a
>>>> better way to address this, we'd be very interested in your
>>>> suggestion.
>>>
>>> I definitely acknowledge that there's a problem with foreign flushing due
>>> to bdev inodes that needs to be fixed because as Tejun wrote, bdev inodes
>>> break the fundamental assumptions of foreign flushes. But so far I haven't
>>> seen a solution which I'd definitely support.
>>>
>>> Having data from production to understand what exactly is the practical
>>> problem will hopefully help us understand how the solution should look
>>> like. Is it just that we trigger too many flushes and that slows down
>>> things? Or is even one foreign flush of the whole wb due to bdev inode
>>> making performance bad? Is it that we have many foreign pages and foreign
>>> flushes fail to properly write them? These are questions we need to answer
>>> to decide how the solution should look like. And the solution could be
>>> anything from just limiting number of foreign flushes for a wb to 1
>>> (because there really isn't big point in more of them being queued) to more
>>> involved foreign page tracking for bdev inodes so that flushes can be more
>>> targetted.
>>
>> Hi Jan,
>>
>> Thanks for your feedback.
>>
>> All the foreign flushes we've observed were triggered by bdev inodes.
>> What we've observed is that the bdev inode has only around 10 MB of
>> dirty pages, but it triggers writeback of around 40 GB of dirty pages
>> across the entire wb. The actual foreign pages account for less than
>> 10 MB, yet flushing them triggers writeback of the whole wb, and this
>> happens repeatedly.
>>
>> This appears to cause significant write amplification.
>>
>> Limiting foreign flushes to one may not be sufficient in this case,
>> since even a single flush can trigger a large amount of unnecessary
>> writeback. I'm currently working on tracking foreign inodes so that
>> foreign flushes can target those inodes specifically.
>
> OK, good, these are already some tangible data to use :). So in that case
> I'd maybe just:
>
> 1) Avoid adding 'wb' as foreign just because it owns bdev inode where we
> are dirtying a page.
> 2) Track that we have some foreign pages in this particular bdev inode.
> 3) In balance dirty pages queue flushes of bdev inodes where we have
> foreign pages (besides queueing standard foreign flushes).
>
> I didn't put too much thought into how to integrate this in an elegant way
> with the current writeback infrastructure but something like this shouldn't
> be that hard and it should help you with the foreign flush issues you
> observe.
>
>> The impact on production workloads is not easy to measure directly,
>> which is another challenge we're facing. We've had reports of slow
>> dirty page writeback from production workloads. Although we haven't
>> established the full causal chain yet, we suspect foreign flushes may
>> be contributing to the problem.
Jan, thanks for the suggestion. This approach makes more sense.
I'll implement it and send a patch after testing.
>
> Just being able to measure that "with these patches we don't need to flush
> 40GB of unrelated data in foreign flushes in our production workloads" is
> enough to convince me they bring some tangible benefit for your workload
> :).
>
> Honza
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
next prev parent reply other threads:[~2026-09-14 2:18 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 3:43 [PATCH v3] writeback, memcg: skip foreign dirty tracking for bdev inodes Julian Sun
2026-09-09 11:03 ` Jan Kara
2026-09-09 13:18 ` [External] " Julian Sun
2026-09-11 11:20 ` Jan Kara
2026-09-11 11:39 ` Julian Sun
2026-09-11 15:39 ` Jan Kara
2026-09-14 2:18 ` Julian Sun [this message]
2026-09-09 18:34 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e55a7422-1871-4a03-80bd-09cedf9630dd@bytedance.com \
--to=sunjunchao@bytedance.com \
--cc=akpm@linux-foundation.org \
--cc=axboe@kernel.dk \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jack@suse.cz \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox