From: Julian Sun <sunjunchao@bytedance.com>
To: linux-block@vger.kernel.org, cgroups@vger.kernel.org,
linux-mm@kvack.org, linux-fsdevel@vger.kernel.org
Cc: axboe@kernel.dk, hannes@cmpxchg.org, mhocko@kernel.org,
roman.gushchin@linux.dev, shakeel.butt@linux.dev,
muchun.song@linux.dev, willy@infradead.org, jack@suse.cz,
tj@kernel.org, akpm@linux-foundation.org
Subject: Re: [PATCH v8 0/3] memcg,writeback: flush foreign bdev mappings separately
Date: Sun, 27 Sep 2026 19:48:26 +0800 [thread overview]
Message-ID: <b5da5c30-7c8b-4c60-aa3b-d2e3603e82c1@bytedance.com> (raw)
In-Reply-To: <20260922125015.3435215-1-sunjunchao@bytedance.com>
On 9/22/26 8:50 PM, Julian Sun wrote:
Hi,
Do you have any further thoughts on this series?
Should I rename bdev_flush_by_dev() to bdev_writeback_by_dev()
and send the next revision?
> Hi,
>
> This series avoids owner-wide foreign writeback triggered by dirtying
> shared bdev inodes. Instead, it tracks bdev targets separately and
> schedules writeback of their mappings.
>
> Problem
> =======
>
> Bdev inodes are commonly shared by many memcgs. Dirtying their pages
> can cause the owner wb to be tracked as foreign, triggering writeback
> of the entire wb rather than just the bdev mapping.
>
> In production, almost all the foreign dirtying we traced came from bdev
> inodes. Across our observations, a bdev mapping had as little as about
> 10 MiB of dirty pages, yet foreign flushes targeted the entire owner
> wb. On one machine, we observed 599 such works queued on one wb, with
> an average request budget of about 43 GiB per work.
>
> Measured impact
> ===============
>
> The synthetic test used a VM with 4 vCPUs, 8 GiB RAM, ext4 and cgroup
> v2, with device write bandwidth capped at 200 MiB/s. Background fio
> repeatedly overwrote a 256 MiB file using buffered I/O, while
> foreground fio performed sequential direct writes. Four memcgs
> generated metadata and buffered writes to trigger foreign flushes.
>
> We measured completed block writes to the background fio file,
> including final sync, and foreground fio bandwidth, latency and
> completion time.
>
> Across three runs per kernel, mean background file write I/O decreased
> by 72.7%, foreground bandwidth increased by 23.0%, and completion time
> decreased by 18.8%.
>
> Metric Baseline Patched Change
> Background file write I/O (GiB) 6.154 1.680 -72.7%
> Foreground bandwidth (MiB/s) 150.29 184.81 +23.0%
> Foreground completion time (s) 109.203 88.652 -18.8%
> Foreground mean latency (ms) 212.84 172.74 -18.8%
>
> How it helps
> ============
>
> Frequent writeback of the file overwritten by background fio reduces
> the opportunity to coalesce buffered writes in the page cache.
> Several overwrites of the same page can otherwise result in a single
> write of its latest contents. Flushing between updates instead writes
> the same page repeatedly, generating more device I/O for the same
> application writes.
>
> This additional I/O makes the device reach its performance limit
> more easily and competes with foreground I/O for device bandwidth.
>
> The benefit is expected to be most pronounced when these conditions
> occur together on the same device:
>
> 1. Limited device bandwidth, as with an HDD, makes the additional
> writeback I/O more likely to saturate the device.
> 2. Multiple cgroups perform buffered overwrites. Frequent foreign
> writeback reduces write coalescing and increases device I/O.
> 3. Concurrent direct I/O bypasses the page cache and competes directly
> with writeback I/O for device bandwidth.
>
> Approach
> ========
>
> Following Jan's suggestion, this series records foreign bdev targets
> separately and flushes their mappings instead of their owner wbs. This
> preserves a way for dirty throttling to initiate bdev writeback while
> avoiding owner-wide writeback triggered by these records. The existing
> foreign-wb mechanism remains unchanged for other inodes.
>
> Tracking uses oldest-first replacement and expiry after
> dirty_expire_interval. Further dirtying during an ongoing flush can
> trigger another flush after it finishes.
>
> Tracking is best effort: each memcg keeps a bounded set of device
> numbers without persistent device or inode references. Recording runs
> under mapping->i_pages with IRQs disabled, while dropping a device
> reference can sleep. A memcg that never enters dirty throttling could
> also retain an unused record and pin the device object indefinitely.
> Recording only dev_t avoids these lifetime constraints. Device removal
> and device-number reuse may cause an unintended best-effort flush,
> which is an accepted trade-off.
>
> Writeback submission may block for a long time, with up to four work
> items outstanding per memcg. A dedicated unbound workqueue provides
> a separate max_active budget so these flushes do not exhaust active
> slots on a shared workqueue and delay unrelated work. If the workqueue
> cannot be allocated at boot, bdev inodes continue to use the existing
> foreign-wb path.
>
> This primarily affects filesystems that keep metadata in the bdev page
> cache, such as ext4 and other buffer_head users, rather than the usual
> metadata writeback paths of XFS and Btrfs. This also covers buffered
> writes to raw block devices. The decision still does not account for
> how much of the memcg's dirty memory belongs to the bdev: even a single
> dirty bitmap block can trigger writeback of the entire bdev mapping
> when the memcg enters dirty throttling.
>
> Changes since v7:
> - Update comments and commit messages as suggested by Tejun Heo.
>
> Changes since v6:
> - Drop open_mutex and the openers check from bdev_flush_by_dev(), and
> add a !CONFIG_BLOCK stub.
> - Follow the existing foreign-wb tracking policy: use timestamps for
> oldest-first replacement and expiry, and allow dirtying during an
> in-flight flush to re-arm the record.
> - Replace frn_lock with an atomic in-flight flag.
> - Drop WQ_MEM_RECLAIM and explain the separate max_active budget.
> - Expand the dev_t lifetime rationale and filesystem scope, update
> the foreign-dirty overview, and drop the Fixes tag.
>
> Changes since v5:
> - Fix repeated overwriting of slot 0 in patch 2 and add comments.
>
> Changes since v4:
> - Add before-and-after performance comparisons to the cover letter.
>
> Changes since v3:
> - Flush foreign-dirtied bdev inode mappings separately, as suggested
> by Jan.
>
> Thanks.
>
> Julian Sun (3):
> block: introduce bdev_flush_by_dev()
> memcg,writeback: flush foreign bdev mappings separately from owner wbs
> writeback: record bdev targets in foreign writeback tracepoints
>
> block/bdev.c | 18 ++++++
> include/linux/blkdev.h | 4 ++
> include/linux/memcontrol.h | 12 ++++
> include/trace/events/writeback.h | 22 ++++---
> mm/memcontrol.c | 123 +++++++++++++++++++++++++++++++++++----
> 5 files changed, 160 insertions(+), 19 deletions(-)
>
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
next prev parent reply other threads:[~2026-09-27 11:48 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 12:50 [PATCH v8 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-22 12:50 ` [PATCH v8 1/3] block: introduce bdev_flush_by_dev() Julian Sun
2026-09-24 10:57 ` Jan Kara
2026-09-24 12:06 ` Jan Kara
2026-09-22 12:50 ` [PATCH v8 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
2026-09-24 12:02 ` Jan Kara
2026-09-25 6:04 ` Julian Sun
2026-09-22 12:50 ` [PATCH v8 3/3] writeback: record bdev targets in foreign writeback tracepoints Julian Sun
2026-09-24 10:57 ` Jan Kara
2026-09-22 22:02 ` [PATCH v8 0/3] memcg,writeback: flush foreign bdev mappings separately Tejun Heo
2026-09-27 11:48 ` Julian Sun [this message]
2026-09-27 20:26 ` Andrew Morton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b5da5c30-7c8b-4c60-aa3b-d2e3603e82c1@bytedance.com \
--to=sunjunchao@bytedance.com \
--cc=akpm@linux-foundation.org \
--cc=axboe@kernel.dk \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jack@suse.cz \
--cc=linux-block@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=muchun.song@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=tj@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox