Linux cgroups development
 help / color / mirror / Atom feed
* [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately
@ 2026-09-17  6:57 Julian Sun
  2026-09-17  6:57 ` [PATCH v6 1/3] block: introduce bdev_flush_by_dev() Julian Sun
                   ` (4 more replies)
  0 siblings, 5 replies; 10+ messages in thread
From: Julian Sun @ 2026-09-17  6:57 UTC (permalink / raw)
  To: linux-block, cgroups, linux-mm, linux-fsdevel
  Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
	willy, jack, tj, akpm

Hi,

This series avoids owner-wide foreign writeback triggered by dirtying
shared bdev inodes. Instead, it tracks bdev targets separately and
schedules writeback of their mappings.

Problem
=======

Bdev inodes are commonly shared by many memcgs. Dirtying their pages
can cause the owner wb to be tracked as foreign, triggering writeback
of the entire wb rather than just the bdev mapping.

In production, almost all the foreign dirtying we traced came from bdev
inodes. Across our observations, a bdev mapping had as little as about
10 MiB of dirty pages, yet foreign flushes targeted the entire owner
wb. On one machine, we observed 599 such works queued on one wb, with
an average request budget of about 43 GiB per work.

Measured impact
===============

The synthetic test used a VM with 4 vCPUs, 8 GiB RAM, ext4 and cgroup
v2, with device write bandwidth capped at 200 MiB/s. Background fio
repeatedly overwrote a 256 MiB file using buffered I/O, while
foreground fio performed sequential direct writes. Four memcgs
generated metadata and buffered writes to trigger foreign flushes.

We measured completed block writes to the background fio file,
including final sync, and foreground fio bandwidth, latency and
completion time.

Across three runs per kernel, mean background file write I/O decreased
by 72.7%, foreground bandwidth increased by 23.0%, and completion time
decreased by 18.8%.

Metric                         Baseline    Patched     Change
Background file write I/O (GiB)    6.154      1.680     -72.7%
Foreground bandwidth (MiB/s)      150.29     184.81     +23.0%
Foreground completion time (s)   109.203     88.652     -18.8%
Foreground mean latency (ms)      212.84     172.74     -18.8%

How it helps
============

Frequent writeback of the file overwritten by background fio reduces
the opportunity to coalesce buffered writes in the page cache.
Several overwrites of the same page can otherwise result in a single
write of its latest contents. Flushing between updates instead writes
the same page repeatedly, generating more device I/O for the same
application writes.

This additional I/O makes the device reach its performance limit
more easily and competes with foreground I/O for device bandwidth.

The benefit is expected to be most pronounced when these conditions
occur together on the same device:

1. Limited device bandwidth, as with an HDD, makes the additional
   writeback I/O more likely to saturate the device.
2. Multiple cgroups perform buffered overwrites. Frequent foreign
   writeback reduces write coalescing and increases device I/O.
3. Concurrent direct I/O bypasses the page cache and competes directly
   with writeback I/O for device bandwidth.

Approach
========

Following Jan's suggestion, this series records foreign bdev targets
separately and flushes their mappings instead of their owner wbs. This
preserves a way for dirty throttling to initiate bdev writeback while
avoiding owner-wide writeback triggered by these records. The existing
foreign-wb mechanism remains unchanged for other inodes.

Tracking is best effort: each memcg keeps a bounded set of device
numbers without persistent device or inode references. Closed devices
and devices with a busy open_mutex are skipped. Device removal and
device-number reuse may race with lookup.

Flushes run asynchronously on a dedicated workqueue to isolate these
frequently triggered tasks from existing workqueue users. If workqueue
allocation fails, the existing foreign-wb mechanism remains in use.

Changes since v5:
- Fix repeated overwriting of slot 0 in patch 2 and add comments.

Changes since v4:
- Add before-and-after performance comparisons to the cover letter.

Changes since v3:
- Flush foreign-dirtied bdev inode mappings separately, as suggested
  by Jan.

Thanks.

Julian Sun (3):
  block: introduce bdev_flush_by_dev()
  memcg,writeback: flush foreign bdev mappings separately from owner wbs
  writeback: record bdev targets in foreign writeback tracepoints

 block/bdev.c                     |  24 ++++++
 include/linux/blkdev.h           |   1 +
 include/linux/memcontrol.h       |  15 ++++
 include/trace/events/writeback.h |  22 ++++--
 mm/memcontrol.c                  | 124 ++++++++++++++++++++++++++++---
 5 files changed, 168 insertions(+), 18 deletions(-)

-- 
2.39.5

^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH v6 1/3] block: introduce bdev_flush_by_dev()
  2026-09-17  6:57 [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
@ 2026-09-17  6:57 ` Julian Sun
  2026-09-20 12:37   ` Tejun Heo
  2026-09-17  7:01 ` [PATCH v6 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
                   ` (3 subsequent siblings)
  4 siblings, 1 reply; 10+ messages in thread
From: Julian Sun @ 2026-09-17  6:57 UTC (permalink / raw)
  To: linux-block, cgroups, linux-mm, linux-fsdevel
  Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
	willy, jack, tj, akpm

Add bdev_flush_by_dev() to submit page-cache writeback for a device
identified by dev_t without requiring callers to hold a persistent
device reference.

A subsequent patch will use this helper to flush foreign bdev mappings.

Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
 block/bdev.c           | 24 ++++++++++++++++++++++++
 include/linux/blkdev.h |  1 +
 2 files changed, 25 insertions(+)

diff --git a/block/bdev.c b/block/bdev.c
index cd8323083740..da1fc56926fd 100644
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -264,6 +264,30 @@ int sync_blockdev_nowait(struct block_device *bdev)
 }
 EXPORT_SYMBOL_GPL(sync_blockdev_nowait);
 
+/**
+ * bdev_flush_by_dev - attempt writeback of a block device's page cache
+ * @dev: target block device number
+ *
+ * Skip closed devices or devices whose open_mutex is busy.
+ * May sleep while submitting writeback; does not wait for all I/O to complete.
+ */
+void bdev_flush_by_dev(dev_t dev)
+{
+	struct block_device *bdev;
+
+	bdev = blkdev_get_no_open(dev, false);
+	if (!bdev)
+		return;
+
+	/* Prevent the device from closing between the check and writeback. */
+	if (mutex_trylock(&bdev->bd_disk->open_mutex)) {
+		if (atomic_read(&bdev->bd_openers))
+			sync_blockdev_nowait(bdev);
+		mutex_unlock(&bdev->bd_disk->open_mutex);
+	}
+	blkdev_put_no_open(bdev);
+}
+
 /*
  * Write out and wait upon all the dirty data associated with a block
  * device via its mapping.  Does not take the superblock lock.
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 4f7905c3412b..7a5f3e2f17d6 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1710,6 +1710,7 @@ void invalidate_bdev(struct block_device *bdev);
 int sync_blockdev(struct block_device *bdev);
 int sync_blockdev_range(struct block_device *bdev, loff_t lstart, loff_t lend);
 int sync_blockdev_nowait(struct block_device *bdev);
+void bdev_flush_by_dev(dev_t dev);
 void sync_bdevs(bool wait);
 void bdev_statx(const struct path *path, struct kstat *stat, u32 request_mask);
 void printk_all_partitions(void);
-- 
2.39.5


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH v6 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
  2026-09-17  6:57 [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
  2026-09-17  6:57 ` [PATCH v6 1/3] block: introduce bdev_flush_by_dev() Julian Sun
@ 2026-09-17  7:01 ` Julian Sun
  2026-09-20 12:37   ` Tejun Heo
  2026-09-17  7:01 ` [PATCH v6 3/3] writeback: record bdev targets in foreign writeback tracepoints Julian Sun
                   ` (2 subsequent siblings)
  4 siblings, 1 reply; 10+ messages in thread
From: Julian Sun @ 2026-09-17  7:01 UTC (permalink / raw)
  To: linux-block, cgroups, linux-mm, linux-fsdevel
  Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
	willy, jack, tj, akpm

Bdev inodes are routinely shared by multiple memcgs. Recording their
owner wb as foreign can flush unrelated file data when a source memcg
enters dirty throttling.

Record bdev device numbers separately and queue best-effort work on
memcg_bdev_frn_flusher to write only their mappings. Keep the existing
foreign-wb mechanism for non-bdev inodes.

Bdev foreign flushes can be triggered frequently. Use a dedicated
workqueue to avoid interfering with tasks on existing workqueues. If
allocation fails, retain the existing foreign-wb path.

Use fixed per-memcg slots, with a shared lock protecting device records
and in-flight state. Tracking is best effort: record only dev_t without
holding device or inode references. Skip closed devices or devices whose
open_mutex is busy. Device removal and device-number reuse may race with
lookup.

Fixes: 97b27821b485 ("writeback, memcg: Implement foreign dirty flushing")
Suggested-by: Jan Kara <jack@suse.cz>
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
 include/linux/memcontrol.h |  15 +++++
 mm/memcontrol.c            | 117 ++++++++++++++++++++++++++++++++++---
 2 files changed, 124 insertions(+), 8 deletions(-)

diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 7d1c0ce189a8..3cd90913773b 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -176,6 +176,18 @@ struct memcg_cgwb_frn {
 	struct wb_completion done;	/* tracks in-flight foreign writebacks */
 };
 
+/*
+ * frn_lock protects dev and inflight.
+ * No extra memcg reference is taken: this work is embedded in the memcg,
+ * and mem_cgroup_css_free() waits for it to finish before freeing the memcg.
+ */
+struct bdev_frn_flush_ctx {
+	struct work_struct work;
+	dev_t dev;
+	bool inflight;
+	struct mem_cgroup *memcg;
+};
+
 /*
  * Bucket for arbitrarily byte-sized objects charged to a memory
  * cgroup. The bucket can be reparented in one piece when the cgroup
@@ -279,6 +291,9 @@ struct mem_cgroup {
 #ifdef CONFIG_CGROUP_WRITEBACK
 	struct wb_domain cgwb_domain;
 	struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
+	struct bdev_frn_flush_ctx bdev_frn[MEMCG_CGWB_FRN_CNT];
+	/* Nests inside mapping->i_pages in the dirty tracking path. */
+	spinlock_t frn_lock;
 #endif
 
 #ifdef CONFIG_LRU_GEN_WALKS_MMU
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1271d390b617..55bd5d100caa 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -39,6 +39,7 @@
 #include <linux/smp.h>
 #include <linux/page-flags.h>
 #include <linux/backing-dev.h>
+#include <linux/blkdev.h>
 #include <linux/bit_spinlock.h>
 #include <linux/rcupdate.h>
 #include <linux/limits.h>
@@ -104,6 +105,7 @@ static struct kmem_cache *memcg_pn_cachep;
 
 #ifdef CONFIG_CGROUP_WRITEBACK
 static DECLARE_WAIT_QUEUE_HEAD(memcg_cgwb_frn_waitq);
+static struct workqueue_struct *memcg_bdev_frn_wq __ro_after_init;
 #endif
 
 static inline bool task_is_dying(void)
@@ -3845,6 +3847,52 @@ void mem_cgroup_wb_stats(struct bdi_writeback *wb, unsigned long *pfilepages,
 	}
 }
 
+/*
+ * Bdev inodes are commonly shared by many memcgs. Flushing their owner wb
+ * can write unrelated file data, so record device numbers separately and
+ * flush only the bdev mappings. Tracking is best effort: record only dev_t
+ * without holding device or inode references, so it may race with device
+ * removal and re-addition.
+ */
+static void mem_cgroup_track_foreign_bdev(struct mem_cgroup *memcg, dev_t dev)
+{
+	struct bdev_frn_flush_ctx *ctx;
+	unsigned long flags;
+	int slot = -1;
+	int i;
+
+	/*
+	 * __folio_mark_dirty() takes mapping->i_pages with xa_lock_irqsave()
+	 * before reaching this helper. Use irqsave here as well so frn_lock's
+	 * IRQ protection does not depend on that outer locking.
+	 */
+	spin_lock_irqsave(&memcg->frn_lock, flags);
+	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+		if (memcg->bdev_frn[i].dev == dev)
+			goto out;
+	}
+
+	/*
+	 * Tracking is best effort, so losing hints when all slots are occupied
+	 * is expected.
+	 */
+	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+		ctx = &memcg->bdev_frn[i];
+		if (ctx->inflight)
+			continue;
+		if (slot < 0)
+			slot = i;
+		if (!ctx->dev) {
+			slot = i;
+			break;
+		}
+	}
+	if (slot >= 0)
+		memcg->bdev_frn[slot].dev = dev;
+out:
+	spin_unlock_irqrestore(&memcg->frn_lock, flags);
+}
+
 /*
  * Foreign dirty flushing
  *
@@ -3875,17 +3923,17 @@ void mem_cgroup_wb_stats(struct bdi_writeback *wb, unsigned long *pfilepages,
  * most recent foreign dirtying events and initiating remote flushes on
  * them when local writeback isn't enough to keep the memory clean enough.
  *
- * The following two functions implement such mechanism.  When a foreign
- * page - a page whose memcg and writeback ownerships don't match - is
- * dirtied, mem_cgroup_track_foreign_dirty() records the inode owning
- * bdi_writeback on the page owning memcg.  When balance_dirty_pages()
- * decides that the memcg needs to sleep due to high dirty ratio, it calls
+ * For non-bdev inodes, when a foreign page - a page whose memcg and
+ * writeback ownerships don't match - is dirtied,
+ * mem_cgroup_track_foreign_dirty() records the inode owning bdi_writeback
+ * on the page owning memcg. When balance_dirty_pages() decides that the
+ * memcg needs to sleep due to high dirty ratio, it calls
  * mem_cgroup_flush_foreign() which queues writeback on the recorded
  * foreign bdi_writebacks which haven't expired.  Both the numbers of
  * recorded bdi_writebacks and concurrent in-flight foreign writebacks are
  * limited to MEMCG_CGWB_FRN_CNT.
  *
- * The mechanism only remembers IDs and doesn't hold any object references.
+ * These wb records only remember IDs and don't hold any object references.
  * As being wrong occasionally doesn't matter, updates and accesses to the
  * records are lockless and racy.
  */
@@ -3898,9 +3946,17 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
 	u64 oldest_at = now;
 	int oldest = -1;
 	int i;
+	struct address_space *mapping = folio_mapping(folio);
+	struct inode *bdev_inode = mapping ? mapping->host : NULL;
 
 	trace_track_foreign_dirty(folio, wb);
 
+	if (memcg_bdev_frn_wq && bdev_inode &&
+	    sb_is_blkdev_sb(bdev_inode->i_sb)) {
+		mem_cgroup_track_foreign_bdev(memcg, bdev_inode->i_rdev);
+		return;
+	}
+
 	/*
 	 * Pick the slot to use.  If there is already a slot for @wb, keep
 	 * using it.  If not replace the oldest one which isn't being
@@ -3941,6 +3997,25 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
 	}
 }
 
+static void bdev_frn_flush_work(struct work_struct *work)
+{
+	struct bdev_frn_flush_ctx *ctx =
+		container_of(work, struct bdev_frn_flush_ctx, work);
+	unsigned long flags;
+
+	bdev_flush_by_dev(ctx->dev);
+
+	/*
+	 * The dirty tracking path takes frn_lock while holding mapping->i_pages.
+	 * Disable local IRQs here to avoid deadlocks with I/O completion
+	 * handlers that also take mapping->i_pages.
+	 */
+	spin_lock_irqsave(&ctx->memcg->frn_lock, flags);
+	ctx->inflight = false;
+	ctx->dev = 0;
+	spin_unlock_irqrestore(&ctx->memcg->frn_lock, flags);
+}
+
 /* issue foreign writeback flushes for recorded foreign dirtying events */
 void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
 {
@@ -3949,6 +4024,18 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
 	u64 now = jiffies_64;
 	int i;
 
+	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+		struct bdev_frn_flush_ctx *ctx = &memcg->bdev_frn[i];
+		unsigned long flags;
+
+		spin_lock_irqsave(&memcg->frn_lock, flags);
+		if (!ctx->inflight && ctx->dev) {
+			ctx->inflight = true;
+			queue_work(memcg_bdev_frn_wq, &ctx->work);
+		}
+		spin_unlock_irqrestore(&memcg->frn_lock, flags);
+	}
+
 	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
 		struct memcg_cgwb_frn *frn = &memcg->cgwb_frn[i];
 
@@ -4202,9 +4289,15 @@ static struct mem_cgroup *mem_cgroup_alloc(struct mem_cgroup *parent)
 	memcg->kmemcg_id = -1;
 #ifdef CONFIG_CGROUP_WRITEBACK
 	INIT_LIST_HEAD(&memcg->cgwb_list);
-	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++)
+	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+		struct bdev_frn_flush_ctx *ctx = &memcg->bdev_frn[i];
+
 		memcg->cgwb_frn[i].done =
 			__WB_COMPLETION_INIT(&memcg_cgwb_frn_waitq);
+		INIT_WORK(&ctx->work, bdev_frn_flush_work);
+		ctx->memcg = memcg;
+	}
+	spin_lock_init(&memcg->frn_lock);
 #endif
 	lru_gen_init_memcg(memcg);
 	return memcg;
@@ -4386,8 +4479,10 @@ static void mem_cgroup_css_free(struct cgroup_subsys_state *css)
 	int __maybe_unused i;
 
 #ifdef CONFIG_CGROUP_WRITEBACK
-	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++)
+	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
 		wb_wait_for_completion(&memcg->cgwb_frn[i].done);
+		flush_work(&memcg->bdev_frn[i].work);
+	}
 #endif
 	if (cgroup_subsys_on_dfl(memory_cgrp_subsys) && !cgroup_memory_nosocket)
 		static_branch_dec(&memcg_sockets_enabled_key);
@@ -5703,6 +5798,12 @@ int __init mem_cgroup_init(void)
 	memcg_wq = alloc_workqueue("memcg", WQ_PERCPU, 0);
 	WARN_ON(!memcg_wq);
 
+#ifdef CONFIG_CGROUP_WRITEBACK
+	memcg_bdev_frn_wq = alloc_workqueue("memcg_bdev_frn_flusher",
+					    WQ_UNBOUND | WQ_MEM_RECLAIM, 0);
+	WARN_ON(!memcg_bdev_frn_wq);
+#endif
+
 	for_each_possible_cpu(cpu) {
 		INIT_WORK(&per_cpu_ptr(&memcg_stock, cpu)->work,
 			  drain_local_memcg_stock);
-- 
2.39.5


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH v6 3/3] writeback: record bdev targets in foreign writeback tracepoints
  2026-09-17  6:57 [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
  2026-09-17  6:57 ` [PATCH v6 1/3] block: introduce bdev_flush_by_dev() Julian Sun
  2026-09-17  7:01 ` [PATCH v6 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
@ 2026-09-17  7:01 ` Julian Sun
  2026-09-17 22:53 ` [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Andrew Morton
  2026-09-21 11:55 ` [PATCH v7 " Julian Sun
  4 siblings, 0 replies; 10+ messages in thread
From: Julian Sun @ 2026-09-17  7:01 UTC (permalink / raw)
  To: linux-block, cgroups, linux-mm, linux-fsdevel
  Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
	willy, jack, tj, akpm

Add dev to track_foreign_dirty and flush_foreign so traces can
distinguish foreign-wb tracking from the bdev-only path. Append the
field to each event and print it as major:minor; zero denotes the
normal foreign-wb path.

Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
 include/trace/events/writeback.h | 22 ++++++++++++++--------
 mm/memcontrol.c                  |  9 ++++++---
 2 files changed, 20 insertions(+), 11 deletions(-)

diff --git a/include/trace/events/writeback.h b/include/trace/events/writeback.h
index 13ee076ccd16..ffd1b8df4231 100644
--- a/include/trace/events/writeback.h
+++ b/include/trace/events/writeback.h
@@ -273,9 +273,9 @@ TRACE_EVENT(inode_switch_wbs,
 
 TRACE_EVENT(track_foreign_dirty,
 
-	TP_PROTO(struct folio *folio, struct bdi_writeback *wb),
+	TP_PROTO(struct folio *folio, struct bdi_writeback *wb, dev_t dev),
 
-	TP_ARGS(folio, wb),
+	TP_ARGS(folio, wb, dev),
 
 	TP_STRUCT__entry(
 		__array(char,		name, 32)
@@ -284,6 +284,7 @@ TRACE_EVENT(track_foreign_dirty,
 		__field(u64,		cgroup_ino)
 		__field(u64,		page_cgroup_ino)
 		__field(unsigned int,	memcg_id)
+		__field(dev_t,		dev)
 	),
 
 	TP_fast_assign(
@@ -295,34 +296,37 @@ TRACE_EVENT(track_foreign_dirty,
 		__entry->ino		= inode ? inode->i_ino : 0;
 		__entry->memcg_id	= wb->memcg_css->id;
 		__entry->cgroup_ino	= __trace_wb_assign_cgroup(wb);
+		__entry->dev		= dev;
 
 		rcu_read_lock();
 		__entry->page_cgroup_ino = cgroup_ino(folio_memcg(folio)->css.cgroup);
 		rcu_read_unlock();
 	),
 
-	TP_printk("bdi %s[%llu]: ino=%llu memcg_id=%u cgroup_ino=%llu page_cgroup_ino=%llu",
+	TP_printk("bdi %s[%llu]: ino=%llu memcg_id=%u cgroup_ino=%llu page_cgroup_ino=%llu dev=%u:%u",
 		__entry->name,
 		__entry->bdi_id,
 		__entry->ino,
 		__entry->memcg_id,
 		__entry->cgroup_ino,
-		__entry->page_cgroup_ino
+		__entry->page_cgroup_ino,
+		MAJOR(__entry->dev), MINOR(__entry->dev)
 	)
 );
 
 TRACE_EVENT(flush_foreign,
 
 	TP_PROTO(struct bdi_writeback *wb, unsigned int frn_bdi_id,
-		 unsigned int frn_memcg_id),
+		 unsigned int frn_memcg_id, dev_t dev),
 
-	TP_ARGS(wb, frn_bdi_id, frn_memcg_id),
+	TP_ARGS(wb, frn_bdi_id, frn_memcg_id, dev),
 
 	TP_STRUCT__entry(
 		__array(char,		name, 32)
 		__field(u64,		cgroup_ino)
 		__field(unsigned int,	frn_bdi_id)
 		__field(unsigned int,	frn_memcg_id)
+		__field(dev_t,		dev)
 	),
 
 	TP_fast_assign(
@@ -330,13 +334,15 @@ TRACE_EVENT(flush_foreign,
 		__entry->cgroup_ino	= __trace_wb_assign_cgroup(wb);
 		__entry->frn_bdi_id	= frn_bdi_id;
 		__entry->frn_memcg_id	= frn_memcg_id;
+		__entry->dev		= dev;
 	),
 
-	TP_printk("bdi %s: cgroup_ino=%llu frn_bdi_id=%u frn_memcg_id=%u",
+	TP_printk("bdi %s: cgroup_ino=%llu frn_bdi_id=%u frn_memcg_id=%u dev=%u:%u",
 		__entry->name,
 		__entry->cgroup_ino,
 		__entry->frn_bdi_id,
-		__entry->frn_memcg_id
+		__entry->frn_memcg_id,
+		MAJOR(__entry->dev), MINOR(__entry->dev)
 	)
 );
 #endif
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 55bd5d100caa..f634874bbcef 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -3949,14 +3949,15 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
 	struct address_space *mapping = folio_mapping(folio);
 	struct inode *bdev_inode = mapping ? mapping->host : NULL;
 
-	trace_track_foreign_dirty(folio, wb);
-
 	if (memcg_bdev_frn_wq && bdev_inode &&
 	    sb_is_blkdev_sb(bdev_inode->i_sb)) {
+		trace_track_foreign_dirty(folio, wb, bdev_inode->i_rdev);
 		mem_cgroup_track_foreign_bdev(memcg, bdev_inode->i_rdev);
 		return;
 	}
 
+	trace_track_foreign_dirty(folio, wb, 0);
+
 	/*
 	 * Pick the slot to use.  If there is already a slot for @wb, keep
 	 * using it.  If not replace the oldest one which isn't being
@@ -4031,6 +4032,7 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
 		spin_lock_irqsave(&memcg->frn_lock, flags);
 		if (!ctx->inflight && ctx->dev) {
 			ctx->inflight = true;
+			trace_flush_foreign(wb, 0, 0, ctx->dev);
 			queue_work(memcg_bdev_frn_wq, &ctx->work);
 		}
 		spin_unlock_irqrestore(&memcg->frn_lock, flags);
@@ -4048,7 +4050,8 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
 		if (time_after64(frn->at, now - intv) &&
 		    atomic_read(&frn->done.cnt) == 1) {
 			frn->at = 0;
-			trace_flush_foreign(wb, frn->bdi_id, frn->memcg_id);
+			trace_flush_foreign(wb, frn->bdi_id, frn->memcg_id,
+					    0);
 			cgroup_writeback_by_id(frn->bdi_id, frn->memcg_id,
 					       WB_REASON_FOREIGN_FLUSH,
 					       &frn->done);
-- 
2.39.5


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* Re: [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately
  2026-09-17  6:57 [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
                   ` (2 preceding siblings ...)
  2026-09-17  7:01 ` [PATCH v6 3/3] writeback: record bdev targets in foreign writeback tracepoints Julian Sun
@ 2026-09-17 22:53 ` Andrew Morton
  2026-09-18  4:57   ` Julian Sun
  2026-09-21 11:55 ` [PATCH v7 " Julian Sun
  4 siblings, 1 reply; 10+ messages in thread
From: Andrew Morton @ 2026-09-17 22:53 UTC (permalink / raw)
  To: Julian Sun
  Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
	mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
	tj

On Thu, 17 Sep 2026 14:57:56 +0800 Julian Sun <sunjunchao@bytedance.com> wrote:

> Hi,
> 
> This series avoids owner-wide foreign writeback triggered by dirtying
> shared bdev inodes. Instead, it tracks bdev targets separately and
> schedules writeback of their mappings.
> 
> Problem
> =======
> 
> Bdev inodes are commonly shared by many memcgs. Dirtying their pages
> can cause the owner wb to be tracked as foreign, triggering writeback
> of the entire wb rather than just the bdev mapping.
> 
> In production, almost all the foreign dirtying we traced came from bdev
> inodes. Across our observations, a bdev mapping had as little as about
> 10 MiB of dirty pages, yet foreign flushes targeted the entire owner
> wb. On one machine, we observed 599 such works queued on one wb, with
> an average request budget of about 43 GiB per work.
> 
> Measured impact
> ===============
> 
> The synthetic test used a VM with 4 vCPUs, 8 GiB RAM, ext4 and cgroup
> v2, with device write bandwidth capped at 200 MiB/s. Background fio
> repeatedly overwrote a 256 MiB file using buffered I/O, while
> foreground fio performed sequential direct writes. Four memcgs
> generated metadata and buffered writes to trigger foreign flushes.
> 
> We measured completed block writes to the background fio file,
> including final sync, and foreground fio bandwidth, latency and
> completion time.
> 
> Across three runs per kernel, mean background file write I/O decreased
> by 72.7%, foreground bandwidth increased by 23.0%, and completion time
> decreased by 18.8%.
> 
> Metric                         Baseline    Patched     Change
> Background file write I/O (GiB)    6.154      1.680     -72.7%
> Foreground bandwidth (MiB/s)      150.29     184.81     +23.0%
> Foreground completion time (s)   109.203     88.652     -18.8%
> Foreground mean latency (ms)      212.84     172.74     -18.8%

Thanks.

As I understand it, this is basically ext4-specific.

I don't think btrfs or xfs mess with the bdev address_space at all? 
But google tells me that "roughly 70% to 80% of all Linux machines use
ext4 as their primary or root file system", so there is that.

> Approach
> ========
> 
> Following Jan's suggestion, this series records foreign bdev targets
> separately and flushes their mappings instead of their owner wbs. This
> preserves a way for dirty throttling to initiate bdev writeback while
> avoiding owner-wide writeback triggered by these records. The existing
> foreign-wb mechanism remains unchanged for other inodes.
> 
> Tracking is best effort: each memcg keeps a bounded set of device
> numbers without persistent device or inode references. Closed devices
> and devices with a busy open_mutex are skipped. Device removal and
> device-number reuse may race with lookup.

This part looks plain nasty.  Why are we messing with dev_t's and
risking these races?

At the very least, this description should explain the reasoning behind
this decision at some length.

Surely it's cleaner and safer to grab a ref on something (the bdev
inode?) and hang onto that object.  Use it for these operations, let it
go at the appropriate time.  Clearly there's something wrong with that
approach, but what?

> Flushes run asynchronously on a dedicated workqueue to isolate these
> frequently triggered tasks from existing workqueue users. If workqueue
> allocation fails, the existing foreign-wb mechanism remains in use.

Unclear what this means.  If a kmalloc/etc fails then we fall back to
the current (mainline) behavior?  Fair enough, failure of small
kmallocs are so rare.  The main problem is testing the failure-path
code!


Also, and most importantly, what the heck is "frn"?  Would the world
end if you did s/frn/foreign/g?


^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately
  2026-09-17 22:53 ` [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Andrew Morton
@ 2026-09-18  4:57   ` Julian Sun
  2026-09-18  5:12     ` Andrew Morton
  0 siblings, 1 reply; 10+ messages in thread
From: Julian Sun @ 2026-09-18  4:57 UTC (permalink / raw)
  To: Andrew Morton
  Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
	mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
	tj

On 9/18/26 6:53 AM, Andrew Morton wrote:
> On Thu, 17 Sep 2026 14:57:56 +0800 Julian Sun <sunjunchao@bytedance.com> wrote:
> 
>> Hi,
>>
>>
>> Measured impact
>> ===============
>>
>> The synthetic test used a VM with 4 vCPUs, 8 GiB RAM, ext4 and cgroup
>> v2, with device write bandwidth capped at 200 MiB/s. Background fio
>> repeatedly overwrote a 256 MiB file using buffered I/O, while
>> foreground fio performed sequential direct writes. Four memcgs
>> generated metadata and buffered writes to trigger foreign flushes.
>>
>> We measured completed block writes to the background fio file,
>> including final sync, and foreground fio bandwidth, latency and
>> completion time.
>>
>> Across three runs per kernel, mean background file write I/O decreased
>> by 72.7%, foreground bandwidth increased by 23.0%, and completion time
>> decreased by 18.8%.
>>
>> Metric                         Baseline    Patched     Change
>> Background file write I/O (GiB)    6.154      1.680     -72.7%
>> Foreground bandwidth (MiB/s)      150.29     184.81     +23.0%
>> Foreground completion time (s)   109.203     88.652     -18.8%
>> Foreground mean latency (ms)      212.84     172.74     -18.8%
> 
> Thanks.
> 
> As I understand it, this is basically ext4-specific.
> 
> I don't think btrfs or xfs mess with the bdev address_space at all? 
> But google tells me that "roughly 70% to 80% of all Linux machines use
> ext4 as their primary or root file system", so there is that.

Yes, most of our systems still use ext4.
> 
>> Approach
>> ========
>>
>> Following Jan's suggestion, this series records foreign bdev targets
>> separately and flushes their mappings instead of their owner wbs. This
>> preserves a way for dirty throttling to initiate bdev writeback while
>> avoiding owner-wide writeback triggered by these records. The existing
>> foreign-wb mechanism remains unchanged for other inodes.
>>
>> Tracking is best effort: each memcg keeps a bounded set of device
>> numbers without persistent device or inode references. Closed devices
>> and devices with a busy open_mutex are skipped. Device removal and
>> device-number reuse may race with lookup.
> 
> This part looks plain nasty.  Why are we messing with dev_t's and
> risking these races?
> 
> At the very least, this description should explain the reasoning behind
> this decision at some length.
> 
> Surely it's cleaner and safer to grab a ref on something (the bdev
> inode?) and hang onto that object.  Use it for these operations, let it
> go at the appropriate time.  Clearly there's something wrong with that
> approach, but what?

This follows the existing foreign-writeback mechanism's best-effort design:
it records IDs without holding any references and tolerates occasional
stale hints. I followed the same approach for bdev tracking by recording
dev_t.

Taking a device reference would introduce additional lifetime management.
If we acquire a reference during foreign tracking but the memcg does not
enter dirty throttling for a long time, no foreign flush would consume
the record and release the reference, then delaying final device-object
release. A timeout mechanism would only limit this delay, not eliminate it.

There is also a locking issue: dropping a record happens under mapping->i_pages,
but dropping the last reference may sleep. We would therefore need to restructure
the release path, adding concurrency and lifetime-management complexity.

Recording only dev_t follows the existing mechanism's best-effort design.
Device-number reuse may cause a stale record to trigger writeback of the new device's
own mapping. In my view, this occasional unnecessary writeback is an acceptable
trade-off for avoiding the complexity above.
> 
>> Flushes run asynchronously on a dedicated workqueue to isolate these
>> frequently triggered tasks from existing workqueue users. If workqueue
>> allocation fails, the existing foreign-wb mechanism remains in use.
> 
> Unclear what this means.  If a kmalloc/etc fails then we fall back to
> the current (mainline) behavior?  Fair enough, failure of small
> kmallocs are so rare.  The main problem is testing the failure-path
> code!

Yes, if allocation of memcg_bdev_frn_wq in mem_cgroup_init() fails,
foreign writeback falls back to the existing mechanism. The workqueue is 
allocated once during initialization. The new tracking and work-
queueing paths use preallocated per-memcg slots and perform no dynamic
allocation at runtime. please see patch 2 for more details.

I tested this failure path by forcing the workqueue pointer to NULL
and the results were as expected.> 
> 
> Also, and most importantly, what the heck is "frn"?  Would the world
> end if you did s/frn/foreign/g?

Emmm, the naming follows the existing code. I'll rename the new uses
of "frn" to "foreign" in the next version. Renaming the existing code
would be a separate cleanup patch; I'd prefer not to mix that into this
fix series.

Thanks,
-- 
Julian Sun <sunjunchao@bytedance.com>

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately
  2026-09-18  4:57   ` Julian Sun
@ 2026-09-18  5:12     ` Andrew Morton
  0 siblings, 0 replies; 10+ messages in thread
From: Andrew Morton @ 2026-09-18  5:12 UTC (permalink / raw)
  To: Julian Sun
  Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
	mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
	tj

On Fri, 18 Sep 2026 12:57:03 +0800 Julian Sun <sunjunchao@bytedance.com> wrote:

> > Also, and most importantly, what the heck is "frn"?  Would the world
> > end if you did s/frn/foreign/g?
> 
> Emmm, the naming follows the existing code. I'll rename the new uses
> of "frn" to "foreign" in the next version. Renaming the existing code
> would be a separate cleanup patch; I'd prefer not to mix that into this
> fix series.

Well crap.  Let's stick with "frn" then.  Renaming everything can be
argued about much later.

Shall consider the rest of your email soon, thanks.

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH v6 1/3] block: introduce bdev_flush_by_dev()
  2026-09-17  6:57 ` [PATCH v6 1/3] block: introduce bdev_flush_by_dev() Julian Sun
@ 2026-09-20 12:37   ` Tejun Heo
  0 siblings, 0 replies; 10+ messages in thread
From: Tejun Heo @ 2026-09-20 12:37 UTC (permalink / raw)
  To: Julian Sun
  Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
	mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
	tj, akpm

Hello, Julian.

On Thu, Sep 17, 2026 at 02:57:57PM +0800, Julian Sun wrote:
> +void bdev_flush_by_dev(dev_t dev)
> +{
> +	struct block_device *bdev;
> +
> +	bdev = blkdev_get_no_open(dev, false);
> +	if (!bdev)
> +		return;
> +
> +	/* Prevent the device from closing between the check and writeback. */
> +	if (mutex_trylock(&bdev->bd_disk->open_mutex)) {
> +		if (atomic_read(&bdev->bd_openers))
> +			sync_blockdev_nowait(bdev);
> +		mutex_unlock(&bdev->bd_disk->open_mutex);
> +	}
> +	blkdev_put_no_open(bdev);
> +}

This holds the disk's open_mutex across the whole writepages pass, so every
throttle-triggered flush blocks opens, closes and partition rescans on the
disk for as long as submission takes, which can be a while on a congested
device. bdev_release() syncs before taking the mutex for the same reason.

I don't think the mutex or the openers check is needed. If the device is
closed underneath, the last close syncs and truncates the mapping and
truncate waits for writeback, so a racing no-wait flush either finishes
first or finds nothing left. Holding just the reference from
blkdev_get_no_open() should be enough.

> +void bdev_flush_by_dev(dev_t dev);

The neighbors have !CONFIG_BLOCK stubs. The only caller is under
CGROUP_WRITEBACK so it builds either way, just noting for consistency.

Thanks.

--
tejun

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH v6 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
  2026-09-17  7:01 ` [PATCH v6 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
@ 2026-09-20 12:37   ` Tejun Heo
  0 siblings, 0 replies; 10+ messages in thread
From: Tejun Heo @ 2026-09-20 12:37 UTC (permalink / raw)
  To: Julian Sun
  Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
	mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
	tj, akpm

Hello, Julian.

On Thu, Sep 17, 2026 at 03:01:18PM +0800, Julian Sun wrote:
> Bdev foreign flushes can be triggered frequently. Use a dedicated
> workqueue to avoid interfering with tasks on existing workqueues. If
> allocation fails, retain the existing foreign-wb path.

Work items on an unbound workqueue don't interfere with each other beyond
max_active, so this isn't a reason for a dedicated workqueue. Wanting a
separate max_active pool for up to four long-blocking items per memcg is,
if that's the intent. Can you update the rationale?

> Use fixed per-memcg slots, with a shared lock protecting device records
> and in-flight state. Tracking is best effort: record only dev_t without
> holding device or inode references. Skip closed devices or devices whose
> open_mutex is busy. Device removal and device-number reuse may race with
> lookup.

Andrew asked why dev_t rather than a reference. The reasons you gave in the
thread belong here: the record is written under i_pages with IRQs off,
dropping a bdev reference can sleep, and a memcg that never throttles would
pin the device indefinitely.

It'd also help to say who this affects. Only filesystems that keep metadata
in the bdev page cache, so ext4 and the other buffer_head users but not
xfs, btrfs or f2fs. And that the decision is still blind to how much of the
memcg's dirty memory is actually in the bdev, so one dirty bitmap block
still flushes the whole mapping each time the memcg throttles.

> Fixes: 97b27821b485 ("writeback, memcg: Implement foreign dirty flushing")

This is new development rather than a fix. Maybe drop the tag? Patch 1
doesn't carry one, so a stable pick of this patch alone wouldn't build
anyway.

> +static void mem_cgroup_track_foreign_bdev(struct mem_cgroup *memcg, dev_t dev)

This lands above the "Foreign dirty flushing" overview, which now says
nothing about the bdev path. Can you move it below the overview and add a
sentence there?

> +	/*
> +	 * Tracking is best effort, so losing hints when all slots are occupied
> +	 * is expected.
> +	 */
> +	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
> +		ctx = &memcg->bdev_frn[i];
> +		if (ctx->inflight)
> +			continue;
> +		if (slot < 0)
> +			slot = i;
> +		if (!ctx->dev) {
> +			slot = i;
> +			break;
> +		}
> +	}

When all four are occupied this always evicts slot 0, and the comment says
the new hint is dropped, which isn't what happens.

Why not follow the cgwb_frn style? A timestamp per slot gives oldest-first
replacement here, expiry in mem_cgroup_flush_foreign() so a record from
long ago doesn't fire a flush when the memcg finally throttles, and
re-arming when the mapping is dirtied while a flush is in flight, which the
worker currently forgets when it clears the slot. It could also drop
frn_lock. An atomic in-flight flag on the work gives the same guarantee
done.cnt does, and a torn read of dev or the stamp only costs a wrong
best-effort flush, which is already the deal.

> +	if (memcg_bdev_frn_wq && bdev_inode &&
> +	    sb_is_blkdev_sb(bdev_inode->i_sb)) {
> +		mem_cgroup_track_foreign_bdev(memcg, bdev_inode->i_rdev);
> +		return;
> +	}

Patch 3 moves the trace call below this. Might as well put it there here.

> +	for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
> +		struct bdev_frn_flush_ctx *ctx = &memcg->bdev_frn[i];
> +		unsigned long flags;
> +
> +		spin_lock_irqsave(&memcg->frn_lock, flags);
> +		if (!ctx->inflight && ctx->dev) {
> +			ctx->inflight = true;
> +			queue_work(memcg_bdev_frn_wq, &ctx->work);
> +		}
> +		spin_unlock_irqrestore(&memcg->frn_lock, flags);
> +	}

If the lock stays, one acquisition around the loop is enough.

> +#ifdef CONFIG_CGROUP_WRITEBACK
> +	memcg_bdev_frn_wq = alloc_workqueue("memcg_bdev_frn_flusher",
> +					    WQ_UNBOUND | WQ_MEM_RECLAIM, 0);
> +	WARN_ON(!memcg_bdev_frn_wq);
> +#endif

WQ_MEM_RECLAIM is for work items that memory reclaim depends on to make
forward progress. Nothing waits on these flushes. The throttled task just
sleeps and re-evaluates, and the bdi flusher remains the path that's
guaranteed to make progress. Can you drop the flag?

Thanks.

--
tejun

^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH v7 0/3] memcg,writeback: flush foreign bdev mappings separately
  2026-09-17  6:57 [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
                   ` (3 preceding siblings ...)
  2026-09-17 22:53 ` [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Andrew Morton
@ 2026-09-21 11:55 ` Julian Sun
  4 siblings, 0 replies; 10+ messages in thread
From: Julian Sun @ 2026-09-21 11:55 UTC (permalink / raw)
  To: linux-block, cgroups, linux-mm, linux-fsdevel
  Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
	willy, jack, tj, akpm

Hi,

This series avoids owner-wide foreign writeback triggered by dirtying
shared bdev inodes. Instead, it tracks bdev targets separately and
schedules writeback of their mappings.

Problem
=======

Bdev inodes are commonly shared by many memcgs. Dirtying their pages
can cause the owner wb to be tracked as foreign, triggering writeback
of the entire wb rather than just the bdev mapping.

In production, almost all the foreign dirtying we traced came from bdev
inodes. Across our observations, a bdev mapping had as little as about
10 MiB of dirty pages, yet foreign flushes targeted the entire owner
wb. On one machine, we observed 599 such works queued on one wb, with
an average request budget of about 43 GiB per work.

Measured impact
===============

The synthetic test used a VM with 4 vCPUs, 8 GiB RAM, ext4 and cgroup
v2, with device write bandwidth capped at 200 MiB/s. Background fio
repeatedly overwrote a 256 MiB file using buffered I/O, while
foreground fio performed sequential direct writes. Four memcgs
generated metadata and buffered writes to trigger foreign flushes.

We measured completed block writes to the background fio file,
including final sync, and foreground fio bandwidth, latency and
completion time.

Across three runs per kernel, mean background file write I/O decreased
by 72.7%, foreground bandwidth increased by 23.0%, and completion time
decreased by 18.8%.

Metric                         Baseline    Patched     Change
Background file write I/O (GiB)    6.154      1.680     -72.7%
Foreground bandwidth (MiB/s)      150.29     184.81     +23.0%
Foreground completion time (s)   109.203     88.652     -18.8%
Foreground mean latency (ms)      212.84     172.74     -18.8%

How it helps
============

Frequent writeback of the file overwritten by background fio reduces
the opportunity to coalesce buffered writes in the page cache.
Several overwrites of the same page can otherwise result in a single
write of its latest contents. Flushing between updates instead writes
the same page repeatedly, generating more device I/O for the same
application writes.

This additional I/O makes the device reach its performance limit
more easily and competes with foreground I/O for device bandwidth.

The benefit is expected to be most pronounced when these conditions
occur together on the same device:

1. Limited device bandwidth, as with an HDD, makes the additional
   writeback I/O more likely to saturate the device.
2. Multiple cgroups perform buffered overwrites. Frequent foreign
   writeback reduces write coalescing and increases device I/O.
3. Concurrent direct I/O bypasses the page cache and competes directly
   with writeback I/O for device bandwidth.

Approach
========

Following Jan's suggestion, this series records foreign bdev targets
separately and flushes their mappings instead of their owner wbs. This
preserves a way for dirty throttling to initiate bdev writeback while
avoiding owner-wide writeback triggered by these records. The existing
foreign-wb mechanism remains unchanged for other inodes.

Tracking is best effort: each memcg keeps a bounded set of device
numbers without persistent device or inode references. Recording runs
under mapping->i_pages with IRQs disabled, while dropping a device
reference can sleep. A memcg that never enters dirty throttling could
also retain an unused record and pin the device object indefinitely.
Recording only dev_t avoids these lifetime constraints. Device removal
and device-number reuse may cause an unintended best-effort flush,
which is an accepted trade-off.

Writeback submission may block for a long time, with up to four work
items outstanding per memcg. A dedicated unbound workqueue provides
a separate max_active budget so these flushes do not exhaust active
slots on a shared workqueue and delay unrelated work. If allocation
fails, the existing foreign-wb mechanism remains in use.

This primarily affects filesystems that keep metadata in the bdev page
cache, such as ext4 and other buffer_head users, rather than the usual
metadata writeback paths of XFS and Btrfs. The decision still does not
account for how much of the memcg's dirty memory belongs to the bdev:
even a single dirty bitmap block can trigger writeback of the entire
bdev mapping when the memcg enters dirty throttling.

Changes since v6:
- Drop open_mutex and the openers check from bdev_flush_by_dev(), and
  add a !CONFIG_BLOCK stub.
- Follow the existing foreign-wb tracking policy: use timestamps for
  oldest-first replacement and expiry, and allow dirtying during an
  in-flight flush to re-arm the record.
- Replace frn_lock with an atomic in-flight flag.
- Drop WQ_MEM_RECLAIM and explain the separate max_active budget.
- Guard bdev work submission when workqueue allocation has failed.
- Expand the dev_t lifetime rationale and filesystem scope, update
  the foreign-dirty overview, and drop the Fixes tag.

Changes since v5:
- Fix repeated overwriting of slot 0 in patch 2 and add comments.

Changes since v4:
- Add before-and-after performance comparisons to the cover letter.

Changes since v3:
- Flush foreign-dirtied bdev inode mappings separately, as suggested
  by Jan.

Thanks.

Julian Sun (3):
  block: introduce bdev_flush_by_dev()
  memcg,writeback: flush foreign bdev mappings separately from owner wbs
  writeback: record bdev targets in foreign writeback tracepoints

 block/bdev.c                     |  18 +++++
 include/linux/blkdev.h           |   4 +
 include/linux/memcontrol.h       |  12 +++
 include/trace/events/writeback.h |  22 ++++--
 mm/memcontrol.c                  | 124 ++++++++++++++++++++++++++++---
 5 files changed, 160 insertions(+), 20 deletions(-)

-- 
2.39.5

^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-09-21 11:55 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17  6:57 [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-17  6:57 ` [PATCH v6 1/3] block: introduce bdev_flush_by_dev() Julian Sun
2026-09-20 12:37   ` Tejun Heo
2026-09-17  7:01 ` [PATCH v6 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
2026-09-20 12:37   ` Tejun Heo
2026-09-17  7:01 ` [PATCH v6 3/3] writeback: record bdev targets in foreign writeback tracepoints Julian Sun
2026-09-17 22:53 ` [PATCH v6 0/3] memcg,writeback: flush foreign bdev mappings separately Andrew Morton
2026-09-18  4:57   ` Julian Sun
2026-09-18  5:12     ` Andrew Morton
2026-09-21 11:55 ` [PATCH v7 " Julian Sun

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox