* [PATCH v7 1/3] block: introduce bdev_flush_by_dev()
2026-09-21 12:06 [PATCH v7 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
@ 2026-09-21 12:06 ` Julian Sun
2026-09-21 12:06 ` [PATCH v7 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
2026-09-21 12:06 ` [PATCH v7 3/3] writeback: record bdev targets in foreign writeback tracepoints Julian Sun
2 siblings, 0 replies; 6+ messages in thread
From: Julian Sun @ 2026-09-21 12:06 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Add bdev_flush_by_dev() to submit page-cache writeback for a device
identified by dev_t without requiring callers to hold a persistent
device reference.
A subsequent patch will use this helper to flush foreign bdev mappings.
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
block/bdev.c | 18 ++++++++++++++++++
include/linux/blkdev.h | 4 ++++
2 files changed, 22 insertions(+)
diff --git a/block/bdev.c b/block/bdev.c
index cd8323083740..5cbf01a13a33 100644
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -264,6 +264,24 @@ int sync_blockdev_nowait(struct block_device *bdev)
}
EXPORT_SYMBOL_GPL(sync_blockdev_nowait);
+/**
+ * bdev_flush_by_dev - attempt writeback of a block device's page cache
+ * @dev: target block device number
+ *
+ * May sleep while submitting writeback; does not wait for all I/O to complete.
+ */
+void bdev_flush_by_dev(dev_t dev)
+{
+ struct block_device *bdev;
+
+ bdev = blkdev_get_no_open(dev, false);
+ if (!bdev)
+ return;
+
+ sync_blockdev_nowait(bdev);
+ blkdev_put_no_open(bdev);
+}
+
/*
* Write out and wait upon all the dirty data associated with a block
* device via its mapping. Does not take the superblock lock.
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 4f7905c3412b..27afe3ff6fb5 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1710,6 +1710,7 @@ void invalidate_bdev(struct block_device *bdev);
int sync_blockdev(struct block_device *bdev);
int sync_blockdev_range(struct block_device *bdev, loff_t lstart, loff_t lend);
int sync_blockdev_nowait(struct block_device *bdev);
+void bdev_flush_by_dev(dev_t dev);
void sync_bdevs(bool wait);
void bdev_statx(const struct path *path, struct kstat *stat, u32 request_mask);
void printk_all_partitions(void);
@@ -1726,6 +1727,9 @@ static inline int sync_blockdev_nowait(struct block_device *bdev)
{
return 0;
}
+static inline void bdev_flush_by_dev(dev_t dev)
+{
+}
static inline void sync_bdevs(bool wait)
{
}
--
2.39.5
^ permalink raw reply related [flat|nested] 6+ messages in thread* [PATCH v7 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-21 12:06 [PATCH v7 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-21 12:06 ` [PATCH v7 1/3] block: introduce bdev_flush_by_dev() Julian Sun
@ 2026-09-21 12:06 ` Julian Sun
2026-09-21 17:55 ` Tejun Heo
2026-09-21 12:06 ` [PATCH v7 3/3] writeback: record bdev targets in foreign writeback tracepoints Julian Sun
2 siblings, 1 reply; 6+ messages in thread
From: Julian Sun @ 2026-09-21 12:06 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Bdev inodes are routinely shared by multiple memcgs. Recording their
owner wb as foreign can flush unrelated file data when a source memcg
enters dirty throttling.
Record bdev device numbers separately and queue work on
memcg_bdev_frn_flusher to write only their mappings. Keep the existing
foreign-wb mechanism for non-bdev inodes.
Use fixed per-memcg slots. Tracking is best effort: record only dev_t
without holding device or inode references. Records are updated under
the mapping->i_pages lock with IRQs disabled, but dropping a bdev
reference can sleep. Also, since foreign flushes are triggered by dirty
throttling, a memcg that never throttles could retain an unused record
and pin the device object indefinitely. Recording only dev_t avoids
these lifetime constraints, but device removal and device-number reuse
may race with lookup. Such races are expected under the best-effort
semantics.
Bdev writeback submission can block for a long time on congested devices,
with up to four work items outstanding per memcg. Use a dedicated unbound
workqueue to give these flushes a separate max_active budget, so they do
not exhaust a shared workqueue's active slots and delay unrelated work.
If allocation fails, retain the existing foreign-wb path.
This primarily affects filesystems that keep metadata in the bdev page
cache, such as ext4 and other buffer_head users, rather than the usual
metadata writeback paths of XFS and Btrfs. The decision still does not
account for how much of the memcg's dirty memory belongs to the bdev.
Even a single dirty bitmap block can trigger writeback of the entire
bdev mapping when the memcg enters dirty throttling.
Suggested-by: Jan Kara <jack@suse.cz>
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
include/linux/memcontrol.h | 12 ++++
mm/memcontrol.c | 117 +++++++++++++++++++++++++++++++++----
2 files changed, 119 insertions(+), 10 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 7d1c0ce189a8..30d0c5d0da14 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -176,6 +176,17 @@ struct memcg_cgwb_frn {
struct wb_completion done; /* tracks in-flight foreign writebacks */
};
+/*
+ * No extra memcg reference is taken: this work is embedded in the memcg,
+ * and mem_cgroup_css_free() waits for it to finish before freeing the memcg.
+ */
+struct bdev_frn_flush_ctx {
+ struct work_struct work;
+ dev_t dev; /* for bdev inode, device hint */
+ u64 at; /* last recorded dirtying time in jiffies */
+ atomic_t inflight;
+};
+
/*
* Bucket for arbitrarily byte-sized objects charged to a memory
* cgroup. The bucket can be reparented in one piece when the cgroup
@@ -279,6 +290,7 @@ struct mem_cgroup {
#ifdef CONFIG_CGROUP_WRITEBACK
struct wb_domain cgwb_domain;
struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
+ struct bdev_frn_flush_ctx bdev_frn[MEMCG_CGWB_FRN_CNT];
#endif
#ifdef CONFIG_LRU_GEN_WALKS_MMU
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 856a7d07586c..494540cf6278 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -39,6 +39,7 @@
#include <linux/smp.h>
#include <linux/page-flags.h>
#include <linux/backing-dev.h>
+#include <linux/blkdev.h>
#include <linux/bit_spinlock.h>
#include <linux/rcupdate.h>
#include <linux/limits.h>
@@ -104,6 +105,7 @@ static struct kmem_cache *memcg_pn_cachep;
#ifdef CONFIG_CGROUP_WRITEBACK
static DECLARE_WAIT_QUEUE_HEAD(memcg_cgwb_frn_waitq);
+static struct workqueue_struct *memcg_bdev_frn_wq __ro_after_init;
#endif
static inline bool task_is_dying(void)
@@ -3875,20 +3877,74 @@ void mem_cgroup_wb_stats(struct bdi_writeback *wb, unsigned long *pfilepages,
* most recent foreign dirtying events and initiating remote flushes on
* them when local writeback isn't enough to keep the memory clean enough.
*
- * The following two functions implement such mechanism. When a foreign
- * page - a page whose memcg and writeback ownerships don't match - is
- * dirtied, mem_cgroup_track_foreign_dirty() records the inode owning
- * bdi_writeback on the page owning memcg. When balance_dirty_pages()
- * decides that the memcg needs to sleep due to high dirty ratio, it calls
+ * For non-bdev inodes, when a foreign page - a page whose memcg and
+ * writeback ownerships don't match - is dirtied,
+ * mem_cgroup_track_foreign_dirty() records the inode owning bdi_writeback
+ * on the page owning memcg. When balance_dirty_pages() decides that the
+ * memcg needs to sleep due to high dirty ratio, it calls
* mem_cgroup_flush_foreign() which queues writeback on the recorded
* foreign bdi_writebacks which haven't expired. Both the numbers of
* recorded bdi_writebacks and concurrent in-flight foreign writebacks are
* limited to MEMCG_CGWB_FRN_CNT.
*
- * The mechanism only remembers IDs and doesn't hold any object references.
- * As being wrong occasionally doesn't matter, updates and accesses to the
- * records are lockless and racy.
+ * Bdev inodes are commonly shared by many memcgs. Flushing their owner wb
+ * can write unrelated file data, so record device numbers separately and
+ * flush only the bdev mappings. Tracking is best effort: record only dev_t
+ * without holding device or inode references, so it may race with device
+ * removal and re-addition. A resulting flush of the wrong device is
+ * acceptable under these best-effort semantics.
+ *
+ * These wb/bdev_inodes records only remember IDs and don't hold any object
+ * references. As being wrong occasionally doesn't matter, updates and
+ * accesses to the records are lockless and racy.
*/
+
+static void mem_cgroup_track_foreign_bdev(struct mem_cgroup *memcg, dev_t dev)
+{
+ struct bdev_frn_flush_ctx *ctx;
+ int i;
+ int oldest = -1;
+ u64 now = get_jiffies_64();
+ u64 oldest_at = now;
+
+ /*
+ * Pick the slot to use. If there is already a slot for @dev, keep
+ * using it. If not replace the oldest one which isn't being
+ * written out.
+ */
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+ ctx = &memcg->bdev_frn[i];
+ if (ctx->dev == dev)
+ break;
+ if (atomic_read(&ctx->inflight))
+ continue;
+ if (time_before64(ctx->at, oldest_at)) {
+ oldest = i;
+ oldest_at = ctx->at;
+ }
+ }
+
+ if (i < MEMCG_CGWB_FRN_CNT) {
+ /*
+ * Re-using an existing one. Update timestamp lazily to
+ * avoid making the cacheline hot. We want them to be
+ * reasonably up-to-date and significantly shorter than
+ * dirty_expire_interval as that's what expires the record.
+ * Use the shorter of 1s and dirty_expire_interval / 8.
+ */
+ unsigned long update_intv =
+ min_t(unsigned long, HZ,
+ msecs_to_jiffies(dirty_expire_interval * 10) / 8);
+
+ if (time_before64(ctx->at, now - update_intv))
+ ctx->at = now;
+ } else if (oldest >= 0) {
+ /* replace the oldest free one */
+ memcg->bdev_frn[oldest].dev = dev;
+ memcg->bdev_frn[oldest].at = now;
+ }
+}
+
void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
struct bdi_writeback *wb)
{
@@ -3898,6 +3954,14 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
u64 oldest_at = now;
int oldest = -1;
int i;
+ struct address_space *mapping = folio_mapping(folio);
+ struct inode *bdev_inode = mapping ? mapping->host : NULL;
+
+ if (memcg_bdev_frn_wq && bdev_inode &&
+ sb_is_blkdev_sb(bdev_inode->i_sb)) {
+ mem_cgroup_track_foreign_bdev(memcg, bdev_inode->i_rdev);
+ return;
+ }
trace_track_foreign_dirty(folio, wb);
@@ -3941,6 +4005,16 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
}
}
+static void bdev_frn_flush_work(struct work_struct *work)
+{
+ struct bdev_frn_flush_ctx *ctx =
+ container_of(work, struct bdev_frn_flush_ctx, work);
+
+ bdev_flush_by_dev(ctx->dev);
+
+ atomic_set(&ctx->inflight, 0);
+}
+
/* issue foreign writeback flushes for recorded foreign dirtying events */
void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
{
@@ -3949,6 +4023,16 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
u64 now = jiffies_64;
int i;
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+ struct bdev_frn_flush_ctx *ctx = &memcg->bdev_frn[i];
+
+ if (memcg_bdev_frn_wq && time_after64(ctx->at, now - intv) &&
+ atomic_cmpxchg(&ctx->inflight, 0, 1) == 0) {
+ ctx->at = 0;
+ queue_work(memcg_bdev_frn_wq, &ctx->work);
+ }
+ }
+
for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
struct memcg_cgwb_frn *frn = &memcg->cgwb_frn[i];
@@ -4202,9 +4286,14 @@ static struct mem_cgroup *mem_cgroup_alloc(struct mem_cgroup *parent)
memcg->kmemcg_id = -1;
#ifdef CONFIG_CGROUP_WRITEBACK
INIT_LIST_HEAD(&memcg->cgwb_list);
- for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++)
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+ struct bdev_frn_flush_ctx *ctx = &memcg->bdev_frn[i];
+
memcg->cgwb_frn[i].done =
__WB_COMPLETION_INIT(&memcg_cgwb_frn_waitq);
+ INIT_WORK(&ctx->work, bdev_frn_flush_work);
+ atomic_set(&ctx->inflight, 0);
+ }
#endif
lru_gen_init_memcg(memcg);
return memcg;
@@ -4386,8 +4475,10 @@ static void mem_cgroup_css_free(struct cgroup_subsys_state *css)
int __maybe_unused i;
#ifdef CONFIG_CGROUP_WRITEBACK
- for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++)
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
wb_wait_for_completion(&memcg->cgwb_frn[i].done);
+ flush_work(&memcg->bdev_frn[i].work);
+ }
#endif
if (cgroup_subsys_on_dfl(memory_cgrp_subsys) && !cgroup_memory_nosocket)
static_branch_dec(&memcg_sockets_enabled_key);
@@ -5703,6 +5794,12 @@ int __init mem_cgroup_init(void)
memcg_wq = alloc_workqueue("memcg", WQ_PERCPU, 0);
WARN_ON(!memcg_wq);
+#ifdef CONFIG_CGROUP_WRITEBACK
+ memcg_bdev_frn_wq = alloc_workqueue("memcg_bdev_frn_flusher",
+ WQ_UNBOUND, 0);
+ WARN_ON(!memcg_bdev_frn_wq);
+#endif
+
for_each_possible_cpu(cpu) {
INIT_WORK(&per_cpu_ptr(&memcg_stock, cpu)->work,
drain_local_memcg_stock);
--
2.39.5
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH v7 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-21 12:06 ` [PATCH v7 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
@ 2026-09-21 17:55 ` Tejun Heo
2026-09-22 9:52 ` Julian Sun
0 siblings, 1 reply; 6+ messages in thread
From: Tejun Heo @ 2026-09-21 17:55 UTC (permalink / raw)
To: Julian Sun
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
tj, akpm
Hello, Julian.
This is an AI review. The series was built here, and the bdev reference
across a racing last close and disk removal, the slot races, and the
css_free ordering were traced and look correct. Two behavioral notes and
some description and comment fixups.
On Mon, Sep 21, 2026 at 08:06:57PM +0800, Julian Sun wrote:
> Use fixed per-memcg slots. Tracking is best effort: record only dev_t
The slot policy this version adopted, oldest-first replacement, expiry
after dirty_expire_interval, and re-arming when the mapping is dirtied
while a flush is in flight, isn't stated anywhere. Can you add a sentence?
> If allocation fails, retain the existing foreign-wb path.
Andrew found this unclear on v6 and the subject is still implied. Something
like "If the workqueue can't be allocated at boot, bdev inodes keep using
the existing foreign-wb path."?
> This primarily affects filesystems that keep metadata in the bdev page
> cache, such as ext4 and other buffer_head users, rather than the usual
The branch keys on sb_is_blkdev_sb(), so buffered writers of raw block
devices are covered too whenever the bdev inode's wb belongs to another
cgroup. Worth a mention.
> +struct bdev_frn_flush_ctx {
> + struct work_struct work;
> + dev_t dev; /* for bdev inode, device hint */
> + u64 at; /* last recorded dirtying time in jiffies */
> + atomic_t inflight;
> +};
It sits next to memcg_cgwb_frn and is memcg's, so maybe memcg_bdev_frn?
"for bdev inode, device hint" doesn't say what the value is. Following the
sibling's comments, "dev_t of the foreign bdev inode", and @inflight could
use one too.
> + * For non-bdev inodes, when a foreign page - a page whose memcg and
> + * writeback ownerships don't match - is dirtied,
> + * mem_cgroup_track_foreign_dirty() records the inode owning bdi_writeback
> + * on the page owning memcg. When balance_dirty_pages() decides that the
> + * memcg needs to sleep due to high dirty ratio, it calls
> * mem_cgroup_flush_foreign() which queues writeback on the recorded
> * foreign bdi_writebacks which haven't expired. Both the numbers of
> * recorded bdi_writebacks and concurrent in-flight foreign writebacks are
> * limited to MEMCG_CGWB_FRN_CNT.
The flush sentence and the slot limit apply to both paths now, so the "For
non-bdev inodes" scoping reads wrong. Maybe keep the original paragraph and
let the new one state only the difference: separate MEMCG_CGWB_FRN_CNT
slots keyed by dev_t, flushing only the bdev mapping, same expiry.
> + * These wb/bdev_inodes records only remember IDs and don't hold any object
> + * references. As being wrong occasionally doesn't matter, updates and
> + * accesses to the records are lockless and racy.
The paragraph above already says the bdev records hold no references, and
"wb/bdev_inodes records" doesn't parse. "Both kinds of records only
remember IDs ..." would do.
> + if (memcg_bdev_frn_wq && bdev_inode &&
> + sb_is_blkdev_sb(bdev_inode->i_sb)) {
> + mem_cgroup_track_foreign_bdev(memcg, bdev_inode->i_rdev);
> + return;
> + }
>
> trace_track_foreign_dirty(folio, wb);
With this patch alone, foreign dirtying of a bdev folio no longer hits
track_foreign_dirty at all and patch 3 adds the call back inside the
branch. Can you keep the tracepoint firing on the bdev path here so that
patch 3 only adds the argument? Also, @bdev_inode can't be NULL:
__folio_mark_dirty() checked folio->mapping under the same lock and
folio_account_dirtied() already dereferenced mapping->host.
> +static void bdev_frn_flush_work(struct work_struct *work)
> +{
> + struct bdev_frn_flush_ctx *ctx =
> + container_of(work, struct bdev_frn_flush_ctx, work);
> +
> + bdev_flush_by_dev(ctx->dev);
> +
> + atomic_set(&ctx->inflight, 0);
> +}
The in-flight window is now just the submission of the bdev mapping, and
mem_cgroup_flush_foreign() runs on every throttle-loop iteration, so a
memcg that keeps dirtying and throttling requeues the flush back to back.
A folio the flush redirties, e.g. a buffer locked by a jbd2 checkpoint
write, re-stamps the slot on its own. The old record stayed in flight for
the whole owner-wb work. With a journal the write cadence of hot metadata
is still bounded by the commit interval, but for nojournal ext4 and other
buffer_head users it's bounded only by the throttle pause, for every
tenant of the filesystem while any tenant throttles. Not measured. The
description covers the target size but not the cadence.
Each pass also goes through wbc_attach_fdatawrite_inode() and
wbc_detach_inode(), so it votes in the bdev inode's wb-switch detection
like the owner-wb flush did, at the higher pass rate. Whether the bdev
inode ends up switching wbs more often is an open question.
> + if (memcg_bdev_frn_wq && time_after64(ctx->at, now - intv) &&
> + atomic_cmpxchg(&ctx->inflight, 0, 1) == 0) {
Records are only created when the workqueue exists and @at stays 0
otherwise, so the wq test can't be false with a live record. Fine as
documentation, but the cover lists it as a fix.
> + INIT_WORK(&ctx->work, bdev_frn_flush_work);
> + atomic_set(&ctx->inflight, 0);
memcg is zero-allocated, so the atomic_set() isn't needed.
On patch 1, the kerneldoc's "does not wait for all I/O to complete"
understates it, the flush waits for nothing.
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v7 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-21 17:55 ` Tejun Heo
@ 2026-09-22 9:52 ` Julian Sun
0 siblings, 0 replies; 6+ messages in thread
From: Julian Sun @ 2026-09-22 9:52 UTC (permalink / raw)
To: Tejun Heo
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
akpm
On 9/22/26 1:55 AM, Tejun Heo wrote:
Hi, Tejun.
Thanks for your review.
>
>> +static void bdev_frn_flush_work(struct work_struct *work)
>> +{
>> + struct bdev_frn_flush_ctx *ctx =
>> + container_of(work, struct bdev_frn_flush_ctx, work);
>> +
>> + bdev_flush_by_dev(ctx->dev);
>> +
>> + atomic_set(&ctx->inflight, 0);
>> +}
>
> The in-flight window is now just the submission of the bdev mapping, and
> mem_cgroup_flush_foreign() runs on every throttle-loop iteration, so a
> memcg that keeps dirtying and throttling requeues the flush back to back.
> A folio the flush redirties, e.g. a buffer locked by a jbd2 checkpoint
> write, re-stamps the slot on its own. The old record stayed in flight for
> the whole owner-wb work. With a journal the write cadence of hot metadata
> is still bounded by the commit interval, but for nojournal ext4 and other
> buffer_head users it's bounded only by the throttle pause, for every
> tenant of the filesystem while any tenant throttles. Not measured. The
> description covers the target size but not the cadence.
The description, comment and tracepoint suggestions make sense;
I'll address those.
Dirtying or redirtying the same device during a flush can refresh its
recorded timestamp. This is expected. Another flush for the same
slot cannot be queued until bdev_flush_by_dev() returns and inflight
is cleared. A subsequent call to mem_cgroup_flush_foreign() can queue
the work again if the record is still valid and inflight has been cleared.>
> Each pass also goes through wbc_attach_fdatawrite_inode() and
> wbc_detach_inode(), so it votes in the bdev inode's wb-switch detection
> like the owner-wb flush did, at the higher pass rate. Whether the bdev
> inode ends up switching wbs more often is an open question.
>
>> + if (memcg_bdev_frn_wq && time_after64(ctx->at, now - intv) &&
>> + atomic_cmpxchg(&ctx->inflight, 0, 1) == 0) {
>
> Records are only created when the workqueue exists and @at stays 0
> otherwise, so the wq test can't be false with a live record. Fine as
> documentation, but the cover lists it as a fix.
>
>> + INIT_WORK(&ctx->work, bdev_frn_flush_work);
>> + atomic_set(&ctx->inflight, 0);
>
> memcg is zero-allocated, so the atomic_set() isn't needed.
>
> On patch 1, the kerneldoc's "does not wait for all I/O to complete"
> understates it, the flush waits for nothing.
>
> Thanks.
>
> --
> tejun
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v7 3/3] writeback: record bdev targets in foreign writeback tracepoints
2026-09-21 12:06 [PATCH v7 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-21 12:06 ` [PATCH v7 1/3] block: introduce bdev_flush_by_dev() Julian Sun
2026-09-21 12:06 ` [PATCH v7 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
@ 2026-09-21 12:06 ` Julian Sun
2 siblings, 0 replies; 6+ messages in thread
From: Julian Sun @ 2026-09-21 12:06 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Add dev to track_foreign_dirty and flush_foreign so traces can
distinguish foreign-wb tracking from the bdev-only path. Append the
field to each event and print it as major:minor; zero denotes the
normal foreign-wb path.
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
include/trace/events/writeback.h | 22 ++++++++++++++--------
mm/memcontrol.c | 7 +++++--
2 files changed, 19 insertions(+), 10 deletions(-)
diff --git a/include/trace/events/writeback.h b/include/trace/events/writeback.h
index 13ee076ccd16..ffd1b8df4231 100644
--- a/include/trace/events/writeback.h
+++ b/include/trace/events/writeback.h
@@ -273,9 +273,9 @@ TRACE_EVENT(inode_switch_wbs,
TRACE_EVENT(track_foreign_dirty,
- TP_PROTO(struct folio *folio, struct bdi_writeback *wb),
+ TP_PROTO(struct folio *folio, struct bdi_writeback *wb, dev_t dev),
- TP_ARGS(folio, wb),
+ TP_ARGS(folio, wb, dev),
TP_STRUCT__entry(
__array(char, name, 32)
@@ -284,6 +284,7 @@ TRACE_EVENT(track_foreign_dirty,
__field(u64, cgroup_ino)
__field(u64, page_cgroup_ino)
__field(unsigned int, memcg_id)
+ __field(dev_t, dev)
),
TP_fast_assign(
@@ -295,34 +296,37 @@ TRACE_EVENT(track_foreign_dirty,
__entry->ino = inode ? inode->i_ino : 0;
__entry->memcg_id = wb->memcg_css->id;
__entry->cgroup_ino = __trace_wb_assign_cgroup(wb);
+ __entry->dev = dev;
rcu_read_lock();
__entry->page_cgroup_ino = cgroup_ino(folio_memcg(folio)->css.cgroup);
rcu_read_unlock();
),
- TP_printk("bdi %s[%llu]: ino=%llu memcg_id=%u cgroup_ino=%llu page_cgroup_ino=%llu",
+ TP_printk("bdi %s[%llu]: ino=%llu memcg_id=%u cgroup_ino=%llu page_cgroup_ino=%llu dev=%u:%u",
__entry->name,
__entry->bdi_id,
__entry->ino,
__entry->memcg_id,
__entry->cgroup_ino,
- __entry->page_cgroup_ino
+ __entry->page_cgroup_ino,
+ MAJOR(__entry->dev), MINOR(__entry->dev)
)
);
TRACE_EVENT(flush_foreign,
TP_PROTO(struct bdi_writeback *wb, unsigned int frn_bdi_id,
- unsigned int frn_memcg_id),
+ unsigned int frn_memcg_id, dev_t dev),
- TP_ARGS(wb, frn_bdi_id, frn_memcg_id),
+ TP_ARGS(wb, frn_bdi_id, frn_memcg_id, dev),
TP_STRUCT__entry(
__array(char, name, 32)
__field(u64, cgroup_ino)
__field(unsigned int, frn_bdi_id)
__field(unsigned int, frn_memcg_id)
+ __field(dev_t, dev)
),
TP_fast_assign(
@@ -330,13 +334,15 @@ TRACE_EVENT(flush_foreign,
__entry->cgroup_ino = __trace_wb_assign_cgroup(wb);
__entry->frn_bdi_id = frn_bdi_id;
__entry->frn_memcg_id = frn_memcg_id;
+ __entry->dev = dev;
),
- TP_printk("bdi %s: cgroup_ino=%llu frn_bdi_id=%u frn_memcg_id=%u",
+ TP_printk("bdi %s: cgroup_ino=%llu frn_bdi_id=%u frn_memcg_id=%u dev=%u:%u",
__entry->name,
__entry->cgroup_ino,
__entry->frn_bdi_id,
- __entry->frn_memcg_id
+ __entry->frn_memcg_id,
+ MAJOR(__entry->dev), MINOR(__entry->dev)
)
);
#endif
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 494540cf6278..6c3083fcc063 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -3959,11 +3959,12 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
if (memcg_bdev_frn_wq && bdev_inode &&
sb_is_blkdev_sb(bdev_inode->i_sb)) {
+ trace_track_foreign_dirty(folio, wb, bdev_inode->i_rdev);
mem_cgroup_track_foreign_bdev(memcg, bdev_inode->i_rdev);
return;
}
- trace_track_foreign_dirty(folio, wb);
+ trace_track_foreign_dirty(folio, wb, 0);
/*
* Pick the slot to use. If there is already a slot for @wb, keep
@@ -4029,6 +4030,7 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
if (memcg_bdev_frn_wq && time_after64(ctx->at, now - intv) &&
atomic_cmpxchg(&ctx->inflight, 0, 1) == 0) {
ctx->at = 0;
+ trace_flush_foreign(wb, 0, 0, ctx->dev);
queue_work(memcg_bdev_frn_wq, &ctx->work);
}
}
@@ -4045,7 +4047,8 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
if (time_after64(frn->at, now - intv) &&
atomic_read(&frn->done.cnt) == 1) {
frn->at = 0;
- trace_flush_foreign(wb, frn->bdi_id, frn->memcg_id);
+ trace_flush_foreign(wb, frn->bdi_id, frn->memcg_id,
+ 0);
cgroup_writeback_by_id(frn->bdi_id, frn->memcg_id,
WB_REASON_FOREIGN_FLUSH,
&frn->done);
--
2.39.5
^ permalink raw reply related [flat|nested] 6+ messages in thread