* [PATCH v9 0/3] memcg,writeback: flush foreign bdev mappings separately
@ 2026-09-25 6:44 Julian Sun
2026-09-25 6:44 ` [PATCH v9 1/3] block: introduce bdev_flush_by_dev() Julian Sun
` (2 more replies)
0 siblings, 3 replies; 19+ messages in thread
From: Julian Sun @ 2026-09-25 6:44 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Hi,
This series avoids owner-wide foreign writeback triggered by dirtying
shared bdev inodes. Instead, it tracks bdev targets separately and
schedules writeback of their mappings.
Problem
=======
Bdev inodes are commonly shared by many memcgs. Dirtying their pages
can cause the owner wb to be tracked as foreign, triggering writeback
of the entire wb rather than just the bdev mapping.
In production, almost all the foreign dirtying we traced came from bdev
inodes. Across our observations, a bdev mapping had as little as about
10 MiB of dirty pages, yet foreign flushes targeted the entire owner
wb. On one machine, we observed 599 such works queued on one wb, with
an average request budget of about 43 GiB per work.
Measured impact
===============
The synthetic test used a VM with 4 vCPUs, 8 GiB RAM, ext4 and cgroup
v2, with device write bandwidth capped at 200 MiB/s. Background fio
repeatedly overwrote a 256 MiB file using buffered I/O, while
foreground fio performed sequential direct writes. Four memcgs
generated metadata and buffered writes to trigger foreign flushes.
We measured completed block writes to the background fio file,
including final sync, and foreground fio bandwidth, latency and
completion time.
Across three runs per kernel, mean background file write I/O decreased
by 72.7%, foreground bandwidth increased by 23.0%, and completion time
decreased by 18.8%.
Metric Baseline Patched Change
Background file write I/O (GiB) 6.154 1.680 -72.7%
Foreground bandwidth (MiB/s) 150.29 184.81 +23.0%
Foreground completion time (s) 109.203 88.652 -18.8%
Foreground mean latency (ms) 212.84 172.74 -18.8%
How it helps
============
Frequent writeback of the file overwritten by background fio reduces
the opportunity to coalesce buffered writes in the page cache.
Several overwrites of the same page can otherwise result in a single
write of its latest contents. Flushing between updates instead writes
the same page repeatedly, generating more device I/O for the same
application writes.
This additional I/O makes the device reach its performance limit
more easily and competes with foreground I/O for device bandwidth.
The benefit is expected to be most pronounced when these conditions
occur together on the same device:
1. Limited device bandwidth, as with an HDD, makes the additional
writeback I/O more likely to saturate the device.
2. Multiple cgroups perform buffered overwrites. Frequent foreign
writeback reduces write coalescing and increases device I/O.
3. Concurrent direct I/O bypasses the page cache and competes directly
with writeback I/O for device bandwidth.
Approach
========
Following Jan's suggestion, this series records foreign bdev targets
separately and flushes their mappings instead of their owner wbs. This
preserves a way for dirty throttling to initiate bdev writeback while
avoiding owner-wide writeback triggered by these records. The existing
foreign-wb mechanism remains unchanged for other inodes.
Tracking uses oldest-first replacement and expiry after
dirty_expire_interval. Further dirtying during an ongoing flush can
trigger another flush after it finishes.
Tracking is best effort: each memcg keeps a bounded set of device
numbers without persistent device or inode references. Recording runs
under mapping->i_pages with IRQs disabled, while dropping a device
reference can sleep. A memcg that never enters dirty throttling could
also retain an unused record and pin the device object indefinitely.
Recording only dev_t avoids these lifetime constraints. Device removal
and device-number reuse may cause an unintended best-effort flush,
which is an accepted trade-off.
Writeback submission may block for a long time, with up to four work
items outstanding per memcg. A dedicated unbound workqueue provides
a separate max_active budget so these flushes do not exhaust active
slots on a shared workqueue and delay unrelated work. If the workqueue
cannot be allocated at boot, bdev inodes continue to use the existing
foreign-wb path.
Flushes use write_inode_now() with WB_SYNC_NONE to skip inodes already
under writeback, avoiding competing writeback passes on the same bdev
inode.
This primarily affects filesystems that keep metadata in the bdev page
cache, such as ext4 and other buffer_head users, rather than the usual
metadata writeback paths of XFS and Btrfs. This also covers buffered
writes to raw block devices. The decision still does not account for
how much of the memcg's dirty memory belongs to the bdev: even a single
dirty bitmap block can trigger writeback of the entire bdev mapping
when the memcg enters dirty throttling.
Changes since v8:
- Use write_inode_now() with WB_SYNC_NONE to avoid competing writeback
passes on the same bdev inode.
- Fix a race in foreign bdev tracking.
- Add separate tracepoints for foreign bdev tracking and flushing.
Changes since v7:
- Update comments and commit messages as suggested by Tejun Heo.
Changes since v6:
- Drop open_mutex and the openers check from bdev_flush_by_dev(), and
add a !CONFIG_BLOCK stub.
- Follow the existing foreign-wb tracking policy: use timestamps for
oldest-first replacement and expiry, and allow dirtying during an
in-flight flush to re-arm the record.
- Replace frn_lock with an atomic in-flight flag.
- Drop WQ_MEM_RECLAIM and explain the separate max_active budget.
- Expand the dev_t lifetime rationale and filesystem scope, update
the foreign-dirty overview, and drop the Fixes tag.
Changes since v5:
- Fix repeated overwriting of slot 0 in patch 2 and add comments.
Changes since v4:
- Add before-and-after performance comparisons to the cover letter.
Changes since v3:
- Flush foreign-dirtied bdev inode mappings separately, as suggested
by Jan.
Thanks.
Julian Sun (3):
block: introduce bdev_flush_by_dev()
memcg,writeback: flush foreign bdev mappings separately from owner wbs
writeback: add tracepoints for foreign bdev writeback
block/bdev.c | 19 +++++
include/linux/blkdev.h | 4 +
include/linux/memcontrol.h | 12 +++
include/trace/events/writeback.h | 67 ++++++++++++++++
mm/memcontrol.c | 127 ++++++++++++++++++++++++++++---
5 files changed, 220 insertions(+), 9 deletions(-)
--
2.39.5
^ permalink raw reply [flat|nested] 19+ messages in thread
* [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-25 6:44 [PATCH v9 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
@ 2026-09-25 6:44 ` Julian Sun
2026-09-25 6:52 ` Christoph Hellwig
2026-09-25 6:44 ` [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
2026-09-25 6:44 ` [PATCH v9 3/3] writeback: add tracepoints for foreign bdev writeback Julian Sun
2 siblings, 1 reply; 19+ messages in thread
From: Julian Sun @ 2026-09-25 6:44 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Add bdev_flush_by_dev() to submit page-cache writeback for a device
identified by dev_t without requiring callers to hold a persistent
device reference.
Use write_inode_now() with WB_SYNC_NONE to skip inodes already under
writeback, avoiding competing writeback passes from multiple callers.
A subsequent patch will use this helper to flush foreign bdev mappings.
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
block/bdev.c | 19 +++++++++++++++++++
include/linux/blkdev.h | 4 ++++
2 files changed, 23 insertions(+)
diff --git a/block/bdev.c b/block/bdev.c
index cd8323083740..8925f81f16ae 100644
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -264,6 +264,25 @@ int sync_blockdev_nowait(struct block_device *bdev)
}
EXPORT_SYMBOL_GPL(sync_blockdev_nowait);
+/**
+ * bdev_flush_by_dev - attempt writeback of a block device's page cache
+ * @dev: target block device number
+ *
+ * May sleep while submitting writeback; does not wait for I/O completion.
+ */
+void bdev_flush_by_dev(dev_t dev)
+{
+ struct block_device *bdev;
+
+ bdev = blkdev_get_no_open(dev, false);
+ if (!bdev)
+ return;
+
+ /* Use the I_SYNC check to avoid concurrent writeback. */
+ write_inode_now(BD_INODE(bdev), 0);
+ blkdev_put_no_open(bdev);
+}
+
/*
* Write out and wait upon all the dirty data associated with a block
* device via its mapping. Does not take the superblock lock.
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 4f7905c3412b..27afe3ff6fb5 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1710,6 +1710,7 @@ void invalidate_bdev(struct block_device *bdev);
int sync_blockdev(struct block_device *bdev);
int sync_blockdev_range(struct block_device *bdev, loff_t lstart, loff_t lend);
int sync_blockdev_nowait(struct block_device *bdev);
+void bdev_flush_by_dev(dev_t dev);
void sync_bdevs(bool wait);
void bdev_statx(const struct path *path, struct kstat *stat, u32 request_mask);
void printk_all_partitions(void);
@@ -1726,6 +1727,9 @@ static inline int sync_blockdev_nowait(struct block_device *bdev)
{
return 0;
}
+static inline void bdev_flush_by_dev(dev_t dev)
+{
+}
static inline void sync_bdevs(bool wait)
{
}
--
2.39.5
^ permalink raw reply related [flat|nested] 19+ messages in thread
* [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-25 6:44 [PATCH v9 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-25 6:44 ` [PATCH v9 1/3] block: introduce bdev_flush_by_dev() Julian Sun
@ 2026-09-25 6:44 ` Julian Sun
2026-09-25 6:55 ` Christoph Hellwig
2026-09-25 6:44 ` [PATCH v9 3/3] writeback: add tracepoints for foreign bdev writeback Julian Sun
2 siblings, 1 reply; 19+ messages in thread
From: Julian Sun @ 2026-09-25 6:44 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Bdev inodes are routinely shared by multiple memcgs. Recording their
owner wb as foreign can flush unrelated file data when a source memcg
enters dirty throttling.
Record bdev device numbers separately and queue work on
memcg_bdev_frn_flusher to write only their mappings. Keep the existing
foreign-wb mechanism for non-bdev inodes.
Use fixed per-memcg slots. Tracking uses oldest-first replacement and
expiry after dirty_expire_interval. Further dirtying during an ongoing
flush can trigger another flush after it finishes.
Tracking is best effort: record only dev_t without holding device or
inode references. Records are updated under the mapping->i_pages lock
with IRQs disabled, but dropping a bdev reference can sleep. Also, since
foreign flushes are triggered by dirty throttling, a memcg that never
throttles could retain an unused record and pin the device object
indefinitely. Recording only dev_t avoids these lifetime constraints,
but device removal and device-number reuse may race with lookup. Such
races are expected under the best-effort semantics.
Bdev writeback submission can block for a long time on congested devices,
with up to four work items outstanding per memcg. Use a dedicated unbound
workqueue to give these flushes a separate max_active budget, so they do
not exhaust a shared workqueue's active slots and delay unrelated work.
If the workqueue cannot be allocated at boot, bdev inodes continue to use
the existing foreign-wb path.
This primarily affects filesystems that keep metadata in the bdev page
cache, such as ext4 and other buffer_head users, rather than the usual
metadata writeback paths of XFS and Btrfs. This also covers buffered writes
to raw block devices. The decision still does not account for how much of
the memcg's dirty memory belongs to the bdev. Even a single dirty bitmap
block can trigger writeback of the entire bdev mapping when the memcg
enters dirty throttling.
Suggested-by: Jan Kara <jack@suse.cz>
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
include/linux/memcontrol.h | 12 ++++
mm/memcontrol.c | 126 ++++++++++++++++++++++++++++++++++---
2 files changed, 129 insertions(+), 9 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 7d1c0ce189a8..f13aa529fa15 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -176,6 +176,17 @@ struct memcg_cgwb_frn {
struct wb_completion done; /* tracks in-flight foreign writebacks */
};
+/*
+ * No extra memcg reference is taken: this work is embedded in the memcg,
+ * and mem_cgroup_css_free() waits for it to finish before freeing the memcg.
+ */
+struct memcg_bdev_frn {
+ struct work_struct work;
+ dev_t dev; /* dev_t of the foreign bdev inode */
+ u64 at; /* last recorded dirtying time in jiffies */
+ atomic_t inflight; /* slot replacement or queued/running flush */
+};
+
/*
* Bucket for arbitrarily byte-sized objects charged to a memory
* cgroup. The bucket can be reparented in one piece when the cgroup
@@ -279,6 +290,7 @@ struct mem_cgroup {
#ifdef CONFIG_CGROUP_WRITEBACK
struct wb_domain cgwb_domain;
struct memcg_cgwb_frn cgwb_frn[MEMCG_CGWB_FRN_CNT];
+ struct memcg_bdev_frn bdev_frn[MEMCG_CGWB_FRN_CNT];
#endif
#ifdef CONFIG_LRU_GEN_WALKS_MMU
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 856a7d07586c..74677078f9b4 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -39,6 +39,7 @@
#include <linux/smp.h>
#include <linux/page-flags.h>
#include <linux/backing-dev.h>
+#include <linux/blkdev.h>
#include <linux/bit_spinlock.h>
#include <linux/rcupdate.h>
#include <linux/limits.h>
@@ -104,6 +105,7 @@ static struct kmem_cache *memcg_pn_cachep;
#ifdef CONFIG_CGROUP_WRITEBACK
static DECLARE_WAIT_QUEUE_HEAD(memcg_cgwb_frn_waitq);
+static struct workqueue_struct *memcg_bdev_frn_wq __ro_after_init;
#endif
static inline bool task_is_dying(void)
@@ -3875,20 +3877,81 @@ void mem_cgroup_wb_stats(struct bdi_writeback *wb, unsigned long *pfilepages,
* most recent foreign dirtying events and initiating remote flushes on
* them when local writeback isn't enough to keep the memory clean enough.
*
- * The following two functions implement such mechanism. When a foreign
- * page - a page whose memcg and writeback ownerships don't match - is
- * dirtied, mem_cgroup_track_foreign_dirty() records the inode owning
- * bdi_writeback on the page owning memcg. When balance_dirty_pages()
+ * When a foreign page - a page whose memcg and writeback ownerships don't
+ * match - is dirtied, mem_cgroup_track_foreign_dirty() records the inode
+ * owning bdi_writeback on the page owning memcg. When balance_dirty_pages()
* decides that the memcg needs to sleep due to high dirty ratio, it calls
* mem_cgroup_flush_foreign() which queues writeback on the recorded
* foreign bdi_writebacks which haven't expired. Both the numbers of
* recorded bdi_writebacks and concurrent in-flight foreign writebacks are
* limited to MEMCG_CGWB_FRN_CNT.
*
- * The mechanism only remembers IDs and doesn't hold any object references.
- * As being wrong occasionally doesn't matter, updates and accesses to the
- * records are lockless and racy.
+ * Bdev inodes are commonly shared by many memcgs. Flushing their owner wb
+ * can write unrelated file data, so each memcg tracks foreign bdevs
+ * separately in MEMCG_CGWB_FRN_CNT slots keyed by dev_t. These records
+ * use the same expiry policy, but trigger writeback of only the bdev
+ * mappings.
+ *
+ * Both kinds of records only remember IDs and don't hold any object
+ * references. As being wrong occasionally doesn't matter, updates and
+ * accesses to the records are lockless and racy.
*/
+
+static void mem_cgroup_track_foreign_bdev(struct mem_cgroup *memcg, dev_t dev)
+{
+ struct memcg_bdev_frn *frn;
+ int i;
+ int oldest;
+ u64 now = get_jiffies_64();
+ u64 oldest_at;
+
+retry:
+ oldest = -1;
+ oldest_at = now;
+
+ /*
+ * Pick the slot to use. If there is already a slot for @dev, keep
+ * using it. If not replace the oldest one which isn't being
+ * written out.
+ */
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+ frn = &memcg->bdev_frn[i];
+ if (frn->dev == dev)
+ break;
+ if (atomic_read(&frn->inflight))
+ continue;
+ if (time_before64(frn->at, oldest_at)) {
+ oldest = i;
+ oldest_at = frn->at;
+ }
+ }
+
+ if (i < MEMCG_CGWB_FRN_CNT) {
+ /*
+ * Re-using an existing one. Update timestamp lazily to
+ * avoid making the cacheline hot. We want them to be
+ * reasonably up-to-date and significantly shorter than
+ * dirty_expire_interval as that's what expires the record.
+ * Use the shorter of 1s and dirty_expire_interval / 8.
+ */
+ unsigned long update_intv =
+ min_t(unsigned long, HZ,
+ msecs_to_jiffies(dirty_expire_interval * 10) / 8);
+
+ if (time_before64(frn->at, now - update_intv))
+ frn->at = now;
+ } else if (oldest >= 0) {
+ frn = &memcg->bdev_frn[oldest];
+ /* Reserve the slot: it may have become busy after the inflight check. */
+ if (atomic_cmpxchg(&frn->inflight, 0, 1) != 0)
+ goto retry;
+
+ frn->dev = dev;
+ frn->at = now;
+ atomic_set_release(&frn->inflight, 0);
+ }
+}
+
void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
struct bdi_writeback *wb)
{
@@ -3898,6 +3961,14 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
u64 oldest_at = now;
int oldest = -1;
int i;
+ struct address_space *mapping = folio_mapping(folio);
+ struct inode *inode = mapping->host;
+
+ if (memcg_bdev_frn_wq && sb_is_blkdev_sb(inode->i_sb)) {
+ trace_track_foreign_dirty(folio, wb);
+ mem_cgroup_track_foreign_bdev(memcg, inode->i_rdev);
+ return;
+ }
trace_track_foreign_dirty(folio, wb);
@@ -3941,6 +4012,16 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
}
}
+static void bdev_frn_flush_work(struct work_struct *work)
+{
+ struct memcg_bdev_frn *frn =
+ container_of(work, struct memcg_bdev_frn, work);
+
+ bdev_flush_by_dev(frn->dev);
+
+ atomic_set(&frn->inflight, 0);
+}
+
/* issue foreign writeback flushes for recorded foreign dirtying events */
void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
{
@@ -3949,6 +4030,21 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
u64 now = jiffies_64;
int i;
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+ struct memcg_bdev_frn *frn = &memcg->bdev_frn[i];
+
+ /* Keep the workqueue availability check explicit. */
+ if (memcg_bdev_frn_wq && time_after64(frn->at, now - intv) &&
+ atomic_cmpxchg(&frn->inflight, 0, 1) == 0) {
+ /*
+ * Clear now so dirtying during writeback can refresh
+ * the timestamp for a later flush.
+ */
+ frn->at = 0;
+ queue_work(memcg_bdev_frn_wq, &frn->work);
+ }
+ }
+
for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
struct memcg_cgwb_frn *frn = &memcg->cgwb_frn[i];
@@ -4202,9 +4298,13 @@ static struct mem_cgroup *mem_cgroup_alloc(struct mem_cgroup *parent)
memcg->kmemcg_id = -1;
#ifdef CONFIG_CGROUP_WRITEBACK
INIT_LIST_HEAD(&memcg->cgwb_list);
- for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++)
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
+ struct memcg_bdev_frn *frn = &memcg->bdev_frn[i];
+
memcg->cgwb_frn[i].done =
__WB_COMPLETION_INIT(&memcg_cgwb_frn_waitq);
+ INIT_WORK(&frn->work, bdev_frn_flush_work);
+ }
#endif
lru_gen_init_memcg(memcg);
return memcg;
@@ -4386,8 +4486,10 @@ static void mem_cgroup_css_free(struct cgroup_subsys_state *css)
int __maybe_unused i;
#ifdef CONFIG_CGROUP_WRITEBACK
- for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++)
+ for (i = 0; i < MEMCG_CGWB_FRN_CNT; i++) {
wb_wait_for_completion(&memcg->cgwb_frn[i].done);
+ flush_work(&memcg->bdev_frn[i].work);
+ }
#endif
if (cgroup_subsys_on_dfl(memory_cgrp_subsys) && !cgroup_memory_nosocket)
static_branch_dec(&memcg_sockets_enabled_key);
@@ -5703,6 +5805,12 @@ int __init mem_cgroup_init(void)
memcg_wq = alloc_workqueue("memcg", WQ_PERCPU, 0);
WARN_ON(!memcg_wq);
+#ifdef CONFIG_CGROUP_WRITEBACK
+ memcg_bdev_frn_wq = alloc_workqueue("memcg_bdev_frn_flusher",
+ WQ_UNBOUND, 0);
+ WARN_ON(!memcg_bdev_frn_wq);
+#endif
+
for_each_possible_cpu(cpu) {
INIT_WORK(&per_cpu_ptr(&memcg_stock, cpu)->work,
drain_local_memcg_stock);
--
2.39.5
^ permalink raw reply related [flat|nested] 19+ messages in thread
* [PATCH v9 3/3] writeback: add tracepoints for foreign bdev writeback
2026-09-25 6:44 [PATCH v9 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-25 6:44 ` [PATCH v9 1/3] block: introduce bdev_flush_by_dev() Julian Sun
2026-09-25 6:44 ` [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
@ 2026-09-25 6:44 ` Julian Sun
2 siblings, 0 replies; 19+ messages in thread
From: Julian Sun @ 2026-09-25 6:44 UTC (permalink / raw)
To: linux-block, cgroups, linux-mm, linux-fsdevel
Cc: axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
Add track_foreign_bdev_dirty and flush_foreign_bdev to trace foreign
bdev dirty tracking and flush requests separately from the existing
foreign-wb path. Record the target device number as major:minor while
leaving the existing tracepoint interfaces unchanged.
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
---
include/trace/events/writeback.h | 67 ++++++++++++++++++++++++++++++++
mm/memcontrol.c | 3 +-
2 files changed, 69 insertions(+), 1 deletion(-)
diff --git a/include/trace/events/writeback.h b/include/trace/events/writeback.h
index 13ee076ccd16..4a85883fb8a0 100644
--- a/include/trace/events/writeback.h
+++ b/include/trace/events/writeback.h
@@ -339,6 +339,73 @@ TRACE_EVENT(flush_foreign,
__entry->frn_memcg_id
)
);
+
+TRACE_EVENT(track_foreign_bdev_dirty,
+
+ TP_PROTO(struct folio *folio, struct bdi_writeback *wb, dev_t dev),
+
+ TP_ARGS(folio, wb, dev),
+
+ TP_STRUCT__entry(
+ __array(char, name, 32)
+ __field(u64, bdi_id)
+ __field(u64, ino)
+ __field(u64, cgroup_ino)
+ __field(u64, page_cgroup_ino)
+ __field(unsigned int, memcg_id)
+ __field(dev_t, dev)
+ ),
+
+ TP_fast_assign(
+ struct inode *inode = folio_mapping(folio)->host;
+
+ strscpy_pad(__entry->name, bdi_dev_name(wb->bdi), 32);
+ __entry->bdi_id = wb->bdi->id;
+ __entry->ino = inode->i_ino;
+ __entry->memcg_id = wb->memcg_css->id;
+ __entry->cgroup_ino = __trace_wb_assign_cgroup(wb);
+ __entry->dev = dev;
+
+ rcu_read_lock();
+ __entry->page_cgroup_ino = cgroup_ino(folio_memcg(folio)->css.cgroup);
+ rcu_read_unlock();
+ ),
+
+ TP_printk("bdi %s[%llu]: ino=%llu memcg_id=%u cgroup_ino=%llu page_cgroup_ino=%llu dev=%u:%u",
+ __entry->name,
+ __entry->bdi_id,
+ __entry->ino,
+ __entry->memcg_id,
+ __entry->cgroup_ino,
+ __entry->page_cgroup_ino,
+ MAJOR(__entry->dev), MINOR(__entry->dev)
+ )
+);
+
+TRACE_EVENT(flush_foreign_bdev,
+
+ TP_PROTO(struct bdi_writeback *wb, dev_t dev),
+
+ TP_ARGS(wb, dev),
+
+ TP_STRUCT__entry(
+ __array(char, name, 32)
+ __field(u64, cgroup_ino)
+ __field(dev_t, dev)
+ ),
+
+ TP_fast_assign(
+ strscpy_pad(__entry->name, bdi_dev_name(wb->bdi), 32);
+ __entry->cgroup_ino = __trace_wb_assign_cgroup(wb);
+ __entry->dev = dev;
+ ),
+
+ TP_printk("bdi %s: cgroup_ino=%llu dev=%u:%u",
+ __entry->name,
+ __entry->cgroup_ino,
+ MAJOR(__entry->dev), MINOR(__entry->dev)
+ )
+);
#endif
DECLARE_EVENT_CLASS(writeback_write_inode_template,
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 74677078f9b4..0108b4177833 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -3965,7 +3965,7 @@ void mem_cgroup_track_foreign_dirty_slowpath(struct folio *folio,
struct inode *inode = mapping->host;
if (memcg_bdev_frn_wq && sb_is_blkdev_sb(inode->i_sb)) {
- trace_track_foreign_dirty(folio, wb);
+ trace_track_foreign_bdev_dirty(folio, wb, inode->i_rdev);
mem_cgroup_track_foreign_bdev(memcg, inode->i_rdev);
return;
}
@@ -4041,6 +4041,7 @@ void mem_cgroup_flush_foreign(struct bdi_writeback *wb)
* the timestamp for a later flush.
*/
frn->at = 0;
+ trace_flush_foreign_bdev(wb, frn->dev);
queue_work(memcg_bdev_frn_wq, &frn->work);
}
}
--
2.39.5
^ permalink raw reply related [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-25 6:44 ` [PATCH v9 1/3] block: introduce bdev_flush_by_dev() Julian Sun
@ 2026-09-25 6:52 ` Christoph Hellwig
2026-09-25 10:34 ` Jan Kara
0 siblings, 1 reply; 19+ messages in thread
From: Christoph Hellwig @ 2026-09-25 6:52 UTC (permalink / raw)
To: Julian Sun
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
tj, akpm
On Fri, Sep 25, 2026 at 02:44:42PM +0800, Julian Sun wrote:
> Add bdev_flush_by_dev() to submit page-cache writeback for a device
> identified by dev_t without requiring callers to hold a persistent
> device reference.
This isn't really a flush is it? This is writeback.
> + *
> + * May sleep while submitting writeback; does not wait for I/O completion.
> + */
> +void bdev_flush_by_dev(dev_t dev)
> +{
> + struct block_device *bdev;
> +
> + bdev = blkdev_get_no_open(dev, false);
> + if (!bdev)
> + return;
Please don't add more blkdev_get_no_open. You can get from an
bdevfs inode to the struct block_device trivially using
I_BDEV instead of doing another lookup. Or just grab a reference
to the the inode, which would seem even easier.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-25 6:44 ` [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
@ 2026-09-25 6:55 ` Christoph Hellwig
2026-09-25 13:00 ` Julian Sun
0 siblings, 1 reply; 19+ messages in thread
From: Christoph Hellwig @ 2026-09-25 6:55 UTC (permalink / raw)
To: Julian Sun
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
tj, akpm
> +/*
> + * No extra memcg reference is taken: this work is embedded in the memcg,
> + * and mem_cgroup_css_free() waits for it to finish before freeing the memcg.
> + */
> +struct memcg_bdev_frn {
No reason not to spell out foreign here..
> + /* Reserve the slot: it may have become busy after the inflight check. */
Overly long line.
> + struct address_space *mapping = folio_mapping(folio);
> + struct inode *inode = mapping->host;
> +
> + if (memcg_bdev_frn_wq && sb_is_blkdev_sb(inode->i_sb)) {
Why do we need a NULL check for the workqueue here?
> memcg_wq = alloc_workqueue("memcg", WQ_PERCPU, 0);
> WARN_ON(!memcg_wq);
>
> +#ifdef CONFIG_CGROUP_WRITEBACK
> + memcg_bdev_frn_wq = alloc_workqueue("memcg_bdev_frn_flusher",
> + WQ_UNBOUND, 0);
> + WARN_ON(!memcg_bdev_frn_wq);
> +#endif
Instead please fail the boot. If we can't alloc a workqueue at
boot time, we're toast anyway.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-25 6:52 ` Christoph Hellwig
@ 2026-09-25 10:34 ` Jan Kara
2026-09-28 6:24 ` Christoph Hellwig
0 siblings, 1 reply; 19+ messages in thread
From: Jan Kara @ 2026-09-25 10:34 UTC (permalink / raw)
To: Christoph Hellwig
Cc: Julian Sun, linux-block, cgroups, linux-mm, linux-fsdevel, axboe,
hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song, willy,
jack, tj, akpm
On Thu 24-09-26 23:52:33, Christoph Hellwig wrote:
> On Fri, Sep 25, 2026 at 02:44:42PM +0800, Julian Sun wrote:
> > Add bdev_flush_by_dev() to submit page-cache writeback for a device
> > identified by dev_t without requiring callers to hold a persistent
> > device reference.
>
> This isn't really a flush is it? This is writeback.
So bdev_writeback_by_dev()? Fine by me.
> > + *
> > + * May sleep while submitting writeback; does not wait for I/O completion.
> > + */
> > +void bdev_flush_by_dev(dev_t dev)
> > +{
> > + struct block_device *bdev;
> > +
> > + bdev = blkdev_get_no_open(dev, false);
> > + if (!bdev)
> > + return;
>
> Please don't add more blkdev_get_no_open. You can get from an
> bdevfs inode to the struct block_device trivially using
> I_BDEV instead of doing another lookup. Or just grab a reference
> to the the inode, which would seem even easier.
Well, Julian is explaining that in the cover letter. We don't want memcgs
to hold bdev references just for foreign flush tracking. There's no easy
way to get rid of them so they could pin bdevs for a long time leading to
strange artifacts.
As you write, we could just hold a reference to the bdev inode. Those get
unhashed when the device dies (__del_gendisk()) so memcgs would be just
wasting some memory by holding these inodes alive. But still, transitioning
from inode to proper bdev reference verifying bdev is still alive will add
a bit of hairy code (essentially what blkdev_get_no_open() does plus
verification inode is still hashed). Plus when replacing foreign flush
entries you're under irqsafe xa_lock so doing iput() from there is kind of
a nogo.
So I think tracking devices by dev_t is a good way of dealing with these
problems. But then when flushing you have to transition from dev_t to
struct block_device and blkdev_get_no_open() is the canonical way of doing
that (and note this use is just internal to block/bdev.c).
So how would you prefer to deal with that?
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-25 6:55 ` Christoph Hellwig
@ 2026-09-25 13:00 ` Julian Sun
2026-09-28 6:18 ` Christoph Hellwig
0 siblings, 1 reply; 19+ messages in thread
From: Julian Sun @ 2026-09-25 13:00 UTC (permalink / raw)
To: Christoph Hellwig
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, jack,
tj, akpm
On 9/25/26 2:55 PM, Christoph Hellwig wrote:
>> +/*
>> + * No extra memcg reference is taken: this work is embedded in the memcg,
>> + * and mem_cgroup_css_free() waits for it to finish before freeing the memcg.
>> + */
>> +struct memcg_bdev_frn {
>
> No reason not to spell out foreign here..
Yeah..Andrew and I discussed the naming in another thread. The idea was to stay
consistent with the existing naming and keep this patchset focused on the functional
changes, leaving a broader naming cleanup to a separate patch.
https://lore.kernel.org/linux-fsdevel/20260917065759.2643940-1-sunjunchao@bytedance.com/T/#m0d196c3b8376c9b51276f6c5d26ab9f29e170169
>
>> + /* Reserve the slot: it may have become busy after the inflight check. */
>
> Overly long line.
>
>> + struct address_space *mapping = folio_mapping(folio);
>> + struct inode *inode = mapping->host;
>> +
>> + if (memcg_bdev_frn_wq && sb_is_blkdev_sb(inode->i_sb)) {
>
> Why do we need a NULL check for the workqueue here?
>
>> memcg_wq = alloc_workqueue("memcg", WQ_PERCPU, 0);
>> WARN_ON(!memcg_wq);
>>
>> +#ifdef CONFIG_CGROUP_WRITEBACK
>> + memcg_bdev_frn_wq = alloc_workqueue("memcg_bdev_frn_flusher",
>> + WQ_UNBOUND, 0);
>> + WARN_ON(!memcg_bdev_frn_wq);
>> +#endif
>
> Instead please fail the boot. If we can't alloc a workqueue at
> boot time, we're toast anyway.
I followed the existing memcg_wq allocation handling here, which also
warns and continues. Would it make sense to change both allocations to
fail the boot in a separate patch, so the policy is consistent?
If you prefer to address both in this series, I'll make that change
in the next revision.
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs
2026-09-25 13:00 ` Julian Sun
@ 2026-09-28 6:18 ` Christoph Hellwig
0 siblings, 0 replies; 19+ messages in thread
From: Christoph Hellwig @ 2026-09-28 6:18 UTC (permalink / raw)
To: Julian Sun
Cc: Christoph Hellwig, linux-block, cgroups, linux-mm, linux-fsdevel,
axboe, hannes, mhocko, roman.gushchin, shakeel.butt, muchun.song,
willy, jack, tj, akpm
On Fri, Sep 25, 2026 at 09:00:28PM +0800, Julian Sun wrote:
> Yeah..Andrew and I discussed the naming in another thread. The idea was to stay
> consistent with the existing naming and keep this patchset focused on the functional
> changes, leaving a broader naming cleanup to a separate patch.
>
> https://lore.kernel.org/linux-fsdevel/20260917065759.2643940-1-sunjunchao@bytedance.com/T/#m0d196c3b8376c9b51276f6c5d26ab9f29e170169
Let's not add more of this, even if fixing up the previous ones belongs
into a separate thread.
> >
> > Instead please fail the boot. If we can't alloc a workqueue at
> > boot time, we're toast anyway.
> I followed the existing memcg_wq allocation handling here, which also
> warns and continues. Would it make sense to change both allocations to
> fail the boot in a separate patch, so the policy is consistent?
> If you prefer to address both in this series, I'll make that change
> in the next revision.
Leaving dangling pointers during catastropic boot failures and then
dealing with them at runtime isn't a good idea. And yes, the code
above should eventually be fixed up as well.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-25 10:34 ` Jan Kara
@ 2026-09-28 6:24 ` Christoph Hellwig
2026-09-28 9:51 ` Julian Sun
2026-10-02 9:23 ` Julian Sun
0 siblings, 2 replies; 19+ messages in thread
From: Christoph Hellwig @ 2026-09-28 6:24 UTC (permalink / raw)
To: Jan Kara
Cc: Christoph Hellwig, Julian Sun, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Fri, Sep 25, 2026 at 12:34:52PM +0200, Jan Kara wrote:
> On Thu 24-09-26 23:52:33, Christoph Hellwig wrote:
> > On Fri, Sep 25, 2026 at 02:44:42PM +0800, Julian Sun wrote:
> > > Add bdev_flush_by_dev() to submit page-cache writeback for a device
> > > identified by dev_t without requiring callers to hold a persistent
> > > device reference.
> >
> > This isn't really a flush is it? This is writeback.
>
> So bdev_writeback_by_dev()? Fine by me.
Yes.
> Well, Julian is explaining that in the cover letter. We don't want memcgs
> to hold bdev references just for foreign flush tracking. There's no easy
> way to get rid of them so they could pin bdevs for a long time leading to
> strange artifacts.
I have to admit I didn't get it when flying over the cover letter.
But this also really belongs into the patch where it is more obvious,
or even better into code comments.
>
> As you write, we could just hold a reference to the bdev inode. Those get
> unhashed when the device dies (__del_gendisk()) so memcgs would be just
> wasting some memory by holding these inodes alive. But still, transitioning
> from inode to proper bdev reference verifying bdev is still alive will add
> a bit of hairy code (essentially what blkdev_get_no_open() does plus
> verification inode is still hashed). Plus when replacing foreign flush
> entries you're under irqsafe xa_lock so doing iput() from there is kind of
> a nogo.
>
> So I think tracking devices by dev_t is a good way of dealing with these
> problems. But then when flushing you have to transition from dev_t to
> struct block_device and blkdev_get_no_open() is the canonical way of doing
> that (and note this use is just internal to block/bdev.c).
I'd really prefer not to grow more blkdev_get_no_open users.
But the more I look at this I wonder why we even bother. Trying to
attribute individual bits of dirty metadata to cgroups is pretty much
insane. Why don't we bypass memcg accounting for buffer_heads like we
always did for the XFS buffer cache, and what btrfs switched to last
year? See commit b55102826d7d ("btrfs: set AS_KERNEL_FILE on the
btree_inode") for the btrfs side.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-28 6:24 ` Christoph Hellwig
@ 2026-09-28 9:51 ` Julian Sun
2026-10-02 11:10 ` Jan Kara
2026-10-02 9:23 ` Julian Sun
1 sibling, 1 reply; 19+ messages in thread
From: Julian Sun @ 2026-09-28 9:51 UTC (permalink / raw)
To: Christoph Hellwig, Jan Kara
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, tj,
akpm, Boris Burkov
On 9/28/26 2:24 PM, Christoph Hellwig wrote:
Hi, Christoph.
Thanks for your review and suggestions.
> On Fri, Sep 25, 2026 at 12:34:52PM +0200, Jan Kara wrote:
>> On Thu 24-09-26 23:52:33, Christoph Hellwig wrote:
>>> On Fri, Sep 25, 2026 at 02:44:42PM +0800, Julian Sun wrote:
>>>> Add bdev_flush_by_dev() to submit page-cache writeback for a device
>>>> identified by dev_t without requiring callers to hold a persistent
>>>> device reference.
>>>
>>> This isn't really a flush is it? This is writeback.
>>
>> So bdev_writeback_by_dev()? Fine by me.
>
> Yes.
>
>> Well, Julian is explaining that in the cover letter. We don't want memcgs
>> to hold bdev references just for foreign flush tracking. There's no easy
>> way to get rid of them so they could pin bdevs for a long time leading to
>> strange artifacts.
>
> I have to admit I didn't get it when flying over the cover letter.
> But this also really belongs into the patch where it is more obvious,
> or even better into code comments.
>
>>
>> As you write, we could just hold a reference to the bdev inode. Those get
>> unhashed when the device dies (__del_gendisk()) so memcgs would be just
>> wasting some memory by holding these inodes alive. But still, transitioning
>> from inode to proper bdev reference verifying bdev is still alive will add
>> a bit of hairy code (essentially what blkdev_get_no_open() does plus
>> verification inode is still hashed). Plus when replacing foreign flush
>> entries you're under irqsafe xa_lock so doing iput() from there is kind of
>> a nogo.
>>
>> So I think tracking devices by dev_t is a good way of dealing with these
>> problems. But then when flushing you have to transition from dev_t to
>> struct block_device and blkdev_get_no_open() is the canonical way of doing
>> that (and note this use is just internal to block/bdev.c).
>
> I'd really prefer not to grow more blkdev_get_no_open users.
>
> But the more I look at this I wonder why we even bother. Trying to
> attribute individual bits of dirty metadata to cgroups is pretty much
> insane. Why don't we bypass memcg accounting for buffer_heads like we
> always did for the XFS buffer cache, and what btrfs switched to last
> year? See commit b55102826d7d ("btrfs: set AS_KERNEL_FILE on the
> btree_inode") for the btrfs side.
That sounds reasonable. However, IMHO, the main difference is that ext4
stores this metadata in the bdev page cache rather than a separate metadata
cache as btrfs and XFS do. The same bdev->bd_mapping is also used for buffered
reads of the raw block device from userspace. Setting AS_KERNEL_FILE on that
mapping would therefore also charge newly allocated page-cache folios from
those reads to the root memcg rather than the reader's memcg, bypassing its
memory limits.
If we pass a flag from ext4's metadata buffer-cache allocation sites and
check it in grow_buffers() to charge new folios to the root memcg, accounting
them as kernel-file pages would still require persistent per-folio state.
When a folio is removed from the page cache, the allocation context is gone,
so we cannot tell whether it was counted as a kernel-file folio.
Separating ext4's metadata cache from the bdev page cache would address this,
but that seems like a much larger, longer-term change.
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-28 6:24 ` Christoph Hellwig
2026-09-28 9:51 ` Julian Sun
@ 2026-10-02 9:23 ` Julian Sun
2026-10-05 8:27 ` Christoph Hellwig
1 sibling, 1 reply; 19+ messages in thread
From: Julian Sun @ 2026-10-02 9:23 UTC (permalink / raw)
To: Christoph Hellwig, Jan Kara
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, tj,
akpm, Boris Burkov
On 9/28/26 2:24 PM, Christoph Hellwig wrote:
> On Fri, Sep 25, 2026 at 12:34:52PM +0200, Jan Kara wrote:
>> On Thu 24-09-26 23:52:33, Christoph Hellwig wrote:
>>> On Fri, Sep 25, 2026 at 02:44:42PM +0800, Julian Sun wrote:
>>>> Add bdev_flush_by_dev() to submit page-cache writeback for a device
>>>> identified by dev_t without requiring callers to hold a persistent
>>>> device reference.
>>>
>>> This isn't really a flush is it? This is writeback.
>>
>> So bdev_writeback_by_dev()? Fine by me.
>
> Yes.
>
>> Well, Julian is explaining that in the cover letter. We don't want memcgs
>> to hold bdev references just for foreign flush tracking. There's no easy
>> way to get rid of them so they could pin bdevs for a long time leading to
>> strange artifacts.
>
> I have to admit I didn't get it when flying over the cover letter.
> But this also really belongs into the patch where it is more obvious,
> or even better into code comments.
>
>>
>> As you write, we could just hold a reference to the bdev inode. Those get
>> unhashed when the device dies (__del_gendisk()) so memcgs would be just
>> wasting some memory by holding these inodes alive. But still, transitioning
>> from inode to proper bdev reference verifying bdev is still alive will add
>> a bit of hairy code (essentially what blkdev_get_no_open() does plus
>> verification inode is still hashed). Plus when replacing foreign flush
>> entries you're under irqsafe xa_lock so doing iput() from there is kind of
>> a nogo.
>>
>> So I think tracking devices by dev_t is a good way of dealing with these
>> problems. But then when flushing you have to transition from dev_t to
>> struct block_device and blkdev_get_no_open() is the canonical way of doing
>> that (and note this use is just internal to block/bdev.c).
>
> I'd really prefer not to grow more blkdev_get_no_open users.
Hi Christoph,
Could you elaborate on your concern about adding another
blkdev_get_no_open() user here? Is it specific to this helper, or to
submitting writeback with a temporary bdev reference without opening
the device?
We only retain an identifier while tracking, and acquire a reference
when performing writeback. Looking up the device through another path
would still require handling the same lifetime constraints. Is there
any cleanup or rework we could do first to address your concerns?
>
> But the more I look at this I wonder why we even bother. Trying to
> attribute individual bits of dirty metadata to cgroups is pretty much
> insane. Why don't we bypass memcg accounting for buffer_heads like we
> always did for the XFS buffer cache, and what btrfs switched to last
> year? See commit b55102826d7d ("btrfs: set AS_KERNEL_FILE on the
> btree_inode") for the btrfs side.
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-09-28 9:51 ` Julian Sun
@ 2026-10-02 11:10 ` Jan Kara
2026-10-05 8:26 ` Christoph Hellwig
0 siblings, 1 reply; 19+ messages in thread
From: Jan Kara @ 2026-10-02 11:10 UTC (permalink / raw)
To: Julian Sun
Cc: Christoph Hellwig, Jan Kara, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Mon 28-09-26 17:51:09, Julian Sun wrote:
> On 9/28/26 2:24 PM, Christoph Hellwig wrote:
> >> So I think tracking devices by dev_t is a good way of dealing with these
> >> problems. But then when flushing you have to transition from dev_t to
> >> struct block_device and blkdev_get_no_open() is the canonical way of doing
> >> that (and note this use is just internal to block/bdev.c).
> >
> > I'd really prefer not to grow more blkdev_get_no_open users.
> >
> > But the more I look at this I wonder why we even bother. Trying to
> > attribute individual bits of dirty metadata to cgroups is pretty much
> > insane. Why don't we bypass memcg accounting for buffer_heads like we
> > always did for the XFS buffer cache, and what btrfs switched to last
> > year? See commit b55102826d7d ("btrfs: set AS_KERNEL_FILE on the
> > btree_inode") for the btrfs side.
>
> That sounds reasonable. However, IMHO, the main difference is that ext4
> stores this metadata in the bdev page cache rather than a separate metadata
> cache as btrfs and XFS do. The same bdev->bd_mapping is also used for buffered
> reads of the raw block device from userspace. Setting AS_KERNEL_FILE on that
> mapping would therefore also charge newly allocated page-cache folios from
> those reads to the root memcg rather than the reader's memcg, bypassing its
> memory limits.
>
> If we pass a flag from ext4's metadata buffer-cache allocation sites and
> check it in grow_buffers() to charge new folios to the root memcg, accounting
> them as kernel-file pages would still require persistent per-folio state.
> When a folio is removed from the page cache, the allocation context is gone,
> so we cannot tell whether it was counted as a kernel-file folio.
>
> Separating ext4's metadata cache from the bdev page cache would address this,
> but that seems like a much larger, longer-term change.
FWIW I share Julian's concern here. It would be fine to use AS_KERNEL_FILE
for ext4 metadata but I think just unconditionally setting AS_KERNEL_FILE
for bdev mappings will cause issues because some users may be using bdevs
directly for their workloads and they could still expect proper memcg
accounting to work in that case.
We could set AS_KERNEL_FILE when opening bdev for a filesystem (and remove
it when releasing bdev) but that would have to make sure there are no
folios in the bdev mapping when changing the flag as otherwise the
accounting would go wrong. Looks it might be doable but getting all the
cornercases right will be hairy and overall not very appealing to me...
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-10-02 11:10 ` Jan Kara
@ 2026-10-05 8:26 ` Christoph Hellwig
2026-10-05 16:25 ` Jan Kara
0 siblings, 1 reply; 19+ messages in thread
From: Christoph Hellwig @ 2026-10-05 8:26 UTC (permalink / raw)
To: Jan Kara
Cc: Julian Sun, Christoph Hellwig, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Fri, Oct 02, 2026 at 01:10:42PM +0200, Jan Kara wrote:
> FWIW I share Julian's concern here. It would be fine to use AS_KERNEL_FILE
> for ext4 metadata but I think just unconditionally setting AS_KERNEL_FILE
> for bdev mappings will cause issues because some users may be using bdevs
> directly for their workloads and they could still expect proper memcg
> accounting to work in that case.
Agreed that it should not set unconditionally.
> We could set AS_KERNEL_FILE when opening bdev for a filesystem (and remove
> it when releasing bdev) but that would have to make sure there are no
> folios in the bdev mapping when changing the flag as otherwise the
> accounting would go wrong. Looks it might be doable but getting all the
> cornercases right will be hairy and overall not very appealing to me...
I'd rather not support special case writeback code just for this legacy
fs abuses bdev buffer cache case. And we basically need to tear
down pagecache at unmount anyway, as i_blkbits can change, so while
we do need to be careful, I don't think it really is a major issue.
We might be able to restrict to setting it when
CONFIG_BLK_DEV_WRITE_MOUNTED is disabled to avoid the problem of non-fs
shared mmap writers, as anyone using a modern kernel and cgroups really
should have that disabled.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-10-02 9:23 ` Julian Sun
@ 2026-10-05 8:27 ` Christoph Hellwig
0 siblings, 0 replies; 19+ messages in thread
From: Christoph Hellwig @ 2026-10-05 8:27 UTC (permalink / raw)
To: Julian Sun
Cc: Christoph Hellwig, Jan Kara, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Fri, Oct 02, 2026 at 05:23:54PM +0800, Julian Sun wrote:
> Could you elaborate on your concern about adding another
> blkdev_get_no_open() user here? Is it specific to this helper, or to
> submitting writeback with a temporary bdev reference without opening
> the device?
Both, really.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-10-05 8:26 ` Christoph Hellwig
@ 2026-10-05 16:25 ` Jan Kara
2026-10-08 8:00 ` Julian Sun
0 siblings, 1 reply; 19+ messages in thread
From: Jan Kara @ 2026-10-05 16:25 UTC (permalink / raw)
To: Christoph Hellwig
Cc: Jan Kara, Julian Sun, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Mon 05-10-26 01:26:28, Christoph Hellwig wrote:
> On Fri, Oct 02, 2026 at 01:10:42PM +0200, Jan Kara wrote:
> > FWIW I share Julian's concern here. It would be fine to use AS_KERNEL_FILE
> > for ext4 metadata but I think just unconditionally setting AS_KERNEL_FILE
> > for bdev mappings will cause issues because some users may be using bdevs
> > directly for their workloads and they could still expect proper memcg
> > accounting to work in that case.
>
> Agreed that it should not set unconditionally.
>
> > We could set AS_KERNEL_FILE when opening bdev for a filesystem (and remove
> > it when releasing bdev) but that would have to make sure there are no
> > folios in the bdev mapping when changing the flag as otherwise the
> > accounting would go wrong. Looks it might be doable but getting all the
> > cornercases right will be hairy and overall not very appealing to me...
>
> I'd rather not support special case writeback code just for this legacy
> fs abuses bdev buffer cache case. And we basically need to tear
> down pagecache at unmount anyway, as i_blkbits can change, so while
> we do need to be careful, I don't think it really is a major issue.
OK, after some more thought yes, I think we can make that work.
> We might be able to restrict to setting it when
> CONFIG_BLK_DEV_WRITE_MOUNTED is disabled to avoid the problem of non-fs
> shared mmap writers, as anyone using a modern kernel and cgroups really
> should have that disabled.
I don't think that's really needed. If such writes are happening, they are
*very* limited and done by a system administrator only (as they are very
dangerous) so the fact the writes will get charged to the root cgroup is
not an issue.
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-10-05 16:25 ` Jan Kara
@ 2026-10-08 8:00 ` Julian Sun
2026-10-08 8:51 ` Jan Kara
0 siblings, 1 reply; 19+ messages in thread
From: Julian Sun @ 2026-10-08 8:00 UTC (permalink / raw)
To: Jan Kara, Christoph Hellwig
Cc: linux-block, cgroups, linux-mm, linux-fsdevel, axboe, hannes,
mhocko, roman.gushchin, shakeel.butt, muchun.song, willy, tj,
akpm, Boris Burkov
On 10/6/26 12:25 AM, Jan Kara wrote:
> On Mon 05-10-26 01:26:28, Christoph Hellwig wrote:
>> On Fri, Oct 02, 2026 at 01:10:42PM +0200, Jan Kara wrote:
>>> FWIW I share Julian's concern here. It would be fine to use AS_KERNEL_FILE
>>> for ext4 metadata but I think just unconditionally setting AS_KERNEL_FILE
>>> for bdev mappings will cause issues because some users may be using bdevs
>>> directly for their workloads and they could still expect proper memcg
>>> accounting to work in that case.
>>
>> Agreed that it should not set unconditionally.
>>
>>> We could set AS_KERNEL_FILE when opening bdev for a filesystem (and remove
>>> it when releasing bdev) but that would have to make sure there are no
>>> folios in the bdev mapping when changing the flag as otherwise the
>>> accounting would go wrong. Looks it might be doable but getting all the
>>> cornercases right will be hairy and overall not very appealing to me...
>>
>> I'd rather not support special case writeback code just for this legacy
>> fs abuses bdev buffer cache case. And we basically need to tear
>> down pagecache at unmount anyway, as i_blkbits can change, so while
>> we do need to be careful, I don't think it really is a major issue.
>
> OK, after some more thought yes, I think we can make that work.
>
>> We might be able to restrict to setting it when
>> CONFIG_BLK_DEV_WRITE_MOUNTED is disabled to avoid the problem of non-fs
>> shared mmap writers, as anyone using a modern kernel and cgroups really
>> should have that disabled.
>
> I don't think that's really needed. If such writes are happening, they are
> *very* limited and done by a system administrator only (as they are very
> dangerous) so the fact the writes will get charged to the root cgroup is
> not an issue.
Hi, Jan, Christoph.
Thanks for your suggestions.
One remaining concern: while ext4 is mounted, buffered reads of the raw
block device would also charge page-cache allocations to the root memcg,
bypassing the reader's memcg memory limits. Is that acceptable?
I've attached a preliminary diff. Does this approach look reasonable?
If so, I'll split this into a formal patch series.
diff --git a/block/bdev.c b/block/bdev.c
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -79,7 +79,7 @@ static void bdev_write_inode(struct bloc
}
/* Kill _all_ buffers and pagecache , dirty or not.. */
-static void kill_bdev(struct block_device *bdev)
+void kill_bdev(struct block_device *bdev)
{
struct address_space *mapping = bdev->bd_mapping;
@@ -89,6 +89,7 @@ static void kill_bdev(struct block_devic
invalidate_bh_lrus();
truncate_inode_pages(mapping, 0);
}
+EXPORT_SYMBOL_GPL(kill_bdev);
/* Invalidate clean unused buffers and pagecache. */
void invalidate_bdev(struct block_device *bdev)
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1706,6 +1706,7 @@ bool disk_live(struct gendisk *disk);
unsigned int block_size(struct block_device *bdev);
#ifdef CONFIG_BLOCK
+void kill_bdev(struct block_device *bdev);
void invalidate_bdev(struct block_device *bdev);
int sync_blockdev(struct block_device *bdev);
int sync_blockdev_range(struct block_device *bdev, loff_t lstart, loff_t lend);
diff --git a/fs/ext4/super.c b/fs/ext4/super.c
--- a/fs/ext4/super.c
+++ b/fs/ext4/super.c
@@ -1279,6 +1279,56 @@ static void ext4_flex_groups_free(struct
}
}
+/*
+ * Account ext4 metadata cache to the root cgroup. This filesystem-wide
+ * metadata is shared across users, so it is not naturally attributable
+ * to a single user cgroup.
+ *
+ * A bdev inode can contain metadata pages charged to many memory cgroups.
+ * Dirtying these pages can trigger foreign writeback of its entire owner wb,
+ * including unrelated file data. Premature writeback of buffered overwrites
+ * reduces write coalescing, increasing disk traffic and consuming bandwidth
+ * that could serve foreground direct I/O. This can hurt performance,
+ * particularly on bandwidth-limited devices such as HDDs.
+ *
+ * This does not depend on CONFIG_BLK_DEV_WRITE_MOUNTED. Writing directly to
+ * a mounted block device is dangerous and intended for system administration.
+ * Charging the cache from these writes to the root cgroup is acceptable.
+ */
+static void ext4_bdev_set_kernel_file(struct block_device *bdev, bool enable)
+{
+ struct address_space *mapping = bdev->bd_mapping;
+ struct inode *inode = mapping->host;
+
+ inode_lock(inode);
+ filemap_invalidate_lock(mapping);
+
+ if (!!test_bit(AS_KERNEL_FILE, &mapping->flags) == enable)
+ goto out;
+
+ /*
+ * Folio removal uses the mapping's current accounting mode. Remove old
+ * folios before changing it, excluding new raw-I/O cache allocations.
+ */
+ sync_blockdev(bdev);
+ kill_bdev(bdev);
+
+ if (enable)
+ set_bit(AS_KERNEL_FILE, &mapping->flags);
+ else
+ clear_bit(AS_KERNEL_FILE, &mapping->flags);
+out:
+ filemap_invalidate_unlock(mapping);
+ inode_unlock(inode);
+}
+
static void ext4_put_super(struct super_block *sb)
{
struct ext4_sb_info *sbi = EXT4_SB(sb);
@@ -1387,6 +1437,7 @@ static void ext4_put_super(struct super_
#if IS_ENABLED(CONFIG_UNICODE)
utf8_unload(sb->s_encoding);
#endif
+ ext4_bdev_set_kernel_file(sb->s_bdev, false);
kfree(sbi);
}
@@ -5852,6 +5903,8 @@ static int ext4_fill_super(struct super_
if (ctx->spec & EXT4_SPEC_s_sb_block)
sbi->s_sb_block = ctx->s_sb_block;
+ ext4_bdev_set_kernel_file(sb->s_bdev, true);
+
ret = __ext4_fill_super(fc, sb);
if (ret < 0)
goto free_sbi;
@@ -5877,6 +5930,7 @@ static int ext4_fill_super(struct super_
return 0;
free_sbi:
+ ext4_bdev_set_kernel_file(sb->s_bdev, false);
ext4_free_sbi(sbi);
fc->s_fs_info = NULL;
return ret;
>
> Honza
Thanks,
--
Julian Sun <sunjunchao@bytedance.com>
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-10-08 8:00 ` Julian Sun
@ 2026-10-08 8:51 ` Jan Kara
2026-10-08 9:39 ` Christoph Hellwig
0 siblings, 1 reply; 19+ messages in thread
From: Jan Kara @ 2026-10-08 8:51 UTC (permalink / raw)
To: Julian Sun
Cc: Jan Kara, Christoph Hellwig, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Thu 08-10-26 16:00:13, Julian Sun wrote:
> On 10/6/26 12:25 AM, Jan Kara wrote:
> > On Mon 05-10-26 01:26:28, Christoph Hellwig wrote:
> >> On Fri, Oct 02, 2026 at 01:10:42PM +0200, Jan Kara wrote:
> >>> FWIW I share Julian's concern here. It would be fine to use AS_KERNEL_FILE
> >>> for ext4 metadata but I think just unconditionally setting AS_KERNEL_FILE
> >>> for bdev mappings will cause issues because some users may be using bdevs
> >>> directly for their workloads and they could still expect proper memcg
> >>> accounting to work in that case.
> >>
> >> Agreed that it should not set unconditionally.
> >>
> >>> We could set AS_KERNEL_FILE when opening bdev for a filesystem (and remove
> >>> it when releasing bdev) but that would have to make sure there are no
> >>> folios in the bdev mapping when changing the flag as otherwise the
> >>> accounting would go wrong. Looks it might be doable but getting all the
> >>> cornercases right will be hairy and overall not very appealing to me...
> >>
> >> I'd rather not support special case writeback code just for this legacy
> >> fs abuses bdev buffer cache case. And we basically need to tear
> >> down pagecache at unmount anyway, as i_blkbits can change, so while
> >> we do need to be careful, I don't think it really is a major issue.
> >
> > OK, after some more thought yes, I think we can make that work.
> >
> >> We might be able to restrict to setting it when
> >> CONFIG_BLK_DEV_WRITE_MOUNTED is disabled to avoid the problem of non-fs
> >> shared mmap writers, as anyone using a modern kernel and cgroups really
> >> should have that disabled.
> >
> > I don't think that's really needed. If such writes are happening, they are
> > *very* limited and done by a system administrator only (as they are very
> > dangerous) so the fact the writes will get charged to the root cgroup is
> > not an issue.
>
> Hi, Jan, Christoph.
>
> Thanks for your suggestions.
>
> One remaining concern: while ext4 is mounted, buffered reads of the raw
> block device would also charge page-cache allocations to the root memcg,
> bypassing the reader's memcg memory limits. Is that acceptable?
Yes. If you have a setup where unprivileged user can read raw block device
with the filesystem mounted, your security model is very doubtful. I don't
think anybody sane can do that. Also setting up of loop devices and
mounting filesystem on top is privileged operation and if a user can do it,
memcg accounting is the least of your troubles. Of course even privileged
process can be constrained in a cgroup and this containment will not
account bdev pages properly in the case you describe but even then I have
hard time coming up with a tool that would read a significant portion of
the device (e.g. udev will be reading mounted bdevs but only small portions
of them) and even if it does, this would be a clean page cache which is
easy to reclaim.
So I cannot think of a practical setup where this charging behavior could be
a problem. Of course I could be proven wrong by some user complaining and
then we'd have to revert and go back to the drawing board.
> I've attached a preliminary diff. Does this approach look reasonable?
> If so, I'll split this into a formal patch series.
In principle, the mechanism looks good to me but I'd place it in
fs_bdev_file_open_by_dev() & fs_bdev_file_release() so that all block
device based filesystems are handled. ext4 is the most used filesystem that
has this problem with bdev memcg accounting but in principle exactly the
same problem happens for other filesystems using buffer heads so I'd prefer
handling it in the generic code instead of fixing them one-by-one as people
complain...
Honza
>
> diff --git a/block/bdev.c b/block/bdev.c
> --- a/block/bdev.c
> +++ b/block/bdev.c
> @@ -79,7 +79,7 @@ static void bdev_write_inode(struct bloc
> }
>
> /* Kill _all_ buffers and pagecache , dirty or not.. */
> -static void kill_bdev(struct block_device *bdev)
> +void kill_bdev(struct block_device *bdev)
> {
> struct address_space *mapping = bdev->bd_mapping;
>
> @@ -89,6 +89,7 @@ static void kill_bdev(struct block_devic
> invalidate_bh_lrus();
> truncate_inode_pages(mapping, 0);
> }
> +EXPORT_SYMBOL_GPL(kill_bdev);
>
> /* Invalidate clean unused buffers and pagecache. */
> void invalidate_bdev(struct block_device *bdev)
> diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
> --- a/include/linux/blkdev.h
> +++ b/include/linux/blkdev.h
> @@ -1706,6 +1706,7 @@ bool disk_live(struct gendisk *disk);
> unsigned int block_size(struct block_device *bdev);
>
> #ifdef CONFIG_BLOCK
> +void kill_bdev(struct block_device *bdev);
> void invalidate_bdev(struct block_device *bdev);
> int sync_blockdev(struct block_device *bdev);
> int sync_blockdev_range(struct block_device *bdev, loff_t lstart, loff_t lend);
> diff --git a/fs/ext4/super.c b/fs/ext4/super.c
> --- a/fs/ext4/super.c
> +++ b/fs/ext4/super.c
> @@ -1279,6 +1279,56 @@ static void ext4_flex_groups_free(struct
> }
> }
>
> +/*
> + * Account ext4 metadata cache to the root cgroup. This filesystem-wide
> + * metadata is shared across users, so it is not naturally attributable
> + * to a single user cgroup.
> + *
> + * A bdev inode can contain metadata pages charged to many memory cgroups.
> + * Dirtying these pages can trigger foreign writeback of its entire owner wb,
> + * including unrelated file data. Premature writeback of buffered overwrites
> + * reduces write coalescing, increasing disk traffic and consuming bandwidth
> + * that could serve foreground direct I/O. This can hurt performance,
> + * particularly on bandwidth-limited devices such as HDDs.
> + *
> + * This does not depend on CONFIG_BLK_DEV_WRITE_MOUNTED. Writing directly to
> + * a mounted block device is dangerous and intended for system administration.
> + * Charging the cache from these writes to the root cgroup is acceptable.
> + */
> +static void ext4_bdev_set_kernel_file(struct block_device *bdev, bool enable)
> +{
> + struct address_space *mapping = bdev->bd_mapping;
> + struct inode *inode = mapping->host;
> +
> + inode_lock(inode);
> + filemap_invalidate_lock(mapping);
> +
> + if (!!test_bit(AS_KERNEL_FILE, &mapping->flags) == enable)
> + goto out;
> +
> + /*
> + * Folio removal uses the mapping's current accounting mode. Remove old
> + * folios before changing it, excluding new raw-I/O cache allocations.
> + */
> + sync_blockdev(bdev);
> + kill_bdev(bdev);
> +
> + if (enable)
> + set_bit(AS_KERNEL_FILE, &mapping->flags);
> + else
> + clear_bit(AS_KERNEL_FILE, &mapping->flags);
> +out:
> + filemap_invalidate_unlock(mapping);
> + inode_unlock(inode);
> +}
> +
> static void ext4_put_super(struct super_block *sb)
> {
> struct ext4_sb_info *sbi = EXT4_SB(sb);
> @@ -1387,6 +1437,7 @@ static void ext4_put_super(struct super_
> #if IS_ENABLED(CONFIG_UNICODE)
> utf8_unload(sb->s_encoding);
> #endif
> + ext4_bdev_set_kernel_file(sb->s_bdev, false);
> kfree(sbi);
> }
>
> @@ -5852,6 +5903,8 @@ static int ext4_fill_super(struct super_
> if (ctx->spec & EXT4_SPEC_s_sb_block)
> sbi->s_sb_block = ctx->s_sb_block;
>
> + ext4_bdev_set_kernel_file(sb->s_bdev, true);
> +
> ret = __ext4_fill_super(fc, sb);
> if (ret < 0)
> goto free_sbi;
> @@ -5877,6 +5930,7 @@ static int ext4_fill_super(struct super_
> return 0;
>
> free_sbi:
> + ext4_bdev_set_kernel_file(sb->s_bdev, false);
> ext4_free_sbi(sbi);
> fc->s_fs_info = NULL;
> return ret;
>
>
> >
> > Honza
>
> Thanks,
> --
> Julian Sun <sunjunchao@bytedance.com>
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v9 1/3] block: introduce bdev_flush_by_dev()
2026-10-08 8:51 ` Jan Kara
@ 2026-10-08 9:39 ` Christoph Hellwig
0 siblings, 0 replies; 19+ messages in thread
From: Christoph Hellwig @ 2026-10-08 9:39 UTC (permalink / raw)
To: Jan Kara
Cc: Julian Sun, Christoph Hellwig, linux-block, cgroups, linux-mm,
linux-fsdevel, axboe, hannes, mhocko, roman.gushchin,
shakeel.butt, muchun.song, willy, tj, akpm, Boris Burkov
On Thu, Oct 08, 2026 at 10:51:08AM +0200, Jan Kara wrote:
> In principle, the mechanism looks good to me but I'd place it in
> fs_bdev_file_open_by_dev() & fs_bdev_file_release() so that all block
> device based filesystems are handled. ext4 is the most used filesystem that
> has this problem with bdev memcg accounting but in principle exactly the
> same problem happens for other filesystems using buffer heads so I'd prefer
> handling it in the generic code instead of fixing them one-by-one as people
> complain...
Yes, this should be in the common code.
^ permalink raw reply [flat|nested] 19+ messages in thread
end of thread, other threads:[~2026-10-08 9:39 UTC | newest]
Thread overview: 19+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 6:44 [PATCH v9 0/3] memcg,writeback: flush foreign bdev mappings separately Julian Sun
2026-09-25 6:44 ` [PATCH v9 1/3] block: introduce bdev_flush_by_dev() Julian Sun
2026-09-25 6:52 ` Christoph Hellwig
2026-09-25 10:34 ` Jan Kara
2026-09-28 6:24 ` Christoph Hellwig
2026-09-28 9:51 ` Julian Sun
2026-10-02 11:10 ` Jan Kara
2026-10-05 8:26 ` Christoph Hellwig
2026-10-05 16:25 ` Jan Kara
2026-10-08 8:00 ` Julian Sun
2026-10-08 8:51 ` Jan Kara
2026-10-08 9:39 ` Christoph Hellwig
2026-10-02 9:23 ` Julian Sun
2026-10-05 8:27 ` Christoph Hellwig
2026-09-25 6:44 ` [PATCH v9 2/3] memcg,writeback: flush foreign bdev mappings separately from owner wbs Julian Sun
2026-09-25 6:55 ` Christoph Hellwig
2026-09-25 13:00 ` Julian Sun
2026-09-28 6:18 ` Christoph Hellwig
2026-09-25 6:44 ` [PATCH v9 3/3] writeback: add tracepoints for foreign bdev writeback Julian Sun
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox