Linux filesystem development
 help / color / mirror / Atom feed
* [PATCH] writeback: bound cleanup_offline_cgwb() rescans by rotating b_attached
@ 2026-09-09 18:50 Patrick Lu (Anthropic)
  2026-09-09 20:04 ` Tejun Heo
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Patrick Lu (Anthropic) @ 2026-09-09 18:50 UTC (permalink / raw)
  To: Alexander Viro, Christian Brauner, Jan Kara, Roman Gushchin,
	Tejun Heo, Matthew Wilcox (Oracle), Andrew Morton
  Cc: Dennis Zhou, linux-fsdevel, linux-kernel, stable,
	Patrick Lu (Anthropic)

cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
per call and is called again until the dying wb is drained, but every
call walks wb->b_attached from the head. Inodes already prepared (they
stay on the list with I_WB_SWITCH set until the switch worker runs) and
inodes that cannot be switched (I_FREEING, I_WILL_FREE, !SB_ACTIVE,
DAX, already on the target wb) stay at the head, so each pass rescans a
growing prefix under wb->list_lock and a full drain is quadratic in the
number of attached inodes. With ~17M inodes attached to one dying cgwb
we have seen this end in soft lockups, with CPUs reported stuck for
21-48s.

Move every scanned inode to the tail of b_attached, so the next pass
starts where the previous one stopped and the drain becomes linear.
b_attached is unordered and isw_prepare_wbs_switch() is its only
walker, so nobody else sees the reorder. b_dirty_time is ordered by
expiry for move_expired_inodes() and keeps its current scan.

Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Patrick Lu (Anthropic) <perf.patrick.lu@gmail.com>
---
Seen in production on a 6.18-based kernel: with ~17M inodes attached
to one dying cgwb, a node spent 36 minutes in back-to-back
wb->list_lock holds by the cleanup scanner (~6ms each, ~46% of wall
time, starving writeback on that wb); with this patch the same workload
drains in ~30 seconds. Also seen on stock Amazon Linux 2023 6.12.68 as
isw workers spinning on the list_lock in inode_switch_wbs_work_fn()
while cleanup_offline_cgwbs_workfn() runs.

Tested with a QEMU A/B setup at 100k attached inodes and patched vs
unpatched on production-class hardware at ~17M attached inodes.

Josef Bacik's patch making the drain loop report a Tasks-RCU quiescent
state [1] fixes BPF/ftrace detach stalls behind the same drain; this
patch bounds the walk itself. The two are independent.

[1] https://lore.kernel.org/linux-mm/20260909-cgwb-tasks-rcu-qs-v1-1-967a7754771f@toxicpanda.com/
---
 fs/fs-writeback.c | 39 ++++++++++++++++++++++++++++-----------
 1 file changed, 28 insertions(+), 11 deletions(-)

diff --git a/fs/fs-writeback.c b/fs/fs-writeback.c
index e744f9f9d43f..69a452b12b12 100644
--- a/fs/fs-writeback.c
+++ b/fs/fs-writeback.c
@@ -725,21 +725,37 @@ static void inode_switch_wbs(struct inode *inode, int new_wb_id)
 
 static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb,
 				   struct inode_switch_wbs_context *isw,
-				   struct list_head *list, int *nr)
+				   struct list_head *list, bool rotate, int *nr)
 {
-	struct inode *inode;
+	struct inode *inode, *tmp;
+	LIST_HEAD(scanned);
+	bool full = false;
+
+	list_for_each_entry_safe(inode, tmp, list, i_io_list) {
+		/*
+		 * Rotate scanned inodes to the tail so the next scan resumes
+		 * at unscanned ones instead of re-walking an ever-growing
+		 * prefix of prepared and skipped inodes.  b_dirty_time is
+		 * expiry-ordered and so must not be rotated.
+		 */
+		if (rotate)
+			list_move_tail(&inode->i_io_list, &scanned);
 
-	list_for_each_entry(inode, list, i_io_list) {
 		if (!inode_prepare_wbs_switch(inode, new_wb))
 			continue;
 
 		isw->inodes[*nr] = inode;
 		(*nr)++;
 
-		if (*nr >= WB_MAX_INODES_PER_ISW - 1)
-			return true;
+		if (*nr >= WB_MAX_INODES_PER_ISW - 1) {
+			full = true;
+			break;
+		}
 	}
-	return false;
+	if (rotate)
+		list_splice_tail(&scanned, list);
+
+	return full;
 }
 
 /**
@@ -747,8 +763,9 @@ static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb,
  * @wb: target wb
  *
  * Switch all inodes attached to @wb to a nearest living ancestor's wb in order
- * to eventually release the dying @wb.  Returns %true if not all inodes were
- * switched and the function has to be restarted.
+ * to eventually release the dying @wb.  Returns %true if the scan stopped
+ * early after making progress; the caller should call again to continue
+ * draining.
  */
 bool cleanup_offline_cgwb(struct bdi_writeback *wb)
 {
@@ -783,13 +800,13 @@ bool cleanup_offline_cgwb(struct bdi_writeback *wb)
 	 * bandwidth restrictions, as writeback of inode metadata is not
 	 * accounted for.
 	 */
-	restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, &nr);
+	restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_attached, true, &nr);
 	if (!restart)
 		restart = isw_prepare_wbs_switch(new_wb, isw, &wb->b_dirty_time,
-						 &nr);
+						 false, &nr);
 	spin_unlock(&wb->list_lock);
 
-	/* no attached inodes? bail out */
+	/* nothing to switch? bail out */
 	if (nr == 0) {
 		atomic_dec(&isw_nr_in_flight);
 		wb_put(new_wb);

---
base-commit: e14d4302cbd0de773960bec33c2281508c8d8855
change-id: 20260909-wb-cgwb-rotate-f17a75facfdc

Best regards,
--  
Patrick Lu (Anthropic) <perf.patrick.lu@gmail.com>


^ permalink raw reply related	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-11  9:35 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-09 18:50 [PATCH] writeback: bound cleanup_offline_cgwb() rescans by rotating b_attached Patrick Lu (Anthropic)
2026-09-09 20:04 ` Tejun Heo
2026-09-09 21:19 ` Roman Gushchin
2026-09-10 11:24 ` Jan Kara
2026-09-10 23:47   ` Patrick Lu (Anthropic)
2026-09-11  9:35     ` Jan Kara

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox