All of lore.kernel.org
 help / color / mirror / Atom feed
From: Andrew Morton <akpm@linux-foundation.org>
To: "Patrick Lu (Anthropic)" <perf.patrick.lu@gmail.com>
Cc: Alexander Viro <viro@zeniv.linux.org.uk>,
	Christian Brauner <brauner@kernel.org>, Jan Kara <jack@suse.cz>,
	Dennis Zhou <dennis@kernel.org>,
	Roman Gushchin <roman.gushchin@linux.dev>,
	Tejun Heo <tj@kernel.org>,
	"Matthew Wilcox (Oracle)" <willy@infradead.org>,
	linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org,
	stable@vger.kernel.org
Subject: Re: [PATCH v2] writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
Date: Fri, 11 Sep 2026 12:32:01 -0700	[thread overview]
Message-ID: <20260911123201.8fd2cc28abcbf6bb794b4ce8@linux-foundation.org> (raw)
In-Reply-To: <20260911-wb-cgwb-rotate-v2-1-a9ab253a1295@gmail.com>

On Fri, 11 Sep 2026 18:49:49 +0000 "Patrick Lu (Anthropic)" <perf.patrick.lu@gmail.com> wrote:

> cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
> per call and is called again until the dying wb is drained, but every
> call walks wb->b_attached and then wb->b_dirty_time from the same end.
> Inodes already prepared (they stay on the list with I_WB_SWITCH set
> until the switch worker runs) and inodes that cannot be switched
> (I_FREEING, I_WILL_FREE, !SB_ACTIVE, DAX, already on the target wb)
> stay where they are, so each pass rescans a growing run of them under
> wb->list_lock and a full drain is quadratic in the number of inodes on
> the list. With ~17M inodes attached to one dying cgwb we saw this end
> in soft lockups, with CPUs reported stuck for 21-48s.

Well.  "quadratic" is a trigger word around here.  Even if it's O(n),
someone will hit it.  Thanks for working on this.

> Walk both lists from the oldest end and move every scanned inode to
> the newest end, so the next pass starts where the previous one stopped
> and the drain becomes linear. b_attached is unordered, so nobody sees
> the reorder there. b_dirty_time is ordered by dirtied_when, but the
> oldest unscanned inode stays at the end move_expired_inodes() picks
> from, sync takes the whole list regardless of order, and prepared
> inodes leave the list as soon as the switch work runs and get a new
> dirtied_time_when on the new wb anyway, so the only inodes left out of
> order are the ones that can never switch (DAX), and only on the dying
> wb.
> 
> Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
> Cc: stable@vger.kernel.org

It sounds like this.  Maintainers, please lmk if you feel backporting is not
justified.  It's a five-year-old thing.

> ---
> Seen in production on a 6.18-based kernel: with ~17M inodes attached
> to one dying cgwb, a node spent 36 minutes in back-to-back
> wb->list_lock holds by the cleanup scanner (~6ms each, ~46% of wall
> time, starving writeback on that wb); with v1 of this patch the same
> workload drains in ~30 seconds. Also seen on stock Amazon Linux 2023
> 6.12.68 as isw workers spinning on the list_lock in
> inode_switch_wbs_work_fn() while cleanup_offline_cgwbs_workfn() runs.

This is super-important info and it deserves to be above the ---.  In
fact it deserves to become the first paragraph.


      reply	other threads:[~2026-09-11 19:32 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-11 18:49 [PATCH v2] writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes Patrick Lu (Anthropic)
2026-09-11 19:32 ` Andrew Morton [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260911123201.8fd2cc28abcbf6bb794b4ce8@linux-foundation.org \
    --to=akpm@linux-foundation.org \
    --cc=brauner@kernel.org \
    --cc=dennis@kernel.org \
    --cc=jack@suse.cz \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=perf.patrick.lu@gmail.com \
    --cc=roman.gushchin@linux.dev \
    --cc=stable@vger.kernel.org \
    --cc=tj@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.