Linux block layer
 help / color / mirror / Atom feed
From: Julian Sun <sunjunchao@bytedance.com>
To: linux-block@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	gfs2@lists.linux.dev, linux-security-module@vger.kernel.org
Cc: jack@suse.cz, agruenba@redhat.com, mic@digikod.net,
	gnoack@google.com, paul@paul-moore.com, jmorris@namei.org,
	serge@hallyn.com, aleksa@amutable.com, legion@kernel.org,
	djwong@kernel.org, ebiggers@kernel.org, sandeen@redhat.com
Subject: [PATCH 0/7] fs: preserve superblock inode walk positions across lock drops
Date: Wed,  9 Sep 2026 17:01:05 +0800	[thread overview]
Message-ID: <20260909090112.790006-1-sunjunchao@bytedance.com> (raw)

[Motivation]

We observed hung tasks in production during disk hotplug operations. A
kernel thread spends a long time in evict_inodes() while holding s_umount,
blocking other users of that lock and causing further stalls.

The problem is that evict_inodes() restarts its walk from the head of
sb->s_inodes every time it reschedules. When a large number of referenced
inodes remain near the head, each restart scans those inodes again without
making progress through that part of the list. The repeated scans can
delay eviction long enough to trigger hung-task reports. This is the same
problem that [1] attempted to address.

[Approach]

This series introduces sb_for_each_inodes() for two purposes:

  1. Consolidate open-coded s_inodes walks behind a common entry point.
  2. Retain each walk's position across drops of s_inode_list_lock.

The iterator mechanism follows the approach used by cgroup task iteration,
such as css_task_iter_next(). Active iterators are registered on a separate
list, sb->s_inodes_iters. Before removing an inode from s_inodes, the
removal path advances any iterator whose next position points to that
inode. These updates are protected by s_inode_list_lock, so a walker can
drop the lock and later resume from its saved position.

Existing walkers, such as drop_pagecache_sb() and add_dquot_ref(), already
contain their own position-preserving logic: they carry an inode reference
across iterations so that they can resume after dropping the list lock.
Moving that responsibility into sb_for_each_inodes() simplifies these
callers and lets their callbacks focus on the per-inode work.

Patch 1 removes trailing whitespace from include/linux/fs.h.
Patch 2 introduces sb_for_each_inodes().
The remaining patches convert existing walks to the new interface.
remove_dquot_ref() and nr_blockdev_pages() are left unchanged: their
walks are simple and do not require the inode->i_lock locking imposed
by the callback interface. Converting them would add an unnecessary
lock/unlock overhead for every inode.

[Testing]

I tested this series with approximately 20 hours of xfstests case
execution, repeatedly running the auto group on ext4 and XFS, no new
issues were observed. And with this patch applied, the hung task that
previously occurred on every run no longer occurs.

[1] https://lore.kernel.org/all/20241118114508.1405494-1-yebin@huaweicloud.com/

Julian Sun (7):
  fs: remove trailing whitespace from include/linux/fs.h
  fs: introduce sb_for_each_inodes().
  block: use sb_for_each_inodes() in sync_bdevs()
  fs: use sb_for_each_inodes() API.
  gfs2: use sb_for_each_inodes() for cooperative eviction
  quota: use sb_for_each_inodes() in add_dquot_ref()
  landlock: use sb_for_each_inodes() when detaching a superblock

 block/bdev.c                   |  85 +++++++++----------
 fs/drop_caches.c               |  44 +++++-----
 fs/gfs2/ops_fstype.c           |  38 ++++-----
 fs/inode.c                     | 142 +++++++++++++++++++++++--------
 fs/quota/dquot.c               |  72 ++++++----------
 fs/super.c                     |   1 +
 include/linux/fs.h             |  29 +++++--
 include/linux/fs/super_types.h |   3 +-
 security/landlock/fs.c         | 150 +++++++++++++--------------------
 9 files changed, 295 insertions(+), 269 deletions(-)

-- 
2.39.5


             reply	other threads:[~2026-09-09  9:01 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-09  9:01 Julian Sun [this message]
2026-09-09  9:01 ` [PATCH 1/7] fs: remove trailing whitespace from include/linux/fs.h Julian Sun
2026-09-10 16:52   ` Jan Kara
2026-09-09  9:01 ` [PATCH 2/7] fs: introduce sb_for_each_inodes() Julian Sun
2026-09-10 17:47   ` Jan Kara
2026-09-11  3:34     ` [External] " Julian Sun
2026-09-11  3:35     ` Julian Sun
2026-09-09  9:01 ` [PATCH 3/7] block: use sb_for_each_inodes() in sync_bdevs() Julian Sun
2026-09-09  9:01 ` [PATCH 4/7] fs: use sb_for_each_inodes() API Julian Sun
2026-09-09  9:01 ` [PATCH 5/7] gfs2: use sb_for_each_inodes() for cooperative eviction Julian Sun
2026-09-09  9:01 ` [PATCH 6/7] quota: use sb_for_each_inodes() in add_dquot_ref() Julian Sun
2026-09-09  9:01 ` [PATCH 7/7] landlock: use sb_for_each_inodes() when detaching a superblock Julian Sun
2026-09-09 12:49 ` [PATCH 0/7] fs: preserve superblock inode walk positions across lock drops Jan Kara
2026-09-09 13:08   ` [External] " Julian Sun

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260909090112.790006-1-sunjunchao@bytedance.com \
    --to=sunjunchao@bytedance.com \
    --cc=agruenba@redhat.com \
    --cc=aleksa@amutable.com \
    --cc=djwong@kernel.org \
    --cc=ebiggers@kernel.org \
    --cc=gfs2@lists.linux.dev \
    --cc=gnoack@google.com \
    --cc=jack@suse.cz \
    --cc=jmorris@namei.org \
    --cc=legion@kernel.org \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-security-module@vger.kernel.org \
    --cc=mic@digikod.net \
    --cc=paul@paul-moore.com \
    --cc=sandeen@redhat.com \
    --cc=serge@hallyn.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox