Linux EXT4 FS development
 help / color / mirror / Atom feed
* [PATCH v5] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan()
@ 2026-10-01  7:43 Qiliang Yuan
  2026-10-01  7:53 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Qiliang Yuan @ 2026-10-01  7:43 UTC (permalink / raw)
  To: Theodore Ts'o, Jan Kara
  Cc: linux-ext4, linux-kernel, stable, Qiliang Yuan, Jan Kara

do_shrink_slab() reads count_objects() once per invocation to derive
a one-shot scan budget, then calls scan_objects() repeatedly until
that budget is exhausted or scan_objects() returns SHRINK_STOP.
include/linux/shrinker.h documents that scan_objects() "should track
its actual progress" in sc->nr_scanned, so do_shrink_slab() can tell
when there is nothing left to examine and stop early.

jbd2_journal_shrink_scan() never updates sc->nr_scanned, so it
defaults to the full sc->nr_to_scan on every call. do_shrink_slab()
therefore always believes a full batch was examined, regardless of
what jbd2_journal_shrink_checkpoint_list() actually did, and keeps
calling scan_objects() until the budget derived from the (possibly
stale) percpu checkpoint count is drained, even after the checkpoint
list has already been fully drained.

jbd2_journal_shrink_checkpoint_list() already tracks examined
buffers accurately: it decrements its nr_to_scan in/out parameter
for every journal_head walked, whether or not it gets freed. Derive
sc->nr_scanned from the difference, and return SHRINK_STOP once it
comes back zero, since that only happens when the checkpoint list
had nothing left to walk. Key this off sc->nr_scanned rather than
nr_shrunk: a batch can find every buffer in its transactions busy
and free none, while later transactions may still hold buffers
whose writeback has completed; sc->nr_scanned reflects that real
work either way.

Tested by creating 20000 small files, running sync, and letting
checkpointing settle for a few seconds (with systemd-journald and
cron stopped, since their own writes to the same filesystem keep
re-populating the checkpoint list and mask the effect being
measured), then triggering "echo 2 > /proc/sys/vm/drop_caches" while
tracing jbd2_shrink_scan_enter/exit:

                      total scan_objects()   calls with
                      calls                  nr_scanned == 0
  before this patch   22                     10 (45%)
  after this patch     8                      1 (12%)

Fixes: 4ba3fcdde7e3 ("jbd2,ext4: add a shrinker to release checkpointed buffers")
Cc: stable@vger.kernel.org
Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
Reviewed-by: Jan Kara <jack@suse.cz>
---
V4 -> V5:
- Revert to keying SHRINK_STOP off sc->nr_scanned (buffers actually
  examined) rather than nr_shrunk (buffers actually freed), as v1
  originally did. Jan Kara rejected the nr_shrunk-based v4, pointing
  out that failing to reclaim the buffers in one batch doesn't mean
  later transactions won't have freeable ones; Sashiko AI review
  independently flagged the same issue. Restore Jan Kara's v1
  Reviewed-by: the diff in jbd2_journal_shrink_scan() is unchanged
  from what he reviewed there (only the explanatory comment reads
  differently).
- Redesign the test: the old methodology (no sync, buffers kept busy
  by design) exercises exactly the busy-but-productive case Jan Kara
  describes, so it isn't a useful before/after comparison for this
  change. Replace it with one that lets the checkpoint list actually
  drain via sync, then measures scan_objects() calls wasted after
  draining due to percpu-counter staleness.

V3 -> V4:
- Shorten the SHRINK_STOP comment; no functional change (the trigger
  has been nr_shrunk == 0 since v2, not sc->nr_scanned).
- Carry forward Jan Kara's v3 Reviewed-by.

V2 -> V3:
- Rebase onto v7.3-rc5. It already includes Max Kellermann's
  "jbd2: bound shrinker scans by examined checkpoint buffers"
  (15cb16496446), which independently fixes the busy-buffer
  accounting in journal_shrink_one_cp_list()/
  jbd2_journal_shrink_checkpoint_list() that v2 also touched. Drop
  the now-redundant checkpoint.c changes; this revision only touches
  journal.c, reusing the accurate nr_to_scan tracking Kellermann's
  fix already provides.
- Add a Fixes: tag for the commit that introduced
  jbd2_journal_shrink_scan() without ever setting sc->nr_scanned.
- Re-measure test data against the new baseline (176 -> 8 calls,
  vs the old baseline's 168 -> 7).

V1 -> V2:
- Count examined buffers, not just freed ones, in
  journal_shrink_one_cp_list() (Sashiko AI review finding).
- Trigger SHRINK_STOP on nr_shrunk == 0 instead of
  sc->nr_scanned == 0.

v4: https://lore.kernel.org/r/20260930-fix-jbd2-shrink-scan-nr-scanned-v4-1-eb9a3ccc2524@gmail.com
v3: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v3-1-916caaa53420@gmail.com
v2: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v2-1-7e1efec3afdf@gmail.com
v1: https://lore.kernel.org/r/20260928-fix-jbd2-shrink-scan-nr-scanned-v1-1-e6f4016ec699@gmail.com
---
 fs/jbd2/journal.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/fs/jbd2/journal.c b/fs/jbd2/journal.c
index 00f5a98f3d4fe..eac5f84a1afc5 100644
--- a/fs/jbd2/journal.c
+++ b/fs/jbd2/journal.c
@@ -1263,10 +1263,22 @@ static unsigned long jbd2_journal_shrink_scan(struct shrinker *shrink,
 	trace_jbd2_shrink_scan_enter(journal, sc->nr_to_scan, count);
 
 	nr_shrunk = jbd2_journal_shrink_checkpoint_list(journal, &nr_to_scan);
+	sc->nr_scanned = sc->nr_to_scan - nr_to_scan;
 
 	count = percpu_counter_read_positive(&journal->j_checkpoint_jh_count);
 	trace_jbd2_shrink_scan_exit(journal, nr_to_scan, nr_shrunk, count);
 
+	/*
+	 * Key SHRINK_STOP off sc->nr_scanned (buffers actually examined),
+	 * not nr_shrunk (buffers actually freed): a batch can find every
+	 * buffer in its transactions busy and free none, while later
+	 * transactions may still hold buffers whose writeback has
+	 * completed. sc->nr_scanned only comes back 0 when the checkpoint
+	 * list had nothing left to walk.
+	 */
+	if (sc->nr_scanned == 0)
+		return SHRINK_STOP;
+
 	return nr_shrunk;
 }
 

---
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
change-id: 20260928-fix-jbd2-shrink-scan-nr-scanned-c3098cd901e2

Best regards,
-- 
Qiliang Yuan <odys.yuan@gmail.com>


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH v5] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan()
  2026-10-01  7:43 [PATCH v5] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan() Qiliang Yuan
@ 2026-10-01  7:53 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-01  7:53 UTC (permalink / raw)
  To: Qiliang Yuan; +Cc: linux-ext4, tytso

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] jbd2: shrinker busy-loops rescanning the same busy checkpoint items
--

commit 22efbdb58b6c52d305510bf987737313bb851b41
Author: Qiliang Yuan <odys.yuan@gmail.com>

jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan()

This patch updates jbd2_journal_shrink_scan() to report the actual number of
scanned buffers back to the shrinker subsystem. It sets sc->nr_scanned based
on how many buffers jbd2_journal_shrink_checkpoint_list() examined, and uses
this value to determine whether to return SHRINK_STOP.

> diff --git a/fs/jbd2/journal.c b/fs/jbd2/journal.c
> index 00f5a98f3d4fe..eac5f84a1afc5 100644
> --- a/fs/jbd2/journal.c
> +++ b/fs/jbd2/journal.c
> @@ -1263,10 +1263,22 @@ static unsigned long jbd2_journal_shrink_scan(struct shrinker *shrink,
>  	trace_jbd2_shrink_scan_enter(journal, sc->nr_to_scan, count);
>  
>  	nr_shrunk = jbd2_journal_shrink_checkpoint_list(journal, &nr_to_scan);
> +	sc->nr_scanned = sc->nr_to_scan - nr_to_scan;
>  
>  	count = percpu_counter_read_positive(&journal->j_checkpoint_jh_count);
>  	trace_jbd2_shrink_scan_exit(journal, nr_to_scan, nr_shrunk, count);
>  
> +	/*
> +	 * Key SHRINK_STOP off sc->nr_scanned (buffers actually examined),
> +	 * not nr_shrunk (buffers actually freed): a batch can find every
> +	 * buffer in its transactions busy and free none, while later
> +	 * transactions may still hold buffers whose writeback has
> +	 * completed. sc->nr_scanned only comes back 0 when the checkpoint
> +	 * list had nothing left to walk.
> +	 */
> +	if (sc->nr_scanned == 0)

[Severity: High]
Does this condition cause do_shrink_slab() to busy-loop when the checkpoint
list is fully exhausted but contains a few busy items?

If the shrinker budget is large (for example, due to a stale percpu counter)
but the checkpoint list has only a small number of busy items (e.g., 5),
jbd2_journal_shrink_checkpoint_list() will scan them and set sc->nr_scanned
to 5. Upon reaching the end of the list, it resets
journal->j_shrink_transaction to NULL.

Because sc->nr_scanned is 5, this fails to return SHRINK_STOP.

Will do_shrink_slab() then subtract 5 from its budget and loop immediately,
restarting from the beginning of the list and rescanning those exact same 5
busy items? If this repeats until the large budget is completely drained in
increments of 5, could it cause CPU busy-looping and j_list_lock contention
in the memory reclaim path?

> +		return SHRINK_STOP;
> +
>  	return nr_shrunk;
>  }
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001-fix-jbd2-shrink-scan-nr-scanned-v5-1-5348b51efce2@gmail.com?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-01  7:53 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-01  7:43 [PATCH v5] jbd2: fix shrinker scan budget accounting in jbd2_journal_shrink_scan() Qiliang Yuan
2026-10-01  7:53 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox