Linux block layer
 help / color / mirror / Atom feed
* Re: [syzbot] [block?] INFO: task hung in read_cache_folio (6) (extid 9db0864859224b833108)
@ 2026-09-28  7:46 clkernel
  0 siblings, 0 replies; 2+ messages in thread
From: clkernel @ 2026-09-28  7:46 UTC (permalink / raw)
  To: axboe, linux-block, linux-kernel, syzkaller-bugs, linux-mm

On Fri, 03 Jul 2026, syzbot wrote:
> syzbot found the following issue on:
> ...
> folio_wait_bit_common+0x6b8/0xa30 mm/filemap.c:1324
> folio_put_wait_locked mm/filemap.c:1493 [inline]
> do_read_cache_folio+0x23c/0x5a8 mm/filemap.c:4089
> read_cache_folio+0x68/0x88 mm/filemap.c:4139
> read_part_sector+0xcc/0x708 block/partitions/core.c:724

A second observation of the same wait site: on x86-64 servers we see 
reader threads (a database reader in a busy server workload) parked 
D-state in folio_wait_bit_common's io_schedule, entered through the same 
folio_put_wait_locked -> do_read_cache_folio -> read_cache_folio path 
this report shows. Observed on 7.0.0-30, 7.0.0-31 and 7.2.5, reproduced N=3.

Three measurements, per repro:

1. the pinned folio's lock is FREE while the waiter sleeps: a fresh read 
of the exact pinned pages answers in 4-108 ms while the thread stays D 
in the same wait;
2. each later lock/unlock of that folio wakes the waiter for one ~4KB 
read: the blocked read's offsets crawl one page per unlock cycle (~one 
page per 10 s under churn), so the syscall can sit in the same wait for 
minutes to days;
3. the discriminator: read /proc/<tid>/syscall while the waiter sits 
inside the wait, then dd the same page range fresh. The dd returns 
immediately; the waiter does not move.

Root cause (established from source, then observed directly on a local 
build; measured): folio_wake_bit() wakes the hashed folio waitqueue 
through __wake_up_locked_key(), i.e. __wake_up_common(..., 
nr_exclusive=1). The walk in kernel/sched/wait.c stops at the first 
WQ_FLAG_EXCLUSIVE entry it wakes:

if (ret && (flags & WQ_FLAG_EXCLUSIVE) && !--nr_exclusive)
break;

folio_wait_bit_common() queued every waiter with 
__add_wait_queue_entry_tail(), so a non-exclusive waiter could sit 
behind an exclusive folio_lock() waiter on the same folio. When the fill 
completed, folio_end_read()'s wake woke the exclusive waiter and 
stopped. The shared/DROP waiter behind it was never visited: 
WQ_FLAG_WOKEN never set, although the wake itself ran. PG_locked was 
already clear and the folio uptodate (so a fresh read of that page 
answers through the uptodate fast path, measurement 1), and PG_waiters 
stayed set, so the sleeper only ran when some later lock/unlock of the 
same folio address reached the queue (measurement 2: one page per wake, 
crawling until traffic arrived).

Same mechanism explains this report's hang: the waiters are the DROP 
ones (folio_put_wait_locked), which drop their folio reference before 
sleeping and re-lookup only after a wake that arrives late or never.

Observation on a debug build, same queue after folio_wake_bit() returned 
(a read-only walk counting matching non-exclusive entries left 
unvisited): 363 events per 10-minute run of readers racing 
punch/fault/truncate churn under reclaim pressure on an unmodified 
queue; 1 per run after queueing non-exclusive waiters at the head side 
only; 0 after the same change at softleaf_entry_wait_on_locked (the 
other non-exclusive enqueue on this queue).

The fix (mm/filemap.c, attached): queue non-exclusive waiters with 
__add_wait_queue() (wait.h's own default head-side insert) and keep 
exclusive folio_lock waiters on __add_wait_queue_entry_tail(). With that 
order every non-exclusive full-match waiter is visited before any walk 
break can fire, which restores the "wake all shared waiters, then one 
exclusive" contract of nr_exclusive=1; exclusive waiters keep their FIFO 
tail and their one-at-a-time wake. Same change at both non-exclusive 
enqueue sites (folio_wait_bit_common and softleaf_entry_wait_on_locked).

One known sibling waits on a follow-up: __folio_lock_async's entry 
(io_uring buffered reads with ki_waitq) enqueues non-exclusive at the 
tail. Same class, unreachable from the workloads above.

Reproducer shape: threads pread a file while others block in folio_lock 
on the same pages (hole punch, truncate, or mmap faults) under reclaim 
churn. On an unmodified queue the waitqueue walk above counts the 
skipped waiters within seconds. The harness, traces or configs are yours 
on request.

The same change cherry-picks clean onto v7.2.6 (mm/filemap.o builds 
there). Happy to send a [PATCH] for stable@ if the fix is wanted on the 
7.2.x queue.

---
mm/filemap.c | 16 +++++++++++++---
1 file changed, 13 insertions(+), 3 deletions(-)

diff --git a/mm/filemap.c b/mm/filemap.c
index 00fd89cf6f55..383701407843 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -1294,8 +1294,18 @@ static inline int folio_wait_bit_common(struct 
folio *folio, int bit_nr,
*/
spin_lock_irq(&q->lock);
folio_set_waiters(folio);
- if (!folio_trylock_flag(folio, bit_nr, wait))
- __add_wait_queue_entry_tail(q, wait);
+ if (!folio_trylock_flag(folio, bit_nr, wait)) {
+ /*
+ * A non-exclusive (shared/DROP) waiter must queue ahead of
+ * exclusive folio_lock waiters: the nr_exclusive=1 wake walk
+ * stops at the first exclusive match, so a non-exclusive
+ * waiter behind one is never visited and never woken.
+ */
+ if (behavior == EXCLUSIVE)
+ __add_wait_queue_entry_tail(q, wait);
+ else
+ __add_wait_queue(q, wait);
+ }
spin_unlock_irq(&q->lock);
/*
@@ -1429,7 +1439,7 @@ void softleaf_entry_wait_on_locked(softleaf_t 
entry, spinlock_t *ptl)
spin_lock_irq(&q->lock);
folio_set_waiters(folio);
if (!folio_trylock_flag(folio, PG_locked, wait))
- __add_wait_queue_entry_tail(q, wait);
+ __add_wait_queue(q, wait);
spin_unlock_irq(&q->lock);
/*


^ permalink raw reply related	[flat|nested] 2+ messages in thread
* Re: [syzbot] [block?] INFO: task hung in read_cache_folio (6) (extid 9db0864859224b833108)
@ 2026-09-28  7:45 clkernel
  0 siblings, 0 replies; 2+ messages in thread
From: clkernel @ 2026-09-28  7:45 UTC (permalink / raw)
  To: axboe, linux-block, linux-kernel, syzkaller-bugs, linux-mm

On Fri, 03 Jul 2026, syzbot wrote:
> syzbot found the following issue on:
> ...
> folio_wait_bit_common+0x6b8/0xa30 mm/filemap.c:1324
> folio_put_wait_locked mm/filemap.c:1493 [inline]
> do_read_cache_folio+0x23c/0x5a8 mm/filemap.c:4089
> read_cache_folio+0x68/0x88 mm/filemap.c:4139
> read_part_sector+0xcc/0x708 block/partitions/core.c:724

A second observation of the same wait site: on x86-64 servers we see 
reader threads (a database reader in a busy server workload) parked 
D-state in folio_wait_bit_common's io_schedule, entered through the same 
folio_put_wait_locked -> do_read_cache_folio -> read_cache_folio path 
this report shows. Observed on 7.0.0-30, 7.0.0-31 and 7.2.5, reproduced N=3.

Three measurements, per repro:

1. the pinned folio's lock is FREE while the waiter sleeps: a fresh read 
of the exact pinned pages answers in 4-108 ms while the thread stays D 
in the same wait;
2. each later lock/unlock of that folio wakes the waiter for one ~4KB 
read: the blocked read's offsets crawl one page per unlock cycle (~one 
page per 10 s under churn), so the syscall can sit in the same wait for 
minutes to days;
3. the discriminator: read /proc/<tid>/syscall while the waiter sits 
inside the wait, then dd the same page range fresh. The dd returns 
immediately; the waiter does not move.

Root cause (established from source, then observed directly on a local 
build; measured): folio_wake_bit() wakes the hashed folio waitqueue 
through __wake_up_locked_key(), i.e. __wake_up_common(..., 
nr_exclusive=1). The walk in kernel/sched/wait.c stops at the first 
WQ_FLAG_EXCLUSIVE entry it wakes:

if (ret && (flags & WQ_FLAG_EXCLUSIVE) && !--nr_exclusive)
break;

folio_wait_bit_common() queued every waiter with 
__add_wait_queue_entry_tail(), so a non-exclusive waiter could sit 
behind an exclusive folio_lock() waiter on the same folio. When the fill 
completed, folio_end_read()'s wake woke the exclusive waiter and 
stopped. The shared/DROP waiter behind it was never visited: 
WQ_FLAG_WOKEN never set, although the wake itself ran. PG_locked was 
already clear and the folio uptodate (so a fresh read of that page 
answers through the uptodate fast path, measurement 1), and PG_waiters 
stayed set, so the sleeper only ran when some later lock/unlock of the 
same folio address reached the queue (measurement 2: one page per wake, 
crawling until traffic arrived).

Same mechanism explains this report's hang: the waiters are the DROP 
ones (folio_put_wait_locked), which drop their folio reference before 
sleeping and re-lookup only after a wake that arrives late or never.

Observation on a debug build, same queue after folio_wake_bit() returned 
(a read-only walk counting matching non-exclusive entries left 
unvisited): 363 events per 10-minute run of readers racing 
punch/fault/truncate churn under reclaim pressure on an unmodified 
queue; 1 per run after queueing non-exclusive waiters at the head side 
only; 0 after the same change at softleaf_entry_wait_on_locked (the 
other non-exclusive enqueue on this queue).

The fix (mm/filemap.c, attached): queue non-exclusive waiters with 
__add_wait_queue() (wait.h's own default head-side insert) and keep 
exclusive folio_lock waiters on __add_wait_queue_entry_tail(). With that 
order every non-exclusive full-match waiter is visited before any walk 
break can fire, which restores the "wake all shared waiters, then one 
exclusive" contract of nr_exclusive=1; exclusive waiters keep their FIFO 
tail and their one-at-a-time wake. Same change at both non-exclusive 
enqueue sites (folio_wait_bit_common and softleaf_entry_wait_on_locked).

One known sibling waits on a follow-up: __folio_lock_async's entry 
(io_uring buffered reads with ki_waitq) enqueues non-exclusive at the 
tail. Same class, unreachable from the workloads above.

Reproducer shape: threads pread a file while others block in folio_lock 
on the same pages (hole punch, truncate, or mmap faults) under reclaim 
churn. On an unmodified queue the waitqueue walk above counts the 
skipped waiters within seconds. The harness, traces or configs are yours 
on request.

The same change cherry-picks clean onto v7.2.6 (mm/filemap.o builds 
there). Happy to send a [PATCH] for stable@ if the fix is wanted on the 
7.2.x queue.

---
mm/filemap.c | 16 +++++++++++++---
1 file changed, 13 insertions(+), 3 deletions(-)

diff --git a/mm/filemap.c b/mm/filemap.c
index 00fd89cf6f55..383701407843 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -1294,8 +1294,18 @@ static inline int folio_wait_bit_common(struct 
folio *folio, int bit_nr,
*/
spin_lock_irq(&q->lock);
folio_set_waiters(folio);
- if (!folio_trylock_flag(folio, bit_nr, wait))
- __add_wait_queue_entry_tail(q, wait);
+ if (!folio_trylock_flag(folio, bit_nr, wait)) {
+ /*
+ * A non-exclusive (shared/DROP) waiter must queue ahead of
+ * exclusive folio_lock waiters: the nr_exclusive=1 wake walk
+ * stops at the first exclusive match, so a non-exclusive
+ * waiter behind one is never visited and never woken.
+ */
+ if (behavior == EXCLUSIVE)
+ __add_wait_queue_entry_tail(q, wait);
+ else
+ __add_wait_queue(q, wait);
+ }
spin_unlock_irq(&q->lock);
/*
@@ -1429,7 +1439,7 @@ void softleaf_entry_wait_on_locked(softleaf_t 
entry, spinlock_t *ptl)
spin_lock_irq(&q->lock);
folio_set_waiters(folio);
if (!folio_trylock_flag(folio, PG_locked, wait))
- __add_wait_queue_entry_tail(q, wait);
+ __add_wait_queue(q, wait);
spin_unlock_irq(&q->lock);
/*


^ permalink raw reply related	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-28  7:46 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28  7:46 [syzbot] [block?] INFO: task hung in read_cache_folio (6) (extid 9db0864859224b833108) clkernel
  -- strict thread matches above, loose matches on Subject: below --
2026-09-28  7:45 clkernel

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox