Linux EXT4 FS development
 help / color / mirror / Atom feed
* [PATCH e2fsprogs 0/2] e2fsck: fix a self-deadlock that hangs fsck on corrupt images
@ 2026-09-22 11:04 Matthias Goergens
  2026-09-22 11:04 ` [PATCH e2fsprogs 1/2] libext2fs: fix self-deadlock in flush_cached_blocks() write-error retry Matthias Goergens
  2026-09-22 11:04 ` [PATCH e2fsprogs 2/2] tests: add f_cache_mtx_deadlock for the flush_cached_blocks() retry lock Matthias Goergens
  0 siblings, 2 replies; 4+ messages in thread
From: Matthias Goergens @ 2026-09-22 11:04 UTC (permalink / raw)
  To: Theodore Ts'o; +Cc: linux-ext4

e2fsck can hang forever on a corrupt image.  The hang is a self-deadlock
in libext2fs: flush_cached_blocks() releases CACHE_MTX around the
write_error callback, re-acquires it, and then jumps to a label above
the loop whose first statement acquires it again.  The mutex is not
recursive, so the first flush that reports a write error through a
registered handler blocks the thread that already holds the lock.

Patch 1 removes the redundant acquisition.  Patch 2 adds a regression
test.

This matters beyond the fuzzer that found it: the deadlocked process
sits at 0% CPU and does not respond to SIGTERM, so an init script
waiting on fsck waits forever, and read-only checking (-fn) is enough
to reach it.

Patch 2 departs from the usual f_* shape twice, both times because the
bug is a hang rather than a wrong answer: the e2fsck run is wrapped in
timeout(1), or a failure would stop the suite indefinitely rather than
fail, and the transcript is not compared, because it is a thousand
lines of repeated write errors that say nothing about this bug.  Happy
to drop the test or shape it differently if you would rather not have
either of those in the f_* tests.

Matthias Goergens (2):
  libext2fs: fix self-deadlock in flush_cached_blocks() write-error
    retry
  tests: add f_cache_mtx_deadlock for the flush_cached_blocks() retry
    lock

 lib/ext2fs/unix_io.c                |   2 +-
 tests/f_cache_mtx_deadlock/expect   |   2 ++
 tests/f_cache_mtx_deadlock/image.gz | Bin 0 -> 695 bytes
 tests/f_cache_mtx_deadlock/name     |   1 +
 tests/f_cache_mtx_deadlock/script   |  44 ++++++++++++++++++++++++++++
 5 files changed, 48 insertions(+), 1 deletion(-)
 create mode 100644 tests/f_cache_mtx_deadlock/expect
 create mode 100644 tests/f_cache_mtx_deadlock/image.gz
 create mode 100644 tests/f_cache_mtx_deadlock/name
 create mode 100644 tests/f_cache_mtx_deadlock/script


base-commit: 8fd79523d051d5ea881237f27eb1c664486e0078
-- 
2.55.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-22 18:18 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-22 11:04 [PATCH e2fsprogs 0/2] e2fsck: fix a self-deadlock that hangs fsck on corrupt images Matthias Goergens
2026-09-22 11:04 ` [PATCH e2fsprogs 1/2] libext2fs: fix self-deadlock in flush_cached_blocks() write-error retry Matthias Goergens
2026-09-22 18:18   ` Darrick J. Wong
2026-09-22 11:04 ` [PATCH e2fsprogs 2/2] tests: add f_cache_mtx_deadlock for the flush_cached_blocks() retry lock Matthias Goergens

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox