Linux EXT4 FS development
 help / color / mirror / Atom feed
From: Matthias Goergens <matthias.goergens@gmail.com>
To: Theodore Ts'o <tytso@mit.edu>
Cc: linux-ext4@vger.kernel.org
Subject: [PATCH e2fsprogs 0/2] e2fsck: fix a self-deadlock that hangs fsck on corrupt images
Date: Tue, 22 Sep 2026 19:04:33 +0800	[thread overview]
Message-ID: <20260922110435.1528332-1-matthias.goergens@gmail.com> (raw)

e2fsck can hang forever on a corrupt image.  The hang is a self-deadlock
in libext2fs: flush_cached_blocks() releases CACHE_MTX around the
write_error callback, re-acquires it, and then jumps to a label above
the loop whose first statement acquires it again.  The mutex is not
recursive, so the first flush that reports a write error through a
registered handler blocks the thread that already holds the lock.

Patch 1 removes the redundant acquisition.  Patch 2 adds a regression
test.

This matters beyond the fuzzer that found it: the deadlocked process
sits at 0% CPU and does not respond to SIGTERM, so an init script
waiting on fsck waits forever, and read-only checking (-fn) is enough
to reach it.

Patch 2 departs from the usual f_* shape twice, both times because the
bug is a hang rather than a wrong answer: the e2fsck run is wrapped in
timeout(1), or a failure would stop the suite indefinitely rather than
fail, and the transcript is not compared, because it is a thousand
lines of repeated write errors that say nothing about this bug.  Happy
to drop the test or shape it differently if you would rather not have
either of those in the f_* tests.

Matthias Goergens (2):
  libext2fs: fix self-deadlock in flush_cached_blocks() write-error
    retry
  tests: add f_cache_mtx_deadlock for the flush_cached_blocks() retry
    lock

 lib/ext2fs/unix_io.c                |   2 +-
 tests/f_cache_mtx_deadlock/expect   |   2 ++
 tests/f_cache_mtx_deadlock/image.gz | Bin 0 -> 695 bytes
 tests/f_cache_mtx_deadlock/name     |   1 +
 tests/f_cache_mtx_deadlock/script   |  44 ++++++++++++++++++++++++++++
 5 files changed, 48 insertions(+), 1 deletion(-)
 create mode 100644 tests/f_cache_mtx_deadlock/expect
 create mode 100644 tests/f_cache_mtx_deadlock/image.gz
 create mode 100644 tests/f_cache_mtx_deadlock/name
 create mode 100644 tests/f_cache_mtx_deadlock/script


base-commit: 8fd79523d051d5ea881237f27eb1c664486e0078
-- 
2.55.0


             reply	other threads:[~2026-09-22 11:04 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 11:04 Matthias Goergens [this message]
2026-09-22 11:04 ` [PATCH e2fsprogs 1/2] libext2fs: fix self-deadlock in flush_cached_blocks() write-error retry Matthias Goergens
2026-09-22 18:18   ` Darrick J. Wong
2026-09-22 11:04 ` [PATCH e2fsprogs 2/2] tests: add f_cache_mtx_deadlock for the flush_cached_blocks() retry lock Matthias Goergens

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922110435.1528332-1-matthias.goergens@gmail.com \
    --to=matthias.goergens@gmail.com \
    --cc=linux-ext4@vger.kernel.org \
    --cc=tytso@mit.edu \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox