Linux EXT4 FS development
 help / color / mirror / Atom feed
From: Mike Small <smallm@sdf.org>
To: "Darrick J. Wong" <djwong@kernel.org>
Cc: Dan Streetman <ddstreet@ieee.org>,
	Nandakumar Raghavan <naraghavan@linux.microsoft.com>,
	linux-ext4@vger.kernel.org, tytso@mit.edu, adilger@dilger.ca,
	srivatsa@csail.mit.edu
Subject: Re: [PATCH v3] e2fsck: take flock(LOCK_EX) on whole-disk device during filesystem check
Date: Mon, 05 Oct 2026 15:35:11 +0000	[thread overview]
Message-ID: <chxv77gma8g.fsf@sdf.org> (raw)
In-Reply-To: <20260924003642.GE6239@frogsfrogsfrogs> (Darrick J. Wong's message of "Wed, 23 Sep 2026 17:36:42 -0700")

"Darrick J. Wong" <djwong@kernel.org> writes:

> On Wed, Sep 23, 2026 at 03:55:23PM -0400, Dan Streetman wrote:
>> 
>> 
>> On Sun, 20 Sep 2026, Nandakumar Raghavan wrote:
>> 
>> > On Thu, Sep 10, 2026 at 05:04:41AM -0700, Nandakumar Raghavan wrote:
>> > > During journal replay, e2fsck writes the primary superblock back to disk
>> > > in multiple I/O operations. The payload lands before the checksum, leaving
>> > > a transient window where the on-disk superblock has a bad checksum.
>> > > 
>> > > If udevd processes a change uevent during this window, libblkid probes the
>> > > primary superblock, finds a checksum mismatch, and concludes the partition
>> > > has no recognisable filesystem. udev then fires a remove event, wiping all
>> > > symlinks in /dev/disk/by-label/ and /dev/disk/by-uuid/. Any mount unit
>> > > that depends on those symlinks will fail.
>> > > 
>> > > udevd already serialises its own partition probes against whole-disk device
>> > > access using flock(LOCK_SH|LOCK_NB); if EAGAIN is returned it requeues the
>> > > event. Take advantage of this protocol by acquiring flock(LOCK_EX) on the
>> > > whole-disk device before opening the filesystem. This forces udevd to defer
>> > > all probes on that disk until e2fsck exits and the lock is released, by
>> > > which point the filesystem is fully consistent.
...
>> > Gentle ping on this patch.
>> > 
>> > I would appreciate any feedback.
>> > 
>> 
>> Can you clarify why this should go into only fsck.ext4? Doesn't this
>> problem exist for other filesystems too?
>> 
>> I sent an earlier email as well with links to:
>> 
>> 1) fsck used to lock the device, but it surfaced a bug in udevd
>> https://bugs.freedesktop.org/show_bug.cgi?id=79576
>> 
>> 2) because of the bug, fsck stopped locking the device
>> https://github.com/util-linux/util-linux/commit/3bbdae633f4a1dda5f95ee6c61f18a1c8ef12250
>> 
>> 3) the systemd-udevd bug was fixed
>> https://github.com/systemd/systemd/commit/5d354e525a5
>> 
>> To me, it makes more sense for the locking that already exists in fsck
>> to get updated (or reverted) to lock the entire device, using the
>> existing -l param (or maybe a new param like --lock-device,
>> --udevd-lock, etc., if util-linux maintainers don't want to change -l
>> behavior).
>> 
>> Do you see an issue with doing the locking there instead of here in
>> fsck.ext4?
>
> /sbin/fsck (aka the dispatch wrapper program) doesn't necessarily know
> which block device(s) are going to be opened by a the fsck.$FSTYP
> program that it creates.  It might be able to infer that by opening any
> parameter and performing the udev locking protocol after checking if
> what it opened is a block device, but that wouldn't work for (say) a
> fsck.XXX program for a multi-device filesystem wherein you only need to
> specify one device and it will find the others.
>
> That said, this patchset also doesn't handle multi-device ext4
> filesystems (i.e. external jbd2 journal device) because the author
> doesn't want to do that.  In their defense, the udev flock()ing protocol
> requires one to determine if an opened block device is a partition; if
> it is, then it requires opening and locking the parent bdev (e.g. sdf1
> -> sdf) instead of locking the original device.  This makes it way more
> complicated for multi-device filesystems because now the client has to
> detect multiple partitions coming from the same underlying device and
> handle that appropriately.  I don't know why the protocol designers made
> that choice.
>
> I can run "trace-cmd record -e 'flock*'" to observe the locking
> interactions with scsi disk partitions, but for whatever reason I don't
> see any flocking going on if I use kpartx to create the partitions with
> device-mapper.  No idea why that is.
>
> --D

Systemd-udevd will not take its shared lock when the device in the
uevent starts with "dm-", "md", or "drbd". See udev_get_whole_disk() in
src/udev/udev-worker.c and how that's used by worker_lock_whole_disk().
Maybe that explains you not seeing the flocks in the second case. Would
the partition device names (or what their symlinks expand to?) look like
/^dm-*/?

Regards,
Mike Small

  parent reply	other threads:[~2026-10-05 15:38 UTC|newest]

Thread overview: 23+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-20 14:32 [PATCH] e2fsck: take flock(LOCK_EX) on whole-disk device during filesystem check Nandakumar Raghavan
2026-08-15  5:29 ` Nandakumar Raghavan
2026-08-18 22:44   ` Andreas Dilger
2026-08-24 13:24     ` Nandakumar Raghavan
2026-08-24 16:15       ` [PATCH v2] " Nandakumar Raghavan
2026-08-24 22:20         ` Andreas Dilger
2026-08-24 22:40         ` Darrick J. Wong
2026-08-26 14:04           ` Nandakumar Raghavan
2026-08-27  7:33             ` Andreas Dilger
2026-08-27 12:20               ` Nandakumar Raghavan
2026-08-27 21:01             ` Theodore Tso
2026-09-02 12:17               ` Nandakumar Raghavan
2026-09-10 12:04               ` [PATCH v3] " Nandakumar Raghavan
2026-09-21  4:38                 ` Nandakumar Raghavan
2026-09-23 19:55                   ` Dan Streetman
2026-09-24  0:36                     ` Darrick J. Wong
2026-09-24 20:19                       ` Dan Streetman
2026-09-24 21:03                         ` Darrick J. Wong
2026-09-29 13:27                           ` Nandakumar Raghavan
2026-10-05 15:35                       ` Mike Small [this message]
2026-10-05 21:53                         ` Darrick J. Wong
2026-09-29 12:37                     ` Nandakumar Raghavan
2026-08-28 21:30         ` [PATCH v2] " ddstreet

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=chxv77gma8g.fsf@sdf.org \
    --to=smallm@sdf.org \
    --cc=adilger@dilger.ca \
    --cc=ddstreet@ieee.org \
    --cc=djwong@kernel.org \
    --cc=linux-ext4@vger.kernel.org \
    --cc=naraghavan@linux.microsoft.com \
    --cc=srivatsa@csail.mit.edu \
    --cc=tytso@mit.edu \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox