From: Dave Chinner <dgc@kernel.org>
To: John Garry <john.g.garry@oracle.com>
Cc: John Garry <john.garry@linux.dev>,
"Darrick J. Wong" <djwong@kernel.org>,
Shin'ichiro Kawasaki <shinichiro.kawasaki@wdc.com>,
"linux-xfs@vger.kernel.org" <linux-xfs@vger.kernel.org>
Subject: Re: [bug report] fstests generic/774 hang again
Date: Thu, 17 Sep 2026 08:23:58 +1000 [thread overview]
Message-ID: <aqsW_iNwYAjYldVf@dread> (raw)
In-Reply-To: <9bf23063-e6e5-4595-8ff9-27383c99e5a6@oracle.com>
On Wed, Sep 16, 2026 at 11:23:16AM +0100, John Garry wrote:
> On 15/09/2026 16:41, John Garry wrote:
> > On 9/15/26 15:50, Darrick J. Wong wrote:
> > > > thanks
> > > >
> > > > About the kernel code, I am wondering if using IOMAP_DIO_FORCE_WAIT for
> > > > CoW-based atomics could help avoid this issue as we seem to be
> > > > bogged down
> > > > in lock contention.
> > > Which lock is being contended, anyway? It looks like the ILOCK?
> >
> > Yeah, I think so Re. ilock. I had some perf data illustrating this from
> > last year, which I can't seem to find, so I will re-generate it.
>
> Here is perf call graph snippet when running fio with 72x threads issuing
> 4KB atomic writes on 1MB file:
>
> --50.46%--xfs_file_dio_write_atomic
> |
> |--37.19%--iomap_dio_rw
> | |
> | --37.14%--__iomap_dio_rw
> | |
> | --36.50%--iomap_iter
> | |
> | --36.49%--xfs_atomic_write_cow_iomap_next
> | |
> | |--27.34%--xfs_ilock
> | | |
> | | --27.34%--down_write
> | | |
> | | --27.28%--rwsem_down_write_slowpath
> | | |
> | | |--24.65%--osq_lock
> | | |
> | | --2.35%--rwsem_spin_on_owner
> | |
> | |--7.62%--xfs_trans_alloc_inode
> | | |
> | | --7.59%--xfs_ilock
Where's the other half of the lock contention?
If this is all there is, it's because of concurrent submission of
serialising IO, and xfs_atomic_write_cow_iomap_begin() is the first
place in the IO submission path that we hit a serialising lock. That
indicates a problem with high level IO serialisation, not a problem
with the ILOCK.
> Notice how much time we spend getting that ilock.
>
> We could try something like this:
>
> ----8<------
>
> [PATCH] xfs: serialize CoW-based atomic writes
>
> We have had reports of system hangs when testing Cow-based atomic writes
> for large block sizes.
>
> Analysis has shown large contention on the the per-inode ilock,
> superficially in creating the mapping in
> xfs_atomic_write_cow_iomap_begin().
>
> Serialize atomic writes to reduce this contention. In some cases this may
> reduce performance, but high performance has not been a requirement so far
> for CoW-based atomic writes.
>
> Signed-off-by: John Garry <john.garry@linux.dev>
>
> diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
> index d8202da15aca..ff48757cf0f0 100644
> --- a/fs/xfs/xfs_file.c
> +++ b/fs/xfs/xfs_file.c
> @@ -819,10 +819,12 @@ xfs_file_dio_write_atomic(
> * HW offload should be faster, so try that first if it is already
> * known that the write length is not too large.
> */
> - if (ocount > xfs_inode_buftarg(ip)->bt_awu_max)
> + if (ocount > xfs_inode_buftarg(ip)->bt_awu_max) {
> dops = &xfs_atomic_write_cow_iomap_ops;
> - else
> + dio_flags |= IOMAP_DIO_FORCE_WAIT;
> + } else {
> dops = &xfs_direct_write_iomap_ops;
> + }
>
> retry:
> ret = xfs_ilock_iocb_for_write(iocb, &iolock);
> @@ -840,6 +842,8 @@ xfs_file_dio_write_atomic(
> }
>
> trace_xfs_file_direct_write(iocb, from);
> + if (dio_flags & IOMAP_DIO_FORCE_WAIT)
> + inode_dio_wait(VFS_I(ip));
Based on my comment above, this seems like the wrong approach to me.
We serialise concurrent DIO write submission when necessary by
taking the IOLOCK_EXCL instead of IOLOCK_SHARED (e.g. EOF
extension). inode_dio_wait() waits for completion of all DIO (read
and write), but we don't need to wait for completions if all we need
to do is serialise IO submission.
Hence if we know we are going to do a software COW for the write, we
should take IOLOCK_EXCL to serialise submission up high. Waiting
until we are deep into the IO path and serialising on the
internal metadata ILOCK is much, much harder to do efficiently and
correctly.
i.e. we need to take the ILOCK vs transaction reservation vs unbound
user-controlled concurrency deadlock problem out of
the equation. Using the IOLOCK achieves this because IO submission
blocks on the IOLOCK rather than on the ILOCK whilst holding a
transaction reservation.
> I also notice that in xfs_atomic_write_cow_iomap_next() we drop the ilock
> for calling xfs_trans_alloc_inode() (which immediately grabs the same lock).
> It would seem that could be improved.
Doubt it - holding the ILOCK across transaction reservation
guarantees journal space deadlocks will occur. That's why the
suboptimal pattern of 'lock - check - unlock - trans_alloc - lock -
check' exists throughout XFS inode modification paths.
FWIW, that's one of the anit-patterns that the patchset I posted
recently begins to address:
https://lore.kernel.org/linux-xfs/20260819001442.1451892-1-dgc@kernel.org/
It creates consistently correct patterns in the IO path by changes all the
individual alloc-lock-check-mod-commit-unlock loops to an "atomic"
alloc-lock, check-mod-roll loop, commit-unlock pattern.
IOWs, multi-part extent modifications are all done under a single
lock/unlock pair, and so appear atomic to all other extent map and
modification operations. This makes many of the nasty IO corner
cases that lead to sub-optimal behaviour (like data corruption or
lock contention in COW workloads due to continual bouncing of the
ILOCK between submission and completion contexts) largely go away...
-Dave.
--
Dave Chinner
dgc@kernel.org
next prev parent reply other threads:[~2026-09-16 22:24 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 9:34 [bug report] fstests generic/774 hang again Shin'ichiro Kawasaki
2026-09-15 10:00 ` John Garry
2026-09-15 11:48 ` Shin'ichiro Kawasaki
2026-09-15 14:48 ` John Garry
2026-09-15 14:50 ` Darrick J. Wong
2026-09-15 15:41 ` John Garry
2026-09-16 10:23 ` John Garry
2026-09-16 22:23 ` Dave Chinner [this message]
2026-09-17 8:28 ` John Garry
2026-09-17 21:20 ` Dave Chinner
2026-09-17 6:30 ` Shin'ichiro Kawasaki
2026-09-16 2:53 ` Shin'ichiro Kawasaki
2026-09-16 21:45 ` Dave Chinner
2026-09-17 6:52 ` Shin'ichiro Kawasaki
2026-09-17 10:45 ` Shin'ichiro Kawasaki
2026-09-17 21:24 ` Dave Chinner
2026-09-19 11:53 ` Shin'ichiro Kawasaki
2026-09-21 22:05 ` Dave Chinner
2026-09-25 1:32 ` Shin'ichiro Kawasaki
2026-09-27 21:30 ` Dave Chinner
2026-09-28 2:45 ` Darrick J. Wong
2026-09-22 0:28 ` Darrick J. Wong
2026-09-25 1:38 ` Shin'ichiro Kawasaki
2026-09-25 23:03 ` Darrick J. Wong
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqsW_iNwYAjYldVf@dread \
--to=dgc@kernel.org \
--cc=djwong@kernel.org \
--cc=john.g.garry@oracle.com \
--cc=john.garry@linux.dev \
--cc=linux-xfs@vger.kernel.org \
--cc=shinichiro.kawasaki@wdc.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.