From: John Garry <john.g.garry@oracle.com>
To: John Garry <john.garry@linux.dev>, "Darrick J. Wong" <djwong@kernel.org>
Cc: Shin'ichiro Kawasaki <shinichiro.kawasaki@wdc.com>,
"linux-xfs@vger.kernel.org" <linux-xfs@vger.kernel.org>
Subject: Re: [bug report] fstests generic/774 hang again
Date: Wed, 16 Sep 2026 11:23:16 +0100 [thread overview]
Message-ID: <9bf23063-e6e5-4595-8ff9-27383c99e5a6@oracle.com> (raw)
In-Reply-To: <db4ab609-af16-4c5d-aa98-c0655625a2df@linux.dev>
On 15/09/2026 16:41, John Garry wrote:
> On 9/15/26 15:50, Darrick J. Wong wrote:
>>> thanks
>>>
>>> About the kernel code, I am wondering if using IOMAP_DIO_FORCE_WAIT for
>>> CoW-based atomics could help avoid this issue as we seem to be bogged
>>> down
>>> in lock contention.
>> Which lock is being contended, anyway? It looks like the ILOCK?
>
> Yeah, I think so Re. ilock. I had some perf data illustrating this from
> last year, which I can't seem to find, so I will re-generate it.
Here is perf call graph snippet when running fio with 72x threads
issuing 4KB atomic writes on 1MB file:
--50.46%--xfs_file_dio_write_atomic
|
|--37.19%--iomap_dio_rw
| |
| --37.14%--__iomap_dio_rw
| |
| --36.50%--iomap_iter
| |
| --36.49%--xfs_atomic_write_cow_iomap_next
| |
| |--27.34%--xfs_ilock
| | |
| | --27.34%--down_write
| | |
| | --27.28%--rwsem_down_write_slowpath
| | |
| | |--24.65%--osq_lock
| | |
| | --2.35%--rwsem_spin_on_owner
| |
| |--7.62%--xfs_trans_alloc_inode
| | |
| | --7.59%--xfs_ilock
Notice how much time we spend getting that ilock.
We could try something like this:
----8<------
[PATCH] xfs: serialize CoW-based atomic writes
We have had reports of system hangs when testing Cow-based atomic writes
for large block sizes.
Analysis has shown large contention on the the per-inode ilock,
superficially in creating the mapping in
xfs_atomic_write_cow_iomap_begin().
Serialize atomic writes to reduce this contention. In some cases this
may reduce performance, but high performance has not been a requirement
so far for CoW-based atomic writes.
Signed-off-by: John Garry <john.garry@linux.dev>
diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
index d8202da15aca..ff48757cf0f0 100644
--- a/fs/xfs/xfs_file.c
+++ b/fs/xfs/xfs_file.c
@@ -819,10 +819,12 @@ xfs_file_dio_write_atomic(
* HW offload should be faster, so try that first if it is already
* known that the write length is not too large.
*/
- if (ocount > xfs_inode_buftarg(ip)->bt_awu_max)
+ if (ocount > xfs_inode_buftarg(ip)->bt_awu_max) {
dops = &xfs_atomic_write_cow_iomap_ops;
- else
+ dio_flags |= IOMAP_DIO_FORCE_WAIT;
+ } else {
dops = &xfs_direct_write_iomap_ops;
+ }
retry:
ret = xfs_ilock_iocb_for_write(iocb, &iolock);
@@ -840,6 +842,8 @@ xfs_file_dio_write_atomic(
}
trace_xfs_file_direct_write(iocb, from);
+ if (dio_flags & IOMAP_DIO_FORCE_WAIT)
+ inode_dio_wait(VFS_I(ip));
if (mapping_stable_writes(iocb->ki_filp->f_mapping))
dio_flags |= IOMAP_DIO_BOUNCE;
ret = iomap_dio_rw(iocb, from, dops, &xfs_dio_write_ops, dio_flags,
@@ -853,6 +857,7 @@ xfs_file_dio_write_atomic(
*/
if (ret == -ENOPROTOOPT && dops == &xfs_direct_write_iomap_ops) {
xfs_iunlock(ip, iolock);
+ dio_flags |= IOMAP_DIO_FORCE_WAIT;
dops = &xfs_atomic_write_cow_iomap_ops;
goto retry;
}
---->8------
This is far from ideal, but I think that it should stops the hangs -
having the kernel hang is intolerable.
I also notice that in xfs_atomic_write_cow_iomap_next() we drop the
ilock for calling xfs_trans_alloc_inode() (which immediately grabs the
same lock). It would seem that could be improved.
next prev parent reply other threads:[~2026-09-16 10:23 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 9:34 [bug report] fstests generic/774 hang again Shin'ichiro Kawasaki
2026-09-15 10:00 ` John Garry
2026-09-15 11:48 ` Shin'ichiro Kawasaki
2026-09-15 14:48 ` John Garry
2026-09-15 14:50 ` Darrick J. Wong
2026-09-15 15:41 ` John Garry
2026-09-16 10:23 ` John Garry [this message]
2026-09-16 22:23 ` Dave Chinner
2026-09-17 8:28 ` John Garry
2026-09-17 21:20 ` Dave Chinner
2026-09-17 6:30 ` Shin'ichiro Kawasaki
2026-09-16 2:53 ` Shin'ichiro Kawasaki
2026-09-16 21:45 ` Dave Chinner
2026-09-17 6:52 ` Shin'ichiro Kawasaki
2026-09-17 10:45 ` Shin'ichiro Kawasaki
2026-09-17 21:24 ` Dave Chinner
2026-09-19 11:53 ` Shin'ichiro Kawasaki
2026-09-21 22:05 ` Dave Chinner
2026-09-25 1:32 ` Shin'ichiro Kawasaki
2026-09-27 21:30 ` Dave Chinner
2026-09-28 2:45 ` Darrick J. Wong
2026-09-22 0:28 ` Darrick J. Wong
2026-09-25 1:38 ` Shin'ichiro Kawasaki
2026-09-25 23:03 ` Darrick J. Wong
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9bf23063-e6e5-4595-8ff9-27383c99e5a6@oracle.com \
--to=john.g.garry@oracle.com \
--cc=djwong@kernel.org \
--cc=john.garry@linux.dev \
--cc=linux-xfs@vger.kernel.org \
--cc=shinichiro.kawasaki@wdc.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox