All of lore.kernel.org
 help / color / mirror / Atom feed
From: Dave Chinner <dgc@kernel.org>
To: Shin'ichiro Kawasaki <shinichiro.kawasaki@wdc.com>
Cc: "linux-xfs@vger.kernel.org" <linux-xfs@vger.kernel.org>,
	"Darrick J. Wong" <djwong@kernel.org>,
	John Garry <john.garry@linux.dev>
Subject: Re: [bug report] fstests generic/774 hang again
Date: Thu, 17 Sep 2026 07:45:04 +1000	[thread overview]
Message-ID: <aqsN4GOXthkF8rOI@dread> (raw)
In-Reply-To: <aqkP9CTUvlC2YCIb@shinmob>

On Tue, Sep 15, 2026 at 06:34:50PM +0900, Shin'ichiro Kawasaki wrote:
> I observe the fstests test case generic/774 hangs, when I run it for xfs on 8GiB
> TCMU fileio devices. Actually I once reported this hang symptom in last October,
> and Darrick and John kindly took some actions [1]. Since then, the hang
> disappeared and I have not observed the hang almost one year. However, my test
> system started reporting the hang again last week.
> 
> [1] https://lore.kernel.org/linux-xfs/cmk52aqexackyz65phxgme55a3tdrermo3o4skr4lo4pwvvvcp@jmcblnfikbp2/#t
> 
> FYI, here I attache the kernel message [2]. I observed the hang last week with
> the kernel on the xfs/for-next branch at the commit 0ca15a1a1151 ("xfs: remove
> an extra cast in xfs_file_compat_ioctl"). Today, I tried again with the kernel
> at current xfs/for-next branch tip 4266ffdfd7cb ("xfs: remove an extra cast in
> xfs_file_compat_ioctl"), and observed the hang again. The hang can be recreated
> in stable manner by repeating the test case 20 times or so.
> 
> I also tried other test cases in atmoicwrites group, and observed no failure.
> 
>  generic/765        [not run] write atomic not supported by this block device
>  generic/767         11s
>  generic/768         13s
>  generic/769         13s
>  generic/770         32s
>  generic/773        [not run] write atomic not supported by this block device
>  generic/774        (skipped)
>  generic/775         293s
>  generic/776        [notrun] write atomic not supported by this block device
>  generic/778         48s
>  xfs/838            [not run] External volumes not in use, skipped this test
>  xfs/839            [not run] XFS error injection requires CONFIG_XFS_DEBUG
>  xfs/840            [not run] write atomic not supported by this block device
> 
> Action for fix will be appreciated. If I can do anything on my test node, please
> let me know.
.....

A bunch of threads waiting on the ILOCK here:

 down_write_nested+0x1c0/0x1f0
 xfs_reflink_end_atomic_cow+0x2f3/0x560 [xfs]
 xfs_dio_write_end_io+0x4b7/0x650 [xfs]
 iomap_dio_complete+0x140/0xb20
 iomap_dio_complete_work+0x58/0x90
 process_one_work+0x947/0x1760

And the holder:

 schedule+0xe5/0x2e0
 xlog_grant_head_wait+0x175/0xac0 [xfs]
 xlog_grant_head_check+0x312/0x3f0 [xfs]
 xfs_log_regrant+0x380/0x7d0 [xfs]
 xfs_trans_roll+0x2d9/0x420 [xfs]
 xfs_defer_trans_roll+0x11e/0x4b0 [xfs]
 xfs_defer_finish_noroll+0x460/0xe70 [xfs]
 xfs_trans_commit+0xfc/0x180 [xfs]
 xfs_reflink_end_atomic_cow+0x3b2/0x560 [xfs]
 xfs_dio_write_end_io+0x4b7/0x650 [xfs]
 iomap_dio_complete+0x140/0xb20
 iomap_dio_complete_work+0x58/0x90
 process_one_work+0x947/0x1760

is waiting on log space whilst holding the ILOCK.

This looks to me like the test runs out of log space because of all
the IO completions holding transaction reservations waiting on the
ILOCK, whilst the ILOCK holder can't get enough log space to regrant
on transaction roll to continue the transaction.

Oh:

STATIC void
xfs_calc_default_atomic_ioend_reservation(
        struct xfs_mount        *mp,
        struct xfs_trans_resv   *resp)
{
        /* Pick a default that will scale reasonably for the log size. */
        resp->tr_atomic_ioend = resp->tr_itruncate;
}

Which means each reservation is probably holding hundreds of KB to
MB of journal reservation. If the journal is 64MB in size, then a
few dozen o these might be all that is needed to consume all the
reservation space.

Hmmm - there are at least 30 tasks stuck in
xfs_reflink_end_atomic_cow(), another 30 stuck in
xfs_vn_update_time() holding reservations, and significant number
xfs_atomic_write_cow_iomap_begin()->xfs_trans_alloc_inode() holding
tr_write reservations.

IOWs, smells of the journal being run out of reservation space, and
no new space being able to be freed because all the dirty inodes
in the journal that need to be written back to free up space are
locked waiting for journal space to come free....

How big is the journal in the filesystem being tested? If you
increase the size of the journal, does it go away? Can you get a
dump of the transaction reservation sizes for the filesystem in
question (the xfs_db logres command can do this, IIRC) and then run
the math on the reservations held from the dump of all the blocked
tasks holding locks to see if this matches the size of the journal
in the fs?

Cheers,

Dave.
-- 
Dave Chinner
dgc@kernel.org

  parent reply	other threads:[~2026-09-16 21:45 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-15  9:34 [bug report] fstests generic/774 hang again Shin'ichiro Kawasaki
2026-09-15 10:00 ` John Garry
2026-09-15 11:48   ` Shin'ichiro Kawasaki
2026-09-15 14:48     ` John Garry
2026-09-15 14:50       ` Darrick J. Wong
2026-09-15 15:41         ` John Garry
2026-09-16 10:23           ` John Garry
2026-09-16 22:23             ` Dave Chinner
2026-09-17  8:28               ` John Garry
2026-09-17 21:20                 ` Dave Chinner
2026-09-17  6:30             ` Shin'ichiro Kawasaki
2026-09-16  2:53     ` Shin'ichiro Kawasaki
2026-09-16 21:45 ` Dave Chinner [this message]
2026-09-17  6:52   ` Shin'ichiro Kawasaki
2026-09-17 10:45     ` Shin'ichiro Kawasaki
2026-09-17 21:24       ` Dave Chinner
2026-09-19 11:53         ` Shin'ichiro Kawasaki
2026-09-21 22:05           ` Dave Chinner
2026-09-25  1:32             ` Shin'ichiro Kawasaki
2026-09-27 21:30               ` Dave Chinner
2026-09-28  2:45                 ` Darrick J. Wong
2026-09-22  0:28   ` Darrick J. Wong
2026-09-25  1:38     ` Shin'ichiro Kawasaki
2026-09-25 23:03       ` Darrick J. Wong

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqsN4GOXthkF8rOI@dread \
    --to=dgc@kernel.org \
    --cc=djwong@kernel.org \
    --cc=john.garry@linux.dev \
    --cc=linux-xfs@vger.kernel.org \
    --cc=shinichiro.kawasaki@wdc.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.