* [PATCH 01/12] xfs: adjust checks for cross-AG reaping issues
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
@ 2026-09-25 4:10 ` Darrick J. Wong
2026-09-25 5:33 ` Christoph Hellwig
2026-09-25 4:11 ` [PATCH 02/12] xfs: actually check XFS_DIFLAG2_METADATA for metadata btree data forks Darrick J. Wong
` (11 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:10 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
While analyzing the reaping code for the previous patch, I studied the
xfs_fsblock_t code to reconstruct how this code works. AGs have a
superblock at the start of each AG which means that regions in the
xfsb_bitmap (aka xfs_fsblock_t bitmap) cannot cross an AG boundary, so
we don't need a loop here like we do for xfs_rtblock_t bitmaps.
However, the checks here are not quite right -- there can be quite a lot
of reserved space in COW forks, which means that the bitmap extent
actually can exceed XFS_MAX_BMBT_EXTLEN. We also don't have an explicit
debugging check that we never cross AGs, so add that here.
Cc: <stable@vger.kernel.org> # v6.8
Fixes: 66da11280f7ecd ("xfs: reintroduce reaping of file metadata blocks to xrep_reap_extents")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
fs/xfs/scrub/reap.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/fs/xfs/scrub/reap.c b/fs/xfs/scrub/reap.c
index 1980baa6e5eed7..2cb626afa0099e 100644
--- a/fs/xfs/scrub/reap.c
+++ b/fs/xfs/scrub/reap.c
@@ -796,9 +796,9 @@ xreap_fsmeta_extent(
xfs_agblock_t agbno_next = agbno + len;
int error = 0;
- ASSERT(len <= XFS_MAX_BMBT_EXTLEN);
ASSERT(sc->ip != NULL);
ASSERT(!sc->sa.pag);
+ ASSERT(agno == XFS_FSB_TO_AGNO(sc->mp, fsbno + len - 1));
/*
* We're reaping blocks after repairing file metadata, which means that
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 02/12] xfs: actually check XFS_DIFLAG2_METADATA for metadata btree data forks
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
2026-09-25 4:10 ` [PATCH 01/12] xfs: adjust checks for cross-AG reaping issues Darrick J. Wong
@ 2026-09-25 4:11 ` Darrick J. Wong
2026-09-25 5:33 ` Christoph Hellwig
2026-09-25 4:11 ` [PATCH 03/12] xfs: fix xchk_stats_estimate_bufsize to allow for the trailing null Darrick J. Wong
` (10 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:11 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
LOLLM noticed that xrep_dinode_check_dfork doesn't require that the
inode has the METADATA flag set in flags2. Add that check.
Cc: <stable@vger.kernel.org> # v6.14
Fixes: 702c90f4516223 ("xfs: support file data forks containing metadata btrees")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
---
fs/xfs/scrub/inode_repair.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/fs/xfs/scrub/inode_repair.c b/fs/xfs/scrub/inode_repair.c
index fab4015f6a9506..b047abd28bd341 100644
--- a/fs/xfs/scrub/inode_repair.c
+++ b/fs/xfs/scrub/inode_repair.c
@@ -1012,6 +1012,11 @@ xrep_dinode_bad_metabt_fork(
if (whichfork != XFS_DATA_FORK)
return true;
+ if (!xfs_has_metadir(sc->mp))
+ return true;
+ if (!(dip->di_flags2 & cpu_to_be64(XFS_DIFLAG2_METADATA)))
+ return true;
+
switch (be16_to_cpu(dip->di_metatype)) {
case XFS_METAFILE_RTRMAP:
return xrep_dinode_bad_rtrmapbt_fork(sc, dip, dfork_size);
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 03/12] xfs: fix xchk_stats_estimate_bufsize to allow for the trailing null
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
2026-09-25 4:10 ` [PATCH 01/12] xfs: adjust checks for cross-AG reaping issues Darrick J. Wong
2026-09-25 4:11 ` [PATCH 02/12] xfs: actually check XFS_DIFLAG2_METADATA for metadata btree data forks Darrick J. Wong
@ 2026-09-25 4:11 ` Darrick J. Wong
2026-09-25 5:34 ` Christoph Hellwig
2026-09-25 4:11 ` [PATCH 04/12] xfs: cross-reference the primary superblock Darrick J. Wong
` (9 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:11 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
LOLLM complains that xchk_stats_estimate_bufsize underestimates the
buffer size by 1 byte because scnprintf wants to store a null byte after
the newline in each line. Each call to scnprintf overwrites the null
left by the previous call, but we still need the null byte for the last
line in the file.
Cc: <stable@vger.kernel.org> # v6.6
Fixes: d7a74cad8f4513 ("xfs: track usage statistics of online fsck")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
---
fs/xfs/scrub/stats.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/fs/xfs/scrub/stats.c b/fs/xfs/scrub/stats.c
index 3339cae4b39d14..da2a40a91156f4 100644
--- a/fs/xfs/scrub/stats.c
+++ b/fs/xfs/scrub/stats.c
@@ -168,7 +168,8 @@ xchk_stats_estimate_bufsize(
ret += field_width + 1;
}
- return ret;
+ /* null byte for scnprintf */
+ return ret + 1;
}
/* Clear all counters. */
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 04/12] xfs: cross-reference the primary superblock
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (2 preceding siblings ...)
2026-09-25 4:11 ` [PATCH 03/12] xfs: fix xchk_stats_estimate_bufsize to allow for the trailing null Darrick J. Wong
@ 2026-09-25 4:11 ` Darrick J. Wong
2026-09-25 5:35 ` Christoph Hellwig
2026-09-25 4:11 ` [PATCH 05/12] xfs: write the rt super immediately during repair Darrick J. Wong
` (8 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:11 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
Make sure we cross-reference the primary superblock against the rest of
the AG 0 space metadata. Drop the unused @bp parameter of
xchk_superblock_xref to avoid complaints about uninitialized variables.
Cc: <stable@vger.kernel.org> # v4.16
Fixes: 52dc4b44af7419 ("xfs: cross-reference with the bnobt")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
| 15 +++++++--------
1 file changed, 7 insertions(+), 8 deletions(-)
--git a/fs/xfs/scrub/agheader.c b/fs/xfs/scrub/agheader.c
index fa5d32ec020aec..1dff6cb1b93628 100644
--- a/fs/xfs/scrub/agheader.c
+++ b/fs/xfs/scrub/agheader.c
@@ -35,8 +35,7 @@ xchk_setup_agheader(
/* Cross-reference with the other btrees. */
STATIC void
xchk_superblock_xref(
- struct xfs_scrub *sc,
- struct xfs_buf *bp)
+ struct xfs_scrub *sc)
{
struct xfs_mount *mp = sc->mp;
xfs_agnumber_t agno = sc->sm->sm_agno;
@@ -106,16 +105,12 @@ xchk_superblock(
struct xfs_dsb *sb;
struct xfs_perag *pag;
size_t sblen;
- xfs_agnumber_t agno;
+ xfs_agnumber_t agno = sc->sm->sm_agno;
uint32_t v2_ok;
__be32 features_mask;
int error;
__be16 vernum_mask;
- agno = sc->sm->sm_agno;
- if (agno == 0)
- return 0;
-
/*
* Grab an active reference to the perag structure. If we can't get
* it, we're racing with something that's tearing down the AG, so
@@ -125,6 +120,9 @@ xchk_superblock(
if (!pag)
return -ENOENT;
+ if (agno == 0)
+ goto out_xref;
+
error = xfs_sb_read_secondary(mp, sc->tp, agno, &bp);
/*
* The superblock verifier can return several different error codes
@@ -430,7 +428,8 @@ xchk_superblock(
if (memchr_inv((char *)sb + sblen, 0, BBTOB(bp->b_length) - sblen))
xchk_block_set_corrupt(sc, bp);
- xchk_superblock_xref(sc, bp);
+out_xref:
+ xchk_superblock_xref(sc);
out_pag:
xfs_perag_put(pag);
return error;
^ permalink raw reply related [flat|nested] 28+ messages in thread* Re: [PATCH 04/12] xfs: cross-reference the primary superblock
2026-09-25 4:11 ` [PATCH 04/12] xfs: cross-reference the primary superblock Darrick J. Wong
@ 2026-09-25 5:35 ` Christoph Hellwig
0 siblings, 0 replies; 28+ messages in thread
From: Christoph Hellwig @ 2026-09-25 5:35 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: cem, stable, linux-xfs
On Thu, Sep 24, 2026 at 09:11:32PM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
>
> Make sure we cross-reference the primary superblock against the rest of
> the AG 0 space metadata. Drop the unused @bp parameter of
> xchk_superblock_xref to avoid complaints about uninitialized variables.
Who complains about unused function paramters? That's not going to end
up well for methods tables..
But the patch itself looks good:
Reviewed-by: Christoph Hellwig <hch@lst.de>
^ permalink raw reply [flat|nested] 28+ messages in thread
* [PATCH 05/12] xfs: write the rt super immediately during repair
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (3 preceding siblings ...)
2026-09-25 4:11 ` [PATCH 04/12] xfs: cross-reference the primary superblock Darrick J. Wong
@ 2026-09-25 4:11 ` Darrick J. Wong
2026-09-25 5:36 ` Christoph Hellwig
2026-09-25 4:12 ` [PATCH 06/12] xfs: read rt superblock from disk during scrub Darrick J. Wong
` (7 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:11 UTC (permalink / raw)
To: cem, djwong; +Cc: linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
xrep_rgsuperblock has an unfortunate flaw -- even if the repair
transaction commits and is persisted to disk, a crash before the log
bwrite()s the buffer to disk means that the filesystem will not mount
after recovery. The xfs_mount keeps the rtsb buffer live for the
duration of the mount (scrubbers hold a file open on the fs) so let's
write the buffer out immediately to shorten that window.
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
fs/xfs/scrub/rgsuper.c | 20 ++++++++++++++++++--
1 file changed, 18 insertions(+), 2 deletions(-)
diff --git a/fs/xfs/scrub/rgsuper.c b/fs/xfs/scrub/rgsuper.c
index 6e2abe5dc27cd5..2db92cb89cc614 100644
--- a/fs/xfs/scrub/rgsuper.c
+++ b/fs/xfs/scrub/rgsuper.c
@@ -91,12 +91,28 @@ xrep_rgsuperblock(
struct xfs_scrub *sc)
{
struct xfs_buf *sb_bp;
+ struct xfs_buf *rtsb_bp;
+ int error;
ASSERT(rtg_rgno(sc->sr.rtg) == 0);
sb_bp = xfs_trans_getsb(sc->tp);
xfs_log_sb(sc->tp);
- xfs_log_rtsb(sc->tp, sb_bp);
- return 0;
+ rtsb_bp = xfs_log_rtsb(sc->tp, sb_bp);
+ if (!rtsb_bp)
+ return 0;
+
+ /* synchronous transaction to flush/release the buffer log item */
+ xfs_trans_set_sync(sc->tp);
+ error = xrep_trans_commit(sc);
+ if (error)
+ return error;
+
+ /* write the rt super out immediately */
+ xfs_buf_lock(rtsb_bp);
+ xfs_buf_hold(rtsb_bp);
+ error = xfs_bwrite(rtsb_bp);
+ xfs_buf_relse(rtsb_bp);
+ return error;
}
#endif /* CONFIG_XFS_ONLINE_REPAIR */
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 06/12] xfs: read rt superblock from disk during scrub
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (4 preceding siblings ...)
2026-09-25 4:11 ` [PATCH 05/12] xfs: write the rt super immediately during repair Darrick J. Wong
@ 2026-09-25 4:12 ` Darrick J. Wong
2026-09-25 5:37 ` Christoph Hellwig
2026-09-25 4:12 ` [PATCH 07/12] xfs: write the primary super immediately during repair Darrick J. Wong
` (6 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:12 UTC (permalink / raw)
To: cem, djwong; +Cc: linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
The rtgroup superblock scrubber isn't that useful right now because the
disk contents could have been corrupted after mount time. Perform an
uncached re-read of the rtsuper buffer to perform a fresh verification,
and schedule a repair if it doesn't pass muster.
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
fs/xfs/scrub/rgsuper.c | 13 +++++++++++--
1 file changed, 11 insertions(+), 2 deletions(-)
diff --git a/fs/xfs/scrub/rgsuper.c b/fs/xfs/scrub/rgsuper.c
index 2db92cb89cc614..102bd92e9003db 100644
--- a/fs/xfs/scrub/rgsuper.c
+++ b/fs/xfs/scrub/rgsuper.c
@@ -46,6 +46,7 @@ int
xchk_rgsuperblock(
struct xfs_scrub *sc)
{
+ struct xfs_buf *bp = NULL;
xfs_rgnumber_t rgno = sc->sm->sm_agno;
unsigned int flags;
int error;
@@ -78,9 +79,17 @@ xchk_rgsuperblock(
return error;
/*
- * Since we already validated the rt superblock at mount time, we don't
- * need to check its contents again. All we need is to cross-reference.
+ * Read the rt super from disk in case it's been corrupted since mount
+ * time. Crashing with a bad rt super may prevent remount, so we want
+ * to fix these things ASAP.
*/
+ error = xfs_buf_read_uncached(sc->mp->m_rtdev_targp, XFS_RTSB_DADDR,
+ sc->mp->m_sb.sb_blocksize >> BBSHIFT, &bp,
+ &xfs_rtsb_buf_ops);
+ if (!xchk_process_rt_error(sc, 0, 0, &error))
+ return error;
+ xfs_buf_relse(bp);
+
xchk_rgsuperblock_xref(sc);
return 0;
}
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 07/12] xfs: write the primary super immediately during repair
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (5 preceding siblings ...)
2026-09-25 4:12 ` [PATCH 06/12] xfs: read rt superblock from disk during scrub Darrick J. Wong
@ 2026-09-25 4:12 ` Darrick J. Wong
2026-09-25 5:38 ` Christoph Hellwig
2026-09-25 4:12 ` [PATCH 08/12] xfs: read primary superblock from disk during scrub Darrick J. Wong
` (5 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:12 UTC (permalink / raw)
To: cem, djwong; +Cc: linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
xrep_superblock has an unfortunate flaw -- it won't even try to repair
the primary super block, even though we have a live copy in memory that
we could log and write to disk right now. We should do this, because a
crash will render the filesystem unmountable on recovery.
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
| 3 +--
| 31 +++++++++++++++++++++++++++----
2 files changed, 28 insertions(+), 6 deletions(-)
--git a/fs/xfs/scrub/agheader.c b/fs/xfs/scrub/agheader.c
index 1dff6cb1b93628..43f3c3a5060b51 100644
--- a/fs/xfs/scrub/agheader.c
+++ b/fs/xfs/scrub/agheader.c
@@ -56,8 +56,7 @@ xchk_superblock_xref(
xchk_xref_is_only_owned_by(sc, agbno, 1, &XFS_RMAP_OINFO_FS);
xchk_xref_is_not_shared(sc, agbno, 1);
xchk_xref_is_not_cow_staging(sc, agbno, 1);
-
- /* scrub teardown will take care of sc->sa for us */
+ xchk_ag_free(sc, &sc->sa);
}
/*
--git a/fs/xfs/scrub/agheader_repair.c b/fs/xfs/scrub/agheader_repair.c
index a66b611588c47f..3c1947326d4448 100644
--- a/fs/xfs/scrub/agheader_repair.c
+++ b/fs/xfs/scrub/agheader_repair.c
@@ -44,12 +44,35 @@ xrep_superblock(
struct xfs_mount *mp = sc->mp;
struct xfs_buf *bp;
xfs_agnumber_t agno;
- int error;
+ int error = 0;
- /* Don't try to repair AG 0's sb; let xfs_repair deal with it. */
agno = sc->sm->sm_agno;
- if (agno == 0)
- return -EOPNOTSUPP;
+ if (agno == 0) {
+ /* Last chance to abort before we start committing fixes. */
+ if (xchk_should_terminate(sc, &error))
+ return error;
+
+ bp = xfs_trans_getsb(sc->tp);
+
+ /* Format the incore superblock into the primary sb buffer */
+ xfs_buf_zero(bp, 0, BBTOB(bp->b_length));
+ xfs_sb_to_disk(bp->b_addr, &mp->m_sb);
+
+ xfs_log_sb(sc->tp);
+
+ /* synchronous transaction to flush/release the buffer log item */
+ xfs_trans_set_sync(sc->tp);
+ error = xrep_trans_commit(sc);
+ if (error)
+ return error;
+
+ /* write the super out immediately */
+ xfs_buf_lock(bp);
+ xfs_buf_hold(bp);
+ error = xfs_bwrite(bp);
+ xfs_buf_relse(bp);
+ return error;
+ }
error = xfs_sb_get_secondary(mp, sc->tp, agno, &bp);
if (error)
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 08/12] xfs: read primary superblock from disk during scrub
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (6 preceding siblings ...)
2026-09-25 4:12 ` [PATCH 07/12] xfs: write the primary super immediately during repair Darrick J. Wong
@ 2026-09-25 4:12 ` Darrick J. Wong
2026-09-25 5:39 ` Christoph Hellwig
2026-09-25 4:12 ` [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending Darrick J. Wong
` (4 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:12 UTC (permalink / raw)
To: cem, djwong; +Cc: linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
The primary superblock scrubber isn't that useful right now because the
disk contents could have been corrupted after mount time. Perform an
uncached re-read of the super buffer to perform a fresh verification,
and schedule a repair if it doesn't pass muster.
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
| 34 +++++++++++++++++++++++++++++++++-
1 file changed, 33 insertions(+), 1 deletion(-)
--git a/fs/xfs/scrub/agheader.c b/fs/xfs/scrub/agheader.c
index 43f3c3a5060b51..a765d3e5e0686b 100644
--- a/fs/xfs/scrub/agheader.c
+++ b/fs/xfs/scrub/agheader.c
@@ -119,8 +119,40 @@ xchk_superblock(
if (!pag)
return -ENOENT;
- if (agno == 0)
+ if (agno == 0) {
+ /*
+ * Reread the primary super from disk in case it's become
+ * corrupted enough since mount time to fail the verifier.
+ * Crashing now with a bad primary super will prevent remount
+ * (which might recover the primary super from the log) so we
+ * want to trigger a repair to write the incore primary
+ * superblock out to disk ASAP.
+ *
+ * Note that we don't check the geometry (like we do for a
+ * secondary super) because the incore copy is the source of
+ * truth while the filesystem is mounted.
+ *
+ * If the read races with a write, we end up writing the
+ * primary super unnecessarily, but that's not a big deal.
+ */
+ error = xfs_buf_read_uncached(sc->mp->m_ddev_targp,
+ XFS_SB_DADDR, BTOBB(mp->m_sb.sb_sectsize), &bp,
+ &xfs_sb_buf_ops);
+ switch (error) {
+ case -EINVAL: /* also -EWRONGFS */
+ case -ENOSYS:
+ case -EFBIG:
+ error = -EFSCORRUPTED;
+ fallthrough;
+ default:
+ break;
+ }
+ if (!xchk_process_error(sc, agno, XFS_SB_BLOCK(mp), &error))
+ goto out_pag;
+
+ xfs_buf_relse(bp);
goto out_xref;
+ }
error = xfs_sb_read_secondary(mp, sc->tp, agno, &bp);
/*
^ permalink raw reply related [flat|nested] 28+ messages in thread* Re: [PATCH 08/12] xfs: read primary superblock from disk during scrub
2026-09-25 4:12 ` [PATCH 08/12] xfs: read primary superblock from disk during scrub Darrick J. Wong
@ 2026-09-25 5:39 ` Christoph Hellwig
0 siblings, 0 replies; 28+ messages in thread
From: Christoph Hellwig @ 2026-09-25 5:39 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: cem, linux-xfs
On Thu, Sep 24, 2026 at 09:12:35PM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
>
> The primary superblock scrubber isn't that useful right now because the
> disk contents could have been corrupted after mount time. Perform an
> uncached re-read of the super buffer to perform a fresh verification,
> and schedule a repair if it doesn't pass muster.
Can we have a test exercising this?
The patch itself looks good assuming we have something that actually
exercises this new code:
Reviewed-by: Christoph Hellwig <hch@lst.de>
^ permalink raw reply [flat|nested] 28+ messages in thread
* [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (7 preceding siblings ...)
2026-09-25 4:12 ` [PATCH 08/12] xfs: read primary superblock from disk during scrub Darrick J. Wong
@ 2026-09-25 4:12 ` Darrick J. Wong
2026-09-25 5:41 ` Christoph Hellwig
2026-09-25 4:13 ` [PATCH 10/12] xfs: jump out of xchk_bmap_check_rmap if fatal signals Darrick J. Wong
` (3 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:12 UTC (permalink / raw)
To: cem, djwong; +Cc: hch, stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
I noticed that if you run generic/476 for long enough on a zoned
filesystem that it runs out of space in the zoned section, fsstress gets
stuck in this loop trying to reserve space, failing, and kicking off the
gc. The garbage collector in turn is continuously running (so we don't
break out of the loop) but the process is no longer responsive to
signals and can't be killed.
Fix this by backing out to userspace for any pending signal. This
should be safe because space reservation is usually the first step in
any modification to a zoned file.
Cc: <hch@lst.de>
Cc: <stable@vger.kernel.org> # v6.15
Fixes: 0bb2193056b596 ("xfs: add support for zoned space reservations")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
---
fs/xfs/xfs_zone_space_resv.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/fs/xfs/xfs_zone_space_resv.c b/fs/xfs/xfs_zone_space_resv.c
index 7aa3c74fb2e01a..8edc82025abbd5 100644
--- a/fs/xfs/xfs_zone_space_resv.c
+++ b/fs/xfs/xfs_zone_space_resv.c
@@ -157,6 +157,11 @@ xfs_zoned_reserve_available(
if (error != -ENOSPC)
break;
+ if (signal_pending(current)) {
+ error = -ERESTARTSYS;
+ break;
+ }
+
/*
* Make sure to start GC if it is not running already. As we
* check the rtavailable count when filling up zones, GC is
^ permalink raw reply related [flat|nested] 28+ messages in thread* Re: [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending
2026-09-25 4:12 ` [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending Darrick J. Wong
@ 2026-09-25 5:41 ` Christoph Hellwig
2026-09-25 19:18 ` Darrick J. Wong
0 siblings, 1 reply; 28+ messages in thread
From: Christoph Hellwig @ 2026-09-25 5:41 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: cem, stable, linux-xfs, Hans Holmberg, Damien Le Moal
On Thu, Sep 24, 2026 at 09:12:51PM -0700, Darrick J. Wong wrote:
> From: Darrick J. Wong <djwong@kernel.org>
>
> I noticed that if you run generic/476 for long enough on a zoned
> filesystem that it runs out of space in the zoned section, fsstress gets
> stuck in this loop trying to reserve space, failing, and kicking off the
> gc.
I guess this is not directly fixed here, right?
> The garbage collector in turn is continuously running (so we don't
> break out of the loop) but the process is no longer responsive to
> signals and can't be killed.
>
> Fix this by backing out to userspace for any pending signal. This
> should be safe because space reservation is usually the first step in
> any modification to a zoned file.
>
> Cc: <hch@lst.de>
> Cc: <stable@vger.kernel.org> # v6.15
> Fixes: 0bb2193056b596 ("xfs: add support for zoned space reservations")
> Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
> ---
> fs/xfs/xfs_zone_space_resv.c | 5 +++++
> 1 file changed, 5 insertions(+)
>
>
> diff --git a/fs/xfs/xfs_zone_space_resv.c b/fs/xfs/xfs_zone_space_resv.c
> index 7aa3c74fb2e01a..8edc82025abbd5 100644
> --- a/fs/xfs/xfs_zone_space_resv.c
> +++ b/fs/xfs/xfs_zone_space_resv.c
> @@ -157,6 +157,11 @@ xfs_zoned_reserve_available(
> if (error != -ENOSPC)
> break;
>
> + if (signal_pending(current)) {
> + error = -ERESTARTSYS;
> + break;
> + }
I don't think normal fs read/write semantics allow for -ERESTARTSYS
on arbitrary signals. So this should probably be limited to
fatal_signal_pending().
^ permalink raw reply [flat|nested] 28+ messages in thread* Re: [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending
2026-09-25 5:41 ` Christoph Hellwig
@ 2026-09-25 19:18 ` Darrick J. Wong
2026-09-25 19:37 ` Darrick J. Wong
0 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 19:18 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: cem, stable, linux-xfs, Hans Holmberg, Damien Le Moal
On Thu, Sep 24, 2026 at 10:41:54PM -0700, Christoph Hellwig wrote:
> On Thu, Sep 24, 2026 at 09:12:51PM -0700, Darrick J. Wong wrote:
> > From: Darrick J. Wong <djwong@kernel.org>
> >
> > I noticed that if you run generic/476 for long enough on a zoned
> > filesystem that it runs out of space in the zoned section, fsstress gets
> > stuck in this loop trying to reserve space, failing, and kicking off the
> > gc.
>
> I guess this is not directly fixed here, right?
Correct. There's enough space freeing activity going on (truncate,
punch, unlink, etc) so that the gc actually does clear out zones; but
there are also enough writers consuming more space that they all end up
in this loop at some point, and some of the threads never manage to get
what they want from rtavailable.
> > The garbage collector in turn is continuously running (so we don't
> > break out of the loop) but the process is no longer responsive to
> > signals and can't be killed.
> >
> > Fix this by backing out to userspace for any pending signal. This
> > should be safe because space reservation is usually the first step in
> > any modification to a zoned file.
> >
> > Cc: <hch@lst.de>
> > Cc: <stable@vger.kernel.org> # v6.15
> > Fixes: 0bb2193056b596 ("xfs: add support for zoned space reservations")
> > Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
> > ---
> > fs/xfs/xfs_zone_space_resv.c | 5 +++++
> > 1 file changed, 5 insertions(+)
> >
> >
> > diff --git a/fs/xfs/xfs_zone_space_resv.c b/fs/xfs/xfs_zone_space_resv.c
> > index 7aa3c74fb2e01a..8edc82025abbd5 100644
> > --- a/fs/xfs/xfs_zone_space_resv.c
> > +++ b/fs/xfs/xfs_zone_space_resv.c
> > @@ -157,6 +157,11 @@ xfs_zoned_reserve_available(
> > if (error != -ENOSPC)
> > break;
> >
> > + if (signal_pending(current)) {
> > + error = -ERESTARTSYS;
> > + break;
> > + }
>
> I don't think normal fs read/write semantics allow for -ERESTARTSYS
> on arbitrary signals. So this should probably be limited to
> fatal_signal_pending().
Ok. That at least means I can ^C the g/476 test processes. :)
--D
^ permalink raw reply [flat|nested] 28+ messages in thread* Re: [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending
2026-09-25 19:18 ` Darrick J. Wong
@ 2026-09-25 19:37 ` Darrick J. Wong
0 siblings, 0 replies; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 19:37 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: cem, stable, linux-xfs, Hans Holmberg, Damien Le Moal
On Fri, Sep 25, 2026 at 12:18:30PM -0700, Darrick J. Wong wrote:
> On Thu, Sep 24, 2026 at 10:41:54PM -0700, Christoph Hellwig wrote:
> > On Thu, Sep 24, 2026 at 09:12:51PM -0700, Darrick J. Wong wrote:
> > > From: Darrick J. Wong <djwong@kernel.org>
> > >
> > > I noticed that if you run generic/476 for long enough on a zoned
> > > filesystem that it runs out of space in the zoned section, fsstress gets
> > > stuck in this loop trying to reserve space, failing, and kicking off the
> > > gc.
> >
> > I guess this is not directly fixed here, right?
>
> Correct. There's enough space freeing activity going on (truncate,
> punch, unlink, etc) so that the gc actually does clear out zones; but
> there are also enough writers consuming more space that they all end up
> in this loop at some point, and some of the threads never manage to get
> what they want from rtavailable.
>
> > > The garbage collector in turn is continuously running (so we don't
> > > break out of the loop) but the process is no longer responsive to
> > > signals and can't be killed.
> > >
> > > Fix this by backing out to userspace for any pending signal. This
> > > should be safe because space reservation is usually the first step in
> > > any modification to a zoned file.
> > >
> > > Cc: <hch@lst.de>
> > > Cc: <stable@vger.kernel.org> # v6.15
> > > Fixes: 0bb2193056b596 ("xfs: add support for zoned space reservations")
> > > Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
> > > ---
> > > fs/xfs/xfs_zone_space_resv.c | 5 +++++
> > > 1 file changed, 5 insertions(+)
> > >
> > >
> > > diff --git a/fs/xfs/xfs_zone_space_resv.c b/fs/xfs/xfs_zone_space_resv.c
> > > index 7aa3c74fb2e01a..8edc82025abbd5 100644
> > > --- a/fs/xfs/xfs_zone_space_resv.c
> > > +++ b/fs/xfs/xfs_zone_space_resv.c
> > > @@ -157,6 +157,11 @@ xfs_zoned_reserve_available(
> > > if (error != -ENOSPC)
> > > break;
> > >
> > > + if (signal_pending(current)) {
> > > + error = -ERESTARTSYS;
> > > + break;
> > > + }
> >
> > I don't think normal fs read/write semantics allow for -ERESTARTSYS
> > on arbitrary signals. So this should probably be limited to
> > fatal_signal_pending().
>
> Ok. That at least means I can ^C the g/476 test processes. :)
Heh. Rereading my notes it was always possible to kill -9 the process,
but the difficulty with fsstress is that it uses SIGTERM sets a global
flag to tell its threads to exit their file IO loops. Since that's not
a fatal action, fatal_signal_pending() returns false and so the thread
remains stuck looping in the kernel.
--D
^ permalink raw reply [flat|nested] 28+ messages in thread
* [PATCH 10/12] xfs: jump out of xchk_bmap_check_rmap if fatal signals
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (8 preceding siblings ...)
2026-09-25 4:12 ` [PATCH 09/12] xfs: break out of zoned reservation loop if signals are pending Darrick J. Wong
@ 2026-09-25 4:13 ` Darrick J. Wong
2026-09-25 5:42 ` Christoph Hellwig
2026-09-25 4:13 ` [PATCH 11/12] xfs: mark overfull dabtree node blocks as corrupt Darrick J. Wong
` (2 subsequent siblings)
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:13 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
LOLLM observes that we never check for fatal signals inside the loop
that checks file fork mappings against the rmapbt. This can lead to
lengthy delays in responding to a ^C, so add some bailout checks.
Cc: <stable@vger.kernel.org> # v4.17
Fixes: 5e777b62b0bcb6 ("xfs: bmap scrubber should do rmap xref with bmap for sparse files")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
---
fs/xfs/scrub/bmap.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/fs/xfs/scrub/bmap.c b/fs/xfs/scrub/bmap.c
index 4f3c7f681bd921..428be9ecdcd4d4 100644
--- a/fs/xfs/scrub/bmap.c
+++ b/fs/xfs/scrub/bmap.c
@@ -626,6 +626,10 @@ xchk_bmap_check_rmap(
struct xfs_ifork *ifp;
struct xfs_scrub *sc = sbcri->sc;
bool have_map;
+ int error = 0;
+
+ if (xchk_should_terminate(sc, &error))
+ return error;
/* Is this even the right fork? */
if (rec->rm_owner != I_INO(sc->ip))
@@ -657,6 +661,9 @@ xchk_bmap_check_rmap(
*/
check_rec = *rec;
while (have_map) {
+ if (xchk_should_terminate(sc, &error))
+ return error;
+
if (irec.br_startoff != check_rec.rm_offset)
xchk_fblock_set_corrupt(sc, sbcri->whichfork,
check_rec.rm_offset);
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 11/12] xfs: mark overfull dabtree node blocks as corrupt
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (9 preceding siblings ...)
2026-09-25 4:13 ` [PATCH 10/12] xfs: jump out of xchk_bmap_check_rmap if fatal signals Darrick J. Wong
@ 2026-09-25 4:13 ` Darrick J. Wong
2026-09-25 5:42 ` Christoph Hellwig
2026-09-25 4:13 ` [PATCH 12/12] xfs: check fork type against dabtree block magic number Darrick J. Wong
2026-10-08 13:32 ` [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Carlos Maiolino
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:13 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
LOLLM observes that we need to range-check the dabtree node record count
before using it to index an array. Fix that.
Cc: <stable@vger.kernel.org> # v4.15
Fixes: 7c4a07a424c18d ("xfs: scrub directory/attribute btrees")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
---
fs/xfs/scrub/dabtree.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/fs/xfs/scrub/dabtree.c b/fs/xfs/scrub/dabtree.c
index d4e5afbde85613..72cd880a5745e7 100644
--- a/fs/xfs/scrub/dabtree.c
+++ b/fs/xfs/scrub/dabtree.c
@@ -439,6 +439,10 @@ xchk_da_btree_block(
node = blk->bp->b_addr;
xfs_da3_node_hdr_from_disk(ip->i_mount, &nodehdr, node);
btree = nodehdr.btree;
+ if (nodehdr.count > dargs->geo->node_ents) {
+ xchk_da_set_corrupt(ds, level);
+ goto out_freebp;
+ }
*pmaxrecs = nodehdr.count;
blk->hashval = be32_to_cpu(btree[*pmaxrecs - 1].hashval);
if (level == 0) {
^ permalink raw reply related [flat|nested] 28+ messages in thread* [PATCH 12/12] xfs: check fork type against dabtree block magic number
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (10 preceding siblings ...)
2026-09-25 4:13 ` [PATCH 11/12] xfs: mark overfull dabtree node blocks as corrupt Darrick J. Wong
@ 2026-09-25 4:13 ` Darrick J. Wong
2026-09-25 5:42 ` Christoph Hellwig
2026-10-08 13:32 ` [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Carlos Maiolino
12 siblings, 1 reply; 28+ messages in thread
From: Darrick J. Wong @ 2026-09-25 4:13 UTC (permalink / raw)
To: cem, djwong; +Cc: stable, linux-xfs
From: Darrick J. Wong <djwong@kernel.org>
LOLLM complains that we fail to flag dabtree blocks that should never be
associated with the file fork -- data forks can't point to xattr leaves,
and attr forks can't point to dirent blocks. Add the missing checks.
Cc: <stable@vger.kernel.org> # v4.15
Fixes: 7c4a07a424c18d ("xfs: scrub directory/attribute btrees")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
---
fs/xfs/scrub/dabtree.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
diff --git a/fs/xfs/scrub/dabtree.c b/fs/xfs/scrub/dabtree.c
index 72cd880a5745e7..639228601cc6bc 100644
--- a/fs/xfs/scrub/dabtree.c
+++ b/fs/xfs/scrub/dabtree.c
@@ -410,6 +410,10 @@ xchk_da_btree_block(
XFS_BLFT_ATTR_LEAF_BUF);
blk->magic = XFS_ATTR_LEAF_MAGIC;
blk->hashval = xfs_attr_leaf_lasthash(blk->bp, pmaxrecs);
+ if (dargs->whichfork != XFS_ATTR_FORK) {
+ xchk_da_set_corrupt(ds, level);
+ goto out_freebp;
+ }
if (ds->tree_level != 0)
xchk_da_set_corrupt(ds, level);
break;
@@ -419,6 +423,10 @@ xchk_da_btree_block(
XFS_BLFT_DIR_LEAFN_BUF);
blk->magic = XFS_DIR2_LEAFN_MAGIC;
blk->hashval = xfs_dir2_leaf_lasthash(ip, blk->bp, pmaxrecs);
+ if (dargs->whichfork != XFS_DATA_FORK) {
+ xchk_da_set_corrupt(ds, level);
+ goto out_freebp;
+ }
if (ds->tree_level != 0)
xchk_da_set_corrupt(ds, level);
break;
@@ -428,6 +436,10 @@ xchk_da_btree_block(
XFS_BLFT_DIR_LEAF1_BUF);
blk->magic = XFS_DIR2_LEAF1_MAGIC;
blk->hashval = xfs_dir2_leaf_lasthash(ip, blk->bp, pmaxrecs);
+ if (dargs->whichfork != XFS_DATA_FORK) {
+ xchk_da_set_corrupt(ds, level);
+ goto out_freebp;
+ }
if (ds->tree_level != 0)
xchk_da_set_corrupt(ds, level);
break;
^ permalink raw reply related [flat|nested] 28+ messages in thread* Re: [PATCHSET] xfs: LLM-inspired bug fixes, part 18
2026-09-25 4:10 [PATCHSET] xfs: LLM-inspired bug fixes, part 18 Darrick J. Wong
` (11 preceding siblings ...)
2026-09-25 4:13 ` [PATCH 12/12] xfs: check fork type against dabtree block magic number Darrick J. Wong
@ 2026-10-08 13:32 ` Carlos Maiolino
12 siblings, 0 replies; 28+ messages in thread
From: Carlos Maiolino @ 2026-10-08 13:32 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: hch, stable, linux-xfs
On Thu, 24 Sep 2026 21:10:40 -0700, Darrick J. Wong wrote:
> Here's a eighteenth batch of xfs fixes resulting from a LLaMma. Mwa
> mwa mwa...
>
> If you're going to start using this code, I strongly recommend pulling
> from my git trees, which are linked below.
>
> With a bit of luck, this should all go splendidly.
> Comments and questions are, as always, welcome.
>
> [...]
Applied to for-next, thanks!
[01/12] xfs: adjust checks for cross-AG reaping issues
commit: 4c6668ecebc61bb0a53d4c6e95509bc960cf6bbc
[02/12] xfs: actually check XFS_DIFLAG2_METADATA for metadata btree data forks
commit: 436ae68743a6974021fefa21fd3093d9bf8e66e4
[03/12] xfs: fix xchk_stats_estimate_bufsize to allow for the trailing null
commit: 6b8d60784b61a17bd591cb4c6e8aa9ed29c24f5c
[04/12] xfs: cross-reference the primary superblock
commit: 0c61016006ab08f8346f9e7b1e38559bd4ea499e
[05/12] xfs: write the rt super immediately during repair
commit: 33dfad5a7bbf5c7a4626404eaf3dcdaf94fed0e9
[06/12] xfs: read rt superblock from disk during scrub
commit: cd16b947cac1463d6f3eeab17247a9db7aa10c38
[07/12] xfs: write the primary super immediately during repair
commit: 02722664508c1f7368663db6fc2410541c75379d
[08/12] xfs: read primary superblock from disk during scrub
commit: 612d2e7506c3553b1deedc694846d1c0364ae446
[09/12] xfs: break out of zoned reservation loop if signals are pending
(no commit info)
[10/12] xfs: jump out of xchk_bmap_check_rmap if fatal signals
commit: 49d300a51442bd64b627645b7b0b1cd66e492905
[11/12] xfs: mark overfull dabtree node blocks as corrupt
commit: f61c91a2e5994fa81141d6c06dadae1fed206018
[12/12] xfs: check fork type against dabtree block magic number
commit: 43030b72a11a532e7e3c21ac84a8625425b4201b
Best regards,
--
Carlos Maiolino <cem@kernel.org>
^ permalink raw reply [flat|nested] 28+ messages in thread