Linux XFS filesystem development
 help / color / mirror / Atom feed
* support for RT data checksums
@ 2026-09-24  9:59 Christoph Hellwig
  2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
                   ` (21 more replies)
  0 siblings, 22 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Hi all,

data checksums provide an additional safeguard against silent data loss.

In classic XFS they were hard to support because they need to be
atomically updated with the written data.  The zoned allocator solves
that problem because it always writes out of place, and the checksums
can be committed at the same as the metadata linking the newly written
file data into place.  In theory, a conventional allocator could be used
in combination with the always_cow option, but there are few upsides of
this compared to using the zoned allocator.

Data checksums are stored in per-realtime group files in the metadir,
similar to other modern RT metadata.  Unlike the checksum design in btrfs
or some other file system, the checksums are associated with the
physical blocks, and not with logical data in files.  This reduces the
mapping overhead, and significantly reduces the write amplification,
and also avoids duplicate checksums for reflinked files (although those
are not yet supported with the zoned allocator anyway).

The initial version provides two checksums algorithms: crc32c and crc64.
Both of those are cyclic redundancy check algorithms which provide known
good detection of bit flips that is better than general purpose hash
functions.  Both are not cryptographic hashes and thus do not provide any
kind of protection against intentional tampering with the data.
The crc32c parameters exactly match those use for xfs metadata checksums,
and also those used by the default btrfs checksum, and the NVMe PI
formats using crc32c.  The crc64 parameters exactly match those using
the NVMe PI formats using crc64.  crc32c provides reasonable assurance
for today's hardware, but might prove limiting for extremely large data
sets, crc64 fills that void, but probably warrants using > 4k file system
block sizes to amortize the overhead.

In this series the checksums are only used to check data integrity
and report issues with it.  That on it's own is a bit of a lame story,
but it is important as a building block for two additional features
under development:

 - exposing the checksum to userspace through io_uring with
   IORING_RW_ATTR_FLAG_PI.  A prototype of this exists, but it needs
   a bit more rework of common code than I'd like for the initial
   review.  As a side effect this support will also support salvaging
   data with bad checksums for analysis in userspace.
 - retrying reads through a different replica from RAID devices that
   provide it.  This has been proposed before and requires some
   hairy block layer infrastructure.  I have a prototype for MD-based
   mirrors that can be extended to other use cases.  The same mechanism
   can also be used to retry reads for the already checksum protected
   XFS metadata.

The performance drop when using data checksums is between not measurable
to about 2% for most workloads on HDD and SSD.  On fast enough SSDs
single threaded large reads can see up to 10% slow down as the
checksum validation runs on a strict per-cpu workqueue through the block
layer in-task bio completion.  If needed, different completion methods
that distribute the I/O completions could be added.

This series is based on the lazy bounce series queued up in a special
block branch, the prep series just send out and fixes queue up in
different maintainer tree, so it is best to use the git branch:

    git://git.infradead.org/users/hch/xfs.git xfs-crc

Gitweb:

    https://git.infradead.org/?p=users/hch/xfs.git;a=shortlog;h=refs/heads/xfs-crc

Note that the series will need a rebase on top of the fsverity work,
but the code points have been chosen to hopefully not conflict.

Diffstat:
 block/bio-integrity-fs.c       |    1 
 fs/iomap/bio.c                 |    3 
 fs/iomap/direct-io.c           |    2 
 fs/iomap/internal.h            |   21 ++
 fs/iomap/ioend.c               |   77 +++++++++
 fs/xfs/Kconfig                 |    1 
 fs/xfs/Makefile                |    2 
 fs/xfs/libxfs/xfs_cksum.h      |    7 
 fs/xfs/libxfs/xfs_format.h     |   46 +++++
 fs/xfs/libxfs/xfs_fs.h         |    6 
 fs/xfs/libxfs/xfs_health.h     |    4 
 fs/xfs/libxfs/xfs_log_format.h |    1 
 fs/xfs/libxfs/xfs_ondisk.h     |    4 
 fs/xfs/libxfs/xfs_rtbitmap.c   |   54 ++++--
 fs/xfs/libxfs/xfs_rtbitmap.h   |    3 
 fs/xfs/libxfs/xfs_rtcsumfile.c |   94 ++++++++++++
 fs/xfs/libxfs/xfs_rtcsumfile.h |  150 +++++++++++++++++++
 fs/xfs/libxfs/xfs_rtgroup.c    |   10 +
 fs/xfs/libxfs/xfs_rtgroup.h    |   26 +++
 fs/xfs/libxfs/xfs_sb.c         |   80 ++++++++++
 fs/xfs/libxfs/xfs_sb.h         |    1 
 fs/xfs/libxfs/xfs_shared.h     |    1 
 fs/xfs/libxfs/xfs_trans_resv.c |   23 ++
 fs/xfs/libxfs/xfs_trans_resv.h |    5 
 fs/xfs/scrub/agheader.c        |    5 
 fs/xfs/xfs_aops.c              |    4 
 fs/xfs/xfs_buf.c               |  130 +++++++++++++---
 fs/xfs/xfs_buf.h               |    4 
 fs/xfs/xfs_buf_item.c          |   10 +
 fs/xfs/xfs_buf_item.h          |    4 
 fs/xfs/xfs_buf_item_recover.c  |    7 
 fs/xfs/xfs_file.c              |   21 ++
 fs/xfs/xfs_inode.h             |    8 -
 fs/xfs/xfs_ioend.c             |  159 ++++++++++++++++++--
 fs/xfs/xfs_ioend.h             |    2 
 fs/xfs/xfs_iomap.c             |   72 ++++++++-
 fs/xfs/xfs_iomap.h             |    4 
 fs/xfs/xfs_iops.c              |   10 +
 fs/xfs/xfs_message.c           |    4 
 fs/xfs/xfs_message.h           |    1 
 fs/xfs/xfs_mount.h             |    9 +
 fs/xfs/xfs_platform.h          |    1 
 fs/xfs/xfs_reflink.c           |    2 
 fs/xfs/xfs_rtalloc.c           |   11 +
 fs/xfs/xfs_rtcsum.c            |  317 +++++++++++++++++++++++++++++++++++++++++
 fs/xfs/xfs_rtcsum.h            |   23 ++
 fs/xfs/xfs_super.c             |   14 +
 fs/xfs/xfs_sysfs.c             |    2 
 fs/xfs/xfs_trace.h             |    3 
 fs/xfs/xfs_verify_media.c      |  169 ++++++++++++++++++---
 fs/xfs/xfs_zone_alloc.c        |   40 ++++-
 fs/xfs/xfs_zone_gc.c           |   65 ++++++--
 fs/xfs/xfs_zone_priv.h         |    8 +
 include/linux/iomap.h          |   31 +++-
 54 files changed, 1619 insertions(+), 143 deletions(-)

^ permalink raw reply	[flat|nested] 69+ messages in thread

* [PATCH 01/21] block: export fs_bio_integrity_verify
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 20:29   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
                   ` (20 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel, Damien Le Moal

From: Damien Le Moal <dlemoal@kernel.org>

XFS will need to call this in the iomap_write_ops.read_folio_range
implementation.

Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 block/bio-integrity-fs.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/block/bio-integrity-fs.c b/block/bio-integrity-fs.c
index c8e91ada8ca6..20d0a3d52932 100644
--- a/block/bio-integrity-fs.c
+++ b/block/bio-integrity-fs.c
@@ -74,6 +74,7 @@ int fs_bio_integrity_verify(struct bio *bio, struct bvec_iter *data_iter)
 		bio_integrity_bytes(bi, data_iter->bi_size >> SECTOR_SHIFT);
 	return blk_status_to_errno(bio_integrity_verify(bio, data_iter));
 }
+EXPORT_SYMBOL_GPL(fs_bio_integrity_verify);
 
 static int __init fs_bio_integrity_init(void)
 {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 02/21] iomap: add support for data checksumming
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
  2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 21:39   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
                   ` (19 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Add a new refcounted structure for a data checksum attached to
an iomap_ioend, and helpers to allocate, split and access it.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/iomap/bio.c        |  3 +-
 fs/iomap/direct-io.c  |  2 +-
 fs/iomap/internal.h   | 21 +++++++++---
 fs/iomap/ioend.c      | 77 ++++++++++++++++++++++++++++++++++++++++---
 include/linux/iomap.h | 31 ++++++++++++++++-
 5 files changed, 122 insertions(+), 12 deletions(-)

diff --git a/fs/iomap/bio.c b/fs/iomap/bio.c
index d46c2f8ea18c..7b8fddd6d061 100644
--- a/fs/iomap/bio.c
+++ b/fs/iomap/bio.c
@@ -151,7 +151,8 @@ int iomap_bio_read_folio_range(const struct iomap_iter *iter,
 
 	if (!bio ||
 	    bio_end_sector(bio) != iomap_sector(&iter->iomap, iter->pos) ||
-	    bio->bi_iter.bi_size > iomap_max_bio_size(&iter->iomap) - plen ||
+	    bio->bi_iter.bi_size >
+			iomap_max_bio_size(iter->inode, &iter->iomap) - plen ||
 	    !bio_add_folio(bio, folio, plen, offset_in_folio(folio, iter->pos)))
 		iomap_read_alloc_bio(iter, ctx, plen);
 	return 0;
diff --git a/fs/iomap/direct-io.c b/fs/iomap/direct-io.c
index 41fdc90a9094..7cf70fb0165a 100644
--- a/fs/iomap/direct-io.c
+++ b/fs/iomap/direct-io.c
@@ -344,7 +344,7 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter,
 		struct iomap_dio *dio, loff_t pos, unsigned int alignment,
 		blk_opf_t op)
 {
-	unsigned int maxsize = iomap_max_bio_size(&iter->iomap);
+	unsigned int maxsize = iomap_max_bio_size(iter->inode, &iter->iomap);
 	unsigned int nr_vecs;
 	struct bio *bio;
 	ssize_t ret;
diff --git a/fs/iomap/internal.h b/fs/iomap/internal.h
index 74e898b196dc..7ac5e300e010 100644
--- a/fs/iomap/internal.h
+++ b/fs/iomap/internal.h
@@ -4,17 +4,28 @@
 
 #define IOEND_BATCH_SIZE	4096
 
+static inline unsigned int max_csum_io_size(const struct inode *inode,
+		const struct iomap *iomap)
+{
+	return IOMAP_CSUM_MAX_SIZE << (inode->i_blkbits - iomap->csum_shift);
+}
+
 /*
  * Normally we can build bios as big as the data structure supports.
  *
- * But for integrity protected I/O we need to respect the maximum size of the
- * single contiguous allocation for the integrity buffer.
+ * But for checksum or integrity protected I/O we need to respect the maximum
+ * size of the single contiguous allocation for the checksum/integrity buffer.
  */
-static inline size_t iomap_max_bio_size(const struct iomap *iomap)
+static inline size_t iomap_max_bio_size(const struct inode *inode,
+		const struct iomap *iomap)
 {
+	size_t max = BIO_MAX_SIZE;
+
 	if (iomap->flags & IOMAP_F_INTEGRITY)
-		return max_integrity_io_size(bdev_limits(iomap->bdev));
-	return BIO_MAX_SIZE;
+		max = min(max, max_integrity_io_size(bdev_limits(iomap->bdev)));
+	if (iomap->csum_shift)
+		max = min(max, max_csum_io_size(inode, iomap));
+	return max;
 }
 
 u32 iomap_finish_ioend_buffered_read(struct iomap_ioend *ioend);
diff --git a/fs/iomap/ioend.c b/fs/iomap/ioend.c
index bbebecc31670..877f0f470947 100644
--- a/fs/iomap/ioend.c
+++ b/fs/iomap/ioend.c
@@ -14,6 +14,7 @@
 struct bio_set iomap_ioend_bioset;
 EXPORT_SYMBOL_GPL(iomap_ioend_bioset);
 static struct bio_set iomap_ioend_split_bioset;
+static mempool_t ioend_csum_pool;
 
 struct iomap_ioend *iomap_init_ioend(struct inode *inode,
 		struct bio *bio, loff_t file_offset, u16 ioend_flags)
@@ -25,6 +26,7 @@ struct iomap_ioend *iomap_init_ioend(struct inode *inode,
 	ioend->io_parent = NULL;
 	INIT_LIST_HEAD(&ioend->io_list);
 	ioend->io_flags = ioend_flags;
+	ioend->io_csum_shift = 0;
 	ioend->io_bvec_offset = bio->bi_iter.bi_offset;
 	ioend->io_inode = inode;
 	ioend->io_offset = file_offset;
@@ -32,10 +34,55 @@ struct iomap_ioend *iomap_init_ioend(struct inode *inode,
 	ioend->io_sector = bio->bi_iter.bi_sector;
 	ioend->io_vi = NULL;
 	ioend->io_private = NULL;
+	ioend->io_csum = NULL;
 	return ioend;
 }
 EXPORT_SYMBOL_GPL(iomap_init_ioend);
 
+static void *__iomap_csum_alloc(struct iomap_ioend *ioend)
+{
+	size_t csum_size = iomap_csum_size(ioend);
+
+	WARN_ON_ONCE(csum_size > IOMAP_CSUM_MAX_SIZE);
+	if (csum_size <= sizeof(ioend->io_csum_inline)) {
+		ioend->io_flags |= IOMAP_IOEND_CSUM_INLINE;
+		return ioend->io_csum_inline;
+	}
+
+	ioend->io_csum_alloc = kmalloc(csum_size,
+		GFP_NOWAIT | __GFP_NOMEMALLOC | __GFP_NORETRY);
+	if (!ioend->io_csum_alloc) {
+		struct page *page = mempool_alloc(&ioend_csum_pool, GFP_NOFS);
+
+		ioend->io_flags |= IOMAP_IOEND_CSUM_MEMPOOL;
+		ioend->io_csum_alloc = page_address(page);
+	}
+	return ioend->io_csum_alloc;
+}
+
+void *iomap_csum_alloc(struct iomap_ioend *ioend, u8 csum_shift)
+{
+	ioend->io_csum_shift = csum_shift;
+	ioend->io_csum = __iomap_csum_alloc(ioend);
+	return ioend->io_csum;
+}
+EXPORT_SYMBOL_GPL(iomap_csum_alloc);
+
+void iomap_csum_free(struct iomap_ioend *ioend)
+{
+	if (ioend->io_flags & IOMAP_IOEND_CSUM_MEMPOOL) {
+		mempool_free(virt_to_page(ioend->io_csum_alloc),
+				&ioend_csum_pool);
+	} else if (!(ioend->io_flags & IOMAP_IOEND_CSUM_INLINE)) {
+		kfree(ioend->io_csum_alloc);
+	}
+
+	ioend->io_flags &=
+		~(IOMAP_IOEND_CSUM_MEMPOOL | IOMAP_IOEND_CSUM_INLINE);
+	ioend->io_csum_shift = 0;
+}
+EXPORT_SYMBOL_GPL(iomap_csum_free);
+
 /*
  * We're now finished for good with this ioend structure.  Update the folio
  * state, release holds on bios, and finally free up memory.  Do not use the
@@ -178,7 +225,7 @@ static bool iomap_can_add_to_ioend(struct iomap_writepage_ctx *wpc, loff_t pos,
 	struct iomap_ioend *ioend = wpc->wb_ctx;
 
 	if (ioend->io_bio.bi_iter.bi_size >
-	    iomap_max_bio_size(&wpc->iomap) - map_len)
+	    iomap_max_bio_size(wpc->inode, &wpc->iomap) - map_len)
 		return false;
 	if (ioend_flags & IOMAP_IOEND_BOUNDARY)
 		return false;
@@ -321,7 +368,7 @@ EXPORT_SYMBOL_GPL(iomap_ioend_integrity_verify);
 
 static u32 iomap_finish_ioend(struct iomap_ioend *ioend, int error)
 {
-	if (ioend->io_parent) {
+	if (ioend->io_flags & IOMAP_IOEND_CHAINED) {
 		struct bio *bio = &ioend->io_bio;
 
 		ioend = ioend->io_parent;
@@ -334,6 +381,9 @@ static u32 iomap_finish_ioend(struct iomap_ioend *ioend, int error)
 	if (!atomic_dec_and_test(&ioend->io_remaining))
 		return 0;
 
+	if (iomap_has_csum(ioend))
+		iomap_csum_free(ioend);
+
 	if (ioend->io_flags & IOMAP_IOEND_DIRECT)
 		return iomap_finish_ioend_direct(ioend);
 	if (bio_op(&ioend->io_bio) == REQ_OP_READ)
@@ -409,6 +459,8 @@ static bool iomap_ioend_can_merge(struct iomap_ioend *ioend,
 	if (ioend->io_sector + (ioend->io_size >> SECTOR_SHIFT) !=
 	    next->io_sector)
 		return false;
+	if (ioend->io_csum || next->io_csum)
+		return false;
 	return true;
 }
 
@@ -498,8 +550,9 @@ struct iomap_ioend *iomap_split_ioend(struct iomap_ioend *ioend,
 	split->bi_end_io = bio->bi_end_io;
 
 	split_ioend = iomap_init_ioend(ioend->io_inode, split, ioend->io_offset,
-			ioend->io_flags);
+			ioend->io_flags | IOMAP_IOEND_CHAINED);
 	split_ioend->io_parent = ioend;
+	split_ioend->io_csum_shift = ioend->io_csum_shift;
 
 	atomic_inc(&ioend->io_remaining);
 	ioend->io_offset += split_ioend->io_size;
@@ -508,6 +561,12 @@ struct iomap_ioend *iomap_split_ioend(struct iomap_ioend *ioend,
 	split_ioend->io_sector = ioend->io_sector;
 	if (!is_append)
 		ioend->io_sector += (split_ioend->io_size >> SECTOR_SHIFT);
+
+	if (iomap_has_csum(ioend)) {
+		split_ioend->io_csum = ioend->io_csum;
+		ioend->io_csum += iomap_csum_size(split_ioend);
+	}
+
 	return split_ioend;
 }
 EXPORT_SYMBOL_GPL(iomap_split_ioend);
@@ -547,7 +606,9 @@ void iomap_bounce_read(struct iomap_ioend *orig_ioend, unsigned int minsize,
 		bio->bi_iter.bi_sector = sector;
 
 		ioend = iomap_init_ioend(inode, bio, file_offset,
-				orig_ioend->io_flags);
+				orig_ioend->io_flags &
+					~(IOMAP_IOEND_CSUM_MEMPOOL |
+					  IOMAP_IOEND_CSUM_INLINE));
 
 		total_len -= bio->bi_iter.bi_size;
 		file_offset += bio->bi_iter.bi_size;
@@ -593,6 +654,8 @@ void iomap_bounce_read_end_io(struct iomap_ioend *ioend, struct bio *orig_bio,
 	else
 		iomap_ioend_unbounce(iomap_ioend_from_bio(orig_bio), ioend);
 
+	if (iomap_has_csum(ioend))
+		iomap_csum_free(ioend);
 	bio_free_folios(&ioend->io_bio);
 	if (bio_integrity(&ioend->io_bio))
 		fs_bio_integrity_free(&ioend->io_bio);
@@ -617,8 +680,14 @@ static int __init iomap_ioend_init(void)
 			   BIOSET_NEED_BVECS);
 	if (error)
 		goto out_exit_ioend_bioset;
+	error = mempool_init_page_pool(&ioend_csum_pool, BIO_POOL_SIZE,
+			get_order(IOMAP_CSUM_MAX_SIZE));
+	if (error)
+		goto out_exit_ioend_split_bioset;
 	return 0;
 
+out_exit_ioend_split_bioset:
+	bioset_exit(&iomap_ioend_split_bioset);
 out_exit_ioend_bioset:
 	bioset_exit(&iomap_ioend_bioset);
 	return error;
diff --git a/include/linux/iomap.h b/include/linux/iomap.h
index 59718f73c15a..e7db13ea0f70 100644
--- a/include/linux/iomap.h
+++ b/include/linux/iomap.h
@@ -133,6 +133,7 @@ struct iomap {
 	u64			length;	/* length of mapping, bytes */
 	u16			type;	/* type of mapping */
 	u16			flags;	/* flags for mapping */
+	u8			csum_shift; /* ilog() of csum size */
 	struct block_device	*bdev;	/* block device for I/O */
 	struct dax_device	*dax_dev; /* dax_dev for dax operations */
 	void			*inline_data;
@@ -489,6 +490,12 @@ sector_t iomap_bmap(struct address_space *mapping, sector_t bno,
 #else
 #define IOMAP_IOEND_INTEGRITY		0
 #endif /* CONFIG_BLK_DEV_INTEGRITY */
+/* chained ioend that has io_parent */
+#define IOMAP_IOEND_CHAINED		(1U << 6)
+/* using io_csum_inline */
+#define IOMAP_IOEND_CSUM_INLINE		(1U << 7)
+/* io_csum is backed by a mempool */
+#define IOMAP_IOEND_CSUM_MEMPOOL	(1U << 8)
 
 /*
  * Flags that if set on either ioend prevent the merge of two ioends.
@@ -522,14 +529,20 @@ static inline u16 iomap_ioend_flags(const struct iomap *iomap)
 struct iomap_ioend {
 	struct list_head	io_list;	/* next ioend in chain */
 	u16			io_flags;	/* IOMAP_IOEND_* */
+	u8			io_csum_shift;	/* ilog(2) of csum size */
 	u32			io_bvec_offset;	/* offset into first bvec */
 	struct inode		*io_inode;	/* file being written to */
 	size_t			io_size;	/* size of the extent */
 	atomic_t		io_remaining;	/* completetion defer count */
 	int			io_error;	/* stashed away status */
-	struct iomap_ioend	*io_parent;	/* parent for completions */
+	union {
+		struct iomap_ioend *io_parent;	/* parent for completions */
+		void		*io_csum_alloc; /* original csum allocation. */
+		u8		io_csum_inline[sizeof(void *)];
+	};
 	loff_t			io_offset;	/* offset in the file */
 	sector_t		io_sector;	/* start sector of ioend */
+	void			*io_csum;	/* data checksum */
 	void			*io_private;	/* file system private data */
 	struct fsverity_info	*io_vi;		/* fsverity info */
 	struct bio		io_bio;		/* MUST BE LAST! */
@@ -547,6 +560,22 @@ static inline struct iomap_ioend *iomap_ioend_from_bio(struct bio *bio)
 	.bi_offset	= (_ioend)->io_bvec_offset,	\
 }
 
+#define IOMAP_CSUM_MAX_SIZE	SZ_64K
+
+static inline bool iomap_has_csum(const struct iomap_ioend *ioend)
+{
+	return ioend->io_csum_shift > 0;
+}
+
+static inline size_t iomap_csum_size(const struct iomap_ioend *ioend)
+{
+	return DIV_ROUND_UP(ioend->io_size, i_blocksize(ioend->io_inode)) <<
+			ioend->io_csum_shift;
+}
+
+void *iomap_csum_alloc(struct iomap_ioend *ioend, u8 csum_shift);
+void iomap_csum_free(struct iomap_ioend *ioend);
+
 struct iomap_writeback_ops {
 	/*
 	 * Performs writeback on the passed in range
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
  2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
  2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 21:43   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
                   ` (18 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Add a new helper that reads a buffer asynchronously.  This is similar
to readahead, but doesn't become a no-op under memory or I/O congestion
and returns the buffer to be read.

The intended use is to kick off a read of data checksum buffers at
roughly the same time as the data read so that they are available
in the I/O completion handler.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_buf.c   | 130 ++++++++++++++++++++++++++++++++++++---------
 fs/xfs/xfs_buf.h   |   4 ++
 fs/xfs/xfs_trace.h |   3 ++
 3 files changed, 113 insertions(+), 24 deletions(-)

diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c
index eee491c01d8c..966c06bffeff 100644
--- a/fs/xfs/xfs_buf.c
+++ b/fs/xfs/xfs_buf.c
@@ -627,6 +627,40 @@ _xfs_buf_read(
 	return xfs_buf_iowait(bp);
 }
 
+/*
+ * If we've had a read error, then the contents of the buffer are invalid and
+ * should not be used.  To ensure that a followup read tries to pull the buffer
+ * from disk again, we clear the XBF_DONE flag and mark the buffer stale.
+ * This ensures that anyone who has a current reference to the buffer will
+ * interpret it's contents correctly and future cache lookups will also treat it
+ * as an empty, uninitialised buffer.
+ */
+static int
+xfs_buf_read_error(
+	struct xfs_buf		*bp,
+	xfs_failaddr_t		fa,
+	int			error)
+{
+	/*
+	 * Check against log shutdown for error reporting because metadata
+	 * writeback may require a read first and we need to report errors in
+	 * metadata writeback until the log is shut down.
+	 * High level transaction read functions already check against mount
+	 * shutdown, anyway, so we only need to be concerned about low level IO
+	 * interactions here.
+	 */
+	if (!xlog_is_shutdown(bp->b_mount->m_log))
+		xfs_buf_ioerror_alert(bp, fa);
+	xfs_buf_clear_flags(bp, XBF_DONE);
+	xfs_buf_stale(bp);
+	xfs_buf_relse(bp);
+
+	/* bad CRC means corrupted metadata */
+	if (error == -EFSBADCRC)
+		return -EFSCORRUPTED;
+	return error;
+}
+
 int
 xfs_buf_read_map(
 	struct xfs_buftarg	*target,
@@ -699,39 +733,87 @@ xfs_buf_read_map(
 	}
 
 	if (error)
-		goto out_ioerror;
-
+		return xfs_buf_read_error(bp, fa, error);
 	*bpp = bp;
 	return 0;
+}
 
-out_ioerror:
+int
+xfs_buf_read_async_wait(
+	struct xfs_buf		*bp)
+{
 	/*
-	 * Check against log shutdown for error reporting because metadata
-	 * writeback may require a read first and we need to report errors in
-	 * metadata writeback until the log is shut down.  High level
-	 * transaction read functions already check against mount shutdown, so
-	 * we only need to be concerned about low level/ IO interactions here.
+	 * Protect against the case where the checksum read is slower than the
+	 * data read.
 	 */
-	if (!xlog_is_shutdown(target->bt_mount->m_log))
-		xfs_buf_ioerror_alert(bp, fa);
+	if ((READ_ONCE(bp->b_flags) & (XBF_DONE | XBF_STALE)) == XBF_DONE &&
+	    !bp->b_error) {
+		trace_xfs_buf_read_async_wait(bp, 0, _RET_IP_);
+		return 0;
+	}
+
+	/* xfs_buf_find_lock can't return an error with 0 flags */
+	xfs_buf_find_lock(bp, 0);
+	trace_xfs_buf_read_async_lock(bp, 0, _RET_IP_);
+	if (bp->b_error)
+		return xfs_buf_read_error(bp, __builtin_return_address(0),
+				bp->b_error);
+	ASSERT(bp->b_ops);
+	xfs_buf_clear_flags(bp, XBF_READ);
+	xfs_buf_unlock(bp);
+	return 0;
+}
+
+/*
+ * Kick off an asynchronous read.  Unlike readahead, this returns a reference
+ * to the buffer, and reliably reads the data instead of skipping the read on
+ * memory pressure.
+ *
+ * The buffer may be locked when I/O is kicked off, but the I/O completion
+ * handler will unlock it.  The caller needs to lock itself if need to prevent
+ * concurrent access or to synchronize with I/O completion.
+ */
+int
+xfs_buf_read_async(
+	struct xfs_buftarg	*btp,
+	xfs_daddr_t		daddr,
+	size_t			numblks,
+	const struct xfs_buf_ops *ops,
+	struct xfs_buf		**bpp)
+{
+	DEFINE_SINGLE_BUF_MAP(map, daddr, numblks);
+	struct xfs_buf		*bp;
+	int			error;
+
+	ASSERT(!xfs_buftarg_is_mem(btp));
+
+	error = xfs_find_get_buf(btp, &map, 1, XBF_READ, &bp);
+	if (error)
+		return error;
 
 	/*
-	 * If we've had a read error, then the contents of the buffer are
-	 * invalid and should not be used. To ensure that a followup read tries
-	 * to pull the buffer from disk again, we clear the XBF_DONE flag and
-	 * mark the buffer stale. This ensures that anyone who has a current
-	 * reference to the buffer will interpret it's contents correctly and
-	 * future cache lookups will also treat it as an empty, uninitialised
-	 * buffer.
+	 * Do a lockless fast path check for a valid uptodate buffer and avoid
+	 * locking entirely in this case.
 	 */
-	xfs_buf_clear_flags(bp, XBF_DONE);
-	xfs_buf_stale(bp);
-	xfs_buf_relse(bp);
+	if ((READ_ONCE(bp->b_flags) & (XBF_DONE | XBF_STALE)) == XBF_DONE)
+		goto done;
 
-	/* bad CRC means corrupted metadata */
-	if (error == -EFSBADCRC)
-		return -EFSCORRUPTED;
-	return error;
+	/* xfs_buf_find_lock can't return an error with 0 flags */
+	xfs_buf_find_lock(bp, 0);
+	if (bp->b_flags & XBF_DONE) {
+		xfs_buf_unlock(bp);
+		goto done;
+	}
+	trace_xfs_buf_read_async(bp, 0, _RET_IP_);
+	XFS_STATS_INC(btp->bt_mount, xb_get_read);
+	xfs_buf_hold(bp);
+	bp->b_ops = ops;
+	xfs_buf_clear_flags(bp, XBF_WRITE);
+	xfs_buf_set_flags(bp, XBF_READ | XBF_ASYNC);
+	xfs_buf_submit(bp);
+done:
+	*bpp = bp;
+	return 0;
 }
 
 /*
diff --git a/fs/xfs/xfs_buf.h b/fs/xfs/xfs_buf.h
index a4729253b56f..1b352ec91aa6 100644
--- a/fs/xfs/xfs_buf.h
+++ b/fs/xfs/xfs_buf.h
@@ -255,6 +255,10 @@ xfs_buf_readahead(
 	return xfs_buf_readahead_map(target, &map, 1, ops);
 }
 
+int xfs_buf_read_async(struct xfs_buftarg *btp, xfs_daddr_t daddr,
+		size_t numblks, const struct xfs_buf_ops *ops,
+		struct xfs_buf **bpp);
+int xfs_buf_read_async_wait(struct xfs_buf *bp);
 int xfs_buf_get_uncached(struct xfs_buftarg *target, size_t numblks,
 		struct xfs_buf **bpp);
 int xfs_buf_read_uncached(struct xfs_buftarg *target, xfs_daddr_t daddr,
diff --git a/fs/xfs/xfs_trace.h b/fs/xfs/xfs_trace.h
index 2af9a1429ae9..afaabd3adc73 100644
--- a/fs/xfs/xfs_trace.h
+++ b/fs/xfs/xfs_trace.h
@@ -840,6 +840,9 @@ DEFINE_EVENT(xfs_buf_flags_class, name, \
 	TP_ARGS(bp, flags, caller_ip))
 DEFINE_BUF_FLAGS_EVENT(xfs_buf_get);
 DEFINE_BUF_FLAGS_EVENT(xfs_buf_read);
+DEFINE_BUF_FLAGS_EVENT(xfs_buf_read_async);
+DEFINE_BUF_FLAGS_EVENT(xfs_buf_read_async_wait);
+DEFINE_BUF_FLAGS_EVENT(xfs_buf_read_async_lock);
 DEFINE_BUF_FLAGS_EVENT(xfs_buf_readahead);
 
 TRACE_EVENT(xfs_buf_ioerror,
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (2 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 21:44   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
                   ` (17 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Translate from a disk address to the realtime group and group-relative
block numbers.  This will be needed by the data checksumming code.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_rtgroup.h | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)

diff --git a/fs/xfs/libxfs/xfs_rtgroup.h b/fs/xfs/libxfs/xfs_rtgroup.h
index fca2eb74908c..f26e324f0de3 100644
--- a/fs/xfs/libxfs/xfs_rtgroup.h
+++ b/fs/xfs/libxfs/xfs_rtgroup.h
@@ -390,4 +390,24 @@ xfs_rtgroup_raw_size(
 	return g->blocks;
 }
 
+static inline xfs_rgnumber_t
+xfs_daddr_to_rgno(struct xfs_mount *mp, xfs_daddr_t d)
+{
+	struct xfs_groups	*g = &mp->m_groups[XG_TYPE_RTG];
+	xfs_rfsblock_t		rbno = XFS_BB_TO_FSBT(mp, d) - g->start_fsb;
+
+	ASSERT(xfs_has_rtgroups(mp));
+	return div_u64(rbno, xfs_rtgroup_raw_size(mp));
+}
+
+static inline xfs_rgblock_t
+xfs_daddr_to_rgbno(struct xfs_mount *mp, xfs_daddr_t d)
+{
+	struct xfs_groups	*g = &mp->m_groups[XG_TYPE_RTG];
+	xfs_rfsblock_t		rbno = XFS_BB_TO_FSBT(mp, d) - g->start_fsb;
+
+	ASSERT(xfs_has_rtgroups(mp));
+	return do_div(rbno, xfs_rtgroup_raw_size(mp));
+}
+
 #endif /* __LIBXFS_RTGROUP_H */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (3 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 21:49   ` Darrick J. Wong
  2026-10-08 11:46   ` Anuj gupta
  2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
                   ` (16 subsequent siblings)
  21 siblings, 2 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Add a flag so that the shadow CIL buffer for a buffer log item is always
sizes to the maximum to prevent reallocations.  This will be used for the
RT checksum item, where we know that we are going to fill it up very soon,
and (almost) sequentially, so there is no point in doing a constant
realloc cycle when more data is added to it.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_trans_resv.c |  2 +-
 fs/xfs/libxfs/xfs_trans_resv.h |  1 +
 fs/xfs/xfs_buf_item.c          | 10 ++++++++++
 fs/xfs/xfs_buf_item.h          |  4 +++-
 4 files changed, 15 insertions(+), 2 deletions(-)

diff --git a/fs/xfs/libxfs/xfs_trans_resv.c b/fs/xfs/libxfs/xfs_trans_resv.c
index 3151e97ca8ff..5382ece51812 100644
--- a/fs/xfs/libxfs/xfs_trans_resv.c
+++ b/fs/xfs/libxfs/xfs_trans_resv.c
@@ -53,7 +53,7 @@ xfs_buf_log_overhead(void)
  * will be changed in a transaction.  size is used to tell how many
  * bytes should be reserved per item.
  */
-STATIC uint
+uint
 xfs_calc_buf_res(
 	uint		nbufs,
 	uint		size)
diff --git a/fs/xfs/libxfs/xfs_trans_resv.h b/fs/xfs/libxfs/xfs_trans_resv.h
index 336279e0fc61..1804e821f382 100644
--- a/fs/xfs/libxfs/xfs_trans_resv.h
+++ b/fs/xfs/libxfs/xfs_trans_resv.h
@@ -96,6 +96,7 @@ struct xfs_trans_resv {
 #define	XFS_ITRUNCATE_LOG_COUNT_REFLINK	8
 #define	XFS_WRITE_LOG_COUNT_REFLINK	8
 
+uint xfs_calc_buf_res(uint nbufs, uint size);
 void xfs_trans_resv_calc(struct xfs_mount *mp, struct xfs_trans_resv *resp);
 uint xfs_allocfree_block_count(struct xfs_mount *mp, uint num_ops);
 
diff --git a/fs/xfs/xfs_buf_item.c b/fs/xfs/xfs_buf_item.c
index 1a4ef34af8d5..644c3fb18310 100644
--- a/fs/xfs/xfs_buf_item.c
+++ b/fs/xfs/xfs_buf_item.c
@@ -252,6 +252,16 @@ xfs_buf_item_size(
 		offset += BBTOB(bp->b_maps[i].bm_len);
 	}
 
+	/*
+	 * For buffers with the prealloc flag, always size the allocation size
+	 * to the maximum as per the log reservation.  This avoids constant
+	 * realloc cycles for buffers that are filled sequentially in rapid
+	 * pace.  Note that the nvecs calaculation is kept from the regular
+	 * look as the buffer item formatting expects it.
+	 */
+	if (bip->bli_flags & XFS_BLI_PREALLOC)
+		*nbytes = xfs_calc_buf_res(bip->bli_format_count, bp->b_length);
+
 	/*
 	 * Round up the buffer size required to minimise the number of memory
 	 * allocations that need to be done as this item grows when relogged by
diff --git a/fs/xfs/xfs_buf_item.h b/fs/xfs/xfs_buf_item.h
index 3159325dd17b..ddc8ecc4683e 100644
--- a/fs/xfs/xfs_buf_item.h
+++ b/fs/xfs/xfs_buf_item.h
@@ -20,6 +20,7 @@ struct xfs_mount;
 #define XFS_BLI_STALE_INODE	(1u << 5)
 #define	XFS_BLI_INODE_BUF	(1u << 6)
 #define	XFS_BLI_ORDERED		(1u << 7)
+#define	XFS_BLI_PREALLOC	(1u << 8)
 
 #define XFS_BLI_FLAGS \
 	{ XFS_BLI_HOLD,		"HOLD" }, \
@@ -29,7 +30,8 @@ struct xfs_mount;
 	{ XFS_BLI_INODE_ALLOC_BUF, "INODE_ALLOC" }, \
 	{ XFS_BLI_STALE_INODE,	"STALE_INODE" }, \
 	{ XFS_BLI_INODE_BUF,	"INODE_BUF" }, \
-	{ XFS_BLI_ORDERED,	"ORDERED" }
+	{ XFS_BLI_ORDERED,	"ORDERED" }, \
+	{ XFS_BLI_PREALLOC,	"PREALLOC" }
 
 /*
  * This is the in core log item structure used to track information
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (4 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 22:03   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
                   ` (15 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

The upcoming RT data checksum feature will use larger than FSB blocks.
Prepare xfs_rtfile_initialize_blocks to pass the number of FSBs per
RT blocks, and to pass bmapi_flags to ask for contiguous allocation.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_rtbitmap.c | 44 +++++++++++++++++++++---------------
 fs/xfs/libxfs/xfs_rtbitmap.h |  3 ++-
 fs/xfs/xfs_rtalloc.c         |  4 ++--
 3 files changed, 30 insertions(+), 21 deletions(-)

diff --git a/fs/xfs/libxfs/xfs_rtbitmap.c b/fs/xfs/libxfs/xfs_rtbitmap.c
index 01536f4fb386..db6a22b4506a 100644
--- a/fs/xfs/libxfs/xfs_rtbitmap.c
+++ b/fs/xfs/libxfs/xfs_rtbitmap.c
@@ -1346,6 +1346,7 @@ xfs_rtfile_alloc_blocks(
 	struct xfs_inode	*ip,
 	xfs_fileoff_t		offset_fsb,
 	xfs_filblks_t		count_fsb,
+	uint32_t		bmapi_flags,
 	struct xfs_bmbt_irec	*map)
 {
 	struct xfs_mount	*mp = ip->i_mount;
@@ -1367,7 +1368,7 @@ xfs_rtfile_alloc_blocks(
 		goto out_trans_cancel;
 
 	error = xfs_bmapi_write(tp, ip, offset_fsb, count_fsb,
-			XFS_BMAPI_METADATA, 0, map, &nmap);
+			XFS_BMAPI_METADATA | bmapi_flags, 0, map, &nmap);
 	if (error)
 		goto out_trans_cancel;
 
@@ -1405,34 +1406,43 @@ xfs_rtfile_initialize_block(
 	struct xfs_rtgroup	*rtg,
 	enum xfs_rtg_inodes	type,
 	xfs_fsblock_t		fsbno,
-	void			*data)
+	xfs_filblks_t		nblks,
+	void			**data)
 {
 	struct xfs_mount	*mp = rtg_mount(rtg);
 	struct xfs_inode	*ip = rtg->rtg_inodes[type];
+	size_t			len = XFS_FSB_TO_B(mp, nblks);
+	size_t			copylen = len;
+	struct xfs_trans_res	tres = M_RES(mp)->tr_growrtzero;
 	struct xfs_trans	*tp;
 	struct xfs_buf		*bp;
-	const size_t		copylen = mp->m_blockwsize << XFS_WORDLOG;
 	int			error;
 
-	error = xfs_trans_alloc(mp, &M_RES(mp)->tr_growrtzero, 0, 0, 0, &tp);
+	tres.tr_logres *= nblks;
+	error = xfs_trans_alloc(mp, &tres, 0, 0, 0, &tp);
 	if (error)
 		return error;
 	xfs_ilock(ip, XFS_ILOCK_EXCL);
 	xfs_trans_ijoin(tp, ip, XFS_ILOCK_EXCL);
 
 	error = xfs_trans_get_buf(tp, mp->m_ddev_targp,
-			XFS_FSB_TO_DADDR(mp, fsbno), mp->m_bsize, 0, &bp);
+			XFS_FSB_TO_DADDR(mp, fsbno), BTOBB(len), 0, &bp);
 	if (error) {
 		xfs_trans_cancel(tp);
 		return error;
 	}
 
+	if (xfs_has_rtgroups(mp))
+		copylen -= sizeof(struct xfs_rtbuf_blkinfo);
+
 	xfs_rtfile_initialize_buf(rtg, type, bp, tp);
-	if (data)
-		memcpy(xfs_rtblock_payload(bp), data, copylen);
-	else
+	if (*data) {
+		memcpy(xfs_rtblock_payload(bp), *data, copylen);
+		*data += copylen;
+	} else {
 		memset(xfs_rtblock_payload(bp), 0, copylen);
-	xfs_trans_log_buf(tp, bp, 0, mp->m_sb.sb_blocksize - 1);
+	}
+	xfs_trans_log_buf(tp, bp, 0, len - 1);
 	return xfs_trans_commit(tp);
 }
 
@@ -1447,33 +1457,31 @@ xfs_rtfile_initialize_blocks(
 	enum xfs_rtg_inodes	type,
 	xfs_fileoff_t		offset_fsb,	/* offset to start from */
 	xfs_fileoff_t		end_fsb,	/* offset to allocate to */
+	xfs_filblks_t		bsize,
+	uint32_t		bmapi_flags,
 	void			*data)		/* data to fill the blocks */
 {
-	struct xfs_mount	*mp = rtg_mount(rtg);
-	const size_t		copylen = mp->m_blockwsize << XFS_WORDLOG;
-
 	while (offset_fsb < end_fsb) {
 		struct xfs_bmbt_irec	map;
 		xfs_filblks_t		i;
 		int			error;
 
 		error = xfs_rtfile_alloc_blocks(rtg->rtg_inodes[type],
-				offset_fsb, end_fsb - offset_fsb, &map);
+				offset_fsb, end_fsb - offset_fsb, bmapi_flags,
+				&map);
 		if (error)
 			return error;
 
 		/*
-		 * Now we need to clear the allocated blocks.
+		 * Now we need to clear or initialize the allocated blocks.
 		 *
 		 * Do this one block per transaction, to keep it simple.
 		 */
-		for (i = 0; i < map.br_blockcount; i++) {
+		for (i = 0; i < map.br_blockcount; i += bsize) {
 			error = xfs_rtfile_initialize_block(rtg, type,
-					map.br_startblock + i, data);
+					map.br_startblock + i, bsize, &data);
 			if (error)
 				return error;
-			if (data)
-				data += copylen;
 		}
 
 		offset_fsb = map.br_startoff + map.br_blockcount;
diff --git a/fs/xfs/libxfs/xfs_rtbitmap.h b/fs/xfs/libxfs/xfs_rtbitmap.h
index 750d74fbf4ed..e9e3378d15aa 100644
--- a/fs/xfs/libxfs/xfs_rtbitmap.h
+++ b/fs/xfs/libxfs/xfs_rtbitmap.h
@@ -410,7 +410,8 @@ void xfs_rtfile_initialize_buf(struct xfs_rtgroup *rtg,
 		struct xfs_trans *tp);
 int xfs_rtfile_initialize_blocks(struct xfs_rtgroup *rtg,
 		enum xfs_rtg_inodes type, xfs_fileoff_t offset_fsb,
-		xfs_fileoff_t end_fsb, void *data);
+		xfs_fileoff_t end_fsb, xfs_filblks_t bsize,
+		uint32_t bmapi_flags, void *data);
 int xfs_rtbitmap_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
 		struct xfs_trans *tp, bool init);
 int xfs_rtsummary_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
diff --git a/fs/xfs/xfs_rtalloc.c b/fs/xfs/xfs_rtalloc.c
index 84efe5a8fb11..78a1c066c7eb 100644
--- a/fs/xfs/xfs_rtalloc.c
+++ b/fs/xfs/xfs_rtalloc.c
@@ -1191,11 +1191,11 @@ xfs_growfs_rt_alloc_blocks(
 	}
 
 	error = xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_BITMAP, orbmblocks,
-			nmp->m_sb.sb_rbmblocks, NULL);
+			nmp->m_sb.sb_rbmblocks, 1, 0, NULL);
 	if (error)
 		goto out_free;
 	error = xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_SUMMARY, orsumblocks,
-			nmp->m_rsumblocks, NULL);
+			nmp->m_rsumblocks, 1, 0, NULL);
 out_free:
 	kfree(nmp);
 	return error;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (5 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
                   ` (14 subsequent siblings)
  21 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

xfs_open_zone_put will soon grow code that could sleep.  The only
caller that currently doesn't allow that is xfs_free_open_zones, but
that can easily drop zi_open_zones_lock because at this point nothing
else is looking at the open zone list.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_zone_alloc.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c
index 8cdd473d8144..bbdff9abe2b2 100644
--- a/fs/xfs/xfs_zone_alloc.c
+++ b/fs/xfs/xfs_zone_alloc.c
@@ -1007,7 +1007,10 @@ xfs_free_open_zones(
 	while ((oz = list_first_entry_or_null(&zi->zi_open_zones,
 			struct xfs_open_zone, oz_entry))) {
 		list_del(&oz->oz_entry);
+		spin_unlock(&zi->zi_open_zones_lock);
+
 		xfs_open_zone_put(oz);
+		spin_lock(&zi->zi_open_zones_lock);
 	}
 	spin_unlock(&zi->zi_open_zones_lock);
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 08/21] xfs: define the RT data checksum on-disk format
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (6 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 22:13   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
                   ` (13 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Add the on-disk format for the new RT data checksum format.

Keyed off a new read-only compat feature flag, this adds new fields to
the superblock to indicate the checksum algorithm used and the size of
the blocks containing the checksums.  These new fields reuse the
previously reserved padding to make efficient use of the space in the
on-disk superblock.

Data checksums are only supported on zoned RT devices, because they
require out of places writes to safely update the checksums for file
overwrites and a data/metadata split to be able to store the checksums
for a group in a file without causing recursion.  This means they can't
be supported directly on the data device at all, and only when using
the always_cow mode on regular RT devices, but that has no benefit
over the zoned allocator which is designed for out of place writes.

The initially supported data checksum algorithms are crc32c and crc64 as
specified by NVMe.  Both have extremely fast kernel implementations and
the strong data protection guarantees offered by CRC-style algorithms.
Both also happen to be support by NVMe for protection information so that
the userspace PI passthrough support (once extended to files on file
systems) can be reused to expose the checksums to applications and thus
provide true end-to-end data integrity.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_format.h     | 43 +++++++++++++++++--
 fs/xfs/libxfs/xfs_log_format.h |  1 +
 fs/xfs/libxfs/xfs_ondisk.h     |  4 +-
 fs/xfs/libxfs/xfs_sb.c         | 75 ++++++++++++++++++++++++++++++++++
 fs/xfs/libxfs/xfs_sb.h         |  1 +
 fs/xfs/scrub/agheader.c        |  5 +++
 fs/xfs/xfs_mount.h             |  7 ++++
 7 files changed, 132 insertions(+), 4 deletions(-)

diff --git a/fs/xfs/libxfs/xfs_format.h b/fs/xfs/libxfs/xfs_format.h
index dd0ed046fbe9..1be3d21910a7 100644
--- a/fs/xfs/libxfs/xfs_format.h
+++ b/fs/xfs/libxfs/xfs_format.h
@@ -179,7 +179,9 @@ typedef struct xfs_sb {
 	xfs_rgnumber_t	sb_rgcount;	/* number of realtime groups */
 	xfs_rtxlen_t	sb_rgextents;	/* size of a realtime group in rtx */
 	uint8_t		sb_rgblklog;    /* rt group number shift */
-	uint8_t		sb_pad[7];	/* zeroes */
+	uint8_t		sb_rtcsum_type;	/* RT device data checksum type */
+	uint8_t		sb_rtcsum_blklog; /* log2 of rtcsum bsize */
+	uint8_t		sb_pad[5];	/* zero */
 	xfs_rfsblock_t	sb_rtstart;	/* start of internal RT section (FSB) */
 	xfs_filblks_t	sb_rtreserved;	/* reserved (zoned) RT blocks */
 
@@ -272,7 +274,9 @@ struct xfs_dsb {
 	__be32		sb_rgcount;	/* # of realtime groups */
 	__be32		sb_rgextents;	/* size of rtgroup in rtx */
 	__u8		sb_rgblklog;    /* rt group number shift */
-	__u8		sb_pad[7];	/* zeroes */
+	__u8		sb_rtcsum_type;	/* RT device data checksum type */
+	__u8		sb_rtcsum_blklog; /* log2 of rtcsum bsize */
+	__u8		sb_pad[5];	/* zero */
 	__be64		sb_rtstart;	/* start of internal RT section (FSB) */
 	__be64		sb_rtreserved;	/* reserved (zoned) RT blocks */
 
@@ -374,6 +378,8 @@ xfs_sb_has_compat_feature(
 #define XFS_SB_FEAT_RO_COMPAT_RMAPBT   (1 << 1)		/* reverse map btree */
 #define XFS_SB_FEAT_RO_COMPAT_REFLINK  (1 << 2)		/* reflinked files */
 #define XFS_SB_FEAT_RO_COMPAT_INOBTCNT (1 << 3)		/* inobt block counts */
+#define XFS_SB_FEAT_RO_COMPAT_RTCSUM     (1 << 5)	/* RT data checksums */
+
 #define XFS_SB_FEAT_RO_COMPAT_ALL \
 		(XFS_SB_FEAT_RO_COMPAT_FINOBT | \
 		 XFS_SB_FEAT_RO_COMPAT_RMAPBT | \
@@ -866,6 +872,7 @@ enum xfs_metafile_type {
 	XFS_METAFILE_RTSUMMARY,		/* rt summary */
 	XFS_METAFILE_RTRMAP,		/* rt rmap */
 	XFS_METAFILE_RTREFCOUNT,	/* rt refcount */
+	XFS_METAFILE_RTCSUM,		/* rt data checksums */
 
 	XFS_METAFILE_MAX
 } __packed;
@@ -879,7 +886,8 @@ enum xfs_metafile_type {
 	{ XFS_METAFILE_RTBITMAP,	"rtbitmap" }, \
 	{ XFS_METAFILE_RTSUMMARY,	"rtsummary" }, \
 	{ XFS_METAFILE_RTRMAP,		"rtrmap" }, \
-	{ XFS_METAFILE_RTREFCOUNT,	"rtrefcount" }
+	{ XFS_METAFILE_RTREFCOUNT,	"rtrefcount" }, \
+	{ XFS_METAFILE_RTCSUM,		"rtcsum", }
 
 /*
  * On-disk inode structure.
@@ -1318,6 +1326,7 @@ static inline bool xfs_dinode_is_metadir(const struct xfs_dinode *dip)
  */
 #define XFS_RTBITMAP_MAGIC	0x424D505A	/* BMPZ */
 #define XFS_RTSUMMARY_MAGIC	0x53554D59	/* SUMY */
+#define XFS_RTCSUM_MAGIC	0x4353554D	/* CSUM */
 
 struct xfs_rtbuf_blkinfo {
 	__be32		rt_magic;	/* validity check on block */
@@ -2027,4 +2036,32 @@ struct xfs_acl {
 #define SGI_ACL_FILE_SIZE	(sizeof(SGI_ACL_FILE)-1)
 #define SGI_ACL_DEFAULT_SIZE	(sizeof(SGI_ACL_DEFAULT)-1)
 
+/*
+ * Size of a RT data checksum block.  Data reads must be contained in a single
+ * block, so this should be fairly large.
+ *
+ * The default is 32k, matching the default inode cluster size and the maximum
+ * memory allocation the Linux MM can handle in the fast path.  64k is primarily
+ * there so that his value never needs to be below the FSB size, even for 64k
+ * blocks.
+ */
+#define XFS_RTCSUM_BSIZE_LOG_MIN	15
+#define XFS_RTCSUM_BSIZE_LOG_MAX	16
+
+/*
+ * Data checksum types.
+ */
+#define XFS_CSUM_TYPE_NONE	0u
+#define XFS_CSUM_TYPE_CRC32C	1u
+#define XFS_CSUM_TYPE_CRC64	2u
+#define XFS_CSUM_TYPE_MAX	3u
+
+/*
+ * On-disk data checksums.
+ */
+union xfs_disk_csum {
+	__le32			crc32c;
+	__le64			crc64;
+};
+
 #endif /* __XFS_FORMAT_H__ */
diff --git a/fs/xfs/libxfs/xfs_log_format.h b/fs/xfs/libxfs/xfs_log_format.h
index a4e1b3eb425c..b1037b77338b 100644
--- a/fs/xfs/libxfs/xfs_log_format.h
+++ b/fs/xfs/libxfs/xfs_log_format.h
@@ -581,6 +581,7 @@ enum xfs_blft {
 	XFS_BLFT_SB_BUF,
 	XFS_BLFT_RTBITMAP_BUF,
 	XFS_BLFT_RTSUMMARY_BUF,
+	XFS_BLFT_RTCSUM_BUF,
 	XFS_BLFT_MAX_BUF = (1 << XFS_BLFT_BITS),
 };
 
diff --git a/fs/xfs/libxfs/xfs_ondisk.h b/fs/xfs/libxfs/xfs_ondisk.h
index 23cde1248f01..17ab9366b3b9 100644
--- a/fs/xfs/libxfs/xfs_ondisk.h
+++ b/fs/xfs/libxfs/xfs_ondisk.h
@@ -284,7 +284,9 @@ xfs_check_ondisk_structs(void)
 	XFS_CHECK_SB_OFFSET(sb_rgcount,			272);
 	XFS_CHECK_SB_OFFSET(sb_rgextents,		276);
 	XFS_CHECK_SB_OFFSET(sb_rgblklog,		280);
-	XFS_CHECK_SB_OFFSET(sb_pad,			281);
+	XFS_CHECK_SB_OFFSET(sb_rtcsum_type,		281);
+	XFS_CHECK_SB_OFFSET(sb_rtcsum_blklog,		282);
+	XFS_CHECK_SB_OFFSET(sb_pad,			283);
 	XFS_CHECK_SB_OFFSET(sb_rtstart,			288);
 	XFS_CHECK_SB_OFFSET(sb_rtreserved,		296);
 
diff --git a/fs/xfs/libxfs/xfs_sb.c b/fs/xfs/libxfs/xfs_sb.c
index f0341adbb879..3a470aec6c0c 100644
--- a/fs/xfs/libxfs/xfs_sb.c
+++ b/fs/xfs/libxfs/xfs_sb.c
@@ -487,6 +487,40 @@ xfs_validate_sb_zoned(
 	return 0;
 }
 
+static int
+xfs_validate_sb_csum(
+	struct xfs_mount	*mp,
+	struct xfs_sb		*sbp)
+{
+	unsigned int		rtcsum_bsize = 1u << sbp->sb_rtcsum_blklog;
+
+	if (!(sbp->sb_features_incompat & XFS_SB_FEAT_INCOMPAT_ZONED)) {
+		xfs_warn(mp, "data checksum required the zone allocator");
+		return -EINVAL;
+	}
+	if (sbp->sb_rtcsum_type >= XFS_CSUM_TYPE_MAX) {
+		xfs_warn(mp, "invalid data checksum type: 0x%x",
+			sbp->sb_rtcsum_type);
+		return -EINVAL;
+	}
+	if (sbp->sb_rtcsum_blklog < XFS_RTCSUM_BSIZE_LOG_MIN ||
+	    sbp->sb_rtcsum_blklog > XFS_RTCSUM_BSIZE_LOG_MAX) {
+		xfs_warn(mp,
+"invalid data checksum block log: %u (min %u/max %u)",
+			sbp->sb_rtcsum_blklog,
+			XFS_RTCSUM_BSIZE_LOG_MIN,
+			XFS_RTCSUM_BSIZE_LOG_MAX);
+		return -EINVAL;
+	}
+	if (rtcsum_bsize < sbp->sb_blocksize) {
+		xfs_warn(mp,
+"checksum block size must not be smaller than file system block size: %u/%u",
+			rtcsum_bsize, sbp->sb_blocksize);
+		return -EINVAL;
+	}
+	return 0;
+}
+
 /* Check the validity of the SB. */
 STATIC int
 xfs_validate_sb_common(
@@ -580,6 +614,17 @@ xfs_validate_sb_common(
 			if (error)
 				return error;
 		}
+		if (sbp->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
+			error = xfs_validate_sb_csum(mp, sbp);
+			if (error)
+				return error;
+		} else {
+			if (sbp->sb_rtcsum_type || sbp->sb_rtcsum_blklog) {
+				xfs_warn(mp,
+"rtcsum superblock fields must be zero for non-RTCSUM file systems.");
+				return -EINVAL;
+			}
+		}
 	} else if (sbp->sb_qflags & (XFS_PQUOTA_ENFD | XFS_GQUOTA_ENFD |
 				XFS_PQUOTA_CHKD | XFS_GQUOTA_CHKD)) {
 			xfs_notice(mp,
@@ -900,6 +945,14 @@ __xfs_sb_from_disk(
 		to->sb_rtstart = 0;
 		to->sb_rtreserved = 0;
 	}
+
+	if (to->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
+		to->sb_rtcsum_type = from->sb_rtcsum_type;
+		to->sb_rtcsum_blklog = from->sb_rtcsum_blklog;
+	} else {
+		to->sb_rtcsum_type = XFS_CSUM_TYPE_NONE;
+		to->sb_rtcsum_blklog = 0;
+	}
 }
 
 void
@@ -1071,6 +1124,11 @@ xfs_sb_to_disk(
 		to->sb_rtstart = cpu_to_be64(from->sb_rtstart);
 		to->sb_rtreserved = cpu_to_be64(from->sb_rtreserved);
 	}
+
+	if (from->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
+		to->sb_rtcsum_type = from->sb_rtcsum_type;
+		to->sb_rtcsum_blklog = from->sb_rtcsum_blklog;
+	}
 }
 
 /*
@@ -1243,6 +1301,20 @@ xfs_mount_sb_set_rextsize(
 	xfs_sb_mount_rextsize(mp, sbp);
 }
 
+uint8_t
+xfs_data_csum_shift(
+	uint8_t			csum)
+{
+	switch (csum) {
+	case XFS_CSUM_TYPE_CRC32C:
+		return 2;
+	case XFS_CSUM_TYPE_CRC64:
+		return 3;
+	default:
+		return 0;
+	}
+}
+
 /*
  * xfs_mount_common
  *
@@ -1311,6 +1383,9 @@ xfs_sb_mount_common(
 	mp->m_bsize = XFS_FSB_TO_BB(mp, 1);
 	mp->m_alloc_set_aside = xfs_alloc_set_aside(mp);
 	mp->m_ag_max_usable = xfs_alloc_ag_max_usable(mp);
+
+	mp->m_rtcsum_shift = xfs_data_csum_shift(mp->m_sb.sb_rtcsum_type);
+	mp->m_rtcsum_bsize = 1u << mp->m_sb.sb_rtcsum_blklog;
 }
 
 /*
diff --git a/fs/xfs/libxfs/xfs_sb.h b/fs/xfs/libxfs/xfs_sb.h
index 34d0dd374e9b..16f300c12e37 100644
--- a/fs/xfs/libxfs/xfs_sb.h
+++ b/fs/xfs/libxfs/xfs_sb.h
@@ -20,6 +20,7 @@ extern void	xfs_sb_mount_common(struct xfs_mount *mp, struct xfs_sb *sbp);
 void		xfs_sb_mount_rextsize(struct xfs_mount *mp, struct xfs_sb *sbp);
 void		xfs_mount_sb_set_rextsize(struct xfs_mount *mp,
 			struct xfs_sb *sbp, xfs_agblock_t rextsize);
+uint8_t		xfs_data_csum_shift(uint8_t csum);
 extern void	xfs_sb_from_disk(struct xfs_sb *to, struct xfs_dsb *from);
 extern void	xfs_sb_to_disk(struct xfs_dsb *to, struct xfs_sb *from);
 extern void	xfs_sb_quota_from_disk(struct xfs_sb *sbp);
diff --git a/fs/xfs/scrub/agheader.c b/fs/xfs/scrub/agheader.c
index 1fa66aa68e16..316a3085f95e 100644
--- a/fs/xfs/scrub/agheader.c
+++ b/fs/xfs/scrub/agheader.c
@@ -416,6 +416,11 @@ xchk_superblock(
 
 		if (memchr_inv(sb->sb_pad, 0, sizeof(sb->sb_pad)))
 			xchk_block_set_corrupt(sc, bp);
+
+		if (sb->sb_rtcsum_type != mp->m_sb.sb_rtcsum_type)
+			xchk_block_set_corrupt(sc, bp);
+		if (sb->sb_rtcsum_blklog != mp->m_sb.sb_rtcsum_blklog)
+			xchk_block_set_corrupt(sc, bp);
 	}
 
 	/* Everything else must be zero. */
diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
index 894ff2f4ecbd..fa86697f463a 100644
--- a/fs/xfs/xfs_mount.h
+++ b/fs/xfs/xfs_mount.h
@@ -191,6 +191,8 @@ typedef struct xfs_mount {
 	uint8_t			m_agno_log;	/* log #ag's */
 	uint8_t			m_sectbb_log;	/* sectlog - BBSHIFT */
 	int8_t			m_rtxblklog;	/* log2 of rextsize, if possible */
+	uint8_t			m_rtcsum_shift;	/* log2 of RT data csum size */
+	uint32_t		m_rtcsum_bsize;	/* rtcsum block size in bytes */
 
 	uint			m_blockmask;	/* sb_blocksize-1 */
 	uint			m_blockwsize;	/* sb_blocksize in words */
@@ -461,6 +463,11 @@ __XFS_HAS_FEAT(metadir, METADIR)
 __XFS_HAS_FEAT(zoned, ZONED)
 __XFS_HAS_FEAT(nolifetime, NOLIFETIME)
 
+static inline bool xfs_has_rtcsum(const struct xfs_mount *mp)
+{
+	return mp->m_sb.sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM;
+}
+
 static inline bool xfs_has_rtgroups(const struct xfs_mount *mp)
 {
 	/* all metadir file systems also allow rtgroups */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 09/21] xfs: add support for per-RTG csum files
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (7 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 22:24   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
                   ` (12 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Add the definitions for another per-RTG file that stores data checksums
for the RTG.  The file is fully preallocated at mkfs/growfs time, and
thus bmap lookups for it can be performed without taking locks.

Checksums are organized in fixed size large (initially 32Kib or 64KiB)
blocks to reduce the lookup and read overhead compare to using the
smaller file system block size.

Each block uses the standard RT file header for self-describing metadata
and can thus reuse the buf_ops including the verifier.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/Makefile                |   1 +
 fs/xfs/libxfs/xfs_cksum.h      |   7 +-
 fs/xfs/libxfs/xfs_health.h     |   4 +-
 fs/xfs/libxfs/xfs_rtbitmap.c   |  10 +++
 fs/xfs/libxfs/xfs_rtcsumfile.c |  94 +++++++++++++++++++++
 fs/xfs/libxfs/xfs_rtcsumfile.h | 150 +++++++++++++++++++++++++++++++++
 fs/xfs/libxfs/xfs_rtgroup.c    |  10 +++
 fs/xfs/libxfs/xfs_rtgroup.h    |   6 ++
 fs/xfs/libxfs/xfs_shared.h     |   1 +
 fs/xfs/xfs_buf_item_recover.c  |   7 ++
 fs/xfs/xfs_platform.h          |   1 +
 fs/xfs/xfs_rtalloc.c           |   7 ++
 12 files changed, 296 insertions(+), 2 deletions(-)
 create mode 100644 fs/xfs/libxfs/xfs_rtcsumfile.c
 create mode 100644 fs/xfs/libxfs/xfs_rtcsumfile.h

diff --git a/fs/xfs/Makefile b/fs/xfs/Makefile
index 399a207f2d0e..79ea4136fbba 100644
--- a/fs/xfs/Makefile
+++ b/fs/xfs/Makefile
@@ -64,6 +64,7 @@ xfs-y				+= $(addprefix libxfs/, \
 xfs-$(CONFIG_XFS_RT)		+= $(addprefix libxfs/, \
 				   xfs_rtbitmap.o \
 				   xfs_rtgroup.o \
+				   xfs_rtcsumfile.o \
 				   xfs_zones.o \
 				   )
 
diff --git a/fs/xfs/libxfs/xfs_cksum.h b/fs/xfs/libxfs/xfs_cksum.h
index 999a290cfd72..315b85ff78ae 100644
--- a/fs/xfs/libxfs/xfs_cksum.h
+++ b/fs/xfs/libxfs/xfs_cksum.h
@@ -2,7 +2,12 @@
 #ifndef _XFS_CKSUM_H
 #define _XFS_CKSUM_H 1
 
-#define XFS_CRC_SEED	(~(uint32_t)0)
+/*
+ * crc32c() does not include the inversion at the beginning and end, while
+ * crc64_nvme() does.
+ */
+#define XFS_CRC_SEED		(~(uint32_t)0)
+#define XFS_CRC64_SEED		0
 
 /*
  * Calculate the intermediate checksum for a buffer that has the CRC field
diff --git a/fs/xfs/libxfs/xfs_health.h b/fs/xfs/libxfs/xfs_health.h
index 1d45cf5789e8..349093e75268 100644
--- a/fs/xfs/libxfs/xfs_health.h
+++ b/fs/xfs/libxfs/xfs_health.h
@@ -72,6 +72,7 @@ struct xfs_rtgroup;
 #define XFS_SICK_RG_SUMMARY	(1 << 2)  /* rt groups summary */
 #define XFS_SICK_RG_RMAPBT	(1 << 3)  /* reverse mappings */
 #define XFS_SICK_RG_REFCNTBT	(1 << 4)  /* reference counts */
+#define XFS_SICK_RG_CSUM	(1 << 5)  /* data checksums */
 
 /* Observable health issues for AG metadata. */
 #define XFS_SICK_AG_SB		(1 << 0)  /* superblock */
@@ -119,7 +120,8 @@ struct xfs_rtgroup;
 				 XFS_SICK_RG_BITMAP | \
 				 XFS_SICK_RG_SUMMARY | \
 				 XFS_SICK_RG_RMAPBT | \
-				 XFS_SICK_RG_REFCNTBT)
+				 XFS_SICK_RG_REFCNTBT | \
+				 XFS_SICK_RG_CSUM)
 
 #define XFS_SICK_AG_PRIMARY	(XFS_SICK_AG_SB | \
 				 XFS_SICK_AG_AGF | \
diff --git a/fs/xfs/libxfs/xfs_rtbitmap.c b/fs/xfs/libxfs/xfs_rtbitmap.c
index db6a22b4506a..12e67e1abd3b 100644
--- a/fs/xfs/libxfs/xfs_rtbitmap.c
+++ b/fs/xfs/libxfs/xfs_rtbitmap.c
@@ -125,9 +125,18 @@ const struct xfs_buf_ops xfs_rtsummary_buf_ops = {
 	.verify_struct	= xfs_rtbuf_verify,
 };
 
+const struct xfs_buf_ops xfs_rtcsum_buf_ops = {
+	.name		= "xfs_rtcsum",
+	.magic		= { 0, cpu_to_be32(XFS_RTCSUM_MAGIC) },
+	.verify_read	= xfs_rtbuf_verify_read,
+	.verify_write	= xfs_rtbuf_verify_write,
+	.verify_struct	= xfs_rtbuf_verify,
+};
+
 static const struct xfs_buf_ops *xfs_rtblock_buf_ops[XFS_RTGI_MAX] = {
 	[XFS_RTGI_SUMMARY]	= &xfs_rtsummary_buf_ops,
 	[XFS_RTGI_BITMAP]	= &xfs_rtbitmap_buf_ops,
+	[XFS_RTGI_CSUM]		= &xfs_rtcsum_buf_ops,
 };
 
 const struct xfs_buf_ops *
@@ -143,6 +152,7 @@ xfs_rtblock_ops(
 static enum xfs_blft xfs_rtblock_buf_types[XFS_RTGI_MAX] = {
 	[XFS_RTGI_SUMMARY]	= XFS_BLFT_RTSUMMARY_BUF,
 	[XFS_RTGI_BITMAP]	= XFS_BLFT_RTBITMAP_BUF,
+	[XFS_RTGI_CSUM]		= XFS_BLFT_RTCSUM_BUF,
 };
 
 /* Release cached rt bitmap and summary buffers. */
diff --git a/fs/xfs/libxfs/xfs_rtcsumfile.c b/fs/xfs/libxfs/xfs_rtcsumfile.c
new file mode 100644
index 000000000000..fc71c52beafd
--- /dev/null
+++ b/fs/xfs/libxfs/xfs_rtcsumfile.c
@@ -0,0 +1,94 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Christoph Hellwig.
+ */
+#include "xfs_platform.h"
+#include "xfs_fs.h"
+#include "xfs_format.h"
+#include "xfs_log_format.h"
+#include "xfs_shared.h"
+#include "xfs_trans_resv.h"
+#include "xfs_bit.h"
+#include "xfs_mount.h"
+#include "xfs_inode.h"
+#include "xfs_bmap.h"
+#include "xfs_rtbitmap.h"
+#include "xfs_bmap_btree.h"
+#include "xfs_trans.h"
+#include "xfs_error.h"
+#include "xfs_health.h"
+#include "xfs_rtcsumfile.h"
+
+int
+xfs_rtcsum_bmap(
+	struct xfs_rtgroup	*rtg,
+	xfs_rgblock_t		rgbno,
+	xfs_daddr_t		*daddr)
+{
+	struct xfs_mount	*mp = rtg_mount(rtg);
+	struct xfs_inode	*csumip = rtg_csum(rtg);
+	struct xfs_ifork	*ifp = &csumip->i_df;
+	unsigned int		csum_block = xfs_rgb_to_rtcsumblock(mp, rgbno);
+	xfs_fileoff_t		start_fsb =
+		XFS_B_TO_FSB(mp, mp->m_rtcsum_bsize) * csum_block;
+	struct xfs_iext_cursor	icur;
+	struct xfs_bmbt_irec	got;
+
+	ASSERT(!xfs_need_iread_extents(ifp));
+
+	if (XFS_IS_CORRUPT(mp, ifp->if_nextents != 1))
+		goto sick;
+
+	/*
+	 * We can do an unlocked lookup here because the bmap btree for the
+	 * csum files is immutable once created.
+	 */
+	if (XFS_IS_CORRUPT(mp, !xfs_iext_lookup_extent(csumip, ifp, start_fsb,
+			&icur, &got)))
+		goto sick;
+	if (XFS_IS_CORRUPT(mp, got.br_startoff > start_fsb))
+		goto sick;
+
+	start_fsb -= got.br_startoff;
+	*daddr = XFS_FSB_TO_DADDR(mp, got.br_startblock + start_fsb);
+	return 0;
+sick:
+	xfs_rtginode_mark_sick(rtg, XFS_RTGI_CSUM);
+	return -EFSCORRUPTED;
+}
+
+xfs_off_t
+xfs_rtcsum_file_size(
+	struct xfs_rtgroup	*rtg)
+{
+	struct xfs_mount	*mp = rtg_mount(rtg);
+	uint64_t		raw_size;
+
+	raw_size = (xfs_off_t)rtg_blocks(rtg) << mp->m_rtcsum_shift;
+	return DIV_ROUND_UP_ULL(raw_size, xfs_rtcsum_payload_size(mp)) *
+			mp->m_rtcsum_bsize;
+}
+
+int
+xfs_rtcsum_alloc_blocks(
+	struct xfs_rtgroup	*rtg)
+{
+	struct xfs_mount	*mp = rtg_mount(rtg);
+
+	return xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_CSUM, 0,
+			XFS_B_TO_FSB(mp, rtg_csum(rtg)->i_disk_size),
+			XFS_B_TO_FSB(mp, mp->m_rtcsum_bsize),
+			XFS_BMAPI_CONTIG, NULL);
+}
+
+int
+xfs_rtcsum_create(
+	struct xfs_rtgroup	*rtg,
+	struct xfs_inode	*ip,
+	struct xfs_trans	*tp,
+	bool			init)
+{
+	ip->i_disk_size = xfs_rtcsum_file_size(rtg);
+	xfs_trans_log_inode(tp, ip, XFS_ILOG_CORE);
+	return 0;
+}
diff --git a/fs/xfs/libxfs/xfs_rtcsumfile.h b/fs/xfs/libxfs/xfs_rtcsumfile.h
new file mode 100644
index 000000000000..8aacac05adc3
--- /dev/null
+++ b/fs/xfs/libxfs/xfs_rtcsumfile.h
@@ -0,0 +1,150 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _XFS_RTCSUMFILE_H
+#define _XFS_RTCSUMFILE_H
+
+#include "xfs_rtgroup.h"
+
+/*
+ * Maximum size of a checksum buffer for writes, used for the log reservation.
+ */
+#define XFS_RTCSUM_MAX_WRITE	SZ_64K
+
+/*
+ * Size of the actual payload in the RT data checksum block.  This excludes the
+ * self-describing metadata header.
+ */
+static inline unsigned int
+xfs_rtcsum_payload_size(
+	struct xfs_mount	*mp)
+{
+	return mp->m_rtcsum_bsize - sizeof(struct xfs_rtbuf_blkinfo);
+}
+
+/* Convert data length in logical blocks to checksum length in bytes. */
+static inline unsigned int
+xfs_extlen_to_rtcsum_len(
+	struct xfs_mount	*mp,
+	xfs_extlen_t		nb)
+{
+	return nb << mp->m_rtcsum_shift;
+}
+
+/* Convert checksum length in bytes to data length in logical blocks. */
+static inline xfs_extlen_t
+xfs_rtcsum_len_to_extlen(
+	struct xfs_mount	*mp,
+	unsigned int		csum_len)
+{
+	return csum_len >> mp->m_rtcsum_shift;
+}
+
+/* Convert an rgbno to the csum byte position in the csum file. */
+static inline xfs_off_t
+xfs_rgb_to_rtcsumpos(
+	struct xfs_mount	*mp,
+	xfs_rtblock_t		rgbno)
+{
+	return (xfs_off_t)rgbno << mp->m_rtcsum_shift;
+}
+
+/* Convert an rgbno to the csum block index in the csum file. */
+static inline unsigned int
+xfs_rgb_to_rtcsumblock(
+	struct xfs_mount	*mp,
+	xfs_rtblock_t		rgbno)
+{
+	return div_u64(xfs_rgb_to_rtcsumpos(mp, rgbno),
+			xfs_rtcsum_payload_size(mp));
+}
+
+/* Convert an rgbno to a the checksum offset within an rt csum block. */
+static inline unsigned int
+xfs_rgb_to_rtcsumoff(
+	struct xfs_mount	*mp,
+	xfs_rgblock_t		rgbno)
+{
+	uint32_t		off;
+
+	div_u64_rem(xfs_rgb_to_rtcsumpos(mp, rgbno),
+			xfs_rtcsum_payload_size(mp), &off);
+	return sizeof(struct xfs_rtbuf_blkinfo) + off;
+}
+
+/* Convert an rtbno to a the checksum offset within an rt csum block. */
+static inline unsigned int
+xfs_rtb_to_rtcsumoff(
+	struct xfs_mount	*mp,
+	xfs_rtblock_t		fsbno)
+{
+	return xfs_rgb_to_rtcsumoff(mp, xfs_rtb_to_rgbno(mp, fsbno));
+}
+
+int xfs_rtcsum_bmap(struct xfs_rtgroup *rtg, xfs_rgblock_t rgbno,
+		xfs_daddr_t *daddr);
+xfs_off_t xfs_rtcsum_file_size(struct xfs_rtgroup *rtg);
+int xfs_rtcsum_alloc_blocks(struct xfs_rtgroup *rtg);
+int xfs_rtcsum_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
+		struct xfs_trans *tp, bool init);
+
+static inline unsigned int
+xfs_rtcsum_bufs_per_rtg(
+	struct xfs_rtgroup	*rtg)
+{
+	struct xfs_mount	*mp = rtg_mount(rtg);
+
+	if (xfs_has_rtcsum(mp))
+		return div_u64(xfs_rtcsum_file_size(rtg), mp->m_rtcsum_bsize);
+	return 0;
+}
+
+static inline xfs_filblks_t
+xfs_rtcsum_max_len(
+	struct xfs_mount	*mp,
+	xfs_fsblock_t		fsbno)
+{
+	return xfs_rtcsum_len_to_extlen(mp,
+			mp->m_rtcsum_bsize - xfs_rtb_to_rtcsumoff(mp, fsbno));
+}
+
+union xfs_csum {
+	uint32_t		crc32c;
+	uint64_t		crc64;
+};
+
+static __always_inline void
+xfs_csum_seed(
+	struct xfs_mount	*mp,
+	union xfs_csum		*csum)
+{
+	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
+		csum->crc32c = XFS_CRC_SEED;
+	else
+		csum->crc64 = XFS_CRC64_SEED;
+}
+
+static __always_inline void
+xfs_csum_gen(
+	struct xfs_mount	*mp,
+	void			*data,
+	unsigned int		len,
+	union xfs_csum		*csum)
+{
+	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
+		csum->crc32c = crc32c(csum->crc32c, data, len);
+	else
+		csum->crc64 = crc64_nvme(csum->crc64, data, len);
+}
+
+static __always_inline void
+xfs_csum_finalize(
+	struct xfs_mount	*mp,
+	union xfs_disk_csum	*to,
+	union xfs_csum		*csum)
+{
+	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
+		to->crc32c = cpu_to_le32(~csum->crc32c);
+	else
+		to->crc64 = cpu_to_le64(csum->crc64);
+}
+
+#endif /* _XFS_RTCSUMFILE_H */
diff --git a/fs/xfs/libxfs/xfs_rtgroup.c b/fs/xfs/libxfs/xfs_rtgroup.c
index fe7222bbe449..22ad71798b62 100644
--- a/fs/xfs/libxfs/xfs_rtgroup.c
+++ b/fs/xfs/libxfs/xfs_rtgroup.c
@@ -35,6 +35,7 @@
 #include "xfs_metadir.h"
 #include "xfs_rtrmap_btree.h"
 #include "xfs_rtrefcount_btree.h"
+#include "xfs_rtcsumfile.h"
 
 /* Find the first usable fsblock in this rtgroup. */
 static inline uint32_t
@@ -394,6 +395,15 @@ static const struct xfs_rtginode_ops xfs_rtginode_ops[XFS_RTGI_MAX] = {
 		.enabled	= xfs_has_reflink,
 		.create		= xfs_rtrefcountbt_create,
 	},
+	[XFS_RTGI_CSUM] = {
+		.name		= "csum",
+		.metafile_type	= XFS_METAFILE_RTCSUM,
+		.sick		= XFS_SICK_RG_CSUM,
+		.fmt_mask	= (1U << XFS_DINODE_FMT_EXTENTS) |
+				  (1U << XFS_DINODE_FMT_BTREE),
+		.enabled	= xfs_has_rtcsum,
+		.create		= xfs_rtcsum_create,
+	},
 };
 
 /* Return the shortname of this rtgroup inode. */
diff --git a/fs/xfs/libxfs/xfs_rtgroup.h b/fs/xfs/libxfs/xfs_rtgroup.h
index f26e324f0de3..5cca0f4dd25a 100644
--- a/fs/xfs/libxfs/xfs_rtgroup.h
+++ b/fs/xfs/libxfs/xfs_rtgroup.h
@@ -16,6 +16,7 @@ enum xfs_rtg_inodes {
 	XFS_RTGI_SUMMARY,	/* allocation summary */
 	XFS_RTGI_RMAP,		/* rmap btree inode */
 	XFS_RTGI_REFCOUNT,	/* refcount btree inode */
+	XFS_RTGI_CSUM,		/* data checksum inode */
 
 	XFS_RTGI_MAX,
 };
@@ -109,6 +110,11 @@ static inline struct xfs_inode *rtg_refcount(const struct xfs_rtgroup *rtg)
 	return rtg->rtg_inodes[XFS_RTGI_REFCOUNT];
 }
 
+static inline struct xfs_inode *rtg_csum(const struct xfs_rtgroup *rtg)
+{
+	return rtg->rtg_inodes[XFS_RTGI_CSUM];
+}
+
 /* Passive rtgroup references */
 static inline struct xfs_rtgroup *
 xfs_rtgroup_get(
diff --git a/fs/xfs/libxfs/xfs_shared.h b/fs/xfs/libxfs/xfs_shared.h
index b1e0d9bc1f7d..8a60f2c56fe7 100644
--- a/fs/xfs/libxfs/xfs_shared.h
+++ b/fs/xfs/libxfs/xfs_shared.h
@@ -40,6 +40,7 @@ extern const struct xfs_buf_ops xfs_refcountbt_buf_ops;
 extern const struct xfs_buf_ops xfs_rmapbt_buf_ops;
 extern const struct xfs_buf_ops xfs_rtbitmap_buf_ops;
 extern const struct xfs_buf_ops xfs_rtsummary_buf_ops;
+extern const struct xfs_buf_ops xfs_rtcsum_buf_ops;
 extern const struct xfs_buf_ops xfs_rtbuf_ops;
 extern const struct xfs_buf_ops xfs_rtsb_buf_ops;
 extern const struct xfs_buf_ops xfs_rtrefcountbt_buf_ops;
diff --git a/fs/xfs/xfs_buf_item_recover.c b/fs/xfs/xfs_buf_item_recover.c
index 57929f115055..ed74dcdbe483 100644
--- a/fs/xfs/xfs_buf_item_recover.c
+++ b/fs/xfs/xfs_buf_item_recover.c
@@ -414,6 +414,13 @@ xlog_recover_validate_buf_type(
 		}
 		bp->b_ops = xfs_rtblock_ops(mp, XFS_RTGI_SUMMARY);
 		break;
+	case XFS_BLFT_RTCSUM_BUF:
+		if (xfs_has_rtgroups(mp) && magic32 != XFS_RTCSUM_MAGIC) {
+			warnmsg = "Bad rtcsum magic!";
+			break;
+		}
+		bp->b_ops = &xfs_rtcsum_buf_ops;
+		break;
 #endif /* CONFIG_XFS_RT */
 	default:
 		xfs_warn(mp, "Unknown buffer type %d!",
diff --git a/fs/xfs/xfs_platform.h b/fs/xfs/xfs_platform.h
index 5d542e95fe44..c478d8935f61 100644
--- a/fs/xfs/xfs_platform.h
+++ b/fs/xfs/xfs_platform.h
@@ -16,6 +16,7 @@
 #include <linux/slab.h>
 #include <linux/vmalloc.h>
 #include <linux/crc32c.h>
+#include <linux/crc64.h>
 #include <linux/module.h>
 #include <linux/mutex.h>
 #include <linux/file.h>
diff --git a/fs/xfs/xfs_rtalloc.c b/fs/xfs/xfs_rtalloc.c
index 78a1c066c7eb..e2b6113772b6 100644
--- a/fs/xfs/xfs_rtalloc.c
+++ b/fs/xfs/xfs_rtalloc.c
@@ -32,6 +32,7 @@
 #include "xfs_error.h"
 #include "xfs_trace.h"
 #include "xfs_rtrefcount_btree.h"
+#include "xfs_rtcsumfile.h"
 #include "xfs_reflink.h"
 #include "xfs_zone_alloc.h"
 
@@ -898,6 +899,12 @@ xfs_growfs_rt_zoned(
 	xfs_rtbxlen_t		freed_rtx;
 	int			error;
 
+	if (xfs_has_rtcsum(mp)) {
+		error = xfs_rtcsum_alloc_blocks(rtg);
+		if (error)
+			return error;
+	}
+
 	/*
 	 * Calculate new sb and mount fields for this round.  Also ensure the
 	 * rtg_extents value is uptodate as the rtbitmap code relies on it.
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (8 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-24 22:30   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
                   ` (11 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

The data checksum is logged in its own transaction, and only logs
transactions buffers.  While the maximum size of a checksummed data write
is the same as that of a single checksum buffer, the Zone Append based
zoned write path can't guarantee alignment, so it might be spread over
up to three buffers.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_trans_resv.c | 21 +++++++++++++++++++++
 fs/xfs/libxfs/xfs_trans_resv.h |  4 ++++
 2 files changed, 25 insertions(+)

diff --git a/fs/xfs/libxfs/xfs_trans_resv.c b/fs/xfs/libxfs/xfs_trans_resv.c
index 5382ece51812..8854253afba0 100644
--- a/fs/xfs/libxfs/xfs_trans_resv.c
+++ b/fs/xfs/libxfs/xfs_trans_resv.c
@@ -11,6 +11,7 @@
 #include "xfs_log_format.h"
 #include "xfs_trans_resv.h"
 #include "xfs_mount.h"
+#include "xfs_rtcsumfile.h"
 #include "xfs_da_format.h"
 #include "xfs_da_btree.h"
 #include "xfs_inode.h"
@@ -1233,6 +1234,22 @@ xfs_calc_qm_dqalloc_reservation_minlogsize(
 	return xfs_calc_qm_dqalloc_reservation(mp, true);
 }
 
+/*
+ * Log data checksums for a write.
+ *
+ * Must cover a checksum for each FSB of data written, and the checksums can
+ * span the FSB-sized checksum buffers at both ends.
+ */
+unsigned int
+xfs_calc_csum_reservation(
+	struct xfs_mount	*mp,
+	unsigned int		csum_len)
+{
+	return xfs_calc_buf_res(
+			howmany(csum_len, xfs_rtcsum_payload_size(mp)) + 1,
+			mp->m_rtcsum_bsize);
+}
+
 /*
  * Syncing the incore super block changes to disk.
  *     the super block to reflect the changes: sector size
@@ -1354,6 +1371,10 @@ xfs_trans_resv_calc(
 
 	xfs_calc_namespace_reservations(mp, resp);
 
+	resp->tr_csum.tr_logres =
+		xfs_calc_csum_reservation(mp, XFS_RTCSUM_MAX_WRITE);
+	resp->tr_csum.tr_logcount = XFS_DEFAULT_LOG_COUNT;
+
 	/*
 	 * The following transactions are logged in logical format with
 	 * a default log count.
diff --git a/fs/xfs/libxfs/xfs_trans_resv.h b/fs/xfs/libxfs/xfs_trans_resv.h
index 1804e821f382..127db4da31c1 100644
--- a/fs/xfs/libxfs/xfs_trans_resv.h
+++ b/fs/xfs/libxfs/xfs_trans_resv.h
@@ -49,6 +49,7 @@ struct xfs_trans_resv {
 	struct xfs_trans_res	tr_sb;		/* modify superblock */
 	struct xfs_trans_res	tr_fsyncts;	/* update timestamps on fsync */
 	struct xfs_trans_res	tr_atomic_ioend; /* untorn write completion */
+	struct xfs_trans_res	tr_csum;	/* data checksums in metafile */
 };
 
 /* shorthand way of accessing reservation structure */
@@ -122,6 +123,9 @@ unsigned int xfs_calc_itruncate_reservation_minlogsize(struct xfs_mount *mp);
 unsigned int xfs_calc_write_reservation_minlogsize(struct xfs_mount *mp);
 unsigned int xfs_calc_qm_dqalloc_reservation_minlogsize(struct xfs_mount *mp);
 
+unsigned int xfs_calc_csum_reservation(struct xfs_mount *mp,
+		unsigned int csum_len);
+
 xfs_extlen_t xfs_calc_max_atomic_write_fsblocks(struct xfs_mount *mp);
 xfs_extlen_t xfs_calc_atomic_write_log_geometry(struct xfs_mount *mp,
 		xfs_extlen_t blockcount, unsigned int *new_logres);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 11/21] xfs: core RT data checksum support
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (9 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-25 23:20   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
                   ` (10 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Support reading and writing of data checksum buffers, and generating
and verifying the checksum.

All data checksum buffers for an open zone are pre-allocated at zone open
time, so that we never have to read in a partially written buffer as part
of a data write, which would otherwise impose very expensive seeks and
stall writes.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/Kconfig          |   1 +
 fs/xfs/Makefile         |   1 +
 fs/xfs/xfs_inode.h      |   5 +
 fs/xfs/xfs_rtcsum.c     | 317 ++++++++++++++++++++++++++++++++++++++++
 fs/xfs/xfs_rtcsum.h     |  23 +++
 fs/xfs/xfs_super.c      |  14 ++
 fs/xfs/xfs_sysfs.c      |   2 +
 fs/xfs/xfs_zone_alloc.c |  32 +++-
 fs/xfs/xfs_zone_priv.h  |   8 +
 9 files changed, 399 insertions(+), 4 deletions(-)
 create mode 100644 fs/xfs/xfs_rtcsum.c
 create mode 100644 fs/xfs/xfs_rtcsum.h

diff --git a/fs/xfs/Kconfig b/fs/xfs/Kconfig
index b99da294e9a3..424a04b507a6 100644
--- a/fs/xfs/Kconfig
+++ b/fs/xfs/Kconfig
@@ -106,6 +106,7 @@ config XFS_RT
 	bool "XFS Realtime subvolume support"
 	depends on XFS_FS
 	default BLK_DEV_ZONED
+	select CRC64
 	help
 	  If you say Y here you will be able to mount and use XFS filesystems
 	  which contain a realtime subvolume.  The realtime subvolume is a
diff --git a/fs/xfs/Makefile b/fs/xfs/Makefile
index 79ea4136fbba..0d57bf0701ec 100644
--- a/fs/xfs/Makefile
+++ b/fs/xfs/Makefile
@@ -142,6 +142,7 @@ xfs-$(CONFIG_XFS_QUOTA)		+= xfs_dquot.o \
 
 # xfs_rtbitmap is shared with libxfs
 xfs-$(CONFIG_XFS_RT)		+= xfs_rtalloc.o \
+				   xfs_rtcsum.o \
 				   xfs_zone_alloc.o \
 				   xfs_zone_gc.o \
 				   xfs_zone_info.o \
diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h
index 34c1038ebfcd..9ed2fbfe86ff 100644
--- a/fs/xfs/xfs_inode.h
+++ b/fs/xfs/xfs_inode.h
@@ -376,6 +376,11 @@ static inline bool xfs_inode_can_sw_atomic_write(const struct xfs_inode *ip)
 	return xfs_can_sw_atomic_write(ip->i_mount);
 }
 
+static inline bool xfs_is_rtcsum_inode(const struct xfs_inode *ip)
+{
+	return xfs_has_rtcsum(ip->i_mount) && XFS_IS_REALTIME_INODE(ip);
+}
+
 /*
  * In-core inode flags.
  */
diff --git a/fs/xfs/xfs_rtcsum.c b/fs/xfs/xfs_rtcsum.c
new file mode 100644
index 000000000000..7cdc5a5029eb
--- /dev/null
+++ b/fs/xfs/xfs_rtcsum.c
@@ -0,0 +1,317 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Christoph Hellwig.
+ */
+#include "xfs_platform.h"
+#include "xfs_fs.h"
+#include "xfs_format.h"
+#include "xfs_log_format.h"
+#include "xfs_shared.h"
+#include "xfs_trans_resv.h"
+#include "xfs_bit.h"
+#include "xfs_mount.h"
+#include "xfs_inode.h"
+#include "xfs_bmap.h"
+#include "xfs_rtgroup.h"
+#include "xfs_rtbitmap.h"
+#include "xfs_bmap_btree.h"
+#include "xfs_trans.h"
+#include "xfs_buf_item.h"
+#include "xfs_trans_space.h"
+#include "xfs_error.h"
+#include "xfs_health.h"
+#include "xfs_rtcsum.h"
+#include "xfs_zone_priv.h"
+#include <linux/iomap.h>
+
+static_assert(IOMAP_CSUM_MAX_SIZE <= XFS_RTCSUM_MAX_WRITE);
+
+static const char *xfs_data_csum_names[XFS_CSUM_TYPE_MAX] = {
+	[XFS_CSUM_TYPE_CRC32C]	= "crc32c",
+	[XFS_CSUM_TYPE_CRC64]	= "crc64",
+};
+
+int
+xfs_csum_verify(
+	struct xfs_mount	*mp,
+	struct bio		*bio,
+	struct bvec_iter	*iter,
+	void			*csum_buf,
+	xfs_fsblock_t		bno,
+	bool			verbose)
+{
+	unsigned int		bsize = mp->m_sb.sb_blocksize;
+	unsigned int		csum_size = 1u << mp->m_rtcsum_shift;
+	unsigned int		offset = 0;
+	union xfs_csum		csum;
+	union xfs_disk_csum	dsum;
+
+	do {
+		struct bio_vec	bv = mp_bvec_iter_bvec(bio->bi_io_vec, *iter);
+
+		if (offset == 0)
+			xfs_csum_seed(mp, &csum);
+		bv.bv_len = min(bv.bv_len, bsize - offset);
+		xfs_csum_gen(mp, bvec_virt(&bv), bv.bv_len, &csum);
+		offset += bv.bv_len;
+		if (offset == bsize) {
+			xfs_csum_finalize(mp, &dsum, &csum);
+			if (unlikely(memcmp(&dsum, csum_buf, csum_size) != 0))
+				goto mismatch;
+			bno++;
+			csum_buf += csum_size;
+			offset = 0;
+		}
+		bio_advance_iter_single(bio, iter, bv.bv_len);
+	} while (iter->bi_size);
+
+	return 0;
+
+mismatch:
+	if (verbose) {
+		xfs_warn_ratelimited(mp,
+"data csum mismatch for rtblock 0x%llx: 0x%*phN (expected 0x%*phN)",
+			bno, csum_size, &csum, csum_size, csum_buf);
+	}
+	return -EIO;
+}
+
+void
+xfs_csum_generate(
+	struct xfs_mount	*mp,
+	struct bio		*bio,
+	void			*csum_buf)
+{
+	struct bvec_iter	iter = bio->bi_iter;
+	unsigned int		bsize = mp->m_sb.sb_blocksize;
+	unsigned int		csum_size = 1u << mp->m_rtcsum_shift;
+	unsigned int		offset = 0;
+	union xfs_csum		csum;
+
+	do {
+		struct bio_vec	bv = mp_bvec_iter_bvec(bio->bi_io_vec, iter);
+
+		if (offset == 0)
+			xfs_csum_seed(mp, &csum);
+		bv.bv_len = min(bv.bv_len, bsize - offset);
+		xfs_csum_gen(mp, bvec_virt(&bv), bv.bv_len, &csum);
+		offset += bv.bv_len;
+		if (offset == bsize) {
+			xfs_csum_finalize(mp, csum_buf, &csum);
+			csum_buf += csum_size;
+			offset = 0;
+		}
+		bio_advance_iter_single(bio, &iter, bv.bv_len);
+	} while (iter.bi_size);
+}
+
+/*
+ * Read the checksum buffer for @rtg/@csum_off and return it unlocked.
+ *
+ * We don't need to lock read access to the buffer because checksums will not
+ * change until the @rtg is reset.
+ */
+int
+xfs_rtcsum_read_async(
+	struct xfs_mount	*mp,
+	xfs_rtblock_t		fsbno,
+	struct xfs_buf		**bpp)
+{
+	xfs_daddr_t		csum_daddr;
+	struct xfs_rtgroup	*rtg;
+	int			error;
+
+	rtg = xfs_rtgroup_get(mp, xfs_rtb_to_rgno(mp, fsbno));
+	if (!rtg)
+		return -EFSCORRUPTED;
+	error = xfs_rtcsum_bmap(rtg, xfs_rtb_to_rgbno(mp, fsbno), &csum_daddr);
+	if (!error)
+		error = xfs_buf_read_async(mp->m_ddev_targp, csum_daddr,
+				BTOBB(mp->m_rtcsum_bsize), &xfs_rtcsum_buf_ops,
+				bpp);
+	xfs_rtgroup_put(rtg);
+	return error;
+}
+
+/*
+ * When opening a zone for writing, do a speculative buf_get for each csum
+ * buffer.  This ensures we usually have a buffer in-memory when we actually
+ * start writing to it.
+ *
+ * Without this we'd have to read the buffer from disk, as we don't know if
+ * anyone has already written to it by the time we get to the buffer due to
+ * completion reordering.
+ *
+ * When reopening a partially written zone at mount time, just read ahead
+ * the entire csums for the zone.
+ */
+int
+xfs_rtcsum_open_zone(
+	struct xfs_open_zone	*oz)
+{
+	struct xfs_rtgroup	*rtg = oz->oz_rtg;
+	struct xfs_mount	*mp = rtg_mount(rtg);
+	xfs_rgblock_t		rgbno = 0;
+	int			i, error;
+
+	for (i = 0; i < oz->oz_nr_csum_bufs; i++) {
+		xfs_daddr_t	csum_daddr;
+		struct xfs_buf	*bp;
+
+		error = xfs_rtcsum_bmap(rtg, rgbno, &csum_daddr);
+		if (error)
+			goto out_error;
+
+		if (oz->oz_allocated) {
+			error = xfs_buf_read_async(mp->m_ddev_targp, csum_daddr,
+					BTOBB(mp->m_rtcsum_bsize),
+					&xfs_rtcsum_buf_ops, &bp);
+		} else {
+			error = xfs_buf_get(mp->m_ddev_targp, csum_daddr,
+					BTOBB(mp->m_rtcsum_bsize), &bp);
+			if (!error) {
+				bp->b_flags = XBF_DONE;
+				xfs_rtfile_initialize_buf(rtg, XFS_RTGI_CSUM,
+						bp, NULL);
+			}
+			xfs_buf_unlock(bp);
+		}
+		if (error)
+			goto out_error;
+		oz->oz_csum_bufs[i] = bp;
+		rgbno += (xfs_rtcsum_payload_size(mp) /
+			  (1u << mp->m_rtcsum_shift));
+	}
+
+	return 0;
+
+out_error:
+	while (--i >= 0)
+		xfs_buf_rele(oz->oz_csum_bufs[i]);
+	return error;
+}
+
+void
+xfs_rtcsum_free_zone(
+	struct xfs_open_zone	*oz)
+{
+	unsigned int		i;
+
+	for (i = 0; i < oz->oz_nr_csum_bufs; i++)
+		xfs_buf_rele(oz->oz_csum_bufs[i]);
+}
+
+static int
+xfs_rtcsum_log_one(
+	struct xfs_open_zone	*oz,
+	xfs_rgblock_t		rgbno,
+	const void		*csum,
+	xfs_extlen_t		nfsb,
+	struct xfs_trans	*tp,
+	xfs_extlen_t		*nr)
+{
+	struct xfs_rtgroup	*rtg = oz->oz_rtg;
+	struct xfs_mount	*mp = rtg_mount(rtg);
+	unsigned int		boff = xfs_rgb_to_rtcsumoff(mp, rgbno);
+	struct xfs_buf		*bp;
+	unsigned int		len;
+	int			error;
+
+	bp = oz->oz_csum_bufs[xfs_rgb_to_rtcsumblock(mp, rgbno)];
+	error = xfs_buf_read_async_wait(bp);
+	if (error) {
+		if (xfs_metadata_is_sick(error))
+			xfs_rtginode_mark_sick(rtg, XFS_RTGI_CSUM);
+		return error;
+	}
+
+	xfs_buf_hold(bp);
+	xfs_buf_lock(bp);
+	xfs_trans_bjoin(tp, bp);
+	xfs_trans_buf_set_type(tp, bp, XFS_BLFT_RTCSUM_BUF);
+
+	*nr = min(nfsb,
+		xfs_rtcsum_len_to_extlen(mp, mp->m_rtcsum_bsize - boff));
+	len = xfs_extlen_to_rtcsum_len(mp, *nr);
+	memcpy(bp->b_addr + boff, csum, len);
+
+	/*
+	 * Avoid realloc churn as we're filling the buffer (almost) sequentially
+	 * and rather fast.
+	 */
+	bp->b_log_item->bli_flags |= XFS_BLI_PREALLOC;
+	xfs_trans_log_buf(tp, bp, boff, boff + len - 1);
+	return 0;
+}
+
+int
+xfs_rtcsum_log(
+	struct xfs_open_zone	*oz,
+	xfs_daddr_t		daddr,
+	unsigned int		data_len,
+	const void		*csum)
+{
+	struct xfs_mount	*mp = rtg_mount(oz->oz_rtg);
+	xfs_rgblock_t		rgbno = xfs_daddr_to_rgbno(mp, daddr);
+	xfs_extlen_t		nfsb = XFS_B_TO_FSB(mp, data_len);
+	struct xfs_trans_res	tres = M_RES(mp)->tr_csum;
+	struct xfs_trans	*tp;
+	int			error;
+
+	tres.tr_logres = xfs_calc_csum_reservation(mp,
+				xfs_extlen_to_rtcsum_len(mp, nfsb));
+	ASSERT(tres.tr_logres <= M_RES(mp)->tr_csum.tr_logres);
+
+	error = xfs_trans_alloc(mp, &tres, 0, 0, 0, &tp);
+	if (error)
+		return error;
+
+	do {
+		xfs_extlen_t nr;
+
+		error = xfs_rtcsum_log_one(oz, rgbno, csum, nfsb, tp, &nr);
+		if (error)
+			goto out_trans_cancel;
+		rgbno += nr;
+		nfsb -= nr;
+		csum += xfs_extlen_to_rtcsum_len(mp, nr);
+	} while (nfsb);
+
+	return xfs_trans_commit(tp);
+
+out_trans_cancel:
+	xfs_trans_cancel(tp);
+	return error;
+}
+
+int
+xfs_rtcsum_mount(
+	struct xfs_mount	*mp)
+{
+	/*
+	 * Supporting highmem would require less efficient loops for the
+	 * checksumming.  We could do that with a separate code path, but
+	 * there isn't much of a point in the extra test coverage required just
+	 * for obsolete 32-bit platforms.
+	 */
+	if (IS_ENABLED(CONFIG_HIGHMEM)) {
+		xfs_alert(mp,
+"CONFIG_HIGHMEM not supported with data checksums.");
+		return -EINVAL;
+	}
+
+	/*
+	 * We can't support data checksums on DAX because the direct access
+	 * doesn't allow computing and verifying the checksums.
+	 */
+	if (xfs_has_dax_always(mp)) {
+		xfs_warn(mp, "dax=always not supported with data checksums");
+		return -EINVAL;
+	}
+	mp->m_features |= XFS_FEAT_DAX_NEVER;
+
+	xfs_info(mp, "using %s for RT device data checksums",
+		xfs_data_csum_names[mp->m_sb.sb_rtcsum_type]);
+
+	return 0;
+}
diff --git a/fs/xfs/xfs_rtcsum.h b/fs/xfs/xfs_rtcsum.h
new file mode 100644
index 000000000000..bd16f5a99451
--- /dev/null
+++ b/fs/xfs/xfs_rtcsum.h
@@ -0,0 +1,23 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _XFS_RTCSUM_H
+#define _XFS_RTCSUM_H
+
+#include "xfs_rtcsumfile.h"
+
+void xfs_csum_generate(struct xfs_mount *mp, struct bio *bio,
+		void *csum_buf);
+int xfs_csum_verify(struct xfs_mount *mp, struct bio *bio,
+		struct bvec_iter *iter, void *csum_buf,
+		xfs_fsblock_t bno, bool verbose);
+
+int xfs_rtcsum_read_async(struct xfs_mount *mp, xfs_rtblock_t fsbno,
+		struct xfs_buf **bpp);
+int xfs_rtcsum_log(struct xfs_open_zone *oz, xfs_daddr_t daddr,
+		unsigned int data_len, const void *csum);
+
+int xfs_rtcsum_open_zone(struct xfs_open_zone *oz);
+void xfs_rtcsum_free_zone(struct xfs_open_zone *oz);
+
+int xfs_rtcsum_mount(struct xfs_mount *mp);
+
+#endif /* _XFS_RTCSUM_H */
diff --git a/fs/xfs/xfs_super.c b/fs/xfs/xfs_super.c
index fce1d2905c94..5a6769239f35 100644
--- a/fs/xfs/xfs_super.c
+++ b/fs/xfs/xfs_super.c
@@ -50,6 +50,7 @@
 #include "xfs_rtalloc.h"
 #include "xfs_zone_alloc.h"
 #include "xfs_healthmon.h"
+#include "xfs_rtcsum.h"
 #include "scrub/stats.h"
 #include "scrub/rcbag_btree.h"
 
@@ -1953,6 +1954,19 @@ xfs_fs_fill_super(
 		}
 	}
 
+	if (xfs_has_rtcsum(mp)) {
+		if (!IS_ENABLED(CONFIG_XFS_RT) || !xfs_has_zoned(mp)) {
+			xfs_alert(mp,
+	"zoned allocator required for RT data checksums");
+			error = -EINVAL;
+			goto out_filestream_unmount;
+		}
+
+		error = xfs_rtcsum_mount(mp);
+		if (error)
+			goto out_filestream_unmount;
+	}
+
 	if (xfs_has_reflink(mp)) {
 		if (xfs_has_realtime(mp) &&
 		    !xfs_reflink_supports_rextsize(mp, mp->m_sb.sb_rextsize)) {
diff --git a/fs/xfs/xfs_sysfs.c b/fs/xfs/xfs_sysfs.c
index e77917ac179d..c1ae9b7c4cc1 100644
--- a/fs/xfs/xfs_sysfs.c
+++ b/fs/xfs/xfs_sysfs.c
@@ -401,6 +401,8 @@ static bool
 xfs_has_read_bounce(
 	struct xfs_mount	*mp)
 {
+	if (xfs_has_rtcsum(mp))
+		return true;
 	if (bdev_has_integrity_csum(mp->m_ddev_targp->bt_bdev))
 		return true;
 	if (mp->m_rtdev_targp &&
diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c
index bbdff9abe2b2..9d9a713684b9 100644
--- a/fs/xfs/xfs_zone_alloc.c
+++ b/fs/xfs/xfs_zone_alloc.c
@@ -26,6 +26,7 @@
 #include "xfs_zones.h"
 #include "xfs_trace.h"
 #include "xfs_mru_cache.h"
+#include "xfs_rtcsum.h"
 #include <linux/bio-integrity.h>
 
 static void
@@ -42,8 +43,10 @@ void
 xfs_open_zone_put(
 	struct xfs_open_zone	*oz)
 {
-	if (atomic_dec_and_test(&oz->oz_ref))
+	if (atomic_dec_and_test(&oz->oz_ref)) {
+		xfs_rtcsum_free_zone(oz);
 		call_rcu(&oz->oz_rcu, xfs_open_zone_free_rcu);
+	}
 }
 
 static inline uint32_t
@@ -416,9 +419,12 @@ xfs_init_open_zone(
 	enum rw_hint		write_hint,
 	bool			is_gc)
 {
+	unsigned int		nr_csum_bufs = xfs_rtcsum_bufs_per_rtg(rtg);
 	struct xfs_open_zone	*oz;
+	int			error;
 
-	oz = kzalloc_obj(*oz, GFP_NOFS | __GFP_NOFAIL);
+	oz = kzalloc_flex(*oz, oz_csum_bufs, nr_csum_bufs,
+			GFP_NOFS | __GFP_NOFAIL);
 	spin_lock_init(&oz->oz_alloc_lock);
 	atomic_set(&oz->oz_ref, 1);
 	oz->oz_rtg = rtg;
@@ -426,6 +432,14 @@ xfs_init_open_zone(
 	oz->oz_written = write_pointer;
 	oz->oz_write_hint = write_hint;
 	oz->oz_is_gc = is_gc;
+	oz->oz_nr_csum_bufs = nr_csum_bufs;
+	if (nr_csum_bufs) {
+		error = xfs_rtcsum_open_zone(oz);
+		if (error) {
+			kfree(oz);
+			return ERR_PTR(error);
+		}
+	}
 
 	/*
 	 * All dereferences of rtg->rtg_open_zone hold the ILOCK for the rmap
@@ -449,6 +463,7 @@ xfs_open_zone(
 {
 	struct xfs_zone_info	*zi = mp->m_zone_info;
 	XA_STATE		(xas, &mp->m_groups[XG_TYPE_RTG].xa, 0);
+	struct xfs_open_zone	*oz;
 	struct xfs_group	*xg;
 
 	/*
@@ -469,7 +484,14 @@ xfs_open_zone(
 	xas_unlock(&xas);
 
 	set_current_state(TASK_RUNNING);
-	return xfs_init_open_zone(to_rtg(xg), 0, write_hint, is_gc);
+	oz = xfs_init_open_zone(to_rtg(xg), 0, write_hint, is_gc);
+	if (IS_ERR(oz)) {
+		xfs_rtgroup_rele(to_rtg(xg));
+		xfs_force_shutdown(mp, SHUTDOWN_CORRUPT_ONDISK);
+		return NULL;
+	}
+
+	return oz;
 }
 
 static struct xfs_open_zone *
@@ -1144,9 +1166,11 @@ xfs_init_zone(
 		/* zone is open */
 		struct xfs_open_zone *oz;
 
-		atomic_inc(&rtg_group(rtg)->xg_active_ref);
 		oz = xfs_init_open_zone(rtg, write_pointer, WRITE_LIFE_NOT_SET,
 				false);
+		if (IS_ERR(oz))
+			return PTR_ERR(oz);
+		atomic_inc(&rtg_group(rtg)->xg_active_ref);
 		list_add_tail(&oz->oz_entry, &zi->zi_open_zones);
 		zi->zi_nr_open_zones++;
 
diff --git a/fs/xfs/xfs_zone_priv.h b/fs/xfs/xfs_zone_priv.h
index fcb57506d8e6..322a031e6ba0 100644
--- a/fs/xfs/xfs_zone_priv.h
+++ b/fs/xfs/xfs_zone_priv.h
@@ -42,6 +42,14 @@ struct xfs_open_zone {
 	struct xfs_rtgroup	*oz_rtg;
 
 	struct rcu_head		oz_rcu;
+
+	/*
+	 * Checksum buffers.  All checksum buffers for an open zone are pinned
+	 * into memory so that we never need to read them in to log checksums
+	 * for a write.
+	 */
+	unsigned int		oz_nr_csum_bufs;
+	struct xfs_buf		*oz_csum_bufs[] __counted_by(oz_nr_csum_bufs);
 };
 
 /*
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 12/21] xfs: data checksums require stable writes
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (10 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-25 23:21   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
                   ` (9 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Call mapping_set_stable_writes based on the data checksum flag
so that the page cache doesn't change data in-flight as that could
corrupt the checksum.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_inode.h | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h
index 9ed2fbfe86ff..72582ea9afcc 100644
--- a/fs/xfs/xfs_inode.h
+++ b/fs/xfs/xfs_inode.h
@@ -619,7 +619,8 @@ int	xfs_break_layouts(struct inode *inode, uint *iolock,
 
 static inline void xfs_update_stable_writes(struct xfs_inode *ip)
 {
-	if (bdev_stable_writes(xfs_inode_buftarg(ip)->bt_bdev))
+	if (xfs_is_rtcsum_inode(ip) ||
+	    bdev_stable_writes(xfs_inode_buftarg(ip)->bt_bdev))
 		mapping_set_stable_writes(VFS_I(ip)->i_mapping);
 	else
 		mapping_clear_stable_writes(VFS_I(ip)->i_mapping);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 13/21] xfs: require file system block size alignment when using data checksums
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (11 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-25 23:24   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
                   ` (8 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

The checksums cover a whole block, so we can't read or update parts of a
block.  Report the requirement and enforce it for direct I/O reads.
Direct I/O writes already require file system block size alignment when
using the zoned allocator, and buffered I/O never does sub-block I/O.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_file.c  | 14 +++++++++++++-
 fs/xfs/xfs_ioend.c | 13 +++++++++++--
 fs/xfs/xfs_iops.c  | 10 +++++++++-
 3 files changed, 33 insertions(+), 4 deletions(-)

diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
index 6f25879b6510..5b25f33527c0 100644
--- a/fs/xfs/xfs_file.c
+++ b/fs/xfs/xfs_file.c
@@ -29,6 +29,7 @@
 #include "xfs_zone_alloc.h"
 #include "xfs_error.h"
 #include "xfs_errortag.h"
+#include "xfs_rtcsum.h"
 
 #include <linux/dax.h>
 #include <linux/falloc.h>
@@ -270,9 +271,20 @@ xfs_file_dio_read(
 	if (ret)
 		return ret;
 	if (mapping_stable_writes(iocb->ki_filp->f_mapping)) {
+		unsigned int		dio_flags = 0;
+
+		/*
+		 * Each checksums covers a whole file system block, and thus
+		 * sub-fsblock reads are not supported for file systems using
+		 * data checksums.
+		 */
+		if (xfs_is_rtcsum_inode(ip))
+			dio_flags |= IOMAP_DIO_FSBLOCK_ALIGNED;
 		ret = iomap_dio_rw(iocb, to, &xfs_read_iomap_ops,
-				&xfs_dio_read_bounce_ops, 0, NULL, 0);
+				&xfs_dio_read_bounce_ops, dio_flags, NULL, 0);
 	} else {
+		ASSERT(!xfs_is_rtcsum_inode(ip));
+
 		ret = iomap_dio_read_simple(iocb, to, xfs_read_iomap_begin);
 		if (ret == -ENOTBLK)
 			ret = iomap_dio_rw(iocb, to, &xfs_read_iomap_ops, NULL,
diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c
index f0e01ac34de8..54bd0995ac29 100644
--- a/fs/xfs/xfs_ioend.c
+++ b/fs/xfs/xfs_ioend.c
@@ -44,6 +44,15 @@ xfs_bounce_submit_ioend(
 	submit_bio(&ioend->io_bio);
 }
 
+static unsigned int
+xfs_read_bounce_minsize(
+	struct iomap_ioend	*ioend)
+{
+	if (xfs_is_rtcsum_inode(XFS_I(ioend->io_inode)))
+		return i_blocksize(ioend->io_inode);
+	return bdev_logical_block_size(ioend->io_bio.bi_bdev);
+}
+
 static void
 xfs_end_bio_bounced(
 	struct bio		*bio)
@@ -86,7 +95,7 @@ xfs_read_bounce_and_resubmit(
 		.bi_offset	= ioend->io_bvec_offset,
 	};
 	bio->bi_end_io = xfs_end_bio_bounced;
-	iomap_bounce_read(ioend, bdev_logical_block_size(bio->bi_bdev),
+	iomap_bounce_read(ioend, xfs_read_bounce_minsize(ioend),
 			xfs_bounce_submit_ioend);
 	memalloc_nofs_restore(nofs_flag);
 }
@@ -134,7 +143,7 @@ xfs_ioend_submit_read(
 	ioend = iomap_init_ioend(inode, bio, file_offset, ioend_flags);
 	if ((ioend_flags & IOMAP_IOEND_DIRECT) &&
 	    READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_ALWAYS) {
-		iomap_bounce_read(ioend, bdev_logical_block_size(bio->bi_bdev),
+		iomap_bounce_read(ioend, xfs_read_bounce_minsize(ioend),
 				xfs_bounce_submit_ioend);
 		return;
 	}
diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
index d1306e723899..a5f01e2e3a67 100644
--- a/fs/xfs/xfs_iops.c
+++ b/fs/xfs/xfs_iops.c
@@ -581,6 +581,15 @@ xfs_report_dioalign(
 	stat->result_mask |= STATX_DIOALIGN | STATX_DIO_READ_ALIGN;
 	stat->dio_mem_align = bdev_dma_alignment(bdev) + 1;
 
+	/*
+	 * Each checksums covers a whole file system block, and thus sub-fsblock
+	 * reads are not supported for file systems using data checksums.
+	 */
+	if (xfs_is_rtcsum_inode(ip))
+		stat->dio_read_offset_align = xfs_inode_alloc_unitsize(ip);
+	else
+		stat->dio_read_offset_align = bdev_logical_block_size(bdev);
+
 	/*
 	 * For COW inodes, we can only perform out of place writes of entire
 	 * allocation units (blocks or RT extents).
@@ -591,7 +600,6 @@ xfs_report_dioalign(
 	 * alignment in dio_offset_align, and the smaller read alignment in
 	 * dio_read_offset_align.
 	 */
-	stat->dio_read_offset_align = bdev_logical_block_size(bdev);
 	if (xfs_is_cow_inode(ip))
 		stat->dio_offset_align = xfs_inode_alloc_unitsize(ip);
 	else
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 14/21] xfs: add support for reading with data checksums
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (12 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  0:42   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
                   ` (7 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

All reads from files with data checksums have the returned iomaps for
data blocks limited to be inside a single RT csum file block, so that
each data read only needs to deal with a single checksum buffer.

All reads on checksummed files need to use ioends so that the checksum
can be verified from process context.  The ioend submission path looks
up the checksum buffer and kicks of an asynchronous read of it.  The
completion path waits for the buffer if needed and verifies the checksum.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_aops.c  |   4 +-
 fs/xfs/xfs_ioend.c | 108 +++++++++++++++++++++++++++++++++++++++------
 fs/xfs/xfs_iomap.c |  23 ++++++++--
 3 files changed, 114 insertions(+), 21 deletions(-)

diff --git a/fs/xfs/xfs_aops.c b/fs/xfs/xfs_aops.c
index c30e688cfc9f..931795316de4 100644
--- a/fs/xfs/xfs_aops.c
+++ b/fs/xfs/xfs_aops.c
@@ -599,9 +599,7 @@ static inline const struct iomap_read_ops *
 xfs_get_iomap_read_ops(
 	const struct address_space	*mapping)
 {
-	struct xfs_inode		*ip = XFS_I(mapping->host);
-
-	if (bdev_has_integrity_csum(xfs_inode_buftarg(ip)->bt_bdev))
+	if (mapping_stable_writes(mapping))
 		return &xfs_iomap_read_ops;
 	return &iomap_bio_read_ops;
 }
diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c
index 54bd0995ac29..7570a1b915c0 100644
--- a/fs/xfs/xfs_ioend.c
+++ b/fs/xfs/xfs_ioend.c
@@ -14,12 +14,72 @@
 #include "xfs_trace.h"
 #include "xfs_bmap_util.h"
 #include "xfs_reflink.h"
+#include "xfs_rtcsum.h"
 #include "xfs_zone_alloc.h"
 #include "xfs_ioend.h"
 #include "xfs_error.h"
 #include "xfs_errortag.h"
 #include <linux/bio-integrity.h>
 
+static bool
+xfs_rtcsum_prepare_read(
+	struct iomap_ioend	*ioend)
+{
+	struct xfs_inode	*ip = XFS_I(ioend->io_inode);
+	struct xfs_mount	*mp = ip->i_mount;
+	struct xfs_buf		*bp;
+	int			error;
+
+	error = -EIO;
+	if (WARN_ON_ONCE(ioend->io_bio.bi_iter.bi_idx))
+		goto fail;
+
+	error = xfs_rtcsum_read_async(mp,
+			xfs_daddr_to_rtb(mp, ioend->io_sector), &bp);
+	if (error)
+		goto fail;
+	ioend->io_private = bp;
+	return true;
+
+fail:
+	ioend->io_bio.bi_status = errno_to_blk_status(error);
+	bio_endio(&ioend->io_bio);
+	return false;
+}
+
+static int
+xfs_rtcsum_verify_ioend(
+	struct iomap_ioend	*ioend,
+	int			error)
+{
+	struct xfs_inode	*ip = XFS_I(ioend->io_inode);
+	struct xfs_mount	*mp = ip->i_mount;
+	xfs_rtblock_t		bno = xfs_daddr_to_rtb(mp, ioend->io_sector);
+	unsigned int		bsize = mp->m_sb.sb_blocksize;
+	struct xfs_buf		*bp = ioend->io_private;
+	struct bvec_iter	iter = {
+		.bi_size	= roundup(ioend->io_size, bsize),
+		.bi_offset	= ioend->io_bvec_offset,
+	};
+
+	/* No bp for early xfs_rtcsum_prepare_read failures. */
+	if (!bp)
+		return error;
+
+	if (error)
+		goto out_rele;
+	error = xfs_buf_read_async_wait(bp);
+	if (error)
+		goto out_rele;
+
+	error = xfs_csum_verify(mp, &ioend->io_bio, &iter,
+				bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno,
+				true);
+out_rele:
+	xfs_buf_rele(bp);
+	return error;
+}
+
 static void
 xfs_dio_bounce_end_io(
 	struct bio		*bio)
@@ -30,6 +90,9 @@ xfs_dio_bounce_end_io(
 
 	if ((ioend->io_flags & IOMAP_IOEND_INTEGRITY) && !bio->bi_status)
 		error = iomap_ioend_integrity_verify(ioend);
+	if (xfs_is_rtcsum_inode(XFS_I(ioend->io_inode)))
+		error = xfs_rtcsum_verify_ioend(ioend, error);
+
 	iomap_bounce_read_end_io(ioend, orig_bio, error);
 }
 
@@ -39,6 +102,9 @@ xfs_bounce_submit_ioend(
 {
 	if (ioend->io_flags & IOMAP_IOEND_INTEGRITY)
 		fs_bio_integrity_alloc(&ioend->io_bio);
+	if (xfs_is_rtcsum_inode(XFS_I(ioend->io_inode)) &&
+	    !xfs_rtcsum_prepare_read(ioend))
+		return;
 	ioend->io_bio.bi_end_io = xfs_dio_bounce_end_io;
 	bio_set_flag(&ioend->io_bio, BIO_COMPLETE_IN_TASK);
 	submit_bio(&ioend->io_bio);
@@ -108,25 +174,36 @@ xfs_end_io_read(
 	struct xfs_inode	*ip = XFS_I(ioend->io_inode);
 	struct xfs_mount	*mp = ip->i_mount;
 	int			error = blk_status_to_errno(bio->bi_status);
+	bool			is_csum_error = false;
 
 	if (!error && (ioend->io_flags & IOMAP_IOEND_INTEGRITY)) {
 		error = iomap_ioend_integrity_verify(ioend);
-		if ((ioend->io_flags & IOMAP_IOEND_DIRECT) &&
-		    READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_LAZY) {
-			/*
-			 * We only really need to retry for guard tag errors,
-			 * but right now we can't distinguish them from other
-			 * (i.e, reftag) errors.
-			 */
-			if (error ||
-			    XFS_TEST_ERROR(mp, XFS_ERRTAG_BOUNCE_REREAD)) {
-				xfs_read_bounce_and_resubmit(ioend);
-				return;
-			}
-		}
+		/*
+		 * We only really need to retry for guard tag errors, but right
+		 * now we can't distinguish them from other (i.e, reftag) errors.
+		 */
+		if (error)
+			is_csum_error = true;
 	}
 
-	iomap_finish_ioends(ioend, error);
+	if (xfs_is_rtcsum_inode(ip)) {
+		error = xfs_rtcsum_verify_ioend(ioend, error);
+		if (error && !bio->bi_status)
+			is_csum_error = true;
+	}
+
+	/*
+	 * If we saw a checksum failure on a direct I/O read that uses lazy
+	 * bouncing, resubmit the read using a bounce buffer so that we can
+	 * guarantee this was not caused by the user corrupting the buffer.
+	 */
+	if ((ioend->io_flags & IOMAP_IOEND_DIRECT) &&
+	    READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_LAZY &&
+	    (is_csum_error ||
+	     (!error && XFS_TEST_ERROR(mp, XFS_ERRTAG_BOUNCE_REREAD))))
+		xfs_read_bounce_and_resubmit(ioend);
+	else
+		iomap_finish_ioends(ioend, error);
 }
 
 void
@@ -148,6 +225,9 @@ xfs_ioend_submit_read(
 		return;
 	}
 
+	if (xfs_is_rtcsum_inode(ip) && !xfs_rtcsum_prepare_read(ioend))
+		return;
+
 	if (ioend_flags & IOMAP_IOEND_INTEGRITY)
 		fs_bio_integrity_alloc(bio);
 	bio->bi_end_io = xfs_end_io_read;
diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 6701be9325ef..0e4396e52809 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -32,6 +32,7 @@
 #include "xfs_rtbitmap.h"
 #include "xfs_icache.h"
 #include "xfs_zone_alloc.h"
+#include "xfs_rtcsum.h"
 
 #define XFS_ALLOC_ALIGN(mp, off) \
 	(((off) >> mp->m_allocsize_log) << mp->m_allocsize_log)
@@ -166,6 +167,8 @@ xfs_bmbt_to_iomap(
 	}
 
 	iomap->validity_cookie = sequence_cookie;
+	if (xfs_is_rtcsum_inode(ip))
+		iomap->csum_shift = mp->m_rtcsum_shift;
 	return 0;
 }
 
@@ -2227,16 +2230,28 @@ xfs_read_iomap_begin(
 		return error;
 	error = xfs_bmapi_read(ip, offset_fsb, end_fsb - offset_fsb, &imap,
 			       &nimaps, 0);
-	if (!error && ((flags & IOMAP_REPORT) || IS_DAX(inode)))
+	if (error)
+		goto out_unlock;
+
+	if ((flags & IOMAP_REPORT) || IS_DAX(inode)) {
 		error = xfs_reflink_trim_around_shared(ip, &imap, &shared);
+		if (error)
+			goto out_unlock;
+	} else if (!isnullstartblock(imap.br_startblock) &&
+		   xfs_is_rtcsum_inode(ip)) {
+		imap.br_blockcount = min(imap.br_blockcount,
+				xfs_rtcsum_max_len(mp, imap.br_startblock));
+	}
+
 	seq = xfs_iomap_inode_sequence(ip, shared ? IOMAP_F_SHARED : 0);
 	xfs_iunlock(ip, lockmode);
-
-	if (error)
-		return error;
 	trace_xfs_iomap_found(ip, offset, length, XFS_DATA_FORK, &imap);
 	return xfs_bmbt_to_iomap(ip, iomap, &imap, flags,
 				 shared ? IOMAP_F_SHARED : 0, seq);
+
+out_unlock:
+	xfs_iunlock(ip, lockmode);
+	return error;
 }
 
 static DEFINE_IOMAP_ITER_NEXT(xfs_read_iomap_next, xfs_read_iomap_begin);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 15/21] xfs: add support for writing with data checksums
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (13 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:01   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
                   ` (6 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

All write to checksummed files need to use ioends so that the checksum
can be verified from process context.

The ioend submission path allocates the csum buffer and attaches it to
the ioend before generating the checksum from the file data.  The I/O
completion then logs the checksums into the buffers for the LBAs that
were written.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_file.c       |  7 +++---
 fs/xfs/xfs_ioend.c      | 38 ++++++++++++++++++++++++++++++++
 fs/xfs/xfs_ioend.h      |  2 ++
 fs/xfs/xfs_iomap.c      | 49 ++++++++++++++++++++++++++++++++++++-----
 fs/xfs/xfs_iomap.h      |  4 +++-
 fs/xfs/xfs_reflink.c    |  2 +-
 fs/xfs/xfs_zone_alloc.c |  5 +++++
 7 files changed, 97 insertions(+), 10 deletions(-)

diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
index 5b25f33527c0..a38191760f81 100644
--- a/fs/xfs/xfs_file.c
+++ b/fs/xfs/xfs_file.c
@@ -1073,7 +1073,8 @@ xfs_file_buffered_write(
 
 	trace_xfs_file_buffered_write(iocb, from);
 	ret = iomap_file_buffered_write(iocb, from,
-			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
+			&xfs_buffered_write_iomap_ops,
+			xfs_get_iomap_write_ops(ip),
 			NULL);
 
 	/*
@@ -1154,8 +1155,8 @@ xfs_file_buffered_write_zoned(
 retry:
 	trace_xfs_file_buffered_write(iocb, from);
 	ret = iomap_file_buffered_write(iocb, from,
-			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
-			&ac);
+			&xfs_buffered_write_iomap_ops,
+			xfs_get_iomap_write_ops(ip), &ac);
 	if (ret == -ENOSPC && !cleared_space) {
 		/*
 		 * Kick off writeback to convert delalloc space and release the
diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c
index 7570a1b915c0..7b9c82e3c449 100644
--- a/fs/xfs/xfs_ioend.c
+++ b/fs/xfs/xfs_ioend.c
@@ -235,6 +235,37 @@ xfs_ioend_submit_read(
 	submit_bio(bio);
 }
 
+int
+xfs_ioend_submit_read_sync(
+	struct bio		*bio,
+	struct inode		*inode,
+	loff_t			file_offset,
+	u16			ioend_flags)
+{
+	struct xfs_inode	*ip = XFS_I(inode);
+	struct iomap_ioend	*ioend;
+	struct bvec_iter	saved_iter;
+	int			error;
+
+	ASSERT(!(ioend_flags & IOMAP_IOEND_DIRECT));
+
+	ioend = iomap_init_ioend(inode, bio, file_offset, ioend_flags);
+	if (xfs_is_rtcsum_inode(ip) && !xfs_rtcsum_prepare_read(ioend))
+		return blk_status_to_errno(bio->bi_status);
+	if (ioend_flags & IOMAP_IOEND_INTEGRITY)
+		fs_bio_integrity_alloc(bio);
+	saved_iter = bio->bi_iter;
+	error = submit_bio_wait(bio);
+	if (bio_integrity(bio)) {
+		if (!error)
+			error = fs_bio_integrity_verify(bio, &saved_iter);
+		fs_bio_integrity_free(bio);
+	}
+	if (xfs_is_rtcsum_inode(ip))
+		error = xfs_rtcsum_verify_ioend(ioend, error);
+	return error;
+}
+
 static void
 xfs_end_ioend_write_zoned(
 	struct iomap_ioend	*ioend)
@@ -261,6 +292,13 @@ xfs_end_ioend_write_zoned(
 		goto done;
 	}
 
+	if (xfs_is_rtcsum_inode(ip)) {
+		error = xfs_rtcsum_log(oz, ioend->io_sector, ioend->io_size,
+				ioend->io_csum);
+		if (error)
+			goto done;
+	}
+
 	error = xfs_zoned_end_io(ip, ioend->io_offset, ioend->io_size,
 			ioend->io_sector, oz, NULLFSBLOCK);
 	if (error)
diff --git a/fs/xfs/xfs_ioend.h b/fs/xfs/xfs_ioend.h
index 7c2a1ea3e6ed..f01aa208c61f 100644
--- a/fs/xfs/xfs_ioend.h
+++ b/fs/xfs/xfs_ioend.h
@@ -14,5 +14,7 @@ static inline bool xfs_ioend_is_append(struct iomap_ioend *ioend)
 void xfs_end_bio(struct bio *bio);
 void xfs_ioend_submit_read(struct inode *inode, struct bio *bio,
 		loff_t file_offset, u16 ioend_flags);
+int xfs_ioend_submit_read_sync(struct bio *bio, struct inode *inode,
+		loff_t file_offset, u16 ioend_flags);
 
 #endif /* __XFS_IOEND_H */
diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 0e4396e52809..75d02a32a36a 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -33,6 +33,7 @@
 #include "xfs_icache.h"
 #include "xfs_zone_alloc.h"
 #include "xfs_rtcsum.h"
+#include "xfs_ioend.h"
 
 #define XFS_ALLOC_ALIGN(mp, off) \
 	(((off) >> mp->m_allocsize_log) << mp->m_allocsize_log)
@@ -93,10 +94,45 @@ xfs_iomap_valid(
 	return true;
 }
 
-const struct iomap_write_ops xfs_iomap_write_ops = {
+static const struct iomap_write_ops xfs_iomap_write_ops = {
 	.iomap_valid		= xfs_iomap_valid,
 };
 
+static int
+xfs_csum_read_folio_range(
+	const struct iomap_iter	*iter,
+	struct folio		*folio,
+	loff_t			pos,
+	size_t			len)
+{
+	const struct iomap	*srcmap = iomap_iter_srcmap(iter);
+	unsigned int		ioend_flags = iomap_ioend_flags(&iter->iomap);
+	struct bio		*bio;
+	int			error;
+
+	bio = bio_alloc_bioset(srcmap->bdev, 1, REQ_OP_READ, GFP_NOFS,
+			&iomap_ioend_bioset);
+	bio->bi_iter.bi_sector = iomap_sector(srcmap, pos);
+	bio_add_folio_nofail(bio, folio, len, offset_in_folio(folio, pos));
+	error = xfs_ioend_submit_read_sync(bio, iter->inode, pos, ioend_flags);
+	bio_put(bio);
+	return error;
+}
+
+static const struct iomap_write_ops xfs_iomap_csum_write_ops = {
+	.iomap_valid		= xfs_iomap_valid,
+	.read_folio_range	= xfs_csum_read_folio_range,
+};
+
+const struct iomap_write_ops *
+xfs_get_iomap_write_ops(
+	struct xfs_inode	*ip)
+{
+	if (xfs_is_rtcsum_inode(ip))
+		return &xfs_iomap_csum_write_ops;
+	return &xfs_iomap_write_ops;
+}
+
 int
 xfs_bmbt_to_iomap(
 	struct xfs_inode	*ip,
@@ -1617,6 +1653,9 @@ xfs_zoned_fill_srcmap(
 	 * There is a data fork mapping, only map until the end of it.
 	 */
 	xfs_trim_extent(&smap, offset_fsb, *end_fsb - offset_fsb);
+	if (xfs_is_rtcsum_inode(ip))
+		smap.br_blockcount = min(smap.br_blockcount,
+			xfs_rtcsum_max_len(ip->i_mount, smap.br_startblock));
 	*end_fsb = min(*end_fsb, smap.br_startoff + smap.br_blockcount);
 	return xfs_bmbt_to_iomap(ip, srcmap, &smap, flags, 0,
 			xfs_iomap_inode_sequence(ip, 0));
@@ -2415,8 +2454,8 @@ xfs_zero_range(
 		return dax_zero_range(inode, pos, len, did_zero,
 				      &xfs_dax_write_iomap_ops);
 	return iomap_zero_range(inode, pos, len, did_zero,
-			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
-			ac);
+			&xfs_buffered_write_iomap_ops,
+			xfs_get_iomap_write_ops(ip), ac);
 }
 
 int
@@ -2432,6 +2471,6 @@ xfs_truncate_page(
 		return dax_truncate_page(inode, pos, did_zero,
 					&xfs_dax_write_iomap_ops);
 	return iomap_truncate_page(inode, pos, did_zero,
-			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
-			ac);
+			&xfs_buffered_write_iomap_ops,
+			xfs_get_iomap_write_ops(ip), ac);
 }
diff --git a/fs/xfs/xfs_iomap.h b/fs/xfs/xfs_iomap.h
index f2520a9b3a13..bb35e58d31ee 100644
--- a/fs/xfs/xfs_iomap.h
+++ b/fs/xfs/xfs_iomap.h
@@ -43,6 +43,7 @@ xfs_iomap_set_anon_write(
 	iomap->flags = IOMAP_F_ANON_WRITE | IOMAP_F_DIRTY;
 	if (bdev_has_integrity_csum(iomap->bdev))
 		iomap->flags |= IOMAP_F_INTEGRITY;
+	iomap->csum_shift = ip->i_mount->m_rtcsum_shift;
 }
 
 static inline xfs_filblks_t
@@ -69,6 +70,8 @@ int xfs_read_iomap_begin(struct inode *inode, loff_t offset,
 		loff_t length, unsigned flags, struct iomap *iomap,
 		struct iomap *srcmap);
 
+const struct iomap_write_ops *xfs_get_iomap_write_ops(struct xfs_inode *ip);
+
 extern const struct iomap_ops xfs_buffered_write_iomap_ops;
 extern const struct iomap_ops xfs_direct_write_iomap_ops;
 extern const struct iomap_ops xfs_zoned_direct_write_iomap_ops;
@@ -77,6 +80,5 @@ extern const struct iomap_ops xfs_seek_iomap_ops;
 extern const struct iomap_ops xfs_xattr_iomap_ops;
 extern const struct iomap_ops xfs_dax_write_iomap_ops;
 extern const struct iomap_ops xfs_atomic_write_cow_iomap_ops;
-extern const struct iomap_write_ops xfs_iomap_write_ops;
 
 #endif /* __XFS_IOMAP_H__*/
diff --git a/fs/xfs/xfs_reflink.c b/fs/xfs/xfs_reflink.c
index 480136136635..6edbe12777ac 100644
--- a/fs/xfs/xfs_reflink.c
+++ b/fs/xfs/xfs_reflink.c
@@ -1918,7 +1918,7 @@ xfs_reflink_unshare(
 	else
 		error = iomap_file_unshare(inode, offset, len,
 				&xfs_buffered_write_iomap_ops,
-				&xfs_iomap_write_ops);
+				xfs_get_iomap_write_ops(ip));
 	if (error)
 		goto out;
 
diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c
index 9d9a713684b9..71cd35352e2e 100644
--- a/fs/xfs/xfs_zone_alloc.c
+++ b/fs/xfs/xfs_zone_alloc.c
@@ -27,6 +27,7 @@
 #include "xfs_trace.h"
 #include "xfs_mru_cache.h"
 #include "xfs_rtcsum.h"
+#include "xfs_rtcsum.h"
 #include <linux/bio-integrity.h>
 
 static void
@@ -934,6 +935,10 @@ xfs_zone_alloc_and_submit(
 
 	if (ioend->io_flags & IOMAP_IOEND_INTEGRITY)
 		fs_bio_integrity_generate(&ioend->io_bio);
+	if (xfs_is_rtcsum_inode(ip)) {
+		xfs_csum_generate(mp, &ioend->io_bio,
+				iomap_csum_alloc(ioend, mp->m_rtcsum_shift));
+	}
 
 	/*
 	 * If we don't have a locally cached zone in this write context, see if
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 16/21] xfs: add data checksum support to zoned garbage collection
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (14 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:06   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
                   ` (5 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Transfer the checksum from the old location to the new one.  The
implementation mirrors that of user data reads and writes.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_zone_gc.c | 65 ++++++++++++++++++++++++++++++++++++--------
 1 file changed, 53 insertions(+), 12 deletions(-)

diff --git a/fs/xfs/xfs_zone_gc.c b/fs/xfs/xfs_zone_gc.c
index 5fdcf98a2133..1cebd2ee2136 100644
--- a/fs/xfs/xfs_zone_gc.c
+++ b/fs/xfs/xfs_zone_gc.c
@@ -21,6 +21,7 @@
 #include "xfs_zone_alloc.h"
 #include "xfs_zone_priv.h"
 #include "xfs_zones.h"
+#include "xfs_rtcsum.h"
 #include "xfs_trace.h"
 
 /*
@@ -106,6 +107,10 @@ struct xfs_gc_bio {
 	/* Realtime group currently being reclaimed */
 	struct xfs_rtgroup		*victim_rtg;
 
+	/* Buffer for data checksums */
+	struct xfs_buf			*csum_bp;
+	void				*csum_buf;
+
 	/* Bio used for reads and writes, including the bvec used by it */
 	struct bio			bio;	/* must be last */
 };
@@ -660,6 +665,20 @@ xfs_zone_gc_alloc_blocks(
 	return true;
 }
 
+static void
+xfs_zone_gc_free_chunk(
+	struct xfs_gc_bio	*chunk)
+{
+	atomic_dec(&chunk->victim_rtg->rtg_gccount);
+	xfs_rtgroup_rele(chunk->victim_rtg);
+	list_del(&chunk->entry);
+	xfs_open_zone_put(chunk->oz);
+	if (chunk->csum_bp)
+		xfs_buf_rele(chunk->csum_bp);
+	xfs_irele(chunk->ip);
+	bio_put(&chunk->bio);
+}
+
 static void
 xfs_zone_gc_add_data(
 	struct xfs_gc_bio	*chunk)
@@ -725,6 +744,11 @@ xfs_zone_gc_start_chunk(
 	if (!xfs_zone_gc_iter_irec(mp, iter, &irec, &ip))
 		return false;
 
+	if (xfs_has_rtcsum(mp)) {
+		irec.rm_blockcount = min(irec.rm_blockcount,
+			xfs_rtcsum_max_len(mp, irec.rm_startblock));
+	}
+
 	if (!xfs_zone_gc_alloc_blocks(data, &irec.rm_blockcount, &daddr,
 			&is_seq)) {
 		xfs_irele(ip);
@@ -748,6 +772,7 @@ xfs_zone_gc_start_chunk(
 	chunk->data = data;
 	chunk->oz = data->oz;
 	chunk->victim_rtg = iter->victim_rtg;
+	chunk->csum_bp = NULL;
 	atomic_inc(&rtg_group(chunk->victim_rtg)->xg_active_ref);
 	atomic_inc(&chunk->victim_rtg->rtg_gccount);
 
@@ -764,22 +789,22 @@ xfs_zone_gc_start_chunk(
 	list_add_tail(&chunk->entry, &data->reading);
 	xfs_zone_gc_iter_advance(iter, irec.rm_blockcount);
 
+	if (xfs_is_rtcsum_inode(ip)) {
+		int error;
+
+		error = xfs_rtcsum_read_async(mp, chunk->old_startblock,
+				&chunk->csum_bp);
+		if (error) {
+			xfs_force_shutdown(mp, SHUTDOWN_META_IO_ERROR);
+			xfs_zone_gc_free_chunk(chunk);
+			return false;
+		}
+	}
+
 	submit_bio(bio);
 	return true;
 }
 
-static void
-xfs_zone_gc_free_chunk(
-	struct xfs_gc_bio	*chunk)
-{
-	atomic_dec(&chunk->victim_rtg->rtg_gccount);
-	xfs_rtgroup_rele(chunk->victim_rtg);
-	list_del(&chunk->entry);
-	xfs_open_zone_put(chunk->oz);
-	xfs_irele(chunk->ip);
-	bio_put(&chunk->bio);
-}
-
 static void
 xfs_zone_gc_submit_write(
 	struct xfs_zone_gc_data	*data,
@@ -832,6 +857,9 @@ xfs_zone_gc_split_write(
 	split_chunk->old_startblock = chunk->old_startblock;
 	split_chunk->new_daddr = chunk->new_daddr;
 	split_chunk->oz = chunk->oz;
+	split_chunk->csum_bp = chunk->csum_bp;
+	if (split_chunk->csum_bp)
+		xfs_buf_hold(split_chunk->csum_bp);
 	atomic_inc(&chunk->oz->oz_ref);
 
 	split_chunk->victim_rtg = chunk->victim_rtg;
@@ -919,6 +947,19 @@ xfs_zone_gc_finish_chunk(
 
 	if (chunk->is_seq)
 		chunk->new_daddr = chunk->bio.bi_iter.bi_sector;
+
+	if (xfs_is_rtcsum_inode(ip)) {
+		unsigned int	boff;
+
+		error = xfs_buf_read_async_wait(chunk->csum_bp);
+		if (error)
+			goto free;
+
+		boff = xfs_rtb_to_rtcsumoff(mp, chunk->old_startblock);
+		error = xfs_rtcsum_log(chunk->oz, chunk->new_daddr, chunk->len,
+				chunk->csum_bp->b_addr + boff);
+	}
+
 	error = xfs_zoned_end_io(ip, chunk->offset, chunk->len,
 			chunk->new_daddr, chunk->oz, chunk->old_startblock);
 free:
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 17/21] xfs: verify data checksums during media verification
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (15 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:19   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
                   ` (4 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Wire up reading and verifying data checksums during media verification.
This is very similar to the file read path in that it kicks of an async
read for the checksum buffer before reading the data, and then validating
once both are read in.

To support this, split the per-read logic in xfs_verify_media into a
separate helpers for the data checksums vs no checksum cases.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_verify_media.c | 131 +++++++++++++++++++++++++++++++-------
 1 file changed, 107 insertions(+), 24 deletions(-)

diff --git a/fs/xfs/xfs_verify_media.c b/fs/xfs/xfs_verify_media.c
index 71f4d6c832a9..46a19405c9e5 100644
--- a/fs/xfs/xfs_verify_media.c
+++ b/fs/xfs/xfs_verify_media.c
@@ -22,6 +22,7 @@
 #include "xfs_rtrmap_btree.h"
 #include "xfs_health.h"
 #include "xfs_healthmon.h"
+#include "xfs_rtcsum.h"
 #include "xfs_trace.h"
 #include "xfs_verify_media.h"
 
@@ -261,6 +262,95 @@ xfs_verify_media_error(
 	}
 }
 
+static int
+xfs_submit_verify_bio(
+	struct xfs_mount	*mp,
+	struct xfs_verify_media	*me,
+	struct xfs_buftarg	*btp,
+	struct folio		*folio,
+	xfs_daddr_t		*daddr,
+	uint64_t		*bbcount)
+{
+	unsigned int		bio_bbcount;
+	int			error;
+
+	bio_bbcount = min(*bbcount, folio_size(folio) >> SECTOR_SHIFT);
+	error = bdev_rw_virt(btp->bt_bdev, *daddr, folio_address(folio),
+			bio_bbcount << SECTOR_SHIFT,
+			REQ_OP_READ);
+	if (error) {
+		xfs_verify_media_error(mp, me, btp, *daddr, bio_bbcount, error);
+		return 1;
+	}
+
+	*daddr += bio_bbcount;
+	*bbcount -= bio_bbcount;
+	return 0;
+}
+
+static int
+xfs_submit_verify_bio_csum(
+	struct xfs_mount	*mp,
+	struct xfs_verify_media	*me,
+	struct xfs_buftarg	*btp,
+	struct folio		*folio,
+	xfs_daddr_t		*daddr,
+	uint64_t		*bbcount)
+{
+	struct xfs_buf		*csum_bp = NULL;
+	unsigned int		bio_bbcount;
+	struct bvec_iter	saved_iter;
+	xfs_fsblock_t		bno, end;
+	xfs_filblks_t		len;
+	struct bio		bio;
+	struct bio_vec		bv;
+	int			error;
+
+	bno = xfs_daddr_to_rtb(mp, *daddr);
+	end = xfs_daddr_to_rtb(mp, *daddr + *bbcount);
+	len = min(end - bno, XFS_B_TO_FSBT(mp, folio_size(folio)));
+	len = min(len, xfs_rtcsum_max_len(mp, bno));
+
+	error = xfs_rtcsum_read_async(mp, bno, &csum_bp);
+	if (error)
+		return error;
+
+	*daddr = xfs_rtb_to_daddr(mp, bno);
+	bio_bbcount = XFS_FSB_TO_BB(mp, len);
+
+	bio_init(&bio, btp->bt_bdev, &bv, 1, REQ_OP_READ);
+	bio.bi_iter.bi_sector = *daddr;
+	bio_add_folio_nofail(&bio, folio,
+			min(bio_bbcount << SECTOR_SHIFT, folio_size(folio)), 0);
+	saved_iter = bio.bi_iter;
+
+	error = submit_bio_wait(&bio);
+	if (error)
+		goto out_media_error;
+
+	error = xfs_buf_read_async_wait(csum_bp);
+	if (error)
+		goto out_buf_rele;
+
+	error = xfs_csum_verify(mp, &bio, &saved_iter,
+			csum_bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno,
+			false);
+	if (error)
+		goto out_media_error;
+
+	*daddr += bio_bbcount;
+	*bbcount -= bio_bbcount;
+
+out_buf_rele:
+	xfs_buf_rele(csum_bp);
+	bio_uninit(&bio);
+	return error;
+out_media_error:
+	xfs_verify_media_error(mp, me, btp, *daddr, bio_bbcount, error);
+	error = 1;
+	goto out_buf_rele;
+}
+
 /* Verify the media of an xfs device by submitting read requests to the disk. */
 static int
 xfs_verify_media(
@@ -310,18 +400,13 @@ xfs_verify_media(
 		return 0;
 
 	/*
-	 * There are three ranges involved here:
-	 *
-	 *  - [me->me_start_daddr, me->me_end_daddr) is the range that the
-	 *    user wants to verify.  end_daddr can be beyond the end of the
-	 *    disk; we'll constrain it to the end if necessary.
+	 * [me->me_start_daddr, me->me_end_daddr) is the range that the user
+	 * wants to verify.  end_daddr can be beyond the end of the disk; we'll
+	 * constrain it to the end if necessary.
 	 *
-	 *  - [daddr, me->me_end_daddr) is the range that we have not yet
-	 *    verified.  We update daddr after each successful read.
-	 *    me->me_start_daddr is set to daddr before returning.
-	 *
-	 *  - [daddr, daddr + bio_bbcount) is the range that we're currently
-	 *    verifying.
+	 * [daddr, me->me_end_daddr) is the range that we have not yet verified.
+	 * We update daddr after each successful read.  me->me_start_daddr is
+	 * set to daddr before returning.
 	 */
 	daddr = me->me_start_daddr;
 	bbcount = min_t(sector_t, me->me_end_daddr, btp->bt_nr_sectors) -
@@ -334,22 +419,20 @@ xfs_verify_media(
 	trace_xfs_verify_media(mp, me, btp->bt_dev, daddr, bbcount, folio);
 
 	for (;;) {
-		unsigned int	bio_bbcount;
-
-		bio_bbcount = min(bbcount, folio_size(folio) >> SECTOR_SHIFT);
-		error = bdev_rw_virt(btp->bt_bdev, daddr, folio_address(folio),
-				bio_bbcount << SECTOR_SHIFT,
-				REQ_OP_READ);
+		if (IS_ENABLED(CONFIG_XFS_RT) &&
+		    me->me_dev == XFS_DEV_RT &&
+		    xfs_has_rtcsum(mp)) {
+			error = xfs_submit_verify_bio_csum(mp, me, btp, folio,
+					&daddr, &bbcount);
+		} else {
+			error = xfs_submit_verify_bio(mp, me, btp, folio,
+					&daddr, &bbcount);
+		}
 		if (error) {
-			xfs_verify_media_error(mp, me, btp, daddr, bio_bbcount,
-					error);
-			error = 0;
+			if (error == 1)
+				error = 0;
 			break;
 		}
-
-		daddr += bio_bbcount;
-		bbcount -= bio_bbcount;
-
 		if (bbcount == 0)
 			break;
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 18/21] xfs: don't try to verify checksums on empty zones
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (16 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:25   ` Darrick J. Wong
  2026-10-08 11:43   ` Anuj gupta
  2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
                   ` (3 subsequent siblings)
  21 siblings, 2 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

xfs_scrub can sometimes send XFS_IOC_VERIFY_MEDIA ioctls for ranges
that have never been written since the last zone reset, which will
lead to checksum verification failures.

Protect against this by checking that the range is valid.  If the report
flag is set, a verification failure could lead to health reports and
the file system being marked corrupt, so lock out zone racing reset
completions for this case as well.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_verify_media.c | 44 ++++++++++++++++++++++++++++++++++++---
 1 file changed, 41 insertions(+), 3 deletions(-)

diff --git a/fs/xfs/xfs_verify_media.c b/fs/xfs/xfs_verify_media.c
index 46a19405c9e5..fc739f7ffa4e 100644
--- a/fs/xfs/xfs_verify_media.c
+++ b/fs/xfs/xfs_verify_media.c
@@ -25,6 +25,8 @@
 #include "xfs_rtcsum.h"
 #include "xfs_trace.h"
 #include "xfs_verify_media.h"
+#include "xfs_zone_alloc.h"
+#include "xfs_zone_priv.h"
 
 #include <linux/fserror.h>
 
@@ -288,6 +290,43 @@ xfs_submit_verify_bio(
 	return 0;
 }
 
+static int
+xfs_csum_verify_metafile(
+	struct xfs_mount	*mp,
+	struct bio		*bio,
+	struct bvec_iter	*saved_iter,
+	void			*csum_buf,
+	xfs_fsblock_t		bno)
+{
+	struct xfs_rtgroup	*rtg;
+	int			error = 0;
+
+	rtg = xfs_rtgroup_get(mp, xfs_rtb_to_rgno(mp, bno));
+	if (!rtg)
+		return -EFSCORRUPTED;
+
+	/*
+	 * Only validate the checksums for valid data, as data never written
+	 * will not have valid checksums.  We need to hold the ilock on the rmap
+	 * inode to prevent freeing of blocks and thus a zone reset to happen
+	 * underneath us.
+	 *
+	 * Note that this still relies on cooperating userspace, as there also
+	 * can be blocks that were written but never recorded after an unclean
+	 * shutdown, which this check does not catch.  It purely tries to deal
+	 * with races vs the previous FSMAP output used by xfs_scrub.
+	 */
+	xfs_ilock(rtg_rmap(rtg), XFS_ILOCK_SHARED);
+	if (!xa_get_mark(&mp->m_groups[XG_TYPE_RTG].xa, rtg_rgno(rtg),
+			XFS_RTG_FREE)) {
+		error = xfs_csum_verify(mp, bio, saved_iter, csum_buf, bno,
+				false);
+	}
+	xfs_iunlock(rtg_rmap(rtg), XFS_ILOCK_SHARED);
+	xfs_rtgroup_put(rtg);
+	return error;
+}
+
 static int
 xfs_submit_verify_bio_csum(
 	struct xfs_mount	*mp,
@@ -332,9 +371,8 @@ xfs_submit_verify_bio_csum(
 	if (error)
 		goto out_buf_rele;
 
-	error = xfs_csum_verify(mp, &bio, &saved_iter,
-			csum_bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno,
-			false);
+	error = xfs_csum_verify_metafile(mp, &bio, &saved_iter,
+			csum_bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno);
 	if (error)
 		goto out_media_error;
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (17 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:26   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
                   ` (2 subsequent siblings)
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Report the RT data checksum flag and the checksum algorithm to userspace.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_fs.h | 6 +++++-
 fs/xfs/libxfs/xfs_sb.c | 5 +++++
 2 files changed, 10 insertions(+), 1 deletion(-)

diff --git a/fs/xfs/libxfs/xfs_fs.h b/fs/xfs/libxfs/xfs_fs.h
index 185f09f327c0..afe83ff23ef3 100644
--- a/fs/xfs/libxfs/xfs_fs.h
+++ b/fs/xfs/libxfs/xfs_fs.h
@@ -191,7 +191,10 @@ struct xfs_fsop_geom {
 	__u32		rgcount;	/* number of realtime groups	*/
 	__u64		rtstart;	/* start of internal rt section */
 	__u64		rtreserved;	/* RT (zoned) reserved blocks	*/
-	__u64		reserved[14];	/* reserved space		*/
+	__u8		rtcsum_type;	/* RT data checksum type	*/
+	__u8		rtcsum_blklog;	/* log2 of rtcsum bsize		*/
+	__u8		reserved_pad[6];/* reserved space		*/
+	__u64		reserved[13];	/* reserved space		*/
 };
 
 #define XFS_FSOP_GEOM_SICK_COUNTERS	(1 << 0)  /* summary counters */
@@ -250,6 +253,7 @@ typedef struct xfs_fsop_resblks {
 #define XFS_FSOP_GEOM_FLAGS_PARENT	(1 << 25) /* linux parent pointers */
 #define XFS_FSOP_GEOM_FLAGS_METADIR	(1 << 26) /* metadata directories */
 #define XFS_FSOP_GEOM_FLAGS_ZONED	(1 << 27) /* zoned rt device */
+#define XFS_FSOP_GEOM_FLAGS_DATA_CSUM	(1 << 28) /* data checksums */
 
 /*
  * Minimum and maximum sizes need for growth checks.
diff --git a/fs/xfs/libxfs/xfs_sb.c b/fs/xfs/libxfs/xfs_sb.c
index 3a470aec6c0c..506caf3e5201 100644
--- a/fs/xfs/libxfs/xfs_sb.c
+++ b/fs/xfs/libxfs/xfs_sb.c
@@ -1660,6 +1660,11 @@ xfs_fs_geometry(
 		geo->flags |= XFS_FSOP_GEOM_FLAGS_METADIR;
 	if (xfs_has_zoned(mp))
 		geo->flags |= XFS_FSOP_GEOM_FLAGS_ZONED;
+	if (xfs_has_rtcsum(mp)) {
+		geo->flags |= XFS_FSOP_GEOM_FLAGS_DATA_CSUM;
+		geo->rtcsum_type = mp->m_sb.sb_rtcsum_type;
+		geo->rtcsum_blklog = mp->m_sb.sb_rtcsum_blklog;
+	}
 	geo->rtsectsize = sbp->sb_blocksize;
 	geo->dirblocksize = xfs_dir2_dirblock_bytes(sbp);
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (18 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:27   ` Darrick J. Wong
  2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
  2026-09-24 22:52 ` support for " Dave Chinner
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/xfs_message.c | 4 ++++
 fs/xfs/xfs_message.h | 1 +
 fs/xfs/xfs_mount.h   | 2 ++
 fs/xfs/xfs_rtcsum.c  | 2 +-
 4 files changed, 8 insertions(+), 1 deletion(-)

diff --git a/fs/xfs/xfs_message.c b/fs/xfs/xfs_message.c
index 0243e509a468..53e8ed976a9b 100644
--- a/fs/xfs/xfs_message.c
+++ b/fs/xfs/xfs_message.c
@@ -149,6 +149,10 @@ xfs_warn_experimental(
 			.opstate	= XFS_OPSTATE_WARNED_LARP,
 			.name		= "logged extended attributes",
 		},
+		[XFS_EXPERIMENTAL_CSUM] = {
+			.opstate	= XFS_OPSTATE_WARNED_CSUM,
+			.name		= "data checksum",
+		},
 	};
 	ASSERT(feat >= 0 && feat < XFS_EXPERIMENTAL_MAX);
 	BUILD_BUG_ON(ARRAY_SIZE(features) != XFS_EXPERIMENTAL_MAX);
diff --git a/fs/xfs/xfs_message.h b/fs/xfs/xfs_message.h
index 811b885f41c3..d858e93e5426 100644
--- a/fs/xfs/xfs_message.h
+++ b/fs/xfs/xfs_message.h
@@ -93,6 +93,7 @@ void xfs_buf_alert_ratelimited(struct xfs_buf *bp, const char *rlmsg,
 enum xfs_experimental_feat {
 	XFS_EXPERIMENTAL_SHRINK,
 	XFS_EXPERIMENTAL_LARP,
+	XFS_EXPERIMENTAL_CSUM,
 
 	XFS_EXPERIMENTAL_MAX,
 };
diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
index fa86697f463a..c62eb42e490a 100644
--- a/fs/xfs/xfs_mount.h
+++ b/fs/xfs/xfs_mount.h
@@ -586,6 +586,8 @@ __XFS_HAS_FEAT(nouuid, NOUUID)
  */
 #define XFS_OPSTATE_BLOCKGC_ENABLED	6
 
+/* Kernel has logged a warning about checksums */
+#define XFS_OPSTATE_WARNED_CSUM		8
 /* Kernel has logged a warning about shrink being used on this fs. */
 #define XFS_OPSTATE_WARNED_SHRINK	9
 /* Kernel has logged a warning about logged xattr updates being used. */
diff --git a/fs/xfs/xfs_rtcsum.c b/fs/xfs/xfs_rtcsum.c
index 7cdc5a5029eb..47a7b67f7412 100644
--- a/fs/xfs/xfs_rtcsum.c
+++ b/fs/xfs/xfs_rtcsum.c
@@ -310,8 +310,8 @@ xfs_rtcsum_mount(
 	}
 	mp->m_features |= XFS_FEAT_DAX_NEVER;
 
+	xfs_warn_experimental(mp, XFS_EXPERIMENTAL_CSUM);
 	xfs_info(mp, "using %s for RT device data checksums",
 		xfs_data_csum_names[mp->m_sb.sb_rtcsum_type]);
-
 	return 0;
 }
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* [PATCH 21/21] xfs: enable RT data checksums
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (19 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
@ 2026-09-24  9:59 ` Christoph Hellwig
  2026-09-29  1:27   ` Darrick J. Wong
  2026-09-24 22:52 ` support for " Dave Chinner
  21 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-24  9:59 UTC (permalink / raw)
  To: Carlos Maiolino
  Cc: Darrick J . Wong, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

All mandatory pieces are in place now, allow mounting.

Signed-off-by: Christoph Hellwig <hch@lst.de>
---
 fs/xfs/libxfs/xfs_format.h | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/fs/xfs/libxfs/xfs_format.h b/fs/xfs/libxfs/xfs_format.h
index 1be3d21910a7..8b269fd3ed9f 100644
--- a/fs/xfs/libxfs/xfs_format.h
+++ b/fs/xfs/libxfs/xfs_format.h
@@ -384,7 +384,8 @@ xfs_sb_has_compat_feature(
 		(XFS_SB_FEAT_RO_COMPAT_FINOBT | \
 		 XFS_SB_FEAT_RO_COMPAT_RMAPBT | \
 		 XFS_SB_FEAT_RO_COMPAT_REFLINK| \
-		 XFS_SB_FEAT_RO_COMPAT_INOBTCNT)
+		 XFS_SB_FEAT_RO_COMPAT_INOBTCNT | \
+		 XFS_SB_FEAT_RO_COMPAT_RTCSUM)
 #define XFS_SB_FEAT_RO_COMPAT_UNKNOWN	~XFS_SB_FEAT_RO_COMPAT_ALL
 static inline bool
 xfs_sb_has_ro_compat_feature(
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 69+ messages in thread

* Re: [PATCH 01/21] block: export fs_bio_integrity_verify
  2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
@ 2026-09-24 20:29   ` Darrick J. Wong
  0 siblings, 0 replies; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 20:29 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel, Damien Le Moal

On Thu, Sep 24, 2026 at 11:59:33AM +0200, Christoph Hellwig wrote:
> From: Damien Le Moal <dlemoal@kernel.org>
> 
> XFS will need to call this in the iomap_write_ops.read_folio_range
> implementation.
> 
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> Signed-off-by: Christoph Hellwig <hch@lst.de>

I'm not the arbiter of exports, but to me this looks ok,
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> ---
>  block/bio-integrity-fs.c | 1 +
>  1 file changed, 1 insertion(+)
> 
> diff --git a/block/bio-integrity-fs.c b/block/bio-integrity-fs.c
> index c8e91ada8ca6..20d0a3d52932 100644
> --- a/block/bio-integrity-fs.c
> +++ b/block/bio-integrity-fs.c
> @@ -74,6 +74,7 @@ int fs_bio_integrity_verify(struct bio *bio, struct bvec_iter *data_iter)
>  		bio_integrity_bytes(bi, data_iter->bi_size >> SECTOR_SHIFT);
>  	return blk_status_to_errno(bio_integrity_verify(bio, data_iter));
>  }
> +EXPORT_SYMBOL_GPL(fs_bio_integrity_verify);
>  
>  static int __init fs_bio_integrity_init(void)
>  {
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 02/21] iomap: add support for data checksumming
  2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
@ 2026-09-24 21:39   ` Darrick J. Wong
  2026-09-25  5:53     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 21:39 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:34AM +0200, Christoph Hellwig wrote:
> Add a new refcounted structure for a data checksum attached to
> an iomap_ioend, and helpers to allocate, split and access it.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/iomap/bio.c        |  3 +-
>  fs/iomap/direct-io.c  |  2 +-
>  fs/iomap/internal.h   | 21 +++++++++---
>  fs/iomap/ioend.c      | 77 ++++++++++++++++++++++++++++++++++++++++---
>  include/linux/iomap.h | 31 ++++++++++++++++-
>  5 files changed, 122 insertions(+), 12 deletions(-)
> 
> diff --git a/fs/iomap/bio.c b/fs/iomap/bio.c
> index d46c2f8ea18c..7b8fddd6d061 100644
> --- a/fs/iomap/bio.c
> +++ b/fs/iomap/bio.c
> @@ -151,7 +151,8 @@ int iomap_bio_read_folio_range(const struct iomap_iter *iter,
>  
>  	if (!bio ||
>  	    bio_end_sector(bio) != iomap_sector(&iter->iomap, iter->pos) ||
> -	    bio->bi_iter.bi_size > iomap_max_bio_size(&iter->iomap) - plen ||
> +	    bio->bi_iter.bi_size >
> +			iomap_max_bio_size(iter->inode, &iter->iomap) - plen ||
>  	    !bio_add_folio(bio, folio, plen, offset_in_folio(folio, iter->pos)))
>  		iomap_read_alloc_bio(iter, ctx, plen);
>  	return 0;
> diff --git a/fs/iomap/direct-io.c b/fs/iomap/direct-io.c
> index 41fdc90a9094..7cf70fb0165a 100644
> --- a/fs/iomap/direct-io.c
> +++ b/fs/iomap/direct-io.c
> @@ -344,7 +344,7 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter,
>  		struct iomap_dio *dio, loff_t pos, unsigned int alignment,
>  		blk_opf_t op)
>  {
> -	unsigned int maxsize = iomap_max_bio_size(&iter->iomap);
> +	unsigned int maxsize = iomap_max_bio_size(iter->inode, &iter->iomap);
>  	unsigned int nr_vecs;
>  	struct bio *bio;
>  	ssize_t ret;
> diff --git a/fs/iomap/internal.h b/fs/iomap/internal.h
> index 74e898b196dc..7ac5e300e010 100644
> --- a/fs/iomap/internal.h
> +++ b/fs/iomap/internal.h
> @@ -4,17 +4,28 @@
>  
>  #define IOEND_BATCH_SIZE	4096
>  
> +static inline unsigned int max_csum_io_size(const struct inode *inode,
> +		const struct iomap *iomap)
> +{
> +	return IOMAP_CSUM_MAX_SIZE << (inode->i_blkbits - iomap->csum_shift);
> +}
> +
>  /*
>   * Normally we can build bios as big as the data structure supports.
>   *
> - * But for integrity protected I/O we need to respect the maximum size of the
> - * single contiguous allocation for the integrity buffer.
> + * But for checksum or integrity protected I/O we need to respect the maximum
> + * size of the single contiguous allocation for the checksum/integrity buffer.
>   */
> -static inline size_t iomap_max_bio_size(const struct iomap *iomap)
> +static inline size_t iomap_max_bio_size(const struct inode *inode,
> +		const struct iomap *iomap)
>  {
> +	size_t max = BIO_MAX_SIZE;
> +
>  	if (iomap->flags & IOMAP_F_INTEGRITY)
> -		return max_integrity_io_size(bdev_limits(iomap->bdev));
> -	return BIO_MAX_SIZE;
> +		max = min(max, max_integrity_io_size(bdev_limits(iomap->bdev)));
> +	if (iomap->csum_shift)
> +		max = min(max, max_csum_io_size(inode, iomap));
> +	return max;
>  }
>  
>  u32 iomap_finish_ioend_buffered_read(struct iomap_ioend *ioend);
> diff --git a/fs/iomap/ioend.c b/fs/iomap/ioend.c
> index bbebecc31670..877f0f470947 100644
> --- a/fs/iomap/ioend.c
> +++ b/fs/iomap/ioend.c
> @@ -14,6 +14,7 @@
>  struct bio_set iomap_ioend_bioset;
>  EXPORT_SYMBOL_GPL(iomap_ioend_bioset);
>  static struct bio_set iomap_ioend_split_bioset;
> +static mempool_t ioend_csum_pool;
>  
>  struct iomap_ioend *iomap_init_ioend(struct inode *inode,
>  		struct bio *bio, loff_t file_offset, u16 ioend_flags)
> @@ -25,6 +26,7 @@ struct iomap_ioend *iomap_init_ioend(struct inode *inode,
>  	ioend->io_parent = NULL;
>  	INIT_LIST_HEAD(&ioend->io_list);
>  	ioend->io_flags = ioend_flags;
> +	ioend->io_csum_shift = 0;
>  	ioend->io_bvec_offset = bio->bi_iter.bi_offset;
>  	ioend->io_inode = inode;
>  	ioend->io_offset = file_offset;
> @@ -32,10 +34,55 @@ struct iomap_ioend *iomap_init_ioend(struct inode *inode,
>  	ioend->io_sector = bio->bi_iter.bi_sector;
>  	ioend->io_vi = NULL;
>  	ioend->io_private = NULL;
> +	ioend->io_csum = NULL;
>  	return ioend;
>  }
>  EXPORT_SYMBOL_GPL(iomap_init_ioend);
>  
> +static void *__iomap_csum_alloc(struct iomap_ioend *ioend)
> +{
> +	size_t csum_size = iomap_csum_size(ioend);
> +
> +	WARN_ON_ONCE(csum_size > IOMAP_CSUM_MAX_SIZE);
> +	if (csum_size <= sizeof(ioend->io_csum_inline)) {
> +		ioend->io_flags |= IOMAP_IOEND_CSUM_INLINE;
> +		return ioend->io_csum_inline;
> +	}
> +
> +	ioend->io_csum_alloc = kmalloc(csum_size,
> +		GFP_NOWAIT | __GFP_NOMEMALLOC | __GFP_NORETRY);

Odd indenting here -- two tabs?

> +	if (!ioend->io_csum_alloc) {
> +		struct page *page = mempool_alloc(&ioend_csum_pool, GFP_NOFS);
> +
> +		ioend->io_flags |= IOMAP_IOEND_CSUM_MEMPOOL;
> +		ioend->io_csum_alloc = page_address(page);
> +	}
> +	return ioend->io_csum_alloc;
> +}
> +
> +void *iomap_csum_alloc(struct iomap_ioend *ioend, u8 csum_shift)
> +{
> +	ioend->io_csum_shift = csum_shift;
> +	ioend->io_csum = __iomap_csum_alloc(ioend);
> +	return ioend->io_csum;
> +}
> +EXPORT_SYMBOL_GPL(iomap_csum_alloc);

AFAICT, io_csum_inline is a small amount of memory in the ioend itself
to store checksums for the blocks being written, and io_csum_alloc is a
dynamically allocated blob if the checksums don't fit inline?  And
io_csum points to wherever that data lives?

Hrm.  What if you want to split an ioend, I guess the child's io_csum
points to somewhere inside io_parent->io_csum?

(Perhaps it would help to point this out in the struct definition?)

> +void iomap_csum_free(struct iomap_ioend *ioend)
> +{
> +	if (ioend->io_flags & IOMAP_IOEND_CSUM_MEMPOOL) {
> +		mempool_free(virt_to_page(ioend->io_csum_alloc),
> +				&ioend_csum_pool);
> +	} else if (!(ioend->io_flags & IOMAP_IOEND_CSUM_INLINE)) {
> +		kfree(ioend->io_csum_alloc);
> +	}
> +
> +	ioend->io_flags &=
> +		~(IOMAP_IOEND_CSUM_MEMPOOL | IOMAP_IOEND_CSUM_INLINE);
> +	ioend->io_csum_shift = 0;
> +}
> +EXPORT_SYMBOL_GPL(iomap_csum_free);
> +
>  /*
>   * We're now finished for good with this ioend structure.  Update the folio
>   * state, release holds on bios, and finally free up memory.  Do not use the
> @@ -178,7 +225,7 @@ static bool iomap_can_add_to_ioend(struct iomap_writepage_ctx *wpc, loff_t pos,
>  	struct iomap_ioend *ioend = wpc->wb_ctx;
>  
>  	if (ioend->io_bio.bi_iter.bi_size >
> -	    iomap_max_bio_size(&wpc->iomap) - map_len)
> +	    iomap_max_bio_size(wpc->inode, &wpc->iomap) - map_len)
>  		return false;
>  	if (ioend_flags & IOMAP_IOEND_BOUNDARY)
>  		return false;
> @@ -321,7 +368,7 @@ EXPORT_SYMBOL_GPL(iomap_ioend_integrity_verify);
>  
>  static u32 iomap_finish_ioend(struct iomap_ioend *ioend, int error)
>  {
> -	if (ioend->io_parent) {
> +	if (ioend->io_flags & IOMAP_IOEND_CHAINED) {
>  		struct bio *bio = &ioend->io_bio;
>  
>  		ioend = ioend->io_parent;
> @@ -334,6 +381,9 @@ static u32 iomap_finish_ioend(struct iomap_ioend *ioend, int error)
>  	if (!atomic_dec_and_test(&ioend->io_remaining))
>  		return 0;
>  
> +	if (iomap_has_csum(ioend))
> +		iomap_csum_free(ioend);
> +
>  	if (ioend->io_flags & IOMAP_IOEND_DIRECT)
>  		return iomap_finish_ioend_direct(ioend);
>  	if (bio_op(&ioend->io_bio) == REQ_OP_READ)
> @@ -409,6 +459,8 @@ static bool iomap_ioend_can_merge(struct iomap_ioend *ioend,
>  	if (ioend->io_sector + (ioend->io_size >> SECTOR_SHIFT) !=
>  	    next->io_sector)
>  		return false;
> +	if (ioend->io_csum || next->io_csum)
> +		return false;
>  	return true;
>  }
>  
> @@ -498,8 +550,9 @@ struct iomap_ioend *iomap_split_ioend(struct iomap_ioend *ioend,
>  	split->bi_end_io = bio->bi_end_io;
>  
>  	split_ioend = iomap_init_ioend(ioend->io_inode, split, ioend->io_offset,
> -			ioend->io_flags);
> +			ioend->io_flags | IOMAP_IOEND_CHAINED);
>  	split_ioend->io_parent = ioend;
> +	split_ioend->io_csum_shift = ioend->io_csum_shift;
>  
>  	atomic_inc(&ioend->io_remaining);
>  	ioend->io_offset += split_ioend->io_size;
> @@ -508,6 +561,12 @@ struct iomap_ioend *iomap_split_ioend(struct iomap_ioend *ioend,
>  	split_ioend->io_sector = ioend->io_sector;
>  	if (!is_append)
>  		ioend->io_sector += (split_ioend->io_size >> SECTOR_SHIFT);
> +
> +	if (iomap_has_csum(ioend)) {
> +		split_ioend->io_csum = ioend->io_csum;
> +		ioend->io_csum += iomap_csum_size(split_ioend);
> +	}
> +
>  	return split_ioend;
>  }
>  EXPORT_SYMBOL_GPL(iomap_split_ioend);
> @@ -547,7 +606,9 @@ void iomap_bounce_read(struct iomap_ioend *orig_ioend, unsigned int minsize,
>  		bio->bi_iter.bi_sector = sector;
>  
>  		ioend = iomap_init_ioend(inode, bio, file_offset,
> -				orig_ioend->io_flags);
> +				orig_ioend->io_flags &
> +					~(IOMAP_IOEND_CSUM_MEMPOOL |
> +					  IOMAP_IOEND_CSUM_INLINE));
>  
>  		total_len -= bio->bi_iter.bi_size;
>  		file_offset += bio->bi_iter.bi_size;
> @@ -593,6 +654,8 @@ void iomap_bounce_read_end_io(struct iomap_ioend *ioend, struct bio *orig_bio,
>  	else
>  		iomap_ioend_unbounce(iomap_ioend_from_bio(orig_bio), ioend);
>  
> +	if (iomap_has_csum(ioend))
> +		iomap_csum_free(ioend);
>  	bio_free_folios(&ioend->io_bio);
>  	if (bio_integrity(&ioend->io_bio))
>  		fs_bio_integrity_free(&ioend->io_bio);
> @@ -617,8 +680,14 @@ static int __init iomap_ioend_init(void)
>  			   BIOSET_NEED_BVECS);
>  	if (error)
>  		goto out_exit_ioend_bioset;
> +	error = mempool_init_page_pool(&ioend_csum_pool, BIO_POOL_SIZE,
> +			get_order(IOMAP_CSUM_MAX_SIZE));
> +	if (error)
> +		goto out_exit_ioend_split_bioset;
>  	return 0;
>  
> +out_exit_ioend_split_bioset:
> +	bioset_exit(&iomap_ioend_split_bioset);
>  out_exit_ioend_bioset:
>  	bioset_exit(&iomap_ioend_bioset);
>  	return error;
> diff --git a/include/linux/iomap.h b/include/linux/iomap.h
> index 59718f73c15a..e7db13ea0f70 100644
> --- a/include/linux/iomap.h
> +++ b/include/linux/iomap.h
> @@ -133,6 +133,7 @@ struct iomap {
>  	u64			length;	/* length of mapping, bytes */
>  	u16			type;	/* type of mapping */
>  	u16			flags;	/* flags for mapping */
> +	u8			csum_shift; /* ilog() of csum size */
>  	struct block_device	*bdev;	/* block device for I/O */
>  	struct dax_device	*dax_dev; /* dax_dev for dax operations */
>  	void			*inline_data;
> @@ -489,6 +490,12 @@ sector_t iomap_bmap(struct address_space *mapping, sector_t bno,
>  #else
>  #define IOMAP_IOEND_INTEGRITY		0
>  #endif /* CONFIG_BLK_DEV_INTEGRITY */
> +/* chained ioend that has io_parent */
> +#define IOMAP_IOEND_CHAINED		(1U << 6)
> +/* using io_csum_inline */
> +#define IOMAP_IOEND_CSUM_INLINE		(1U << 7)
> +/* io_csum is backed by a mempool */
> +#define IOMAP_IOEND_CSUM_MEMPOOL	(1U << 8)
>  
>  /*
>   * Flags that if set on either ioend prevent the merge of two ioends.
> @@ -522,14 +529,20 @@ static inline u16 iomap_ioend_flags(const struct iomap *iomap)
>  struct iomap_ioend {
>  	struct list_head	io_list;	/* next ioend in chain */
>  	u16			io_flags;	/* IOMAP_IOEND_* */
> +	u8			io_csum_shift;	/* ilog(2) of csum size */
>  	u32			io_bvec_offset;	/* offset into first bvec */
>  	struct inode		*io_inode;	/* file being written to */
>  	size_t			io_size;	/* size of the extent */
>  	atomic_t		io_remaining;	/* completetion defer count */
>  	int			io_error;	/* stashed away status */
> -	struct iomap_ioend	*io_parent;	/* parent for completions */
> +	union {
> +		struct iomap_ioend *io_parent;	/* parent for completions */
> +		void		*io_csum_alloc; /* original csum allocation. */
> +		u8		io_csum_inline[sizeof(void *)];
> +	};
>  	loff_t			io_offset;	/* offset in the file */
>  	sector_t		io_sector;	/* start sector of ioend */
> +	void			*io_csum;	/* data checksum */
>  	void			*io_private;	/* file system private data */
>  	struct fsverity_info	*io_vi;		/* fsverity info */
>  	struct bio		io_bio;		/* MUST BE LAST! */
> @@ -547,6 +560,22 @@ static inline struct iomap_ioend *iomap_ioend_from_bio(struct bio *bio)
>  	.bi_offset	= (_ioend)->io_bvec_offset,	\
>  }
>  
> +#define IOMAP_CSUM_MAX_SIZE	SZ_64K
> +
> +static inline bool iomap_has_csum(const struct iomap_ioend *ioend)
> +{
> +	return ioend->io_csum_shift > 0;
> +}
> +
> +static inline size_t iomap_csum_size(const struct iomap_ioend *ioend)
> +{
> +	return DIV_ROUND_UP(ioend->io_size, i_blocksize(ioend->io_inode)) <<
> +			ioend->io_csum_shift;
> +}
> +
> +void *iomap_csum_alloc(struct iomap_ioend *ioend, u8 csum_shift);
> +void iomap_csum_free(struct iomap_ioend *ioend);
> +
>  struct iomap_writeback_ops {
>  	/*
>  	 * Performs writeback on the passed in range
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API
  2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
@ 2026-09-24 21:43   ` Darrick J. Wong
  2026-09-25  5:54     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 21:43 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:35AM +0200, Christoph Hellwig wrote:
> Add a new helper that reads a buffer asynchronously.  This is similar
> to readahead, but doesn't become a no-op under memory or I/O congestion
> and returns the buffer to be read.
> 
> The intended use is to kick off a read of data checksum buffers at
> roughly the same time as the data read so that they are available
> in the I/O completion handler.

Does there need to be a "wait until this async-read buffer reaches
XBF_DONE" function too?  Or how do callers do that?

--D

> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_buf.c   | 130 ++++++++++++++++++++++++++++++++++++---------
>  fs/xfs/xfs_buf.h   |   4 ++
>  fs/xfs/xfs_trace.h |   3 ++
>  3 files changed, 113 insertions(+), 24 deletions(-)
> 
> diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c
> index eee491c01d8c..966c06bffeff 100644
> --- a/fs/xfs/xfs_buf.c
> +++ b/fs/xfs/xfs_buf.c
> @@ -627,6 +627,40 @@ _xfs_buf_read(
>  	return xfs_buf_iowait(bp);
>  }
>  
> +/*
> + * If we've had a read error, then the contents of the buffer are invalid and
> + * should not be used.  To ensure that a followup read tries to pull the buffer
> + * from disk again, we clear the XBF_DONE flag and mark the buffer stale.
> + * This ensures that anyone who has a current reference to the buffer will
> + * interpret it's contents correctly and future cache lookups will also treat it
> + * as an empty, uninitialised buffer.
> + */
> +static int
> +xfs_buf_read_error(
> +	struct xfs_buf		*bp,
> +	xfs_failaddr_t		fa,
> +	int			error)
> +{
> +	/*
> +	 * Check against log shutdown for error reporting because metadata
> +	 * writeback may require a read first and we need to report errors in
> +	 * metadata writeback until the log is shut down.
> +	 * High level transaction read functions already check against mount
> +	 * shutdown, anyway, so we only need to be concerned about low level IO
> +	 * interactions here.
> +	 */
> +	if (!xlog_is_shutdown(bp->b_mount->m_log))
> +		xfs_buf_ioerror_alert(bp, fa);
> +	xfs_buf_clear_flags(bp, XBF_DONE);
> +	xfs_buf_stale(bp);
> +	xfs_buf_relse(bp);
> +
> +	/* bad CRC means corrupted metadata */
> +	if (error == -EFSBADCRC)
> +		return -EFSCORRUPTED;
> +	return error;
> +}
> +
>  int
>  xfs_buf_read_map(
>  	struct xfs_buftarg	*target,
> @@ -699,39 +733,87 @@ xfs_buf_read_map(
>  	}
>  
>  	if (error)
> -		goto out_ioerror;
> -
> +		return xfs_buf_read_error(bp, fa, error);
>  	*bpp = bp;
>  	return 0;
> +}
>  
> -out_ioerror:
> +int
> +xfs_buf_read_async_wait(
> +	struct xfs_buf		*bp)
> +{
>  	/*
> -	 * Check against log shutdown for error reporting because metadata
> -	 * writeback may require a read first and we need to report errors in
> -	 * metadata writeback until the log is shut down.  High level
> -	 * transaction read functions already check against mount shutdown, so
> -	 * we only need to be concerned about low level/ IO interactions here.
> +	 * Protect against the case where the checksum read is slower than the
> +	 * data read.
>  	 */
> -	if (!xlog_is_shutdown(target->bt_mount->m_log))
> -		xfs_buf_ioerror_alert(bp, fa);
> +	if ((READ_ONCE(bp->b_flags) & (XBF_DONE | XBF_STALE)) == XBF_DONE &&
> +	    !bp->b_error) {
> +		trace_xfs_buf_read_async_wait(bp, 0, _RET_IP_);
> +		return 0;
> +	}
> +
> +	/* xfs_buf_find_lock can't return an error with 0 flags */
> +	xfs_buf_find_lock(bp, 0);
> +	trace_xfs_buf_read_async_lock(bp, 0, _RET_IP_);
> +	if (bp->b_error)
> +		return xfs_buf_read_error(bp, __builtin_return_address(0),
> +				bp->b_error);
> +	ASSERT(bp->b_ops);
> +	xfs_buf_clear_flags(bp, XBF_READ);
> +	xfs_buf_unlock(bp);
> +	return 0;
> +}
> +
> +/*
> + * Kick off an asynchronous read.  Unlike readahead, this returns a reference
> + * to the buffer, and reliably reads the data instead of skipping the read on
> + * memory pressure.
> + *
> + * The buffer may be locked when I/O is kicked off, but the I/O completion
> + * handler will unlock it.  The caller needs to lock itself if need to prevent
> + * concurrent access or to synchronize with I/O completion.
> + */
> +int
> +xfs_buf_read_async(
> +	struct xfs_buftarg	*btp,
> +	xfs_daddr_t		daddr,
> +	size_t			numblks,
> +	const struct xfs_buf_ops *ops,
> +	struct xfs_buf		**bpp)
> +{
> +	DEFINE_SINGLE_BUF_MAP(map, daddr, numblks);
> +	struct xfs_buf		*bp;
> +	int			error;
> +
> +	ASSERT(!xfs_buftarg_is_mem(btp));
> +
> +	error = xfs_find_get_buf(btp, &map, 1, XBF_READ, &bp);
> +	if (error)
> +		return error;
>  
>  	/*
> -	 * If we've had a read error, then the contents of the buffer are
> -	 * invalid and should not be used. To ensure that a followup read tries
> -	 * to pull the buffer from disk again, we clear the XBF_DONE flag and
> -	 * mark the buffer stale. This ensures that anyone who has a current
> -	 * reference to the buffer will interpret it's contents correctly and
> -	 * future cache lookups will also treat it as an empty, uninitialised
> -	 * buffer.
> +	 * Do a lockless fast path check for a valid uptodate buffer and avoid
> +	 * locking entirely in this case.
>  	 */
> -	xfs_buf_clear_flags(bp, XBF_DONE);
> -	xfs_buf_stale(bp);
> -	xfs_buf_relse(bp);
> +	if ((READ_ONCE(bp->b_flags) & (XBF_DONE | XBF_STALE)) == XBF_DONE)
> +		goto done;
>  
> -	/* bad CRC means corrupted metadata */
> -	if (error == -EFSBADCRC)
> -		return -EFSCORRUPTED;
> -	return error;
> +	/* xfs_buf_find_lock can't return an error with 0 flags */
> +	xfs_buf_find_lock(bp, 0);
> +	if (bp->b_flags & XBF_DONE) {
> +		xfs_buf_unlock(bp);
> +		goto done;
> +	}
> +	trace_xfs_buf_read_async(bp, 0, _RET_IP_);
> +	XFS_STATS_INC(btp->bt_mount, xb_get_read);
> +	xfs_buf_hold(bp);
> +	bp->b_ops = ops;
> +	xfs_buf_clear_flags(bp, XBF_WRITE);
> +	xfs_buf_set_flags(bp, XBF_READ | XBF_ASYNC);
> +	xfs_buf_submit(bp);
> +done:
> +	*bpp = bp;
> +	return 0;
>  }
>  
>  /*
> diff --git a/fs/xfs/xfs_buf.h b/fs/xfs/xfs_buf.h
> index a4729253b56f..1b352ec91aa6 100644
> --- a/fs/xfs/xfs_buf.h
> +++ b/fs/xfs/xfs_buf.h
> @@ -255,6 +255,10 @@ xfs_buf_readahead(
>  	return xfs_buf_readahead_map(target, &map, 1, ops);
>  }
>  
> +int xfs_buf_read_async(struct xfs_buftarg *btp, xfs_daddr_t daddr,
> +		size_t numblks, const struct xfs_buf_ops *ops,
> +		struct xfs_buf **bpp);
> +int xfs_buf_read_async_wait(struct xfs_buf *bp);
>  int xfs_buf_get_uncached(struct xfs_buftarg *target, size_t numblks,
>  		struct xfs_buf **bpp);
>  int xfs_buf_read_uncached(struct xfs_buftarg *target, xfs_daddr_t daddr,
> diff --git a/fs/xfs/xfs_trace.h b/fs/xfs/xfs_trace.h
> index 2af9a1429ae9..afaabd3adc73 100644
> --- a/fs/xfs/xfs_trace.h
> +++ b/fs/xfs/xfs_trace.h
> @@ -840,6 +840,9 @@ DEFINE_EVENT(xfs_buf_flags_class, name, \
>  	TP_ARGS(bp, flags, caller_ip))
>  DEFINE_BUF_FLAGS_EVENT(xfs_buf_get);
>  DEFINE_BUF_FLAGS_EVENT(xfs_buf_read);
> +DEFINE_BUF_FLAGS_EVENT(xfs_buf_read_async);
> +DEFINE_BUF_FLAGS_EVENT(xfs_buf_read_async_wait);
> +DEFINE_BUF_FLAGS_EVENT(xfs_buf_read_async_lock);
>  DEFINE_BUF_FLAGS_EVENT(xfs_buf_readahead);
>  
>  TRACE_EVENT(xfs_buf_ioerror,
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers
  2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
@ 2026-09-24 21:44   ` Darrick J. Wong
  0 siblings, 0 replies; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 21:44 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:36AM +0200, Christoph Hellwig wrote:
> Translate from a disk address to the realtime group and group-relative
> block numbers.  This will be needed by the data checksumming code.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>

Looks fine to me.
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> ---
>  fs/xfs/libxfs/xfs_rtgroup.h | 20 ++++++++++++++++++++
>  1 file changed, 20 insertions(+)
> 
> diff --git a/fs/xfs/libxfs/xfs_rtgroup.h b/fs/xfs/libxfs/xfs_rtgroup.h
> index fca2eb74908c..f26e324f0de3 100644
> --- a/fs/xfs/libxfs/xfs_rtgroup.h
> +++ b/fs/xfs/libxfs/xfs_rtgroup.h
> @@ -390,4 +390,24 @@ xfs_rtgroup_raw_size(
>  	return g->blocks;
>  }
>  
> +static inline xfs_rgnumber_t
> +xfs_daddr_to_rgno(struct xfs_mount *mp, xfs_daddr_t d)
> +{
> +	struct xfs_groups	*g = &mp->m_groups[XG_TYPE_RTG];
> +	xfs_rfsblock_t		rbno = XFS_BB_TO_FSBT(mp, d) - g->start_fsb;
> +
> +	ASSERT(xfs_has_rtgroups(mp));
> +	return div_u64(rbno, xfs_rtgroup_raw_size(mp));
> +}
> +
> +static inline xfs_rgblock_t
> +xfs_daddr_to_rgbno(struct xfs_mount *mp, xfs_daddr_t d)
> +{
> +	struct xfs_groups	*g = &mp->m_groups[XG_TYPE_RTG];
> +	xfs_rfsblock_t		rbno = XFS_BB_TO_FSBT(mp, d) - g->start_fsb;
> +
> +	ASSERT(xfs_has_rtgroups(mp));
> +	return do_div(rbno, xfs_rtgroup_raw_size(mp));
> +}
> +
>  #endif /* __LIBXFS_RTGROUP_H */
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC
  2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
@ 2026-09-24 21:49   ` Darrick J. Wong
  2026-09-25  5:57     ` Christoph Hellwig
  2026-10-08 11:46   ` Anuj gupta
  1 sibling, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 21:49 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:37AM +0200, Christoph Hellwig wrote:
> Add a flag so that the shadow CIL buffer for a buffer log item is always
> sizes to the maximum to prevent reallocations.  This will be used for the

  sized

> RT checksum item, where we know that we are going to fill it up very soon,
> and (almost) sequentially, so there is no point in doing a constant
> realloc cycle when more data is added to it.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/libxfs/xfs_trans_resv.c |  2 +-
>  fs/xfs/libxfs/xfs_trans_resv.h |  1 +
>  fs/xfs/xfs_buf_item.c          | 10 ++++++++++
>  fs/xfs/xfs_buf_item.h          |  4 +++-
>  4 files changed, 15 insertions(+), 2 deletions(-)
> 
> diff --git a/fs/xfs/libxfs/xfs_trans_resv.c b/fs/xfs/libxfs/xfs_trans_resv.c
> index 3151e97ca8ff..5382ece51812 100644
> --- a/fs/xfs/libxfs/xfs_trans_resv.c
> +++ b/fs/xfs/libxfs/xfs_trans_resv.c
> @@ -53,7 +53,7 @@ xfs_buf_log_overhead(void)
>   * will be changed in a transaction.  size is used to tell how many
>   * bytes should be reserved per item.
>   */
> -STATIC uint
> +uint
>  xfs_calc_buf_res(
>  	uint		nbufs,
>  	uint		size)
> diff --git a/fs/xfs/libxfs/xfs_trans_resv.h b/fs/xfs/libxfs/xfs_trans_resv.h
> index 336279e0fc61..1804e821f382 100644
> --- a/fs/xfs/libxfs/xfs_trans_resv.h
> +++ b/fs/xfs/libxfs/xfs_trans_resv.h
> @@ -96,6 +96,7 @@ struct xfs_trans_resv {
>  #define	XFS_ITRUNCATE_LOG_COUNT_REFLINK	8
>  #define	XFS_WRITE_LOG_COUNT_REFLINK	8
>  
> +uint xfs_calc_buf_res(uint nbufs, uint size);
>  void xfs_trans_resv_calc(struct xfs_mount *mp, struct xfs_trans_resv *resp);
>  uint xfs_allocfree_block_count(struct xfs_mount *mp, uint num_ops);
>  
> diff --git a/fs/xfs/xfs_buf_item.c b/fs/xfs/xfs_buf_item.c
> index 1a4ef34af8d5..644c3fb18310 100644
> --- a/fs/xfs/xfs_buf_item.c
> +++ b/fs/xfs/xfs_buf_item.c
> @@ -252,6 +252,16 @@ xfs_buf_item_size(
>  		offset += BBTOB(bp->b_maps[i].bm_len);
>  	}
>  
> +	/*
> +	 * For buffers with the prealloc flag, always size the allocation size
> +	 * to the maximum as per the log reservation.  This avoids constant
> +	 * realloc cycles for buffers that are filled sequentially in rapid
> +	 * pace.  Note that the nvecs calaculation is kept from the regular

	                              calculation

> +	 * look as the buffer item formatting expects it.

"...is kept from the regular look as the buffer item formatting expects
it" ?

I don't understand that.  Is the nvecs calculation kept as the buffer
item formatting code expects it, even though we're allocating more
shadow buffer space?

--D

> +	 */
> +	if (bip->bli_flags & XFS_BLI_PREALLOC)
> +		*nbytes = xfs_calc_buf_res(bip->bli_format_count, bp->b_length);
> +
>  	/*
>  	 * Round up the buffer size required to minimise the number of memory
>  	 * allocations that need to be done as this item grows when relogged by
> diff --git a/fs/xfs/xfs_buf_item.h b/fs/xfs/xfs_buf_item.h
> index 3159325dd17b..ddc8ecc4683e 100644
> --- a/fs/xfs/xfs_buf_item.h
> +++ b/fs/xfs/xfs_buf_item.h
> @@ -20,6 +20,7 @@ struct xfs_mount;
>  #define XFS_BLI_STALE_INODE	(1u << 5)
>  #define	XFS_BLI_INODE_BUF	(1u << 6)
>  #define	XFS_BLI_ORDERED		(1u << 7)
> +#define	XFS_BLI_PREALLOC	(1u << 8)
>  
>  #define XFS_BLI_FLAGS \
>  	{ XFS_BLI_HOLD,		"HOLD" }, \
> @@ -29,7 +30,8 @@ struct xfs_mount;
>  	{ XFS_BLI_INODE_ALLOC_BUF, "INODE_ALLOC" }, \
>  	{ XFS_BLI_STALE_INODE,	"STALE_INODE" }, \
>  	{ XFS_BLI_INODE_BUF,	"INODE_BUF" }, \
> -	{ XFS_BLI_ORDERED,	"ORDERED" }
> +	{ XFS_BLI_ORDERED,	"ORDERED" }, \
> +	{ XFS_BLI_PREALLOC,	"PREALLOC" }
>  
>  /*
>   * This is the in core log item structure used to track information
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks
  2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
@ 2026-09-24 22:03   ` Darrick J. Wong
  2026-09-25  5:58     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 22:03 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:38AM +0200, Christoph Hellwig wrote:
> The upcoming RT data checksum feature will use larger than FSB blocks.
> Prepare xfs_rtfile_initialize_blocks to pass the number of FSBs per
> RT blocks, and to pass bmapi_flags to ask for contiguous allocation.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/libxfs/xfs_rtbitmap.c | 44 +++++++++++++++++++++---------------
>  fs/xfs/libxfs/xfs_rtbitmap.h |  3 ++-
>  fs/xfs/xfs_rtalloc.c         |  4 ++--
>  3 files changed, 30 insertions(+), 21 deletions(-)
> 
> diff --git a/fs/xfs/libxfs/xfs_rtbitmap.c b/fs/xfs/libxfs/xfs_rtbitmap.c
> index 01536f4fb386..db6a22b4506a 100644
> --- a/fs/xfs/libxfs/xfs_rtbitmap.c
> +++ b/fs/xfs/libxfs/xfs_rtbitmap.c
> @@ -1346,6 +1346,7 @@ xfs_rtfile_alloc_blocks(
>  	struct xfs_inode	*ip,
>  	xfs_fileoff_t		offset_fsb,
>  	xfs_filblks_t		count_fsb,
> +	uint32_t		bmapi_flags,
>  	struct xfs_bmbt_irec	*map)
>  {
>  	struct xfs_mount	*mp = ip->i_mount;
> @@ -1367,7 +1368,7 @@ xfs_rtfile_alloc_blocks(
>  		goto out_trans_cancel;
>  
>  	error = xfs_bmapi_write(tp, ip, offset_fsb, count_fsb,
> -			XFS_BMAPI_METADATA, 0, map, &nmap);
> +			XFS_BMAPI_METADATA | bmapi_flags, 0, map, &nmap);
>  	if (error)
>  		goto out_trans_cancel;
>  
> @@ -1405,34 +1406,43 @@ xfs_rtfile_initialize_block(
>  	struct xfs_rtgroup	*rtg,
>  	enum xfs_rtg_inodes	type,
>  	xfs_fsblock_t		fsbno,
> -	void			*data)
> +	xfs_filblks_t		nblks,
> +	void			**data)
>  {
>  	struct xfs_mount	*mp = rtg_mount(rtg);
>  	struct xfs_inode	*ip = rtg->rtg_inodes[type];
> +	size_t			len = XFS_FSB_TO_B(mp, nblks);
> +	size_t			copylen = len;
> +	struct xfs_trans_res	tres = M_RES(mp)->tr_growrtzero;
>  	struct xfs_trans	*tp;
>  	struct xfs_buf		*bp;
> -	const size_t		copylen = mp->m_blockwsize << XFS_WORDLOG;
>  	int			error;
>  
> -	error = xfs_trans_alloc(mp, &M_RES(mp)->tr_growrtzero, 0, 0, 0, &tp);
> +	tres.tr_logres *= nblks;

Hmm.  Is it safe to multiply the log reservation by an arbitrary
block count?  I would think we'd want *some* guarantee that we can't
create a transaction that's larger than the log can support.

It might suffice to put in a safeguard like:

	/* Log should always be able to handle 64k of logged buffers */
	ASSERT(nblks <= XFS_B_TO_FSB(mp, SZ_64K));

> +	error = xfs_trans_alloc(mp, &tres, 0, 0, 0, &tp);
>  	if (error)
>  		return error;
>  	xfs_ilock(ip, XFS_ILOCK_EXCL);
>  	xfs_trans_ijoin(tp, ip, XFS_ILOCK_EXCL);
>  
>  	error = xfs_trans_get_buf(tp, mp->m_ddev_targp,
> -			XFS_FSB_TO_DADDR(mp, fsbno), mp->m_bsize, 0, &bp);
> +			XFS_FSB_TO_DADDR(mp, fsbno), BTOBB(len), 0, &bp);
>  	if (error) {
>  		xfs_trans_cancel(tp);
>  		return error;
>  	}
>  
> +	if (xfs_has_rtgroups(mp))
> +		copylen -= sizeof(struct xfs_rtbuf_blkinfo);

This is how we maintain copylen as the amount of non-header data to copy
out of *data, correct?  I suppose that means that the checksum file
blocks also have a header?

--D

> +
>  	xfs_rtfile_initialize_buf(rtg, type, bp, tp);
> -	if (data)
> -		memcpy(xfs_rtblock_payload(bp), data, copylen);
> -	else
> +	if (*data) {
> +		memcpy(xfs_rtblock_payload(bp), *data, copylen);
> +		*data += copylen;
> +	} else {
>  		memset(xfs_rtblock_payload(bp), 0, copylen);
> -	xfs_trans_log_buf(tp, bp, 0, mp->m_sb.sb_blocksize - 1);
> +	}
> +	xfs_trans_log_buf(tp, bp, 0, len - 1);
>  	return xfs_trans_commit(tp);
>  }
>  
> @@ -1447,33 +1457,31 @@ xfs_rtfile_initialize_blocks(
>  	enum xfs_rtg_inodes	type,
>  	xfs_fileoff_t		offset_fsb,	/* offset to start from */
>  	xfs_fileoff_t		end_fsb,	/* offset to allocate to */
> +	xfs_filblks_t		bsize,
> +	uint32_t		bmapi_flags,
>  	void			*data)		/* data to fill the blocks */
>  {
> -	struct xfs_mount	*mp = rtg_mount(rtg);
> -	const size_t		copylen = mp->m_blockwsize << XFS_WORDLOG;
> -
>  	while (offset_fsb < end_fsb) {
>  		struct xfs_bmbt_irec	map;
>  		xfs_filblks_t		i;
>  		int			error;
>  
>  		error = xfs_rtfile_alloc_blocks(rtg->rtg_inodes[type],
> -				offset_fsb, end_fsb - offset_fsb, &map);
> +				offset_fsb, end_fsb - offset_fsb, bmapi_flags,
> +				&map);
>  		if (error)
>  			return error;
>  
>  		/*
> -		 * Now we need to clear the allocated blocks.
> +		 * Now we need to clear or initialize the allocated blocks.
>  		 *
>  		 * Do this one block per transaction, to keep it simple.
>  		 */
> -		for (i = 0; i < map.br_blockcount; i++) {
> +		for (i = 0; i < map.br_blockcount; i += bsize) {
>  			error = xfs_rtfile_initialize_block(rtg, type,
> -					map.br_startblock + i, data);
> +					map.br_startblock + i, bsize, &data);
>  			if (error)
>  				return error;
> -			if (data)
> -				data += copylen;
>  		}
>  
>  		offset_fsb = map.br_startoff + map.br_blockcount;
> diff --git a/fs/xfs/libxfs/xfs_rtbitmap.h b/fs/xfs/libxfs/xfs_rtbitmap.h
> index 750d74fbf4ed..e9e3378d15aa 100644
> --- a/fs/xfs/libxfs/xfs_rtbitmap.h
> +++ b/fs/xfs/libxfs/xfs_rtbitmap.h
> @@ -410,7 +410,8 @@ void xfs_rtfile_initialize_buf(struct xfs_rtgroup *rtg,
>  		struct xfs_trans *tp);
>  int xfs_rtfile_initialize_blocks(struct xfs_rtgroup *rtg,
>  		enum xfs_rtg_inodes type, xfs_fileoff_t offset_fsb,
> -		xfs_fileoff_t end_fsb, void *data);
> +		xfs_fileoff_t end_fsb, xfs_filblks_t bsize,
> +		uint32_t bmapi_flags, void *data);
>  int xfs_rtbitmap_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
>  		struct xfs_trans *tp, bool init);
>  int xfs_rtsummary_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
> diff --git a/fs/xfs/xfs_rtalloc.c b/fs/xfs/xfs_rtalloc.c
> index 84efe5a8fb11..78a1c066c7eb 100644
> --- a/fs/xfs/xfs_rtalloc.c
> +++ b/fs/xfs/xfs_rtalloc.c
> @@ -1191,11 +1191,11 @@ xfs_growfs_rt_alloc_blocks(
>  	}
>  
>  	error = xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_BITMAP, orbmblocks,
> -			nmp->m_sb.sb_rbmblocks, NULL);
> +			nmp->m_sb.sb_rbmblocks, 1, 0, NULL);
>  	if (error)
>  		goto out_free;
>  	error = xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_SUMMARY, orsumblocks,
> -			nmp->m_rsumblocks, NULL);
> +			nmp->m_rsumblocks, 1, 0, NULL);
>  out_free:
>  	kfree(nmp);
>  	return error;
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 08/21] xfs: define the RT data checksum on-disk format
  2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
@ 2026-09-24 22:13   ` Darrick J. Wong
  2026-09-25  0:04     ` Eric Biggers
  2026-09-25  6:01     ` Christoph Hellwig
  0 siblings, 2 replies; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 22:13 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:40AM +0200, Christoph Hellwig wrote:
> Add the on-disk format for the new RT data checksum format.
> 
> Keyed off a new read-only compat feature flag, this adds new fields to
> the superblock to indicate the checksum algorithm used and the size of
> the blocks containing the checksums.  These new fields reuse the
> previously reserved padding to make efficient use of the space in the
> on-disk superblock.
> 
> Data checksums are only supported on zoned RT devices, because they
> require out of places writes to safely update the checksums for file
> overwrites and a data/metadata split to be able to store the checksums
> for a group in a file without causing recursion.  This means they can't
> be supported directly on the data device at all, and only when using
> the always_cow mode on regular RT devices, but that has no benefit
> over the zoned allocator which is designed for out of place writes.
> 
> The initially supported data checksum algorithms are crc32c and crc64 as
> specified by NVMe.  Both have extremely fast kernel implementations and
> the strong data protection guarantees offered by CRC-style algorithms.
> Both also happen to be support by NVMe for protection information so that
> the userspace PI passthrough support (once extended to files on file
> systems) can be reused to expose the checksums to applications and thus
> provide true end-to-end data integrity.

And just to play the peanut gallery here, xxhash? <duck>

> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/libxfs/xfs_format.h     | 43 +++++++++++++++++--
>  fs/xfs/libxfs/xfs_log_format.h |  1 +
>  fs/xfs/libxfs/xfs_ondisk.h     |  4 +-
>  fs/xfs/libxfs/xfs_sb.c         | 75 ++++++++++++++++++++++++++++++++++
>  fs/xfs/libxfs/xfs_sb.h         |  1 +
>  fs/xfs/scrub/agheader.c        |  5 +++
>  fs/xfs/xfs_mount.h             |  7 ++++
>  7 files changed, 132 insertions(+), 4 deletions(-)
> 
> diff --git a/fs/xfs/libxfs/xfs_format.h b/fs/xfs/libxfs/xfs_format.h
> index dd0ed046fbe9..1be3d21910a7 100644
> --- a/fs/xfs/libxfs/xfs_format.h
> +++ b/fs/xfs/libxfs/xfs_format.h
> @@ -179,7 +179,9 @@ typedef struct xfs_sb {
>  	xfs_rgnumber_t	sb_rgcount;	/* number of realtime groups */
>  	xfs_rtxlen_t	sb_rgextents;	/* size of a realtime group in rtx */
>  	uint8_t		sb_rgblklog;    /* rt group number shift */
> -	uint8_t		sb_pad[7];	/* zeroes */
> +	uint8_t		sb_rtcsum_type;	/* RT device data checksum type */
> +	uint8_t		sb_rtcsum_blklog; /* log2 of rtcsum bsize */
> +	uint8_t		sb_pad[5];	/* zero */
>  	xfs_rfsblock_t	sb_rtstart;	/* start of internal RT section (FSB) */
>  	xfs_filblks_t	sb_rtreserved;	/* reserved (zoned) RT blocks */
>  
> @@ -272,7 +274,9 @@ struct xfs_dsb {
>  	__be32		sb_rgcount;	/* # of realtime groups */
>  	__be32		sb_rgextents;	/* size of rtgroup in rtx */
>  	__u8		sb_rgblklog;    /* rt group number shift */
> -	__u8		sb_pad[7];	/* zeroes */
> +	__u8		sb_rtcsum_type;	/* RT device data checksum type */
> +	__u8		sb_rtcsum_blklog; /* log2 of rtcsum bsize */
> +	__u8		sb_pad[5];	/* zero */
>  	__be64		sb_rtstart;	/* start of internal RT section (FSB) */
>  	__be64		sb_rtreserved;	/* reserved (zoned) RT blocks */
>  
> @@ -374,6 +378,8 @@ xfs_sb_has_compat_feature(
>  #define XFS_SB_FEAT_RO_COMPAT_RMAPBT   (1 << 1)		/* reverse map btree */
>  #define XFS_SB_FEAT_RO_COMPAT_REFLINK  (1 << 2)		/* reflinked files */
>  #define XFS_SB_FEAT_RO_COMPAT_INOBTCNT (1 << 3)		/* inobt block counts */
> +#define XFS_SB_FEAT_RO_COMPAT_RTCSUM     (1 << 5)	/* RT data checksums */
> +
>  #define XFS_SB_FEAT_RO_COMPAT_ALL \
>  		(XFS_SB_FEAT_RO_COMPAT_FINOBT | \
>  		 XFS_SB_FEAT_RO_COMPAT_RMAPBT | \
> @@ -866,6 +872,7 @@ enum xfs_metafile_type {
>  	XFS_METAFILE_RTSUMMARY,		/* rt summary */
>  	XFS_METAFILE_RTRMAP,		/* rt rmap */
>  	XFS_METAFILE_RTREFCOUNT,	/* rt refcount */
> +	XFS_METAFILE_RTCSUM,		/* rt data checksums */
>  
>  	XFS_METAFILE_MAX
>  } __packed;
> @@ -879,7 +886,8 @@ enum xfs_metafile_type {
>  	{ XFS_METAFILE_RTBITMAP,	"rtbitmap" }, \
>  	{ XFS_METAFILE_RTSUMMARY,	"rtsummary" }, \
>  	{ XFS_METAFILE_RTRMAP,		"rtrmap" }, \
> -	{ XFS_METAFILE_RTREFCOUNT,	"rtrefcount" }
> +	{ XFS_METAFILE_RTREFCOUNT,	"rtrefcount" }, \
> +	{ XFS_METAFILE_RTCSUM,		"rtcsum", }
>  
>  /*
>   * On-disk inode structure.
> @@ -1318,6 +1326,7 @@ static inline bool xfs_dinode_is_metadir(const struct xfs_dinode *dip)
>   */
>  #define XFS_RTBITMAP_MAGIC	0x424D505A	/* BMPZ */
>  #define XFS_RTSUMMARY_MAGIC	0x53554D59	/* SUMY */
> +#define XFS_RTCSUM_MAGIC	0x4353554D	/* CSUM */
>  
>  struct xfs_rtbuf_blkinfo {
>  	__be32		rt_magic;	/* validity check on block */
> @@ -2027,4 +2036,32 @@ struct xfs_acl {
>  #define SGI_ACL_FILE_SIZE	(sizeof(SGI_ACL_FILE)-1)
>  #define SGI_ACL_DEFAULT_SIZE	(sizeof(SGI_ACL_DEFAULT)-1)
>  
> +/*
> + * Size of a RT data checksum block.  Data reads must be contained in a single
> + * block, so this should be fairly large.
> + *
> + * The default is 32k, matching the default inode cluster size and the maximum
> + * memory allocation the Linux MM can handle in the fast path.  64k is primarily
> + * there so that his value never needs to be below the FSB size, even for 64k

                    this

> + * blocks.
> + */
> +#define XFS_RTCSUM_BSIZE_LOG_MIN	15
> +#define XFS_RTCSUM_BSIZE_LOG_MAX	16
> +
> +/*
> + * Data checksum types.
> + */
> +#define XFS_CSUM_TYPE_NONE	0u
> +#define XFS_CSUM_TYPE_CRC32C	1u
> +#define XFS_CSUM_TYPE_CRC64	2u
> +#define XFS_CSUM_TYPE_MAX	3u
> +
> +/*
> + * On-disk data checksums.
> + */
> +union xfs_disk_csum {
> +	__le32			crc32c;
> +	__le64			crc64;
> +};

I'm curious where this type will lead since its size is 8 bytes...
but we'll see.

> +
>  #endif /* __XFS_FORMAT_H__ */
> diff --git a/fs/xfs/libxfs/xfs_log_format.h b/fs/xfs/libxfs/xfs_log_format.h
> index a4e1b3eb425c..b1037b77338b 100644
> --- a/fs/xfs/libxfs/xfs_log_format.h
> +++ b/fs/xfs/libxfs/xfs_log_format.h
> @@ -581,6 +581,7 @@ enum xfs_blft {
>  	XFS_BLFT_SB_BUF,
>  	XFS_BLFT_RTBITMAP_BUF,
>  	XFS_BLFT_RTSUMMARY_BUF,
> +	XFS_BLFT_RTCSUM_BUF,
>  	XFS_BLFT_MAX_BUF = (1 << XFS_BLFT_BITS),
>  };
>  
> diff --git a/fs/xfs/libxfs/xfs_ondisk.h b/fs/xfs/libxfs/xfs_ondisk.h
> index 23cde1248f01..17ab9366b3b9 100644
> --- a/fs/xfs/libxfs/xfs_ondisk.h
> +++ b/fs/xfs/libxfs/xfs_ondisk.h
> @@ -284,7 +284,9 @@ xfs_check_ondisk_structs(void)
>  	XFS_CHECK_SB_OFFSET(sb_rgcount,			272);
>  	XFS_CHECK_SB_OFFSET(sb_rgextents,		276);
>  	XFS_CHECK_SB_OFFSET(sb_rgblklog,		280);
> -	XFS_CHECK_SB_OFFSET(sb_pad,			281);
> +	XFS_CHECK_SB_OFFSET(sb_rtcsum_type,		281);
> +	XFS_CHECK_SB_OFFSET(sb_rtcsum_blklog,		282);
> +	XFS_CHECK_SB_OFFSET(sb_pad,			283);
>  	XFS_CHECK_SB_OFFSET(sb_rtstart,			288);
>  	XFS_CHECK_SB_OFFSET(sb_rtreserved,		296);
>  
> diff --git a/fs/xfs/libxfs/xfs_sb.c b/fs/xfs/libxfs/xfs_sb.c
> index f0341adbb879..3a470aec6c0c 100644
> --- a/fs/xfs/libxfs/xfs_sb.c
> +++ b/fs/xfs/libxfs/xfs_sb.c
> @@ -487,6 +487,40 @@ xfs_validate_sb_zoned(
>  	return 0;
>  }
>  
> +static int
> +xfs_validate_sb_csum(
> +	struct xfs_mount	*mp,
> +	struct xfs_sb		*sbp)
> +{
> +	unsigned int		rtcsum_bsize = 1u << sbp->sb_rtcsum_blklog;
> +
> +	if (!(sbp->sb_features_incompat & XFS_SB_FEAT_INCOMPAT_ZONED)) {
> +		xfs_warn(mp, "data checksum required the zone allocator");
> +		return -EINVAL;
> +	}
> +	if (sbp->sb_rtcsum_type >= XFS_CSUM_TYPE_MAX) {
> +		xfs_warn(mp, "invalid data checksum type: 0x%x",
> +			sbp->sb_rtcsum_type);
> +		return -EINVAL;
> +	}
> +	if (sbp->sb_rtcsum_blklog < XFS_RTCSUM_BSIZE_LOG_MIN ||
> +	    sbp->sb_rtcsum_blklog > XFS_RTCSUM_BSIZE_LOG_MAX) {
> +		xfs_warn(mp,
> +"invalid data checksum block log: %u (min %u/max %u)",
> +			sbp->sb_rtcsum_blklog,
> +			XFS_RTCSUM_BSIZE_LOG_MIN,
> +			XFS_RTCSUM_BSIZE_LOG_MAX);
> +		return -EINVAL;
> +	}
> +	if (rtcsum_bsize < sbp->sb_blocksize) {
> +		xfs_warn(mp,
> +"checksum block size must not be smaller than file system block size: %u/%u",
> +			rtcsum_bsize, sbp->sb_blocksize);
> +		return -EINVAL;
> +	}
> +	return 0;
> +}
> +
>  /* Check the validity of the SB. */
>  STATIC int
>  xfs_validate_sb_common(
> @@ -580,6 +614,17 @@ xfs_validate_sb_common(
>  			if (error)
>  				return error;
>  		}
> +		if (sbp->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
> +			error = xfs_validate_sb_csum(mp, sbp);
> +			if (error)
> +				return error;
> +		} else {
> +			if (sbp->sb_rtcsum_type || sbp->sb_rtcsum_blklog) {
> +				xfs_warn(mp,
> +"rtcsum superblock fields must be zero for non-RTCSUM file systems.");
> +				return -EINVAL;
> +			}
> +		}
>  	} else if (sbp->sb_qflags & (XFS_PQUOTA_ENFD | XFS_GQUOTA_ENFD |
>  				XFS_PQUOTA_CHKD | XFS_GQUOTA_CHKD)) {
>  			xfs_notice(mp,
> @@ -900,6 +945,14 @@ __xfs_sb_from_disk(
>  		to->sb_rtstart = 0;
>  		to->sb_rtreserved = 0;
>  	}
> +
> +	if (to->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
> +		to->sb_rtcsum_type = from->sb_rtcsum_type;
> +		to->sb_rtcsum_blklog = from->sb_rtcsum_blklog;
> +	} else {
> +		to->sb_rtcsum_type = XFS_CSUM_TYPE_NONE;
> +		to->sb_rtcsum_blklog = 0;
> +	}
>  }
>  
>  void
> @@ -1071,6 +1124,11 @@ xfs_sb_to_disk(
>  		to->sb_rtstart = cpu_to_be64(from->sb_rtstart);
>  		to->sb_rtreserved = cpu_to_be64(from->sb_rtreserved);
>  	}
> +
> +	if (from->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
> +		to->sb_rtcsum_type = from->sb_rtcsum_type;
> +		to->sb_rtcsum_blklog = from->sb_rtcsum_blklog;
> +	}
>  }
>  
>  /*
> @@ -1243,6 +1301,20 @@ xfs_mount_sb_set_rextsize(
>  	xfs_sb_mount_rextsize(mp, sbp);
>  }
>  
> +uint8_t
> +xfs_data_csum_shift(
> +	uint8_t			csum)
> +{
> +	switch (csum) {
> +	case XFS_CSUM_TYPE_CRC32C:
> +		return 2;
> +	case XFS_CSUM_TYPE_CRC64:
> +		return 3;
> +	default:
> +		return 0;
> +	}
> +}
> +
>  /*
>   * xfs_mount_common
>   *
> @@ -1311,6 +1383,9 @@ xfs_sb_mount_common(
>  	mp->m_bsize = XFS_FSB_TO_BB(mp, 1);
>  	mp->m_alloc_set_aside = xfs_alloc_set_aside(mp);
>  	mp->m_ag_max_usable = xfs_alloc_ag_max_usable(mp);
> +
> +	mp->m_rtcsum_shift = xfs_data_csum_shift(mp->m_sb.sb_rtcsum_type);
> +	mp->m_rtcsum_bsize = 1u << mp->m_sb.sb_rtcsum_blklog;
>  }
>  
>  /*
> diff --git a/fs/xfs/libxfs/xfs_sb.h b/fs/xfs/libxfs/xfs_sb.h
> index 34d0dd374e9b..16f300c12e37 100644
> --- a/fs/xfs/libxfs/xfs_sb.h
> +++ b/fs/xfs/libxfs/xfs_sb.h
> @@ -20,6 +20,7 @@ extern void	xfs_sb_mount_common(struct xfs_mount *mp, struct xfs_sb *sbp);
>  void		xfs_sb_mount_rextsize(struct xfs_mount *mp, struct xfs_sb *sbp);
>  void		xfs_mount_sb_set_rextsize(struct xfs_mount *mp,
>  			struct xfs_sb *sbp, xfs_agblock_t rextsize);
> +uint8_t		xfs_data_csum_shift(uint8_t csum);
>  extern void	xfs_sb_from_disk(struct xfs_sb *to, struct xfs_dsb *from);
>  extern void	xfs_sb_to_disk(struct xfs_dsb *to, struct xfs_sb *from);
>  extern void	xfs_sb_quota_from_disk(struct xfs_sb *sbp);
> diff --git a/fs/xfs/scrub/agheader.c b/fs/xfs/scrub/agheader.c
> index 1fa66aa68e16..316a3085f95e 100644
> --- a/fs/xfs/scrub/agheader.c
> +++ b/fs/xfs/scrub/agheader.c
> @@ -416,6 +416,11 @@ xchk_superblock(
>  
>  		if (memchr_inv(sb->sb_pad, 0, sizeof(sb->sb_pad)))
>  			xchk_block_set_corrupt(sc, bp);
> +
> +		if (sb->sb_rtcsum_type != mp->m_sb.sb_rtcsum_type)
> +			xchk_block_set_corrupt(sc, bp);
> +		if (sb->sb_rtcsum_blklog != mp->m_sb.sb_rtcsum_blklog)
> +			xchk_block_set_corrupt(sc, bp);
>  	}
>  
>  	/* Everything else must be zero. */
> diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
> index 894ff2f4ecbd..fa86697f463a 100644
> --- a/fs/xfs/xfs_mount.h
> +++ b/fs/xfs/xfs_mount.h
> @@ -191,6 +191,8 @@ typedef struct xfs_mount {
>  	uint8_t			m_agno_log;	/* log #ag's */
>  	uint8_t			m_sectbb_log;	/* sectlog - BBSHIFT */
>  	int8_t			m_rtxblklog;	/* log2 of rextsize, if possible */
> +	uint8_t			m_rtcsum_shift;	/* log2 of RT data csum size */
> +	uint32_t		m_rtcsum_bsize;	/* rtcsum block size in bytes */
>  
>  	uint			m_blockmask;	/* sb_blocksize-1 */
>  	uint			m_blockwsize;	/* sb_blocksize in words */
> @@ -461,6 +463,11 @@ __XFS_HAS_FEAT(metadir, METADIR)
>  __XFS_HAS_FEAT(zoned, ZONED)
>  __XFS_HAS_FEAT(nolifetime, NOLIFETIME)
>  
> +static inline bool xfs_has_rtcsum(const struct xfs_mount *mp)
> +{
> +	return mp->m_sb.sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM;

Shouldn't this be an __XFS_HAS_FEAT item that gets set in
xfs_sb_version_to_features?

I think the only place why we need to look at the ondisk superblock
fields are the xfs_sb_version_hasXXXX calls in secondary_sb_whack in
xfs_repair (and a few other places in userspace).

--D

> +}
> +
>  static inline bool xfs_has_rtgroups(const struct xfs_mount *mp)
>  {
>  	/* all metadir file systems also allow rtgroups */
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 09/21] xfs: add support for per-RTG csum files
  2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
@ 2026-09-24 22:24   ` Darrick J. Wong
  2026-09-25  6:10     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 22:24 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:41AM +0200, Christoph Hellwig wrote:
> Add the definitions for another per-RTG file that stores data checksums
> for the RTG.  The file is fully preallocated at mkfs/growfs time, and
> thus bmap lookups for it can be performed without taking locks.
> 
> Checksums are organized in fixed size large (initially 32Kib or 64KiB)
> blocks to reduce the lookup and read overhead compare to using the
> smaller file system block size.
> 
> Each block uses the standard RT file header for self-describing metadata
> and can thus reuse the buf_ops including the verifier.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/Makefile                |   1 +
>  fs/xfs/libxfs/xfs_cksum.h      |   7 +-
>  fs/xfs/libxfs/xfs_health.h     |   4 +-
>  fs/xfs/libxfs/xfs_rtbitmap.c   |  10 +++
>  fs/xfs/libxfs/xfs_rtcsumfile.c |  94 +++++++++++++++++++++
>  fs/xfs/libxfs/xfs_rtcsumfile.h | 150 +++++++++++++++++++++++++++++++++
>  fs/xfs/libxfs/xfs_rtgroup.c    |  10 +++
>  fs/xfs/libxfs/xfs_rtgroup.h    |   6 ++
>  fs/xfs/libxfs/xfs_shared.h     |   1 +
>  fs/xfs/xfs_buf_item_recover.c  |   7 ++
>  fs/xfs/xfs_platform.h          |   1 +
>  fs/xfs/xfs_rtalloc.c           |   7 ++
>  12 files changed, 296 insertions(+), 2 deletions(-)
>  create mode 100644 fs/xfs/libxfs/xfs_rtcsumfile.c
>  create mode 100644 fs/xfs/libxfs/xfs_rtcsumfile.h
> 
> diff --git a/fs/xfs/Makefile b/fs/xfs/Makefile
> index 399a207f2d0e..79ea4136fbba 100644
> --- a/fs/xfs/Makefile
> +++ b/fs/xfs/Makefile
> @@ -64,6 +64,7 @@ xfs-y				+= $(addprefix libxfs/, \
>  xfs-$(CONFIG_XFS_RT)		+= $(addprefix libxfs/, \
>  				   xfs_rtbitmap.o \
>  				   xfs_rtgroup.o \
> +				   xfs_rtcsumfile.o \
>  				   xfs_zones.o \
>  				   )
>  
> diff --git a/fs/xfs/libxfs/xfs_cksum.h b/fs/xfs/libxfs/xfs_cksum.h
> index 999a290cfd72..315b85ff78ae 100644
> --- a/fs/xfs/libxfs/xfs_cksum.h
> +++ b/fs/xfs/libxfs/xfs_cksum.h
> @@ -2,7 +2,12 @@
>  #ifndef _XFS_CKSUM_H
>  #define _XFS_CKSUM_H 1
>  
> -#define XFS_CRC_SEED	(~(uint32_t)0)
> +/*
> + * crc32c() does not include the inversion at the beginning and end, while
> + * crc64_nvme() does.
> + */
> +#define XFS_CRC_SEED		(~(uint32_t)0)
> +#define XFS_CRC64_SEED		0
>  
>  /*
>   * Calculate the intermediate checksum for a buffer that has the CRC field
> diff --git a/fs/xfs/libxfs/xfs_health.h b/fs/xfs/libxfs/xfs_health.h
> index 1d45cf5789e8..349093e75268 100644
> --- a/fs/xfs/libxfs/xfs_health.h
> +++ b/fs/xfs/libxfs/xfs_health.h
> @@ -72,6 +72,7 @@ struct xfs_rtgroup;
>  #define XFS_SICK_RG_SUMMARY	(1 << 2)  /* rt groups summary */
>  #define XFS_SICK_RG_RMAPBT	(1 << 3)  /* reverse mappings */
>  #define XFS_SICK_RG_REFCNTBT	(1 << 4)  /* reference counts */
> +#define XFS_SICK_RG_CSUM	(1 << 5)  /* data checksums */
>  
>  /* Observable health issues for AG metadata. */
>  #define XFS_SICK_AG_SB		(1 << 0)  /* superblock */
> @@ -119,7 +120,8 @@ struct xfs_rtgroup;
>  				 XFS_SICK_RG_BITMAP | \
>  				 XFS_SICK_RG_SUMMARY | \
>  				 XFS_SICK_RG_RMAPBT | \
> -				 XFS_SICK_RG_REFCNTBT)
> +				 XFS_SICK_RG_REFCNTBT | \
> +				 XFS_SICK_RG_CSUM)
>  
>  #define XFS_SICK_AG_PRIMARY	(XFS_SICK_AG_SB | \
>  				 XFS_SICK_AG_AGF | \
> diff --git a/fs/xfs/libxfs/xfs_rtbitmap.c b/fs/xfs/libxfs/xfs_rtbitmap.c
> index db6a22b4506a..12e67e1abd3b 100644
> --- a/fs/xfs/libxfs/xfs_rtbitmap.c
> +++ b/fs/xfs/libxfs/xfs_rtbitmap.c
> @@ -125,9 +125,18 @@ const struct xfs_buf_ops xfs_rtsummary_buf_ops = {
>  	.verify_struct	= xfs_rtbuf_verify,
>  };
>  
> +const struct xfs_buf_ops xfs_rtcsum_buf_ops = {
> +	.name		= "xfs_rtcsum",
> +	.magic		= { 0, cpu_to_be32(XFS_RTCSUM_MAGIC) },
> +	.verify_read	= xfs_rtbuf_verify_read,
> +	.verify_write	= xfs_rtbuf_verify_write,
> +	.verify_struct	= xfs_rtbuf_verify,
> +};
> +
>  static const struct xfs_buf_ops *xfs_rtblock_buf_ops[XFS_RTGI_MAX] = {
>  	[XFS_RTGI_SUMMARY]	= &xfs_rtsummary_buf_ops,
>  	[XFS_RTGI_BITMAP]	= &xfs_rtbitmap_buf_ops,
> +	[XFS_RTGI_CSUM]		= &xfs_rtcsum_buf_ops,
>  };
>  
>  const struct xfs_buf_ops *
> @@ -143,6 +152,7 @@ xfs_rtblock_ops(
>  static enum xfs_blft xfs_rtblock_buf_types[XFS_RTGI_MAX] = {
>  	[XFS_RTGI_SUMMARY]	= XFS_BLFT_RTSUMMARY_BUF,
>  	[XFS_RTGI_BITMAP]	= XFS_BLFT_RTBITMAP_BUF,
> +	[XFS_RTGI_CSUM]		= XFS_BLFT_RTCSUM_BUF,
>  };
>  
>  /* Release cached rt bitmap and summary buffers. */
> diff --git a/fs/xfs/libxfs/xfs_rtcsumfile.c b/fs/xfs/libxfs/xfs_rtcsumfile.c
> new file mode 100644
> index 000000000000..fc71c52beafd
> --- /dev/null
> +++ b/fs/xfs/libxfs/xfs_rtcsumfile.c
> @@ -0,0 +1,94 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * Copyright (c) 2026 Christoph Hellwig.
> + */
> +#include "xfs_platform.h"
> +#include "xfs_fs.h"
> +#include "xfs_format.h"
> +#include "xfs_log_format.h"
> +#include "xfs_shared.h"
> +#include "xfs_trans_resv.h"
> +#include "xfs_bit.h"
> +#include "xfs_mount.h"
> +#include "xfs_inode.h"
> +#include "xfs_bmap.h"
> +#include "xfs_rtbitmap.h"
> +#include "xfs_bmap_btree.h"
> +#include "xfs_trans.h"
> +#include "xfs_error.h"
> +#include "xfs_health.h"
> +#include "xfs_rtcsumfile.h"
> +
> +int
> +xfs_rtcsum_bmap(
> +	struct xfs_rtgroup	*rtg,
> +	xfs_rgblock_t		rgbno,
> +	xfs_daddr_t		*daddr)
> +{
> +	struct xfs_mount	*mp = rtg_mount(rtg);
> +	struct xfs_inode	*csumip = rtg_csum(rtg);
> +	struct xfs_ifork	*ifp = &csumip->i_df;
> +	unsigned int		csum_block = xfs_rgb_to_rtcsumblock(mp, rgbno);
> +	xfs_fileoff_t		start_fsb =
> +		XFS_B_TO_FSB(mp, mp->m_rtcsum_bsize) * csum_block;
> +	struct xfs_iext_cursor	icur;
> +	struct xfs_bmbt_irec	got;
> +
> +	ASSERT(!xfs_need_iread_extents(ifp));
> +
> +	if (XFS_IS_CORRUPT(mp, ifp->if_nextents != 1))
> +		goto sick;
> +
> +	/*
> +	 * We can do an unlocked lookup here because the bmap btree for the
> +	 * csum files is immutable once created.
> +	 */

This might be problematic if we ever want online repair to be able to
rebuild the checksum file data fork at runtime.  That would probably
cause catastrophic loss of file data integrity guarantees, but that
might be better than the filesystem dying.

> +	if (XFS_IS_CORRUPT(mp, !xfs_iext_lookup_extent(csumip, ifp, start_fsb,
> +			&icur, &got)))
> +		goto sick;
> +	if (XFS_IS_CORRUPT(mp, got.br_startoff > start_fsb))
> +		goto sick;
> +
> +	start_fsb -= got.br_startoff;
> +	*daddr = XFS_FSB_TO_DADDR(mp, got.br_startblock + start_fsb);
> +	return 0;
> +sick:
> +	xfs_rtginode_mark_sick(rtg, XFS_RTGI_CSUM);
> +	return -EFSCORRUPTED;
> +}
> +
> +xfs_off_t
> +xfs_rtcsum_file_size(
> +	struct xfs_rtgroup	*rtg)
> +{
> +	struct xfs_mount	*mp = rtg_mount(rtg);
> +	uint64_t		raw_size;
> +
> +	raw_size = (xfs_off_t)rtg_blocks(rtg) << mp->m_rtcsum_shift;
> +	return DIV_ROUND_UP_ULL(raw_size, xfs_rtcsum_payload_size(mp)) *
> +			mp->m_rtcsum_bsize;
> +}
> +
> +int
> +xfs_rtcsum_alloc_blocks(
> +	struct xfs_rtgroup	*rtg)
> +{
> +	struct xfs_mount	*mp = rtg_mount(rtg);
> +
> +	return xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_CSUM, 0,
> +			XFS_B_TO_FSB(mp, rtg_csum(rtg)->i_disk_size),
> +			XFS_B_TO_FSB(mp, mp->m_rtcsum_bsize),
> +			XFS_BMAPI_CONTIG, NULL);
> +}
> +
> +int
> +xfs_rtcsum_create(
> +	struct xfs_rtgroup	*rtg,
> +	struct xfs_inode	*ip,
> +	struct xfs_trans	*tp,
> +	bool			init)
> +{
> +	ip->i_disk_size = xfs_rtcsum_file_size(rtg);
> +	xfs_trans_log_inode(tp, ip, XFS_ILOG_CORE);
> +	return 0;
> +}
> diff --git a/fs/xfs/libxfs/xfs_rtcsumfile.h b/fs/xfs/libxfs/xfs_rtcsumfile.h
> new file mode 100644
> index 000000000000..8aacac05adc3
> --- /dev/null
> +++ b/fs/xfs/libxfs/xfs_rtcsumfile.h
> @@ -0,0 +1,150 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _XFS_RTCSUMFILE_H
> +#define _XFS_RTCSUMFILE_H
> +
> +#include "xfs_rtgroup.h"
> +
> +/*
> + * Maximum size of a checksum buffer for writes, used for the log reservation.
> + */
> +#define XFS_RTCSUM_MAX_WRITE	SZ_64K
> +
> +/*
> + * Size of the actual payload in the RT data checksum block.  This excludes the
> + * self-describing metadata header.
> + */
> +static inline unsigned int
> +xfs_rtcsum_payload_size(
> +	struct xfs_mount	*mp)
> +{
> +	return mp->m_rtcsum_bsize - sizeof(struct xfs_rtbuf_blkinfo);
> +}
> +
> +/* Convert data length in logical blocks to checksum length in bytes. */
> +static inline unsigned int
> +xfs_extlen_to_rtcsum_len(
> +	struct xfs_mount	*mp,
> +	xfs_extlen_t		nb)
> +{
> +	return nb << mp->m_rtcsum_shift;
> +}
> +
> +/* Convert checksum length in bytes to data length in logical blocks. */
> +static inline xfs_extlen_t
> +xfs_rtcsum_len_to_extlen(
> +	struct xfs_mount	*mp,
> +	unsigned int		csum_len)
> +{
> +	return csum_len >> mp->m_rtcsum_shift;
> +}
> +
> +/* Convert an rgbno to the csum byte position in the csum file. */
> +static inline xfs_off_t
> +xfs_rgb_to_rtcsumpos(

rgbno_to_ ?

> +	struct xfs_mount	*mp,
> +	xfs_rtblock_t		rgbno)

Shouldn't ^^^^ this be an xfs_rgblock_t given the name and comment?

I don't think you support checksums on !rtgroups filesystems, so you'll
never have to deal with a rt block address larger than 2^32.

> +{
> +	return (xfs_off_t)rgbno << mp->m_rtcsum_shift;

Hrm.  What's the difference between this and +xfs_extlen_to_rtcsum_len?
I guess one describes a quantity of fsblocks, whereas another describes
an address?

> +}
> +
> +/* Convert an rgbno to the csum block index in the csum file. */
> +static inline unsigned int
> +xfs_rgb_to_rtcsumblock(
> +	struct xfs_mount	*mp,
> +	xfs_rtblock_t		rgbno)
> +{
> +	return div_u64(xfs_rgb_to_rtcsumpos(mp, rgbno),
> +			xfs_rtcsum_payload_size(mp));
> +}
> +
> +/* Convert an rgbno to a the checksum offset within an rt csum block. */
> +static inline unsigned int
> +xfs_rgb_to_rtcsumoff(
> +	struct xfs_mount	*mp,
> +	xfs_rgblock_t		rgbno)
> +{
> +	uint32_t		off;
> +
> +	div_u64_rem(xfs_rgb_to_rtcsumpos(mp, rgbno),
> +			xfs_rtcsum_payload_size(mp), &off);
> +	return sizeof(struct xfs_rtbuf_blkinfo) + off;
> +}
> +
> +/* Convert an rtbno to a the checksum offset within an rt csum block. */
> +static inline unsigned int
> +xfs_rtb_to_rtcsumoff(
> +	struct xfs_mount	*mp,
> +	xfs_rtblock_t		fsbno)
> +{
> +	return xfs_rgb_to_rtcsumoff(mp, xfs_rtb_to_rgbno(mp, fsbno));
> +}
> +
> +int xfs_rtcsum_bmap(struct xfs_rtgroup *rtg, xfs_rgblock_t rgbno,
> +		xfs_daddr_t *daddr);
> +xfs_off_t xfs_rtcsum_file_size(struct xfs_rtgroup *rtg);
> +int xfs_rtcsum_alloc_blocks(struct xfs_rtgroup *rtg);
> +int xfs_rtcsum_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
> +		struct xfs_trans *tp, bool init);
> +
> +static inline unsigned int
> +xfs_rtcsum_bufs_per_rtg(
> +	struct xfs_rtgroup	*rtg)
> +{
> +	struct xfs_mount	*mp = rtg_mount(rtg);
> +
> +	if (xfs_has_rtcsum(mp))
> +		return div_u64(xfs_rtcsum_file_size(rtg), mp->m_rtcsum_bsize);
> +	return 0;
> +}
> +
> +static inline xfs_filblks_t
> +xfs_rtcsum_max_len(
> +	struct xfs_mount	*mp,
> +	xfs_fsblock_t		fsbno)
> +{
> +	return xfs_rtcsum_len_to_extlen(mp,
> +			mp->m_rtcsum_bsize - xfs_rtb_to_rtcsumoff(mp, fsbno));
> +}

What does this do?  Does it compute the number of checksums you can
write to the rest of a single csum file block given a fsblock address?

> +
> +union xfs_csum {
> +	uint32_t		crc32c;
> +	uint64_t		crc64;
> +};
> +
> +static __always_inline void
> +xfs_csum_seed(
> +	struct xfs_mount	*mp,
> +	union xfs_csum		*csum)
> +{
> +	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
> +		csum->crc32c = XFS_CRC_SEED;
> +	else
> +		csum->crc64 = XFS_CRC64_SEED;

I was expecting a switch() here. ;)

> +}
> +
> +static __always_inline void
> +xfs_csum_gen(
> +	struct xfs_mount	*mp,
> +	void			*data,
> +	unsigned int		len,
> +	union xfs_csum		*csum)
> +{
> +	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
> +		csum->crc32c = crc32c(csum->crc32c, data, len);
> +	else
> +		csum->crc64 = crc64_nvme(csum->crc64, data, len);
> +}
> +
> +static __always_inline void
> +xfs_csum_finalize(
> +	struct xfs_mount	*mp,
> +	union xfs_disk_csum	*to,
> +	union xfs_csum		*csum)
> +{
> +	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
> +		to->crc32c = cpu_to_le32(~csum->crc32c);
> +	else
> +		to->crc64 = cpu_to_le64(csum->crc64);
> +}
> +
> +#endif /* _XFS_RTCSUMFILE_H */
> diff --git a/fs/xfs/libxfs/xfs_rtgroup.c b/fs/xfs/libxfs/xfs_rtgroup.c
> index fe7222bbe449..22ad71798b62 100644
> --- a/fs/xfs/libxfs/xfs_rtgroup.c
> +++ b/fs/xfs/libxfs/xfs_rtgroup.c
> @@ -35,6 +35,7 @@
>  #include "xfs_metadir.h"
>  #include "xfs_rtrmap_btree.h"
>  #include "xfs_rtrefcount_btree.h"
> +#include "xfs_rtcsumfile.h"
>  
>  /* Find the first usable fsblock in this rtgroup. */
>  static inline uint32_t
> @@ -394,6 +395,15 @@ static const struct xfs_rtginode_ops xfs_rtginode_ops[XFS_RTGI_MAX] = {
>  		.enabled	= xfs_has_reflink,
>  		.create		= xfs_rtrefcountbt_create,
>  	},
> +	[XFS_RTGI_CSUM] = {
> +		.name		= "csum",
> +		.metafile_type	= XFS_METAFILE_RTCSUM,
> +		.sick		= XFS_SICK_RG_CSUM,
> +		.fmt_mask	= (1U << XFS_DINODE_FMT_EXTENTS) |
> +				  (1U << XFS_DINODE_FMT_BTREE),
> +		.enabled	= xfs_has_rtcsum,
> +		.create		= xfs_rtcsum_create,
> +	},
>  };
>  
>  /* Return the shortname of this rtgroup inode. */
> diff --git a/fs/xfs/libxfs/xfs_rtgroup.h b/fs/xfs/libxfs/xfs_rtgroup.h
> index f26e324f0de3..5cca0f4dd25a 100644
> --- a/fs/xfs/libxfs/xfs_rtgroup.h
> +++ b/fs/xfs/libxfs/xfs_rtgroup.h
> @@ -16,6 +16,7 @@ enum xfs_rtg_inodes {
>  	XFS_RTGI_SUMMARY,	/* allocation summary */
>  	XFS_RTGI_RMAP,		/* rmap btree inode */
>  	XFS_RTGI_REFCOUNT,	/* refcount btree inode */
> +	XFS_RTGI_CSUM,		/* data checksum inode */
>  
>  	XFS_RTGI_MAX,
>  };
> @@ -109,6 +110,11 @@ static inline struct xfs_inode *rtg_refcount(const struct xfs_rtgroup *rtg)
>  	return rtg->rtg_inodes[XFS_RTGI_REFCOUNT];
>  }
>  
> +static inline struct xfs_inode *rtg_csum(const struct xfs_rtgroup *rtg)
> +{
> +	return rtg->rtg_inodes[XFS_RTGI_CSUM];
> +}
> +
>  /* Passive rtgroup references */
>  static inline struct xfs_rtgroup *
>  xfs_rtgroup_get(
> diff --git a/fs/xfs/libxfs/xfs_shared.h b/fs/xfs/libxfs/xfs_shared.h
> index b1e0d9bc1f7d..8a60f2c56fe7 100644
> --- a/fs/xfs/libxfs/xfs_shared.h
> +++ b/fs/xfs/libxfs/xfs_shared.h
> @@ -40,6 +40,7 @@ extern const struct xfs_buf_ops xfs_refcountbt_buf_ops;
>  extern const struct xfs_buf_ops xfs_rmapbt_buf_ops;
>  extern const struct xfs_buf_ops xfs_rtbitmap_buf_ops;
>  extern const struct xfs_buf_ops xfs_rtsummary_buf_ops;
> +extern const struct xfs_buf_ops xfs_rtcsum_buf_ops;
>  extern const struct xfs_buf_ops xfs_rtbuf_ops;
>  extern const struct xfs_buf_ops xfs_rtsb_buf_ops;
>  extern const struct xfs_buf_ops xfs_rtrefcountbt_buf_ops;
> diff --git a/fs/xfs/xfs_buf_item_recover.c b/fs/xfs/xfs_buf_item_recover.c
> index 57929f115055..ed74dcdbe483 100644
> --- a/fs/xfs/xfs_buf_item_recover.c
> +++ b/fs/xfs/xfs_buf_item_recover.c
> @@ -414,6 +414,13 @@ xlog_recover_validate_buf_type(
>  		}
>  		bp->b_ops = xfs_rtblock_ops(mp, XFS_RTGI_SUMMARY);
>  		break;
> +	case XFS_BLFT_RTCSUM_BUF:
> +		if (xfs_has_rtgroups(mp) && magic32 != XFS_RTCSUM_MAGIC) {
> +			warnmsg = "Bad rtcsum magic!";
> +			break;
> +		}
> +		bp->b_ops = &xfs_rtcsum_buf_ops;
> +		break;
>  #endif /* CONFIG_XFS_RT */
>  	default:
>  		xfs_warn(mp, "Unknown buffer type %d!",
> diff --git a/fs/xfs/xfs_platform.h b/fs/xfs/xfs_platform.h
> index 5d542e95fe44..c478d8935f61 100644
> --- a/fs/xfs/xfs_platform.h
> +++ b/fs/xfs/xfs_platform.h
> @@ -16,6 +16,7 @@
>  #include <linux/slab.h>
>  #include <linux/vmalloc.h>
>  #include <linux/crc32c.h>
> +#include <linux/crc64.h>
>  #include <linux/module.h>
>  #include <linux/mutex.h>
>  #include <linux/file.h>
> diff --git a/fs/xfs/xfs_rtalloc.c b/fs/xfs/xfs_rtalloc.c
> index 78a1c066c7eb..e2b6113772b6 100644
> --- a/fs/xfs/xfs_rtalloc.c
> +++ b/fs/xfs/xfs_rtalloc.c
> @@ -32,6 +32,7 @@
>  #include "xfs_error.h"
>  #include "xfs_trace.h"
>  #include "xfs_rtrefcount_btree.h"
> +#include "xfs_rtcsumfile.h"
>  #include "xfs_reflink.h"
>  #include "xfs_zone_alloc.h"
>  
> @@ -898,6 +899,12 @@ xfs_growfs_rt_zoned(
>  	xfs_rtbxlen_t		freed_rtx;
>  	int			error;
>  
> +	if (xfs_has_rtcsum(mp)) {
> +		error = xfs_rtcsum_alloc_blocks(rtg);
> +		if (error)
> +			return error;
> +	}
> +
>  	/*
>  	 * Calculate new sb and mount fields for this round.  Also ensure the
>  	 * rtg_extents value is uptodate as the rtbitmap code relies on it.
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers
  2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
@ 2026-09-24 22:30   ` Darrick J. Wong
  2026-09-25  6:12     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-24 22:30 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:42AM +0200, Christoph Hellwig wrote:
> The data checksum is logged in its own transaction, and only logs
> transactions buffers.  While the maximum size of a checksummed data write
> is the same as that of a single checksum buffer, the Zone Append based
> zoned write path can't guarantee alignment, so it might be spread over
> up to three buffers.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/libxfs/xfs_trans_resv.c | 21 +++++++++++++++++++++
>  fs/xfs/libxfs/xfs_trans_resv.h |  4 ++++
>  2 files changed, 25 insertions(+)
> 
> diff --git a/fs/xfs/libxfs/xfs_trans_resv.c b/fs/xfs/libxfs/xfs_trans_resv.c
> index 5382ece51812..8854253afba0 100644
> --- a/fs/xfs/libxfs/xfs_trans_resv.c
> +++ b/fs/xfs/libxfs/xfs_trans_resv.c
> @@ -11,6 +11,7 @@
>  #include "xfs_log_format.h"
>  #include "xfs_trans_resv.h"
>  #include "xfs_mount.h"
> +#include "xfs_rtcsumfile.h"
>  #include "xfs_da_format.h"
>  #include "xfs_da_btree.h"
>  #include "xfs_inode.h"
> @@ -1233,6 +1234,22 @@ xfs_calc_qm_dqalloc_reservation_minlogsize(
>  	return xfs_calc_qm_dqalloc_reservation(mp, true);
>  }
>  
> +/*
> + * Log data checksums for a write.
> + *
> + * Must cover a checksum for each FSB of data written, and the checksums can
> + * span the FSB-sized checksum buffers at both ends.
> + */
> +unsigned int
> +xfs_calc_csum_reservation(
> +	struct xfs_mount	*mp,
> +	unsigned int		csum_len)
> +{
> +	return xfs_calc_buf_res(
> +			howmany(csum_len, xfs_rtcsum_payload_size(mp)) + 1,
> +			mp->m_rtcsum_bsize);
> +}
> +
>  /*
>   * Syncing the incore super block changes to disk.
>   *     the super block to reflect the changes: sector size
> @@ -1354,6 +1371,10 @@ xfs_trans_resv_calc(
>  
>  	xfs_calc_namespace_reservations(mp, resp);
>  
> +	resp->tr_csum.tr_logres =
> +		xfs_calc_csum_reservation(mp, XFS_RTCSUM_MAX_WRITE);
> +	resp->tr_csum.tr_logcount = XFS_DEFAULT_LOG_COUNT;

Going back to a question I had in "xfs: prepare
xfs_rtfile_initialize_blocks for larger than FSB blocks", should we be
using tr_csum for the transaction to write out new checksum file block
contents?

Patch itself looks ok though,
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> +
>  	/*
>  	 * The following transactions are logged in logical format with
>  	 * a default log count.
> diff --git a/fs/xfs/libxfs/xfs_trans_resv.h b/fs/xfs/libxfs/xfs_trans_resv.h
> index 1804e821f382..127db4da31c1 100644
> --- a/fs/xfs/libxfs/xfs_trans_resv.h
> +++ b/fs/xfs/libxfs/xfs_trans_resv.h
> @@ -49,6 +49,7 @@ struct xfs_trans_resv {
>  	struct xfs_trans_res	tr_sb;		/* modify superblock */
>  	struct xfs_trans_res	tr_fsyncts;	/* update timestamps on fsync */
>  	struct xfs_trans_res	tr_atomic_ioend; /* untorn write completion */
> +	struct xfs_trans_res	tr_csum;	/* data checksums in metafile */
>  };
>  
>  /* shorthand way of accessing reservation structure */
> @@ -122,6 +123,9 @@ unsigned int xfs_calc_itruncate_reservation_minlogsize(struct xfs_mount *mp);
>  unsigned int xfs_calc_write_reservation_minlogsize(struct xfs_mount *mp);
>  unsigned int xfs_calc_qm_dqalloc_reservation_minlogsize(struct xfs_mount *mp);
>  
> +unsigned int xfs_calc_csum_reservation(struct xfs_mount *mp,
> +		unsigned int csum_len);
> +
>  xfs_extlen_t xfs_calc_max_atomic_write_fsblocks(struct xfs_mount *mp);
>  xfs_extlen_t xfs_calc_atomic_write_log_geometry(struct xfs_mount *mp,
>  		xfs_extlen_t blockcount, unsigned int *new_logres);
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-24  9:59 support for RT data checksums Christoph Hellwig
                   ` (20 preceding siblings ...)
  2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
@ 2026-09-24 22:52 ` Dave Chinner
  2026-09-25  6:27   ` Christoph Hellwig
  21 siblings, 1 reply; 69+ messages in thread
From: Dave Chinner @ 2026-09-24 22:52 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:32AM +0200, Christoph Hellwig wrote:
> Hi all,
> 
> data checksums provide an additional safeguard against silent data loss.
> 
> In classic XFS they were hard to support because they need to be
> atomically updated with the written data.  The zoned allocator solves
> that problem because it always writes out of place, and the checksums
> can be committed at the same as the metadata linking the newly written
> file data into place.  In theory, a conventional allocator could be used
> in combination with the always_cow option, but there are few upsides of
> this compared to using the zoned allocator.
> 
> Data checksums are stored in per-realtime group files in the metadir,
> similar to other modern RT metadata.  Unlike the checksum design in btrfs
> or some other file system, the checksums are associated with the
> physical blocks, and not with logical data in files.  This reduces the
> mapping overhead, and significantly reduces the write amplification,
> and also avoids duplicate checksums for reflinked files (although those
> are not yet supported with the zoned allocator anyway).
> 
> The initial version provides two checksums algorithms: crc32c and crc64.
> Both of those are cyclic redundancy check algorithms which provide known
> good detection of bit flips that is better than general purpose hash
> functions.  Both are not cryptographic hashes and thus do not provide any
> kind of protection against intentional tampering with the data.
> The crc32c parameters exactly match those use for xfs metadata checksums,
> and also those used by the default btrfs checksum, and the NVMe PI
> formats using crc32c.  The crc64 parameters exactly match those using
> the NVMe PI formats using crc64.  crc32c provides reasonable assurance
> for today's hardware, but might prove limiting for extremely large data
> sets, crc64 fills that void, but probably warrants using > 4k file system
> block sizes to amortize the overhead.

Ok, so this really needs a design doc to explain how it all works,
what the new on-disk format is, scope, constraints, etc, as the
first patch in the series (i.e. in
Documentation/filesystems/xfs/data_checksum_design.rst) so that we
have high level descriptions of the functionality being implemented.
Stuff like why certain crc alrgorithms are supported, how we can add
new ones in the future, constraints of doing so, how different sized
checksums are cleanly supported, etc will make doing such things
much easier.

There's new buffer and inode locking in transactions, and there's a
whole new buffer cache interface to "read a buffer", and that is
used to open code reading checksum buffers and joining them to a
transaction rather than using the existing xfs_trans_read_buf...()
interfaces.  That in itself needs careful consideration, and clear
justification for why it must be duplicated to stand outside all the
existing BLI/transaction APIs, especially given all the "use the new
async buf read interface to do sync buffer reads" behaviour across
the patchset that could just use the existing interfaces.

I'd also like to have the format of the new on disk log item format
structures clearly documented (because we're going to have to
validate them) at recovery time, and also have a clear explaination
of the data vs metadata ordering algorithms that ensures that
checksums are always valid in crash+recovery situations, especially
w.r.t. data integrity operations like fsync.

These are the sorts of details that we need to get right, and it's
really hard to extract the actual design intent from the code that
implements it to determine if the algorithms, ordering and recovery
strategies are solid.

Hence I'd like to see the actual design documented first, then we
can understand and review algorithms, etc, and then verify the code
matches the described algorithms, behaviours, etc. checksums are all
about data integrity, so I'd really like to be able to understand
how it is supposed to work and where it doesn't work. Being forced
to understand design and implementation descisions and constraints
by reverse engineering disjoint chunks of code is not an efficient
use of reviewer time...

Cheers,

-Dave.
-- 
Dave Chinner
dgc@kernel.org

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 08/21] xfs: define the RT data checksum on-disk format
  2026-09-24 22:13   ` Darrick J. Wong
@ 2026-09-25  0:04     ` Eric Biggers
  2026-09-25  6:01     ` Christoph Hellwig
  1 sibling, 0 replies; 69+ messages in thread
From: Eric Biggers @ 2026-09-25  0:04 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 03:13:59PM -0700, Darrick J. Wong wrote:
> On Thu, Sep 24, 2026 at 11:59:40AM +0200, Christoph Hellwig wrote:
> > Add the on-disk format for the new RT data checksum format.
> > 
> > Keyed off a new read-only compat feature flag, this adds new fields to
> > the superblock to indicate the checksum algorithm used and the size of
> > the blocks containing the checksums.  These new fields reuse the
> > previously reserved padding to make efficient use of the space in the
> > on-disk superblock.
> > 
> > Data checksums are only supported on zoned RT devices, because they
> > require out of places writes to safely update the checksums for file
> > overwrites and a data/metadata split to be able to store the checksums
> > for a group in a file without causing recursion.  This means they can't
> > be supported directly on the data device at all, and only when using
> > the always_cow mode on regular RT devices, but that has no benefit
> > over the zoned allocator which is designed for out of place writes.
> > 
> > The initially supported data checksum algorithms are crc32c and crc64 as
> > specified by NVMe.  Both have extremely fast kernel implementations and
> > the strong data protection guarantees offered by CRC-style algorithms.
> > Both also happen to be support by NVMe for protection information so that
> > the userspace PI passthrough support (once extended to files on file
> > systems) can be reused to expose the checksums to applications and thus
> > provide true end-to-end data integrity.
> 
> And just to play the peanut gallery here, xxhash? <duck>

lib/xxhash.c is kind of outdated and just supports XXH32 and XXH64 with
generic C code.  CRCs of the same width would generally be better.

(CRCs used to be kind of slow.  But with carryless multiplication
instructions, which the kernel uses on most architectures, they're super
fast.  They also have "unlimited" parallelism, unlike XXH32 and XXH64,
which have a data dependency after each 4 words processed.) 

XXH3 could be more competitive and would also offer 128-bit hashes.  But
it seems it would be quite a bit of work to add XXH3 to the kernel.

Unless someone really wants 128-bit checksums for this feature, I think
the two options proposed here (CRC-32C and CRC64-NVME) sound good.

- Eric

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 02/21] iomap: add support for data checksumming
  2026-09-24 21:39   ` Darrick J. Wong
@ 2026-09-25  5:53     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  5:53 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 02:39:02PM -0700, Darrick J. Wong wrote:
> > +void *iomap_csum_alloc(struct iomap_ioend *ioend, u8 csum_shift)
> > +{
> > +	ioend->io_csum_shift = csum_shift;
> > +	ioend->io_csum = __iomap_csum_alloc(ioend);
> > +	return ioend->io_csum;
> > +}
> > +EXPORT_SYMBOL_GPL(iomap_csum_alloc);
> 
> AFAICT, io_csum_inline is a small amount of memory in the ioend itself
> to store checksums for the blocks being written, and io_csum_alloc is a
> dynamically allocated blob if the checksums don't fit inline?

Yes.  And both can only exist for the initial ioend and not a split one.

> And io_csum points to wherever that data lives?
> 
> Hrm.  What if you want to split an ioend, I guess the child's io_csum
> points to somewhere inside io_parent->io_csum?

Yes.

> (Perhaps it would help to point this out in the struct definition?)

I'll find a good place to better document this, thanks.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API
  2026-09-24 21:43   ` Darrick J. Wong
@ 2026-09-25  5:54     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  5:54 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 02:43:41PM -0700, Darrick J. Wong wrote:
> On Thu, Sep 24, 2026 at 11:59:35AM +0200, Christoph Hellwig wrote:
> > Add a new helper that reads a buffer asynchronously.  This is similar
> > to readahead, but doesn't become a no-op under memory or I/O congestion
> > and returns the buffer to be read.
> > 
> > The intended use is to kick off a read of data checksum buffers at
> > roughly the same time as the data read so that they are available
> > in the I/O completion handler.
> 
> Does there need to be a "wait until this async-read buffer reaches
> XBF_DONE" function too?  Or how do callers do that?

Using the new xfs_buf_read_async_wait function added in this patch.
I'll reword the commit log to make it more clear that it exists and needs
to be used.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC
  2026-09-24 21:49   ` Darrick J. Wong
@ 2026-09-25  5:57     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  5:57 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 02:49:39PM -0700, Darrick J. Wong wrote:
> > +	 * pace.  Note that the nvecs calaculation is kept from the regular
> 
> 	                              calculation
> 
> > +	 * look as the buffer item formatting expects it.
> 
> "...is kept from the regular look as the buffer item formatting expects
> it" ?
> 
> I don't understand that.  Is the nvecs calculation kept as the buffer
> item formatting code expects it, even though we're allocating more
> shadow buffer space?

Yes.  nvecs always need to be correct for the actual formatting at any
given time.  The allocation is the worst case one, and will be underused
until we fill the log item (and even then be very slightly underused,
as it'll only used a single region then).

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks
  2026-09-24 22:03   ` Darrick J. Wong
@ 2026-09-25  5:58     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  5:58 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 03:03:11PM -0700, Darrick J. Wong wrote:
> > +	size_t			copylen = len;
> > +	struct xfs_trans_res	tres = M_RES(mp)->tr_growrtzero;
> >  	struct xfs_trans	*tp;
> >  	struct xfs_buf		*bp;
> > -	const size_t		copylen = mp->m_blockwsize << XFS_WORDLOG;
> >  	int			error;
> >  
> > -	error = xfs_trans_alloc(mp, &M_RES(mp)->tr_growrtzero, 0, 0, 0, &tp);
> > +	tres.tr_logres *= nblks;
> 
> Hmm.  Is it safe to multiply the log reservation by an arbitrary
> block count?  I would think we'd want *some* guarantee that we can't
> create a transaction that's larger than the log can support.
> 
> It might suffice to put in a safeguard like:
> 
> 	/* Log should always be able to handle 64k of logged buffers */
> 	ASSERT(nblks <= XFS_B_TO_FSB(mp, SZ_64K));

Sounds good.  Or recalculate the entire reservation, but that feels a bit
wasteful.

> > +	if (xfs_has_rtgroups(mp))
> > +		copylen -= sizeof(struct xfs_rtbuf_blkinfo);
> 
> This is how we maintain copylen as the amount of non-header data to copy
> out of *data, correct? 

Yes.

> I suppose that means that the checksum file blocks also have a header?

Yes.  We need this to see if the checksums are corrupted, and then with
the follow on series potentially read it from another mirror.

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 08/21] xfs: define the RT data checksum on-disk format
  2026-09-24 22:13   ` Darrick J. Wong
  2026-09-25  0:04     ` Eric Biggers
@ 2026-09-25  6:01     ` Christoph Hellwig
  1 sibling, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  6:01 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 03:13:59PM -0700, Darrick J. Wong wrote:
> And just to play the peanut gallery here, xxhash? <duck>

Way too slow, and doesn't provided the nice properties of CRCs.

With the diff below that adds xxh32/xxh64 to the crc benchmark, the
numbers do not look good (AMD Zen5 mobile in my laptop), all in MB/s:

		4k              16k
crc32c          69027           77384
crc64_nvme      65164           71290
xxh32           13822           13695
xxh64           26820           28194

Note that given how the crc kunit is designed, nvme64 even does two
extra inversion per cycle, although that probably doesn't show up
in the numbers.

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 09/21] xfs: add support for per-RTG csum files
  2026-09-24 22:24   ` Darrick J. Wong
@ 2026-09-25  6:10     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  6:10 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 03:24:42PM -0700, Darrick J. Wong wrote:
> > +	if (XFS_IS_CORRUPT(mp, ifp->if_nextents != 1))
> > +		goto sick;
> > +
> > +	/*
> > +	 * We can do an unlocked lookup here because the bmap btree for the
> > +	 * csum files is immutable once created.
> > +	 */
> 
> This might be problematic if we ever want online repair to be able to
> rebuild the checksum file data fork at runtime.  That would probably
> cause catastrophic loss of file data integrity guarantees, but that
> might be better than the filesystem dying.

repair is still on my todo list, but the plan would be to use a static_call
to introduce looking if/when we have to online repair to add locking.

> > +/* Convert an rgbno to the csum byte position in the csum file. */
> > +static inline xfs_off_t
> > +xfs_rgb_to_rtcsumpos(
> 
> rgbno_to_ ?

Yeah.

> 
> > +	struct xfs_mount	*mp,
> > +	xfs_rtblock_t		rgbno)
> 
> Shouldn't ^^^^ this be an xfs_rgblock_t given the name and comment?

Yes.

> I don't think you support checksums on !rtgroups filesystems, so you'll
> never have to deal with a rt block address larger than 2^32.

Yes, this is zoned only and zoned requires rtgroups.  Even if for some
reason we want to support non-zoned, requiring rtgroups is required for
the per-group file infrastucture.

> 
> > +{
> > +	return (xfs_off_t)rgbno << mp->m_rtcsum_shift;
> 
> Hrm.  What's the difference between this and +xfs_extlen_to_rtcsum_len?
> I guess one describes a quantity of fsblocks, whereas another describes
> an address?

Yes, which implies 32 vs 64bit values.
> > +static inline xfs_filblks_t
> > +xfs_rtcsum_max_len(
> > +	struct xfs_mount	*mp,
> > +	xfs_fsblock_t		fsbno)
> > +{
> > +	return xfs_rtcsum_len_to_extlen(mp,
> > +			mp->m_rtcsum_bsize - xfs_rtb_to_rtcsumoff(mp, fsbno));
> > +}
> 
> What does this do?  Does it compute the number of checksums you can
> write to the rest of a single csum file block given a fsblock address?

Yes, although the limiting factor is reads where the code wants to deal
with a single buffer per read.  For writes the data is logged where
we can trivially loop over buffers.

> > +union xfs_csum {
> > +	uint32_t		crc32c;
> > +	uint64_t		crc64;
> > +};
> > +
> > +static __always_inline void
> > +xfs_csum_seed(
> > +	struct xfs_mount	*mp,
> > +	union xfs_csum		*csum)
> > +{
> > +	if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
> > +		csum->crc32c = XFS_CRC_SEED;
> > +	else
> > +		csum->crc64 = XFS_CRC64_SEED;
> 
> I was expecting a switch() here. ;)

That would require coming up with something for the impossible case,
so I'd rather side-step it.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers
  2026-09-24 22:30   ` Darrick J. Wong
@ 2026-09-25  6:12     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  6:12 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Thu, Sep 24, 2026 at 03:30:21PM -0700, Darrick J. Wong wrote:
> > +	resp->tr_csum.tr_logres =
> > +		xfs_calc_csum_reservation(mp, XFS_RTCSUM_MAX_WRITE);
> > +	resp->tr_csum.tr_logcount = XFS_DEFAULT_LOG_COUNT;
> 
> Going back to a question I had in "xfs: prepare
> xfs_rtfile_initialize_blocks for larger than FSB blocks", should we be
> using tr_csum for the transaction to write out new checksum file block
> contents?

They actually are different reservations.  Preparing the file writes
one csumblock at a time, while the write path can write upto three
right now as it only has an absolute size limit, but no alignment
limit.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-24 22:52 ` support for " Dave Chinner
@ 2026-09-25  6:27   ` Christoph Hellwig
  2026-09-27 22:59     ` Dave Chinner
  0 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-25  6:27 UTC (permalink / raw)
  To: Dave Chinner
  Cc: Christoph Hellwig, Carlos Maiolino, Darrick J . Wong, Jens Axboe,
	Christian Brauner, linux-xfs, linux-fsdevel

On Fri, Sep 25, 2026 at 08:52:58AM +1000, Dave Chinner wrote:
> There's new buffer and inode locking in transactions,

Not sure what is new about that.  Tere is a single new transaction,
which logs a single buffer per transaction.  Not exactly new and
dancy.

> and there's a
> whole new buffer cache interface to "read a buffer", and that is
> used to open code reading checksum buffers and joining them to a
> transaction rather than using the existing xfs_trans_read_buf...()
> interfaces.  That in itself needs careful consideration, and clear
> justification for why it must be duplicated to stand outside all the
> existing BLI/transaction APIs, especially given all the "use the new
> async buf read interface to do sync buffer reads" behaviour across
> the patchset that could just use the existing interfaces.

I'm not sure what to make of this.  The paragraph almost reads like
AI slop to me.  The rationale is pretty clear and documented: it
turns two dependent reads into two reads that work in parallel.

> I'd also like to have the format of the new on disk log item format
> structures clearly documented (because we're going to have to

There is no new log item format, it uses the standard buffer log format.
The buffer payload is somewhat new.  It is is the standard rt format,
a xfs_rtbuf_blkinfo followed by the real payload, which is an array
of checksums.

> validate them) at recovery time, and also have a clear explaination
> of the data vs metadata ordering algorithms that ensures that
> checksums are always valid in crash+recovery situations, especially
> w.r.t. data integrity operations like fsync.

I think I explained it pretty well, but happy to repeat it again:
The zoned write path writes data first, and then records bmap, rmap
and used space tacking in the zone from the I/O completion handler.
The rtcsum code builds on that and only logs that csum from that
same I/O completion handler.  I.e. that data must have reached the
device for the code to log it to be even called, and for devices
with volatile write caches the generic cache flushing must work
(it did not until recently, but the verification of this code found
that bug and it is now fixed upstream).


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 11/21] xfs: core RT data checksum support
  2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
@ 2026-09-25 23:20   ` Darrick J. Wong
  2026-09-26  6:13     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:20 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:43AM +0200, Christoph Hellwig wrote:
> Support reading and writing of data checksum buffers, and generating
> and verifying the checksum.
> 
> All data checksum buffers for an open zone are pre-allocated at zone open
> time, so that we never have to read in a partially written buffer as part
> of a data write, which would otherwise impose very expensive seeks and
> stall writes.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/Kconfig          |   1 +
>  fs/xfs/Makefile         |   1 +
>  fs/xfs/xfs_inode.h      |   5 +
>  fs/xfs/xfs_rtcsum.c     | 317 ++++++++++++++++++++++++++++++++++++++++
>  fs/xfs/xfs_rtcsum.h     |  23 +++
>  fs/xfs/xfs_super.c      |  14 ++
>  fs/xfs/xfs_sysfs.c      |   2 +
>  fs/xfs/xfs_zone_alloc.c |  32 +++-
>  fs/xfs/xfs_zone_priv.h  |   8 +
>  9 files changed, 399 insertions(+), 4 deletions(-)
>  create mode 100644 fs/xfs/xfs_rtcsum.c
>  create mode 100644 fs/xfs/xfs_rtcsum.h
> 
> diff --git a/fs/xfs/Kconfig b/fs/xfs/Kconfig
> index b99da294e9a3..424a04b507a6 100644
> --- a/fs/xfs/Kconfig
> +++ b/fs/xfs/Kconfig
> @@ -106,6 +106,7 @@ config XFS_RT
>  	bool "XFS Realtime subvolume support"
>  	depends on XFS_FS
>  	default BLK_DEV_ZONED
> +	select CRC64
>  	help
>  	  If you say Y here you will be able to mount and use XFS filesystems
>  	  which contain a realtime subvolume.  The realtime subvolume is a
> diff --git a/fs/xfs/Makefile b/fs/xfs/Makefile
> index 79ea4136fbba..0d57bf0701ec 100644
> --- a/fs/xfs/Makefile
> +++ b/fs/xfs/Makefile
> @@ -142,6 +142,7 @@ xfs-$(CONFIG_XFS_QUOTA)		+= xfs_dquot.o \
>  
>  # xfs_rtbitmap is shared with libxfs
>  xfs-$(CONFIG_XFS_RT)		+= xfs_rtalloc.o \
> +				   xfs_rtcsum.o \
>  				   xfs_zone_alloc.o \
>  				   xfs_zone_gc.o \
>  				   xfs_zone_info.o \
> diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h
> index 34c1038ebfcd..9ed2fbfe86ff 100644
> --- a/fs/xfs/xfs_inode.h
> +++ b/fs/xfs/xfs_inode.h
> @@ -376,6 +376,11 @@ static inline bool xfs_inode_can_sw_atomic_write(const struct xfs_inode *ip)
>  	return xfs_can_sw_atomic_write(ip->i_mount);
>  }
>  
> +static inline bool xfs_is_rtcsum_inode(const struct xfs_inode *ip)
> +{
> +	return xfs_has_rtcsum(ip->i_mount) && XFS_IS_REALTIME_INODE(ip);
> +}
> +
>  /*
>   * In-core inode flags.
>   */
> diff --git a/fs/xfs/xfs_rtcsum.c b/fs/xfs/xfs_rtcsum.c
> new file mode 100644
> index 000000000000..7cdc5a5029eb
> --- /dev/null
> +++ b/fs/xfs/xfs_rtcsum.c
> @@ -0,0 +1,317 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * Copyright (c) 2026 Christoph Hellwig.
> + */
> +#include "xfs_platform.h"
> +#include "xfs_fs.h"
> +#include "xfs_format.h"
> +#include "xfs_log_format.h"
> +#include "xfs_shared.h"
> +#include "xfs_trans_resv.h"
> +#include "xfs_bit.h"
> +#include "xfs_mount.h"
> +#include "xfs_inode.h"
> +#include "xfs_bmap.h"
> +#include "xfs_rtgroup.h"
> +#include "xfs_rtbitmap.h"
> +#include "xfs_bmap_btree.h"
> +#include "xfs_trans.h"
> +#include "xfs_buf_item.h"
> +#include "xfs_trans_space.h"
> +#include "xfs_error.h"
> +#include "xfs_health.h"
> +#include "xfs_rtcsum.h"
> +#include "xfs_zone_priv.h"
> +#include <linux/iomap.h>
> +
> +static_assert(IOMAP_CSUM_MAX_SIZE <= XFS_RTCSUM_MAX_WRITE);
> +
> +static const char *xfs_data_csum_names[XFS_CSUM_TYPE_MAX] = {
> +	[XFS_CSUM_TYPE_CRC32C]	= "crc32c",
> +	[XFS_CSUM_TYPE_CRC64]	= "crc64",
> +};
> +
> +int
> +xfs_csum_verify(
> +	struct xfs_mount	*mp,
> +	struct bio		*bio,
> +	struct bvec_iter	*iter,
> +	void			*csum_buf,
> +	xfs_fsblock_t		bno,
> +	bool			verbose)
> +{
> +	unsigned int		bsize = mp->m_sb.sb_blocksize;
> +	unsigned int		csum_size = 1u << mp->m_rtcsum_shift;
> +	unsigned int		offset = 0;
> +	union xfs_csum		csum;
> +	union xfs_disk_csum	dsum;
> +
> +	do {
> +		struct bio_vec	bv = mp_bvec_iter_bvec(bio->bi_io_vec, *iter);
> +
> +		if (offset == 0)
> +			xfs_csum_seed(mp, &csum);
> +		bv.bv_len = min(bv.bv_len, bsize - offset);
> +		xfs_csum_gen(mp, bvec_virt(&bv), bv.bv_len, &csum);
> +		offset += bv.bv_len;
> +		if (offset == bsize) {
> +			xfs_csum_finalize(mp, &dsum, &csum);
> +			if (unlikely(memcmp(&dsum, csum_buf, csum_size) != 0))
> +				goto mismatch;
> +			bno++;
> +			csum_buf += csum_size;
> +			offset = 0;
> +		}
> +		bio_advance_iter_single(bio, iter, bv.bv_len);
> +	} while (iter->bi_size);
> +
> +	return 0;
> +
> +mismatch:
> +	if (verbose) {
> +		xfs_warn_ratelimited(mp,
> +"data csum mismatch for rtblock 0x%llx: 0x%*phN (expected 0x%*phN)",
> +			bno, csum_size, &csum, csum_size, csum_buf);
> +	}
> +	return -EIO;
> +}
> +
> +void
> +xfs_csum_generate(
> +	struct xfs_mount	*mp,
> +	struct bio		*bio,
> +	void			*csum_buf)
> +{
> +	struct bvec_iter	iter = bio->bi_iter;
> +	unsigned int		bsize = mp->m_sb.sb_blocksize;
> +	unsigned int		csum_size = 1u << mp->m_rtcsum_shift;
> +	unsigned int		offset = 0;
> +	union xfs_csum		csum;
> +
> +	do {
> +		struct bio_vec	bv = mp_bvec_iter_bvec(bio->bi_io_vec, iter);
> +
> +		if (offset == 0)
> +			xfs_csum_seed(mp, &csum);
> +		bv.bv_len = min(bv.bv_len, bsize - offset);
> +		xfs_csum_gen(mp, bvec_virt(&bv), bv.bv_len, &csum);
> +		offset += bv.bv_len;
> +		if (offset == bsize) {
> +			xfs_csum_finalize(mp, csum_buf, &csum);
> +			csum_buf += csum_size;
> +			offset = 0;
> +		}
> +		bio_advance_iter_single(bio, &iter, bv.bv_len);
> +	} while (iter.bi_size);
> +}
> +
> +/*
> + * Read the checksum buffer for @rtg/@csum_off and return it unlocked.

      Read the checksum buffer for @fsbno?

> + *
> + * We don't need to lock read access to the buffer because checksums will not
> + * change until the @rtg is reset.
> + */
> +int
> +xfs_rtcsum_read_async(
> +	struct xfs_mount	*mp,
> +	xfs_rtblock_t		fsbno,
> +	struct xfs_buf		**bpp)
> +{
> +	xfs_daddr_t		csum_daddr;
> +	struct xfs_rtgroup	*rtg;
> +	int			error;
> +
> +	rtg = xfs_rtgroup_get(mp, xfs_rtb_to_rgno(mp, fsbno));
> +	if (!rtg)
> +		return -EFSCORRUPTED;
> +	error = xfs_rtcsum_bmap(rtg, xfs_rtb_to_rgbno(mp, fsbno), &csum_daddr);
> +	if (!error)
> +		error = xfs_buf_read_async(mp->m_ddev_targp, csum_daddr,
> +				BTOBB(mp->m_rtcsum_bsize), &xfs_rtcsum_buf_ops,
> +				bpp);
> +	xfs_rtgroup_put(rtg);
> +	return error;
> +}
> +
> +/*
> + * When opening a zone for writing, do a speculative buf_get for each csum
> + * buffer.  This ensures we usually have a buffer in-memory when we actually
> + * start writing to it.
> + *
> + * Without this we'd have to read the buffer from disk, as we don't know if
> + * anyone has already written to it by the time we get to the buffer due to
> + * completion reordering.
> + *
> + * When reopening a partially written zone at mount time, just read ahead
> + * the entire csums for the zone.
> + */
> +int
> +xfs_rtcsum_open_zone(
> +	struct xfs_open_zone	*oz)
> +{
> +	struct xfs_rtgroup	*rtg = oz->oz_rtg;
> +	struct xfs_mount	*mp = rtg_mount(rtg);
> +	xfs_rgblock_t		rgbno = 0;
> +	int			i, error;
> +
> +	for (i = 0; i < oz->oz_nr_csum_bufs; i++) {
> +		xfs_daddr_t	csum_daddr;
> +		struct xfs_buf	*bp;
> +
> +		error = xfs_rtcsum_bmap(rtg, rgbno, &csum_daddr);
> +		if (error)
> +			goto out_error;
> +
> +		if (oz->oz_allocated) {
> +			error = xfs_buf_read_async(mp->m_ddev_targp, csum_daddr,
> +					BTOBB(mp->m_rtcsum_bsize),
> +					&xfs_rtcsum_buf_ops, &bp);
> +		} else {
> +			error = xfs_buf_get(mp->m_ddev_targp, csum_daddr,
> +					BTOBB(mp->m_rtcsum_bsize), &bp);
> +			if (!error) {
> +				bp->b_flags = XBF_DONE;
> +				xfs_rtfile_initialize_buf(rtg, XFS_RTGI_CSUM,
> +						bp, NULL);
> +			}
> +			xfs_buf_unlock(bp);
> +		}
> +		if (error)
> +			goto out_error;
> +		oz->oz_csum_bufs[i] = bp;
> +		rgbno += (xfs_rtcsum_payload_size(mp) /
> +			  (1u << mp->m_rtcsum_shift));
> +	}
> +
> +	return 0;
> +
> +out_error:
> +	while (--i >= 0)
> +		xfs_buf_rele(oz->oz_csum_bufs[i]);

/me wonders if this should null out oz_csum_bufs to avoid the
possibility of dangling pointers?  Though I think the only caller will
free the oz if this function returns error so it might not matter much.

--D

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 12/21] xfs: data checksums require stable writes
  2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
@ 2026-09-25 23:21   ` Darrick J. Wong
  0 siblings, 0 replies; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:21 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:44AM +0200, Christoph Hellwig wrote:
> Call mapping_set_stable_writes based on the data checksum flag
> so that the page cache doesn't change data in-flight as that could
> corrupt the checksum.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>

Makes sense that we need to stabilize the pagecache during writeback.
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> ---
>  fs/xfs/xfs_inode.h | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
> 
> diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h
> index 9ed2fbfe86ff..72582ea9afcc 100644
> --- a/fs/xfs/xfs_inode.h
> +++ b/fs/xfs/xfs_inode.h
> @@ -619,7 +619,8 @@ int	xfs_break_layouts(struct inode *inode, uint *iolock,
>  
>  static inline void xfs_update_stable_writes(struct xfs_inode *ip)
>  {
> -	if (bdev_stable_writes(xfs_inode_buftarg(ip)->bt_bdev))
> +	if (xfs_is_rtcsum_inode(ip) ||
> +	    bdev_stable_writes(xfs_inode_buftarg(ip)->bt_bdev))
>  		mapping_set_stable_writes(VFS_I(ip)->i_mapping);
>  	else
>  		mapping_clear_stable_writes(VFS_I(ip)->i_mapping);
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 13/21] xfs: require file system block size alignment when using data checksums
  2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
@ 2026-09-25 23:24   ` Darrick J. Wong
  2026-09-26  6:15     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:24 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:45AM +0200, Christoph Hellwig wrote:
> The checksums cover a whole block, so we can't read or update parts of a
> block.  Report the requirement and enforce it for direct I/O reads.
> Direct I/O writes already require file system block size alignment when
> using the zoned allocator, and buffered I/O never does sub-block I/O.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_file.c  | 14 +++++++++++++-
>  fs/xfs/xfs_ioend.c | 13 +++++++++++--
>  fs/xfs/xfs_iops.c  | 10 +++++++++-
>  3 files changed, 33 insertions(+), 4 deletions(-)
> 
> diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
> index 6f25879b6510..5b25f33527c0 100644
> --- a/fs/xfs/xfs_file.c
> +++ b/fs/xfs/xfs_file.c
> @@ -29,6 +29,7 @@
>  #include "xfs_zone_alloc.h"
>  #include "xfs_error.h"
>  #include "xfs_errortag.h"
> +#include "xfs_rtcsum.h"
>  
>  #include <linux/dax.h>
>  #include <linux/falloc.h>
> @@ -270,9 +271,20 @@ xfs_file_dio_read(
>  	if (ret)
>  		return ret;
>  	if (mapping_stable_writes(iocb->ki_filp->f_mapping)) {
> +		unsigned int		dio_flags = 0;
> +
> +		/*
> +		 * Each checksums covers a whole file system block, and thus
> +		 * sub-fsblock reads are not supported for file systems using
> +		 * data checksums.
> +		 */
> +		if (xfs_is_rtcsum_inode(ip))
> +			dio_flags |= IOMAP_DIO_FSBLOCK_ALIGNED;
>  		ret = iomap_dio_rw(iocb, to, &xfs_read_iomap_ops,
> -				&xfs_dio_read_bounce_ops, 0, NULL, 0);
> +				&xfs_dio_read_bounce_ops, dio_flags, NULL, 0);
>  	} else {
> +		ASSERT(!xfs_is_rtcsum_inode(ip));
> +
>  		ret = iomap_dio_read_simple(iocb, to, xfs_read_iomap_begin);
>  		if (ret == -ENOTBLK)
>  			ret = iomap_dio_rw(iocb, to, &xfs_read_iomap_ops, NULL,
> diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c
> index f0e01ac34de8..54bd0995ac29 100644
> --- a/fs/xfs/xfs_ioend.c
> +++ b/fs/xfs/xfs_ioend.c
> @@ -44,6 +44,15 @@ xfs_bounce_submit_ioend(
>  	submit_bio(&ioend->io_bio);
>  }
>  
> +static unsigned int
> +xfs_read_bounce_minsize(
> +	struct iomap_ioend	*ioend)
> +{
> +	if (xfs_is_rtcsum_inode(XFS_I(ioend->io_inode)))
> +		return i_blocksize(ioend->io_inode);
> +	return bdev_logical_block_size(ioend->io_bio.bi_bdev);
> +}
> +
>  static void
>  xfs_end_bio_bounced(
>  	struct bio		*bio)
> @@ -86,7 +95,7 @@ xfs_read_bounce_and_resubmit(
>  		.bi_offset	= ioend->io_bvec_offset,
>  	};
>  	bio->bi_end_io = xfs_end_bio_bounced;
> -	iomap_bounce_read(ioend, bdev_logical_block_size(bio->bi_bdev),
> +	iomap_bounce_read(ioend, xfs_read_bounce_minsize(ioend),
>  			xfs_bounce_submit_ioend);
>  	memalloc_nofs_restore(nofs_flag);
>  }
> @@ -134,7 +143,7 @@ xfs_ioend_submit_read(
>  	ioend = iomap_init_ioend(inode, bio, file_offset, ioend_flags);
>  	if ((ioend_flags & IOMAP_IOEND_DIRECT) &&
>  	    READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_ALWAYS) {
> -		iomap_bounce_read(ioend, bdev_logical_block_size(bio->bi_bdev),
> +		iomap_bounce_read(ioend, xfs_read_bounce_minsize(ioend),
>  				xfs_bounce_submit_ioend);
>  		return;
>  	}
> diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
> index d1306e723899..a5f01e2e3a67 100644
> --- a/fs/xfs/xfs_iops.c
> +++ b/fs/xfs/xfs_iops.c
> @@ -581,6 +581,15 @@ xfs_report_dioalign(
>  	stat->result_mask |= STATX_DIOALIGN | STATX_DIO_READ_ALIGN;
>  	stat->dio_mem_align = bdev_dma_alignment(bdev) + 1;
>  
> +	/*
> +	 * Each checksums covers a whole file system block, and thus sub-fsblock
> +	 * reads are not supported for file systems using data checksums.
> +	 */
> +	if (xfs_is_rtcsum_inode(ip))
> +		stat->dio_read_offset_align = xfs_inode_alloc_unitsize(ip);

The allocation unit could be larger than the fsblock size, why is it
necessary to have such large directio reads on a checksummed file?
Though come to think of it zoned mode doesn't allow rtextsize > 1fsb
so this question might be hair-splitting.

I think you could reuse xfs_read_bounce_minsize() here.

--D

> +	else
> +		stat->dio_read_offset_align = bdev_logical_block_size(bdev);
> +
>  	/*
>  	 * For COW inodes, we can only perform out of place writes of entire
>  	 * allocation units (blocks or RT extents).
> @@ -591,7 +600,6 @@ xfs_report_dioalign(
>  	 * alignment in dio_offset_align, and the smaller read alignment in
>  	 * dio_read_offset_align.
>  	 */
> -	stat->dio_read_offset_align = bdev_logical_block_size(bdev);
>  	if (xfs_is_cow_inode(ip))
>  		stat->dio_offset_align = xfs_inode_alloc_unitsize(ip);
>  	else
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 11/21] xfs: core RT data checksum support
  2026-09-25 23:20   ` Darrick J. Wong
@ 2026-09-26  6:13     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-26  6:13 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Fri, Sep 25, 2026 at 04:20:20PM -0700, Darrick J. Wong wrote:
> > +/*
> > + * Read the checksum buffer for @rtg/@csum_off and return it unlocked.
> 
>       Read the checksum buffer for @fsbno?

Yeah.  This went forth and back a few times during development, and it
looks like the comment didn't keep up..

> > +out_error:
> > +	while (--i >= 0)
> > +		xfs_buf_rele(oz->oz_csum_bufs[i]);
> 
> /me wonders if this should null out oz_csum_bufs to avoid the
> possibility of dangling pointers?  Though I think the only caller will
> free the oz if this function returns error so it might not matter much.

Sure, that sounds like an easy enough safe guard.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 13/21] xfs: require file system block size alignment when using data checksums
  2026-09-25 23:24   ` Darrick J. Wong
@ 2026-09-26  6:15     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-26  6:15 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Fri, Sep 25, 2026 at 04:24:56PM -0700, Darrick J. Wong wrote:
> > +	/*
> > +	 * Each checksums covers a whole file system block, and thus sub-fsblock
> > +	 * reads are not supported for file systems using data checksums.
> > +	 */
> > +	if (xfs_is_rtcsum_inode(ip))
> > +		stat->dio_read_offset_align = xfs_inode_alloc_unitsize(ip);
> 
> The allocation unit could be larger than the fsblock size, why is it
> necessary to have such large directio reads on a checksummed file?
> Though come to think of it zoned mode doesn't allow rtextsize > 1fsb
> so this question might be hair-splitting.

Exactly, allocsize is always the blocksize here.

> 
> I think you could reuse xfs_read_bounce_minsize() here.

Not a bad idea.  Let's see if our header mess allows doing that as an
inline, as doing function call for it would be a bit silly.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-25  6:27   ` Christoph Hellwig
@ 2026-09-27 22:59     ` Dave Chinner
  2026-09-28  5:24       ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Dave Chinner @ 2026-09-27 22:59 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Fri, Sep 25, 2026 at 08:27:10AM +0200, Christoph Hellwig wrote:
> On Fri, Sep 25, 2026 at 08:52:58AM +1000, Dave Chinner wrote:
> > There's new buffer and inode locking in transactions,
> 
> Not sure what is new about that.  Tere is a single new transaction,
> which logs a single buffer per transaction.  Not exactly new and
> dancy.

You haven't answered any of my concerns - you're just handwaving
them away and....

> > and there's a
> > whole new buffer cache interface to "read a buffer", and that is
> > used to open code reading checksum buffers and joining them to a
> > transaction rather than using the existing xfs_trans_read_buf...()
> > interfaces.  That in itself needs careful consideration, and clear
> > justification for why it must be duplicated to stand outside all the
> > existing BLI/transaction APIs, especially given all the "use the new
> > async buf read interface to do sync buffer reads" behaviour across
> > the patchset that could just use the existing interfaces.
> 
> I'm not sure what to make of this.  The paragraph almost reads like
> AI slop to me.

... calling the concerns of an experienced engineer "AI slop".

I'm so disappointed right now.

You're better than this, Christoph. You know better than to attack
the person instead of addressing the technical concerns they've
raised. Calling the concerns of an experienced engineer "AI Slop" is
also pretty insulting.

You also know this has a chilling effect - how many people are going
to be willing to say anything negative about your code, if all they
get from it is a bunch of insults in return?

I don't tolerate people who behave like this towards me or other
team members anymore. "But I write lots of code" is just not good
enough anymore, Christoph. You need to do better.

FWIW, let's address the "AI Slop" aspect. I don't need to us an LLM
to review code - I know the XFS code base as well as you do,
Christoph.

I hadn't even considered passing your code to a LLM yet to see how
many holes I can poke in it. Everything I wrote document concerns my
own brain raised  when doing an initial high level scan of the
patchset. I always do that first so I don't waste time (or tokens!)
doing detailed review on something that needs deeper rework before
it is anywhere near ready for merge.

That scan raised lots of questions in my head about the change, and
so I documented some of my concerns and asked for more detail about
how all this new functionality is supposed to work to be documented.
Not just for me right now, but so there is a permanent record of the
design for future reference.

Performing critical analysis and asking for more detail to be
documented is not "AI Slop". A LLM would just plow on through with
whatever misunderstanding it started with and produce slop. It is
very much a human behaviour to stop and ask for more context,
detail, documentation, etc to address missing knoweldge before going
any further.

If you don't want meaningful review of this code, then just say so -
don't insult the people who spend their own time tryng to understand
the new functionality well enough to review it...

> The rationale is pretty clear and documented: it
> turns two dependent reads into two reads that work in parallel.

Maybe that is clear to you, but it certainly isn't from my initial
reading of the patches.

I don't have all the unwritten knowledge that is in your head that
you didn't document. I've never discussed this code with you, I have
no background on how it is supposed to work, etc. I am coming in
-cold-.

I have not been able to find clear, logical descriptions of
the algorithms for reading and writing checksums, the data
intregrity protocols, the crash recovery protocols, etc. At this
point, the only way for me to to understand details like this is to
-reverse engineer the design from implementation-.

That's the problem - you've presented an implementation without an
overall design being documented, and then asking people to review
the implementation without any other context. I need to understand
the design before I review the implementation, as that is the only
way I can perform a decent -implementation- review.

Indeed, people often complain that it is too hard to adequately
review complex changes, and this is one of the reasons: we are
presented with an implementation, and -zero- design/architecture
context from which to understand the implementation that has been
presented. We need to fix that.

> > I'd also like to have the format of the new on disk log item format
> > structures clearly documented (because we're going to have to
> 
> There is no new log item format, it uses the standard buffer log format.
> The buffer payload is somewhat new.  It is is the standard rt format,
> a xfs_rtbuf_blkinfo followed by the real payload, which is an array
> of checksums.

Yes, the buffer payload is new, and it uses a new BLF flag that
indicates it contains regions with some new on-disk format. And
there are interactions with fsync and data integrity requirements.
Document them!

> > validate them) at recovery time, and also have a clear explaination
> > of the data vs metadata ordering algorithms that ensures that
> > checksums are always valid in crash+recovery situations, especially
> > w.r.t. data integrity operations like fsync.
> 
> I think I explained it pretty well,

You didn't, and that's the problem I am asking you to address.

> but happy to repeat it again:
> The zoned write path writes data first, and then records bmap, rmap
> and used space tacking in the zone from the I/O completion handler.
> The rtcsum code builds on that and only logs that csum from that
> same I/O completion handler.  I.e. that data must have reached the
> device for the code to log it to be even called, and for devices
> with volatile write caches the generic cache flushing must work
> (it did not until recently, but the verification of this code found
> that bug and it is now fixed upstream).

So document the design in the first patch in the series! In detail,
in the tree alongside the code. That gives all reviewers the
necessary context to understand what comes next in the patchset.

Indeed, I ask this not just for reviewiers, but for anyone coming
along in the future that needs to understand how data checksums work
because they need to triage a bug, add new functionality, need to
optimise for performance, etc. Just because you understand how it
works and what you wrote in the commit messages, it doesn't mean how
it is supposed to work, why it works, or why it was implemented in
the way it was is obvious to anyone else.

Documenting the design helps -everyone-, not just now, but well into
the future as well.

-Dave.
-- 
Dave Chinner
dgc@kernel.org

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-27 22:59     ` Dave Chinner
@ 2026-09-28  5:24       ` Christoph Hellwig
  2026-09-29 14:11         ` Dave Chinner
  0 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-09-28  5:24 UTC (permalink / raw)
  To: Dave Chinner
  Cc: Christoph Hellwig, Carlos Maiolino, Darrick J . Wong, Jens Axboe,
	Christian Brauner, linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 08:59:31AM +1000, Dave Chinner wrote:
> You haven't answered any of my concerns - you're just handwaving
> them away and....
> 
> > > and there's a
> > > whole new buffer cache interface to "read a buffer", and that is
> > > used to open code reading checksum buffers and joining them to a
> > > transaction rather than using the existing xfs_trans_read_buf...()
> > > interfaces.  That in itself needs careful consideration, and clear
> > > justification for why it must be duplicated to stand outside all the
> > > existing BLI/transaction APIs, especially given all the "use the new
> > > async buf read interface to do sync buffer reads" behaviour across
> > > the patchset that could just use the existing interfaces.
> > 
> > I'm not sure what to make of this.  The paragraph almost reads like
> > AI slop to me.
> 
> ... calling the concerns of an experienced engineer "AI slop".

No, I call your meandering writing style slop.

> I'm so disappointed right now.
> 
> You're better than this, Christoph. You know better than to attack
> the person instead of addressing the technical concerns they've
> raised. Calling the concerns of an experienced engineer "AI Slop" is
> also pretty insulting.

Stop this bullshit.   Replay to technical details in the patches if you
want, or wait for the requested document, but don't write weirdly
halluscinated high-level concerns.

> You also know this has a chilling effect - how many people are going
> to be willing to say anything negative about your code, if all they
> get from it is a bunch of insults in return?

They don't that response to technical concerns.  They get detailed
answers like Darrick did.  But that requires actually expressing
technical concerns.

<lots of ramblings snipped>

> > The buffer payload is somewhat new.  It is is the standard rt format,
> > a xfs_rtbuf_blkinfo followed by the real payload, which is an array
> > of checksums.
> 
> Yes, the buffer payload is new, and it uses a new BLF flag that
> indicates it contains regions with some new on-disk format. And
> there are interactions with fsync and data integrity requirements.
> Document them!

No, as explained before and clearly visible even from the full diff
it does not use any new BLF flag.   See why this discussion is so
hard?

> Documenting the design helps -everyone-, not just now, but well into
> the future as well.

And I've not disagree with this.  But next time you think you need one
just request it, and don't generate pages full of rambling and incorrect
text.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 14/21] xfs: add support for reading with data checksums
  2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
@ 2026-09-29  0:42   ` Darrick J. Wong
  2026-10-05 12:59     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  0:42 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:46AM +0200, Christoph Hellwig wrote:
> All reads from files with data checksums have the returned iomaps for
> data blocks limited to be inside a single RT csum file block, so that
> each data read only needs to deal with a single checksum buffer.
> 
> All reads on checksummed files need to use ioends so that the checksum
> can be verified from process context.  The ioend submission path looks
> up the checksum buffer and kicks of an asynchronous read of it.  The
> completion path waits for the buffer if needed and verifies the checksum.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_aops.c  |   4 +-
>  fs/xfs/xfs_ioend.c | 108 +++++++++++++++++++++++++++++++++++++++------
>  fs/xfs/xfs_iomap.c |  23 ++++++++--
>  3 files changed, 114 insertions(+), 21 deletions(-)
> 
> diff --git a/fs/xfs/xfs_aops.c b/fs/xfs/xfs_aops.c
> index c30e688cfc9f..931795316de4 100644
> --- a/fs/xfs/xfs_aops.c
> +++ b/fs/xfs/xfs_aops.c
> @@ -599,9 +599,7 @@ static inline const struct iomap_read_ops *
>  xfs_get_iomap_read_ops(
>  	const struct address_space	*mapping)
>  {
> -	struct xfs_inode		*ip = XFS_I(mapping->host);
> -
> -	if (bdev_has_integrity_csum(xfs_inode_buftarg(ip)->bt_bdev))
> +	if (mapping_stable_writes(mapping))
>  		return &xfs_iomap_read_ops;
>  	return &iomap_bio_read_ops;
>  }
> diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c
> index 54bd0995ac29..7570a1b915c0 100644
> --- a/fs/xfs/xfs_ioend.c
> +++ b/fs/xfs/xfs_ioend.c
> @@ -14,12 +14,72 @@
>  #include "xfs_trace.h"
>  #include "xfs_bmap_util.h"
>  #include "xfs_reflink.h"
> +#include "xfs_rtcsum.h"
>  #include "xfs_zone_alloc.h"
>  #include "xfs_ioend.h"
>  #include "xfs_error.h"
>  #include "xfs_errortag.h"
>  #include <linux/bio-integrity.h>
>  
> +static bool
> +xfs_rtcsum_prepare_read(
> +	struct iomap_ioend	*ioend)
> +{
> +	struct xfs_inode	*ip = XFS_I(ioend->io_inode);
> +	struct xfs_mount	*mp = ip->i_mount;
> +	struct xfs_buf		*bp;
> +	int			error;
> +
> +	error = -EIO;
> +	if (WARN_ON_ONCE(ioend->io_bio.bi_iter.bi_idx))

What does this warning mean?  That we've somehow already advanced the
bvec iterator?

> +		goto fail;
> +
> +	error = xfs_rtcsum_read_async(mp,
> +			xfs_daddr_to_rtb(mp, ioend->io_sector), &bp);
> +	if (error)
> +		goto fail;
> +	ioend->io_private = bp;
> +	return true;
> +
> +fail:
> +	ioend->io_bio.bi_status = errno_to_blk_status(error);
> +	bio_endio(&ioend->io_bio);
> +	return false;
> +}
> +
> +static int
> +xfs_rtcsum_verify_ioend(
> +	struct iomap_ioend	*ioend,
> +	int			error)
> +{
> +	struct xfs_inode	*ip = XFS_I(ioend->io_inode);
> +	struct xfs_mount	*mp = ip->i_mount;
> +	xfs_rtblock_t		bno = xfs_daddr_to_rtb(mp, ioend->io_sector);
> +	unsigned int		bsize = mp->m_sb.sb_blocksize;
> +	struct xfs_buf		*bp = ioend->io_private;
> +	struct bvec_iter	iter = {
> +		.bi_size	= roundup(ioend->io_size, bsize),
> +		.bi_offset	= ioend->io_bvec_offset,
> +	};
> +
> +	/* No bp for early xfs_rtcsum_prepare_read failures. */
> +	if (!bp)
> +		return error;

Is it possible for error to be zero here?

> +
> +	if (error)
> +		goto out_rele;
> +	error = xfs_buf_read_async_wait(bp);
> +	if (error)
> +		goto out_rele;
> +
> +	error = xfs_csum_verify(mp, &ioend->io_bio, &iter,
> +				bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno,
> +				true);

You only need two tab indent here.

> +out_rele:
> +	xfs_buf_rele(bp);
> +	return error;
> +}
> +
>  static void
>  xfs_dio_bounce_end_io(
>  	struct bio		*bio)
> @@ -30,6 +90,9 @@ xfs_dio_bounce_end_io(
>  
>  	if ((ioend->io_flags & IOMAP_IOEND_INTEGRITY) && !bio->bi_status)
>  		error = iomap_ioend_integrity_verify(ioend);
> +	if (xfs_is_rtcsum_inode(XFS_I(ioend->io_inode)))
> +		error = xfs_rtcsum_verify_ioend(ioend, error);
> +
>  	iomap_bounce_read_end_io(ioend, orig_bio, error);
>  }
>  
> @@ -39,6 +102,9 @@ xfs_bounce_submit_ioend(
>  {
>  	if (ioend->io_flags & IOMAP_IOEND_INTEGRITY)
>  		fs_bio_integrity_alloc(&ioend->io_bio);
> +	if (xfs_is_rtcsum_inode(XFS_I(ioend->io_inode)) &&
> +	    !xfs_rtcsum_prepare_read(ioend))
> +		return;
>  	ioend->io_bio.bi_end_io = xfs_dio_bounce_end_io;
>  	bio_set_flag(&ioend->io_bio, BIO_COMPLETE_IN_TASK);
>  	submit_bio(&ioend->io_bio);
> @@ -108,25 +174,36 @@ xfs_end_io_read(
>  	struct xfs_inode	*ip = XFS_I(ioend->io_inode);
>  	struct xfs_mount	*mp = ip->i_mount;
>  	int			error = blk_status_to_errno(bio->bi_status);
> +	bool			is_csum_error = false;
>  
>  	if (!error && (ioend->io_flags & IOMAP_IOEND_INTEGRITY)) {
>  		error = iomap_ioend_integrity_verify(ioend);
> -		if ((ioend->io_flags & IOMAP_IOEND_DIRECT) &&
> -		    READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_LAZY) {
> -			/*
> -			 * We only really need to retry for guard tag errors,
> -			 * but right now we can't distinguish them from other
> -			 * (i.e, reftag) errors.
> -			 */
> -			if (error ||
> -			    XFS_TEST_ERROR(mp, XFS_ERRTAG_BOUNCE_REREAD)) {
> -				xfs_read_bounce_and_resubmit(ioend);
> -				return;
> -			}
> -		}
> +		/*
> +		 * We only really need to retry for guard tag errors, but right
> +		 * now we can't distinguish them from other (i.e, reftag) errors.
> +		 */
> +		if (error)
> +			is_csum_error = true;
>  	}
>  
> -	iomap_finish_ioends(ioend, error);
> +	if (xfs_is_rtcsum_inode(ip)) {
> +		error = xfs_rtcsum_verify_ioend(ioend, error);
> +		if (error && !bio->bi_status)
> +			is_csum_error = true;
> +	}
> +
> +	/*
> +	 * If we saw a checksum failure on a direct I/O read that uses lazy
> +	 * bouncing, resubmit the read using a bounce buffer so that we can
> +	 * guarantee this was not caused by the user corrupting the buffer.
> +	 */
> +	if ((ioend->io_flags & IOMAP_IOEND_DIRECT) &&
> +	    READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_LAZY &&
> +	    (is_csum_error ||
> +	     (!error && XFS_TEST_ERROR(mp, XFS_ERRTAG_BOUNCE_REREAD))))
> +		xfs_read_bounce_and_resubmit(ioend);

Ok, so now we bounce the read on PI verification errors or fs checksum
verification errors.  Makes sense.

--D

> +	else
> +		iomap_finish_ioends(ioend, error);
>  }
>  
>  void
> @@ -148,6 +225,9 @@ xfs_ioend_submit_read(
>  		return;
>  	}
>  
> +	if (xfs_is_rtcsum_inode(ip) && !xfs_rtcsum_prepare_read(ioend))
> +		return;
> +
>  	if (ioend_flags & IOMAP_IOEND_INTEGRITY)
>  		fs_bio_integrity_alloc(bio);
>  	bio->bi_end_io = xfs_end_io_read;
> diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
> index 6701be9325ef..0e4396e52809 100644
> --- a/fs/xfs/xfs_iomap.c
> +++ b/fs/xfs/xfs_iomap.c
> @@ -32,6 +32,7 @@
>  #include "xfs_rtbitmap.h"
>  #include "xfs_icache.h"
>  #include "xfs_zone_alloc.h"
> +#include "xfs_rtcsum.h"
>  
>  #define XFS_ALLOC_ALIGN(mp, off) \
>  	(((off) >> mp->m_allocsize_log) << mp->m_allocsize_log)
> @@ -166,6 +167,8 @@ xfs_bmbt_to_iomap(
>  	}
>  
>  	iomap->validity_cookie = sequence_cookie;
> +	if (xfs_is_rtcsum_inode(ip))
> +		iomap->csum_shift = mp->m_rtcsum_shift;
>  	return 0;
>  }
>  
> @@ -2227,16 +2230,28 @@ xfs_read_iomap_begin(
>  		return error;
>  	error = xfs_bmapi_read(ip, offset_fsb, end_fsb - offset_fsb, &imap,
>  			       &nimaps, 0);
> -	if (!error && ((flags & IOMAP_REPORT) || IS_DAX(inode)))
> +	if (error)
> +		goto out_unlock;
> +
> +	if ((flags & IOMAP_REPORT) || IS_DAX(inode)) {
>  		error = xfs_reflink_trim_around_shared(ip, &imap, &shared);
> +		if (error)
> +			goto out_unlock;
> +	} else if (!isnullstartblock(imap.br_startblock) &&
> +		   xfs_is_rtcsum_inode(ip)) {
> +		imap.br_blockcount = min(imap.br_blockcount,
> +				xfs_rtcsum_max_len(mp, imap.br_startblock));
> +	}
> +
>  	seq = xfs_iomap_inode_sequence(ip, shared ? IOMAP_F_SHARED : 0);
>  	xfs_iunlock(ip, lockmode);
> -
> -	if (error)
> -		return error;
>  	trace_xfs_iomap_found(ip, offset, length, XFS_DATA_FORK, &imap);
>  	return xfs_bmbt_to_iomap(ip, iomap, &imap, flags,
>  				 shared ? IOMAP_F_SHARED : 0, seq);
> +
> +out_unlock:
> +	xfs_iunlock(ip, lockmode);
> +	return error;
>  }
>  
>  static DEFINE_IOMAP_ITER_NEXT(xfs_read_iomap_next, xfs_read_iomap_begin);
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 15/21] xfs: add support for writing with data checksums
  2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
@ 2026-09-29  1:01   ` Darrick J. Wong
  2026-10-05 13:00     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:01 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:47AM +0200, Christoph Hellwig wrote:
> All write to checksummed files need to use ioends so that the checksum
> can be verified from process context.
> 
> The ioend submission path allocates the csum buffer and attaches it to
> the ioend before generating the checksum from the file data.  The I/O
> completion then logs the checksums into the buffers for the LBAs that
> were written.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_file.c       |  7 +++---
>  fs/xfs/xfs_ioend.c      | 38 ++++++++++++++++++++++++++++++++
>  fs/xfs/xfs_ioend.h      |  2 ++
>  fs/xfs/xfs_iomap.c      | 49 ++++++++++++++++++++++++++++++++++++-----
>  fs/xfs/xfs_iomap.h      |  4 +++-
>  fs/xfs/xfs_reflink.c    |  2 +-
>  fs/xfs/xfs_zone_alloc.c |  5 +++++
>  7 files changed, 97 insertions(+), 10 deletions(-)
> 
> diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
> index 5b25f33527c0..a38191760f81 100644
> --- a/fs/xfs/xfs_file.c
> +++ b/fs/xfs/xfs_file.c
> @@ -1073,7 +1073,8 @@ xfs_file_buffered_write(
>  
>  	trace_xfs_file_buffered_write(iocb, from);
>  	ret = iomap_file_buffered_write(iocb, from,
> -			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
> +			&xfs_buffered_write_iomap_ops,
> +			xfs_get_iomap_write_ops(ip),
>  			NULL);
>  
>  	/*
> @@ -1154,8 +1155,8 @@ xfs_file_buffered_write_zoned(
>  retry:
>  	trace_xfs_file_buffered_write(iocb, from);
>  	ret = iomap_file_buffered_write(iocb, from,
> -			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
> -			&ac);
> +			&xfs_buffered_write_iomap_ops,
> +			xfs_get_iomap_write_ops(ip), &ac);
>  	if (ret == -ENOSPC && !cleared_space) {
>  		/*
>  		 * Kick off writeback to convert delalloc space and release the
> diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c
> index 7570a1b915c0..7b9c82e3c449 100644
> --- a/fs/xfs/xfs_ioend.c
> +++ b/fs/xfs/xfs_ioend.c
> @@ -235,6 +235,37 @@ xfs_ioend_submit_read(
>  	submit_bio(bio);
>  }
>  
> +int
> +xfs_ioend_submit_read_sync(
> +	struct bio		*bio,
> +	struct inode		*inode,
> +	loff_t			file_offset,
> +	u16			ioend_flags)
> +{
> +	struct xfs_inode	*ip = XFS_I(inode);
> +	struct iomap_ioend	*ioend;
> +	struct bvec_iter	saved_iter;
> +	int			error;
> +
> +	ASSERT(!(ioend_flags & IOMAP_IOEND_DIRECT));
> +
> +	ioend = iomap_init_ioend(inode, bio, file_offset, ioend_flags);
> +	if (xfs_is_rtcsum_inode(ip) && !xfs_rtcsum_prepare_read(ioend))
> +		return blk_status_to_errno(bio->bi_status);
> +	if (ioend_flags & IOMAP_IOEND_INTEGRITY)
> +		fs_bio_integrity_alloc(bio);
> +	saved_iter = bio->bi_iter;
> +	error = submit_bio_wait(bio);
> +	if (bio_integrity(bio)) {
> +		if (!error)
> +			error = fs_bio_integrity_verify(bio, &saved_iter);
> +		fs_bio_integrity_free(bio);
> +	}
> +	if (xfs_is_rtcsum_inode(ip))
> +		error = xfs_rtcsum_verify_ioend(ioend, error);
> +	return error;
> +}

Hm, synchronous reads to pull in whatever unaligned parts of the
pagecache aren't yet uptodate?

> +
>  static void
>  xfs_end_ioend_write_zoned(
>  	struct iomap_ioend	*ioend)
> @@ -261,6 +292,13 @@ xfs_end_ioend_write_zoned(
>  		goto done;
>  	}
>  
> +	if (xfs_is_rtcsum_inode(ip)) {
> +		error = xfs_rtcsum_log(oz, ioend->io_sector, ioend->io_size,
> +				ioend->io_csum);
> +		if (error)
> +			goto done;
> +	}
> +

Neat that this is all we need to do -- log the computed checksum to
the rtcsum file prior to remapping the new blocks into the data fork.
That's how we take care of the ordering requirements: if the remap
transaction is written to disk, then we know the previous csum update
transaction has already gone out before that.  Right?  And it's harmless
if the csum update makes it to disk but the remap never does, because
the space is now written, nobody can see it yet, and can only be cleared
by zonegc.  Right?

It would be helpful to document this ordering dependency here explicitly
for the benefit of code spelunkers in a few years.

	/*
	 * Log the checksum updates before remapping the newly written
	 * extents into the data fork.  We must commit the csum update
	 * before the remap transaction to satisfy an ordering
	 * requirement that any read after a write must be able to find
	 * the new data and new checksum; or the old data and the old
	 * checksum.  It's harmless if the system fails after the csum
	 * update but before the remap because nobody will ever see the
	 * newly written extent until the next gc cycle.
	 */
	if (xfs_is_rtcsum_inode(ip)) {
		error = xfs_rtcsum_log(...);

(How does that sound?)

I think this looks right, so if the answers to the questions are all
'yes' and you're ok with the comment, then
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

>  	error = xfs_zoned_end_io(ip, ioend->io_offset, ioend->io_size,
>  			ioend->io_sector, oz, NULLFSBLOCK);
>  	if (error)
> diff --git a/fs/xfs/xfs_ioend.h b/fs/xfs/xfs_ioend.h
> index 7c2a1ea3e6ed..f01aa208c61f 100644
> --- a/fs/xfs/xfs_ioend.h
> +++ b/fs/xfs/xfs_ioend.h
> @@ -14,5 +14,7 @@ static inline bool xfs_ioend_is_append(struct iomap_ioend *ioend)
>  void xfs_end_bio(struct bio *bio);
>  void xfs_ioend_submit_read(struct inode *inode, struct bio *bio,
>  		loff_t file_offset, u16 ioend_flags);
> +int xfs_ioend_submit_read_sync(struct bio *bio, struct inode *inode,
> +		loff_t file_offset, u16 ioend_flags);
>  
>  #endif /* __XFS_IOEND_H */
> diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
> index 0e4396e52809..75d02a32a36a 100644
> --- a/fs/xfs/xfs_iomap.c
> +++ b/fs/xfs/xfs_iomap.c
> @@ -33,6 +33,7 @@
>  #include "xfs_icache.h"
>  #include "xfs_zone_alloc.h"
>  #include "xfs_rtcsum.h"
> +#include "xfs_ioend.h"
>  
>  #define XFS_ALLOC_ALIGN(mp, off) \
>  	(((off) >> mp->m_allocsize_log) << mp->m_allocsize_log)
> @@ -93,10 +94,45 @@ xfs_iomap_valid(
>  	return true;
>  }
>  
> -const struct iomap_write_ops xfs_iomap_write_ops = {
> +static const struct iomap_write_ops xfs_iomap_write_ops = {
>  	.iomap_valid		= xfs_iomap_valid,
>  };
>  
> +static int
> +xfs_csum_read_folio_range(
> +	const struct iomap_iter	*iter,
> +	struct folio		*folio,
> +	loff_t			pos,
> +	size_t			len)
> +{
> +	const struct iomap	*srcmap = iomap_iter_srcmap(iter);
> +	unsigned int		ioend_flags = iomap_ioend_flags(&iter->iomap);
> +	struct bio		*bio;
> +	int			error;
> +
> +	bio = bio_alloc_bioset(srcmap->bdev, 1, REQ_OP_READ, GFP_NOFS,
> +			&iomap_ioend_bioset);
> +	bio->bi_iter.bi_sector = iomap_sector(srcmap, pos);
> +	bio_add_folio_nofail(bio, folio, len, offset_in_folio(folio, pos));
> +	error = xfs_ioend_submit_read_sync(bio, iter->inode, pos, ioend_flags);
> +	bio_put(bio);
> +	return error;
> +}
> +
> +static const struct iomap_write_ops xfs_iomap_csum_write_ops = {
> +	.iomap_valid		= xfs_iomap_valid,
> +	.read_folio_range	= xfs_csum_read_folio_range,
> +};
> +
> +const struct iomap_write_ops *
> +xfs_get_iomap_write_ops(
> +	struct xfs_inode	*ip)
> +{
> +	if (xfs_is_rtcsum_inode(ip))
> +		return &xfs_iomap_csum_write_ops;
> +	return &xfs_iomap_write_ops;
> +}
> +
>  int
>  xfs_bmbt_to_iomap(
>  	struct xfs_inode	*ip,
> @@ -1617,6 +1653,9 @@ xfs_zoned_fill_srcmap(
>  	 * There is a data fork mapping, only map until the end of it.
>  	 */
>  	xfs_trim_extent(&smap, offset_fsb, *end_fsb - offset_fsb);
> +	if (xfs_is_rtcsum_inode(ip))
> +		smap.br_blockcount = min(smap.br_blockcount,
> +			xfs_rtcsum_max_len(ip->i_mount, smap.br_startblock));
>  	*end_fsb = min(*end_fsb, smap.br_startoff + smap.br_blockcount);
>  	return xfs_bmbt_to_iomap(ip, srcmap, &smap, flags, 0,
>  			xfs_iomap_inode_sequence(ip, 0));
> @@ -2415,8 +2454,8 @@ xfs_zero_range(
>  		return dax_zero_range(inode, pos, len, did_zero,
>  				      &xfs_dax_write_iomap_ops);
>  	return iomap_zero_range(inode, pos, len, did_zero,
> -			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
> -			ac);
> +			&xfs_buffered_write_iomap_ops,
> +			xfs_get_iomap_write_ops(ip), ac);
>  }
>  
>  int
> @@ -2432,6 +2471,6 @@ xfs_truncate_page(
>  		return dax_truncate_page(inode, pos, did_zero,
>  					&xfs_dax_write_iomap_ops);
>  	return iomap_truncate_page(inode, pos, did_zero,
> -			&xfs_buffered_write_iomap_ops, &xfs_iomap_write_ops,
> -			ac);
> +			&xfs_buffered_write_iomap_ops,
> +			xfs_get_iomap_write_ops(ip), ac);
>  }
> diff --git a/fs/xfs/xfs_iomap.h b/fs/xfs/xfs_iomap.h
> index f2520a9b3a13..bb35e58d31ee 100644
> --- a/fs/xfs/xfs_iomap.h
> +++ b/fs/xfs/xfs_iomap.h
> @@ -43,6 +43,7 @@ xfs_iomap_set_anon_write(
>  	iomap->flags = IOMAP_F_ANON_WRITE | IOMAP_F_DIRTY;
>  	if (bdev_has_integrity_csum(iomap->bdev))
>  		iomap->flags |= IOMAP_F_INTEGRITY;
> +	iomap->csum_shift = ip->i_mount->m_rtcsum_shift;
>  }
>  
>  static inline xfs_filblks_t
> @@ -69,6 +70,8 @@ int xfs_read_iomap_begin(struct inode *inode, loff_t offset,
>  		loff_t length, unsigned flags, struct iomap *iomap,
>  		struct iomap *srcmap);
>  
> +const struct iomap_write_ops *xfs_get_iomap_write_ops(struct xfs_inode *ip);
> +
>  extern const struct iomap_ops xfs_buffered_write_iomap_ops;
>  extern const struct iomap_ops xfs_direct_write_iomap_ops;
>  extern const struct iomap_ops xfs_zoned_direct_write_iomap_ops;
> @@ -77,6 +80,5 @@ extern const struct iomap_ops xfs_seek_iomap_ops;
>  extern const struct iomap_ops xfs_xattr_iomap_ops;
>  extern const struct iomap_ops xfs_dax_write_iomap_ops;
>  extern const struct iomap_ops xfs_atomic_write_cow_iomap_ops;
> -extern const struct iomap_write_ops xfs_iomap_write_ops;
>  
>  #endif /* __XFS_IOMAP_H__*/
> diff --git a/fs/xfs/xfs_reflink.c b/fs/xfs/xfs_reflink.c
> index 480136136635..6edbe12777ac 100644
> --- a/fs/xfs/xfs_reflink.c
> +++ b/fs/xfs/xfs_reflink.c
> @@ -1918,7 +1918,7 @@ xfs_reflink_unshare(
>  	else
>  		error = iomap_file_unshare(inode, offset, len,
>  				&xfs_buffered_write_iomap_ops,
> -				&xfs_iomap_write_ops);
> +				xfs_get_iomap_write_ops(ip));
>  	if (error)
>  		goto out;
>  
> diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c
> index 9d9a713684b9..71cd35352e2e 100644
> --- a/fs/xfs/xfs_zone_alloc.c
> +++ b/fs/xfs/xfs_zone_alloc.c
> @@ -27,6 +27,7 @@
>  #include "xfs_trace.h"
>  #include "xfs_mru_cache.h"
>  #include "xfs_rtcsum.h"
> +#include "xfs_rtcsum.h"
>  #include <linux/bio-integrity.h>
>  
>  static void
> @@ -934,6 +935,10 @@ xfs_zone_alloc_and_submit(
>  
>  	if (ioend->io_flags & IOMAP_IOEND_INTEGRITY)
>  		fs_bio_integrity_generate(&ioend->io_bio);
> +	if (xfs_is_rtcsum_inode(ip)) {
> +		xfs_csum_generate(mp, &ioend->io_bio,
> +				iomap_csum_alloc(ioend, mp->m_rtcsum_shift));
> +	}
>  
>  	/*
>  	 * If we don't have a locally cached zone in this write context, see if
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 16/21] xfs: add data checksum support to zoned garbage collection
  2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
@ 2026-09-29  1:06   ` Darrick J. Wong
  2026-10-05 13:11     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:06 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:48AM +0200, Christoph Hellwig wrote:
> Transfer the checksum from the old location to the new one.  The
> implementation mirrors that of user data reads and writes.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_zone_gc.c | 65 ++++++++++++++++++++++++++++++++++++--------
>  1 file changed, 53 insertions(+), 12 deletions(-)
> 
> diff --git a/fs/xfs/xfs_zone_gc.c b/fs/xfs/xfs_zone_gc.c
> index 5fdcf98a2133..1cebd2ee2136 100644
> --- a/fs/xfs/xfs_zone_gc.c
> +++ b/fs/xfs/xfs_zone_gc.c
> @@ -21,6 +21,7 @@
>  #include "xfs_zone_alloc.h"
>  #include "xfs_zone_priv.h"
>  #include "xfs_zones.h"
> +#include "xfs_rtcsum.h"
>  #include "xfs_trace.h"
>  
>  /*
> @@ -106,6 +107,10 @@ struct xfs_gc_bio {
>  	/* Realtime group currently being reclaimed */
>  	struct xfs_rtgroup		*victim_rtg;
>  
> +	/* Buffer for data checksums */
> +	struct xfs_buf			*csum_bp;
> +	void				*csum_buf;

Where is csum_buf used?

> +
>  	/* Bio used for reads and writes, including the bvec used by it */
>  	struct bio			bio;	/* must be last */
>  };
> @@ -660,6 +665,20 @@ xfs_zone_gc_alloc_blocks(
>  	return true;
>  }
>  
> +static void
> +xfs_zone_gc_free_chunk(
> +	struct xfs_gc_bio	*chunk)
> +{
> +	atomic_dec(&chunk->victim_rtg->rtg_gccount);
> +	xfs_rtgroup_rele(chunk->victim_rtg);
> +	list_del(&chunk->entry);
> +	xfs_open_zone_put(chunk->oz);
> +	if (chunk->csum_bp)
> +		xfs_buf_rele(chunk->csum_bp);
> +	xfs_irele(chunk->ip);
> +	bio_put(&chunk->bio);
> +}
> +
>  static void
>  xfs_zone_gc_add_data(
>  	struct xfs_gc_bio	*chunk)
> @@ -725,6 +744,11 @@ xfs_zone_gc_start_chunk(
>  	if (!xfs_zone_gc_iter_irec(mp, iter, &irec, &ip))
>  		return false;
>  
> +	if (xfs_has_rtcsum(mp)) {
> +		irec.rm_blockcount = min(irec.rm_blockcount,
> +			xfs_rtcsum_max_len(mp, irec.rm_startblock));
> +	}
> +
>  	if (!xfs_zone_gc_alloc_blocks(data, &irec.rm_blockcount, &daddr,
>  			&is_seq)) {
>  		xfs_irele(ip);
> @@ -748,6 +772,7 @@ xfs_zone_gc_start_chunk(
>  	chunk->data = data;
>  	chunk->oz = data->oz;
>  	chunk->victim_rtg = iter->victim_rtg;
> +	chunk->csum_bp = NULL;
>  	atomic_inc(&rtg_group(chunk->victim_rtg)->xg_active_ref);
>  	atomic_inc(&chunk->victim_rtg->rtg_gccount);
>  
> @@ -764,22 +789,22 @@ xfs_zone_gc_start_chunk(
>  	list_add_tail(&chunk->entry, &data->reading);
>  	xfs_zone_gc_iter_advance(iter, irec.rm_blockcount);
>  
> +	if (xfs_is_rtcsum_inode(ip)) {
> +		int error;
> +
> +		error = xfs_rtcsum_read_async(mp, chunk->old_startblock,
> +				&chunk->csum_bp);
> +		if (error) {
> +			xfs_force_shutdown(mp, SHUTDOWN_META_IO_ERROR);
> +			xfs_zone_gc_free_chunk(chunk);
> +			return false;
> +		}

Ok so this starts the csum block read asynchronously, and then submits
the read bio.  chunk->csum_bp is the checksum buffer of the rtcsum file
of the zone we're cleaning out, right?

> +	}
> +
>  	submit_bio(bio);
>  	return true;
>  }
>  
> -static void
> -xfs_zone_gc_free_chunk(
> -	struct xfs_gc_bio	*chunk)
> -{
> -	atomic_dec(&chunk->victim_rtg->rtg_gccount);
> -	xfs_rtgroup_rele(chunk->victim_rtg);
> -	list_del(&chunk->entry);
> -	xfs_open_zone_put(chunk->oz);
> -	xfs_irele(chunk->ip);
> -	bio_put(&chunk->bio);
> -}
> -
>  static void
>  xfs_zone_gc_submit_write(
>  	struct xfs_zone_gc_data	*data,
> @@ -832,6 +857,9 @@ xfs_zone_gc_split_write(
>  	split_chunk->old_startblock = chunk->old_startblock;
>  	split_chunk->new_daddr = chunk->new_daddr;
>  	split_chunk->oz = chunk->oz;
> +	split_chunk->csum_bp = chunk->csum_bp;
> +	if (split_chunk->csum_bp)
> +		xfs_buf_hold(split_chunk->csum_bp);
>  	atomic_inc(&chunk->oz->oz_ref);
>  
>  	split_chunk->victim_rtg = chunk->victim_rtg;
> @@ -919,6 +947,19 @@ xfs_zone_gc_finish_chunk(
>  
>  	if (chunk->is_seq)
>  		chunk->new_daddr = chunk->bio.bi_iter.bi_sector;
> +
> +	if (xfs_is_rtcsum_inode(ip)) {
> +		unsigned int	boff;
> +
> +		error = xfs_buf_read_async_wait(chunk->csum_bp);
> +		if (error)
> +			goto free;
> +
> +		boff = xfs_rtb_to_rtcsumoff(mp, chunk->old_startblock);
> +		error = xfs_rtcsum_log(chunk->oz, chunk->new_daddr, chunk->len,
> +				chunk->csum_bp->b_addr + boff);

...and then write the same checksum to the new daddr using the old
checksum.  There's no explicit validation of the checksum in the garbage
collector itself, but I gather the read ioend did this for us already,
right?

--D

> +	}
> +
>  	error = xfs_zoned_end_io(ip, chunk->offset, chunk->len,
>  			chunk->new_daddr, chunk->oz, chunk->old_startblock);
>  free:
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 17/21] xfs: verify data checksums during media verification
  2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
@ 2026-09-29  1:19   ` Darrick J. Wong
  2026-10-05 13:13     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:19 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:49AM +0200, Christoph Hellwig wrote:
> Wire up reading and verifying data checksums during media verification.
> This is very similar to the file read path in that it kicks of an async
> read for the checksum buffer before reading the data, and then validating
> once both are read in.
> 
> To support this, split the per-read logic in xfs_verify_media into a
> separate helpers for the data checksums vs no checksum cases.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_verify_media.c | 131 +++++++++++++++++++++++++++++++-------
>  1 file changed, 107 insertions(+), 24 deletions(-)
> 
> diff --git a/fs/xfs/xfs_verify_media.c b/fs/xfs/xfs_verify_media.c
> index 71f4d6c832a9..46a19405c9e5 100644
> --- a/fs/xfs/xfs_verify_media.c
> +++ b/fs/xfs/xfs_verify_media.c
> @@ -22,6 +22,7 @@
>  #include "xfs_rtrmap_btree.h"
>  #include "xfs_health.h"
>  #include "xfs_healthmon.h"
> +#include "xfs_rtcsum.h"
>  #include "xfs_trace.h"
>  #include "xfs_verify_media.h"
>  
> @@ -261,6 +262,95 @@ xfs_verify_media_error(
>  	}
>  }
>  
> +static int
> +xfs_submit_verify_bio(
> +	struct xfs_mount	*mp,
> +	struct xfs_verify_media	*me,
> +	struct xfs_buftarg	*btp,
> +	struct folio		*folio,
> +	xfs_daddr_t		*daddr,
> +	uint64_t		*bbcount)
> +{
> +	unsigned int		bio_bbcount;
> +	int			error;
> +
> +	bio_bbcount = min(*bbcount, folio_size(folio) >> SECTOR_SHIFT);
> +	error = bdev_rw_virt(btp->bt_bdev, *daddr, folio_address(folio),
> +			bio_bbcount << SECTOR_SHIFT,
> +			REQ_OP_READ);
> +	if (error) {
> +		xfs_verify_media_error(mp, me, btp, *daddr, bio_bbcount, error);
> +		return 1;
> +	}
> +
> +	*daddr += bio_bbcount;
> +	*bbcount -= bio_bbcount;
> +	return 0;
> +}
> +
> +static int
> +xfs_submit_verify_bio_csum(

What's the return value convention here?  1 for media error, 0 for
success, or negative errno if we failed to issue the read?

> +	struct xfs_mount	*mp,
> +	struct xfs_verify_media	*me,
> +	struct xfs_buftarg	*btp,
> +	struct folio		*folio,
> +	xfs_daddr_t		*daddr,
> +	uint64_t		*bbcount)
> +{
> +	struct xfs_buf		*csum_bp = NULL;
> +	unsigned int		bio_bbcount;
> +	struct bvec_iter	saved_iter;
> +	xfs_fsblock_t		bno, end;
> +	xfs_filblks_t		len;
> +	struct bio		bio;
> +	struct bio_vec		bv;
> +	int			error;
> +
> +	bno = xfs_daddr_to_rtb(mp, *daddr);
> +	end = xfs_daddr_to_rtb(mp, *daddr + *bbcount);

It occurs to me that xfs_daddr_to_rtb rounds its argument down.  So if
you pass in daddr==0 and bbcount==2, you'll get bno==end==0 and do no
verification.  I would hope that callers won't pass in parameters like
that, but who knows?

So I think this should be:

	end = xfs_daddr_to_rtb(mp,
			*daddr + *bbcount + XFS_FSB_TO_BB(mp, 1) - 1);

Which I admit is a bit gross.

--D

> +	len = min(end - bno, XFS_B_TO_FSBT(mp, folio_size(folio)));
> +	len = min(len, xfs_rtcsum_max_len(mp, bno));
> +
> +	error = xfs_rtcsum_read_async(mp, bno, &csum_bp);
> +	if (error)
> +		return error;
> +
> +	*daddr = xfs_rtb_to_daddr(mp, bno);
> +	bio_bbcount = XFS_FSB_TO_BB(mp, len);
> +
> +	bio_init(&bio, btp->bt_bdev, &bv, 1, REQ_OP_READ);
> +	bio.bi_iter.bi_sector = *daddr;
> +	bio_add_folio_nofail(&bio, folio,
> +			min(bio_bbcount << SECTOR_SHIFT, folio_size(folio)), 0);
> +	saved_iter = bio.bi_iter;
> +
> +	error = submit_bio_wait(&bio);
> +	if (error)
> +		goto out_media_error;
> +
> +	error = xfs_buf_read_async_wait(csum_bp);
> +	if (error)
> +		goto out_buf_rele;
> +
> +	error = xfs_csum_verify(mp, &bio, &saved_iter,
> +			csum_bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno,
> +			false);
> +	if (error)
> +		goto out_media_error;
> +
> +	*daddr += bio_bbcount;
> +	*bbcount -= bio_bbcount;
> +
> +out_buf_rele:
> +	xfs_buf_rele(csum_bp);
> +	bio_uninit(&bio);
> +	return error;
> +out_media_error:
> +	xfs_verify_media_error(mp, me, btp, *daddr, bio_bbcount, error);
> +	error = 1;
> +	goto out_buf_rele;
> +}
> +
>  /* Verify the media of an xfs device by submitting read requests to the disk. */
>  static int
>  xfs_verify_media(
> @@ -310,18 +400,13 @@ xfs_verify_media(
>  		return 0;
>  
>  	/*
> -	 * There are three ranges involved here:
> -	 *
> -	 *  - [me->me_start_daddr, me->me_end_daddr) is the range that the
> -	 *    user wants to verify.  end_daddr can be beyond the end of the
> -	 *    disk; we'll constrain it to the end if necessary.
> +	 * [me->me_start_daddr, me->me_end_daddr) is the range that the user
> +	 * wants to verify.  end_daddr can be beyond the end of the disk; we'll
> +	 * constrain it to the end if necessary.
>  	 *
> -	 *  - [daddr, me->me_end_daddr) is the range that we have not yet
> -	 *    verified.  We update daddr after each successful read.
> -	 *    me->me_start_daddr is set to daddr before returning.
> -	 *
> -	 *  - [daddr, daddr + bio_bbcount) is the range that we're currently
> -	 *    verifying.
> +	 * [daddr, me->me_end_daddr) is the range that we have not yet verified.
> +	 * We update daddr after each successful read.  me->me_start_daddr is
> +	 * set to daddr before returning.
>  	 */
>  	daddr = me->me_start_daddr;
>  	bbcount = min_t(sector_t, me->me_end_daddr, btp->bt_nr_sectors) -
> @@ -334,22 +419,20 @@ xfs_verify_media(
>  	trace_xfs_verify_media(mp, me, btp->bt_dev, daddr, bbcount, folio);
>  
>  	for (;;) {
> -		unsigned int	bio_bbcount;
> -
> -		bio_bbcount = min(bbcount, folio_size(folio) >> SECTOR_SHIFT);
> -		error = bdev_rw_virt(btp->bt_bdev, daddr, folio_address(folio),
> -				bio_bbcount << SECTOR_SHIFT,
> -				REQ_OP_READ);
> +		if (IS_ENABLED(CONFIG_XFS_RT) &&
> +		    me->me_dev == XFS_DEV_RT &&
> +		    xfs_has_rtcsum(mp)) {
> +			error = xfs_submit_verify_bio_csum(mp, me, btp, folio,
> +					&daddr, &bbcount);
> +		} else {
> +			error = xfs_submit_verify_bio(mp, me, btp, folio,
> +					&daddr, &bbcount);
> +		}
>  		if (error) {
> -			xfs_verify_media_error(mp, me, btp, daddr, bio_bbcount,
> -					error);
> -			error = 0;
> +			if (error == 1)
> +				error = 0;
>  			break;
>  		}
> -
> -		daddr += bio_bbcount;
> -		bbcount -= bio_bbcount;
> -
>  		if (bbcount == 0)
>  			break;
>  
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 18/21] xfs: don't try to verify checksums on empty zones
  2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
@ 2026-09-29  1:25   ` Darrick J. Wong
  2026-10-05 13:14     ` Christoph Hellwig
  2026-10-08 11:43   ` Anuj gupta
  1 sibling, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:25 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:50AM +0200, Christoph Hellwig wrote:
> xfs_scrub can sometimes send XFS_IOC_VERIFY_MEDIA ioctls for ranges
> that have never been written since the last zone reset, which will
> lead to checksum verification failures.

...and pointless work. :)

> Protect against this by checking that the range is valid.  If the report
> flag is set, a verification failure could lead to health reports and
> the file system being marked corrupt, so lock out zone racing reset
> completions for this case as well.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_verify_media.c | 44 ++++++++++++++++++++++++++++++++++++---
>  1 file changed, 41 insertions(+), 3 deletions(-)
> 
> diff --git a/fs/xfs/xfs_verify_media.c b/fs/xfs/xfs_verify_media.c
> index 46a19405c9e5..fc739f7ffa4e 100644
> --- a/fs/xfs/xfs_verify_media.c
> +++ b/fs/xfs/xfs_verify_media.c
> @@ -25,6 +25,8 @@
>  #include "xfs_rtcsum.h"
>  #include "xfs_trace.h"
>  #include "xfs_verify_media.h"
> +#include "xfs_zone_alloc.h"
> +#include "xfs_zone_priv.h"
>  
>  #include <linux/fserror.h>
>  
> @@ -288,6 +290,43 @@ xfs_submit_verify_bio(
>  	return 0;
>  }
>  
> +static int
> +xfs_csum_verify_metafile(
> +	struct xfs_mount	*mp,
> +	struct bio		*bio,
> +	struct bvec_iter	*saved_iter,
> +	void			*csum_buf,
> +	xfs_fsblock_t		bno)
> +{
> +	struct xfs_rtgroup	*rtg;
> +	int			error = 0;
> +
> +	rtg = xfs_rtgroup_get(mp, xfs_rtb_to_rgno(mp, bno));
> +	if (!rtg)
> +		return -EFSCORRUPTED;
> +
> +	/*
> +	 * Only validate the checksums for valid data, as data never written
> +	 * will not have valid checksums.  We need to hold the ilock on the rmap
> +	 * inode to prevent freeing of blocks and thus a zone reset to happen
> +	 * underneath us.
> +	 *
> +	 * Note that this still relies on cooperating userspace, as there also
> +	 * can be blocks that were written but never recorded after an unclean
> +	 * shutdown, which this check does not catch.  It purely tries to deal
> +	 * with races vs the previous FSMAP output used by xfs_scrub.

Just a general note -- if the csum update never makes it to disk, then
the remap cannot have been written to disk either.  Therefore, fsmap
will report that as unowned space, and xfs_scrub won't try to
read-verify it, given the xfs_scrub patch you post later to only read rt
blocks for which there are rmap records.

So I think this is only a theoretical concern, right?

The code looks fine to me, so
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> +	 */
> +	xfs_ilock(rtg_rmap(rtg), XFS_ILOCK_SHARED);
> +	if (!xa_get_mark(&mp->m_groups[XG_TYPE_RTG].xa, rtg_rgno(rtg),
> +			XFS_RTG_FREE)) {
> +		error = xfs_csum_verify(mp, bio, saved_iter, csum_buf, bno,
> +				false);
> +	}
> +	xfs_iunlock(rtg_rmap(rtg), XFS_ILOCK_SHARED);
> +	xfs_rtgroup_put(rtg);
> +	return error;
> +}
> +
>  static int
>  xfs_submit_verify_bio_csum(
>  	struct xfs_mount	*mp,
> @@ -332,9 +371,8 @@ xfs_submit_verify_bio_csum(
>  	if (error)
>  		goto out_buf_rele;
>  
> -	error = xfs_csum_verify(mp, &bio, &saved_iter,
> -			csum_bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno,
> -			false);
> +	error = xfs_csum_verify_metafile(mp, &bio, &saved_iter,
> +			csum_bp->b_addr + xfs_rtb_to_rtcsumoff(mp, bno), bno);
>  	if (error)
>  		goto out_media_error;
>  
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM
  2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
@ 2026-09-29  1:26   ` Darrick J. Wong
  0 siblings, 0 replies; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:26 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:51AM +0200, Christoph Hellwig wrote:
> Report the RT data checksum flag and the checksum algorithm to userspace.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>

Looks good to me
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> ---
>  fs/xfs/libxfs/xfs_fs.h | 6 +++++-
>  fs/xfs/libxfs/xfs_sb.c | 5 +++++
>  2 files changed, 10 insertions(+), 1 deletion(-)
> 
> diff --git a/fs/xfs/libxfs/xfs_fs.h b/fs/xfs/libxfs/xfs_fs.h
> index 185f09f327c0..afe83ff23ef3 100644
> --- a/fs/xfs/libxfs/xfs_fs.h
> +++ b/fs/xfs/libxfs/xfs_fs.h
> @@ -191,7 +191,10 @@ struct xfs_fsop_geom {
>  	__u32		rgcount;	/* number of realtime groups	*/
>  	__u64		rtstart;	/* start of internal rt section */
>  	__u64		rtreserved;	/* RT (zoned) reserved blocks	*/
> -	__u64		reserved[14];	/* reserved space		*/
> +	__u8		rtcsum_type;	/* RT data checksum type	*/
> +	__u8		rtcsum_blklog;	/* log2 of rtcsum bsize		*/
> +	__u8		reserved_pad[6];/* reserved space		*/
> +	__u64		reserved[13];	/* reserved space		*/
>  };
>  
>  #define XFS_FSOP_GEOM_SICK_COUNTERS	(1 << 0)  /* summary counters */
> @@ -250,6 +253,7 @@ typedef struct xfs_fsop_resblks {
>  #define XFS_FSOP_GEOM_FLAGS_PARENT	(1 << 25) /* linux parent pointers */
>  #define XFS_FSOP_GEOM_FLAGS_METADIR	(1 << 26) /* metadata directories */
>  #define XFS_FSOP_GEOM_FLAGS_ZONED	(1 << 27) /* zoned rt device */
> +#define XFS_FSOP_GEOM_FLAGS_DATA_CSUM	(1 << 28) /* data checksums */
>  
>  /*
>   * Minimum and maximum sizes need for growth checks.
> diff --git a/fs/xfs/libxfs/xfs_sb.c b/fs/xfs/libxfs/xfs_sb.c
> index 3a470aec6c0c..506caf3e5201 100644
> --- a/fs/xfs/libxfs/xfs_sb.c
> +++ b/fs/xfs/libxfs/xfs_sb.c
> @@ -1660,6 +1660,11 @@ xfs_fs_geometry(
>  		geo->flags |= XFS_FSOP_GEOM_FLAGS_METADIR;
>  	if (xfs_has_zoned(mp))
>  		geo->flags |= XFS_FSOP_GEOM_FLAGS_ZONED;
> +	if (xfs_has_rtcsum(mp)) {
> +		geo->flags |= XFS_FSOP_GEOM_FLAGS_DATA_CSUM;
> +		geo->rtcsum_type = mp->m_sb.sb_rtcsum_type;
> +		geo->rtcsum_blklog = mp->m_sb.sb_rtcsum_blklog;
> +	}
>  	geo->rtsectsize = sbp->sb_blocksize;
>  	geo->dirblocksize = xfs_dir2_dirblock_bytes(sbp);
>  
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums
  2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
@ 2026-09-29  1:27   ` Darrick J. Wong
  0 siblings, 0 replies; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:27 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:52AM +0200, Christoph Hellwig wrote:
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
>  fs/xfs/xfs_message.c | 4 ++++
>  fs/xfs/xfs_message.h | 1 +
>  fs/xfs/xfs_mount.h   | 2 ++
>  fs/xfs/xfs_rtcsum.c  | 2 +-
>  4 files changed, 8 insertions(+), 1 deletion(-)
> 
> diff --git a/fs/xfs/xfs_message.c b/fs/xfs/xfs_message.c
> index 0243e509a468..53e8ed976a9b 100644
> --- a/fs/xfs/xfs_message.c
> +++ b/fs/xfs/xfs_message.c
> @@ -149,6 +149,10 @@ xfs_warn_experimental(
>  			.opstate	= XFS_OPSTATE_WARNED_LARP,
>  			.name		= "logged extended attributes",
>  		},
> +		[XFS_EXPERIMENTAL_CSUM] = {
> +			.opstate	= XFS_OPSTATE_WARNED_CSUM,
> +			.name		= "data checksum",

Nit: checksums (plural)?  Seeing as the xfs_info() below also says
checksums.

> +		},
>  	};
>  	ASSERT(feat >= 0 && feat < XFS_EXPERIMENTAL_MAX);
>  	BUILD_BUG_ON(ARRAY_SIZE(features) != XFS_EXPERIMENTAL_MAX);
> diff --git a/fs/xfs/xfs_message.h b/fs/xfs/xfs_message.h
> index 811b885f41c3..d858e93e5426 100644
> --- a/fs/xfs/xfs_message.h
> +++ b/fs/xfs/xfs_message.h
> @@ -93,6 +93,7 @@ void xfs_buf_alert_ratelimited(struct xfs_buf *bp, const char *rlmsg,
>  enum xfs_experimental_feat {
>  	XFS_EXPERIMENTAL_SHRINK,
>  	XFS_EXPERIMENTAL_LARP,
> +	XFS_EXPERIMENTAL_CSUM,
>  
>  	XFS_EXPERIMENTAL_MAX,
>  };
> diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
> index fa86697f463a..c62eb42e490a 100644
> --- a/fs/xfs/xfs_mount.h
> +++ b/fs/xfs/xfs_mount.h
> @@ -586,6 +586,8 @@ __XFS_HAS_FEAT(nouuid, NOUUID)
>   */
>  #define XFS_OPSTATE_BLOCKGC_ENABLED	6
>  
> +/* Kernel has logged a warning about checksums */
> +#define XFS_OPSTATE_WARNED_CSUM		8
>  /* Kernel has logged a warning about shrink being used on this fs. */
>  #define XFS_OPSTATE_WARNED_SHRINK	9
>  /* Kernel has logged a warning about logged xattr updates being used. */
> diff --git a/fs/xfs/xfs_rtcsum.c b/fs/xfs/xfs_rtcsum.c
> index 7cdc5a5029eb..47a7b67f7412 100644
> --- a/fs/xfs/xfs_rtcsum.c
> +++ b/fs/xfs/xfs_rtcsum.c
> @@ -310,8 +310,8 @@ xfs_rtcsum_mount(
>  	}
>  	mp->m_features |= XFS_FEAT_DAX_NEVER;
>  
> +	xfs_warn_experimental(mp, XFS_EXPERIMENTAL_CSUM);
>  	xfs_info(mp, "using %s for RT device data checksums",
>  		xfs_data_csum_names[mp->m_sb.sb_rtcsum_type]);
> -

Extraneous deletion?

With that dealt with,
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 21/21] xfs: enable RT data checksums
  2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
@ 2026-09-29  1:27   ` Darrick J. Wong
  2026-10-05 13:16     ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Darrick J. Wong @ 2026-09-29  1:27 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Jens Axboe, Christian Brauner, linux-xfs,
	linux-fsdevel

On Thu, Sep 24, 2026 at 11:59:53AM +0200, Christoph Hellwig wrote:
> All mandatory pieces are in place now, allow mounting.
> 
> Signed-off-by: Christoph Hellwig <hch@lst.de>

Looks good.  I'll try to cough up an amendment to the ondisk metadata
documentation to go with this.

Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>

--D

> ---
>  fs/xfs/libxfs/xfs_format.h | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
> 
> diff --git a/fs/xfs/libxfs/xfs_format.h b/fs/xfs/libxfs/xfs_format.h
> index 1be3d21910a7..8b269fd3ed9f 100644
> --- a/fs/xfs/libxfs/xfs_format.h
> +++ b/fs/xfs/libxfs/xfs_format.h
> @@ -384,7 +384,8 @@ xfs_sb_has_compat_feature(
>  		(XFS_SB_FEAT_RO_COMPAT_FINOBT | \
>  		 XFS_SB_FEAT_RO_COMPAT_RMAPBT | \
>  		 XFS_SB_FEAT_RO_COMPAT_REFLINK| \
> -		 XFS_SB_FEAT_RO_COMPAT_INOBTCNT)
> +		 XFS_SB_FEAT_RO_COMPAT_INOBTCNT | \
> +		 XFS_SB_FEAT_RO_COMPAT_RTCSUM)
>  #define XFS_SB_FEAT_RO_COMPAT_UNKNOWN	~XFS_SB_FEAT_RO_COMPAT_ALL
>  static inline bool
>  xfs_sb_has_ro_compat_feature(
> -- 
> 2.53.0
> 
> 

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-28  5:24       ` Christoph Hellwig
@ 2026-09-29 14:11         ` Dave Chinner
  2026-09-30  7:11           ` Dave Chinner
  0 siblings, 1 reply; 69+ messages in thread
From: Dave Chinner @ 2026-09-29 14:11 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 07:24:36AM +0200, Christoph Hellwig wrote:
> On Mon, Sep 28, 2026 at 08:59:31AM +1000, Dave Chinner wrote:
> > You haven't answered any of my concerns - you're just handwaving
> > them away and....
> > 
> > > > and there's a
> > > > whole new buffer cache interface to "read a buffer", and that is
> > > > used to open code reading checksum buffers and joining them to a
> > > > transaction rather than using the existing xfs_trans_read_buf...()
> > > > interfaces.  That in itself needs careful consideration, and clear
> > > > justification for why it must be duplicated to stand outside all the
> > > > existing BLI/transaction APIs, especially given all the "use the new
> > > > async buf read interface to do sync buffer reads" behaviour across
> > > > the patchset that could just use the existing interfaces.
> > > 
> > > I'm not sure what to make of this.  The paragraph almost reads like
> > > AI slop to me.
> > 
> > ... calling the concerns of an experienced engineer "AI slop".
> 
> No, I call your meandering writing style slop.

And you think that makes what you said any better? Talk about not
knowing when to stop digging...

How hard is it to say "I don't really understand your concern - can
you clarify what you are concerned about?" instead of calling it
slop?

That's would have been a constructive response, and the discussion
then goes an entirely different (and far more pleasant) way from
there.

There are three lines in the commit message and two in your reply
repeating the same thing: "it's for issuing the checksum block read
in parallel with the data read".

I know that - it's just async readahead of the checksum block
followed by a blocking operation that waits for the readahead to
complete.

What I said above is that this async read IO pattern already exists
in XFS - btree traversals do this same readahead/blocking read
pattern to pull sibling nodes into memory ahead of time as we search
sideways across levels. And they do it within existing transaction
APIs, too.

That is, we already have an API that allows async_read+blocking_read
pairs that wait for async readahead to complete.  It also gathers
errors - if readahead IO fails, the blocking read will reissue the
read IO and gather the error if it fails again.

Hence I'm wanting to know why you chose to duplicate that code and
place it behind a slightly different API with slightly different
semantics instead of just using the existing code with a couple of
small tweaks? Why is this new code better than reusing the existing
code?

My concerns about the checksum buffer transactionsi are about not
using existing APIs. By open coding them, you've skipped all the
transaction/BLI recursion detection. i.e. you've encoded an
assumption that a buffer will never get relogged within a given
transaction. You've also encoded that external values bound
memcpy()s in to the buffer and logging ranges to within the buffer
length, but the buffer length itself is never checked. 

I don't know enough about the design to validate these assumptions,
and there is no documentation (code, comments or commit messages)
describing why the code is safe the way it is written...

We also use wrapper functions to set BLFT type specific BLI flags
(inode bufs, dquot bufs, ordered bufs, etc) along with the BLFT
type. These also include asserts to ensure that we call those
functions appropriately. The csum logging code open codes these
flags/types, and there are no asserts anywhere to indicate incorrect
usage of the new BLI flags that csum buffers use. Why deviate from
the existing BLI patterns and APIs?

IMO, if you're going to implement new transaction and BLI
interactions, you need to document and explain how it all works.
BLI life cycle bugs are still an ongoing source of crashes and UAFs
and I'd really like to make sure that these changes don't make it
impossible to fix the life cycle issues and UAFs they result in.

> > I'm so disappointed right now.
> > 
> > You're better than this, Christoph. You know better than to attack
> > the person instead of addressing the technical concerns they've
> > raised. Calling the concerns of an experienced engineer "AI Slop" is
> > also pretty insulting.
> 
> Stop this bullshit. Replay to technical details in the patches if you
> want, or wait for the requested document, but don't write weirdly
> halluscinated high-level concerns.

Clearly you haven't understood why I'm disappointed in you.  If you
don't want to deal with "this BS", then -don't be an asshole-. End
of story. 

> > Yes, the buffer payload is new, and it uses a new BLF flag that
> > indicates it contains regions with some new on-disk format. And
> > there are interactions with fsync and data integrity requirements.
> > Document them!
> 
> No, as explained before and clearly visible even from the full diff
> it does not use any new BLF flag.   See why this discussion is so
> hard?

I meant "BLFT", not BLF - it's just a simple typo. There's no need
to be an asshole over a simple typo, especially as the typo doesn't
materially change what I said (i.e. BFLTs are BLF flags...)

> > Documenting the design helps -everyone-, not just now, but well into
> > the future as well.
> 
> And I've not disagree with this.

But you also haven't agreed to write a design doc yet, either.

Your previous response was pretty negative towards my request -
saying "I think I explained it pretty well" is a fair indication
that you aren't going to write one.

So, are you going to write a design doc or not?

> But next time you think you need one
> just request it, and don't generate pages full of rambling and incorrect
> text.

Four paragraphs to request a design doc and it's scope it is hardly
"pages full of rambling".  Why are you being so obnoxious about
being asked for a design doc?

-Dave.
-- 
Dave Chinner
dgc@kernel.org

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-29 14:11         ` Dave Chinner
@ 2026-09-30  7:11           ` Dave Chinner
  2026-10-05 13:53             ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Dave Chinner @ 2026-09-30  7:11 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Wed, Sep 30, 2026 at 12:11:32AM +1000, Dave Chinner wrote:
> My concerns about the checksum buffer transactionsi are about not
> using existing APIs. By open coding them, you've skipped all the
> transaction/BLI recursion detection. i.e. you've encoded an
> assumption that a buffer will never get relogged within a given
> transaction. You've also encoded that external values bound
> memcpy()s in to the buffer and logging ranges to within the buffer
> length, but the buffer length itself is never checked. 
> 
> I don't know enough about the design to validate these assumptions,
> and there is no documentation (code, comments or commit messages)
> describing why the code is safe the way it is written...

After a day having this stuff percolate through my brain, I realised
why the CSUM logging code and the BLI_PREALLOC hack just didn't seem
right.

TL;DR: csum record logging should use the ICREATE item + ordered
buffer model for efficient one-shot updates to the csum buffers.

Long story:

+/*
+ * Size of a RT data checksum block.  Data reads must be contained in a single
+ * block, so this should be fairly large.
+ *
+ * The default is 32k, matching the default inode cluster size and the maximum
+ * memory allocation the Linux MM can handle in the fast path.  64k is primarily
+ * there so that his value never needs to be below the FSB size, even for 64k
+ * blocks.
+ */
+#define XFS_RTCSUM_BSIZE_LOG_MIN	15
+#define XFS_RTCSUM_BSIZE_LOG_MAX	16

My initial thought was that on disk format structures shouldn't be
defined by the limitations of the OS memory allocation, but <shrug>.
It kept nagging at me, though.

I looked more closely at what XFS_BLI_PREALLOC did to try to
understand why it existed.  It triggers a max-sized CIL logvec
structure for the buffer object. For a 32kB buffer logged as a
single contiguous range, this ends up being about 32kB + a logvec
header, plus a log iovec, plus a BLF, plus a couple of ophdrs. So
it's about 32kB + 200-250 bytes.

That means the shadow buffer for a csum buffer is always considered
a costly allocation by the MM subsystem.

Shadow buffers are ephmeral - they disappear from the log item
whenever the CIL flushes. Hence as the item is repeatedly logged,
they often need to be reallocated becaus something else caused the
CIL to flush (e.g. a fsync operation).

Hence for csum BLI, if a CIL flush happens on a partially filled
buffer, a good amount of that shadow buffer will go unused. Then we
allocate another (costly) shadow buffer on the next update. If CIL
flushes happen frequently enough then we will be repeatedly doing
costly allocations for shadow buffers that we don't actually use.

Not ideal - I think that means the original "sized for mm fast path"
intent is really only valid for the read side of the csum
algorithms as implemented by the patchset.

Then it got me wondering, because the more I thought about it, the
more BLI_PREALLOC smelt of premature optimisation. CIL formatting
has lots of other overhead, especially when re....

Duh. Relogging.

The root cause of the performance issues is interaction of repeated
small delta updates to the buffer and the BLI relogging algorithms
that the CIL and AIL require for correct journal/metadata writeback
ordering. 

Worst case:  Start with an empty csum buffer, assuming no header for
simplicity, using crc32 (4 bytes) for each 4kB data block. The
relogging pattern looks like for single record updates (small
independent 4kB data writes):

	dirtys		dirtied range 	CIL formatting memcpy
T0	0-3		0-127		128 bytes
T1	4-7		0-127		128 bytes
...
T32	128-131		0-255		256 bytes
T33	132-135		0-255		256 bytes
....
T8191	32764-32767	0-32767		32768 bytes

If BLI_PREALLOC didn't exist, the shadow buffer would be reallocated
every 32 updates (i.e. every time the dirtied range increases). So,
for this worst case, that's a 1024 reallocs. Yes, I can see why
that's an obvious optimisation target.

However, what BLI_PREALLOC misses is the CIL formatting memcpy
overhead from the same relogging algorithm.

In this case, the first csum is copied into the shadow buffer 8192
times, even though it was only updated once.  A quick calculation of
how much csum data is actually copied in this scenario is:

$ val=0; for i in `seq 128 128 32768`; do let val+=$(( 32 * $i )); echo $i $val ; done |tail -1
32768 134742016
$

~128MB of csum data is memcpy()d from the buffer to the BLI logvec
to fully update a 32kB csum buffer.

Yes, that's worst case, but even if we update 32 records at a
time (128kB IOs on 4kb rtblksz), the relogging still copies 4MB of
csum data to fully log that 32kB csum buffer. It hurts, even on
decent sized write IOs.

Ok, we have a solution to this problem. I created ordered buffers
and one-shot log items to avoid the journalling overhead of static
inode buffer initialisation back in 2013. The ICREATE log item is
the one-shot log item that records a buffer should be initialised,
and the ordered buffer allows the modified buffer to be passed
through the journal to metadta writeback without it's contents being
logged.

Given that csum updates are a small, known size, non-overlapping
one-shot update to a buffer, they fit the same model that
ICREAT+ordered implements. Adding a new CSUM log item made up of a
format header and varible size csum payload region provides the
equivalent of the ICREAT item for journalled inode buffer
initialisation.

At runtime: the buffer is updated, ordered and joined to the
transaction. A CSUM item with the csum payload is built and logged.
The transaction commit pins the buffer in memory until the CSUM item
is in the journal, then the buffer gets pushed into the AIL at the
CSUM item LSN and the CSUM item is freed.

Hence we only need to copy csum items once. We don't need to
optimise away the overhead of relogging buffers with a constantly
growing dirty range. The shadow buffer is always sized by the number
of csum records being updated this transaction, and it never just
pushed to the journal underutilised.

So it appears to me that using a CSUM item and ordered buffers is a
far more CPU, memory and journal space efficient algorithm for
logging the csum updates. I suspect there may also be some potential
ordered buffer locking optimisations we can also do because nothing
ever reads from ordered buffers in the commit code path.

I'm going to wait for the design doc, though, before going any
further - I've spent enough time reverse engineering design and
algorithms from this patchset for now.

-Dave.
-- 
Dave Chinner
dgc@kernel.org

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 14/21] xfs: add support for reading with data checksums
  2026-09-29  0:42   ` Darrick J. Wong
@ 2026-10-05 12:59     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 12:59 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 05:42:04PM -0700, Darrick J. Wong wrote:
> > +	error = -EIO;
> > +	if (WARN_ON_ONCE(ioend->io_bio.bi_iter.bi_idx))
> 
> What does this warning mean?  That we've somehow already advanced the
> bvec iterator?

Yes.  I've added the following comment:

	/*
         * The main I/O end bio is either freshly build by iomap, and has both
         * bi_idx and bi_offset set to zero, or comes from a bvec iter for
         * direct I/O, in which case bi_offset can be non-zero, but bi_idx still
         * must be zero as the bio always points to the first actually used
         * bio_vec in the iov_iter.
         */


> > +	/* No bp for early xfs_rtcsum_prepare_read failures. */
> > +	if (!bp)
> > +		return error;
> 
> Is it possible for error to be zero here?

It shouldn't.  I've added an assert and clarified the comment.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 15/21] xfs: add support for writing with data checksums
  2026-09-29  1:01   ` Darrick J. Wong
@ 2026-10-05 13:00     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 13:00 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 06:01:03PM -0700, Darrick J. Wong wrote:
> > +	if (xfs_is_rtcsum_inode(ip) && !xfs_rtcsum_prepare_read(ioend))
> > +		return blk_status_to_errno(bio->bi_status);
> > +	if (ioend_flags & IOMAP_IOEND_INTEGRITY)
> > +		fs_bio_integrity_alloc(bio);
> > +	saved_iter = bio->bi_iter;
> > +	error = submit_bio_wait(bio);
> > +	if (bio_integrity(bio)) {
> > +		if (!error)
> > +			error = fs_bio_integrity_verify(bio, &saved_iter);
> > +		fs_bio_integrity_free(bio);
> > +	}
> > +	if (xfs_is_rtcsum_inode(ip))
> > +		error = xfs_rtcsum_verify_ioend(ioend, error);
> > +	return error;
> > +}
> 
> Hm, synchronous reads to pull in whatever unaligned parts of the
> pagecache aren't yet uptodate?

Yes.  Very similar to the generic version in
iomap_bio_read_folio_range_sync, but using an ioend and with all the
extra csum magic.

> It would be helpful to document this ordering dependency here explicitly
> for the benefit of code spelunkers in a few years.

I've added it with a minor edit.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 16/21] xfs: add data checksum support to zoned garbage collection
  2026-09-29  1:06   ` Darrick J. Wong
@ 2026-10-05 13:11     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 13:11 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 06:06:39PM -0700, Darrick J. Wong wrote:
> > +	/* Buffer for data checksums */
> > +	struct xfs_buf			*csum_bp;
> > +	void				*csum_buf;
> 
> Where is csum_buf used?

Nowhere, left from an earlier version.

> > +		error = xfs_rtcsum_read_async(mp, chunk->old_startblock,
> > +				&chunk->csum_bp);
> > +		if (error) {
> > +			xfs_force_shutdown(mp, SHUTDOWN_META_IO_ERROR);
> > +			xfs_zone_gc_free_chunk(chunk);
> > +			return false;
> > +		}
> 
> Ok so this starts the csum block read asynchronously, and then submits
> the read bio.  chunk->csum_bp is the checksum buffer of the rtcsum file
> of the zone we're cleaning out, right?

Yes.

> > +		boff = xfs_rtb_to_rtcsumoff(mp, chunk->old_startblock);
> > +		error = xfs_rtcsum_log(chunk->oz, chunk->new_daddr, chunk->len,
> > +				chunk->csum_bp->b_addr + boff);
> 
> ...and then write the same checksum to the new daddr using the old
> checksum.  There's no explicit validation of the checksum in the garbage
> collector itself, but I gather the read ioend did this for us already,
> right?

Exactly.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 17/21] xfs: verify data checksums during media verification
  2026-09-29  1:19   ` Darrick J. Wong
@ 2026-10-05 13:13     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 13:13 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 06:19:36PM -0700, Darrick J. Wong wrote:
> > +xfs_submit_verify_bio_csum(
> 
> What's the return value convention here?  1 for media error, 0 for
> success, or negative errno if we failed to issue the read?

Yeah:

/* returns 0 for success, 1 for media error, or -errno for other errors */


> > +	bno = xfs_daddr_to_rtb(mp, *daddr);
> > +	end = xfs_daddr_to_rtb(mp, *daddr + *bbcount);
> 
> It occurs to me that xfs_daddr_to_rtb rounds its argument down.  So if
> you pass in daddr==0 and bbcount==2, you'll get bno==end==0 and do no
> verification.  I would hope that callers won't pass in parameters like
> that, but who knows?
> 
> So I think this should be:
> 
> 	end = xfs_daddr_to_rtb(mp,
> 			*daddr + *bbcount + XFS_FSB_TO_BB(mp, 1) - 1);
> 
> Which I admit is a bit gross.

Not too horrible.  We should also check for a 0 len after rounding in
case nothing is left.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 18/21] xfs: don't try to verify checksums on empty zones
  2026-09-29  1:25   ` Darrick J. Wong
@ 2026-10-05 13:14     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 13:14 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 06:25:55PM -0700, Darrick J. Wong wrote:
> Just a general note -- if the csum update never makes it to disk, then
> the remap cannot have been written to disk either.  Therefore, fsmap
> will report that as unowned space, and xfs_scrub won't try to
> read-verify it, given the xfs_scrub patch you post later to only read rt
> blocks for which there are rmap records.
> 
> So I think this is only a theoretical concern, right?

At least earlier I could reproduce a case where scrub raced with
a zone reset and tries to verify the reset zone against the now
implicitly zeroed zone.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 21/21] xfs: enable RT data checksums
  2026-09-29  1:27   ` Darrick J. Wong
@ 2026-10-05 13:16     ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 13:16 UTC (permalink / raw)
  To: Darrick J. Wong
  Cc: Christoph Hellwig, Carlos Maiolino, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Sep 28, 2026 at 06:27:55PM -0700, Darrick J. Wong wrote:
> On Thu, Sep 24, 2026 at 11:59:53AM +0200, Christoph Hellwig wrote:
> > All mandatory pieces are in place now, allow mounting.
> > 
> > Signed-off-by: Christoph Hellwig <hch@lst.de>
> 
> Looks good.  I'll try to cough up an amendment to the ondisk metadata
> documentation to go with this.

I was going to do that myself, but if you can afford the time that's
even better as that now means at least two people understand the format,
which always is a good thing.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-09-30  7:11           ` Dave Chinner
@ 2026-10-05 13:53             ` Christoph Hellwig
  2026-10-06  5:31               ` Dave Chinner
  0 siblings, 1 reply; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-05 13:53 UTC (permalink / raw)
  To: Dave Chinner
  Cc: Christoph Hellwig, Carlos Maiolino, Darrick J . Wong, Jens Axboe,
	Christian Brauner, linux-xfs, linux-fsdevel

On Wed, Sep 30, 2026 at 05:11:58PM +1000, Dave Chinner wrote:
> My initial thought was that on disk format structures shouldn't be
> defined by the limitations of the OS memory allocation, but <shrug>.
> It kept nagging at me, though.

Well, it would be nice to be able to design without all the real-life
constraints around us, wouldn't it?  I'm trying to strike a balance
between what would be useful (go big) and what is feasible.  If we
want other values, we can always increase the support range for
newer kernels and tools.

> I looked more closely at what XFS_BLI_PREALLOC did to try to
> understand why it existed.  It triggers a max-sized CIL logvec
> structure for the buffer object. For a 32kB buffer logged as a
> single contiguous range, this ends up being about 32kB + a logvec
> header, plus a log iovec, plus a BLF, plus a couple of ophdrs. So
> it's about 32kB + 200-250 bytes.
> 
> That means the shadow buffer for a csum buffer is always considered
> a costly allocation by the MM subsystem.

It ends up using vmalloc exclusively based on tracing for me,
but that might be different on different systems.

> Hence for csum BLI, if a CIL flush happens on a partially filled
> buffer, a good amount of that shadow buffer will go unused. Then we
> allocate another (costly) shadow buffer on the next update. If CIL
> flushes happen frequently enough then we will be repeatedly doing
> costly allocations for shadow buffers that we don't actually use.

Yes.  But if we don't do this we realloc for every few blocks
written, which is a lot more costly.

> Not ideal - I think that means the original "sized for mm fast path"
> intent is really only valid for the read side of the csum
> algorithms as implemented by the patchset.

It is valid for the xfs_buf backing where we actually hit the folio
allocator.  Which is used both for read and write, but obviously
most workloads tend to hit reads a lot harder than writes.

> Ok, we have a solution to this problem. I created ordered buffers
> and one-shot log items to avoid the journalling overhead of static
> inode buffer initialisation back in 2013. The ICREATE log item is
> the one-shot log item that records a buffer should be initialised,
> and the ordered buffer allows the modified buffer to be passed
> through the journal to metadta writeback without it's contents being
> logged.
> 
> Given that csum updates are a small, known size, non-overlapping
> one-shot update to a buffer, they fit the same model that
> ICREAT+ordered implements. Adding a new CSUM log item made up of a
> format header and varible size csum payload region provides the
> equivalent of the ICREAT item for journalled inode buffer
> initialisation.

I initially looked into intent/done based csums, but we still end
with an allocation per log operation, and a memcpy both into that
and into the buffer, while adding a lot of new log items.

One thing I played with for a while until I realized that the simple
buf item actually provides good enough performance is special
xfs_log_vec that is not included in the main log vec / shadow allocation
but points to external memory.  This obviously only works for fixed
size non-overlapping regions, but then isn't too bad.  This is the
prep work for it, which I recently refreshed:

https://git.infradead.org/?p=users/hch/xfs.git;a=shortlog;h=refs/heads/xlog-ophdr

These can work with the buf_item on-disk format, so I'd rather not
prematurely optimize it, as the prototype shows that I can go to that
any time I want.  And eventually I think I'd want to go there, as it
drastically reduces the memory usage if only the format header and
two ophrs need to be allocated ontop of the backing buffer.  But there's
plenty more important things on the plate for now.

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-10-05 13:53             ` Christoph Hellwig
@ 2026-10-06  5:31               ` Dave Chinner
  2026-10-07 13:46                 ` Christoph Hellwig
  0 siblings, 1 reply; 69+ messages in thread
From: Dave Chinner @ 2026-10-06  5:31 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

On Mon, Oct 05, 2026 at 03:53:20PM +0200, Christoph Hellwig wrote:
> On Wed, Sep 30, 2026 at 05:11:58PM +1000, Dave Chinner wrote:
> > My initial thought was that on disk format structures shouldn't be
> > defined by the limitations of the OS memory allocation, but <shrug>.
> > It kept nagging at me, though.
> 
> Well, it would be nice to be able to design without all the real-life
> constraints around us, wouldn't it?  I'm trying to strike a balance
> between what would be useful (go big) and what is feasible.  If we
> want other values, we can always increase the support range for
> newer kernels and tools.
> 
> > I looked more closely at what XFS_BLI_PREALLOC did to try to
> > understand why it existed.  It triggers a max-sized CIL logvec
> > structure for the buffer object. For a 32kB buffer logged as a
> > single contiguous range, this ends up being about 32kB + a logvec
> > header, plus a log iovec, plus a BLF, plus a couple of ophdrs. So
> > it's about 32kB + 200-250 bytes.
> > 
> > That means the shadow buffer for a csum buffer is always considered
> > a costly allocation by the MM subsystem.
> 
> It ends up using vmalloc exclusively based on tracing for me,
> but that might be different on different systems.

That'll be because the repeated multipage allocations eventually end
up fragmenting memory and so anything over 8-16kB will go straight
to vmalloc.

> > Hence for csum BLI, if a CIL flush happens on a partially filled
> > buffer, a good amount of that shadow buffer will go unused. Then we
> > allocate another (costly) shadow buffer on the next update. If CIL
> > flushes happen frequently enough then we will be repeatedly doing
> > costly allocations for shadow buffers that we don't actually use.
> 
> Yes.  But if we don't do this we realloc for every few blocks
> written, which is a lot more costly.

I'm not sure I understand what you are refering to there. I'm not
talking aobut the 'realloc because logged size grows', I'm talking
aoubt 'realloc because CIL pushes steal the shadow buffer'.

> > Not ideal - I think that means the original "sized for mm fast path"
> > intent is really only valid for the read side of the csum
> > algorithms as implemented by the patchset.
> 
> It is valid for the xfs_buf backing where we actually hit the folio
> allocator.  Which is used both for read and write, but obviously
> most workloads tend to hit reads a lot harder than writes.

Not if the working set fits in cache. A 'read heavy' workload will
often change to 'write heavy' when it fits in cache.....

> > Ok, we have a solution to this problem. I created ordered buffers
> > and one-shot log items to avoid the journalling overhead of static
> > inode buffer initialisation back in 2013. The ICREATE log item is
> > the one-shot log item that records a buffer should be initialised,
> > and the ordered buffer allows the modified buffer to be passed
> > through the journal to metadta writeback without it's contents being
> > logged.
> > 
> > Given that csum updates are a small, known size, non-overlapping
> > one-shot update to a buffer, they fit the same model that
> > ICREAT+ordered implements. Adding a new CSUM log item made up of a
> > format header and varible size csum payload region provides the
> > equivalent of the ICREAT item for journalled inode buffer
> > initialisation.
> 
> I initially looked into intent/done based csums, but we still end
> with an allocation per log operation, and a memcpy both into that
> and into the buffer, while adding a lot of new log items.

I'm not talking about an intent/done based setup - that requires an
intent transaction, then a buffer + done transaction. That's very
different to what I suggested (and what ICREATE implements).

I'm talking about a single one-shot transaction that logs the csums
and that only. The buffer is ordered, so not logged. Single
transaction, generally small in size, never gets relogged, can be
committed asynchronously as long as it is replayed before the file
offset relocation BMBT update in recovery.

The CIL can handle hundreds of thousands of csum objects just fine
(no different to logging hundreds of thousands of inode cores). It
is simpler than intents, too, because they are at checkpoint
completion and never enter the AIL (unlike intents). The only thing
that ends up the AIL is the BLI, and that works like a
INODE_ALLOC_BUF in that it remains at the initial LSN in the AIL
even when it gets updated. i.e. once it is in the AIL, it never gets
moved forward, but it continues to aggregate changes until it gets
written back with the LSN of the latest committed change stamped
into it so recovery does the right thing with it.

So, yeah, I'm definitely not suggesting using intents...

> One thing I played with for a while until I realized that the
> simple buf item actually provides good enough performance is
> special xfs_log_vec that is not included in the main log vec /
> shadow allocation but points to external memory.

That's problematic. The reason delayed logging works is that it
broke the dependency between external memory that log items pointed
at needing to be locked and stable until the external memory was
copied into the iclogs. The disconnection of the objects passed to
xfs_trans_commit() vs xlog_write() whilst keeping the logged data
stable is what allows the CIL to work

The shadow buffer does that decoupling, and it means that there is
no requirement for the original logged item to remain stable, or
even remain in existence whilst the CIL holds onto the infomration
that needs to be journalled. At checkpoint completion, shadow buffer
is also used to do a reverse lookup to the log item to enable
insertion into the AIL.  IOWs, log items are not tracked across
journal checkpoint IO - shadow buffers are, and if you get rid of
shadow buffers for a log item, we have to special case that
everywhere in the LV/checkpoint handling. I dont' think that's a good
idea.

The shadow buffer also avoids the need for the CIL to lock external
objects to copy the data out of them. The lock order is lock external
object -> commit -> read-lock checkpoint - format into CIL -> unlock.
When pushing, the order is write-lock checkpoint -> lock iclog ->
format from CIL into iclog -> unlock iclog -> unlock chkpt. We also
can call xfs_log_force() whilst holding inode locks, putting iclog
locks inside high level object locks.

Hence we really can't lock external objects from the checkpoint side
because of the lock inversion problems they entail. So object
stability is a problem, and ....

> This obviously
> only works for fixed size non-overlapping regions, but then isn't
> too bad. 

"trust me, bro!" is not my idea of maintainable, landmine free
design, especially now with LLMs being able to poke holes in complex
zero-copy/object sharing schemes and exploit them in less than
obvious ways...

> This is the prep work for it, which I recently
> refreshed:
> 
> https://git.infradead.org/?p=users/hch/xfs.git;a=shortlog;h=refs/heads/xlog-ophdr

Not a fan of rewriting xlog_write() -again-, this time to bring back
all the bad old patterns of managing ophdr space itself. We got rid of
that method of managing ophdrs because of all the special accounting
it needs to sprinkle through the logic to get log space consumption
correct. It was complex, difficult to reason about, and a source of
bugs.

Fixing these problems was the one of the main reasons we moved all
the ophdr management and accounting out into the CIL and logvecs to
begin with. Hence I'm not a great fan of going back to the old
way, whatever the reason.

> These can work with the buf_item on-disk format, so I'd rather not
> prematurely optimize it, as the prototype shows that I can go to
> that any time I want.  And eventually I think I'd want to go
> there, as it drastically reduces the memory usage if only the
> format header and two ophrs need to be allocated ontop of the
> backing buffer.  But there's plenty more important things on the
> plate for now.

I'd much prefer we use a method we know works and scales rahter than
create something new that requires punching through abstractions,
can't guarantee stability or lifetime of external objects, and isn't
demonstrated to be necessary to meet performance requirements.

-Dave.
-- 
Dave Chinner
dgc@kernel.org

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: support for RT data checksums
  2026-10-06  5:31               ` Dave Chinner
@ 2026-10-07 13:46                 ` Christoph Hellwig
  0 siblings, 0 replies; 69+ messages in thread
From: Christoph Hellwig @ 2026-10-07 13:46 UTC (permalink / raw)
  To: Dave Chinner
  Cc: Christoph Hellwig, Carlos Maiolino, Darrick J . Wong, Jens Axboe,
	Christian Brauner, linux-xfs, linux-fsdevel

On Tue, Oct 06, 2026 at 04:31:29PM +1100, Dave Chinner wrote:
> I'm not sure I understand what you are refering to there. I'm not
> talking aobut the 'realloc because logged size grows', I'm talking
> aoubt 'realloc because CIL pushes steal the shadow buffer'.

So you mean alloc again, not realloc, ok.

> > with an allocation per log operation, and a memcpy both into that
> > and into the buffer, while adding a lot of new log items.
> 
> I'm not talking about an intent/done based setup - that requires an
> intent transaction, then a buffer + done transaction. That's very
> different to what I suggested (and what ICREATE implements).
> 
> I'm talking about a single one-shot transaction that logs the csums
> and that only. The buffer is ordered, so not logged. Single
> transaction, generally small in size, never gets relogged, can be
> committed asynchronously as long as it is replayed before the file
> offset relocation BMBT update in recovery.

We'd still need memory for the csum values both in the icreate
item and the ordered buffer.

> > One thing I played with for a while until I realized that the
> > simple buf item actually provides good enough performance is
> > special xfs_log_vec that is not included in the main log vec /
> > shadow allocation but points to external memory.
> 
> That's problematic. The reason delayed logging works is that it
> broke the dependency between external memory that log items pointed
> at needing to be locked and stable until the external memory was
> copied into the iclogs. The disconnection of the objects passed to
> xfs_trans_commit() vs xlog_write() whilst keeping the logged data
> stable is what allows the CIL to work

Yes, and for all the current buffers and similar items that's very
important because the data can actually change, and we must not
pick up that change.  It does however not matter much when we know
each bit of the buffer payload is only ever updated once, and thus
the is no need to double buffer or lock the backing memory.
The only important thing left with that is that the buffer as the
backing memory must not be freed.


^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 18/21] xfs: don't try to verify checksums on empty zones
  2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
  2026-09-29  1:25   ` Darrick J. Wong
@ 2026-10-08 11:43   ` Anuj gupta
  1 sibling, 0 replies; 69+ messages in thread
From: Anuj gupta @ 2026-10-08 11:43 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel, Anuj Gupta

> +               error = xfs_csum_verify(mp, bio, saved_iter, csum_buf, bno,
> +                               false);

Could an open zone be verified past oz_allocated here? XFS_RTG_FREE is
cleared on opening, but the verification range is not bounded by
oz_allocated. xfs_zone_rgbno_is_valid() treats only the prefix below
oz_allocated as valid; should this trim the verification range at that
boundary? Otherwise a whole-rtdev XFS_IOC_VERIFY_MEDIA request returns
a spurious userspace error.

^ permalink raw reply	[flat|nested] 69+ messages in thread

* Re: [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC
  2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
  2026-09-24 21:49   ` Darrick J. Wong
@ 2026-10-08 11:46   ` Anuj gupta
  1 sibling, 0 replies; 69+ messages in thread
From: Anuj gupta @ 2026-10-08 11:46 UTC (permalink / raw)
  To: Christoph Hellwig, Anuj Gupta
  Cc: Carlos Maiolino, Darrick J . Wong, Jens Axboe, Christian Brauner,
	linux-xfs, linux-fsdevel

> +       if (bip->bli_flags & XFS_BLI_PREALLOC)
> +               *nbytes = xfs_calc_buf_res(bip->bli_format_count, bp->b_length);
> +

The *nbytes assignment here is immediately overwritten by the
unconditional round_up(bytes, 512) below, so the PREALLOC pre-sizing
never takes effect. Move the PREALLOC block after the round_up line.

^ permalink raw reply	[flat|nested] 69+ messages in thread

end of thread, other threads:[~2026-10-08 11:47 UTC | newest]

Thread overview: 69+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24  9:59 support for RT data checksums Christoph Hellwig
2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
2026-09-24 20:29   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
2026-09-24 21:39   ` Darrick J. Wong
2026-09-25  5:53     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
2026-09-24 21:43   ` Darrick J. Wong
2026-09-25  5:54     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
2026-09-24 21:44   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
2026-09-24 21:49   ` Darrick J. Wong
2026-09-25  5:57     ` Christoph Hellwig
2026-10-08 11:46   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
2026-09-24 22:03   ` Darrick J. Wong
2026-09-25  5:58     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
2026-09-24 22:13   ` Darrick J. Wong
2026-09-25  0:04     ` Eric Biggers
2026-09-25  6:01     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
2026-09-24 22:24   ` Darrick J. Wong
2026-09-25  6:10     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
2026-09-24 22:30   ` Darrick J. Wong
2026-09-25  6:12     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
2026-09-25 23:20   ` Darrick J. Wong
2026-09-26  6:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
2026-09-25 23:21   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
2026-09-25 23:24   ` Darrick J. Wong
2026-09-26  6:15     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
2026-09-29  0:42   ` Darrick J. Wong
2026-10-05 12:59     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
2026-09-29  1:01   ` Darrick J. Wong
2026-10-05 13:00     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
2026-09-29  1:06   ` Darrick J. Wong
2026-10-05 13:11     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
2026-09-29  1:19   ` Darrick J. Wong
2026-10-05 13:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
2026-09-29  1:25   ` Darrick J. Wong
2026-10-05 13:14     ` Christoph Hellwig
2026-10-08 11:43   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
2026-09-29  1:26   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-10-05 13:16     ` Christoph Hellwig
2026-09-24 22:52 ` support for " Dave Chinner
2026-09-25  6:27   ` Christoph Hellwig
2026-09-27 22:59     ` Dave Chinner
2026-09-28  5:24       ` Christoph Hellwig
2026-09-29 14:11         ` Dave Chinner
2026-09-30  7:11           ` Dave Chinner
2026-10-05 13:53             ` Christoph Hellwig
2026-10-06  5:31               ` Dave Chinner
2026-10-07 13:46                 ` Christoph Hellwig

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox