* xfsprogs support for RT data checksums
@ 2026-09-24 10:03 Christoph Hellwig
2026-09-24 10:03 ` [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2 Christoph Hellwig
` (31 more replies)
0 siblings, 32 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Hi all,
this series add support for RT data checksums to xfsprogs.
Much of this should be pretty straight forward, but the repair support
to rebuild the checksum metafiles isn't exactly great code. I wonder
if it makes sense to even run this by default, or only under a special
option. And if we run it, if we should parallelize it and/or come up
with a scheme to prefetch to the file data.
^ permalink raw reply [flat|nested] 57+ messages in thread
* [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-24 20:30 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 02/32] add cpu_to_le64 and le64_to_cpu_helpers Christoph Hellwig
` (30 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Using tabs for indentation messes up man page rendering, so use spaces
for rtstart to match the other fields.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
man/man2/ioctl_xfs_fsgeometry.2 | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/man/man2/ioctl_xfs_fsgeometry.2 b/man/man2/ioctl_xfs_fsgeometry.2
index 037f8e15e415..6d30610b71be 100644
--- a/man/man2/ioctl_xfs_fsgeometry.2
+++ b/man/man2/ioctl_xfs_fsgeometry.2
@@ -50,7 +50,7 @@ struct xfs_fsop_geom {
__u32 sick;
__u32 checked;
__u64 rgextents;
- __u64 rtstart;
+ __u64 rtstart;
__u64 rtreserved;
__u64 reserved[14];
};
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 02/32] add cpu_to_le64 and le64_to_cpu_helpers
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
2026-09-24 10:03 ` [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2 Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-24 20:30 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 03/32] libfrog: add a crc64_nvme implementation Christoph Hellwig
` (29 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
The crc64 handling will need them.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
include/xfs_arch.h | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/include/xfs_arch.h b/include/xfs_arch.h
index d46ae47094ae..348ff0077c99 100644
--- a/include/xfs_arch.h
+++ b/include/xfs_arch.h
@@ -195,6 +195,8 @@ static __inline__ void __swab64s(__u64 *addr)
#define cpu_to_le32(val) ((__force __be32)__swab32((__u32)(val)))
#define le32_to_cpu(val) (__swab32((__force __u32)(__le32)(val)))
+#define cpu_to_le64(val) ((__force __be64)__swab64((__u64)(val)))
+#define le64_to_cpu(val) (__swab64((__force __u64)(__le64)(val)))
#define __constant_cpu_to_le32(val) \
((__force __le32)___constant_swab32((__u32)(val)))
@@ -210,6 +212,8 @@ static __inline__ void __swab64s(__u64 *addr)
#define cpu_to_le32(val) ((__force __le32)(__u32)(val))
#define le32_to_cpu(val) ((__force __u32)(__le32)(val))
+#define cpu_to_le64(val) ((__force __le64)(__u64)(val))
+#define le64_to_cpu(val) ((__force __u64)(__le64)(val))
#define __constant_cpu_to_le32(val) \
((__force __le32)(__u32)(val))
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 03/32] libfrog: add a crc64_nvme implementation
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
2026-09-24 10:03 ` [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2 Christoph Hellwig
2026-09-24 10:03 ` [PATCH 02/32] add cpu_to_le64 and le64_to_cpu_helpers Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-24 10:03 ` [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL Christoph Hellwig
` (28 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Imported from the Linux kernel, using and older version as of commit
commit 067bc8717aee ("lib/crc64: add support for arch-optimized
implementations") which doesn't have the relatively complicated
common optimized crc implementation in the current kernel.
Compared to the kernel, libfrog drops the ECMA CRC64 and only keeps the
NVMe one that is used by XFS RT deata checksums.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libfrog/Makefile | 14 +++++++--
libfrog/crc64.c | 23 ++++++++++++++
libfrog/crc64.h | 25 +++++++++++++++
libfrog/gen_crc64table.c | 67 ++++++++++++++++++++++++++++++++++++++++
4 files changed, 127 insertions(+), 2 deletions(-)
create mode 100644 libfrog/crc64.c
create mode 100644 libfrog/crc64.h
create mode 100644 libfrog/gen_crc64table.c
diff --git a/libfrog/Makefile b/libfrog/Makefile
index c7bcd6a778d7..8cf07622f01f 100644
--- a/libfrog/Makefile
+++ b/libfrog/Makefile
@@ -18,6 +18,7 @@ bitmap.c \
bulkstat.c \
convert.c \
crc32.c \
+crc64.c \
file_exchange.c \
flagmap.c \
fsgeom.c \
@@ -51,6 +52,8 @@ crc32c.h \
crc32cselftest.h \
crc32defs.h \
crc32table.h \
+crc64.h \
+crc64table.h \
dahashselftest.h \
div64.h \
fakelibattr.h \
@@ -80,9 +83,10 @@ zones.h
GETTEXT_PY = \
gettext.py
-LSRCFILES += gen_crc32table.c
+LSRCFILES += gen_crc32table.c gen_crc64table.c
-LDIRT = gen_crc32table crc32table.h
+LDIRT = gen_crc32table crc32table.h \
+ gen_crc64table crc64table.h
ifeq ($(ENABLE_GETTEXT),yes)
HAVE_GETTEXT = True
@@ -120,6 +124,12 @@ crc32table.h: gen_crc32table.c crc32defs.h
@echo " [GENERATE] $@"
$(Q) ./gen_crc32table > crc32table.h
+crc64table.h: gen_crc64table.c
+ @echo " [CC] gen_crc64table"
+ $(Q) $(BUILD_CC) $(BUILD_CFLAGS) -o gen_crc64table $<
+ @echo " [GENERATE] $@"
+ $(Q) ./gen_crc64table > crc64table.h
+
$(GETTEXT_PY): $(GETTEXT_PY).in $(TOPDIR)/include/builddefs
@echo " [SED] $@"
$(Q)$(SED) -e "s|@HAVE_GETTEXT@|$(HAVE_GETTEXT)|g" \
diff --git a/libfrog/crc64.c b/libfrog/crc64.c
new file mode 100644
index 000000000000..280ac77b16cc
--- /dev/null
+++ b/libfrog/crc64.c
@@ -0,0 +1,23 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * 64-bit CRC calculation based on the NVMe NVM Command Set specication:
+ * using least-significant-bit first bit order:
+ *
+ * x^64 + x^63 + x^61 + x^59 + x^58 + x^56 + x^55 + x^52 + x^49 + x^48 + x^47 +
+ * x^46 + x^44 + x^41 + x^37 + x^36 + x^34 + x^32 + x^31 + x^28 + x^26 + x^23 +
+ * x^22 + x^19 + x^16 + x^13 + x^12 + x^10 + x^9 + x^6 + x^4 + x^3 + 1
+ *
+ * Copyright 2018 SUSE Linux.
+ * Author: Coly Li <colyli@suse.de>
+ */
+
+#include <stdint.h>
+#include "libfrog/crc64.h"
+#include "crc64table.h"
+
+uint64_t crc64_nvme_generic(uint64_t crc, const uint8_t *p, size_t len)
+{
+ while (len--)
+ crc = (crc >> 8) ^ crc64nvmetable[(crc & 0xff) ^ *p++];
+ return crc;
+}
diff --git a/libfrog/crc64.h b/libfrog/crc64.h
new file mode 100644
index 000000000000..283bd3949577
--- /dev/null
+++ b/libfrog/crc64.h
@@ -0,0 +1,25 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LIBFROG_CRC64_H
+#define _LIBFROG_CRC64_H
+
+#include <stdint.h>
+#include <stddef.h>
+
+uint64_t crc64_nvme_generic(uint64_t crc, const uint8_t *p, size_t len);
+
+/**
+ * crc64_nvme - Calculate CRC64-NVME
+ * @crc: seed value for computation. 0 for a new CRC calculation, or the
+ * previous crc64 value if computing incrementally.
+ * @p: pointer to buffer over which CRC64 is run
+ * @len: length of buffer @p
+ *
+ * This computes the CRC64 defined in the NVME NVM Command Set Specification,
+ * *including the bitwise inversion at the beginning and end*.
+ */
+static inline uint64_t crc64_nvme(uint64_t crc, const void *p, size_t len)
+{
+ return ~crc64_nvme_generic(~crc, p, len);
+}
+
+#endif /* _LIBFROG_CRC64_H */
diff --git a/libfrog/gen_crc64table.c b/libfrog/gen_crc64table.c
new file mode 100644
index 000000000000..bf88f8c6ffc8
--- /dev/null
+++ b/libfrog/gen_crc64table.c
@@ -0,0 +1,67 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Generate lookup table for the table-driven CRC64 calculation.
+ *
+ * gen_crc64table is executed at build time and generates crc64table.h.
+ * This header is included by crc64.c for the table-driven CRC64 calculation.
+ *
+ * See crc64.c for more information about which specification and polynomial
+ * arithmetic that gen_crc64table.c follows to generate the lookup table.
+ *
+ * Copyright 2018 SUSE Linux.
+ * Author: Coly Li <colyli@suse.de>
+ */
+#include <inttypes.h>
+#include <stdio.h>
+
+#define CRC64_NVME_POLY 0x9A6C9329AC4BC9B5ULL
+
+static uint64_t crc64_nvme_table[256] = {0};
+
+static void generate_reflected_crc64_table(uint64_t table[256], uint64_t poly)
+{
+ uint64_t i, j, c, crc;
+
+ for (i = 0; i < 256; i++) {
+ crc = 0ULL;
+ c = i;
+
+ for (j = 0; j < 8; j++) {
+ if ((crc ^ (c >> j)) & 1)
+ crc = (crc >> 1) ^ poly;
+ else
+ crc >>= 1;
+ }
+ table[i] = crc;
+ }
+}
+
+static void output_table(uint64_t table[256])
+{
+ int i;
+
+ for (i = 0; i < 256; i++) {
+ printf("\t0x%016" PRIx64 "ULL", table[i]);
+ if (i & 0x1)
+ printf(",\n");
+ else
+ printf(", ");
+ }
+ printf("};\n");
+}
+
+static void print_crc64_tables(void)
+{
+ printf("/* this file is generated - do not edit */\n\n");
+ printf("#include <stdint.h>\n");
+
+ printf("\nstatic const uint64_t crc64nvmetable[256] = {\n");
+ output_table(crc64_nvme_table);
+}
+
+int main(int argc, char *argv[])
+{
+ generate_reflected_crc64_table(crc64_nvme_table, CRC64_NVME_POLY);
+ print_crc64_tables();
+ return 0;
+}
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (2 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 03/32] libfrog: add a crc64_nvme implementation Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-24 20:44 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 05/32] libxfs: remove a spurious xfs_rtbitmap.h include in xfs_sb.c Christoph Hellwig
` (27 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_platform.h | 1 +
1 file changed, 1 insertion(+)
diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
index d09fb0aae00f..a61d50de9dc0 100644
--- a/libxfs/xfs_platform.h
+++ b/libxfs/xfs_platform.h
@@ -212,6 +212,7 @@ void inode_init_owner(struct mnt_idmap *idmap, struct inode *inode,
#define round_up(x, y) ((((x)-1) | __round_mask(x, y))+1)
#define round_down(x, y) ((x) & ~__round_mask(x, y))
#define DIV_ROUND_UP(n,d) (((n) + (d) - 1) / (d))
+#define DIV_ROUND_UP_ULL(n,d) DIV_ROUND_UP(n,d)
/*
* Handling for kernel bitmap types.
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 05/32] libxfs: remove a spurious xfs_rtbitmap.h include in xfs_sb.c
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (3 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-25 23:27 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 06/32] libxfs: add SZ_* constants Christoph Hellwig
` (26 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
This fully resyncs with the kernel version.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_sb.c | 1 -
1 file changed, 1 deletion(-)
diff --git a/libxfs/xfs_sb.c b/libxfs/xfs_sb.c
index ea99f8b5eee5..cccbcd153316 100644
--- a/libxfs/xfs_sb.c
+++ b/libxfs/xfs_sb.c
@@ -27,7 +27,6 @@
#include "xfs_rtgroup.h"
#include "xfs_rtrmap_btree.h"
#include "xfs_rtrefcount_btree.h"
-#include "xfs_rtbitmap.h"
/*
* Physical superblock buffer manipulations. Shared with libxfs in userspace.
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 06/32] libxfs: add SZ_* constants
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (4 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 05/32] libxfs: remove a spurious xfs_rtbitmap.h include in xfs_sb.c Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-25 23:27 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 07/32] xfs: remove spurious XBF_DONE clearing on readahead validation failure Christoph Hellwig
` (25 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Import a slightly adapted version of <linux/sizes.h> from the kernel so that
we can use these constants in shared libxfs code.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
include/libxfs.h | 1 +
include/sizes.h | 67 +++++++++++++++++++++++++++++++++++++++++++
libxfs/xfs_platform.h | 1 +
3 files changed, 69 insertions(+)
create mode 100644 include/sizes.h
diff --git a/include/libxfs.h b/include/libxfs.h
index 7f0c22ef4991..34dd267831c0 100644
--- a/include/libxfs.h
+++ b/include/libxfs.h
@@ -18,6 +18,7 @@
#include "platform_defs.h"
#include "xfs.h"
+#include "sizes.h"
#include "list.h"
#include "hlist.h"
#include "cache.h"
diff --git a/include/sizes.h b/include/sizes.h
new file mode 100644
index 000000000000..69075e8f3370
--- /dev/null
+++ b/include/sizes.h
@@ -0,0 +1,67 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+#ifndef __SIZES_H__
+#define __SIZES_H__
+
+#define SZ_1 0x00000001
+#define SZ_2 0x00000002
+#define SZ_4 0x00000004
+#define SZ_8 0x00000008
+#define SZ_16 0x00000010
+#define SZ_32 0x00000020
+#define SZ_64 0x00000040
+#define SZ_128 0x00000080
+#define SZ_256 0x00000100
+#define SZ_512 0x00000200
+
+#define SZ_1K 0x00000400
+#define SZ_2K 0x00000800
+#define SZ_4K 0x00001000
+#define SZ_8K 0x00002000
+#define SZ_16K 0x00004000
+#define SZ_24K 0x00006000
+#define SZ_32K 0x00008000
+#define SZ_64K 0x00010000
+#define SZ_128K 0x00020000
+#define SZ_192K 0x00030000
+#define SZ_256K 0x00040000
+#define SZ_384K 0x00060000
+#define SZ_512K 0x00080000
+
+#define SZ_1M 0x00100000
+#define SZ_2M 0x00200000
+#define SZ_3M 0x00300000
+#define SZ_4M 0x00400000
+#define SZ_6M 0x00600000
+#define SZ_8M 0x00800000
+#define SZ_12M 0x00c00000
+#define SZ_16M 0x01000000
+#define SZ_18M 0x01200000
+#define SZ_24M 0x01800000
+#define SZ_32M 0x02000000
+#define SZ_64M 0x04000000
+#define SZ_128M 0x08000000
+#define SZ_256M 0x10000000
+#define SZ_512M 0x20000000
+
+#define SZ_1G 0x40000000
+#define SZ_2G 0x80000000
+
+#define SZ_4G 0x100000000ULL
+#define SZ_8G 0x200000000ULL
+#define SZ_16G 0x400000000ULL
+#define SZ_32G 0x800000000ULL
+#define SZ_64G 0x1000000000ULL
+#define SZ_128G 0x2000000000ULL
+#define SZ_256G 0x4000000000ULL
+#define SZ_512G 0x8000000000ULL
+
+#define SZ_1T 0x10000000000ULL
+#define SZ_2T 0x20000000000ULL
+#define SZ_4T 0x40000000000ULL
+#define SZ_8T 0x80000000000ULL
+#define SZ_16T 0x100000000000ULL
+#define SZ_32T 0x200000000000ULL
+#define SZ_64T 0x400000000000ULL
+#define SZ_128T 0x800000000000ULL
+
+#endif /* __SIZES_H__ */
diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
index a61d50de9dc0..0aeb80fb8244 100644
--- a/libxfs/xfs_platform.h
+++ b/libxfs/xfs_platform.h
@@ -45,6 +45,7 @@
#include "platform_defs.h"
#include "xfs.h"
+#include "sizes.h"
#include "list.h"
#include "hlist.h"
#include "cache.h"
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 07/32] xfs: remove spurious XBF_DONE clearing on readahead validation failure
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (5 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 06/32] libxfs: add SZ_* constants Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-24 10:03 ` [PATCH 08/32] xfs: hide b_flags manipulation from code outside of xfs_buf.c Christoph Hellwig
` (24 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs, Carlos Maiolino
Source kernel commit: 7a4eae80b5b3db8d68961af3707fd56f2220cc29
Both callers of ->verify_read already do this, so don't duplicate the
flag manipulation.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_dquot_buf.c | 8 +++-----
libxfs/xfs_inode_buf.c | 15 ++++++++-------
2 files changed, 11 insertions(+), 12 deletions(-)
diff --git a/libxfs/xfs_dquot_buf.c b/libxfs/xfs_dquot_buf.c
index 0e59106303c4..b9a38346bdb9 100644
--- a/libxfs/xfs_dquot_buf.c
+++ b/libxfs/xfs_dquot_buf.c
@@ -251,8 +251,8 @@ xfs_dquot_buf_read_verify(
/*
* readahead errors are silent and simply leave the buffer as !done so a real
* read will then be run with the xfs_dquot_buf_ops verifier. See
- * xfs_inode_buf_verify() for why we use EIO and ~XBF_DONE here rather than
- * reporting the failure.
+ * xfs_inode_buf_verify() for why we use EIO here rather than reporting the
+ * failure.
*/
static void
xfs_dquot_buf_readahead_verify(
@@ -261,10 +261,8 @@ xfs_dquot_buf_readahead_verify(
struct xfs_mount *mp = bp->b_mount;
if (!xfs_dquot_buf_verify_crc(mp, bp, true) ||
- xfs_dquot_buf_verify(mp, bp, true) != NULL) {
+ xfs_dquot_buf_verify(mp, bp, true) != NULL)
xfs_buf_ioerror(bp, -EIO);
- bp->b_flags &= ~XBF_DONE;
- }
}
/*
diff --git a/libxfs/xfs_inode_buf.c b/libxfs/xfs_inode_buf.c
index 8ea12f82ad81..6fff9bbd39ab 100644
--- a/libxfs/xfs_inode_buf.c
+++ b/libxfs/xfs_inode_buf.c
@@ -27,12 +27,14 @@
* has not had the inode cores stamped into it. Hence for readahead, the buffer
* may be potentially invalid.
*
- * If the readahead buffer is invalid, we need to mark it with an error and
- * clear the DONE status of the buffer so that a followup read will re-read it
- * from disk. We don't report the error otherwise to avoid warnings during log
- * recovery and we don't get unnecessary panics on debug kernels. We use EIO here
- * because all we want to do is say readahead failed; there is no-one to report
- * the error to, so this will distinguish it from a non-ra verifier failure.
+ * If the readahead buffer is invalid, we need to mark it with an error so that a
+ * followup read will re-read it from disk.
+ *
+ * We don't report the error otherwise to avoid warnings during log recovery and
+ * we don't get unnecessary panics on debug kernels. Use EIO here because all
+ * we want to do is say readahead failed; there is no-one to report the error
+ * to, so this will distinguish it from a non-ra verifier failure.
+ *
* Changes to this readahead error behaviour also need to be reflected in
* xfs_dquot_buf_readahead_verify().
*/
@@ -62,7 +64,6 @@ xfs_inode_buf_verify(
if (unlikely(!di_ok ||
XFS_TEST_ERROR(mp, XFS_ERRTAG_ITOBP_INOTOBP))) {
if (readahead) {
- bp->b_flags &= ~XBF_DONE;
xfs_buf_ioerror(bp, -EIO);
return;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 08/32] xfs: hide b_flags manipulation from code outside of xfs_buf.c
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (6 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 07/32] xfs: remove spurious XBF_DONE clearing on readahead validation failure Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-24 10:03 ` [PATCH 09/32] FIXUP Christoph Hellwig
` (23 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs, Carlos Maiolino
Source kernel commit: 9f224de410d7efafc44ce1880aabd0101aa8688f
Add helpers for the remaining buffer flags manipulation not done in the
core buffer cache code.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_btree_staging.c | 7 +++----
libxfs/xfs_ialloc.c | 2 +-
2 files changed, 4 insertions(+), 5 deletions(-)
diff --git a/libxfs/xfs_btree_staging.c b/libxfs/xfs_btree_staging.c
index c3c7ea54895a..7314dab4bcfb 100644
--- a/libxfs/xfs_btree_staging.c
+++ b/libxfs/xfs_btree_staging.c
@@ -248,11 +248,10 @@ xfs_btree_bload_drop_buf(
return 0;
/*
- * Mark this buffer XBF_DONE (i.e. uptodate) so that a subsequent
- * xfs_buf_read will not pointlessly reread the contents from the disk.
+ * Mark this buffer uptodate so that a subsequent xfs_buf_read will
+ * not pointlessly reread the contents from the disk.
*/
- bp->b_flags |= XBF_DONE;
-
+ xfs_buf_set_uptodate(bp);
xfs_buf_delwri_queue_here(bp, buffers_list);
xfs_buf_relse(bp);
*bpp = NULL;
diff --git a/libxfs/xfs_ialloc.c b/libxfs/xfs_ialloc.c
index aa895c0decbc..c60e2a5ed56b 100644
--- a/libxfs/xfs_ialloc.c
+++ b/libxfs/xfs_ialloc.c
@@ -409,7 +409,7 @@ xfs_ialloc_inode_init(
xfs_trans_ordered_buf(tp, fbuf);
}
} else {
- fbuf->b_flags |= XBF_DONE;
+ xfs_buf_set_uptodate(fbuf);
xfs_buf_delwri_queue(fbuf, buffer_list);
xfs_buf_relse(fbuf);
}
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 09/32] FIXUP
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (7 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 08/32] xfs: hide b_flags manipulation from code outside of xfs_buf.c Christoph Hellwig
@ 2026-09-24 10:03 ` Christoph Hellwig
2026-09-25 23:29 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 10/32] xfs: add error injection for lazy bounce buffering Christoph Hellwig
` (22 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:03 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
---
libxfs/libxfs_io.h | 6 ++++++
libxfs/rdwr.c | 2 --
libxfs/xfs_platform.h | 1 -
3 files changed, 6 insertions(+), 3 deletions(-)
diff --git a/libxfs/libxfs_io.h b/libxfs/libxfs_io.h
index 5562e2928254..bd0d9ec04bea 100644
--- a/libxfs/libxfs_io.h
+++ b/libxfs/libxfs_io.h
@@ -293,4 +293,10 @@ xfs_buftarg_verify_daddr(
return daddr < xfs_buftarg_nr_sectors(btp);
}
+static inline void
+xfs_buf_set_uptodate(
+ struct xfs_buf *bp)
+{
+}
+
#endif /* __LIBXFS_IO_H__ */
diff --git a/libxfs/rdwr.c b/libxfs/rdwr.c
index 90f2d56687ca..9b6cbd4aa158 100644
--- a/libxfs/rdwr.c
+++ b/libxfs/rdwr.c
@@ -1426,8 +1426,6 @@ __xfs_buf_mark_corrupt(
struct xfs_buf *bp,
xfs_failaddr_t fa)
{
- ASSERT(bp->b_flags & XBF_DONE);
-
xfs_buf_corruption_error(bp, fa);
xfs_buf_stale(bp);
}
diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
index 0aeb80fb8244..d9cfcec4c1a0 100644
--- a/libxfs/xfs_platform.h
+++ b/libxfs/xfs_platform.h
@@ -322,7 +322,6 @@ static inline unsigned long long mask64_if_power2(unsigned long b)
/* buffer management */
#define XBF_TRYLOCK 0
-#define XBF_DONE 0
#define xfs_buf_stale(bp) ((bp)->b_flags |= LIBXFS_B_STALE)
#define XFS_BUF_UNDELAYWRITE(bp) ((bp)->b_flags &= ~LIBXFS_B_DIRTY)
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 10/32] xfs: add error injection for lazy bounce buffering
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (8 preceding siblings ...)
2026-09-24 10:03 ` [PATCH 09/32] FIXUP Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 11/32] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
` (21 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 021b211991604298126b47f03abe2ab234048c4c
Add an error injection knob to exercise the lazy bounce buffering
code path, i.e. to inject direct I/O re-read using the bounce buffer.
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_errortag.h | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/libxfs/xfs_errortag.h b/libxfs/xfs_errortag.h
index 6de207fed2d8..2dc441da0333 100644
--- a/libxfs/xfs_errortag.h
+++ b/libxfs/xfs_errortag.h
@@ -75,7 +75,8 @@
#define XFS_ERRTAG_METAFILE_RESV_CRITICAL 45
#define XFS_ERRTAG_FORCE_ZERO_RANGE 46
#define XFS_ERRTAG_ZONE_RESET 47
-#define XFS_ERRTAG_MAX 48
+#define XFS_ERRTAG_BOUNCE_REREAD 48
+#define XFS_ERRTAG_MAX 49
/*
* Random factors for above tags, 1 means always, 2 means 1/2 time, etc.
@@ -137,7 +138,8 @@ XFS_ERRTAG(WRITE_DELAY_MS, write_delay_ms, 3000) \
XFS_ERRTAG(EXCHMAPS_FINISH_ONE, exchmaps_finish_one, 1) \
XFS_ERRTAG(METAFILE_RESV_CRITICAL, metafile_resv_crit, 4) \
XFS_ERRTAG(FORCE_ZERO_RANGE, force_zero_range, 4) \
-XFS_ERRTAG(ZONE_RESET, zone_reset, 1)
+XFS_ERRTAG(ZONE_RESET, zone_reset, 1) \
+XFS_ERRTAG(BOUNCE_REREAD, bounce_reread, XFS_RANDOM_DEFAULT)
#endif /* XFS_ERRTAG */
#endif /* __XFS_ERRORTAG_H_ */
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 11/32] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (9 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 10/32] xfs: add error injection for lazy bounce buffering Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 12/32] FIXUP Christoph Hellwig
` (20 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 128ec674f28f15b8c4697e45a5446c545b1e6f36
Translate from a disk address to the realtime group and group-relative
block numbers. This will be needed by the data checksumming code.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_rtgroup.h | 20 ++++++++++++++++++++
1 file changed, 20 insertions(+)
diff --git a/libxfs/xfs_rtgroup.h b/libxfs/xfs_rtgroup.h
index c0b9f9f2c413..2aa6d4e59cc3 100644
--- a/libxfs/xfs_rtgroup.h
+++ b/libxfs/xfs_rtgroup.h
@@ -386,4 +386,24 @@ xfs_rtgroup_raw_size(
return g->blocks;
}
+static inline xfs_rgnumber_t
+xfs_daddr_to_rgno(struct xfs_mount *mp, xfs_daddr_t d)
+{
+ struct xfs_groups *g = &mp->m_groups[XG_TYPE_RTG];
+ xfs_rfsblock_t rbno = XFS_BB_TO_FSBT(mp, d) - g->start_fsb;
+
+ ASSERT(xfs_has_rtgroups(mp));
+ return div_u64(rbno, xfs_rtgroup_raw_size(mp));
+}
+
+static inline xfs_rgblock_t
+xfs_daddr_to_rgbno(struct xfs_mount *mp, xfs_daddr_t d)
+{
+ struct xfs_groups *g = &mp->m_groups[XG_TYPE_RTG];
+ xfs_rfsblock_t rbno = XFS_BB_TO_FSBT(mp, d) - g->start_fsb;
+
+ ASSERT(xfs_has_rtgroups(mp));
+ return do_div(rbno, xfs_rtgroup_raw_size(mp));
+}
+
#endif /* __LIBXFS_RTGROUP_H */
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 12/32] FIXUP
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (10 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 11/32] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 13/32] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
` (19 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
---
db/convert.c | 15 ---------------
1 file changed, 15 deletions(-)
diff --git a/db/convert.c b/db/convert.c
index 3eec4f224f51..9fde50ac8a91 100644
--- a/db/convert.c
+++ b/db/convert.c
@@ -39,21 +39,6 @@
#define rgnumber_to_bytes(x) \
rgblock_to_bytes((uint64_t)(x) * mp->m_groups[XG_TYPE_RTG].blocks)
-static inline xfs_rgnumber_t
-xfs_daddr_to_rgno(
- struct xfs_mount *mp,
- xfs_daddr_t daddr)
-{
- struct xfs_groups *g = &mp->m_groups[XG_TYPE_RTG];
-
- if (!xfs_has_rtgroups(mp))
- return 0;
-
- if (g->has_daddr_gaps)
- return XFS_BB_TO_FSBT(mp, daddr) / (1 << g->blklog);
- return XFS_BB_TO_FSBT(mp, daddr) / g->blocks;
-}
-
typedef enum {
CT_NONE = -1,
CT_AGBLOCK, /* xfs_agblock_t */
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 13/32] xfs: introduce XFS_BLI_PREALLOC
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (11 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 12/32] FIXUP Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 14/32] xfs: centralize setting of buf_ops/buf_type/magic for rtblocks Christoph Hellwig
` (18 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 4d5a534c06d09083d7b70801a11201ff0b5bacd2
Add a flag so that the shadow CIL buffer for a buffer log item is always
sizes to the maximum to prevent reallocations. This will be used for the
RT checksum item, where we know that we are going to fill it up very soon,
and (almost) sequentially, so there is no point in doing a constant
realloc cycle when more data is added to it.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_trans_resv.c | 2 +-
libxfs/xfs_trans_resv.h | 1 +
2 files changed, 2 insertions(+), 1 deletion(-)
diff --git a/libxfs/xfs_trans_resv.c b/libxfs/xfs_trans_resv.c
index 5b7660b1cd8a..3a9de562ac2b 100644
--- a/libxfs/xfs_trans_resv.c
+++ b/libxfs/xfs_trans_resv.c
@@ -46,7 +46,7 @@ xfs_buf_log_overhead(void)
* will be changed in a transaction. size is used to tell how many
* bytes should be reserved per item.
*/
-STATIC uint
+uint
xfs_calc_buf_res(
uint nbufs,
uint size)
diff --git a/libxfs/xfs_trans_resv.h b/libxfs/xfs_trans_resv.h
index 336279e0fc61..1804e821f382 100644
--- a/libxfs/xfs_trans_resv.h
+++ b/libxfs/xfs_trans_resv.h
@@ -96,6 +96,7 @@ struct xfs_trans_resv {
#define XFS_ITRUNCATE_LOG_COUNT_REFLINK 8
#define XFS_WRITE_LOG_COUNT_REFLINK 8
+uint xfs_calc_buf_res(uint nbufs, uint size);
void xfs_trans_resv_calc(struct xfs_mount *mp, struct xfs_trans_resv *resp);
uint xfs_allocfree_block_count(struct xfs_mount *mp, uint num_ops);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 14/32] xfs: centralize setting of buf_ops/buf_type/magic for rtblocks
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (12 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 13/32] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 15/32] xfs: factor out a xfs_rtfile_initialize_buf helper Christoph Hellwig
` (17 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 1dbf45a26cdbb2a120b3b135cc298ce92ba3f04a
Add two tables for the buf_ops and buf_type, and derive the magic from
the buf_ops to make have a single source of truth for the different RT
block variants. This cleans up the existing code and makes adding another
type of block/file easier.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_rtbitmap.c | 40 +++++++++++++++++++++++-----------------
libxfs/xfs_rtbitmap.h | 15 ++-------------
2 files changed, 25 insertions(+), 30 deletions(-)
diff --git a/libxfs/xfs_rtbitmap.c b/libxfs/xfs_rtbitmap.c
index 6a0fc1032753..dda6ef84aeac 100644
--- a/libxfs/xfs_rtbitmap.c
+++ b/libxfs/xfs_rtbitmap.c
@@ -122,6 +122,26 @@ const struct xfs_buf_ops xfs_rtsummary_buf_ops = {
.verify_struct = xfs_rtbuf_verify,
};
+static const struct xfs_buf_ops *xfs_rtblock_buf_ops[XFS_RTGI_MAX] = {
+ [XFS_RTGI_SUMMARY] = &xfs_rtsummary_buf_ops,
+ [XFS_RTGI_BITMAP] = &xfs_rtbitmap_buf_ops,
+};
+
+const struct xfs_buf_ops *
+xfs_rtblock_ops(
+ struct xfs_mount *mp,
+ enum xfs_rtg_inodes type)
+{
+ if (!xfs_has_rtgroups(mp))
+ return &xfs_rtbuf_ops;
+ return xfs_rtblock_buf_ops[type];
+}
+
+static enum xfs_blft xfs_rtblock_buf_types[XFS_RTGI_MAX] = {
+ [XFS_RTGI_SUMMARY] = XFS_BLFT_RTSUMMARY_BUF,
+ [XFS_RTGI_BITMAP] = XFS_BLFT_RTBITMAP_BUF,
+};
+
/* Release cached rt bitmap and summary buffers. */
void
xfs_rtbuf_cache_relse(
@@ -155,7 +175,6 @@ xfs_rtbuf_get(
xfs_fileoff_t *coffp; /* cached block number */
struct xfs_buf *bp; /* block buffer, result */
struct xfs_bmbt_irec map;
- enum xfs_blft buf_type;
int nmap = 1;
int error;
@@ -163,12 +182,10 @@ xfs_rtbuf_get(
case XFS_RTGI_SUMMARY:
cbpp = &args->sumbp;
coffp = &args->sumoff;
- buf_type = XFS_BLFT_RTSUMMARY_BUF;
break;
case XFS_RTGI_BITMAP:
cbpp = &args->rbmbp;
coffp = &args->rbmoff;
- buf_type = XFS_BLFT_RTBITMAP_BUF;
break;
default:
return -EINVAL;
@@ -219,7 +236,7 @@ xfs_rtbuf_get(
}
}
- xfs_trans_buf_set_type(args->tp, bp, buf_type);
+ xfs_trans_buf_set_type(args->tp, bp, xfs_rtblock_buf_types[type]);
*cbpp = bp;
*coffp = block;
return 0;
@@ -1372,16 +1389,8 @@ xfs_rtfile_initialize_block(
struct xfs_buf *bp;
void *bufdata;
const size_t copylen = mp->m_blockwsize << XFS_WORDLOG;
- enum xfs_blft buf_type;
int error;
- if (type == XFS_RTGI_BITMAP)
- buf_type = XFS_BLFT_RTBITMAP_BUF;
- else if (type == XFS_RTGI_SUMMARY)
- buf_type = XFS_BLFT_RTSUMMARY_BUF;
- else
- return -EINVAL;
-
error = xfs_trans_alloc(mp, &M_RES(mp)->tr_growrtzero, 0, 0, 0, &tp);
if (error)
return error;
@@ -1396,16 +1405,13 @@ xfs_rtfile_initialize_block(
}
bufdata = bp->b_addr;
- xfs_trans_buf_set_type(tp, bp, buf_type);
+ xfs_trans_buf_set_type(tp, bp, xfs_rtblock_buf_types[type]);
bp->b_ops = xfs_rtblock_ops(mp, type);
if (xfs_has_rtgroups(mp)) {
struct xfs_rtbuf_blkinfo *hdr = bp->b_addr;
- if (type == XFS_RTGI_BITMAP)
- hdr->rt_magic = cpu_to_be32(XFS_RTBITMAP_MAGIC);
- else
- hdr->rt_magic = cpu_to_be32(XFS_RTSUMMARY_MAGIC);
+ hdr->rt_magic = bp->b_ops->magic[1];
hdr->rt_owner = cpu_to_be64(I_INO(ip));
hdr->rt_blkno = cpu_to_be64(XFS_FSB_TO_DADDR(mp, fsbno));
hdr->rt_lsn = 0;
diff --git a/libxfs/xfs_rtbitmap.h b/libxfs/xfs_rtbitmap.h
index 22e5d9cd95f4..375cc48e1a53 100644
--- a/libxfs/xfs_rtbitmap.h
+++ b/libxfs/xfs_rtbitmap.h
@@ -354,19 +354,6 @@ xfs_suminfo_add(
return info->old;
}
-static inline const struct xfs_buf_ops *
-xfs_rtblock_ops(
- struct xfs_mount *mp,
- enum xfs_rtg_inodes type)
-{
- if (xfs_has_rtgroups(mp)) {
- if (type == XFS_RTGI_SUMMARY)
- return &xfs_rtsummary_buf_ops;
- return &xfs_rtbitmap_buf_ops;
- }
- return &xfs_rtbuf_ops;
-}
-
/*
* Functions for walking free space rtextents in the realtime bitmap.
*/
@@ -419,6 +406,8 @@ xfs_filblks_t xfs_rtbitmap_blockcount_len(struct xfs_mount *mp,
xfs_filblks_t xfs_rtsummary_blockcount(struct xfs_mount *mp,
unsigned int *rsumlevels);
+const struct xfs_buf_ops *xfs_rtblock_ops(struct xfs_mount *mp,
+ enum xfs_rtg_inodes type);
int xfs_rtfile_initialize_blocks(struct xfs_rtgroup *rtg,
enum xfs_rtg_inodes type, xfs_fileoff_t offset_fsb,
xfs_fileoff_t end_fsb, void *data);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 15/32] xfs: factor out a xfs_rtfile_initialize_buf helper
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (13 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 14/32] xfs: centralize setting of buf_ops/buf_type/magic for rtblocks Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 16/32] xfs: add a xfs_rtblock_payload helper Christoph Hellwig
` (16 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: eb8ab10dddf6e7fbfe602f827f98afe9723e5384
Share the code to initialize the header and buf ops for rtfile blocks
into a single helper.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_rtbitmap.c | 38 ++++++++++++++++++++++++--------------
libxfs/xfs_rtbitmap.h | 3 +++
2 files changed, 27 insertions(+), 14 deletions(-)
diff --git a/libxfs/xfs_rtbitmap.c b/libxfs/xfs_rtbitmap.c
index dda6ef84aeac..245cdd9c22ca 100644
--- a/libxfs/xfs_rtbitmap.c
+++ b/libxfs/xfs_rtbitmap.c
@@ -1375,6 +1375,27 @@ out_trans_cancel:
return error;
}
+void
+xfs_rtfile_initialize_buf(
+ struct xfs_rtgroup *rtg,
+ enum xfs_rtg_inodes type,
+ struct xfs_buf *bp,
+ struct xfs_trans *tp)
+{
+ bp->b_ops = xfs_rtblock_ops(bp->b_mount, type);
+ if (tp)
+ xfs_trans_buf_set_type(tp, bp, xfs_rtblock_buf_types[type]);
+ if (xfs_has_rtgroups(bp->b_mount)) {
+ struct xfs_rtbuf_blkinfo *hdr = bp->b_addr;
+
+ hdr->rt_magic = bp->b_ops->magic[1];
+ hdr->rt_owner = cpu_to_be64(I_INO(rtg->rtg_inodes[type]));
+ hdr->rt_blkno = cpu_to_be64(xfs_buf_daddr(bp));
+ hdr->rt_lsn = 0;
+ uuid_copy(&hdr->rt_uuid, &bp->b_mount->m_sb.sb_meta_uuid);
+ }
+}
+
/* Get a buffer for the block. */
static int
xfs_rtfile_initialize_block(
@@ -1405,21 +1426,10 @@ xfs_rtfile_initialize_block(
}
bufdata = bp->b_addr;
- xfs_trans_buf_set_type(tp, bp, xfs_rtblock_buf_types[type]);
- bp->b_ops = xfs_rtblock_ops(mp, type);
-
- if (xfs_has_rtgroups(mp)) {
- struct xfs_rtbuf_blkinfo *hdr = bp->b_addr;
-
- hdr->rt_magic = bp->b_ops->magic[1];
- hdr->rt_owner = cpu_to_be64(I_INO(ip));
- hdr->rt_blkno = cpu_to_be64(XFS_FSB_TO_DADDR(mp, fsbno));
- hdr->rt_lsn = 0;
- uuid_copy(&hdr->rt_uuid, &mp->m_sb.sb_meta_uuid);
-
- bufdata += sizeof(*hdr);
- }
+ xfs_rtfile_initialize_buf(rtg, type, bp, tp);
+ if (xfs_has_rtgroups(mp))
+ bufdata += sizeof(struct xfs_rtbuf_blkinfo);
if (data)
memcpy(bufdata, data, copylen);
else
diff --git a/libxfs/xfs_rtbitmap.h b/libxfs/xfs_rtbitmap.h
index 375cc48e1a53..4a87e1fd3e99 100644
--- a/libxfs/xfs_rtbitmap.h
+++ b/libxfs/xfs_rtbitmap.h
@@ -408,6 +408,9 @@ xfs_filblks_t xfs_rtsummary_blockcount(struct xfs_mount *mp,
const struct xfs_buf_ops *xfs_rtblock_ops(struct xfs_mount *mp,
enum xfs_rtg_inodes type);
+void xfs_rtfile_initialize_buf(struct xfs_rtgroup *rtg,
+ enum xfs_rtg_inodes type, struct xfs_buf *bp,
+ struct xfs_trans *tp);
int xfs_rtfile_initialize_blocks(struct xfs_rtgroup *rtg,
enum xfs_rtg_inodes type, xfs_fileoff_t offset_fsb,
xfs_fileoff_t end_fsb, void *data);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 16/32] xfs: add a xfs_rtblock_payload helper
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (14 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 15/32] xfs: factor out a xfs_rtfile_initialize_buf helper Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 17/32] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
` (15 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: c63d609c1ed6dd49664e4f822a70abaf2572b35b
Add a helper to calculate the rtblock payload start with or without the
self-describing metadata header.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_rtbitmap.c | 9 ++-------
libxfs/xfs_rtbitmap.h | 27 ++++++++++++---------------
2 files changed, 14 insertions(+), 22 deletions(-)
diff --git a/libxfs/xfs_rtbitmap.c b/libxfs/xfs_rtbitmap.c
index 245cdd9c22ca..39a6b45d9526 100644
--- a/libxfs/xfs_rtbitmap.c
+++ b/libxfs/xfs_rtbitmap.c
@@ -1408,7 +1408,6 @@ xfs_rtfile_initialize_block(
struct xfs_inode *ip = rtg->rtg_inodes[type];
struct xfs_trans *tp;
struct xfs_buf *bp;
- void *bufdata;
const size_t copylen = mp->m_blockwsize << XFS_WORDLOG;
int error;
@@ -1424,16 +1423,12 @@ xfs_rtfile_initialize_block(
xfs_trans_cancel(tp);
return error;
}
- bufdata = bp->b_addr;
xfs_rtfile_initialize_buf(rtg, type, bp, tp);
-
- if (xfs_has_rtgroups(mp))
- bufdata += sizeof(struct xfs_rtbuf_blkinfo);
if (data)
- memcpy(bufdata, data, copylen);
+ memcpy(xfs_rtblock_payload(bp), data, copylen);
else
- memset(bufdata, 0, copylen);
+ memset(xfs_rtblock_payload(bp), 0, copylen);
xfs_trans_log_buf(tp, bp, 0, mp->m_sb.sb_blocksize - 1);
return xfs_trans_commit(tp);
}
diff --git a/libxfs/xfs_rtbitmap.h b/libxfs/xfs_rtbitmap.h
index 4a87e1fd3e99..750d74fbf4ed 100644
--- a/libxfs/xfs_rtbitmap.h
+++ b/libxfs/xfs_rtbitmap.h
@@ -20,6 +20,16 @@ struct xfs_rtalloc_args {
xfs_fileoff_t sumoff; /* summary block number */
};
+/* Return the payload of the buffer after the optional header. */
+static inline void *
+xfs_rtblock_payload(
+ struct xfs_buf *bp)
+{
+ if (!xfs_has_rtgroups(bp->b_mount))
+ return bp->b_addr;
+ return bp->b_addr + sizeof(struct xfs_rtbuf_blkinfo);
+}
+
static inline xfs_rtblock_t
xfs_rtx_to_rtb(
struct xfs_rtgroup *rtg,
@@ -221,14 +231,7 @@ xfs_rbmblock_wordptr(
struct xfs_rtalloc_args *args,
unsigned int index)
{
- struct xfs_mount *mp = args->mp;
- union xfs_rtword_raw *words;
- struct xfs_rtbuf_blkinfo *hdr = args->rbmbp->b_addr;
-
- if (xfs_has_rtgroups(mp))
- words = (union xfs_rtword_raw *)(hdr + 1);
- else
- words = args->rbmbp->b_addr;
+ union xfs_rtword_raw *words = xfs_rtblock_payload(args->rbmbp);
return words + index;
}
@@ -312,13 +315,7 @@ xfs_rsumblock_infoptr(
struct xfs_rtalloc_args *args,
unsigned int index)
{
- union xfs_suminfo_raw *info;
- struct xfs_rtbuf_blkinfo *hdr = args->sumbp->b_addr;
-
- if (xfs_has_rtgroups(args->mp))
- info = (union xfs_suminfo_raw *)(hdr + 1);
- else
- info = args->sumbp->b_addr;
+ union xfs_suminfo_raw *info = xfs_rtblock_payload(args->sumbp);
return info + index;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 17/32] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (15 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 16/32] xfs: add a xfs_rtblock_payload helper Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 18/32] FIXUP Christoph Hellwig
` (14 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: e2ab369189d3a62e28b9f7a1de255a21bc8b1534
The upcoming RT data checksum feature will use larger than FSB blocks.
Prepare xfs_rtfile_initialize_blocks to pass the number of FSBs per
RT blocks, and to pass bmapi_flags to ask for contiguous allocation.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_rtbitmap.c | 44 +++++++++++++++++++++++++------------------
libxfs/xfs_rtbitmap.h | 3 ++-
2 files changed, 28 insertions(+), 19 deletions(-)
diff --git a/libxfs/xfs_rtbitmap.c b/libxfs/xfs_rtbitmap.c
index 39a6b45d9526..3abc495c2a3c 100644
--- a/libxfs/xfs_rtbitmap.c
+++ b/libxfs/xfs_rtbitmap.c
@@ -1343,6 +1343,7 @@ xfs_rtfile_alloc_blocks(
struct xfs_inode *ip,
xfs_fileoff_t offset_fsb,
xfs_filblks_t count_fsb,
+ uint32_t bmapi_flags,
struct xfs_bmbt_irec *map)
{
struct xfs_mount *mp = ip->i_mount;
@@ -1364,7 +1365,7 @@ xfs_rtfile_alloc_blocks(
goto out_trans_cancel;
error = xfs_bmapi_write(tp, ip, offset_fsb, count_fsb,
- XFS_BMAPI_METADATA, 0, map, &nmap);
+ XFS_BMAPI_METADATA | bmapi_flags, 0, map, &nmap);
if (error)
goto out_trans_cancel;
@@ -1402,34 +1403,43 @@ xfs_rtfile_initialize_block(
struct xfs_rtgroup *rtg,
enum xfs_rtg_inodes type,
xfs_fsblock_t fsbno,
- void *data)
+ xfs_filblks_t nblks,
+ void **data)
{
struct xfs_mount *mp = rtg_mount(rtg);
struct xfs_inode *ip = rtg->rtg_inodes[type];
+ size_t len = XFS_FSB_TO_B(mp, nblks);
+ size_t copylen = len;
+ struct xfs_trans_res tres = M_RES(mp)->tr_growrtzero;
struct xfs_trans *tp;
struct xfs_buf *bp;
- const size_t copylen = mp->m_blockwsize << XFS_WORDLOG;
int error;
- error = xfs_trans_alloc(mp, &M_RES(mp)->tr_growrtzero, 0, 0, 0, &tp);
+ tres.tr_logres *= nblks;
+ error = xfs_trans_alloc(mp, &tres, 0, 0, 0, &tp);
if (error)
return error;
xfs_ilock(ip, XFS_ILOCK_EXCL);
xfs_trans_ijoin(tp, ip, XFS_ILOCK_EXCL);
error = xfs_trans_get_buf(tp, mp->m_ddev_targp,
- XFS_FSB_TO_DADDR(mp, fsbno), mp->m_bsize, 0, &bp);
+ XFS_FSB_TO_DADDR(mp, fsbno), BTOBB(len), 0, &bp);
if (error) {
xfs_trans_cancel(tp);
return error;
}
+ if (xfs_has_rtgroups(mp))
+ copylen -= sizeof(struct xfs_rtbuf_blkinfo);
+
xfs_rtfile_initialize_buf(rtg, type, bp, tp);
- if (data)
- memcpy(xfs_rtblock_payload(bp), data, copylen);
- else
+ if (*data) {
+ memcpy(xfs_rtblock_payload(bp), *data, copylen);
+ *data += copylen;
+ } else {
memset(xfs_rtblock_payload(bp), 0, copylen);
- xfs_trans_log_buf(tp, bp, 0, mp->m_sb.sb_blocksize - 1);
+ }
+ xfs_trans_log_buf(tp, bp, 0, len - 1);
return xfs_trans_commit(tp);
}
@@ -1444,33 +1454,31 @@ xfs_rtfile_initialize_blocks(
enum xfs_rtg_inodes type,
xfs_fileoff_t offset_fsb, /* offset to start from */
xfs_fileoff_t end_fsb, /* offset to allocate to */
+ xfs_filblks_t bsize,
+ uint32_t bmapi_flags,
void *data) /* data to fill the blocks */
{
- struct xfs_mount *mp = rtg_mount(rtg);
- const size_t copylen = mp->m_blockwsize << XFS_WORDLOG;
-
while (offset_fsb < end_fsb) {
struct xfs_bmbt_irec map;
xfs_filblks_t i;
int error;
error = xfs_rtfile_alloc_blocks(rtg->rtg_inodes[type],
- offset_fsb, end_fsb - offset_fsb, &map);
+ offset_fsb, end_fsb - offset_fsb, bmapi_flags,
+ &map);
if (error)
return error;
/*
- * Now we need to clear the allocated blocks.
+ * Now we need to clear or initialize the allocated blocks.
*
* Do this one block per transaction, to keep it simple.
*/
- for (i = 0; i < map.br_blockcount; i++) {
+ for (i = 0; i < map.br_blockcount; i += bsize) {
error = xfs_rtfile_initialize_block(rtg, type,
- map.br_startblock + i, data);
+ map.br_startblock + i, bsize, &data);
if (error)
return error;
- if (data)
- data += copylen;
}
offset_fsb = map.br_startoff + map.br_blockcount;
diff --git a/libxfs/xfs_rtbitmap.h b/libxfs/xfs_rtbitmap.h
index 750d74fbf4ed..e9e3378d15aa 100644
--- a/libxfs/xfs_rtbitmap.h
+++ b/libxfs/xfs_rtbitmap.h
@@ -410,7 +410,8 @@ void xfs_rtfile_initialize_buf(struct xfs_rtgroup *rtg,
struct xfs_trans *tp);
int xfs_rtfile_initialize_blocks(struct xfs_rtgroup *rtg,
enum xfs_rtg_inodes type, xfs_fileoff_t offset_fsb,
- xfs_fileoff_t end_fsb, void *data);
+ xfs_fileoff_t end_fsb, xfs_filblks_t bsize,
+ uint32_t bmapi_flags, void *data);
int xfs_rtbitmap_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
struct xfs_trans *tp, bool init);
int xfs_rtsummary_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 18/32] FIXUP
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (16 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 17/32] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 19/32] xfs: define the RT data checksum on-disk format Christoph Hellwig
` (13 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
---
mkfs/proto.c | 4 ++--
repair/rt.c | 4 ++--
2 files changed, 4 insertions(+), 4 deletions(-)
diff --git a/mkfs/proto.c b/mkfs/proto.c
index bdd0fadda517..7ef8cc9d88a1 100644
--- a/mkfs/proto.c
+++ b/mkfs/proto.c
@@ -1099,12 +1099,12 @@ rtfreesp_init(
* First zero the realtime bitmap and summary files.
*/
error = -libxfs_rtfile_initialize_blocks(rtg, XFS_RTGI_BITMAP, 0,
- mp->m_sb.sb_rbmblocks, NULL);
+ mp->m_sb.sb_rbmblocks, 1, 0, NULL);
if (error)
fail(_("Initialization of rtbitmap inode failed"), error);
error = -libxfs_rtfile_initialize_blocks(rtg, XFS_RTGI_SUMMARY, 0,
- mp->m_rsumblocks, NULL);
+ mp->m_rsumblocks, 1, 0, NULL);
if (error)
fail(_("Initialization of rtsummary inode failed"), error);
diff --git a/repair/rt.c b/repair/rt.c
index 3e51c9b5eb4b..b0ff775bd339 100644
--- a/repair/rt.c
+++ b/repair/rt.c
@@ -381,7 +381,7 @@ fill_rtbitmap(
return;
error = -libxfs_rtfile_initialize_blocks(rtg, XFS_RTGI_BITMAP,
- 0, rtg_mount(rtg)->m_sb.sb_rbmblocks,
+ 0, rtg_mount(rtg)->m_sb.sb_rbmblocks, 1, 0,
rt_computed[rtg_rgno(rtg)].bmp);
if (error)
do_error(
@@ -403,7 +403,7 @@ fill_rtsummary(
return;
error = -libxfs_rtfile_initialize_blocks(rtg, XFS_RTGI_SUMMARY,
- 0, rtg_mount(rtg)->m_rsumblocks,
+ 0, rtg_mount(rtg)->m_rsumblocks, 1, 0,
rt_computed[rtg_rgno(rtg)].sum);
if (error)
do_error(
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 19/32] xfs: define the RT data checksum on-disk format
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (17 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 18/32] FIXUP Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 20/32] FIXUP Christoph Hellwig
` (12 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 5f420198c30c87b4273bc18b1d85ce4f340a1d74
Add the on-disk format for the new RT data checksum format.
Keyed off a new read-only compat feature flag, this adds new fields to
the superblock to indicate the checksum algorithm used and the size of
the blocks containing the checksums. These new fields reuse the
previously reserved padding to make efficient use of the space in the
on-disk superblock.
Data checksums are only supported on zoned RT devices, because they
require out of places writes to safely update the checksums for file
overwrites and a data/metadata split to be able to store the checksums
for a group in a file without causing recursion. This means they can't
be supported directly on the data device at all, and only when using
the always_cow mode on regular RT devices, but that has no benefit
over the zoned allocator which is designed for out of place writes.
The initially supported data checksum algorithms are crc32c and crc64 as
specified by NVMe. Both have extremely fast kernel implementations and
the strong data protection guarantees offered by CRC-style algorithms.
Both also happen to be support by NVMe for protection information so that
the userspace PI passthrough support (once extended to files on file
systems) can be reused to expose the checksums to applications and thus
provide true end-to-end data integrity.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_format.h | 43 +++++++++++++++++++++--
libxfs/xfs_log_format.h | 1 +
libxfs/xfs_ondisk.h | 4 ++-
libxfs/xfs_sb.c | 75 +++++++++++++++++++++++++++++++++++++++++
libxfs/xfs_sb.h | 1 +
5 files changed, 120 insertions(+), 4 deletions(-)
diff --git a/libxfs/xfs_format.h b/libxfs/xfs_format.h
index dd0ed046fbe9..1be3d21910a7 100644
--- a/libxfs/xfs_format.h
+++ b/libxfs/xfs_format.h
@@ -179,7 +179,9 @@ typedef struct xfs_sb {
xfs_rgnumber_t sb_rgcount; /* number of realtime groups */
xfs_rtxlen_t sb_rgextents; /* size of a realtime group in rtx */
uint8_t sb_rgblklog; /* rt group number shift */
- uint8_t sb_pad[7]; /* zeroes */
+ uint8_t sb_rtcsum_type; /* RT device data checksum type */
+ uint8_t sb_rtcsum_blklog; /* log2 of rtcsum bsize */
+ uint8_t sb_pad[5]; /* zero */
xfs_rfsblock_t sb_rtstart; /* start of internal RT section (FSB) */
xfs_filblks_t sb_rtreserved; /* reserved (zoned) RT blocks */
@@ -272,7 +274,9 @@ struct xfs_dsb {
__be32 sb_rgcount; /* # of realtime groups */
__be32 sb_rgextents; /* size of rtgroup in rtx */
__u8 sb_rgblklog; /* rt group number shift */
- __u8 sb_pad[7]; /* zeroes */
+ __u8 sb_rtcsum_type; /* RT device data checksum type */
+ __u8 sb_rtcsum_blklog; /* log2 of rtcsum bsize */
+ __u8 sb_pad[5]; /* zero */
__be64 sb_rtstart; /* start of internal RT section (FSB) */
__be64 sb_rtreserved; /* reserved (zoned) RT blocks */
@@ -374,6 +378,8 @@ xfs_sb_has_compat_feature(
#define XFS_SB_FEAT_RO_COMPAT_RMAPBT (1 << 1) /* reverse map btree */
#define XFS_SB_FEAT_RO_COMPAT_REFLINK (1 << 2) /* reflinked files */
#define XFS_SB_FEAT_RO_COMPAT_INOBTCNT (1 << 3) /* inobt block counts */
+#define XFS_SB_FEAT_RO_COMPAT_RTCSUM (1 << 5) /* RT data checksums */
+
#define XFS_SB_FEAT_RO_COMPAT_ALL \
(XFS_SB_FEAT_RO_COMPAT_FINOBT | \
XFS_SB_FEAT_RO_COMPAT_RMAPBT | \
@@ -866,6 +872,7 @@ enum xfs_metafile_type {
XFS_METAFILE_RTSUMMARY, /* rt summary */
XFS_METAFILE_RTRMAP, /* rt rmap */
XFS_METAFILE_RTREFCOUNT, /* rt refcount */
+ XFS_METAFILE_RTCSUM, /* rt data checksums */
XFS_METAFILE_MAX
} __packed;
@@ -879,7 +886,8 @@ enum xfs_metafile_type {
{ XFS_METAFILE_RTBITMAP, "rtbitmap" }, \
{ XFS_METAFILE_RTSUMMARY, "rtsummary" }, \
{ XFS_METAFILE_RTRMAP, "rtrmap" }, \
- { XFS_METAFILE_RTREFCOUNT, "rtrefcount" }
+ { XFS_METAFILE_RTREFCOUNT, "rtrefcount" }, \
+ { XFS_METAFILE_RTCSUM, "rtcsum", }
/*
* On-disk inode structure.
@@ -1318,6 +1326,7 @@ static inline bool xfs_dinode_is_metadir(const struct xfs_dinode *dip)
*/
#define XFS_RTBITMAP_MAGIC 0x424D505A /* BMPZ */
#define XFS_RTSUMMARY_MAGIC 0x53554D59 /* SUMY */
+#define XFS_RTCSUM_MAGIC 0x4353554D /* CSUM */
struct xfs_rtbuf_blkinfo {
__be32 rt_magic; /* validity check on block */
@@ -2027,4 +2036,32 @@ struct xfs_acl {
#define SGI_ACL_FILE_SIZE (sizeof(SGI_ACL_FILE)-1)
#define SGI_ACL_DEFAULT_SIZE (sizeof(SGI_ACL_DEFAULT)-1)
+/*
+ * Size of a RT data checksum block. Data reads must be contained in a single
+ * block, so this should be fairly large.
+ *
+ * The default is 32k, matching the default inode cluster size and the maximum
+ * memory allocation the Linux MM can handle in the fast path. 64k is primarily
+ * there so that his value never needs to be below the FSB size, even for 64k
+ * blocks.
+ */
+#define XFS_RTCSUM_BSIZE_LOG_MIN 15
+#define XFS_RTCSUM_BSIZE_LOG_MAX 16
+
+/*
+ * Data checksum types.
+ */
+#define XFS_CSUM_TYPE_NONE 0u
+#define XFS_CSUM_TYPE_CRC32C 1u
+#define XFS_CSUM_TYPE_CRC64 2u
+#define XFS_CSUM_TYPE_MAX 3u
+
+/*
+ * On-disk data checksums.
+ */
+union xfs_disk_csum {
+ __le32 crc32c;
+ __le64 crc64;
+};
+
#endif /* __XFS_FORMAT_H__ */
diff --git a/libxfs/xfs_log_format.h b/libxfs/xfs_log_format.h
index a4e1b3eb425c..b1037b77338b 100644
--- a/libxfs/xfs_log_format.h
+++ b/libxfs/xfs_log_format.h
@@ -581,6 +581,7 @@ enum xfs_blft {
XFS_BLFT_SB_BUF,
XFS_BLFT_RTBITMAP_BUF,
XFS_BLFT_RTSUMMARY_BUF,
+ XFS_BLFT_RTCSUM_BUF,
XFS_BLFT_MAX_BUF = (1 << XFS_BLFT_BITS),
};
diff --git a/libxfs/xfs_ondisk.h b/libxfs/xfs_ondisk.h
index 23cde1248f01..17ab9366b3b9 100644
--- a/libxfs/xfs_ondisk.h
+++ b/libxfs/xfs_ondisk.h
@@ -284,7 +284,9 @@ xfs_check_ondisk_structs(void)
XFS_CHECK_SB_OFFSET(sb_rgcount, 272);
XFS_CHECK_SB_OFFSET(sb_rgextents, 276);
XFS_CHECK_SB_OFFSET(sb_rgblklog, 280);
- XFS_CHECK_SB_OFFSET(sb_pad, 281);
+ XFS_CHECK_SB_OFFSET(sb_rtcsum_type, 281);
+ XFS_CHECK_SB_OFFSET(sb_rtcsum_blklog, 282);
+ XFS_CHECK_SB_OFFSET(sb_pad, 283);
XFS_CHECK_SB_OFFSET(sb_rtstart, 288);
XFS_CHECK_SB_OFFSET(sb_rtreserved, 296);
diff --git a/libxfs/xfs_sb.c b/libxfs/xfs_sb.c
index cccbcd153316..2d64412363b4 100644
--- a/libxfs/xfs_sb.c
+++ b/libxfs/xfs_sb.c
@@ -484,6 +484,40 @@ xfs_validate_sb_zoned(
return 0;
}
+static int
+xfs_validate_sb_csum(
+ struct xfs_mount *mp,
+ struct xfs_sb *sbp)
+{
+ unsigned int rtcsum_bsize = 1u << sbp->sb_rtcsum_blklog;
+
+ if (!(sbp->sb_features_incompat & XFS_SB_FEAT_INCOMPAT_ZONED)) {
+ xfs_warn(mp, "data checksum required the zone allocator");
+ return -EINVAL;
+ }
+ if (sbp->sb_rtcsum_type >= XFS_CSUM_TYPE_MAX) {
+ xfs_warn(mp, "invalid data checksum type: 0x%x",
+ sbp->sb_rtcsum_type);
+ return -EINVAL;
+ }
+ if (sbp->sb_rtcsum_blklog < XFS_RTCSUM_BSIZE_LOG_MIN ||
+ sbp->sb_rtcsum_blklog > XFS_RTCSUM_BSIZE_LOG_MAX) {
+ xfs_warn(mp,
+"invalid data checksum block log: %u (min %u/max %u)",
+ sbp->sb_rtcsum_blklog,
+ XFS_RTCSUM_BSIZE_LOG_MIN,
+ XFS_RTCSUM_BSIZE_LOG_MAX);
+ return -EINVAL;
+ }
+ if (rtcsum_bsize < sbp->sb_blocksize) {
+ xfs_warn(mp,
+"checksum block size must not be smaller than file system block size: %u/%u",
+ rtcsum_bsize, sbp->sb_blocksize);
+ return -EINVAL;
+ }
+ return 0;
+}
+
/* Check the validity of the SB. */
STATIC int
xfs_validate_sb_common(
@@ -577,6 +611,17 @@ xfs_validate_sb_common(
if (error)
return error;
}
+ if (sbp->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
+ error = xfs_validate_sb_csum(mp, sbp);
+ if (error)
+ return error;
+ } else {
+ if (sbp->sb_rtcsum_type || sbp->sb_rtcsum_blklog) {
+ xfs_warn(mp,
+"rtcsum superblock fields must be zero for non-RTCSUM file systems.");
+ return -EINVAL;
+ }
+ }
} else if (sbp->sb_qflags & (XFS_PQUOTA_ENFD | XFS_GQUOTA_ENFD |
XFS_PQUOTA_CHKD | XFS_GQUOTA_CHKD)) {
xfs_notice(mp,
@@ -897,6 +942,14 @@ __xfs_sb_from_disk(
to->sb_rtstart = 0;
to->sb_rtreserved = 0;
}
+
+ if (to->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
+ to->sb_rtcsum_type = from->sb_rtcsum_type;
+ to->sb_rtcsum_blklog = from->sb_rtcsum_blklog;
+ } else {
+ to->sb_rtcsum_type = XFS_CSUM_TYPE_NONE;
+ to->sb_rtcsum_blklog = 0;
+ }
}
void
@@ -1068,6 +1121,11 @@ xfs_sb_to_disk(
to->sb_rtstart = cpu_to_be64(from->sb_rtstart);
to->sb_rtreserved = cpu_to_be64(from->sb_rtreserved);
}
+
+ if (from->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM) {
+ to->sb_rtcsum_type = from->sb_rtcsum_type;
+ to->sb_rtcsum_blklog = from->sb_rtcsum_blklog;
+ }
}
/*
@@ -1240,6 +1298,20 @@ xfs_mount_sb_set_rextsize(
xfs_sb_mount_rextsize(mp, sbp);
}
+uint8_t
+xfs_data_csum_shift(
+ uint8_t csum)
+{
+ switch (csum) {
+ case XFS_CSUM_TYPE_CRC32C:
+ return 2;
+ case XFS_CSUM_TYPE_CRC64:
+ return 3;
+ default:
+ return 0;
+ }
+}
+
/*
* xfs_mount_common
*
@@ -1308,6 +1380,9 @@ xfs_sb_mount_common(
mp->m_bsize = XFS_FSB_TO_BB(mp, 1);
mp->m_alloc_set_aside = xfs_alloc_set_aside(mp);
mp->m_ag_max_usable = xfs_alloc_ag_max_usable(mp);
+
+ mp->m_rtcsum_shift = xfs_data_csum_shift(mp->m_sb.sb_rtcsum_type);
+ mp->m_rtcsum_bsize = 1u << mp->m_sb.sb_rtcsum_blklog;
}
/*
diff --git a/libxfs/xfs_sb.h b/libxfs/xfs_sb.h
index 34d0dd374e9b..16f300c12e37 100644
--- a/libxfs/xfs_sb.h
+++ b/libxfs/xfs_sb.h
@@ -20,6 +20,7 @@ extern void xfs_sb_mount_common(struct xfs_mount *mp, struct xfs_sb *sbp);
void xfs_sb_mount_rextsize(struct xfs_mount *mp, struct xfs_sb *sbp);
void xfs_mount_sb_set_rextsize(struct xfs_mount *mp,
struct xfs_sb *sbp, xfs_agblock_t rextsize);
+uint8_t xfs_data_csum_shift(uint8_t csum);
extern void xfs_sb_from_disk(struct xfs_sb *to, struct xfs_dsb *from);
extern void xfs_sb_to_disk(struct xfs_dsb *to, struct xfs_sb *from);
extern void xfs_sb_quota_from_disk(struct xfs_sb *sbp);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 20/32] FIXUP
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (18 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 19/32] xfs: define the RT data checksum on-disk format Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 21/32] xfs: add support for per-RTG csum files Christoph Hellwig
` (11 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
---
libxfs/stubs/xfs_mount.h | 7 +++++++
| 6 ++++++
repair/sb.c | 13 +++++++++++++
3 files changed, 26 insertions(+)
diff --git a/libxfs/stubs/xfs_mount.h b/libxfs/stubs/xfs_mount.h
index 5a714333c16e..838746fa16d8 100644
--- a/libxfs/stubs/xfs_mount.h
+++ b/libxfs/stubs/xfs_mount.h
@@ -111,6 +111,8 @@ typedef struct xfs_mount {
uint8_t m_sectbb_log; /* sectorlog - BBSHIFT */
uint8_t m_agno_log; /* log #ag's */
int8_t m_rtxblklog; /* log2 of rextsize, if possible */
+ uint8_t m_rtcsum_shift; /* log2 of RT data csum size */
+ uint32_t m_rtcsum_bsize; /* rtcsum block size in bytes */
uint m_blockmask; /* sb_blocksize-1 */
uint m_blockwsize; /* sb_blocksize in words */
@@ -272,6 +274,11 @@ __XFS_HAS_FEAT(exchange_range, EXCHANGE_RANGE)
__XFS_HAS_FEAT(metadir, METADIR)
__XFS_HAS_FEAT(zoned, ZONED)
+static inline bool xfs_has_rtcsum(const struct xfs_mount *mp)
+{
+ return mp->m_sb.sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM;
+}
+
static inline bool xfs_has_rtgroups(const struct xfs_mount *mp)
{
/* all metadir file systems also allow rtgroups */
--git a/repair/agheader.h b/repair/agheader.h
index c81147b6607d..e11d13f620bf 100644
--- a/repair/agheader.h
+++ b/repair/agheader.h
@@ -91,4 +91,10 @@ static inline bool xfs_sb_version_hasmetadir(const struct xfs_sb *sbp)
(sbp->sb_features_incompat & XFS_SB_FEAT_INCOMPAT_METADIR);
}
+static inline bool xfs_sb_version_rtcsum(const struct xfs_sb *sbp)
+{
+ return xfs_sb_is_v5(sbp) &&
+ (sbp->sb_features_ro_compat & XFS_SB_FEAT_RO_COMPAT_RTCSUM);
+}
+
#endif /* __XFS_REPAIR_AGHEADER_H__ */
diff --git a/repair/sb.c b/repair/sb.c
index ee1cc63fae64..73562c3d061a 100644
--- a/repair/sb.c
+++ b/repair/sb.c
@@ -523,6 +523,19 @@ verify_sb(char *sb_buf, xfs_sb_t *sb, int is_primary_sb)
ret = verify_sb_rtgroups(sb);
if (ret)
return ret;
+
+ if (xfs_sb_version_rtcsum(sb)) {
+ if (sb->sb_rtcsum_type >= XFS_CSUM_TYPE_MAX)
+ return XR_SB_GEO_MISMATCH;
+ if (sb->sb_rtcsum_blklog < XFS_RTCSUM_BSIZE_LOG_MIN ||
+ sb->sb_rtcsum_blklog > XFS_RTCSUM_BSIZE_LOG_MAX)
+ return XR_SB_GEO_MISMATCH;
+ } else {
+ if (sb->sb_rtcsum_type)
+ return XR_SB_GEO_MISMATCH;
+ if (sb->sb_rtcsum_blklog)
+ return XR_SB_GEO_MISMATCH;
+ }
}
return(XR_OK);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 21/32] xfs: add support for per-RTG csum files
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (19 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 20/32] FIXUP Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 22/32] FIXUP Christoph Hellwig
` (10 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: b9f4a597a9dcf00b8f22badb5d81d0f8c5584beb
Add the definitions for another per-RTG file that stores data checksums
for the RTG. The file is fully preallocated at mkfs/growfs time, and
thus bmap lookups for it can be performed without taking locks.
Checksums are organized in fixed size large (initially 32Kib or 64KiB)
blocks to reduce the lookup and read overhead compare to using the
smaller file system block size.
Each block uses the standard RT file header for self-describing metadata
and can thus reuse the buf_ops including the verifier.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_cksum.h | 7 +-
libxfs/xfs_health.h | 4 +-
libxfs/xfs_rtbitmap.c | 10 +++
libxfs/xfs_rtcsumfile.c | 94 +++++++++++++++++++++++++
libxfs/xfs_rtcsumfile.h | 150 ++++++++++++++++++++++++++++++++++++++++
libxfs/xfs_rtgroup.c | 10 +++
libxfs/xfs_rtgroup.h | 6 ++
libxfs/xfs_shared.h | 1 +
8 files changed, 280 insertions(+), 2 deletions(-)
create mode 100644 libxfs/xfs_rtcsumfile.c
create mode 100644 libxfs/xfs_rtcsumfile.h
diff --git a/libxfs/xfs_cksum.h b/libxfs/xfs_cksum.h
index 999a290cfd72..315b85ff78ae 100644
--- a/libxfs/xfs_cksum.h
+++ b/libxfs/xfs_cksum.h
@@ -2,7 +2,12 @@
#ifndef _XFS_CKSUM_H
#define _XFS_CKSUM_H 1
-#define XFS_CRC_SEED (~(uint32_t)0)
+/*
+ * crc32c() does not include the inversion at the beginning and end, while
+ * crc64_nvme() does.
+ */
+#define XFS_CRC_SEED (~(uint32_t)0)
+#define XFS_CRC64_SEED 0
/*
* Calculate the intermediate checksum for a buffer that has the CRC field
diff --git a/libxfs/xfs_health.h b/libxfs/xfs_health.h
index 1d45cf5789e8..349093e75268 100644
--- a/libxfs/xfs_health.h
+++ b/libxfs/xfs_health.h
@@ -72,6 +72,7 @@ struct xfs_rtgroup;
#define XFS_SICK_RG_SUMMARY (1 << 2) /* rt groups summary */
#define XFS_SICK_RG_RMAPBT (1 << 3) /* reverse mappings */
#define XFS_SICK_RG_REFCNTBT (1 << 4) /* reference counts */
+#define XFS_SICK_RG_CSUM (1 << 5) /* data checksums */
/* Observable health issues for AG metadata. */
#define XFS_SICK_AG_SB (1 << 0) /* superblock */
@@ -119,7 +120,8 @@ struct xfs_rtgroup;
XFS_SICK_RG_BITMAP | \
XFS_SICK_RG_SUMMARY | \
XFS_SICK_RG_RMAPBT | \
- XFS_SICK_RG_REFCNTBT)
+ XFS_SICK_RG_REFCNTBT | \
+ XFS_SICK_RG_CSUM)
#define XFS_SICK_AG_PRIMARY (XFS_SICK_AG_SB | \
XFS_SICK_AG_AGF | \
diff --git a/libxfs/xfs_rtbitmap.c b/libxfs/xfs_rtbitmap.c
index 3abc495c2a3c..50f89f668889 100644
--- a/libxfs/xfs_rtbitmap.c
+++ b/libxfs/xfs_rtbitmap.c
@@ -122,9 +122,18 @@ const struct xfs_buf_ops xfs_rtsummary_buf_ops = {
.verify_struct = xfs_rtbuf_verify,
};
+const struct xfs_buf_ops xfs_rtcsum_buf_ops = {
+ .name = "xfs_rtcsum",
+ .magic = { 0, cpu_to_be32(XFS_RTCSUM_MAGIC) },
+ .verify_read = xfs_rtbuf_verify_read,
+ .verify_write = xfs_rtbuf_verify_write,
+ .verify_struct = xfs_rtbuf_verify,
+};
+
static const struct xfs_buf_ops *xfs_rtblock_buf_ops[XFS_RTGI_MAX] = {
[XFS_RTGI_SUMMARY] = &xfs_rtsummary_buf_ops,
[XFS_RTGI_BITMAP] = &xfs_rtbitmap_buf_ops,
+ [XFS_RTGI_CSUM] = &xfs_rtcsum_buf_ops,
};
const struct xfs_buf_ops *
@@ -140,6 +149,7 @@ xfs_rtblock_ops(
static enum xfs_blft xfs_rtblock_buf_types[XFS_RTGI_MAX] = {
[XFS_RTGI_SUMMARY] = XFS_BLFT_RTSUMMARY_BUF,
[XFS_RTGI_BITMAP] = XFS_BLFT_RTBITMAP_BUF,
+ [XFS_RTGI_CSUM] = XFS_BLFT_RTCSUM_BUF,
};
/* Release cached rt bitmap and summary buffers. */
diff --git a/libxfs/xfs_rtcsumfile.c b/libxfs/xfs_rtcsumfile.c
new file mode 100644
index 000000000000..fc71c52beafd
--- /dev/null
+++ b/libxfs/xfs_rtcsumfile.c
@@ -0,0 +1,94 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Christoph Hellwig.
+ */
+#include "xfs_platform.h"
+#include "xfs_fs.h"
+#include "xfs_format.h"
+#include "xfs_log_format.h"
+#include "xfs_shared.h"
+#include "xfs_trans_resv.h"
+#include "xfs_bit.h"
+#include "xfs_mount.h"
+#include "xfs_inode.h"
+#include "xfs_bmap.h"
+#include "xfs_rtbitmap.h"
+#include "xfs_bmap_btree.h"
+#include "xfs_trans.h"
+#include "xfs_error.h"
+#include "xfs_health.h"
+#include "xfs_rtcsumfile.h"
+
+int
+xfs_rtcsum_bmap(
+ struct xfs_rtgroup *rtg,
+ xfs_rgblock_t rgbno,
+ xfs_daddr_t *daddr)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+ struct xfs_inode *csumip = rtg_csum(rtg);
+ struct xfs_ifork *ifp = &csumip->i_df;
+ unsigned int csum_block = xfs_rgb_to_rtcsumblock(mp, rgbno);
+ xfs_fileoff_t start_fsb =
+ XFS_B_TO_FSB(mp, mp->m_rtcsum_bsize) * csum_block;
+ struct xfs_iext_cursor icur;
+ struct xfs_bmbt_irec got;
+
+ ASSERT(!xfs_need_iread_extents(ifp));
+
+ if (XFS_IS_CORRUPT(mp, ifp->if_nextents != 1))
+ goto sick;
+
+ /*
+ * We can do an unlocked lookup here because the bmap btree for the
+ * csum files is immutable once created.
+ */
+ if (XFS_IS_CORRUPT(mp, !xfs_iext_lookup_extent(csumip, ifp, start_fsb,
+ &icur, &got)))
+ goto sick;
+ if (XFS_IS_CORRUPT(mp, got.br_startoff > start_fsb))
+ goto sick;
+
+ start_fsb -= got.br_startoff;
+ *daddr = XFS_FSB_TO_DADDR(mp, got.br_startblock + start_fsb);
+ return 0;
+sick:
+ xfs_rtginode_mark_sick(rtg, XFS_RTGI_CSUM);
+ return -EFSCORRUPTED;
+}
+
+xfs_off_t
+xfs_rtcsum_file_size(
+ struct xfs_rtgroup *rtg)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+ uint64_t raw_size;
+
+ raw_size = (xfs_off_t)rtg_blocks(rtg) << mp->m_rtcsum_shift;
+ return DIV_ROUND_UP_ULL(raw_size, xfs_rtcsum_payload_size(mp)) *
+ mp->m_rtcsum_bsize;
+}
+
+int
+xfs_rtcsum_alloc_blocks(
+ struct xfs_rtgroup *rtg)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+
+ return xfs_rtfile_initialize_blocks(rtg, XFS_RTGI_CSUM, 0,
+ XFS_B_TO_FSB(mp, rtg_csum(rtg)->i_disk_size),
+ XFS_B_TO_FSB(mp, mp->m_rtcsum_bsize),
+ XFS_BMAPI_CONTIG, NULL);
+}
+
+int
+xfs_rtcsum_create(
+ struct xfs_rtgroup *rtg,
+ struct xfs_inode *ip,
+ struct xfs_trans *tp,
+ bool init)
+{
+ ip->i_disk_size = xfs_rtcsum_file_size(rtg);
+ xfs_trans_log_inode(tp, ip, XFS_ILOG_CORE);
+ return 0;
+}
diff --git a/libxfs/xfs_rtcsumfile.h b/libxfs/xfs_rtcsumfile.h
new file mode 100644
index 000000000000..8aacac05adc3
--- /dev/null
+++ b/libxfs/xfs_rtcsumfile.h
@@ -0,0 +1,150 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _XFS_RTCSUMFILE_H
+#define _XFS_RTCSUMFILE_H
+
+#include "xfs_rtgroup.h"
+
+/*
+ * Maximum size of a checksum buffer for writes, used for the log reservation.
+ */
+#define XFS_RTCSUM_MAX_WRITE SZ_64K
+
+/*
+ * Size of the actual payload in the RT data checksum block. This excludes the
+ * self-describing metadata header.
+ */
+static inline unsigned int
+xfs_rtcsum_payload_size(
+ struct xfs_mount *mp)
+{
+ return mp->m_rtcsum_bsize - sizeof(struct xfs_rtbuf_blkinfo);
+}
+
+/* Convert data length in logical blocks to checksum length in bytes. */
+static inline unsigned int
+xfs_extlen_to_rtcsum_len(
+ struct xfs_mount *mp,
+ xfs_extlen_t nb)
+{
+ return nb << mp->m_rtcsum_shift;
+}
+
+/* Convert checksum length in bytes to data length in logical blocks. */
+static inline xfs_extlen_t
+xfs_rtcsum_len_to_extlen(
+ struct xfs_mount *mp,
+ unsigned int csum_len)
+{
+ return csum_len >> mp->m_rtcsum_shift;
+}
+
+/* Convert an rgbno to the csum byte position in the csum file. */
+static inline xfs_off_t
+xfs_rgb_to_rtcsumpos(
+ struct xfs_mount *mp,
+ xfs_rtblock_t rgbno)
+{
+ return (xfs_off_t)rgbno << mp->m_rtcsum_shift;
+}
+
+/* Convert an rgbno to the csum block index in the csum file. */
+static inline unsigned int
+xfs_rgb_to_rtcsumblock(
+ struct xfs_mount *mp,
+ xfs_rtblock_t rgbno)
+{
+ return div_u64(xfs_rgb_to_rtcsumpos(mp, rgbno),
+ xfs_rtcsum_payload_size(mp));
+}
+
+/* Convert an rgbno to a the checksum offset within an rt csum block. */
+static inline unsigned int
+xfs_rgb_to_rtcsumoff(
+ struct xfs_mount *mp,
+ xfs_rgblock_t rgbno)
+{
+ uint32_t off;
+
+ div_u64_rem(xfs_rgb_to_rtcsumpos(mp, rgbno),
+ xfs_rtcsum_payload_size(mp), &off);
+ return sizeof(struct xfs_rtbuf_blkinfo) + off;
+}
+
+/* Convert an rtbno to a the checksum offset within an rt csum block. */
+static inline unsigned int
+xfs_rtb_to_rtcsumoff(
+ struct xfs_mount *mp,
+ xfs_rtblock_t fsbno)
+{
+ return xfs_rgb_to_rtcsumoff(mp, xfs_rtb_to_rgbno(mp, fsbno));
+}
+
+int xfs_rtcsum_bmap(struct xfs_rtgroup *rtg, xfs_rgblock_t rgbno,
+ xfs_daddr_t *daddr);
+xfs_off_t xfs_rtcsum_file_size(struct xfs_rtgroup *rtg);
+int xfs_rtcsum_alloc_blocks(struct xfs_rtgroup *rtg);
+int xfs_rtcsum_create(struct xfs_rtgroup *rtg, struct xfs_inode *ip,
+ struct xfs_trans *tp, bool init);
+
+static inline unsigned int
+xfs_rtcsum_bufs_per_rtg(
+ struct xfs_rtgroup *rtg)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+
+ if (xfs_has_rtcsum(mp))
+ return div_u64(xfs_rtcsum_file_size(rtg), mp->m_rtcsum_bsize);
+ return 0;
+}
+
+static inline xfs_filblks_t
+xfs_rtcsum_max_len(
+ struct xfs_mount *mp,
+ xfs_fsblock_t fsbno)
+{
+ return xfs_rtcsum_len_to_extlen(mp,
+ mp->m_rtcsum_bsize - xfs_rtb_to_rtcsumoff(mp, fsbno));
+}
+
+union xfs_csum {
+ uint32_t crc32c;
+ uint64_t crc64;
+};
+
+static __always_inline void
+xfs_csum_seed(
+ struct xfs_mount *mp,
+ union xfs_csum *csum)
+{
+ if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
+ csum->crc32c = XFS_CRC_SEED;
+ else
+ csum->crc64 = XFS_CRC64_SEED;
+}
+
+static __always_inline void
+xfs_csum_gen(
+ struct xfs_mount *mp,
+ void *data,
+ unsigned int len,
+ union xfs_csum *csum)
+{
+ if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
+ csum->crc32c = crc32c(csum->crc32c, data, len);
+ else
+ csum->crc64 = crc64_nvme(csum->crc64, data, len);
+}
+
+static __always_inline void
+xfs_csum_finalize(
+ struct xfs_mount *mp,
+ union xfs_disk_csum *to,
+ union xfs_csum *csum)
+{
+ if (mp->m_sb.sb_rtcsum_type == XFS_CSUM_TYPE_CRC32C)
+ to->crc32c = cpu_to_le32(~csum->crc32c);
+ else
+ to->crc64 = cpu_to_le64(csum->crc64);
+}
+
+#endif /* _XFS_RTCSUMFILE_H */
diff --git a/libxfs/xfs_rtgroup.c b/libxfs/xfs_rtgroup.c
index c1f8d2cb186a..dcacdf4c9c13 100644
--- a/libxfs/xfs_rtgroup.c
+++ b/libxfs/xfs_rtgroup.c
@@ -33,6 +33,7 @@
#include "xfs_metadir.h"
#include "xfs_rtrmap_btree.h"
#include "xfs_rtrefcount_btree.h"
+#include "xfs_rtcsumfile.h"
/* Find the first usable fsblock in this rtgroup. */
static inline uint32_t
@@ -392,6 +393,15 @@ static const struct xfs_rtginode_ops xfs_rtginode_ops[XFS_RTGI_MAX] = {
.enabled = xfs_has_reflink,
.create = xfs_rtrefcountbt_create,
},
+ [XFS_RTGI_CSUM] = {
+ .name = "csum",
+ .metafile_type = XFS_METAFILE_RTCSUM,
+ .sick = XFS_SICK_RG_CSUM,
+ .fmt_mask = (1U << XFS_DINODE_FMT_EXTENTS) |
+ (1U << XFS_DINODE_FMT_BTREE),
+ .enabled = xfs_has_rtcsum,
+ .create = xfs_rtcsum_create,
+ },
};
/* Return the shortname of this rtgroup inode. */
diff --git a/libxfs/xfs_rtgroup.h b/libxfs/xfs_rtgroup.h
index 2aa6d4e59cc3..4b8e6b526deb 100644
--- a/libxfs/xfs_rtgroup.h
+++ b/libxfs/xfs_rtgroup.h
@@ -16,6 +16,7 @@ enum xfs_rtg_inodes {
XFS_RTGI_SUMMARY, /* allocation summary */
XFS_RTGI_RMAP, /* rmap btree inode */
XFS_RTGI_REFCOUNT, /* refcount btree inode */
+ XFS_RTGI_CSUM, /* data checksum inode */
XFS_RTGI_MAX,
};
@@ -109,6 +110,11 @@ static inline struct xfs_inode *rtg_refcount(const struct xfs_rtgroup *rtg)
return rtg->rtg_inodes[XFS_RTGI_REFCOUNT];
}
+static inline struct xfs_inode *rtg_csum(const struct xfs_rtgroup *rtg)
+{
+ return rtg->rtg_inodes[XFS_RTGI_CSUM];
+}
+
/* Passive rtgroup references */
static inline struct xfs_rtgroup *
xfs_rtgroup_get(
diff --git a/libxfs/xfs_shared.h b/libxfs/xfs_shared.h
index b1e0d9bc1f7d..8a60f2c56fe7 100644
--- a/libxfs/xfs_shared.h
+++ b/libxfs/xfs_shared.h
@@ -40,6 +40,7 @@ extern const struct xfs_buf_ops xfs_refcountbt_buf_ops;
extern const struct xfs_buf_ops xfs_rmapbt_buf_ops;
extern const struct xfs_buf_ops xfs_rtbitmap_buf_ops;
extern const struct xfs_buf_ops xfs_rtsummary_buf_ops;
+extern const struct xfs_buf_ops xfs_rtcsum_buf_ops;
extern const struct xfs_buf_ops xfs_rtbuf_ops;
extern const struct xfs_buf_ops xfs_rtsb_buf_ops;
extern const struct xfs_buf_ops xfs_rtrefcountbt_buf_ops;
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 22/32] FIXUP
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (20 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 21/32] xfs: add support for per-RTG csum files Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 23/32] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
` (9 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
---
libxfs/Makefile | 2 ++
libxfs/xfs_platform.h | 1 +
2 files changed, 3 insertions(+)
diff --git a/libxfs/Makefile b/libxfs/Makefile
index 9795982b0e83..03e50ad42bed 100644
--- a/libxfs/Makefile
+++ b/libxfs/Makefile
@@ -67,6 +67,7 @@ HFILES = \
xfs_rtbitmap.h \
xfs_rtgroup.h \
xfs_rtrmap_btree.h \
+ xfs_rtcsumfile.h \
xfs_sb.h \
xfs_shared.h \
xfs_trans_resv.h \
@@ -139,6 +140,7 @@ CFILES = buf_mem.c \
xfs_rtbitmap.c \
xfs_rtgroup.c \
xfs_rtrmap_btree.c \
+ xfs_rtcsumfile.c \
xfs_sb.c \
xfs_symlink_remote.c \
xfs_trans_inode.c \
diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
index d9cfcec4c1a0..1613dcf07986 100644
--- a/libxfs/xfs_platform.h
+++ b/libxfs/xfs_platform.h
@@ -64,6 +64,7 @@
#include "xfs_fs.h"
#include "libfrog/crc32c.h"
+#include "libfrog/crc64.h"
#include <sys/xattr.h>
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 23/32] xfs: calculate the log reservation for logging data checksum buffers
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (21 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 22/32] FIXUP Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 24/32] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
` (8 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 78980b5ca4f00b0f664396ff5a81f44762f5a07c
The data checksum is logged in its own transaction, and only logs
transactions buffers. While the maximum size of a checksummed data write
is the same as that of a single checksum buffer, the Zone Append based
zoned write path can't guarantee alignment, so it might be spread over
up to three buffers.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_trans_resv.c | 21 +++++++++++++++++++++
libxfs/xfs_trans_resv.h | 4 ++++
2 files changed, 25 insertions(+)
diff --git a/libxfs/xfs_trans_resv.c b/libxfs/xfs_trans_resv.c
index 3a9de562ac2b..cc785b1476cd 100644
--- a/libxfs/xfs_trans_resv.c
+++ b/libxfs/xfs_trans_resv.c
@@ -11,6 +11,7 @@
#include "xfs_log_format.h"
#include "xfs_trans_resv.h"
#include "xfs_mount.h"
+#include "xfs_rtcsumfile.h"
#include "xfs_da_format.h"
#include "xfs_da_btree.h"
#include "xfs_inode.h"
@@ -1226,6 +1227,22 @@ xfs_calc_qm_dqalloc_reservation_minlogsize(
return xfs_calc_qm_dqalloc_reservation(mp, true);
}
+/*
+ * Log data checksums for a write.
+ *
+ * Must cover a checksum for each FSB of data written, and the checksums can
+ * span the FSB-sized checksum buffers at both ends.
+ */
+unsigned int
+xfs_calc_csum_reservation(
+ struct xfs_mount *mp,
+ unsigned int csum_len)
+{
+ return xfs_calc_buf_res(
+ howmany(csum_len, xfs_rtcsum_payload_size(mp)) + 1,
+ mp->m_rtcsum_bsize);
+}
+
/*
* Syncing the incore super block changes to disk.
* the super block to reflect the changes: sector size
@@ -1347,6 +1364,10 @@ xfs_trans_resv_calc(
xfs_calc_namespace_reservations(mp, resp);
+ resp->tr_csum.tr_logres =
+ xfs_calc_csum_reservation(mp, XFS_RTCSUM_MAX_WRITE);
+ resp->tr_csum.tr_logcount = XFS_DEFAULT_LOG_COUNT;
+
/*
* The following transactions are logged in logical format with
* a default log count.
diff --git a/libxfs/xfs_trans_resv.h b/libxfs/xfs_trans_resv.h
index 1804e821f382..127db4da31c1 100644
--- a/libxfs/xfs_trans_resv.h
+++ b/libxfs/xfs_trans_resv.h
@@ -49,6 +49,7 @@ struct xfs_trans_resv {
struct xfs_trans_res tr_sb; /* modify superblock */
struct xfs_trans_res tr_fsyncts; /* update timestamps on fsync */
struct xfs_trans_res tr_atomic_ioend; /* untorn write completion */
+ struct xfs_trans_res tr_csum; /* data checksums in metafile */
};
/* shorthand way of accessing reservation structure */
@@ -122,6 +123,9 @@ unsigned int xfs_calc_itruncate_reservation_minlogsize(struct xfs_mount *mp);
unsigned int xfs_calc_write_reservation_minlogsize(struct xfs_mount *mp);
unsigned int xfs_calc_qm_dqalloc_reservation_minlogsize(struct xfs_mount *mp);
+unsigned int xfs_calc_csum_reservation(struct xfs_mount *mp,
+ unsigned int csum_len);
+
xfs_extlen_t xfs_calc_max_atomic_write_fsblocks(struct xfs_mount *mp);
xfs_extlen_t xfs_calc_atomic_write_log_geometry(struct xfs_mount *mp,
xfs_extlen_t blockcount, unsigned int *new_logres);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 24/32] xfs: report RT data checksum information via XFS_FSOP_GEOM
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (22 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 23/32] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 25/32] xfs: enable RT data checksums Christoph Hellwig
` (7 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: 581da73914b7380f50e435a2260a72ab8601f0e4
Report the RT data checksum flag and the checksum algorithm to userspace.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_fs.h | 6 +++++-
libxfs/xfs_sb.c | 5 +++++
2 files changed, 10 insertions(+), 1 deletion(-)
diff --git a/libxfs/xfs_fs.h b/libxfs/xfs_fs.h
index 185f09f327c0..afe83ff23ef3 100644
--- a/libxfs/xfs_fs.h
+++ b/libxfs/xfs_fs.h
@@ -191,7 +191,10 @@ struct xfs_fsop_geom {
__u32 rgcount; /* number of realtime groups */
__u64 rtstart; /* start of internal rt section */
__u64 rtreserved; /* RT (zoned) reserved blocks */
- __u64 reserved[14]; /* reserved space */
+ __u8 rtcsum_type; /* RT data checksum type */
+ __u8 rtcsum_blklog; /* log2 of rtcsum bsize */
+ __u8 reserved_pad[6];/* reserved space */
+ __u64 reserved[13]; /* reserved space */
};
#define XFS_FSOP_GEOM_SICK_COUNTERS (1 << 0) /* summary counters */
@@ -250,6 +253,7 @@ typedef struct xfs_fsop_resblks {
#define XFS_FSOP_GEOM_FLAGS_PARENT (1 << 25) /* linux parent pointers */
#define XFS_FSOP_GEOM_FLAGS_METADIR (1 << 26) /* metadata directories */
#define XFS_FSOP_GEOM_FLAGS_ZONED (1 << 27) /* zoned rt device */
+#define XFS_FSOP_GEOM_FLAGS_DATA_CSUM (1 << 28) /* data checksums */
/*
* Minimum and maximum sizes need for growth checks.
diff --git a/libxfs/xfs_sb.c b/libxfs/xfs_sb.c
index 2d64412363b4..372009424970 100644
--- a/libxfs/xfs_sb.c
+++ b/libxfs/xfs_sb.c
@@ -1660,6 +1660,11 @@ xfs_fs_geometry(
geo->flags |= XFS_FSOP_GEOM_FLAGS_METADIR;
if (xfs_has_zoned(mp))
geo->flags |= XFS_FSOP_GEOM_FLAGS_ZONED;
+ if (xfs_has_rtcsum(mp)) {
+ geo->flags |= XFS_FSOP_GEOM_FLAGS_DATA_CSUM;
+ geo->rtcsum_type = mp->m_sb.sb_rtcsum_type;
+ geo->rtcsum_blklog = mp->m_sb.sb_rtcsum_blklog;
+ }
geo->rtsectsize = sbp->sb_blocksize;
geo->dirblocksize = xfs_dir2_dirblock_bytes(sbp);
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 25/32] xfs: enable RT data checksums
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (23 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 24/32] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 26/32] man: document the rtcsum geom fields Christoph Hellwig
` (6 subsequent siblings)
31 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Source kernel commit: fedfd9ab32d7f65ef37abb2c06d320b9368d7b6a
All mandatory pieces are in place now, allow mounting.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libxfs/xfs_format.h | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/libxfs/xfs_format.h b/libxfs/xfs_format.h
index 1be3d21910a7..8b269fd3ed9f 100644
--- a/libxfs/xfs_format.h
+++ b/libxfs/xfs_format.h
@@ -384,7 +384,8 @@ xfs_sb_has_compat_feature(
(XFS_SB_FEAT_RO_COMPAT_FINOBT | \
XFS_SB_FEAT_RO_COMPAT_RMAPBT | \
XFS_SB_FEAT_RO_COMPAT_REFLINK| \
- XFS_SB_FEAT_RO_COMPAT_INOBTCNT)
+ XFS_SB_FEAT_RO_COMPAT_INOBTCNT | \
+ XFS_SB_FEAT_RO_COMPAT_RTCSUM)
#define XFS_SB_FEAT_RO_COMPAT_UNKNOWN ~XFS_SB_FEAT_RO_COMPAT_ALL
static inline bool
xfs_sb_has_ro_compat_feature(
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 26/32] man: document the rtcsum geom fields
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (24 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 25/32] xfs: enable RT data checksums Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:31 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 27/32] libfrog: print csum geometry information Christoph Hellwig
` (5 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
man/man2/ioctl_xfs_fsgeometry.2 | 31 +++++++++++++++++++++++++++++--
1 file changed, 29 insertions(+), 2 deletions(-)
diff --git a/man/man2/ioctl_xfs_fsgeometry.2 b/man/man2/ioctl_xfs_fsgeometry.2
index 6d30610b71be..97216df7525b 100644
--- a/man/man2/ioctl_xfs_fsgeometry.2
+++ b/man/man2/ioctl_xfs_fsgeometry.2
@@ -52,7 +52,10 @@ struct xfs_fsop_geom {
__u64 rgextents;
__u64 rtstart;
__u64 rtreserved;
- __u64 reserved[14];
+ __u8 rtcsum_type;
+ __u8 rtcsum_blklog;
+ __u8 reserved_pad[6];
+ __u64 reserved[13];
};
.fi
.in
@@ -159,8 +162,32 @@ This field is meaningful only if the flag
.B XFS_FSOP_GEOM_FLAGS_ZONED
is set.
.PP
+.I rtcsum_type
+Type of RT data checksum. This field is meaningful only if the flag
+.B XFS_FSOP_GEOM_FLAGS_DATA_CSUM
+is set.
+The following values are defined:
+.RS 0.4i
+.TP
+.B XFS_CSUM_TYPE_NONE
+Not data checksum.
+.TP
+.B XFS_CSUM_TYPE_CRC32C
+CRC32c as per NVMe.
+.TP
+.B XFS_CSUM_TYPE_CRC64
+CRC64 as per NVMe.
+.TP
+.RE
+.PP
+.I rtcsum_blklog
+log(2) of the RT data checksum block size.
+This field is meaningful only if the flag
+.PP
+.I reserved_pad
+and
.I reserved
-is set to zero.
+are set to zero.
.SH FILESYSTEM FEATURE FLAGS
Filesystem features are reported to userspace as a combination the following
flags:
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 27/32] libfrog: print csum geometry information
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (25 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 26/32] man: document the rtcsum geom fields Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:32 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs Christoph Hellwig
` (4 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
libfrog/fsgeom.c | 23 +++++++++++++++++++++--
1 file changed, 21 insertions(+), 2 deletions(-)
diff --git a/libfrog/fsgeom.c b/libfrog/fsgeom.c
index ad0476741cf5..324ccdc9b84d 100644
--- a/libfrog/fsgeom.c
+++ b/libfrog/fsgeom.c
@@ -28,6 +28,22 @@ rtdev_name(
return rtname;
}
+static const char *
+csum_name(
+ uint8_t csum)
+{
+ switch (csum) {
+ case XFS_CSUM_TYPE_NONE:
+ return "none";
+ case XFS_CSUM_TYPE_CRC32C:
+ return "crc32";
+ case XFS_CSUM_TYPE_CRC64:
+ return "crc64c";
+ default:
+ return "invalid";
+ }
+};
+
void
xfs_report_geom(
struct xfs_fsop_geom *geo,
@@ -91,7 +107,8 @@ xfs_report_geom(
" =%-22s sectsz=%-5u sunit=%d blks, lazy-count=%d\n"
"realtime =%-22s extsz=%-6d blocks=%lld, rtextents=%lld\n"
" =%-22s rgcount=%-4d rgsize=%u extents\n"
-" =%-22s zoned=%-6d start=%llu reserved=%llu\n"),
+" =%-22s zoned=%-6d start=%llu reserved=%llu\n"
+" =%-22s csum=%-7s csumbsize=%u\n"),
mntpoint, geo->inodesize, geo->agcount, geo->agblocks,
"", geo->sectsize, attrversion, projid32bit,
"", crcs_enabled, finobt_enabled, spinodes, rmapbt_enabled,
@@ -108,7 +125,9 @@ xfs_report_geom(
geo->rtextsize * geo->blocksize, (unsigned long long)geo->rtblocks,
(unsigned long long)geo->rtextents,
"", geo->rgcount, geo->rgextents,
- "", zoned, geo->rtstart, geo->rtreserved);
+ "", zoned, geo->rtstart, geo->rtreserved,
+ "", csum_name(geo->rtcsum_type), geo->rtcsum_blklog ?
+ (1u << geo->rtcsum_blklog) : 0);
}
/* Try to obtain the xfs geometry. On error returns a negative error code. */
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (26 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 27/32] libfrog: print csum geometry information Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:33 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 29/32] mkfs: support RT data checksums Christoph Hellwig
` (3 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
io/stat.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/io/stat.c b/io/stat.c
index e3b3574e0416..c6175298631c 100644
--- a/io/stat.c
+++ b/io/stat.c
@@ -286,8 +286,8 @@ statfs_f(
printf(_("geom.rgcount = %u\n"), fsgeo.rgcount);
printf(_("geom.rtstart = %llu\n"),
(unsigned long long)fsgeo.rtstart);
- printf(_("geom.rtreserved = %llu\n"),
- (unsigned long long)fsgeo.rtreserved);
+ printf(_("geom.rtcsum_type = %u\n"), fsgeo.rtcsum_type);
+ printf(_("geom.rtcsum_blklog = %u\n"), fsgeo.rtcsum_blklog);
}
}
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 29/32] mkfs: support RT data checksums
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (27 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:39 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 30/32] xfs_db: support RT data checksum Christoph Hellwig
` (2 subsequent siblings)
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Parse the -r csum and -r csumbisze options, set the superblock fields
based on that and create the per-RTG csum files when requested.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
man/man8/mkfs.xfs.8.in | 17 ++++++
mkfs/proto.c | 9 +++
mkfs/xfs_mkfs.c | 126 ++++++++++++++++++++++++++++++++++++++++-
3 files changed, 149 insertions(+), 3 deletions(-)
diff --git a/man/man8/mkfs.xfs.8.in b/man/man8/mkfs.xfs.8.in
index 927cd075bf8a..b2da1e76f10f 100644
--- a/man/man8/mkfs.xfs.8.in
+++ b/man/man8/mkfs.xfs.8.in
@@ -1315,6 +1315,23 @@ Controls the amount of space in the realtime section that is reserved for
internal use by garbage collection and reorganization algorithms.
Defaults to 0 if not set.
This option is only valid if the zoned realtime allocator is used.
+.TP
+.BI csum= value
+Controls which data checksum is used.
+Valid options are
+.I none
+,
+.I crc32c
+and
+.I crc64.
+Defaults to
+.I none
+if not set.
+This option is only valid if the zoned realtime allocator is used.
+.TP
+.BI csumbsize= value
+Controls the blocksize of the data checksum files. Valid values are
+32k and 64k. Defaults to 32k or the filesystem block size if larger.
.RE
.PP
.PD 0
diff --git a/mkfs/proto.c b/mkfs/proto.c
index 7ef8cc9d88a1..c232d1171d44 100644
--- a/mkfs/proto.c
+++ b/mkfs/proto.c
@@ -11,6 +11,8 @@
#include <sys/xattr.h>
#include <linux/xattr.h>
#include "libfrog/convert.h"
+#include "libfrog/crc64.h"
+#include "libxfs/xfs_rtcsumfile.h"
#include "proto.h"
/*
@@ -1221,6 +1223,13 @@ rtinit_groups(
fail(_("rtrmap rtsb init failed"), error);
}
+ if (xfs_has_rtcsum(mp)) {
+ error = xfs_rtcsum_alloc_blocks(rtg);
+ if (error)
+ fail(_("Initialization of rtcsum inode failed"),
+ error);
+ }
+
if (!xfs_has_zoned(mp))
rtfreesp_init(rtg);
}
diff --git a/mkfs/xfs_mkfs.c b/mkfs/xfs_mkfs.c
index 334367b81555..9abeca61ee2a 100644
--- a/mkfs/xfs_mkfs.c
+++ b/mkfs/xfs_mkfs.c
@@ -6,15 +6,17 @@
#include "libfrog/util.h"
#include "libxfs.h"
#include <ctype.h>
-#include "libxfs/xfs_zones.h"
#include "xfs_multidisk.h"
#include "libxcmd.h"
#include "libfrog/fsgeom.h"
#include "libfrog/convert.h"
#include "libfrog/crc32cselftest.h"
+#include "libfrog/crc64.h"
#include "libfrog/dahashselftest.h"
#include "libfrog/fsproperties.h"
#include "libfrog/zones.h"
+#include "libxfs/xfs_zones.h"
+#include "libxfs/xfs_rtcsumfile.h"
#include "proto.h"
#include <ini.h>
@@ -145,6 +147,8 @@ enum {
R_ZONED,
R_START,
R_RESERVED,
+ R_CSUM,
+ R_CSUMBSIZE,
R_MAX_OPTS,
};
@@ -777,6 +781,8 @@ static struct opt_params ropts = {
[R_ZONED] = "zoned",
[R_START] = "start",
[R_RESERVED] = "reserved",
+ [R_CSUM] = "csum",
+ [R_CSUMBSIZE] = "csumbsize",
[R_MAX_OPTS] = NULL,
},
.subopt_params = {
@@ -864,6 +870,19 @@ static struct opt_params ropts = {
.maxval = LLONG_MAX,
.defaultval = SUBOPT_NEEDS_VAL,
},
+ { .index = R_CSUM,
+ .conflicts = { { NULL, LAST_CONFLICT } },
+ .minval = 0,
+ .maxval = LLONG_MAX,
+ .defaultval = SUBOPT_NEEDS_VAL,
+ },
+ { .index = R_CSUMBSIZE,
+ .conflicts = { { NULL, LAST_CONFLICT } },
+ .convert = true,
+ .minval = 1u << XFS_RTCSUM_BSIZE_LOG_MIN,
+ .maxval = 1u << XFS_RTCSUM_BSIZE_LOG_MAX,
+ .defaultval = SUBOPT_NEEDS_VAL,
+ },
},
};
@@ -1074,6 +1093,8 @@ struct sb_feat_args {
bool exchrange; /* XFS_SB_FEAT_INCOMPAT_EXCHRANGE */
bool zoned;
bool zone_gaps;
+ uint8_t rtcsum_type;
+ uint8_t rtcsum_blklog;
uint16_t qflags;
};
@@ -1102,6 +1123,7 @@ struct cli_params {
char *rtstart;
uint64_t rtreserved;
char *max_atomic_write;
+ char *rtcsumbsize;
/* parameters where 0 is a valid CLI value */
int dsunit;
@@ -1191,6 +1213,7 @@ struct mkfs_params {
struct sb_feat_args sb_feat;
uint64_t rtstart;
uint64_t rtreserved;
+ uint32_t rtcsumbsize;
uint64_t max_atomic_write;
};
@@ -1229,7 +1252,8 @@ usage( void )
pquota|pqnoenforce]\n\
/* data subvol */ [-d agcount=n,agsize=n,file,name=xxx,size=num,\n\
(sunit=value,swidth=value|su=num,sw=num|noalign),\n\
- sectsize=num,concurrency=num]\n\
+ sectsize=num,concurrency=num,\n\
+ csum=type]\n\
/* force overwrite */ [-f]\n\
/* inode size */ [-i perblock=n|size=num,maxpct=n,attr=0|1|2,\n\
projid32bit=0|1,sparse=0|1,nrext64=0|1,\n\
@@ -1245,7 +1269,8 @@ usage( void )
/* populate from directory */ [-p dirname,atime=0|1]\n\
/* quiet */ [-q]\n\
/* realtime subvol */ [-r extsize=num,size=num,rtdev=xxx,rgcount=n,rgsize=n,\n\
- concurrency=num,zoned=0|1,start=n,reserved=n]\n\
+ concurrency=num,zoned=0|1,start=n,reserved=n,\n\
+ csum=type],csumbsize=n\n\
/* sectorsize */ [-s size=num]\n\
/* version */ [-V]\n\
devicename\n\
@@ -1858,6 +1883,24 @@ set_data_concurrency(
cli->data_concurrency = optnum;
}
+static bool
+set_data_csum(
+ const char *value,
+ uint8_t *csum_val)
+{
+ if (!value)
+ return false;
+ if (!strcmp(value, "none"))
+ *csum_val = XFS_CSUM_TYPE_NONE;
+ else if (!strcmp(value, "crc32c"))
+ *csum_val = XFS_CSUM_TYPE_CRC32C;
+ else if (!strcmp(value, "crc64"))
+ *csum_val = XFS_CSUM_TYPE_CRC64;
+ else
+ return false;
+ return true;
+}
+
static int
data_opts_parser(
struct opt_params *opts,
@@ -2269,6 +2312,13 @@ rtdev_opts_parser(
case R_RESERVED:
cli->rtreserved = getnum(value, opts, subopt);
break;
+ case R_CSUM:
+ if (!set_data_csum(value, &cli->sb_feat.rtcsum_type))
+ return -EINVAL;
+ break;
+ case R_CSUMBSIZE:
+ cli->rtcsumbsize = getstr(value, opts, subopt);
+ break;
default:
return -EINVAL;
}
@@ -3048,6 +3098,11 @@ _("internal RT section only supported in zoned mode\n"));
_("reserved RT blocks only supported in zoned mode\n"));
usage();
}
+ if (cli->sb_feat.rtcsum_type || cli->sb_feat.rtcsum_blklog) {
+ fprintf(stderr,
+_("data checksums not supported without zoned mode\n"));
+ usage();
+ }
}
if (cli->xi->rt.name || cfg->rtstart) {
@@ -4921,6 +4976,11 @@ sb_set_features(
}
if (fp->zone_gaps)
sbp->sb_features_incompat |= XFS_SB_FEAT_INCOMPAT_ZONE_GAPS;
+ if (fp->rtcsum_type) {
+ sbp->sb_features_ro_compat |= XFS_SB_FEAT_RO_COMPAT_RTCSUM;
+ sbp->sb_rtcsum_type = fp->rtcsum_type;
+ sbp->sb_rtcsum_blklog = fp->rtcsum_blklog;
+ }
}
/*
@@ -5243,6 +5303,63 @@ validate_max_atomic_write_ags(
}
}
+static void
+validate_rtcsumbsize(
+ struct mkfs_params *cfg,
+ struct cli_params *cli,
+ struct xfs_mount *mp)
+{
+ unsigned int l;
+
+ if (!cli->sb_feat.rtcsum_type) {
+ if (cli->rtcsumbsize) {
+ fprintf(stderr,
+ _("RT data csum block size options requires RT data csum type.\n"));
+ exit(1);
+ }
+ return;
+ }
+
+ if (!cfg->sb_feat.zoned) {
+ fprintf(stderr,
+ _("RT data checksums require a zoned RT subvolume\n"));
+ exit(1);
+ }
+
+ if (cli->rtcsumbsize) {
+ cfg->rtcsumbsize =
+ getnum(cli->rtcsumbsize, &ropts, R_CSUMBSIZE);
+ } else {
+ cfg->rtcsumbsize =
+ max(1u << XFS_RTCSUM_BSIZE_LOG_MIN, cfg->blocksize);
+ }
+
+ if (!is_power_of_2(cfg->rtcsumbsize)) {
+ fprintf(stderr,
+ _("RT data csum block size of %u bytes is not a power of 2\n"),
+ cfg->rtcsumbsize);
+ exit(1);
+ }
+
+ if (cfg->rtcsumbsize < cfg->blocksize) {
+ fprintf(stderr,
+ _("RT data csum block size of %u bytes is smaller than fsblock size %u.\n"),
+ cfg->rtcsumbsize, cfg->blocksize);
+ exit(1);
+ }
+
+ if (cfg->rtcsumbsize % cfg->blocksize) {
+ fprintf(stderr,
+ _("RT data csum block size of %u bytes not aligned with fsblock size %u.\n"),
+ cfg->rtcsumbsize, cfg->blocksize);
+ exit(1);
+ }
+
+ cfg->sb_feat.rtcsum_blklog = 0;
+ for (l = cfg->rtcsumbsize; l > 1; l >>= 1)
+ cfg->sb_feat.rtcsum_blklog++;
+}
+
static void
calculate_log_size(
struct mkfs_params *cfg,
@@ -5523,6 +5640,8 @@ finish_superblock_setup(
sbp->sb_unit = cfg->dsunit;
sbp->sb_width = cfg->dswidth;
mp->m_features |= libxfs_sb_version_to_features(sbp);
+ mp->m_rtcsum_shift = xfs_data_csum_shift(mp->m_sb.sb_rtcsum_type);
+ mp->m_rtcsum_bsize = 1u << mp->m_sb.sb_rtcsum_blklog;
libxfs_sb_mount_rextsize(mp, sbp);
}
@@ -6290,6 +6409,7 @@ main(
validate_datadev(&cfg, &cli);
validate_logdev(&cfg, &cli);
validate_rtdev(&cfg, &cli, &zt);
+ validate_rtcsumbsize(&cfg, &cli, mp);
calc_stripe_factors(&cfg, &cli, &ft);
/*
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 30/32] xfs_db: support RT data checksum
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (28 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 29/32] mkfs: support RT data checksums Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:43 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 31/32] repair: support RT data checksums Christoph Hellwig
2026-09-24 10:04 ` [PATCH 32/32] xfs_scrub: don't merge over unused space when data checksums are enabled Christoph Hellwig
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Support printing the new super block fields, and content of the per-RTG
rtcsum files.
Note that the superblock fields are printed for all metadir file systems
because they reuse space that previously was padding.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
db/block.c | 10 +++++++++-
db/field.c | 2 ++
db/field.h | 1 +
db/inode.c | 2 ++
db/rtgroup.c | 42 ++++++++++++++++++++++++++++++++++++++++++
db/rtgroup.h | 5 +++++
db/sb.c | 10 ++++++++--
db/type.c | 5 +++++
db/type.h | 1 +
man/man8/xfs_db.8 | 13 ++++++++++++-
10 files changed, 87 insertions(+), 4 deletions(-)
diff --git a/db/block.c b/db/block.c
index 2f1978c41f30..c6dd50dfa764 100644
--- a/db/block.c
+++ b/db/block.c
@@ -263,7 +263,15 @@ dblock_f(
dbprintf(_("no type for file data\n"));
return 0;
}
- nex = nb = type == TYP_DIR2 ? mp->m_dir_geo->fsbcount : 1;
+
+ if (type == TYP_DIR2)
+ nb = mp->m_dir_geo->fsbcount;
+ else if (type == TYP_RTCSUM)
+ nb = XFS_B_TO_FSB(mp, 1u << mp->m_sb.sb_rtcsum_blklog);
+ else
+ nb = 1;
+
+ nex = nb;
bmp = malloc(nb * sizeof(*bmp));
bmap(bno, nb, XFS_DATA_FORK, &nex, bmp);
if (nex == 0) {
diff --git a/db/field.c b/db/field.c
index 62c983fc3d49..275f175493df 100644
--- a/db/field.c
+++ b/db/field.c
@@ -441,6 +441,8 @@ const ftattr_t ftattrtab[] = {
0, NULL, NULL },
{ FLDT_RGSUMMARY, "rgsummary", NULL, (char *)rgsummary_flds,
btblock_size, FTARG_SIZE, NULL, rgsummary_flds },
+ { FLDT_RTCSUM, "rtcsum", NULL, (char *)rtcsum_flds,
+ rtcsumblock_size, FTARG_SIZE, NULL, rtcsum_flds },
{ FLDT_ZZZ, NULL }
};
diff --git a/db/field.h b/db/field.h
index 27df5a0fcd1b..77f84d55cb62 100644
--- a/db/field.h
+++ b/db/field.h
@@ -211,6 +211,7 @@ typedef enum fldt {
FLDT_RGBITMAP,
FLDT_SUMINFO,
FLDT_RGSUMMARY,
+ FLDT_RTCSUM,
FLDT_ZZZ /* mark last entry */
} fldt_t;
diff --git a/db/inode.c b/db/inode.c
index 46f83f18f083..a77595427c56 100644
--- a/db/inode.c
+++ b/db/inode.c
@@ -727,6 +727,8 @@ inode_next_type(void)
return TYP_RTRMAPBT;
case XFS_METAFILE_RTREFCOUNT:
return TYP_RTREFCBT;
+ case XFS_METAFILE_RTCSUM:
+ return TYP_RTCSUM;
default:
return TYP_DATA;
}
diff --git a/db/rtgroup.c b/db/rtgroup.c
index c6b96c9dc79d..666adb2feb33 100644
--- a/db/rtgroup.c
+++ b/db/rtgroup.c
@@ -16,6 +16,8 @@
#include "output.h"
#include "init.h"
#include "rtgroup.h"
+#include "libfrog/crc64.h"
+#include "xfs_rtcsumfile.h"
#define uuid_equal(s,d) (platform_uuid_compare((s),(d)) == 0)
@@ -152,3 +154,43 @@ const field_t rgsummary_hfld[] = {
{ "", FLDT_RGSUMMARY, OI(0), C1, 0, TYP_NONE },
{ NULL }
};
+
+/*
+ * Get the size of a rtcsum block.
+ */
+int
+rtcsumblock_size(
+ void *obj,
+ int startoff,
+ int idx)
+{
+ return bitize(1u << mp->m_sb.sb_rtcsum_blklog);
+}
+
+static int
+rtcsum_count(
+ void *obj,
+ int startoff)
+{
+ return xfs_rtcsum_payload_size(mp) >> mp->m_rtcsum_shift;
+}
+
+#define OFF(f) bitize(offsetof(struct xfs_rtbuf_blkinfo, rt_ ## f))
+const field_t rtcsum_flds[] = {
+ { "magicnum", FLDT_UINT32X, OI(OFF(magic)), C1, 0, TYP_NONE },
+ { "crc", FLDT_CRC, OI(OFF(crc)), C1, 0, TYP_NONE },
+ { "owner", FLDT_INO, OI(OFF(owner)), C1, 0, TYP_NONE },
+ { "bno", FLDT_DFSBNO, OI(OFF(blkno)), C1, 0, TYP_BMAPBTD },
+ { "lsn", FLDT_UINT64X, OI(OFF(lsn)), C1, 0, TYP_NONE },
+ { "uuid", FLDT_UUID, OI(OFF(uuid)), C1, 0, TYP_NONE },
+ /* the checksums are after the blkinfo structure */
+ { "csums", FLDT_SUMINFO, OI(bitize(sizeof(struct xfs_rtbuf_blkinfo))),
+ rtcsum_count, FLD_ARRAY | FLD_COUNT, TYP_DATA },
+ { NULL }
+};
+#undef OFF
+
+const field_t rtcsum_hfld[] = {
+ { "", FLDT_RTCSUM, OI(0), C1, 0, TYP_NONE },
+ { NULL }
+};
diff --git a/db/rtgroup.h b/db/rtgroup.h
index 5b120f2c9a29..b674154fcbad 100644
--- a/db/rtgroup.h
+++ b/db/rtgroup.h
@@ -15,6 +15,11 @@ extern const struct field rgbitmap_hfld[];
extern const struct field rgsummary_flds[];
extern const struct field rgsummary_hfld[];
+extern const struct field rtcsum_flds[];
+extern const struct field rtcsum_hfld[];
+
+int rtcsumblock_size(void *obj, int startoff, int idx);
+
extern void rtsb_init(void);
extern int rtsb_size(void *obj, int startoff, int idx);
diff --git a/db/sb.c b/db/sb.c
index 666207145a86..c565e6535f63 100644
--- a/db/sb.c
+++ b/db/sb.c
@@ -144,8 +144,14 @@ const field_t sb_flds[] = {
FLD_COUNT, TYP_NONE },
{ "pad", FLDT_UINT8X, OI(OFF(pad)), metadirfld_count,
FLD_COUNT, TYP_NONE },
- { "rtstart", FLDT_DRFSBNO, OI(OFF(rtstart)), zonedfld_count, FLD_COUNT, TYP_NONE },
- { "rtreserved", FLDT_UINT64D, OI(OFF(rtreserved)), zonedfld_count, FLD_COUNT, TYP_NONE },
+ { "rtcsum_type", FLDT_UINT8D, OI(OFF(rtcsum_type)), metadirfld_count,
+ FLD_COUNT, TYP_NONE },
+ { "rtcsum_blklog", FLDT_UINT8D, OI(OFF(rtcsum_blklog)), metadirfld_count,
+ FLD_COUNT, TYP_NONE },
+ { "rtstart", FLDT_DRFSBNO, OI(OFF(rtstart)), zonedfld_count,
+ FLD_COUNT, TYP_NONE },
+ { "rtreserved", FLDT_UINT64D, OI(OFF(rtreserved)), zonedfld_count,
+ FLD_COUNT, TYP_NONE },
{ NULL }
};
diff --git a/db/type.c b/db/type.c
index 324f416a49cc..2db5b5f6a54d 100644
--- a/db/type.c
+++ b/db/type.c
@@ -71,6 +71,7 @@ static const typ_t __typtab[] = {
TYP_F_NO_CRC_OFF },
{ TYP_RGBITMAP, NULL },
{ TYP_RGSUMMARY, NULL },
+ { TYP_RTCSUM, NULL },
{ TYP_NONE, NULL }
};
@@ -125,6 +126,8 @@ static const typ_t __typtab_crc[] = {
&xfs_rtbitmap_buf_ops, XFS_RTBUF_CRC_OFF },
{ TYP_RGSUMMARY, "rgsummary", handle_struct, rgsummary_hfld,
&xfs_rtsummary_buf_ops, XFS_RTBUF_CRC_OFF },
+ { TYP_RTCSUM, "rtcsum", handle_struct, rtcsum_hfld,
+ &xfs_rtcsum_buf_ops, XFS_RTBUF_CRC_OFF },
{ TYP_NONE, NULL }
};
@@ -179,6 +182,8 @@ static const typ_t __typtab_spcrc[] = {
&xfs_rtbitmap_buf_ops, XFS_RTBUF_CRC_OFF },
{ TYP_RGSUMMARY, "rgsummary", handle_struct, rgsummary_hfld,
&xfs_rtsummary_buf_ops, XFS_RTBUF_CRC_OFF },
+ { TYP_RTCSUM, "rtcsum", handle_struct, rtcsum_hfld,
+ &xfs_rtcsum_buf_ops, XFS_RTBUF_CRC_OFF },
{ TYP_NONE, NULL }
};
diff --git a/db/type.h b/db/type.h
index a2488a663dbd..d95c83c78d17 100644
--- a/db/type.h
+++ b/db/type.h
@@ -39,6 +39,7 @@ typedef enum typnm
TYP_FINOBT,
TYP_RGBITMAP,
TYP_RGSUMMARY,
+ TYP_RTCSUM,
TYP_NONE
} typnm_t;
diff --git a/man/man8/xfs_db.8 b/man/man8/xfs_db.8
index c2d4065944be..ab89d709b89e 100644
--- a/man/man8/xfs_db.8
+++ b/man/man8/xfs_db.8
@@ -1356,7 +1356,8 @@ The possible data types are:
.BR agf ", " agfl ", " agi ", " attr ", " bmapbta ", " bmapbtd ,
.BR bnobt ", " cntbt ", " data ", " dir ", " dir2 ", " dqblk ,
.BR inobt ", " inode ", " log ", " refcntbt ", " rmapbt ", " rtbitmap ,
-.BR rtsummary ", " sb ", " symlink ", " rtrmapbt ", " rtrefcbt ", and " text .
+.BR rtsummary ", " sb ", " symlink ", " rtrmapbt ", " rtrefcbt ,
+.BR rtcsum ", and " text .
See the TYPES section below for more information on these data types.
.TP
.BI "timelimit [" OPTIONS ]
@@ -2644,6 +2645,16 @@ The first dimension is the size range,
the second dimension is the starting bitmap block number
(adjacent entries are for the same size, adjacent bitmap blocks).
.TP
+.B rtcsum
+For RT data-checksum enabled file systems, there is one data checksum file
+for each realtime group. The data checksum file contains the standard
+RT block self-describing metadata header, and an array of the checksums
+for the realtime group. The size of each entry depends on the checksum
+used: 4 bytes for
+.BR crc32c
+and 8 bytes for
+.BR crc64 .
+.TP
.B sb
There is one sb (superblock) structure per allocation group.
It is the first disk block in the allocation group.
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 31/32] repair: support RT data checksums
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (29 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 30/32] xfs_db: support RT data checksum Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:54 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 32/32] xfs_scrub: don't merge over unused space when data checksums are enabled Christoph Hellwig
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Recognize the rtcsum per-RTG metadir files, and sanity check a few
well known attributes for them.
If a csum file is missing, or had to be nuke, regenerate it by
re-calculating the checksums. While this does lose the protection
of the checksums, this is probably still better than an unmountable
file system.
Note that due to the lack of the regular buffer readahead path,
reading the data and csum buffers for regenerating the csum file
is fully synchronous and thus slow. Hopefully we can come up with
a better version for online repair and never have to use this for
real.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
repair/Makefile | 1 +
repair/dinode.c | 46 ++++++++++++++
repair/phase6.c | 122 +++++++++++++++++++++++++++++++++++-
repair/rt.c | 36 +++++++++++
repair/rt.h | 10 +++
repair/rtcsum.c | 161 ++++++++++++++++++++++++++++++++++++++++++++++++
6 files changed, 375 insertions(+), 1 deletion(-)
create mode 100644 repair/rtcsum.c
diff --git a/repair/Makefile b/repair/Makefile
index fb0b2f96cc91..27eb68e2b5d5 100644
--- a/repair/Makefile
+++ b/repair/Makefile
@@ -73,6 +73,7 @@ CFILES = \
rcbag.c \
rmap.c \
rt.c \
+ rtcsum.c \
rtrefcount_repair.c \
rtrmap_repair.c \
sb.c \
diff --git a/repair/dinode.c b/repair/dinode.c
index 48939f8bd159..5c03c5f50689 100644
--- a/repair/dinode.c
+++ b/repair/dinode.c
@@ -22,6 +22,8 @@
#include "rmap.h"
#include "bmap_repair.h"
#include "rt.h"
+#include "libfrog/crc64.h"
+#include "xfs_rtcsumfile.h"
/* inode types */
enum xr_ino_type {
@@ -41,6 +43,7 @@ enum xr_ino_type {
XR_INO_PQUOTA, /* project quota inode */
XR_INO_RTRMAP, /* realtime rmap */
XR_INO_RTREFC, /* realtime refcount */
+ XR_INO_RTCSUM, /* realtime data checksum */
XR_INO_MAX
};
@@ -61,6 +64,7 @@ static const char *xr_ino_type_name[] = {
[XR_INO_PQUOTA] = N_("project quota"),
[XR_INO_RTRMAP] = N_("realtime rmap"),
[XR_INO_RTREFC] = N_("realtime refcount"),
+ [XR_INO_RTCSUM] = N_("realtime data checksum"),
};
static_assert(ARRAY_SIZE(xr_ino_type_name) == XR_INO_MAX);
@@ -2064,6 +2068,43 @@ _("bad # of extents (%" PRIu64 ") for %s inode %" PRIu64 "\n"),
return 0;
}
+static int
+process_check_rtcsum_inode(
+ struct xfs_mount *mp,
+ struct xfs_dinode *dinoc,
+ xfs_ino_t lino,
+ enum xr_ino_type *type,
+ int *dirty)
+{
+ xfs_fsize_t size = be64_to_cpu(dinoc->di_size);
+ xfs_fsize_t expected_size;
+ int error;
+
+ error = process_check_rt_inode(mp, dinoc, lino, type, dirty,
+ XR_INO_RTCSUM, _("realtime data checksums"));
+ if (error)
+ return error;
+
+ /* rtcsum inodes must be contiguous */
+ if (xfs_dfork_data_extents(dinoc) != 1) {
+ do_warn(
+_("non-contiguous rtcsum inode %" PRIu64 "\n"), lino);
+ return 1;
+ }
+
+ expected_size = (xfs_off_t)mp->m_sb.sb_rgextents << mp->m_rtcsum_shift;
+ expected_size = (expected_size + xfs_rtcsum_payload_size(mp) - 1) /
+ xfs_rtcsum_payload_size(mp) * mp->m_rtcsum_bsize;
+ if (size != expected_size) {
+ do_warn(
+_("unexpected rtcsum file size (%" PRId64 ") for ino %" PRIu64 "\n"),
+ size, lino);
+ return 1;
+ }
+
+ return 0;
+}
+
/*
* If inode is a superblock inode, does type check to make sure is it valid.
* Returns 0 if it's valid, non-zero if it needs to be cleared.
@@ -2130,6 +2171,8 @@ process_check_metadata_inodes(
if (is_rtrefcount_inode(lino))
return process_check_rt_inode(mp, dinoc, lino, type, dirty,
XR_INO_RTREFC, _("realtime refcount btree"));
+ if (is_rtcsum_inode(lino))
+ return process_check_rtcsum_inode(mp, dinoc, lino, type, dirty);
return 0;
}
@@ -2977,6 +3020,7 @@ process_dinode_metafile(
case XR_INO_UQUOTA:
case XR_INO_GQUOTA:
case XR_INO_PQUOTA:
+ case XR_INO_RTCSUM:
/*
* Quota checking and repair doesn't happen until phase7, so
* preserve quota inodes and their contents for later.
@@ -3555,6 +3599,8 @@ _("bad (negative) size %" PRId64 " on inode %" PRIu64 "\n"),
type = XR_INO_RTRMAP;
else if (is_rtrefcount_inode(lino))
type = XR_INO_RTREFC;
+ else if (is_rtcsum_inode(lino))
+ type = XR_INO_RTCSUM;
else
type = XR_INO_DATA;
break;
diff --git a/repair/phase6.c b/repair/phase6.c
index f3951a3d0709..c7bc1cabd0d3 100644
--- a/repair/phase6.c
+++ b/repair/phase6.c
@@ -6,7 +6,6 @@
#include "libxfs.h"
#include "threads.h"
-#include "threads.h"
#include "prefetch.h"
#include "avl.h"
#include "globals.h"
@@ -23,6 +22,8 @@
#include "repair/quotacheck.h"
#include "repair/slab.h"
#include "repair/rmap.h"
+#include "libfrog/crc64.h"
+#include "xfs_rtcsumfile.h"
static xfs_ino_t orphanage_ino;
@@ -703,6 +704,123 @@ ensure_rtgroup_refcountbt(
populate_rtgroup_refcountbt(rtg, est_fdblocks);
}
+/*
+ * Link a metadata directory inode.
+ */
+static int
+metadir_link(
+ struct xfs_inode *dp,
+ struct xfs_inode *ip,
+ const char *path,
+ enum xfs_metafile_type type)
+{
+ struct xfs_metadir_update upd = {
+ .dp = dp,
+ .metafile_type = type,
+ .ip = ip,
+ .path = path,
+ };
+ int error;
+
+ error = xfs_metadir_start_link(&upd);
+ if (error)
+ return error;
+
+ error = xfs_metadir_link(&upd);
+ if (error)
+ return error;
+
+ xfs_trans_log_inode(upd.tp, upd.ip, XFS_ILOG_CORE);
+ return xfs_metadir_commit(&upd);
+}
+
+static void
+ensure_rtgroup_csum(
+ struct xfs_rtgroup *rtg)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+ xfs_rgnumber_t rgno = rtg_rgno(rtg);
+ struct xfs_inode *dp = mp->m_rtdirip;
+ struct xfs_inode *ip;
+ int error;
+
+ if (no_modify) {
+ if (rtcsum_ino_is_bad(rgno))
+ do_warn(_("would reset RG %u csum inode\n"), rgno);;
+ return;
+ }
+
+ if (!rtcsum_ino_is_bad(rgno)) {
+ /*
+ * The /realtime directory has been discarded, but we should be
+ * able to iget the inodes directly.
+ */
+ error = -libxfs_metafile_iget(mp, rtcsum_ino(rgno),
+ XFS_METAFILE_RTCSUM, &ip);
+ if (error) {
+ do_warn(
+_("Could not open RG %u csum inode, error %d\n"), rgno, error);
+ rtcsum_ino_mark_bad(rgno);
+ }
+ }
+
+ if (rtcsum_ino_is_bad(rgno)) {
+ do_warn(_(
+"resetting RG %u csum inode, regenerating data checksums\n"),
+ rgno);
+ error = -libxfs_rtginode_create(rtg, XFS_RTGI_CSUM, false);
+ if (error) {
+ do_warn(
+_("Couldn't create RG %u csum inode, error %d\n"), rgno, error);
+ return;
+ }
+ ip = rtg->rtg_inodes[XFS_RTGI_CSUM];
+ error = -xfs_rtcsum_alloc_blocks(rtg);
+ if (error) {
+ do_warn(
+_("Initialization of RG %u csum inode failed, error %d"), rgno, error);
+ return;
+ }
+
+ calculate_rtgroup_csums(rtg);
+ } else {
+ struct xfs_trans *tp;
+ const char *name;
+
+ /* Erase parent pointers before we create the new link */
+ try_erase_parent_ptrs(ip);
+
+ name = xfs_rtginode_path(rtg_rgno(rtg), XFS_RTGI_CSUM);
+ error = -metadir_link(dp, ip, name, XFS_METAFILE_RTCSUM);
+ kfree(name);
+ if (error) {
+ do_warn(
+_("Couldn't link RG %u csum inode, error %d\n"), rgno, error);
+ return;
+ }
+
+ /*
+ * Reset the link count to 1 because the link above bumped it.
+ */
+ error = -libxfs_trans_alloc_inode(ip, &M_RES(mp)->tr_ichange,
+ 0, 0, false, &tp);
+ if (!error) {
+ set_nlink(VFS_I(ip), 1);
+ libxfs_trans_log_inode(tp, ip, XFS_ILOG_CORE);
+ error = -libxfs_trans_commit(tp);
+ }
+ if (error)
+ do_error(
+_("Couldn't reset link count on RG %i quota inode, error %d\n"),
+ rgno, error);
+ }
+
+ /* Mark the inode in use. */
+ mark_ino_inuse(mp, I_INO(ip), S_IFREG, I_INO(dp));
+ mark_ino_metadata(mp, I_INO(ip));
+ libxfs_irele(ip);
+}
+
/* Initialize a root directory. */
static int
init_fs_root_dir(
@@ -3473,6 +3591,8 @@ _(" - resetting contents of realtime bitmap and summary inodes\n"));
}
ensure_rtgroup_rmapbt(rtg, est_fdblocks);
ensure_rtgroup_refcountbt(rtg, est_fdblocks);
+ if (xfs_has_rtcsum(mp))
+ ensure_rtgroup_csum(rtg);
}
}
diff --git a/repair/rt.c b/repair/rt.c
index b0ff775bd339..4b64a91fbe29 100644
--- a/repair/rt.c
+++ b/repair/rt.c
@@ -26,6 +26,8 @@ struct rtg_computed {
};
struct rtg_computed *rt_computed;
+static xfs_ino_t *rtcsum_inos;
+
static inline void
set_rtword(
struct xfs_mount *mp,
@@ -410,6 +412,27 @@ fill_rtsummary(
_("couldn't re-initialize realtime summary inode, error %d\n"), error);
}
+bool
+rtcsum_ino_is_bad(
+ xfs_rgnumber_t rgno)
+{
+ return rtcsum_inos[rgno] == NULLFSINO;
+}
+
+void
+rtcsum_ino_mark_bad(
+ xfs_rgnumber_t rgno)
+{
+ rtcsum_inos[rgno] = NULLFSINO;
+}
+
+xfs_ino_t
+rtcsum_ino(
+ xfs_rgnumber_t rgno)
+{
+ return rtcsum_inos[rgno];
+}
+
bool
is_rtgroup_inode(
xfs_ino_t ino,
@@ -471,6 +494,13 @@ mark_rtginode(
goto out_corrupt;
}
+ /*
+ * Record the inode numbers of the data checksum inodes, as we don't
+ * just blow those away like other per-RTG metadata.
+ */
+ if (type == XFS_RTGI_CSUM)
+ rtcsum_inos[rtg_rgno(rtg)] = I_INO(ip);
+
/*
* Phase 3 will clear the ondisk inodes of all rt metadata files, but
* it doesn't reset any blocks. Keep the incore inodes loaded so that
@@ -494,6 +524,12 @@ discover_rtgroup_inodes(
int error, err2;
int i;
+ rtcsum_inos = calloc(mp->m_sb.sb_rgcount, sizeof(xfs_ino_t));
+ if (!rtcsum_inos)
+ do_error(_("could not allocate csum ino array\n"));
+ for (i = 0; i < mp->m_sb.sb_rgcount; i++)
+ rtcsum_inos[i] = NULLFSINO;
+
tp = libxfs_trans_alloc_empty(mp);
if (xfs_has_rtgroups(mp) && mp->m_sb.sb_rgcount > 0) {
error = -libxfs_rtginode_load_parent(tp);
diff --git a/repair/rt.h b/repair/rt.h
index e4f3d5d9af31..9400b230de59 100644
--- a/repair/rt.h
+++ b/repair/rt.h
@@ -13,6 +13,10 @@ void check_rtsummary(struct xfs_mount *mp);
void fill_rtbitmap(struct xfs_rtgroup *rtg);
void fill_rtsummary(struct xfs_rtgroup *rtg);
+bool rtcsum_ino_is_bad(xfs_rgnumber_t rgno);
+void rtcsum_ino_mark_bad(xfs_rgnumber_t rgno);
+xfs_ino_t rtcsum_ino(xfs_rgnumber_t rgno);
+
void discover_rtgroup_inodes(struct xfs_mount *mp);
void unload_rtgroup_inodes(struct xfs_mount *mp);
@@ -37,6 +41,10 @@ static inline bool is_rtrefcount_inode(xfs_ino_t ino)
{
return is_rtgroup_inode(ino, XFS_RTGI_REFCOUNT);
}
+static inline bool is_rtcsum_inode(xfs_ino_t ino)
+{
+ return is_rtgroup_inode(ino, XFS_RTGI_CSUM);
+}
void mark_rtgroup_inodes_bad(struct xfs_mount *mp, enum xfs_rtg_inodes type);
bool rtgroup_inodes_were_bad(enum xfs_rtg_inodes type);
@@ -44,4 +52,6 @@ bool rtgroup_inodes_were_bad(enum xfs_rtg_inodes type);
void check_rtsb(struct xfs_mount *mp);
void rewrite_rtsb(struct xfs_mount *mp);
+void calculate_rtgroup_csums(struct xfs_rtgroup *rtg);
+
#endif /* _XFS_REPAIR_RT_H_ */
diff --git a/repair/rtcsum.c b/repair/rtcsum.c
new file mode 100644
index 000000000000..c4d629c18e8d
--- /dev/null
+++ b/repair/rtcsum.c
@@ -0,0 +1,161 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Christoph Hellwig.
+ *
+ * Rebuild data checksums when the csum metafile was lost. This doens't perform
+ * great and is inteded as a last resort.
+ */
+#include "libxfs.h"
+#include "globals.h"
+#include "protos.h"
+#include "rt.h"
+#include "err_protos.h"
+#include "libfrog/crc64.h"
+#include "xfs_platform.h"
+#include "xfs_rtcsumfile.h"
+
+/* default to 1MiB data reads to make the performance only somewhat horrible. */
+#define DATA_BLOCKS_PER_BUF 256
+
+static int
+read_data(
+ struct xfs_buftarg *btp,
+ xfs_daddr_t blkno,
+ unsigned int nblks,
+ void *buf)
+{
+ int fd = btp->bt_bdev_fd;
+ off_t pos = LIBXFS_BBTOOFF64(blkno);
+ size_t len = BBTOB(nblks);
+ ssize_t ret;
+
+ ret = pread(fd, buf, len, pos);
+ if (ret < 0) {
+ ret = errno;
+
+ fprintf(stderr, _("%s: read failed: %s\n"),
+ progname, strerror(ret));
+ return -ret;
+ }
+ if (ret != len) {
+ fprintf(stderr, _("%s: error - read only %zd of %zd bytes\n"),
+ progname, ret, len);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+static int
+calculate_buf_csums(
+ struct xfs_rtgroup *rtg,
+ void *data_buf,
+ xfs_rgblock_t rgbno,
+ unsigned int *nr)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+ unsigned int boff = xfs_rgb_to_rtcsumoff(mp, rgbno);
+ struct xfs_trans_res tres = M_RES(mp)->tr_csum;
+ xfs_daddr_t csum_daddr;
+ void *csum_buf;
+ struct xfs_buf *csum_bp;
+ struct xfs_trans *tp;
+ int error;
+ unsigned int i = 0;
+
+ *nr = min(*nr, xfs_rtcsum_len_to_extlen(mp, mp->m_rtcsum_bsize - boff));
+
+ error = xfs_rtcsum_bmap(rtg, rgbno, &csum_daddr);
+ if (error) {
+ do_error(
+_("Cannot bmap new csum file for RG %u, error %d\n"),
+ rtg_rgno(rtg), error);
+ return error;
+ }
+
+ tres.tr_logres = xfs_calc_csum_reservation(mp,
+ xfs_extlen_to_rtcsum_len(mp, *nr));
+ ASSERT(tres.tr_logres <= M_RES(mp)->tr_csum.tr_logres);
+
+ error = libxfs_trans_alloc(mp, &tres, 0, 0, 0, &tp);
+ if (error) {
+ do_error(
+_("Transaction allocation for csum repair failed: %d\n"),
+ error);
+ return error;
+ }
+
+ error = libxfs_buf_read(mp->m_ddev_targp, csum_daddr,
+ BTOBB(mp->m_rtcsum_bsize), 0, &csum_bp,
+ &xfs_rtcsum_buf_ops);
+ if (error) {
+ do_error(
+_("Cannot get buffer for new csum file for RG %u, error %d\n"),
+ rtg_rgno(rtg), error);
+ return error;
+ }
+
+ libxfs_trans_bjoin(tp, csum_bp);
+ xfs_trans_buf_set_type(tp, csum_bp, XFS_BLFT_RTCSUM_BUF);
+
+ csum_buf = csum_bp->b_addr + boff;
+ for (i = 0; i < *nr; i++) {
+ union xfs_csum csum;
+
+ xfs_csum_seed(mp, &csum);
+ xfs_csum_gen(mp, data_buf + XFS_FSB_TO_B(mp, i),
+ mp->m_sb.sb_blocksize, &csum);
+ if (i == 0)
+ printf("calculated checksum 0x%x\n", csum.crc32c);
+ xfs_csum_finalize(mp, csum_buf, &csum);
+ csum_buf += (1u << mp->m_rtcsum_shift);
+ }
+
+ libxfs_trans_log_buf(tp, csum_bp, boff,
+ boff + xfs_extlen_to_rtcsum_len(mp, *nr) - 1);
+ return -libxfs_trans_commit(tp);
+}
+
+void
+calculate_rtgroup_csums(
+ struct xfs_rtgroup *rtg)
+{
+ struct xfs_mount *mp = rtg_mount(rtg);
+ xfs_rgblock_t rgbno;
+ int error;
+ void *data_buf;
+
+ error = posix_memalign(&data_buf, sysconf(_SC_PAGESIZE),
+ XFS_FSB_TO_B(mp, DATA_BLOCKS_PER_BUF));
+ if (error) {
+ do_warn(
+_("Failed to allocate memory for csum rebuild\n"));
+ return;
+ }
+ for (rgbno = 0; rgbno < rtg_blocks(rtg); rgbno += DATA_BLOCKS_PER_BUF) {
+ unsigned int nr_blocks, done = 0;
+
+ nr_blocks = min(DATA_BLOCKS_PER_BUF, rtg_blocks(rtg) - rgbno);
+ error = read_data(mp->m_rtdev_targp,
+ xfs_gbno_to_daddr(rtg_group(rtg), rgbno),
+ XFS_FSB_TO_BB(mp, nr_blocks),
+ data_buf);
+ if (error) {
+ do_warn(
+_("Reading data at RG %u/%u failed, error %d"),
+ rtg_rgno(rtg), rgbno, error);
+ continue;
+ }
+
+ do {
+ unsigned int n = nr_blocks - done;
+
+ calculate_buf_csums(rtg,
+ data_buf + XFS_FSB_TO_B(mp, done),
+ rgbno + done, &n);
+ done += n;
+ } while (done < nr_blocks);
+ }
+
+ kfree(data_buf);
+}
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* [PATCH 32/32] xfs_scrub: don't merge over unused space when data checksums are enabled
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
` (30 preceding siblings ...)
2026-09-24 10:04 ` [PATCH 31/32] repair: support RT data checksums Christoph Hellwig
@ 2026-09-24 10:04 ` Christoph Hellwig
2026-09-25 23:51 ` Darrick J. Wong
31 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-24 10:04 UTC (permalink / raw)
To: Andrey Albershteyn; +Cc: Darrick J . Wong, linux-xfs
Unused space could be lost space at mount time for which no metadata was
recorded, and we thus might not have valid checksums for it.
Signed-off-by: Christoph Hellwig <hch@lst.de>
---
scrub/read_verify.c | 33 +++++++++++++++++++++++----------
1 file changed, 23 insertions(+), 10 deletions(-)
diff --git a/scrub/read_verify.c b/scrub/read_verify.c
index 311ea1c29008..0c3bec104d16 100644
--- a/scrub/read_verify.c
+++ b/scrub/read_verify.c
@@ -542,20 +542,33 @@ try_read_verify_schedule_io(
/*
* If we have a stashed IO, we haven't changed pools, the error
- * reporting is the same, and the two extents are close,
- * we can combine them.
+ * reporting is the same, and the two extents are close, we can combine
+ * them.
+ *
+ * If the file system has data checksums enabled, only merge fully
+ * contigous ranges, as unused data might have invalid checksums.
*/
- if (rs->rvp == rvp && rs->io_length > 0 &&
- ((start >= rs->io_start && start <= rv_end + locality) ||
- (rs->io_start >= start &&
- rs->io_start <= req_end + locality))) {
- rs->io_start = min(rs->io_start, start);
- rs->io_length = max(req_end, rv_end) - rs->io_start;
-
- return true;
+ if (rs->rvp != rvp || !rs->io_length)
+ return false;
+
+ if (rvp->ctx->mnt.fsgeom.flags & XFS_FSOP_GEOM_FLAGS_DATA_CSUM) {
+ if (start == rv_end)
+ goto merge;
+ if (rs->io_start == req_end)
+ goto merge;
+ } else {
+ if (start >= rs->io_start && start <= rv_end + locality)
+ goto merge;
+ if (rs->io_start >= start && rs->io_start <= req_end + locality)
+ goto merge;
}
return false;
+
+merge:
+ rs->io_start = min(rs->io_start, start);
+ rs->io_length = max(req_end, rv_end) - rs->io_start;
+ return true;
}
/* Did read verification succeed? */
--
2.53.0
^ permalink raw reply related [flat|nested] 57+ messages in thread
* Re: [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2
2026-09-24 10:03 ` [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2 Christoph Hellwig
@ 2026-09-24 20:30 ` Darrick J. Wong
2026-09-26 6:15 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-24 20:30 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:03:51PM +0200, Christoph Hellwig wrote:
> Using tabs for indentation messes up man page rendering, so use spaces
> for rtstart to match the other fields.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
This could go into 7.3.
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> man/man2/ioctl_xfs_fsgeometry.2 | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/man/man2/ioctl_xfs_fsgeometry.2 b/man/man2/ioctl_xfs_fsgeometry.2
> index 037f8e15e415..6d30610b71be 100644
> --- a/man/man2/ioctl_xfs_fsgeometry.2
> +++ b/man/man2/ioctl_xfs_fsgeometry.2
> @@ -50,7 +50,7 @@ struct xfs_fsop_geom {
> __u32 sick;
> __u32 checked;
> __u64 rgextents;
> - __u64 rtstart;
> + __u64 rtstart;
> __u64 rtreserved;
> __u64 reserved[14];
> };
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 02/32] add cpu_to_le64 and le64_to_cpu_helpers
2026-09-24 10:03 ` [PATCH 02/32] add cpu_to_le64 and le64_to_cpu_helpers Christoph Hellwig
@ 2026-09-24 20:30 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-24 20:30 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:03:52PM +0200, Christoph Hellwig wrote:
> The crc64 handling will need them.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
Looks good to me
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> include/xfs_arch.h | 4 ++++
> 1 file changed, 4 insertions(+)
>
> diff --git a/include/xfs_arch.h b/include/xfs_arch.h
> index d46ae47094ae..348ff0077c99 100644
> --- a/include/xfs_arch.h
> +++ b/include/xfs_arch.h
> @@ -195,6 +195,8 @@ static __inline__ void __swab64s(__u64 *addr)
>
> #define cpu_to_le32(val) ((__force __be32)__swab32((__u32)(val)))
> #define le32_to_cpu(val) (__swab32((__force __u32)(__le32)(val)))
> +#define cpu_to_le64(val) ((__force __be64)__swab64((__u64)(val)))
> +#define le64_to_cpu(val) (__swab64((__force __u64)(__le64)(val)))
>
> #define __constant_cpu_to_le32(val) \
> ((__force __le32)___constant_swab32((__u32)(val)))
> @@ -210,6 +212,8 @@ static __inline__ void __swab64s(__u64 *addr)
>
> #define cpu_to_le32(val) ((__force __le32)(__u32)(val))
> #define le32_to_cpu(val) ((__force __u32)(__le32)(val))
> +#define cpu_to_le64(val) ((__force __le64)(__u64)(val))
> +#define le64_to_cpu(val) ((__force __u64)(__le64)(val))
>
> #define __constant_cpu_to_le32(val) \
> ((__force __le32)(__u32)(val))
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL
2026-09-24 10:03 ` [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL Christoph Hellwig
@ 2026-09-24 20:44 ` Darrick J. Wong
2026-09-25 6:13 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-24 20:44 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:03:54PM +0200, Christoph Hellwig wrote:
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
> libxfs/xfs_platform.h | 1 +
> 1 file changed, 1 insertion(+)
>
> diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
> index d09fb0aae00f..a61d50de9dc0 100644
> --- a/libxfs/xfs_platform.h
> +++ b/libxfs/xfs_platform.h
> @@ -212,6 +212,7 @@ void inode_init_owner(struct mnt_idmap *idmap, struct inode *inode,
> #define round_up(x, y) ((((x)-1) | __round_mask(x, y))+1)
> #define round_down(x, y) ((x) & ~__round_mask(x, y))
> #define DIV_ROUND_UP(n,d) (((n) + (d) - 1) / (d))
> +#define DIV_ROUND_UP_ULL(n,d) DIV_ROUND_UP(n,d)
Er... shouldn't there be some sort of u64 cast here?
--D
>
> /*
> * Handling for kernel bitmap types.
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL
2026-09-24 20:44 ` Darrick J. Wong
@ 2026-09-25 6:13 ` Christoph Hellwig
0 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-25 6:13 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 01:44:37PM -0700, Darrick J. Wong wrote:
> > #define DIV_ROUND_UP(n,d) (((n) + (d) - 1) / (d))
> > +#define DIV_ROUND_UP_ULL(n,d) DIV_ROUND_UP(n,d)
>
> Er... shouldn't there be some sort of u64 cast here?
I think the only reason it exists in the kernel is to add the do_div
for 64-bit divisions on 32-bit. It also is a macro in the kernel, so
there is no type-promotion there either.
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 05/32] libxfs: remove a spurious xfs_rtbitmap.h include in xfs_sb.c
2026-09-24 10:03 ` [PATCH 05/32] libxfs: remove a spurious xfs_rtbitmap.h include in xfs_sb.c Christoph Hellwig
@ 2026-09-25 23:27 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:27 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:03:55PM +0200, Christoph Hellwig wrote:
> This fully resyncs with the kernel version.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
Looks fine to me
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> libxfs/xfs_sb.c | 1 -
> 1 file changed, 1 deletion(-)
>
> diff --git a/libxfs/xfs_sb.c b/libxfs/xfs_sb.c
> index ea99f8b5eee5..cccbcd153316 100644
> --- a/libxfs/xfs_sb.c
> +++ b/libxfs/xfs_sb.c
> @@ -27,7 +27,6 @@
> #include "xfs_rtgroup.h"
> #include "xfs_rtrmap_btree.h"
> #include "xfs_rtrefcount_btree.h"
> -#include "xfs_rtbitmap.h"
>
> /*
> * Physical superblock buffer manipulations. Shared with libxfs in userspace.
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 06/32] libxfs: add SZ_* constants
2026-09-24 10:03 ` [PATCH 06/32] libxfs: add SZ_* constants Christoph Hellwig
@ 2026-09-25 23:27 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:27 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:03:56PM +0200, Christoph Hellwig wrote:
> Import a slightly adapted version of <linux/sizes.h> from the kernel so that
> we can use these constants in shared libxfs code.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
Marvelous!
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> include/libxfs.h | 1 +
> include/sizes.h | 67 +++++++++++++++++++++++++++++++++++++++++++
> libxfs/xfs_platform.h | 1 +
> 3 files changed, 69 insertions(+)
> create mode 100644 include/sizes.h
>
> diff --git a/include/libxfs.h b/include/libxfs.h
> index 7f0c22ef4991..34dd267831c0 100644
> --- a/include/libxfs.h
> +++ b/include/libxfs.h
> @@ -18,6 +18,7 @@
> #include "platform_defs.h"
> #include "xfs.h"
>
> +#include "sizes.h"
> #include "list.h"
> #include "hlist.h"
> #include "cache.h"
> diff --git a/include/sizes.h b/include/sizes.h
> new file mode 100644
> index 000000000000..69075e8f3370
> --- /dev/null
> +++ b/include/sizes.h
> @@ -0,0 +1,67 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +#ifndef __SIZES_H__
> +#define __SIZES_H__
> +
> +#define SZ_1 0x00000001
> +#define SZ_2 0x00000002
> +#define SZ_4 0x00000004
> +#define SZ_8 0x00000008
> +#define SZ_16 0x00000010
> +#define SZ_32 0x00000020
> +#define SZ_64 0x00000040
> +#define SZ_128 0x00000080
> +#define SZ_256 0x00000100
> +#define SZ_512 0x00000200
> +
> +#define SZ_1K 0x00000400
> +#define SZ_2K 0x00000800
> +#define SZ_4K 0x00001000
> +#define SZ_8K 0x00002000
> +#define SZ_16K 0x00004000
> +#define SZ_24K 0x00006000
> +#define SZ_32K 0x00008000
> +#define SZ_64K 0x00010000
> +#define SZ_128K 0x00020000
> +#define SZ_192K 0x00030000
> +#define SZ_256K 0x00040000
> +#define SZ_384K 0x00060000
> +#define SZ_512K 0x00080000
> +
> +#define SZ_1M 0x00100000
> +#define SZ_2M 0x00200000
> +#define SZ_3M 0x00300000
> +#define SZ_4M 0x00400000
> +#define SZ_6M 0x00600000
> +#define SZ_8M 0x00800000
> +#define SZ_12M 0x00c00000
> +#define SZ_16M 0x01000000
> +#define SZ_18M 0x01200000
> +#define SZ_24M 0x01800000
> +#define SZ_32M 0x02000000
> +#define SZ_64M 0x04000000
> +#define SZ_128M 0x08000000
> +#define SZ_256M 0x10000000
> +#define SZ_512M 0x20000000
> +
> +#define SZ_1G 0x40000000
> +#define SZ_2G 0x80000000
> +
> +#define SZ_4G 0x100000000ULL
> +#define SZ_8G 0x200000000ULL
> +#define SZ_16G 0x400000000ULL
> +#define SZ_32G 0x800000000ULL
> +#define SZ_64G 0x1000000000ULL
> +#define SZ_128G 0x2000000000ULL
> +#define SZ_256G 0x4000000000ULL
> +#define SZ_512G 0x8000000000ULL
> +
> +#define SZ_1T 0x10000000000ULL
> +#define SZ_2T 0x20000000000ULL
> +#define SZ_4T 0x40000000000ULL
> +#define SZ_8T 0x80000000000ULL
> +#define SZ_16T 0x100000000000ULL
> +#define SZ_32T 0x200000000000ULL
> +#define SZ_64T 0x400000000000ULL
> +#define SZ_128T 0x800000000000ULL
> +
> +#endif /* __SIZES_H__ */
> diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
> index a61d50de9dc0..0aeb80fb8244 100644
> --- a/libxfs/xfs_platform.h
> +++ b/libxfs/xfs_platform.h
> @@ -45,6 +45,7 @@
> #include "platform_defs.h"
> #include "xfs.h"
>
> +#include "sizes.h"
> #include "list.h"
> #include "hlist.h"
> #include "cache.h"
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 09/32] FIXUP
2026-09-24 10:03 ` [PATCH 09/32] FIXUP Christoph Hellwig
@ 2026-09-25 23:29 ` Darrick J. Wong
2026-09-26 6:17 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:29 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:03:59PM +0200, Christoph Hellwig wrote:
> ---
> libxfs/libxfs_io.h | 6 ++++++
> libxfs/rdwr.c | 2 --
> libxfs/xfs_platform.h | 1 -
> 3 files changed, 6 insertions(+), 3 deletions(-)
>
> diff --git a/libxfs/libxfs_io.h b/libxfs/libxfs_io.h
> index 5562e2928254..bd0d9ec04bea 100644
> --- a/libxfs/libxfs_io.h
> +++ b/libxfs/libxfs_io.h
> @@ -293,4 +293,10 @@ xfs_buftarg_verify_daddr(
> return daddr < xfs_buftarg_nr_sectors(btp);
> }
>
> +static inline void
> +xfs_buf_set_uptodate(
> + struct xfs_buf *bp)
> +{
> +}
Should this set LIBXFS_B_UPTODATE?
--D
> +
> #endif /* __LIBXFS_IO_H__ */
> diff --git a/libxfs/rdwr.c b/libxfs/rdwr.c
> index 90f2d56687ca..9b6cbd4aa158 100644
> --- a/libxfs/rdwr.c
> +++ b/libxfs/rdwr.c
> @@ -1426,8 +1426,6 @@ __xfs_buf_mark_corrupt(
> struct xfs_buf *bp,
> xfs_failaddr_t fa)
> {
> - ASSERT(bp->b_flags & XBF_DONE);
> -
> xfs_buf_corruption_error(bp, fa);
> xfs_buf_stale(bp);
> }
> diff --git a/libxfs/xfs_platform.h b/libxfs/xfs_platform.h
> index 0aeb80fb8244..d9cfcec4c1a0 100644
> --- a/libxfs/xfs_platform.h
> +++ b/libxfs/xfs_platform.h
> @@ -322,7 +322,6 @@ static inline unsigned long long mask64_if_power2(unsigned long b)
>
> /* buffer management */
> #define XBF_TRYLOCK 0
> -#define XBF_DONE 0
> #define xfs_buf_stale(bp) ((bp)->b_flags |= LIBXFS_B_STALE)
> #define XFS_BUF_UNDELAYWRITE(bp) ((bp)->b_flags &= ~LIBXFS_B_DIRTY)
>
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 26/32] man: document the rtcsum geom fields
2026-09-24 10:04 ` [PATCH 26/32] man: document the rtcsum geom fields Christoph Hellwig
@ 2026-09-25 23:31 ` Darrick J. Wong
2026-09-26 6:17 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:31 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:16PM +0200, Christoph Hellwig wrote:
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
> man/man2/ioctl_xfs_fsgeometry.2 | 31 +++++++++++++++++++++++++++++--
> 1 file changed, 29 insertions(+), 2 deletions(-)
>
> diff --git a/man/man2/ioctl_xfs_fsgeometry.2 b/man/man2/ioctl_xfs_fsgeometry.2
> index 6d30610b71be..97216df7525b 100644
> --- a/man/man2/ioctl_xfs_fsgeometry.2
> +++ b/man/man2/ioctl_xfs_fsgeometry.2
> @@ -52,7 +52,10 @@ struct xfs_fsop_geom {
> __u64 rgextents;
> __u64 rtstart;
> __u64 rtreserved;
> - __u64 reserved[14];
> + __u8 rtcsum_type;
> + __u8 rtcsum_blklog;
> + __u8 reserved_pad[6];
> + __u64 reserved[13];
> };
> .fi
> .in
> @@ -159,8 +162,32 @@ This field is meaningful only if the flag
> .B XFS_FSOP_GEOM_FLAGS_ZONED
> is set.
> .PP
> +.I rtcsum_type
> +Type of RT data checksum. This field is meaningful only if the flag
> +.B XFS_FSOP_GEOM_FLAGS_DATA_CSUM
> +is set.
> +The following values are defined:
> +.RS 0.4i
> +.TP
> +.B XFS_CSUM_TYPE_NONE
> +Not data checksum.
"Data checksums are not enabled."
> +.TP
> +.B XFS_CSUM_TYPE_CRC32C
> +CRC32c as per NVMe.
Just for my education -- crc32c as per nvme is the same as crc32c
everywhere else right?
--D
> +.TP
> +.B XFS_CSUM_TYPE_CRC64
> +CRC64 as per NVMe.
> +.TP
> +.RE
> +.PP
> +.I rtcsum_blklog
> +log(2) of the RT data checksum block size.
> +This field is meaningful only if the flag
> +.PP
> +.I reserved_pad
> +and
> .I reserved
> -is set to zero.
> +are set to zero.
> .SH FILESYSTEM FEATURE FLAGS
> Filesystem features are reported to userspace as a combination the following
> flags:
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 27/32] libfrog: print csum geometry information
2026-09-24 10:04 ` [PATCH 27/32] libfrog: print csum geometry information Christoph Hellwig
@ 2026-09-25 23:32 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:32 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:17PM +0200, Christoph Hellwig wrote:
> Signed-off-by: Christoph Hellwig <hch@lst.de>
Looks good to me,
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> libfrog/fsgeom.c | 23 +++++++++++++++++++++--
> 1 file changed, 21 insertions(+), 2 deletions(-)
>
> diff --git a/libfrog/fsgeom.c b/libfrog/fsgeom.c
> index ad0476741cf5..324ccdc9b84d 100644
> --- a/libfrog/fsgeom.c
> +++ b/libfrog/fsgeom.c
> @@ -28,6 +28,22 @@ rtdev_name(
> return rtname;
> }
>
> +static const char *
> +csum_name(
> + uint8_t csum)
> +{
> + switch (csum) {
> + case XFS_CSUM_TYPE_NONE:
> + return "none";
> + case XFS_CSUM_TYPE_CRC32C:
> + return "crc32";
> + case XFS_CSUM_TYPE_CRC64:
> + return "crc64c";
> + default:
> + return "invalid";
> + }
> +};
> +
> void
> xfs_report_geom(
> struct xfs_fsop_geom *geo,
> @@ -91,7 +107,8 @@ xfs_report_geom(
> " =%-22s sectsz=%-5u sunit=%d blks, lazy-count=%d\n"
> "realtime =%-22s extsz=%-6d blocks=%lld, rtextents=%lld\n"
> " =%-22s rgcount=%-4d rgsize=%u extents\n"
> -" =%-22s zoned=%-6d start=%llu reserved=%llu\n"),
> +" =%-22s zoned=%-6d start=%llu reserved=%llu\n"
> +" =%-22s csum=%-7s csumbsize=%u\n"),
> mntpoint, geo->inodesize, geo->agcount, geo->agblocks,
> "", geo->sectsize, attrversion, projid32bit,
> "", crcs_enabled, finobt_enabled, spinodes, rmapbt_enabled,
> @@ -108,7 +125,9 @@ xfs_report_geom(
> geo->rtextsize * geo->blocksize, (unsigned long long)geo->rtblocks,
> (unsigned long long)geo->rtextents,
> "", geo->rgcount, geo->rgextents,
> - "", zoned, geo->rtstart, geo->rtreserved);
> + "", zoned, geo->rtstart, geo->rtreserved,
> + "", csum_name(geo->rtcsum_type), geo->rtcsum_blklog ?
> + (1u << geo->rtcsum_blklog) : 0);
> }
>
> /* Try to obtain the xfs geometry. On error returns a negative error code. */
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs
2026-09-24 10:04 ` [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs Christoph Hellwig
@ 2026-09-25 23:33 ` Darrick J. Wong
2026-09-26 6:18 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:33 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:18PM +0200, Christoph Hellwig wrote:
> Signed-off-by: Christoph Hellwig <hch@lst.de>
Seems fine, though I wonder if we should report 1U<<rtcsum_blklog?
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> ---
> io/stat.c | 4 ++--
> 1 file changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/io/stat.c b/io/stat.c
> index e3b3574e0416..c6175298631c 100644
> --- a/io/stat.c
> +++ b/io/stat.c
> @@ -286,8 +286,8 @@ statfs_f(
> printf(_("geom.rgcount = %u\n"), fsgeo.rgcount);
> printf(_("geom.rtstart = %llu\n"),
> (unsigned long long)fsgeo.rtstart);
> - printf(_("geom.rtreserved = %llu\n"),
> - (unsigned long long)fsgeo.rtreserved);
> + printf(_("geom.rtcsum_type = %u\n"), fsgeo.rtcsum_type);
> + printf(_("geom.rtcsum_blklog = %u\n"), fsgeo.rtcsum_blklog);
> }
> }
>
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 29/32] mkfs: support RT data checksums
2026-09-24 10:04 ` [PATCH 29/32] mkfs: support RT data checksums Christoph Hellwig
@ 2026-09-25 23:39 ` Darrick J. Wong
2026-09-26 6:19 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:39 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:19PM +0200, Christoph Hellwig wrote:
> Parse the -r csum and -r csumbisze options, set the superblock fields
csumbsize
> based on that and create the per-RTG csum files when requested.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
> man/man8/mkfs.xfs.8.in | 17 ++++++
> mkfs/proto.c | 9 +++
> mkfs/xfs_mkfs.c | 126 ++++++++++++++++++++++++++++++++++++++++-
> 3 files changed, 149 insertions(+), 3 deletions(-)
>
> diff --git a/man/man8/mkfs.xfs.8.in b/man/man8/mkfs.xfs.8.in
> index 927cd075bf8a..b2da1e76f10f 100644
> --- a/man/man8/mkfs.xfs.8.in
> +++ b/man/man8/mkfs.xfs.8.in
> @@ -1315,6 +1315,23 @@ Controls the amount of space in the realtime section that is reserved for
> internal use by garbage collection and reorganization algorithms.
> Defaults to 0 if not set.
> This option is only valid if the zoned realtime allocator is used.
> +.TP
> +.BI csum= value
> +Controls which data checksum is used.
> +Valid options are
> +.I none
> +,
> +.I crc32c
> +and
> +.I crc64.
> +Defaults to
> +.I none
> +if not set.
> +This option is only valid if the zoned realtime allocator is used.
> +.TP
> +.BI csumbsize= value
> +Controls the blocksize of the data checksum files. Valid values are
> +32k and 64k. Defaults to 32k or the filesystem block size if larger.
> .RE
> .PP
> .PD 0
> diff --git a/mkfs/proto.c b/mkfs/proto.c
> index 7ef8cc9d88a1..c232d1171d44 100644
> --- a/mkfs/proto.c
> +++ b/mkfs/proto.c
> @@ -11,6 +11,8 @@
> #include <sys/xattr.h>
> #include <linux/xattr.h>
> #include "libfrog/convert.h"
> +#include "libfrog/crc64.h"
> +#include "libxfs/xfs_rtcsumfile.h"
> #include "proto.h"
>
> /*
> @@ -1221,6 +1223,13 @@ rtinit_groups(
> fail(_("rtrmap rtsb init failed"), error);
> }
>
> + if (xfs_has_rtcsum(mp)) {
> + error = xfs_rtcsum_alloc_blocks(rtg);
> + if (error)
> + fail(_("Initialization of rtcsum inode failed"),
> + error);
> + }
> +
> if (!xfs_has_zoned(mp))
> rtfreesp_init(rtg);
> }
> diff --git a/mkfs/xfs_mkfs.c b/mkfs/xfs_mkfs.c
> index 334367b81555..9abeca61ee2a 100644
> --- a/mkfs/xfs_mkfs.c
> +++ b/mkfs/xfs_mkfs.c
> @@ -6,15 +6,17 @@
> #include "libfrog/util.h"
> #include "libxfs.h"
> #include <ctype.h>
> -#include "libxfs/xfs_zones.h"
> #include "xfs_multidisk.h"
> #include "libxcmd.h"
> #include "libfrog/fsgeom.h"
> #include "libfrog/convert.h"
> #include "libfrog/crc32cselftest.h"
> +#include "libfrog/crc64.h"
> #include "libfrog/dahashselftest.h"
> #include "libfrog/fsproperties.h"
> #include "libfrog/zones.h"
> +#include "libxfs/xfs_zones.h"
> +#include "libxfs/xfs_rtcsumfile.h"
> #include "proto.h"
> #include <ini.h>
>
> @@ -145,6 +147,8 @@ enum {
> R_ZONED,
> R_START,
> R_RESERVED,
> + R_CSUM,
> + R_CSUMBSIZE,
BTW now that Andrey has merged the config file generator code, you'll
have to add the relevant R_CSUM/R_CSUMBSIZE bits to libfrog/fsgeom.c.
> R_MAX_OPTS,
> };
>
> @@ -777,6 +781,8 @@ static struct opt_params ropts = {
> [R_ZONED] = "zoned",
> [R_START] = "start",
> [R_RESERVED] = "reserved",
> + [R_CSUM] = "csum",
> + [R_CSUMBSIZE] = "csumbsize",
> [R_MAX_OPTS] = NULL,
> },
> .subopt_params = {
> @@ -864,6 +870,19 @@ static struct opt_params ropts = {
> .maxval = LLONG_MAX,
> .defaultval = SUBOPT_NEEDS_VAL,
> },
> + { .index = R_CSUM,
> + .conflicts = { { NULL, LAST_CONFLICT } },
> + .minval = 0,
> + .maxval = LLONG_MAX,
> + .defaultval = SUBOPT_NEEDS_VAL,
> + },
> + { .index = R_CSUMBSIZE,
> + .conflicts = { { NULL, LAST_CONFLICT } },
> + .convert = true,
> + .minval = 1u << XFS_RTCSUM_BSIZE_LOG_MIN,
> + .maxval = 1u << XFS_RTCSUM_BSIZE_LOG_MAX,
> + .defaultval = SUBOPT_NEEDS_VAL,
> + },
> },
> };
>
> @@ -1074,6 +1093,8 @@ struct sb_feat_args {
> bool exchrange; /* XFS_SB_FEAT_INCOMPAT_EXCHRANGE */
> bool zoned;
> bool zone_gaps;
> + uint8_t rtcsum_type;
> + uint8_t rtcsum_blklog;
>
> uint16_t qflags;
> };
> @@ -1102,6 +1123,7 @@ struct cli_params {
> char *rtstart;
> uint64_t rtreserved;
> char *max_atomic_write;
> + char *rtcsumbsize;
>
> /* parameters where 0 is a valid CLI value */
> int dsunit;
> @@ -1191,6 +1213,7 @@ struct mkfs_params {
> struct sb_feat_args sb_feat;
> uint64_t rtstart;
> uint64_t rtreserved;
> + uint32_t rtcsumbsize;
>
> uint64_t max_atomic_write;
> };
> @@ -1229,7 +1252,8 @@ usage( void )
> pquota|pqnoenforce]\n\
> /* data subvol */ [-d agcount=n,agsize=n,file,name=xxx,size=num,\n\
> (sunit=value,swidth=value|su=num,sw=num|noalign),\n\
> - sectsize=num,concurrency=num]\n\
> + sectsize=num,concurrency=num,\n\
> + csum=type]\n\
There's no csum= parameter for -d, or at least I didn't see a D_CSUM
entry above.
> /* force overwrite */ [-f]\n\
> /* inode size */ [-i perblock=n|size=num,maxpct=n,attr=0|1|2,\n\
> projid32bit=0|1,sparse=0|1,nrext64=0|1,\n\
> @@ -1245,7 +1269,8 @@ usage( void )
> /* populate from directory */ [-p dirname,atime=0|1]\n\
> /* quiet */ [-q]\n\
> /* realtime subvol */ [-r extsize=num,size=num,rtdev=xxx,rgcount=n,rgsize=n,\n\
> - concurrency=num,zoned=0|1,start=n,reserved=n]\n\
> + concurrency=num,zoned=0|1,start=n,reserved=n,\n\
> + csum=type],csumbsize=n\n\
> /* sectorsize */ [-s size=num]\n\
> /* version */ [-V]\n\
> devicename\n\
> @@ -1858,6 +1883,24 @@ set_data_concurrency(
> cli->data_concurrency = optnum;
> }
>
> +static bool
> +set_data_csum(
> + const char *value,
> + uint8_t *csum_val)
> +{
> + if (!value)
> + return false;
> + if (!strcmp(value, "none"))
> + *csum_val = XFS_CSUM_TYPE_NONE;
> + else if (!strcmp(value, "crc32c"))
> + *csum_val = XFS_CSUM_TYPE_CRC32C;
> + else if (!strcmp(value, "crc64"))
> + *csum_val = XFS_CSUM_TYPE_CRC64;
> + else
> + return false;
> + return true;
> +}
> +
> static int
> data_opts_parser(
> struct opt_params *opts,
> @@ -2269,6 +2312,13 @@ rtdev_opts_parser(
> case R_RESERVED:
> cli->rtreserved = getnum(value, opts, subopt);
> break;
> + case R_CSUM:
> + if (!set_data_csum(value, &cli->sb_feat.rtcsum_type))
> + return -EINVAL;
> + break;
> + case R_CSUMBSIZE:
> + cli->rtcsumbsize = getstr(value, opts, subopt);
> + break;
> default:
> return -EINVAL;
> }
> @@ -3048,6 +3098,11 @@ _("internal RT section only supported in zoned mode\n"));
> _("reserved RT blocks only supported in zoned mode\n"));
> usage();
> }
> + if (cli->sb_feat.rtcsum_type || cli->sb_feat.rtcsum_blklog) {
> + fprintf(stderr,
> +_("data checksums not supported without zoned mode\n"));
> + usage();
> + }
> }
>
> if (cli->xi->rt.name || cfg->rtstart) {
> @@ -4921,6 +4976,11 @@ sb_set_features(
> }
> if (fp->zone_gaps)
> sbp->sb_features_incompat |= XFS_SB_FEAT_INCOMPAT_ZONE_GAPS;
> + if (fp->rtcsum_type) {
> + sbp->sb_features_ro_compat |= XFS_SB_FEAT_RO_COMPAT_RTCSUM;
> + sbp->sb_rtcsum_type = fp->rtcsum_type;
> + sbp->sb_rtcsum_blklog = fp->rtcsum_blklog;
> + }
> }
>
> /*
> @@ -5243,6 +5303,63 @@ validate_max_atomic_write_ags(
> }
> }
>
> +static void
> +validate_rtcsumbsize(
> + struct mkfs_params *cfg,
> + struct cli_params *cli,
> + struct xfs_mount *mp)
> +{
> + unsigned int l;
> +
> + if (!cli->sb_feat.rtcsum_type) {
> + if (cli->rtcsumbsize) {
> + fprintf(stderr,
> + _("RT data csum block size options requires RT data csum type.\n"));
> + exit(1);
> + }
> + return;
> + }
> +
> + if (!cfg->sb_feat.zoned) {
> + fprintf(stderr,
> + _("RT data checksums require a zoned RT subvolume\n"));
> + exit(1);
> + }
> +
> + if (cli->rtcsumbsize) {
> + cfg->rtcsumbsize =
> + getnum(cli->rtcsumbsize, &ropts, R_CSUMBSIZE);
> + } else {
> + cfg->rtcsumbsize =
> + max(1u << XFS_RTCSUM_BSIZE_LOG_MIN, cfg->blocksize);
> + }
> +
> + if (!is_power_of_2(cfg->rtcsumbsize)) {
> + fprintf(stderr,
> + _("RT data csum block size of %u bytes is not a power of 2\n"),
> + cfg->rtcsumbsize);
> + exit(1);
> + }
> +
> + if (cfg->rtcsumbsize < cfg->blocksize) {
> + fprintf(stderr,
> + _("RT data csum block size of %u bytes is smaller than fsblock size %u.\n"),
> + cfg->rtcsumbsize, cfg->blocksize);
> + exit(1);
> + }
> +
> + if (cfg->rtcsumbsize % cfg->blocksize) {
> + fprintf(stderr,
> + _("RT data csum block size of %u bytes not aligned with fsblock size %u.\n"),
> + cfg->rtcsumbsize, cfg->blocksize);
> + exit(1);
> + }
> +
> + cfg->sb_feat.rtcsum_blklog = 0;
> + for (l = cfg->rtcsumbsize; l > 1; l >>= 1)
> + cfg->sb_feat.rtcsum_blklog++;
log2_roundup?
--D
> +}
> +
> static void
> calculate_log_size(
> struct mkfs_params *cfg,
> @@ -5523,6 +5640,8 @@ finish_superblock_setup(
> sbp->sb_unit = cfg->dsunit;
> sbp->sb_width = cfg->dswidth;
> mp->m_features |= libxfs_sb_version_to_features(sbp);
> + mp->m_rtcsum_shift = xfs_data_csum_shift(mp->m_sb.sb_rtcsum_type);
> + mp->m_rtcsum_bsize = 1u << mp->m_sb.sb_rtcsum_blklog;
> libxfs_sb_mount_rextsize(mp, sbp);
> }
>
> @@ -6290,6 +6409,7 @@ main(
> validate_datadev(&cfg, &cli);
> validate_logdev(&cfg, &cli);
> validate_rtdev(&cfg, &cli, &zt);
> + validate_rtcsumbsize(&cfg, &cli, mp);
> calc_stripe_factors(&cfg, &cli, &ft);
>
> /*
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 30/32] xfs_db: support RT data checksum
2026-09-24 10:04 ` [PATCH 30/32] xfs_db: support RT data checksum Christoph Hellwig
@ 2026-09-25 23:43 ` Darrick J. Wong
2026-09-26 6:21 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:43 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:20PM +0200, Christoph Hellwig wrote:
> Support printing the new super block fields, and content of the per-RTG
> rtcsum files.
>
> Note that the superblock fields are printed for all metadir file systems
> because they reuse space that previously was padding.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
> db/block.c | 10 +++++++++-
> db/field.c | 2 ++
> db/field.h | 1 +
> db/inode.c | 2 ++
> db/rtgroup.c | 42 ++++++++++++++++++++++++++++++++++++++++++
> db/rtgroup.h | 5 +++++
> db/sb.c | 10 ++++++++--
> db/type.c | 5 +++++
> db/type.h | 1 +
> man/man8/xfs_db.8 | 13 ++++++++++++-
> 10 files changed, 87 insertions(+), 4 deletions(-)
>
> diff --git a/db/block.c b/db/block.c
> index 2f1978c41f30..c6dd50dfa764 100644
> --- a/db/block.c
> +++ b/db/block.c
> @@ -263,7 +263,15 @@ dblock_f(
> dbprintf(_("no type for file data\n"));
> return 0;
> }
> - nex = nb = type == TYP_DIR2 ? mp->m_dir_geo->fsbcount : 1;
> +
> + if (type == TYP_DIR2)
> + nb = mp->m_dir_geo->fsbcount;
> + else if (type == TYP_RTCSUM)
> + nb = XFS_B_TO_FSB(mp, 1u << mp->m_sb.sb_rtcsum_blklog);
> + else
> + nb = 1;
> +
> + nex = nb;
> bmp = malloc(nb * sizeof(*bmp));
> bmap(bno, nb, XFS_DATA_FORK, &nex, bmp);
> if (nex == 0) {
> diff --git a/db/field.c b/db/field.c
> index 62c983fc3d49..275f175493df 100644
> --- a/db/field.c
> +++ b/db/field.c
> @@ -441,6 +441,8 @@ const ftattr_t ftattrtab[] = {
> 0, NULL, NULL },
> { FLDT_RGSUMMARY, "rgsummary", NULL, (char *)rgsummary_flds,
> btblock_size, FTARG_SIZE, NULL, rgsummary_flds },
> + { FLDT_RTCSUM, "rtcsum", NULL, (char *)rtcsum_flds,
> + rtcsumblock_size, FTARG_SIZE, NULL, rtcsum_flds },
>
> { FLDT_ZZZ, NULL }
> };
> diff --git a/db/field.h b/db/field.h
> index 27df5a0fcd1b..77f84d55cb62 100644
> --- a/db/field.h
> +++ b/db/field.h
> @@ -211,6 +211,7 @@ typedef enum fldt {
> FLDT_RGBITMAP,
> FLDT_SUMINFO,
> FLDT_RGSUMMARY,
> + FLDT_RTCSUM,
>
> FLDT_ZZZ /* mark last entry */
> } fldt_t;
> diff --git a/db/inode.c b/db/inode.c
> index 46f83f18f083..a77595427c56 100644
> --- a/db/inode.c
> +++ b/db/inode.c
> @@ -727,6 +727,8 @@ inode_next_type(void)
> return TYP_RTRMAPBT;
> case XFS_METAFILE_RTREFCOUNT:
> return TYP_RTREFCBT;
> + case XFS_METAFILE_RTCSUM:
> + return TYP_RTCSUM;
> default:
> return TYP_DATA;
> }
> diff --git a/db/rtgroup.c b/db/rtgroup.c
> index c6b96c9dc79d..666adb2feb33 100644
> --- a/db/rtgroup.c
> +++ b/db/rtgroup.c
> @@ -16,6 +16,8 @@
> #include "output.h"
> #include "init.h"
> #include "rtgroup.h"
> +#include "libfrog/crc64.h"
> +#include "xfs_rtcsumfile.h"
>
> #define uuid_equal(s,d) (platform_uuid_compare((s),(d)) == 0)
>
> @@ -152,3 +154,43 @@ const field_t rgsummary_hfld[] = {
> { "", FLDT_RGSUMMARY, OI(0), C1, 0, TYP_NONE },
> { NULL }
> };
> +
> +/*
> + * Get the size of a rtcsum block.
> + */
> +int
> +rtcsumblock_size(
> + void *obj,
> + int startoff,
> + int idx)
> +{
> + return bitize(1u << mp->m_sb.sb_rtcsum_blklog);
> +}
> +
> +static int
> +rtcsum_count(
> + void *obj,
> + int startoff)
> +{
> + return xfs_rtcsum_payload_size(mp) >> mp->m_rtcsum_shift;
> +}
> +
> +#define OFF(f) bitize(offsetof(struct xfs_rtbuf_blkinfo, rt_ ## f))
> +const field_t rtcsum_flds[] = {
> + { "magicnum", FLDT_UINT32X, OI(OFF(magic)), C1, 0, TYP_NONE },
> + { "crc", FLDT_CRC, OI(OFF(crc)), C1, 0, TYP_NONE },
> + { "owner", FLDT_INO, OI(OFF(owner)), C1, 0, TYP_NONE },
> + { "bno", FLDT_DFSBNO, OI(OFF(blkno)), C1, 0, TYP_BMAPBTD },
> + { "lsn", FLDT_UINT64X, OI(OFF(lsn)), C1, 0, TYP_NONE },
> + { "uuid", FLDT_UUID, OI(OFF(uuid)), C1, 0, TYP_NONE },
> + /* the checksums are after the blkinfo structure */
> + { "csums", FLDT_SUMINFO, OI(bitize(sizeof(struct xfs_rtbuf_blkinfo))),
FLDT_SUMINFO is 4 bytes, this won't work for crc64.
You might want to add a FLDT_CRC32C and FLDT_CRC64 so that the sizes can
be correct, db can convert them from big-endian to host for display, and
the display can be '0x%x' instead of decimal.
--D
> + rtcsum_count, FLD_ARRAY | FLD_COUNT, TYP_DATA },
> + { NULL }
> +};
> +#undef OFF
> +
> +const field_t rtcsum_hfld[] = {
> + { "", FLDT_RTCSUM, OI(0), C1, 0, TYP_NONE },
> + { NULL }
> +};
> diff --git a/db/rtgroup.h b/db/rtgroup.h
> index 5b120f2c9a29..b674154fcbad 100644
> --- a/db/rtgroup.h
> +++ b/db/rtgroup.h
> @@ -15,6 +15,11 @@ extern const struct field rgbitmap_hfld[];
> extern const struct field rgsummary_flds[];
> extern const struct field rgsummary_hfld[];
>
> +extern const struct field rtcsum_flds[];
> +extern const struct field rtcsum_hfld[];
> +
> +int rtcsumblock_size(void *obj, int startoff, int idx);
> +
> extern void rtsb_init(void);
> extern int rtsb_size(void *obj, int startoff, int idx);
>
> diff --git a/db/sb.c b/db/sb.c
> index 666207145a86..c565e6535f63 100644
> --- a/db/sb.c
> +++ b/db/sb.c
> @@ -144,8 +144,14 @@ const field_t sb_flds[] = {
> FLD_COUNT, TYP_NONE },
> { "pad", FLDT_UINT8X, OI(OFF(pad)), metadirfld_count,
> FLD_COUNT, TYP_NONE },
> - { "rtstart", FLDT_DRFSBNO, OI(OFF(rtstart)), zonedfld_count, FLD_COUNT, TYP_NONE },
> - { "rtreserved", FLDT_UINT64D, OI(OFF(rtreserved)), zonedfld_count, FLD_COUNT, TYP_NONE },
> + { "rtcsum_type", FLDT_UINT8D, OI(OFF(rtcsum_type)), metadirfld_count,
> + FLD_COUNT, TYP_NONE },
> + { "rtcsum_blklog", FLDT_UINT8D, OI(OFF(rtcsum_blklog)), metadirfld_count,
> + FLD_COUNT, TYP_NONE },
> + { "rtstart", FLDT_DRFSBNO, OI(OFF(rtstart)), zonedfld_count,
> + FLD_COUNT, TYP_NONE },
> + { "rtreserved", FLDT_UINT64D, OI(OFF(rtreserved)), zonedfld_count,
> + FLD_COUNT, TYP_NONE },
> { NULL }
> };
>
> diff --git a/db/type.c b/db/type.c
> index 324f416a49cc..2db5b5f6a54d 100644
> --- a/db/type.c
> +++ b/db/type.c
> @@ -71,6 +71,7 @@ static const typ_t __typtab[] = {
> TYP_F_NO_CRC_OFF },
> { TYP_RGBITMAP, NULL },
> { TYP_RGSUMMARY, NULL },
> + { TYP_RTCSUM, NULL },
> { TYP_NONE, NULL }
> };
>
> @@ -125,6 +126,8 @@ static const typ_t __typtab_crc[] = {
> &xfs_rtbitmap_buf_ops, XFS_RTBUF_CRC_OFF },
> { TYP_RGSUMMARY, "rgsummary", handle_struct, rgsummary_hfld,
> &xfs_rtsummary_buf_ops, XFS_RTBUF_CRC_OFF },
> + { TYP_RTCSUM, "rtcsum", handle_struct, rtcsum_hfld,
> + &xfs_rtcsum_buf_ops, XFS_RTBUF_CRC_OFF },
> { TYP_NONE, NULL }
> };
>
> @@ -179,6 +182,8 @@ static const typ_t __typtab_spcrc[] = {
> &xfs_rtbitmap_buf_ops, XFS_RTBUF_CRC_OFF },
> { TYP_RGSUMMARY, "rgsummary", handle_struct, rgsummary_hfld,
> &xfs_rtsummary_buf_ops, XFS_RTBUF_CRC_OFF },
> + { TYP_RTCSUM, "rtcsum", handle_struct, rtcsum_hfld,
> + &xfs_rtcsum_buf_ops, XFS_RTBUF_CRC_OFF },
> { TYP_NONE, NULL }
> };
>
> diff --git a/db/type.h b/db/type.h
> index a2488a663dbd..d95c83c78d17 100644
> --- a/db/type.h
> +++ b/db/type.h
> @@ -39,6 +39,7 @@ typedef enum typnm
> TYP_FINOBT,
> TYP_RGBITMAP,
> TYP_RGSUMMARY,
> + TYP_RTCSUM,
> TYP_NONE
> } typnm_t;
>
> diff --git a/man/man8/xfs_db.8 b/man/man8/xfs_db.8
> index c2d4065944be..ab89d709b89e 100644
> --- a/man/man8/xfs_db.8
> +++ b/man/man8/xfs_db.8
> @@ -1356,7 +1356,8 @@ The possible data types are:
> .BR agf ", " agfl ", " agi ", " attr ", " bmapbta ", " bmapbtd ,
> .BR bnobt ", " cntbt ", " data ", " dir ", " dir2 ", " dqblk ,
> .BR inobt ", " inode ", " log ", " refcntbt ", " rmapbt ", " rtbitmap ,
> -.BR rtsummary ", " sb ", " symlink ", " rtrmapbt ", " rtrefcbt ", and " text .
> +.BR rtsummary ", " sb ", " symlink ", " rtrmapbt ", " rtrefcbt ,
> +.BR rtcsum ", and " text .
> See the TYPES section below for more information on these data types.
> .TP
> .BI "timelimit [" OPTIONS ]
> @@ -2644,6 +2645,16 @@ The first dimension is the size range,
> the second dimension is the starting bitmap block number
> (adjacent entries are for the same size, adjacent bitmap blocks).
> .TP
> +.B rtcsum
> +For RT data-checksum enabled file systems, there is one data checksum file
> +for each realtime group. The data checksum file contains the standard
> +RT block self-describing metadata header, and an array of the checksums
> +for the realtime group. The size of each entry depends on the checksum
> +used: 4 bytes for
> +.BR crc32c
> +and 8 bytes for
> +.BR crc64 .
> +.TP
> .B sb
> There is one sb (superblock) structure per allocation group.
> It is the first disk block in the allocation group.
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 32/32] xfs_scrub: don't merge over unused space when data checksums are enabled
2026-09-24 10:04 ` [PATCH 32/32] xfs_scrub: don't merge over unused space when data checksums are enabled Christoph Hellwig
@ 2026-09-25 23:51 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:51 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:22PM +0200, Christoph Hellwig wrote:
> Unused space could be lost space at mount time for which no metadata was
> recorded, and we thus might not have valid checksums for it.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
Pretty straightforward!
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
--D
> scrub/read_verify.c | 33 +++++++++++++++++++++++----------
> 1 file changed, 23 insertions(+), 10 deletions(-)
>
> diff --git a/scrub/read_verify.c b/scrub/read_verify.c
> index 311ea1c29008..0c3bec104d16 100644
> --- a/scrub/read_verify.c
> +++ b/scrub/read_verify.c
> @@ -542,20 +542,33 @@ try_read_verify_schedule_io(
>
> /*
> * If we have a stashed IO, we haven't changed pools, the error
> - * reporting is the same, and the two extents are close,
> - * we can combine them.
> + * reporting is the same, and the two extents are close, we can combine
> + * them.
> + *
> + * If the file system has data checksums enabled, only merge fully
> + * contigous ranges, as unused data might have invalid checksums.
> */
> - if (rs->rvp == rvp && rs->io_length > 0 &&
> - ((start >= rs->io_start && start <= rv_end + locality) ||
> - (rs->io_start >= start &&
> - rs->io_start <= req_end + locality))) {
> - rs->io_start = min(rs->io_start, start);
> - rs->io_length = max(req_end, rv_end) - rs->io_start;
> -
> - return true;
> + if (rs->rvp != rvp || !rs->io_length)
> + return false;
> +
> + if (rvp->ctx->mnt.fsgeom.flags & XFS_FSOP_GEOM_FLAGS_DATA_CSUM) {
> + if (start == rv_end)
> + goto merge;
> + if (rs->io_start == req_end)
> + goto merge;
> + } else {
> + if (start >= rs->io_start && start <= rv_end + locality)
> + goto merge;
> + if (rs->io_start >= start && rs->io_start <= req_end + locality)
> + goto merge;
> }
>
> return false;
> +
> +merge:
> + rs->io_start = min(rs->io_start, start);
> + rs->io_length = max(req_end, rv_end) - rs->io_start;
> + return true;
> }
>
> /* Did read verification succeed? */
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 31/32] repair: support RT data checksums
2026-09-24 10:04 ` [PATCH 31/32] repair: support RT data checksums Christoph Hellwig
@ 2026-09-25 23:54 ` Darrick J. Wong
2026-09-26 6:23 ` Christoph Hellwig
0 siblings, 1 reply; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-25 23:54 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 12:04:21PM +0200, Christoph Hellwig wrote:
> Recognize the rtcsum per-RTG metadir files, and sanity check a few
> well known attributes for them.
>
> If a csum file is missing, or had to be nuke, regenerate it by
or had to be nuked,
> re-calculating the checksums. While this does lose the protection
> of the checksums, this is probably still better than an unmountable
> file system.
>
> Note that due to the lack of the regular buffer readahead path,
> reading the data and csum buffers for regenerating the csum file
> is fully synchronous and thus slow. Hopefully we can come up with
> a better version for online repair and never have to use this for
> real.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
> repair/Makefile | 1 +
> repair/dinode.c | 46 ++++++++++++++
> repair/phase6.c | 122 +++++++++++++++++++++++++++++++++++-
> repair/rt.c | 36 +++++++++++
> repair/rt.h | 10 +++
> repair/rtcsum.c | 161 ++++++++++++++++++++++++++++++++++++++++++++++++
> 6 files changed, 375 insertions(+), 1 deletion(-)
> create mode 100644 repair/rtcsum.c
>
> diff --git a/repair/Makefile b/repair/Makefile
> index fb0b2f96cc91..27eb68e2b5d5 100644
> --- a/repair/Makefile
> +++ b/repair/Makefile
> @@ -73,6 +73,7 @@ CFILES = \
> rcbag.c \
> rmap.c \
> rt.c \
> + rtcsum.c \
> rtrefcount_repair.c \
> rtrmap_repair.c \
> sb.c \
> diff --git a/repair/dinode.c b/repair/dinode.c
> index 48939f8bd159..5c03c5f50689 100644
> --- a/repair/dinode.c
> +++ b/repair/dinode.c
> @@ -22,6 +22,8 @@
> #include "rmap.h"
> #include "bmap_repair.h"
> #include "rt.h"
> +#include "libfrog/crc64.h"
> +#include "xfs_rtcsumfile.h"
>
> /* inode types */
> enum xr_ino_type {
> @@ -41,6 +43,7 @@ enum xr_ino_type {
> XR_INO_PQUOTA, /* project quota inode */
> XR_INO_RTRMAP, /* realtime rmap */
> XR_INO_RTREFC, /* realtime refcount */
> + XR_INO_RTCSUM, /* realtime data checksum */
> XR_INO_MAX
> };
>
> @@ -61,6 +64,7 @@ static const char *xr_ino_type_name[] = {
> [XR_INO_PQUOTA] = N_("project quota"),
> [XR_INO_RTRMAP] = N_("realtime rmap"),
> [XR_INO_RTREFC] = N_("realtime refcount"),
> + [XR_INO_RTCSUM] = N_("realtime data checksum"),
> };
> static_assert(ARRAY_SIZE(xr_ino_type_name) == XR_INO_MAX);
>
> @@ -2064,6 +2068,43 @@ _("bad # of extents (%" PRIu64 ") for %s inode %" PRIu64 "\n"),
> return 0;
> }
>
> +static int
> +process_check_rtcsum_inode(
> + struct xfs_mount *mp,
> + struct xfs_dinode *dinoc,
> + xfs_ino_t lino,
> + enum xr_ino_type *type,
> + int *dirty)
> +{
> + xfs_fsize_t size = be64_to_cpu(dinoc->di_size);
> + xfs_fsize_t expected_size;
> + int error;
> +
> + error = process_check_rt_inode(mp, dinoc, lino, type, dirty,
> + XR_INO_RTCSUM, _("realtime data checksums"));
> + if (error)
> + return error;
> +
> + /* rtcsum inodes must be contiguous */
> + if (xfs_dfork_data_extents(dinoc) != 1) {
> + do_warn(
> +_("non-contiguous rtcsum inode %" PRIu64 "\n"), lino);
> + return 1;
> + }
> +
> + expected_size = (xfs_off_t)mp->m_sb.sb_rgextents << mp->m_rtcsum_shift;
> + expected_size = (expected_size + xfs_rtcsum_payload_size(mp) - 1) /
> + xfs_rtcsum_payload_size(mp) * mp->m_rtcsum_bsize;
round_up()?
> + if (size != expected_size) {
> + do_warn(
> +_("unexpected rtcsum file size (%" PRId64 ") for ino %" PRIu64 "\n"),
> + size, lino);
> + return 1;
> + }
> +
> + return 0;
> +}
> +
> /*
> * If inode is a superblock inode, does type check to make sure is it valid.
> * Returns 0 if it's valid, non-zero if it needs to be cleared.
> @@ -2130,6 +2171,8 @@ process_check_metadata_inodes(
> if (is_rtrefcount_inode(lino))
> return process_check_rt_inode(mp, dinoc, lino, type, dirty,
> XR_INO_RTREFC, _("realtime refcount btree"));
> + if (is_rtcsum_inode(lino))
> + return process_check_rtcsum_inode(mp, dinoc, lino, type, dirty);
> return 0;
> }
>
> @@ -2977,6 +3020,7 @@ process_dinode_metafile(
> case XR_INO_UQUOTA:
> case XR_INO_GQUOTA:
> case XR_INO_PQUOTA:
> + case XR_INO_RTCSUM:
> /*
> * Quota checking and repair doesn't happen until phase7, so
> * preserve quota inodes and their contents for later.
> @@ -3555,6 +3599,8 @@ _("bad (negative) size %" PRId64 " on inode %" PRIu64 "\n"),
> type = XR_INO_RTRMAP;
> else if (is_rtrefcount_inode(lino))
> type = XR_INO_RTREFC;
> + else if (is_rtcsum_inode(lino))
> + type = XR_INO_RTCSUM;
> else
> type = XR_INO_DATA;
> break;
> diff --git a/repair/phase6.c b/repair/phase6.c
> index f3951a3d0709..c7bc1cabd0d3 100644
> --- a/repair/phase6.c
> +++ b/repair/phase6.c
> @@ -6,7 +6,6 @@
>
> #include "libxfs.h"
> #include "threads.h"
> -#include "threads.h"
> #include "prefetch.h"
> #include "avl.h"
> #include "globals.h"
> @@ -23,6 +22,8 @@
> #include "repair/quotacheck.h"
> #include "repair/slab.h"
> #include "repair/rmap.h"
> +#include "libfrog/crc64.h"
> +#include "xfs_rtcsumfile.h"
>
> static xfs_ino_t orphanage_ino;
>
> @@ -703,6 +704,123 @@ ensure_rtgroup_refcountbt(
> populate_rtgroup_refcountbt(rtg, est_fdblocks);
> }
>
> +/*
> + * Link a metadata directory inode.
> + */
> +static int
> +metadir_link(
> + struct xfs_inode *dp,
> + struct xfs_inode *ip,
> + const char *path,
> + enum xfs_metafile_type type)
> +{
> + struct xfs_metadir_update upd = {
> + .dp = dp,
> + .metafile_type = type,
> + .ip = ip,
> + .path = path,
> + };
> + int error;
> +
> + error = xfs_metadir_start_link(&upd);
> + if (error)
> + return error;
> +
> + error = xfs_metadir_link(&upd);
> + if (error)
> + return error;
> +
> + xfs_trans_log_inode(upd.tp, upd.ip, XFS_ILOG_CORE);
> + return xfs_metadir_commit(&upd);
> +}
> +
> +static void
> +ensure_rtgroup_csum(
> + struct xfs_rtgroup *rtg)
> +{
> + struct xfs_mount *mp = rtg_mount(rtg);
> + xfs_rgnumber_t rgno = rtg_rgno(rtg);
> + struct xfs_inode *dp = mp->m_rtdirip;
> + struct xfs_inode *ip;
> + int error;
> +
> + if (no_modify) {
> + if (rtcsum_ino_is_bad(rgno))
> + do_warn(_("would reset RG %u csum inode\n"), rgno);;
> + return;
> + }
> +
> + if (!rtcsum_ino_is_bad(rgno)) {
> + /*
> + * The /realtime directory has been discarded, but we should be
> + * able to iget the inodes directly.
> + */
> + error = -libxfs_metafile_iget(mp, rtcsum_ino(rgno),
> + XFS_METAFILE_RTCSUM, &ip);
> + if (error) {
> + do_warn(
> +_("Could not open RG %u csum inode, error %d\n"), rgno, error);
> + rtcsum_ino_mark_bad(rgno);
> + }
> + }
> +
> + if (rtcsum_ino_is_bad(rgno)) {
> + do_warn(_(
> +"resetting RG %u csum inode, regenerating data checksums\n"),
Nit: Inconsistent _( placement vs. the other errors.
> + rgno);
> + error = -libxfs_rtginode_create(rtg, XFS_RTGI_CSUM, false);
> + if (error) {
> + do_warn(
> +_("Couldn't create RG %u csum inode, error %d\n"), rgno, error);
> + return;
> + }
> + ip = rtg->rtg_inodes[XFS_RTGI_CSUM];
> + error = -xfs_rtcsum_alloc_blocks(rtg);
Does this need the usual -libxfsification macros?
> + if (error) {
> + do_warn(
> +_("Initialization of RG %u csum inode failed, error %d"), rgno, error);
> + return;
> + }
> +
> + calculate_rtgroup_csums(rtg);
> + } else {
> + struct xfs_trans *tp;
> + const char *name;
> +
> + /* Erase parent pointers before we create the new link */
> + try_erase_parent_ptrs(ip);
> +
> + name = xfs_rtginode_path(rtg_rgno(rtg), XFS_RTGI_CSUM);
> + error = -metadir_link(dp, ip, name, XFS_METAFILE_RTCSUM);
> + kfree(name);
> + if (error) {
> + do_warn(
> +_("Couldn't link RG %u csum inode, error %d\n"), rgno, error);
> + return;
> + }
> +
> + /*
> + * Reset the link count to 1 because the link above bumped it.
> + */
> + error = -libxfs_trans_alloc_inode(ip, &M_RES(mp)->tr_ichange,
> + 0, 0, false, &tp);
> + if (!error) {
> + set_nlink(VFS_I(ip), 1);
> + libxfs_trans_log_inode(tp, ip, XFS_ILOG_CORE);
> + error = -libxfs_trans_commit(tp);
> + }
> + if (error)
> + do_error(
> +_("Couldn't reset link count on RG %i quota inode, error %d\n"),
> + rgno, error);
> + }
> +
> + /* Mark the inode in use. */
> + mark_ino_inuse(mp, I_INO(ip), S_IFREG, I_INO(dp));
> + mark_ino_metadata(mp, I_INO(ip));
> + libxfs_irele(ip);
> +}
> +
> /* Initialize a root directory. */
> static int
> init_fs_root_dir(
> @@ -3473,6 +3591,8 @@ _(" - resetting contents of realtime bitmap and summary inodes\n"));
> }
> ensure_rtgroup_rmapbt(rtg, est_fdblocks);
> ensure_rtgroup_refcountbt(rtg, est_fdblocks);
> + if (xfs_has_rtcsum(mp))
> + ensure_rtgroup_csum(rtg);
> }
> }
>
> diff --git a/repair/rt.c b/repair/rt.c
> index b0ff775bd339..4b64a91fbe29 100644
> --- a/repair/rt.c
> +++ b/repair/rt.c
> @@ -26,6 +26,8 @@ struct rtg_computed {
> };
> struct rtg_computed *rt_computed;
>
> +static xfs_ino_t *rtcsum_inos;
> +
> static inline void
> set_rtword(
> struct xfs_mount *mp,
> @@ -410,6 +412,27 @@ fill_rtsummary(
> _("couldn't re-initialize realtime summary inode, error %d\n"), error);
> }
>
> +bool
> +rtcsum_ino_is_bad(
> + xfs_rgnumber_t rgno)
> +{
> + return rtcsum_inos[rgno] == NULLFSINO;
> +}
> +
> +void
> +rtcsum_ino_mark_bad(
> + xfs_rgnumber_t rgno)
> +{
> + rtcsum_inos[rgno] = NULLFSINO;
> +}
> +
> +xfs_ino_t
> +rtcsum_ino(
> + xfs_rgnumber_t rgno)
> +{
> + return rtcsum_inos[rgno];
> +}
> +
> bool
> is_rtgroup_inode(
> xfs_ino_t ino,
> @@ -471,6 +494,13 @@ mark_rtginode(
> goto out_corrupt;
> }
>
> + /*
> + * Record the inode numbers of the data checksum inodes, as we don't
> + * just blow those away like other per-RTG metadata.
> + */
> + if (type == XFS_RTGI_CSUM)
> + rtcsum_inos[rtg_rgno(rtg)] = I_INO(ip);
> +
> /*
> * Phase 3 will clear the ondisk inodes of all rt metadata files, but
> * it doesn't reset any blocks. Keep the incore inodes loaded so that
> @@ -494,6 +524,12 @@ discover_rtgroup_inodes(
> int error, err2;
> int i;
>
> + rtcsum_inos = calloc(mp->m_sb.sb_rgcount, sizeof(xfs_ino_t));
> + if (!rtcsum_inos)
> + do_error(_("could not allocate csum ino array\n"));
> + for (i = 0; i < mp->m_sb.sb_rgcount; i++)
> + rtcsum_inos[i] = NULLFSINO;
> +
> tp = libxfs_trans_alloc_empty(mp);
> if (xfs_has_rtgroups(mp) && mp->m_sb.sb_rgcount > 0) {
> error = -libxfs_rtginode_load_parent(tp);
> diff --git a/repair/rt.h b/repair/rt.h
> index e4f3d5d9af31..9400b230de59 100644
> --- a/repair/rt.h
> +++ b/repair/rt.h
> @@ -13,6 +13,10 @@ void check_rtsummary(struct xfs_mount *mp);
> void fill_rtbitmap(struct xfs_rtgroup *rtg);
> void fill_rtsummary(struct xfs_rtgroup *rtg);
>
> +bool rtcsum_ino_is_bad(xfs_rgnumber_t rgno);
> +void rtcsum_ino_mark_bad(xfs_rgnumber_t rgno);
> +xfs_ino_t rtcsum_ino(xfs_rgnumber_t rgno);
> +
> void discover_rtgroup_inodes(struct xfs_mount *mp);
> void unload_rtgroup_inodes(struct xfs_mount *mp);
>
> @@ -37,6 +41,10 @@ static inline bool is_rtrefcount_inode(xfs_ino_t ino)
> {
> return is_rtgroup_inode(ino, XFS_RTGI_REFCOUNT);
> }
> +static inline bool is_rtcsum_inode(xfs_ino_t ino)
> +{
> + return is_rtgroup_inode(ino, XFS_RTGI_CSUM);
> +}
>
> void mark_rtgroup_inodes_bad(struct xfs_mount *mp, enum xfs_rtg_inodes type);
> bool rtgroup_inodes_were_bad(enum xfs_rtg_inodes type);
> @@ -44,4 +52,6 @@ bool rtgroup_inodes_were_bad(enum xfs_rtg_inodes type);
> void check_rtsb(struct xfs_mount *mp);
> void rewrite_rtsb(struct xfs_mount *mp);
>
> +void calculate_rtgroup_csums(struct xfs_rtgroup *rtg);
> +
> #endif /* _XFS_REPAIR_RT_H_ */
> diff --git a/repair/rtcsum.c b/repair/rtcsum.c
> new file mode 100644
> index 000000000000..c4d629c18e8d
> --- /dev/null
> +++ b/repair/rtcsum.c
> @@ -0,0 +1,161 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * Copyright (c) 2026 Christoph Hellwig.
> + *
> + * Rebuild data checksums when the csum metafile was lost. This doens't perform
> + * great and is inteded as a last resort.
"This doesn't perform great and is intended as..."
> + */
> +#include "libxfs.h"
> +#include "globals.h"
> +#include "protos.h"
> +#include "rt.h"
> +#include "err_protos.h"
> +#include "libfrog/crc64.h"
> +#include "xfs_platform.h"
> +#include "xfs_rtcsumfile.h"
> +
> +/* default to 1MiB data reads to make the performance only somewhat horrible. */
> +#define DATA_BLOCKS_PER_BUF 256
> +
> +static int
> +read_data(
> + struct xfs_buftarg *btp,
> + xfs_daddr_t blkno,
> + unsigned int nblks,
> + void *buf)
> +{
> + int fd = btp->bt_bdev_fd;
> + off_t pos = LIBXFS_BBTOOFF64(blkno);
> + size_t len = BBTOB(nblks);
> + ssize_t ret;
> +
> + ret = pread(fd, buf, len, pos);
> + if (ret < 0) {
> + ret = errno;
> +
> + fprintf(stderr, _("%s: read failed: %s\n"),
> + progname, strerror(ret));
> + return -ret;
> + }
> + if (ret != len) {
> + fprintf(stderr, _("%s: error - read only %zd of %zd bytes\n"),
> + progname, ret, len);
> + return -EIO;
> + }
> +
> + return 0;
> +}
> +
> +static int
> +calculate_buf_csums(
> + struct xfs_rtgroup *rtg,
> + void *data_buf,
> + xfs_rgblock_t rgbno,
> + unsigned int *nr)
> +{
> + struct xfs_mount *mp = rtg_mount(rtg);
> + unsigned int boff = xfs_rgb_to_rtcsumoff(mp, rgbno);
> + struct xfs_trans_res tres = M_RES(mp)->tr_csum;
> + xfs_daddr_t csum_daddr;
> + void *csum_buf;
> + struct xfs_buf *csum_bp;
> + struct xfs_trans *tp;
> + int error;
> + unsigned int i = 0;
> +
> + *nr = min(*nr, xfs_rtcsum_len_to_extlen(mp, mp->m_rtcsum_bsize - boff));
> +
> + error = xfs_rtcsum_bmap(rtg, rgbno, &csum_daddr);
> + if (error) {
> + do_error(
> +_("Cannot bmap new csum file for RG %u, error %d\n"),
> + rtg_rgno(rtg), error);
> + return error;
> + }
> +
> + tres.tr_logres = xfs_calc_csum_reservation(mp,
> + xfs_extlen_to_rtcsum_len(mp, *nr));
> + ASSERT(tres.tr_logres <= M_RES(mp)->tr_csum.tr_logres);
> +
> + error = libxfs_trans_alloc(mp, &tres, 0, 0, 0, &tp);
> + if (error) {
> + do_error(
> +_("Transaction allocation for csum repair failed: %d\n"),
> + error);
> + return error;
> + }
> +
> + error = libxfs_buf_read(mp->m_ddev_targp, csum_daddr,
> + BTOBB(mp->m_rtcsum_bsize), 0, &csum_bp,
> + &xfs_rtcsum_buf_ops);
> + if (error) {
> + do_error(
> +_("Cannot get buffer for new csum file for RG %u, error %d\n"),
> + rtg_rgno(rtg), error);
> + return error;
> + }
> +
> + libxfs_trans_bjoin(tp, csum_bp);
> + xfs_trans_buf_set_type(tp, csum_bp, XFS_BLFT_RTCSUM_BUF);
> +
> + csum_buf = csum_bp->b_addr + boff;
> + for (i = 0; i < *nr; i++) {
> + union xfs_csum csum;
> +
> + xfs_csum_seed(mp, &csum);
> + xfs_csum_gen(mp, data_buf + XFS_FSB_TO_B(mp, i),
> + mp->m_sb.sb_blocksize, &csum);
> + if (i == 0)
> + printf("calculated checksum 0x%x\n", csum.crc32c);
Debug code?
> + xfs_csum_finalize(mp, csum_buf, &csum);
> + csum_buf += (1u << mp->m_rtcsum_shift);
> + }
> +
> + libxfs_trans_log_buf(tp, csum_bp, boff,
> + boff + xfs_extlen_to_rtcsum_len(mp, *nr) - 1);
> + return -libxfs_trans_commit(tp);
> +}
> +
> +void
> +calculate_rtgroup_csums(
> + struct xfs_rtgroup *rtg)
> +{
> + struct xfs_mount *mp = rtg_mount(rtg);
> + xfs_rgblock_t rgbno;
> + int error;
> + void *data_buf;
> +
> + error = posix_memalign(&data_buf, sysconf(_SC_PAGESIZE),
> + XFS_FSB_TO_B(mp, DATA_BLOCKS_PER_BUF));
> + if (error) {
> + do_warn(
> +_("Failed to allocate memory for csum rebuild\n"));
> + return;
> + }
> + for (rgbno = 0; rgbno < rtg_blocks(rtg); rgbno += DATA_BLOCKS_PER_BUF) {
> + unsigned int nr_blocks, done = 0;
> +
> + nr_blocks = min(DATA_BLOCKS_PER_BUF, rtg_blocks(rtg) - rgbno);
> + error = read_data(mp->m_rtdev_targp,
> + xfs_gbno_to_daddr(rtg_group(rtg), rgbno),
> + XFS_FSB_TO_BB(mp, nr_blocks),
> + data_buf);
> + if (error) {
> + do_warn(
> +_("Reading data at RG %u/%u failed, error %d"),
> + rtg_rgno(rtg), rgbno, error);
> + continue;
Er... so a read failure means we just leave broken checksums?
--D
> + }
> +
> + do {
> + unsigned int n = nr_blocks - done;
> +
> + calculate_buf_csums(rtg,
> + data_buf + XFS_FSB_TO_B(mp, done),
> + rgbno + done, &n);
> + done += n;
> + } while (done < nr_blocks);
> + }
> +
> + kfree(data_buf);
> +}
> --
> 2.53.0
>
>
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2
2026-09-24 20:30 ` Darrick J. Wong
@ 2026-09-26 6:15 ` Christoph Hellwig
2026-09-29 10:14 ` Andrey Albershteyn
0 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:15 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Thu, Sep 24, 2026 at 01:30:29PM -0700, Darrick J. Wong wrote:
> On Thu, Sep 24, 2026 at 12:03:51PM +0200, Christoph Hellwig wrote:
> > Using tabs for indentation messes up man page rendering, so use spaces
> > for rtstart to match the other fields.
> >
> > Signed-off-by: Christoph Hellwig <hch@lst.de>
>
> This could go into 7.3.
Andrey, can you pick this up directly, or should I resend it standalone?
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 09/32] FIXUP
2026-09-25 23:29 ` Darrick J. Wong
@ 2026-09-26 6:17 ` Christoph Hellwig
0 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:17 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Fri, Sep 25, 2026 at 04:29:44PM -0700, Darrick J. Wong wrote:
> > +static inline void
> > +xfs_buf_set_uptodate(
> > + struct xfs_buf *bp)
> > +{
> > +}
>
> Should this set LIBXFS_B_UPTODATE?
Old xfsprogs didn't map B_DONE to it either. But maybe we should do
a separate series to align buffer flags and semantics between the
kernel and userspace, including using uptodate instead of done in
the kernel. The repair slob series had some advances in this direction,
but unfortunately is was never split our or followed up upon.
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 26/32] man: document the rtcsum geom fields
2026-09-25 23:31 ` Darrick J. Wong
@ 2026-09-26 6:17 ` Christoph Hellwig
0 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:17 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Fri, Sep 25, 2026 at 04:31:56PM -0700, Darrick J. Wong wrote:
> > +.TP
> > +.B XFS_CSUM_TYPE_CRC32C
> > +CRC32c as per NVMe.
>
> Just for my education -- crc32c as per nvme is the same as crc32c
> everywhere else right?
Where everyone else is xfs meta csums, btrfs, nvme and the recommendation
in the kernel implementation of crc32c - yes.
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs
2026-09-25 23:33 ` Darrick J. Wong
@ 2026-09-26 6:18 ` Christoph Hellwig
0 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:18 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Fri, Sep 25, 2026 at 04:33:51PM -0700, Darrick J. Wong wrote:
> On Thu, Sep 24, 2026 at 12:04:18PM +0200, Christoph Hellwig wrote:
> > Signed-off-by: Christoph Hellwig <hch@lst.de>
>
> Seems fine, though I wonder if we should report 1U<<rtcsum_blklog?
I was trying to save space, but there's arguments either way...
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 29/32] mkfs: support RT data checksums
2026-09-25 23:39 ` Darrick J. Wong
@ 2026-09-26 6:19 ` Christoph Hellwig
0 siblings, 0 replies; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:19 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Fri, Sep 25, 2026 at 04:39:08PM -0700, Darrick J. Wong wrote:
> > R_RESERVED,
> > + R_CSUM,
> > + R_CSUMBSIZE,
>
> BTW now that Andrey has merged the config file generator code, you'll
> have to add the relevant R_CSUM/R_CSUMBSIZE bits to libfrog/fsgeom.c.
Oh, right. Another fun learning experience :)
> > /* data subvol */ [-d agcount=n,agsize=n,file,name=xxx,size=num,\n\
> > (sunit=value,swidth=value|su=num,sw=num|noalign),\n\
> > - sectsize=num,concurrency=num]\n\
> > + sectsize=num,concurrency=num,\n\
> > + csum=type]\n\
>
> There's no csum= parameter for -d, or at least I didn't see a D_CSUM
> entry above.
Correct. This is left from an earlier series where I tried to support
cheksums on the data device using NVMe non-PI metadata, but that didn't
got to well.
> > + cfg->sb_feat.rtcsum_blklog = 0;
> > + for (l = cfg->rtcsumbsize; l > 1; l >>= 1)
> > + cfg->sb_feat.rtcsum_blklog++;
>
> log2_roundup?
Ah, nice. I was looking for a userspace version of the kernels ilog2,
but could not find it.
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 30/32] xfs_db: support RT data checksum
2026-09-25 23:43 ` Darrick J. Wong
@ 2026-09-26 6:21 ` Christoph Hellwig
2026-09-26 18:20 ` Darrick J. Wong
0 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:21 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Fri, Sep 25, 2026 at 04:43:58PM -0700, Darrick J. Wong wrote:
> > + { "csums", FLDT_SUMINFO, OI(bitize(sizeof(struct xfs_rtbuf_blkinfo))),
>
> FLDT_SUMINFO is 4 bytes, this won't work for crc64.
Ouch, yes.
> You might want to add a FLDT_CRC32C and FLDT_CRC64 so that the sizes can
> be correct, db can convert them from big-endian to host for display, and
> the display can be '0x%x' instead of decimal.
The values are little endian, just like the xfs metadata csum.
Do you remember how to dynamically switch between different
interpretations of a field in xfs_db code?
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 31/32] repair: support RT data checksums
2026-09-25 23:54 ` Darrick J. Wong
@ 2026-09-26 6:23 ` Christoph Hellwig
2026-09-27 0:01 ` Darrick J. Wong
0 siblings, 1 reply; 57+ messages in thread
From: Christoph Hellwig @ 2026-09-26 6:23 UTC (permalink / raw)
To: Darrick J. Wong; +Cc: Christoph Hellwig, Andrey Albershteyn, linux-xfs
On Fri, Sep 25, 2026 at 04:54:35PM -0700, Darrick J. Wong wrote:
> > + }
> > + ip = rtg->rtg_inodes[XFS_RTGI_CSUM];
> > + error = -xfs_rtcsum_alloc_blocks(rtg);
>
> Does this need the usual -libxfsification macros?
Probably. Or I should go for another attempt to kill them everywhere...
> > + xfs_csum_seed(mp, &csum);
> > + xfs_csum_gen(mp, data_buf + XFS_FSB_TO_B(mp, i),
> > + mp->m_sb.sb_blocksize, &csum);
> > + if (i == 0)
> > + printf("calculated checksum 0x%x\n", csum.crc32c);
>
> Debug code?
Yes.
> > +
> > + nr_blocks = min(DATA_BLOCKS_PER_BUF, rtg_blocks(rtg) - rgbno);
> > + error = read_data(mp->m_rtdev_targp,
> > + xfs_gbno_to_daddr(rtg_group(rtg), rgbno),
> > + XFS_FSB_TO_BB(mp, nr_blocks),
> > + data_buf);
> > + if (error) {
> > + do_warn(
> > +_("Reading data at RG %u/%u failed, error %d"),
> > + rtg_rgno(rtg), rgbno, error);
> > + continue;
>
> Er... so a read failure means we just leave broken checksums?
A zero-initialized one to be exact, yes. I don't think there is
much more we can do here?
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 30/32] xfs_db: support RT data checksum
2026-09-26 6:21 ` Christoph Hellwig
@ 2026-09-26 18:20 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-26 18:20 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Sat, Sep 26, 2026 at 08:21:15AM +0200, Christoph Hellwig wrote:
> On Fri, Sep 25, 2026 at 04:43:58PM -0700, Darrick J. Wong wrote:
> > > + { "csums", FLDT_SUMINFO, OI(bitize(sizeof(struct xfs_rtbuf_blkinfo))),
> >
> > FLDT_SUMINFO is 4 bytes, this won't work for crc64.
>
> Ouch, yes.
>
> > You might want to add a FLDT_CRC32C and FLDT_CRC64 so that the sizes can
> > be correct, db can convert them from big-endian to host for display, and
> > the display can be '0x%x' instead of decimal.
>
> The values are little endian, just like the xfs metadata csum.
Oh, right, I forgot that :)
> Do you remember how to dynamically switch between different
> interpretations of a field in xfs_db code?
The stupid way is to declare both FLDT types and counting functions
that return nonzero if that checksum type is enabled, and 0 otherwise:
static int
rtcrc32c_count(
void *obj,
int startoff)
{
if (mp->m_sb.sb_rtcsum_type != XFS_CSUM_TYPE_CRC32C)
return 0;
return xfs_rtcsum_payload_size(mp) >> mp->m_rtcsum_shift;
}
static int
rtcrc64_count(
void *obj,
int startoff)
{
if (mp->m_sb.sb_rtcsum_type != XFS_CSUM_TYPE_CRC64)
return 0;
return xfs_rtcsum_payload_size(mp) >> mp->m_rtcsum_shift;
}
{ "crc", FLDT_CRC, OI(bitize(sizeof(struct xfs_rtbuf_blkinfo))),
rtcrc32c_count, FLD_ARRAY | FLD_COUNT, TYP_DATA },
{ "crc64", FLDT_CRC64, OI(bitize(sizeof(struct xfs_rtbuf_blkinfo))),
rtcrc64_count, FLD_ARRAY | FLD_COUNT, TYP_DATA },
FLDT_CRC already exists, so I guess you just have to copy-pasta
FLDT_CRC64 into field.c:
{ FLDT_CRC64, "crc64", fp_crc, "%#x (%s)", SI(bitsz(uint64_t)),
0, NULL, NULL },
Though the iocur_crc_valid() call in fp_crc might cause problems here.
--D
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 31/32] repair: support RT data checksums
2026-09-26 6:23 ` Christoph Hellwig
@ 2026-09-27 0:01 ` Darrick J. Wong
0 siblings, 0 replies; 57+ messages in thread
From: Darrick J. Wong @ 2026-09-27 0:01 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Andrey Albershteyn, linux-xfs
On Sat, Sep 26, 2026 at 08:23:13AM +0200, Christoph Hellwig wrote:
> On Fri, Sep 25, 2026 at 04:54:35PM -0700, Darrick J. Wong wrote:
> > > + }
> > > + ip = rtg->rtg_inodes[XFS_RTGI_CSUM];
> > > + error = -xfs_rtcsum_alloc_blocks(rtg);
> >
> > Does this need the usual -libxfsification macros?
>
> Probably. Or I should go for another attempt to kill them everywhere...
>
> > > + xfs_csum_seed(mp, &csum);
> > > + xfs_csum_gen(mp, data_buf + XFS_FSB_TO_B(mp, i),
> > > + mp->m_sb.sb_blocksize, &csum);
> > > + if (i == 0)
> > > + printf("calculated checksum 0x%x\n", csum.crc32c);
> >
> > Debug code?
>
> Yes.
>
> > > +
> > > + nr_blocks = min(DATA_BLOCKS_PER_BUF, rtg_blocks(rtg) - rgbno);
> > > + error = read_data(mp->m_rtdev_targp,
> > > + xfs_gbno_to_daddr(rtg_group(rtg), rgbno),
> > > + XFS_FSB_TO_BB(mp, nr_blocks),
> > > + data_buf);
> > > + if (error) {
> > > + do_warn(
> > > +_("Reading data at RG %u/%u failed, error %d"),
> > > + rtg_rgno(rtg), rgbno, error);
> > > + continue;
> >
> > Er... so a read failure means we just leave broken checksums?
>
> A zero-initialized one to be exact, yes. I don't think there is
> much more we can do here?
I can't see anything that we can do about it either, but xfs_repair
ought to warn the user that they're going to get read errors.
--D
^ permalink raw reply [flat|nested] 57+ messages in thread
* Re: [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2
2026-09-26 6:15 ` Christoph Hellwig
@ 2026-09-29 10:14 ` Andrey Albershteyn
0 siblings, 0 replies; 57+ messages in thread
From: Andrey Albershteyn @ 2026-09-29 10:14 UTC (permalink / raw)
To: Christoph Hellwig; +Cc: Darrick J. Wong, linux-xfs
On 2026-09-26 08:15:30, Christoph Hellwig wrote:
> On Thu, Sep 24, 2026 at 01:30:29PM -0700, Darrick J. Wong wrote:
> > On Thu, Sep 24, 2026 at 12:03:51PM +0200, Christoph Hellwig wrote:
> > > Using tabs for indentation messes up man page rendering, so use spaces
> > > for rtstart to match the other fields.
> > >
> > > Signed-off-by: Christoph Hellwig <hch@lst.de>
> >
> > This could go into 7.3.
>
> Andrey, can you pick this up directly, or should I resend it standalone?
>
>
I've picked this up, it will be in next for-next
--
- Andrey
^ permalink raw reply [flat|nested] 57+ messages in thread
end of thread, other threads:[~2026-09-29 10:14 UTC | newest]
Thread overview: 57+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 10:03 xfsprogs support for RT data checksums Christoph Hellwig
2026-09-24 10:03 ` [PATCH 01/32] man: fix alignment of the rtstart field in ioctl_xfs_fsgeometry.2 Christoph Hellwig
2026-09-24 20:30 ` Darrick J. Wong
2026-09-26 6:15 ` Christoph Hellwig
2026-09-29 10:14 ` Andrey Albershteyn
2026-09-24 10:03 ` [PATCH 02/32] add cpu_to_le64 and le64_to_cpu_helpers Christoph Hellwig
2026-09-24 20:30 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 03/32] libfrog: add a crc64_nvme implementation Christoph Hellwig
2026-09-24 10:03 ` [PATCH 04/32] libxfs: add DIV_ROUND_UP_ULL Christoph Hellwig
2026-09-24 20:44 ` Darrick J. Wong
2026-09-25 6:13 ` Christoph Hellwig
2026-09-24 10:03 ` [PATCH 05/32] libxfs: remove a spurious xfs_rtbitmap.h include in xfs_sb.c Christoph Hellwig
2026-09-25 23:27 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 06/32] libxfs: add SZ_* constants Christoph Hellwig
2026-09-25 23:27 ` Darrick J. Wong
2026-09-24 10:03 ` [PATCH 07/32] xfs: remove spurious XBF_DONE clearing on readahead validation failure Christoph Hellwig
2026-09-24 10:03 ` [PATCH 08/32] xfs: hide b_flags manipulation from code outside of xfs_buf.c Christoph Hellwig
2026-09-24 10:03 ` [PATCH 09/32] FIXUP Christoph Hellwig
2026-09-25 23:29 ` Darrick J. Wong
2026-09-26 6:17 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 10/32] xfs: add error injection for lazy bounce buffering Christoph Hellwig
2026-09-24 10:04 ` [PATCH 11/32] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
2026-09-24 10:04 ` [PATCH 12/32] FIXUP Christoph Hellwig
2026-09-24 10:04 ` [PATCH 13/32] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
2026-09-24 10:04 ` [PATCH 14/32] xfs: centralize setting of buf_ops/buf_type/magic for rtblocks Christoph Hellwig
2026-09-24 10:04 ` [PATCH 15/32] xfs: factor out a xfs_rtfile_initialize_buf helper Christoph Hellwig
2026-09-24 10:04 ` [PATCH 16/32] xfs: add a xfs_rtblock_payload helper Christoph Hellwig
2026-09-24 10:04 ` [PATCH 17/32] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
2026-09-24 10:04 ` [PATCH 18/32] FIXUP Christoph Hellwig
2026-09-24 10:04 ` [PATCH 19/32] xfs: define the RT data checksum on-disk format Christoph Hellwig
2026-09-24 10:04 ` [PATCH 20/32] FIXUP Christoph Hellwig
2026-09-24 10:04 ` [PATCH 21/32] xfs: add support for per-RTG csum files Christoph Hellwig
2026-09-24 10:04 ` [PATCH 22/32] FIXUP Christoph Hellwig
2026-09-24 10:04 ` [PATCH 23/32] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
2026-09-24 10:04 ` [PATCH 24/32] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
2026-09-24 10:04 ` [PATCH 25/32] xfs: enable RT data checksums Christoph Hellwig
2026-09-24 10:04 ` [PATCH 26/32] man: document the rtcsum geom fields Christoph Hellwig
2026-09-25 23:31 ` Darrick J. Wong
2026-09-26 6:17 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 27/32] libfrog: print csum geometry information Christoph Hellwig
2026-09-25 23:32 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 28/32] xfs_io: report checksum information from fs geometry in statfs Christoph Hellwig
2026-09-25 23:33 ` Darrick J. Wong
2026-09-26 6:18 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 29/32] mkfs: support RT data checksums Christoph Hellwig
2026-09-25 23:39 ` Darrick J. Wong
2026-09-26 6:19 ` Christoph Hellwig
2026-09-24 10:04 ` [PATCH 30/32] xfs_db: support RT data checksum Christoph Hellwig
2026-09-25 23:43 ` Darrick J. Wong
2026-09-26 6:21 ` Christoph Hellwig
2026-09-26 18:20 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 31/32] repair: support RT data checksums Christoph Hellwig
2026-09-25 23:54 ` Darrick J. Wong
2026-09-26 6:23 ` Christoph Hellwig
2026-09-27 0:01 ` Darrick J. Wong
2026-09-24 10:04 ` [PATCH 32/32] xfs_scrub: don't merge over unused space when data checksums are enabled Christoph Hellwig
2026-09-25 23:51 ` Darrick J. Wong
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).