From: Dave Chinner <dgc@kernel.org>
To: Christoph Hellwig <hch@lst.de>
Cc: Carlos Maiolino <cem@kernel.org>,
"Darrick J . Wong" <djwong@kernel.org>,
Jens Axboe <axboe@kernel.dk>,
Christian Brauner <brauner@kernel.org>,
linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: support for RT data checksums
Date: Fri, 25 Sep 2026 08:52:58 +1000 [thread overview]
Message-ID: <arWpyhFyV3ouYAEx@dread> (raw)
In-Reply-To: <20260924100032.2733101-1-hch@lst.de>
On Thu, Sep 24, 2026 at 11:59:32AM +0200, Christoph Hellwig wrote:
> Hi all,
>
> data checksums provide an additional safeguard against silent data loss.
>
> In classic XFS they were hard to support because they need to be
> atomically updated with the written data. The zoned allocator solves
> that problem because it always writes out of place, and the checksums
> can be committed at the same as the metadata linking the newly written
> file data into place. In theory, a conventional allocator could be used
> in combination with the always_cow option, but there are few upsides of
> this compared to using the zoned allocator.
>
> Data checksums are stored in per-realtime group files in the metadir,
> similar to other modern RT metadata. Unlike the checksum design in btrfs
> or some other file system, the checksums are associated with the
> physical blocks, and not with logical data in files. This reduces the
> mapping overhead, and significantly reduces the write amplification,
> and also avoids duplicate checksums for reflinked files (although those
> are not yet supported with the zoned allocator anyway).
>
> The initial version provides two checksums algorithms: crc32c and crc64.
> Both of those are cyclic redundancy check algorithms which provide known
> good detection of bit flips that is better than general purpose hash
> functions. Both are not cryptographic hashes and thus do not provide any
> kind of protection against intentional tampering with the data.
> The crc32c parameters exactly match those use for xfs metadata checksums,
> and also those used by the default btrfs checksum, and the NVMe PI
> formats using crc32c. The crc64 parameters exactly match those using
> the NVMe PI formats using crc64. crc32c provides reasonable assurance
> for today's hardware, but might prove limiting for extremely large data
> sets, crc64 fills that void, but probably warrants using > 4k file system
> block sizes to amortize the overhead.
Ok, so this really needs a design doc to explain how it all works,
what the new on-disk format is, scope, constraints, etc, as the
first patch in the series (i.e. in
Documentation/filesystems/xfs/data_checksum_design.rst) so that we
have high level descriptions of the functionality being implemented.
Stuff like why certain crc alrgorithms are supported, how we can add
new ones in the future, constraints of doing so, how different sized
checksums are cleanly supported, etc will make doing such things
much easier.
There's new buffer and inode locking in transactions, and there's a
whole new buffer cache interface to "read a buffer", and that is
used to open code reading checksum buffers and joining them to a
transaction rather than using the existing xfs_trans_read_buf...()
interfaces. That in itself needs careful consideration, and clear
justification for why it must be duplicated to stand outside all the
existing BLI/transaction APIs, especially given all the "use the new
async buf read interface to do sync buffer reads" behaviour across
the patchset that could just use the existing interfaces.
I'd also like to have the format of the new on disk log item format
structures clearly documented (because we're going to have to
validate them) at recovery time, and also have a clear explaination
of the data vs metadata ordering algorithms that ensures that
checksums are always valid in crash+recovery situations, especially
w.r.t. data integrity operations like fsync.
These are the sorts of details that we need to get right, and it's
really hard to extract the actual design intent from the code that
implements it to determine if the algorithms, ordering and recovery
strategies are solid.
Hence I'd like to see the actual design documented first, then we
can understand and review algorithms, etc, and then verify the code
matches the described algorithms, behaviours, etc. checksums are all
about data integrity, so I'd really like to be able to understand
how it is supposed to work and where it doesn't work. Being forced
to understand design and implementation descisions and constraints
by reverse engineering disjoint chunks of code is not an efficient
use of reviewer time...
Cheers,
-Dave.
--
Dave Chinner
dgc@kernel.org
next prev parent reply other threads:[~2026-09-24 22:53 UTC|newest]
Thread overview: 69+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 9:59 support for RT data checksums Christoph Hellwig
2026-09-24 9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
2026-09-24 20:29 ` Darrick J. Wong
2026-09-24 9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
2026-09-24 21:39 ` Darrick J. Wong
2026-09-25 5:53 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
2026-09-24 21:43 ` Darrick J. Wong
2026-09-25 5:54 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
2026-09-24 21:44 ` Darrick J. Wong
2026-09-24 9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
2026-09-24 21:49 ` Darrick J. Wong
2026-09-25 5:57 ` Christoph Hellwig
2026-10-08 11:46 ` Anuj gupta
2026-09-24 9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
2026-09-24 22:03 ` Darrick J. Wong
2026-09-25 5:58 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
2026-09-24 9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
2026-09-24 22:13 ` Darrick J. Wong
2026-09-25 0:04 ` Eric Biggers
2026-09-25 6:01 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
2026-09-24 22:24 ` Darrick J. Wong
2026-09-25 6:10 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
2026-09-24 22:30 ` Darrick J. Wong
2026-09-25 6:12 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
2026-09-25 23:20 ` Darrick J. Wong
2026-09-26 6:13 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
2026-09-25 23:21 ` Darrick J. Wong
2026-09-24 9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
2026-09-25 23:24 ` Darrick J. Wong
2026-09-26 6:15 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
2026-09-29 0:42 ` Darrick J. Wong
2026-10-05 12:59 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
2026-09-29 1:01 ` Darrick J. Wong
2026-10-05 13:00 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
2026-09-29 1:06 ` Darrick J. Wong
2026-10-05 13:11 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
2026-09-29 1:19 ` Darrick J. Wong
2026-10-05 13:13 ` Christoph Hellwig
2026-09-24 9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
2026-09-29 1:25 ` Darrick J. Wong
2026-10-05 13:14 ` Christoph Hellwig
2026-10-08 11:43 ` Anuj gupta
2026-09-24 9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
2026-09-29 1:26 ` Darrick J. Wong
2026-09-24 9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
2026-09-29 1:27 ` Darrick J. Wong
2026-09-24 9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
2026-09-29 1:27 ` Darrick J. Wong
2026-10-05 13:16 ` Christoph Hellwig
2026-09-24 22:52 ` Dave Chinner [this message]
2026-09-25 6:27 ` support for " Christoph Hellwig
2026-09-27 22:59 ` Dave Chinner
2026-09-28 5:24 ` Christoph Hellwig
2026-09-29 14:11 ` Dave Chinner
2026-09-30 7:11 ` Dave Chinner
2026-10-05 13:53 ` Christoph Hellwig
2026-10-06 5:31 ` Dave Chinner
2026-10-07 13:46 ` Christoph Hellwig
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arWpyhFyV3ouYAEx@dread \
--to=dgc@kernel.org \
--cc=axboe@kernel.dk \
--cc=brauner@kernel.org \
--cc=cem@kernel.org \
--cc=djwong@kernel.org \
--cc=hch@lst.de \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-xfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox