Linux filesystem development
 help / color / mirror / Atom feed
From: Dave Chinner <dgc@kernel.org>
To: Christoph Hellwig <hch@lst.de>
Cc: Carlos Maiolino <cem@kernel.org>,
	"Darrick J . Wong" <djwong@kernel.org>,
	Jens Axboe <axboe@kernel.dk>,
	Christian Brauner <brauner@kernel.org>,
	linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: support for RT data checksums
Date: Mon, 28 Sep 2026 08:59:31 +1000	[thread overview]
Message-ID: <armf05yfzhSBUsUh@dread> (raw)
In-Reply-To: <20260925062710.GA3798@lst.de>

On Fri, Sep 25, 2026 at 08:27:10AM +0200, Christoph Hellwig wrote:
> On Fri, Sep 25, 2026 at 08:52:58AM +1000, Dave Chinner wrote:
> > There's new buffer and inode locking in transactions,
> 
> Not sure what is new about that.  Tere is a single new transaction,
> which logs a single buffer per transaction.  Not exactly new and
> dancy.

You haven't answered any of my concerns - you're just handwaving
them away and....

> > and there's a
> > whole new buffer cache interface to "read a buffer", and that is
> > used to open code reading checksum buffers and joining them to a
> > transaction rather than using the existing xfs_trans_read_buf...()
> > interfaces.  That in itself needs careful consideration, and clear
> > justification for why it must be duplicated to stand outside all the
> > existing BLI/transaction APIs, especially given all the "use the new
> > async buf read interface to do sync buffer reads" behaviour across
> > the patchset that could just use the existing interfaces.
> 
> I'm not sure what to make of this.  The paragraph almost reads like
> AI slop to me.

... calling the concerns of an experienced engineer "AI slop".

I'm so disappointed right now.

You're better than this, Christoph. You know better than to attack
the person instead of addressing the technical concerns they've
raised. Calling the concerns of an experienced engineer "AI Slop" is
also pretty insulting.

You also know this has a chilling effect - how many people are going
to be willing to say anything negative about your code, if all they
get from it is a bunch of insults in return?

I don't tolerate people who behave like this towards me or other
team members anymore. "But I write lots of code" is just not good
enough anymore, Christoph. You need to do better.

FWIW, let's address the "AI Slop" aspect. I don't need to us an LLM
to review code - I know the XFS code base as well as you do,
Christoph.

I hadn't even considered passing your code to a LLM yet to see how
many holes I can poke in it. Everything I wrote document concerns my
own brain raised  when doing an initial high level scan of the
patchset. I always do that first so I don't waste time (or tokens!)
doing detailed review on something that needs deeper rework before
it is anywhere near ready for merge.

That scan raised lots of questions in my head about the change, and
so I documented some of my concerns and asked for more detail about
how all this new functionality is supposed to work to be documented.
Not just for me right now, but so there is a permanent record of the
design for future reference.

Performing critical analysis and asking for more detail to be
documented is not "AI Slop". A LLM would just plow on through with
whatever misunderstanding it started with and produce slop. It is
very much a human behaviour to stop and ask for more context,
detail, documentation, etc to address missing knoweldge before going
any further.

If you don't want meaningful review of this code, then just say so -
don't insult the people who spend their own time tryng to understand
the new functionality well enough to review it...

> The rationale is pretty clear and documented: it
> turns two dependent reads into two reads that work in parallel.

Maybe that is clear to you, but it certainly isn't from my initial
reading of the patches.

I don't have all the unwritten knowledge that is in your head that
you didn't document. I've never discussed this code with you, I have
no background on how it is supposed to work, etc. I am coming in
-cold-.

I have not been able to find clear, logical descriptions of
the algorithms for reading and writing checksums, the data
intregrity protocols, the crash recovery protocols, etc. At this
point, the only way for me to to understand details like this is to
-reverse engineer the design from implementation-.

That's the problem - you've presented an implementation without an
overall design being documented, and then asking people to review
the implementation without any other context. I need to understand
the design before I review the implementation, as that is the only
way I can perform a decent -implementation- review.

Indeed, people often complain that it is too hard to adequately
review complex changes, and this is one of the reasons: we are
presented with an implementation, and -zero- design/architecture
context from which to understand the implementation that has been
presented. We need to fix that.

> > I'd also like to have the format of the new on disk log item format
> > structures clearly documented (because we're going to have to
> 
> There is no new log item format, it uses the standard buffer log format.
> The buffer payload is somewhat new.  It is is the standard rt format,
> a xfs_rtbuf_blkinfo followed by the real payload, which is an array
> of checksums.

Yes, the buffer payload is new, and it uses a new BLF flag that
indicates it contains regions with some new on-disk format. And
there are interactions with fsync and data integrity requirements.
Document them!

> > validate them) at recovery time, and also have a clear explaination
> > of the data vs metadata ordering algorithms that ensures that
> > checksums are always valid in crash+recovery situations, especially
> > w.r.t. data integrity operations like fsync.
> 
> I think I explained it pretty well,

You didn't, and that's the problem I am asking you to address.

> but happy to repeat it again:
> The zoned write path writes data first, and then records bmap, rmap
> and used space tacking in the zone from the I/O completion handler.
> The rtcsum code builds on that and only logs that csum from that
> same I/O completion handler.  I.e. that data must have reached the
> device for the code to log it to be even called, and for devices
> with volatile write caches the generic cache flushing must work
> (it did not until recently, but the verification of this code found
> that bug and it is now fixed upstream).

So document the design in the first patch in the series! In detail,
in the tree alongside the code. That gives all reviewers the
necessary context to understand what comes next in the patchset.

Indeed, I ask this not just for reviewiers, but for anyone coming
along in the future that needs to understand how data checksums work
because they need to triage a bug, add new functionality, need to
optimise for performance, etc. Just because you understand how it
works and what you wrote in the commit messages, it doesn't mean how
it is supposed to work, why it works, or why it was implemented in
the way it was is obvious to anyone else.

Documenting the design helps -everyone-, not just now, but well into
the future as well.

-Dave.
-- 
Dave Chinner
dgc@kernel.org

  reply	other threads:[~2026-09-27 22:59 UTC|newest]

Thread overview: 69+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24  9:59 support for RT data checksums Christoph Hellwig
2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
2026-09-24 20:29   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
2026-09-24 21:39   ` Darrick J. Wong
2026-09-25  5:53     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
2026-09-24 21:43   ` Darrick J. Wong
2026-09-25  5:54     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
2026-09-24 21:44   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
2026-09-24 21:49   ` Darrick J. Wong
2026-09-25  5:57     ` Christoph Hellwig
2026-10-08 11:46   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
2026-09-24 22:03   ` Darrick J. Wong
2026-09-25  5:58     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
2026-09-24 22:13   ` Darrick J. Wong
2026-09-25  0:04     ` Eric Biggers
2026-09-25  6:01     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
2026-09-24 22:24   ` Darrick J. Wong
2026-09-25  6:10     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
2026-09-24 22:30   ` Darrick J. Wong
2026-09-25  6:12     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
2026-09-25 23:20   ` Darrick J. Wong
2026-09-26  6:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
2026-09-25 23:21   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
2026-09-25 23:24   ` Darrick J. Wong
2026-09-26  6:15     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
2026-09-29  0:42   ` Darrick J. Wong
2026-10-05 12:59     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
2026-09-29  1:01   ` Darrick J. Wong
2026-10-05 13:00     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
2026-09-29  1:06   ` Darrick J. Wong
2026-10-05 13:11     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
2026-09-29  1:19   ` Darrick J. Wong
2026-10-05 13:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
2026-09-29  1:25   ` Darrick J. Wong
2026-10-05 13:14     ` Christoph Hellwig
2026-10-08 11:43   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
2026-09-29  1:26   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-10-05 13:16     ` Christoph Hellwig
2026-09-24 22:52 ` support for " Dave Chinner
2026-09-25  6:27   ` Christoph Hellwig
2026-09-27 22:59     ` Dave Chinner [this message]
2026-09-28  5:24       ` Christoph Hellwig
2026-09-29 14:11         ` Dave Chinner
2026-09-30  7:11           ` Dave Chinner
2026-10-05 13:53             ` Christoph Hellwig
2026-10-06  5:31               ` Dave Chinner
2026-10-07 13:46                 ` Christoph Hellwig

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=armf05yfzhSBUsUh@dread \
    --to=dgc@kernel.org \
    --cc=axboe@kernel.dk \
    --cc=brauner@kernel.org \
    --cc=cem@kernel.org \
    --cc=djwong@kernel.org \
    --cc=hch@lst.de \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-xfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox