Linux XFS filesystem development
 help / color / mirror / Atom feed
From: Dave Chinner <dgc@kernel.org>
To: Christoph Hellwig <hch@lst.de>
Cc: Carlos Maiolino <cem@kernel.org>,
	"Darrick J . Wong" <djwong@kernel.org>,
	Jens Axboe <axboe@kernel.dk>,
	Christian Brauner <brauner@kernel.org>,
	linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: support for RT data checksums
Date: Tue, 6 Oct 2026 16:31:29 +1100	[thread overview]
Message-ID: <asSHsVAYevL_VzdF@dread> (raw)
In-Reply-To: <20261005135320.GA29829@lst.de>

On Mon, Oct 05, 2026 at 03:53:20PM +0200, Christoph Hellwig wrote:
> On Wed, Sep 30, 2026 at 05:11:58PM +1000, Dave Chinner wrote:
> > My initial thought was that on disk format structures shouldn't be
> > defined by the limitations of the OS memory allocation, but <shrug>.
> > It kept nagging at me, though.
> 
> Well, it would be nice to be able to design without all the real-life
> constraints around us, wouldn't it?  I'm trying to strike a balance
> between what would be useful (go big) and what is feasible.  If we
> want other values, we can always increase the support range for
> newer kernels and tools.
> 
> > I looked more closely at what XFS_BLI_PREALLOC did to try to
> > understand why it existed.  It triggers a max-sized CIL logvec
> > structure for the buffer object. For a 32kB buffer logged as a
> > single contiguous range, this ends up being about 32kB + a logvec
> > header, plus a log iovec, plus a BLF, plus a couple of ophdrs. So
> > it's about 32kB + 200-250 bytes.
> > 
> > That means the shadow buffer for a csum buffer is always considered
> > a costly allocation by the MM subsystem.
> 
> It ends up using vmalloc exclusively based on tracing for me,
> but that might be different on different systems.

That'll be because the repeated multipage allocations eventually end
up fragmenting memory and so anything over 8-16kB will go straight
to vmalloc.

> > Hence for csum BLI, if a CIL flush happens on a partially filled
> > buffer, a good amount of that shadow buffer will go unused. Then we
> > allocate another (costly) shadow buffer on the next update. If CIL
> > flushes happen frequently enough then we will be repeatedly doing
> > costly allocations for shadow buffers that we don't actually use.
> 
> Yes.  But if we don't do this we realloc for every few blocks
> written, which is a lot more costly.

I'm not sure I understand what you are refering to there. I'm not
talking aobut the 'realloc because logged size grows', I'm talking
aoubt 'realloc because CIL pushes steal the shadow buffer'.

> > Not ideal - I think that means the original "sized for mm fast path"
> > intent is really only valid for the read side of the csum
> > algorithms as implemented by the patchset.
> 
> It is valid for the xfs_buf backing where we actually hit the folio
> allocator.  Which is used both for read and write, but obviously
> most workloads tend to hit reads a lot harder than writes.

Not if the working set fits in cache. A 'read heavy' workload will
often change to 'write heavy' when it fits in cache.....

> > Ok, we have a solution to this problem. I created ordered buffers
> > and one-shot log items to avoid the journalling overhead of static
> > inode buffer initialisation back in 2013. The ICREATE log item is
> > the one-shot log item that records a buffer should be initialised,
> > and the ordered buffer allows the modified buffer to be passed
> > through the journal to metadta writeback without it's contents being
> > logged.
> > 
> > Given that csum updates are a small, known size, non-overlapping
> > one-shot update to a buffer, they fit the same model that
> > ICREAT+ordered implements. Adding a new CSUM log item made up of a
> > format header and varible size csum payload region provides the
> > equivalent of the ICREAT item for journalled inode buffer
> > initialisation.
> 
> I initially looked into intent/done based csums, but we still end
> with an allocation per log operation, and a memcpy both into that
> and into the buffer, while adding a lot of new log items.

I'm not talking about an intent/done based setup - that requires an
intent transaction, then a buffer + done transaction. That's very
different to what I suggested (and what ICREATE implements).

I'm talking about a single one-shot transaction that logs the csums
and that only. The buffer is ordered, so not logged. Single
transaction, generally small in size, never gets relogged, can be
committed asynchronously as long as it is replayed before the file
offset relocation BMBT update in recovery.

The CIL can handle hundreds of thousands of csum objects just fine
(no different to logging hundreds of thousands of inode cores). It
is simpler than intents, too, because they are at checkpoint
completion and never enter the AIL (unlike intents). The only thing
that ends up the AIL is the BLI, and that works like a
INODE_ALLOC_BUF in that it remains at the initial LSN in the AIL
even when it gets updated. i.e. once it is in the AIL, it never gets
moved forward, but it continues to aggregate changes until it gets
written back with the LSN of the latest committed change stamped
into it so recovery does the right thing with it.

So, yeah, I'm definitely not suggesting using intents...

> One thing I played with for a while until I realized that the
> simple buf item actually provides good enough performance is
> special xfs_log_vec that is not included in the main log vec /
> shadow allocation but points to external memory.

That's problematic. The reason delayed logging works is that it
broke the dependency between external memory that log items pointed
at needing to be locked and stable until the external memory was
copied into the iclogs. The disconnection of the objects passed to
xfs_trans_commit() vs xlog_write() whilst keeping the logged data
stable is what allows the CIL to work

The shadow buffer does that decoupling, and it means that there is
no requirement for the original logged item to remain stable, or
even remain in existence whilst the CIL holds onto the infomration
that needs to be journalled. At checkpoint completion, shadow buffer
is also used to do a reverse lookup to the log item to enable
insertion into the AIL.  IOWs, log items are not tracked across
journal checkpoint IO - shadow buffers are, and if you get rid of
shadow buffers for a log item, we have to special case that
everywhere in the LV/checkpoint handling. I dont' think that's a good
idea.

The shadow buffer also avoids the need for the CIL to lock external
objects to copy the data out of them. The lock order is lock external
object -> commit -> read-lock checkpoint - format into CIL -> unlock.
When pushing, the order is write-lock checkpoint -> lock iclog ->
format from CIL into iclog -> unlock iclog -> unlock chkpt. We also
can call xfs_log_force() whilst holding inode locks, putting iclog
locks inside high level object locks.

Hence we really can't lock external objects from the checkpoint side
because of the lock inversion problems they entail. So object
stability is a problem, and ....

> This obviously
> only works for fixed size non-overlapping regions, but then isn't
> too bad. 

"trust me, bro!" is not my idea of maintainable, landmine free
design, especially now with LLMs being able to poke holes in complex
zero-copy/object sharing schemes and exploit them in less than
obvious ways...

> This is the prep work for it, which I recently
> refreshed:
> 
> https://git.infradead.org/?p=users/hch/xfs.git;a=shortlog;h=refs/heads/xlog-ophdr

Not a fan of rewriting xlog_write() -again-, this time to bring back
all the bad old patterns of managing ophdr space itself. We got rid of
that method of managing ophdrs because of all the special accounting
it needs to sprinkle through the logic to get log space consumption
correct. It was complex, difficult to reason about, and a source of
bugs.

Fixing these problems was the one of the main reasons we moved all
the ophdr management and accounting out into the CIL and logvecs to
begin with. Hence I'm not a great fan of going back to the old
way, whatever the reason.

> These can work with the buf_item on-disk format, so I'd rather not
> prematurely optimize it, as the prototype shows that I can go to
> that any time I want.  And eventually I think I'd want to go
> there, as it drastically reduces the memory usage if only the
> format header and two ophrs need to be allocated ontop of the
> backing buffer.  But there's plenty more important things on the
> plate for now.

I'd much prefer we use a method we know works and scales rahter than
create something new that requires punching through abstractions,
can't guarantee stability or lifetime of external objects, and isn't
demonstrated to be necessary to meet performance requirements.

-Dave.
-- 
Dave Chinner
dgc@kernel.org

  reply	other threads:[~2026-10-06  5:31 UTC|newest]

Thread overview: 69+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24  9:59 support for RT data checksums Christoph Hellwig
2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
2026-09-24 20:29   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
2026-09-24 21:39   ` Darrick J. Wong
2026-09-25  5:53     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
2026-09-24 21:43   ` Darrick J. Wong
2026-09-25  5:54     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
2026-09-24 21:44   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
2026-09-24 21:49   ` Darrick J. Wong
2026-09-25  5:57     ` Christoph Hellwig
2026-10-08 11:46   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
2026-09-24 22:03   ` Darrick J. Wong
2026-09-25  5:58     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
2026-09-24 22:13   ` Darrick J. Wong
2026-09-25  0:04     ` Eric Biggers
2026-09-25  6:01     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
2026-09-24 22:24   ` Darrick J. Wong
2026-09-25  6:10     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
2026-09-24 22:30   ` Darrick J. Wong
2026-09-25  6:12     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
2026-09-25 23:20   ` Darrick J. Wong
2026-09-26  6:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
2026-09-25 23:21   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
2026-09-25 23:24   ` Darrick J. Wong
2026-09-26  6:15     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
2026-09-29  0:42   ` Darrick J. Wong
2026-10-05 12:59     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
2026-09-29  1:01   ` Darrick J. Wong
2026-10-05 13:00     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
2026-09-29  1:06   ` Darrick J. Wong
2026-10-05 13:11     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
2026-09-29  1:19   ` Darrick J. Wong
2026-10-05 13:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
2026-09-29  1:25   ` Darrick J. Wong
2026-10-05 13:14     ` Christoph Hellwig
2026-10-08 11:43   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
2026-09-29  1:26   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-10-05 13:16     ` Christoph Hellwig
2026-09-24 22:52 ` support for " Dave Chinner
2026-09-25  6:27   ` Christoph Hellwig
2026-09-27 22:59     ` Dave Chinner
2026-09-28  5:24       ` Christoph Hellwig
2026-09-29 14:11         ` Dave Chinner
2026-09-30  7:11           ` Dave Chinner
2026-10-05 13:53             ` Christoph Hellwig
2026-10-06  5:31               ` Dave Chinner [this message]
2026-10-07 13:46                 ` Christoph Hellwig

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asSHsVAYevL_VzdF@dread \
    --to=dgc@kernel.org \
    --cc=axboe@kernel.dk \
    --cc=brauner@kernel.org \
    --cc=cem@kernel.org \
    --cc=djwong@kernel.org \
    --cc=hch@lst.de \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-xfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox