Linux XFS filesystem development
 help / color / mirror / Atom feed
From: Christoph Hellwig <hch@lst.de>
To: Dave Chinner <dgc@kernel.org>
Cc: Christoph Hellwig <hch@lst.de>, Carlos Maiolino <cem@kernel.org>,
	"Darrick J . Wong" <djwong@kernel.org>,
	Jens Axboe <axboe@kernel.dk>,
	Christian Brauner <brauner@kernel.org>,
	linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: support for RT data checksums
Date: Mon, 5 Oct 2026 15:53:20 +0200	[thread overview]
Message-ID: <20261005135320.GA29829@lst.de> (raw)
In-Reply-To: <ary2PpSp1xTgrd3k@dread>

On Wed, Sep 30, 2026 at 05:11:58PM +1000, Dave Chinner wrote:
> My initial thought was that on disk format structures shouldn't be
> defined by the limitations of the OS memory allocation, but <shrug>.
> It kept nagging at me, though.

Well, it would be nice to be able to design without all the real-life
constraints around us, wouldn't it?  I'm trying to strike a balance
between what would be useful (go big) and what is feasible.  If we
want other values, we can always increase the support range for
newer kernels and tools.

> I looked more closely at what XFS_BLI_PREALLOC did to try to
> understand why it existed.  It triggers a max-sized CIL logvec
> structure for the buffer object. For a 32kB buffer logged as a
> single contiguous range, this ends up being about 32kB + a logvec
> header, plus a log iovec, plus a BLF, plus a couple of ophdrs. So
> it's about 32kB + 200-250 bytes.
> 
> That means the shadow buffer for a csum buffer is always considered
> a costly allocation by the MM subsystem.

It ends up using vmalloc exclusively based on tracing for me,
but that might be different on different systems.

> Hence for csum BLI, if a CIL flush happens on a partially filled
> buffer, a good amount of that shadow buffer will go unused. Then we
> allocate another (costly) shadow buffer on the next update. If CIL
> flushes happen frequently enough then we will be repeatedly doing
> costly allocations for shadow buffers that we don't actually use.

Yes.  But if we don't do this we realloc for every few blocks
written, which is a lot more costly.

> Not ideal - I think that means the original "sized for mm fast path"
> intent is really only valid for the read side of the csum
> algorithms as implemented by the patchset.

It is valid for the xfs_buf backing where we actually hit the folio
allocator.  Which is used both for read and write, but obviously
most workloads tend to hit reads a lot harder than writes.

> Ok, we have a solution to this problem. I created ordered buffers
> and one-shot log items to avoid the journalling overhead of static
> inode buffer initialisation back in 2013. The ICREATE log item is
> the one-shot log item that records a buffer should be initialised,
> and the ordered buffer allows the modified buffer to be passed
> through the journal to metadta writeback without it's contents being
> logged.
> 
> Given that csum updates are a small, known size, non-overlapping
> one-shot update to a buffer, they fit the same model that
> ICREAT+ordered implements. Adding a new CSUM log item made up of a
> format header and varible size csum payload region provides the
> equivalent of the ICREAT item for journalled inode buffer
> initialisation.

I initially looked into intent/done based csums, but we still end
with an allocation per log operation, and a memcpy both into that
and into the buffer, while adding a lot of new log items.

One thing I played with for a while until I realized that the simple
buf item actually provides good enough performance is special
xfs_log_vec that is not included in the main log vec / shadow allocation
but points to external memory.  This obviously only works for fixed
size non-overlapping regions, but then isn't too bad.  This is the
prep work for it, which I recently refreshed:

https://git.infradead.org/?p=users/hch/xfs.git;a=shortlog;h=refs/heads/xlog-ophdr

These can work with the buf_item on-disk format, so I'd rather not
prematurely optimize it, as the prototype shows that I can go to that
any time I want.  And eventually I think I'd want to go there, as it
drastically reduces the memory usage if only the format header and
two ophrs need to be allocated ontop of the backing buffer.  But there's
plenty more important things on the plate for now.

  reply	other threads:[~2026-10-05 13:53 UTC|newest]

Thread overview: 69+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24  9:59 support for RT data checksums Christoph Hellwig
2026-09-24  9:59 ` [PATCH 01/21] block: export fs_bio_integrity_verify Christoph Hellwig
2026-09-24 20:29   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 02/21] iomap: add support for data checksumming Christoph Hellwig
2026-09-24 21:39   ` Darrick J. Wong
2026-09-25  5:53     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 03/21] xfs: add a xfs_buf_read_async buffer cache API Christoph Hellwig
2026-09-24 21:43   ` Darrick J. Wong
2026-09-25  5:54     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 04/21] xfs: add xfs_daddr_to_rgno and xfs_daddr_to_rgbno helpers Christoph Hellwig
2026-09-24 21:44   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 05/21] xfs: introduce XFS_BLI_PREALLOC Christoph Hellwig
2026-09-24 21:49   ` Darrick J. Wong
2026-09-25  5:57     ` Christoph Hellwig
2026-10-08 11:46   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 06/21] xfs: prepare xfs_rtfile_initialize_blocks for larger than FSB blocks Christoph Hellwig
2026-09-24 22:03   ` Darrick J. Wong
2026-09-25  5:58     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 07/21] xfs: relase zi_open_zones_lock over xfs_open_zone_put on unmount Christoph Hellwig
2026-09-24  9:59 ` [PATCH 08/21] xfs: define the RT data checksum on-disk format Christoph Hellwig
2026-09-24 22:13   ` Darrick J. Wong
2026-09-25  0:04     ` Eric Biggers
2026-09-25  6:01     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 09/21] xfs: add support for per-RTG csum files Christoph Hellwig
2026-09-24 22:24   ` Darrick J. Wong
2026-09-25  6:10     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 10/21] xfs: calculate the log reservation for logging data checksum buffers Christoph Hellwig
2026-09-24 22:30   ` Darrick J. Wong
2026-09-25  6:12     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 11/21] xfs: core RT data checksum support Christoph Hellwig
2026-09-25 23:20   ` Darrick J. Wong
2026-09-26  6:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 12/21] xfs: data checksums require stable writes Christoph Hellwig
2026-09-25 23:21   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 13/21] xfs: require file system block size alignment when using data checksums Christoph Hellwig
2026-09-25 23:24   ` Darrick J. Wong
2026-09-26  6:15     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 14/21] xfs: add support for reading with " Christoph Hellwig
2026-09-29  0:42   ` Darrick J. Wong
2026-10-05 12:59     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 15/21] xfs: add support for writing " Christoph Hellwig
2026-09-29  1:01   ` Darrick J. Wong
2026-10-05 13:00     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 16/21] xfs: add data checksum support to zoned garbage collection Christoph Hellwig
2026-09-29  1:06   ` Darrick J. Wong
2026-10-05 13:11     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 17/21] xfs: verify data checksums during media verification Christoph Hellwig
2026-09-29  1:19   ` Darrick J. Wong
2026-10-05 13:13     ` Christoph Hellwig
2026-09-24  9:59 ` [PATCH 18/21] xfs: don't try to verify checksums on empty zones Christoph Hellwig
2026-09-29  1:25   ` Darrick J. Wong
2026-10-05 13:14     ` Christoph Hellwig
2026-10-08 11:43   ` Anuj gupta
2026-09-24  9:59 ` [PATCH 19/21] xfs: report RT data checksum information via XFS_FSOP_GEOM Christoph Hellwig
2026-09-29  1:26   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 20/21] xfs: add an experimental feature warning for RT data checksums Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-09-24  9:59 ` [PATCH 21/21] xfs: enable " Christoph Hellwig
2026-09-29  1:27   ` Darrick J. Wong
2026-10-05 13:16     ` Christoph Hellwig
2026-09-24 22:52 ` support for " Dave Chinner
2026-09-25  6:27   ` Christoph Hellwig
2026-09-27 22:59     ` Dave Chinner
2026-09-28  5:24       ` Christoph Hellwig
2026-09-29 14:11         ` Dave Chinner
2026-09-30  7:11           ` Dave Chinner
2026-10-05 13:53             ` Christoph Hellwig [this message]
2026-10-06  5:31               ` Dave Chinner
2026-10-07 13:46                 ` Christoph Hellwig

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005135320.GA29829@lst.de \
    --to=hch@lst.de \
    --cc=axboe@kernel.dk \
    --cc=brauner@kernel.org \
    --cc=cem@kernel.org \
    --cc=dgc@kernel.org \
    --cc=djwong@kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-xfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox