From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 746DA471406 for ; Thu, 24 Sep 2026 10:07:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.137.202.133 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790244482; cv=none; b=AeaKsNXkJdsr73YlzYmcGddVrjIZRO390BoQSv4kobpMrtseD9rQKD+EqIVPwpXb9ptz9i/xAVN8P7gVFKQkT0SBVb54mcxvAJxmScch6SRfmSflaxipnKK8g3RzrqOvNFdoiprBYouQuKvkxpmJzCgFRo+sfZhkQECUv81Fr70= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790244482; c=relaxed/simple; bh=LWEI1qCZEnTI0QsHSy3QVssXFdfIttIaePzty/seKqg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=VK+cnnhvP7pI6gadSaxuso9tUQhrgA96RwtkFEx1xgfQzQ3W6x+RTGrSjsEt+loX8GqJ3rzCxaau9Cd3rJf9CMVLxHDBZdvAPHt6XDTBIFoza5eXV7qGPJvg6DA2/rfEcp02NaEIXgUaw0yiv6b+9WkO7BhxbmGuOTpTJAgDrUQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=fail (p=none dis=none) header.from=lst.de; spf=none smtp.mailfrom=bombadil.srs.infradead.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b=yiei7/n5; arc=none smtp.client-ip=198.137.202.133 Authentication-Results: smtp.subspace.kernel.org; dmarc=fail (p=none dis=none) header.from=lst.de Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=bombadil.srs.infradead.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=infradead.org header.i=@infradead.org header.b="yiei7/n5" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=bombadil.20210309; h=Content-Transfer-Encoding: MIME-Version:References:In-Reply-To:Message-ID:Date:Subject:Cc:To:From:Sender :Reply-To:Content-Type:Content-ID:Content-Description; bh=US9bllrBa537Zt4t7VldAOGscSuPiqqIOaFMDGVselM=; b=yiei7/n5hIhWUhJO9ie3dQT52Z Dl298E9EqRAv8M7rAHMP8IHhzoAW2Fg0FsBaAMtv/znTUvki7AH0ykw8ybStr+ltcD29eXHFA9noE UOUobb2ETwvI/hXcoFXu5sHNhmkuFcNBzWSbsDsxJQM9N7sksRW4ZFMPGYiZc7WCngkdW9JrdgrMV A0FcIrXaij25Xs1qI8dtnGuQVZS3yA0GL0xQvMgdtmDjVPwVwOBk5+ukuMPYWEu+d9x8pmjAGIRR/ MLff4geGQWQZYxCDpVjUAhskztsRKhGfTet+rOvpMLgkex/p8J67c/8ZoixN/zDUsKME2TpM3M2VA zWCk5Itg==; Received: from 85-127-111-79.dsl.dynamic.surfer.at ([85.127.111.79] helo=localhost) by bombadil.infradead.org with esmtpsa (Exim 4.99.1 #2 (Red Hat Linux)) id 1x9gMW-0000000AfaE-13xs; Thu, 24 Sep 2026 10:07:33 +0000 From: Christoph Hellwig To: Andrey Albershteyn Cc: "Darrick J . Wong" , linux-xfs@vger.kernel.org Subject: [PATCH 31/32] repair: support RT data checksums Date: Thu, 24 Sep 2026 12:04:21 +0200 Message-ID: <20260924100512.2733748-32-hch@lst.de> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260924100512.2733748-1-hch@lst.de> References: <20260924100512.2733748-1-hch@lst.de> Precedence: bulk X-Mailing-List: linux-xfs@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-SRS-Rewrite: SMTP reverse-path rewritten from by bombadil.infradead.org. See http://www.infradead.org/rpr.html Recognize the rtcsum per-RTG metadir files, and sanity check a few well known attributes for them. If a csum file is missing, or had to be nuke, regenerate it by re-calculating the checksums. While this does lose the protection of the checksums, this is probably still better than an unmountable file system. Note that due to the lack of the regular buffer readahead path, reading the data and csum buffers for regenerating the csum file is fully synchronous and thus slow. Hopefully we can come up with a better version for online repair and never have to use this for real. Signed-off-by: Christoph Hellwig --- repair/Makefile | 1 + repair/dinode.c | 46 ++++++++++++++ repair/phase6.c | 122 +++++++++++++++++++++++++++++++++++- repair/rt.c | 36 +++++++++++ repair/rt.h | 10 +++ repair/rtcsum.c | 161 ++++++++++++++++++++++++++++++++++++++++++++++++ 6 files changed, 375 insertions(+), 1 deletion(-) create mode 100644 repair/rtcsum.c diff --git a/repair/Makefile b/repair/Makefile index fb0b2f96cc91..27eb68e2b5d5 100644 --- a/repair/Makefile +++ b/repair/Makefile @@ -73,6 +73,7 @@ CFILES = \ rcbag.c \ rmap.c \ rt.c \ + rtcsum.c \ rtrefcount_repair.c \ rtrmap_repair.c \ sb.c \ diff --git a/repair/dinode.c b/repair/dinode.c index 48939f8bd159..5c03c5f50689 100644 --- a/repair/dinode.c +++ b/repair/dinode.c @@ -22,6 +22,8 @@ #include "rmap.h" #include "bmap_repair.h" #include "rt.h" +#include "libfrog/crc64.h" +#include "xfs_rtcsumfile.h" /* inode types */ enum xr_ino_type { @@ -41,6 +43,7 @@ enum xr_ino_type { XR_INO_PQUOTA, /* project quota inode */ XR_INO_RTRMAP, /* realtime rmap */ XR_INO_RTREFC, /* realtime refcount */ + XR_INO_RTCSUM, /* realtime data checksum */ XR_INO_MAX }; @@ -61,6 +64,7 @@ static const char *xr_ino_type_name[] = { [XR_INO_PQUOTA] = N_("project quota"), [XR_INO_RTRMAP] = N_("realtime rmap"), [XR_INO_RTREFC] = N_("realtime refcount"), + [XR_INO_RTCSUM] = N_("realtime data checksum"), }; static_assert(ARRAY_SIZE(xr_ino_type_name) == XR_INO_MAX); @@ -2064,6 +2068,43 @@ _("bad # of extents (%" PRIu64 ") for %s inode %" PRIu64 "\n"), return 0; } +static int +process_check_rtcsum_inode( + struct xfs_mount *mp, + struct xfs_dinode *dinoc, + xfs_ino_t lino, + enum xr_ino_type *type, + int *dirty) +{ + xfs_fsize_t size = be64_to_cpu(dinoc->di_size); + xfs_fsize_t expected_size; + int error; + + error = process_check_rt_inode(mp, dinoc, lino, type, dirty, + XR_INO_RTCSUM, _("realtime data checksums")); + if (error) + return error; + + /* rtcsum inodes must be contiguous */ + if (xfs_dfork_data_extents(dinoc) != 1) { + do_warn( +_("non-contiguous rtcsum inode %" PRIu64 "\n"), lino); + return 1; + } + + expected_size = (xfs_off_t)mp->m_sb.sb_rgextents << mp->m_rtcsum_shift; + expected_size = (expected_size + xfs_rtcsum_payload_size(mp) - 1) / + xfs_rtcsum_payload_size(mp) * mp->m_rtcsum_bsize; + if (size != expected_size) { + do_warn( +_("unexpected rtcsum file size (%" PRId64 ") for ino %" PRIu64 "\n"), + size, lino); + return 1; + } + + return 0; +} + /* * If inode is a superblock inode, does type check to make sure is it valid. * Returns 0 if it's valid, non-zero if it needs to be cleared. @@ -2130,6 +2171,8 @@ process_check_metadata_inodes( if (is_rtrefcount_inode(lino)) return process_check_rt_inode(mp, dinoc, lino, type, dirty, XR_INO_RTREFC, _("realtime refcount btree")); + if (is_rtcsum_inode(lino)) + return process_check_rtcsum_inode(mp, dinoc, lino, type, dirty); return 0; } @@ -2977,6 +3020,7 @@ process_dinode_metafile( case XR_INO_UQUOTA: case XR_INO_GQUOTA: case XR_INO_PQUOTA: + case XR_INO_RTCSUM: /* * Quota checking and repair doesn't happen until phase7, so * preserve quota inodes and their contents for later. @@ -3555,6 +3599,8 @@ _("bad (negative) size %" PRId64 " on inode %" PRIu64 "\n"), type = XR_INO_RTRMAP; else if (is_rtrefcount_inode(lino)) type = XR_INO_RTREFC; + else if (is_rtcsum_inode(lino)) + type = XR_INO_RTCSUM; else type = XR_INO_DATA; break; diff --git a/repair/phase6.c b/repair/phase6.c index f3951a3d0709..c7bc1cabd0d3 100644 --- a/repair/phase6.c +++ b/repair/phase6.c @@ -6,7 +6,6 @@ #include "libxfs.h" #include "threads.h" -#include "threads.h" #include "prefetch.h" #include "avl.h" #include "globals.h" @@ -23,6 +22,8 @@ #include "repair/quotacheck.h" #include "repair/slab.h" #include "repair/rmap.h" +#include "libfrog/crc64.h" +#include "xfs_rtcsumfile.h" static xfs_ino_t orphanage_ino; @@ -703,6 +704,123 @@ ensure_rtgroup_refcountbt( populate_rtgroup_refcountbt(rtg, est_fdblocks); } +/* + * Link a metadata directory inode. + */ +static int +metadir_link( + struct xfs_inode *dp, + struct xfs_inode *ip, + const char *path, + enum xfs_metafile_type type) +{ + struct xfs_metadir_update upd = { + .dp = dp, + .metafile_type = type, + .ip = ip, + .path = path, + }; + int error; + + error = xfs_metadir_start_link(&upd); + if (error) + return error; + + error = xfs_metadir_link(&upd); + if (error) + return error; + + xfs_trans_log_inode(upd.tp, upd.ip, XFS_ILOG_CORE); + return xfs_metadir_commit(&upd); +} + +static void +ensure_rtgroup_csum( + struct xfs_rtgroup *rtg) +{ + struct xfs_mount *mp = rtg_mount(rtg); + xfs_rgnumber_t rgno = rtg_rgno(rtg); + struct xfs_inode *dp = mp->m_rtdirip; + struct xfs_inode *ip; + int error; + + if (no_modify) { + if (rtcsum_ino_is_bad(rgno)) + do_warn(_("would reset RG %u csum inode\n"), rgno);; + return; + } + + if (!rtcsum_ino_is_bad(rgno)) { + /* + * The /realtime directory has been discarded, but we should be + * able to iget the inodes directly. + */ + error = -libxfs_metafile_iget(mp, rtcsum_ino(rgno), + XFS_METAFILE_RTCSUM, &ip); + if (error) { + do_warn( +_("Could not open RG %u csum inode, error %d\n"), rgno, error); + rtcsum_ino_mark_bad(rgno); + } + } + + if (rtcsum_ino_is_bad(rgno)) { + do_warn(_( +"resetting RG %u csum inode, regenerating data checksums\n"), + rgno); + error = -libxfs_rtginode_create(rtg, XFS_RTGI_CSUM, false); + if (error) { + do_warn( +_("Couldn't create RG %u csum inode, error %d\n"), rgno, error); + return; + } + ip = rtg->rtg_inodes[XFS_RTGI_CSUM]; + error = -xfs_rtcsum_alloc_blocks(rtg); + if (error) { + do_warn( +_("Initialization of RG %u csum inode failed, error %d"), rgno, error); + return; + } + + calculate_rtgroup_csums(rtg); + } else { + struct xfs_trans *tp; + const char *name; + + /* Erase parent pointers before we create the new link */ + try_erase_parent_ptrs(ip); + + name = xfs_rtginode_path(rtg_rgno(rtg), XFS_RTGI_CSUM); + error = -metadir_link(dp, ip, name, XFS_METAFILE_RTCSUM); + kfree(name); + if (error) { + do_warn( +_("Couldn't link RG %u csum inode, error %d\n"), rgno, error); + return; + } + + /* + * Reset the link count to 1 because the link above bumped it. + */ + error = -libxfs_trans_alloc_inode(ip, &M_RES(mp)->tr_ichange, + 0, 0, false, &tp); + if (!error) { + set_nlink(VFS_I(ip), 1); + libxfs_trans_log_inode(tp, ip, XFS_ILOG_CORE); + error = -libxfs_trans_commit(tp); + } + if (error) + do_error( +_("Couldn't reset link count on RG %i quota inode, error %d\n"), + rgno, error); + } + + /* Mark the inode in use. */ + mark_ino_inuse(mp, I_INO(ip), S_IFREG, I_INO(dp)); + mark_ino_metadata(mp, I_INO(ip)); + libxfs_irele(ip); +} + /* Initialize a root directory. */ static int init_fs_root_dir( @@ -3473,6 +3591,8 @@ _(" - resetting contents of realtime bitmap and summary inodes\n")); } ensure_rtgroup_rmapbt(rtg, est_fdblocks); ensure_rtgroup_refcountbt(rtg, est_fdblocks); + if (xfs_has_rtcsum(mp)) + ensure_rtgroup_csum(rtg); } } diff --git a/repair/rt.c b/repair/rt.c index b0ff775bd339..4b64a91fbe29 100644 --- a/repair/rt.c +++ b/repair/rt.c @@ -26,6 +26,8 @@ struct rtg_computed { }; struct rtg_computed *rt_computed; +static xfs_ino_t *rtcsum_inos; + static inline void set_rtword( struct xfs_mount *mp, @@ -410,6 +412,27 @@ fill_rtsummary( _("couldn't re-initialize realtime summary inode, error %d\n"), error); } +bool +rtcsum_ino_is_bad( + xfs_rgnumber_t rgno) +{ + return rtcsum_inos[rgno] == NULLFSINO; +} + +void +rtcsum_ino_mark_bad( + xfs_rgnumber_t rgno) +{ + rtcsum_inos[rgno] = NULLFSINO; +} + +xfs_ino_t +rtcsum_ino( + xfs_rgnumber_t rgno) +{ + return rtcsum_inos[rgno]; +} + bool is_rtgroup_inode( xfs_ino_t ino, @@ -471,6 +494,13 @@ mark_rtginode( goto out_corrupt; } + /* + * Record the inode numbers of the data checksum inodes, as we don't + * just blow those away like other per-RTG metadata. + */ + if (type == XFS_RTGI_CSUM) + rtcsum_inos[rtg_rgno(rtg)] = I_INO(ip); + /* * Phase 3 will clear the ondisk inodes of all rt metadata files, but * it doesn't reset any blocks. Keep the incore inodes loaded so that @@ -494,6 +524,12 @@ discover_rtgroup_inodes( int error, err2; int i; + rtcsum_inos = calloc(mp->m_sb.sb_rgcount, sizeof(xfs_ino_t)); + if (!rtcsum_inos) + do_error(_("could not allocate csum ino array\n")); + for (i = 0; i < mp->m_sb.sb_rgcount; i++) + rtcsum_inos[i] = NULLFSINO; + tp = libxfs_trans_alloc_empty(mp); if (xfs_has_rtgroups(mp) && mp->m_sb.sb_rgcount > 0) { error = -libxfs_rtginode_load_parent(tp); diff --git a/repair/rt.h b/repair/rt.h index e4f3d5d9af31..9400b230de59 100644 --- a/repair/rt.h +++ b/repair/rt.h @@ -13,6 +13,10 @@ void check_rtsummary(struct xfs_mount *mp); void fill_rtbitmap(struct xfs_rtgroup *rtg); void fill_rtsummary(struct xfs_rtgroup *rtg); +bool rtcsum_ino_is_bad(xfs_rgnumber_t rgno); +void rtcsum_ino_mark_bad(xfs_rgnumber_t rgno); +xfs_ino_t rtcsum_ino(xfs_rgnumber_t rgno); + void discover_rtgroup_inodes(struct xfs_mount *mp); void unload_rtgroup_inodes(struct xfs_mount *mp); @@ -37,6 +41,10 @@ static inline bool is_rtrefcount_inode(xfs_ino_t ino) { return is_rtgroup_inode(ino, XFS_RTGI_REFCOUNT); } +static inline bool is_rtcsum_inode(xfs_ino_t ino) +{ + return is_rtgroup_inode(ino, XFS_RTGI_CSUM); +} void mark_rtgroup_inodes_bad(struct xfs_mount *mp, enum xfs_rtg_inodes type); bool rtgroup_inodes_were_bad(enum xfs_rtg_inodes type); @@ -44,4 +52,6 @@ bool rtgroup_inodes_were_bad(enum xfs_rtg_inodes type); void check_rtsb(struct xfs_mount *mp); void rewrite_rtsb(struct xfs_mount *mp); +void calculate_rtgroup_csums(struct xfs_rtgroup *rtg); + #endif /* _XFS_REPAIR_RT_H_ */ diff --git a/repair/rtcsum.c b/repair/rtcsum.c new file mode 100644 index 000000000000..c4d629c18e8d --- /dev/null +++ b/repair/rtcsum.c @@ -0,0 +1,161 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Christoph Hellwig. + * + * Rebuild data checksums when the csum metafile was lost. This doens't perform + * great and is inteded as a last resort. + */ +#include "libxfs.h" +#include "globals.h" +#include "protos.h" +#include "rt.h" +#include "err_protos.h" +#include "libfrog/crc64.h" +#include "xfs_platform.h" +#include "xfs_rtcsumfile.h" + +/* default to 1MiB data reads to make the performance only somewhat horrible. */ +#define DATA_BLOCKS_PER_BUF 256 + +static int +read_data( + struct xfs_buftarg *btp, + xfs_daddr_t blkno, + unsigned int nblks, + void *buf) +{ + int fd = btp->bt_bdev_fd; + off_t pos = LIBXFS_BBTOOFF64(blkno); + size_t len = BBTOB(nblks); + ssize_t ret; + + ret = pread(fd, buf, len, pos); + if (ret < 0) { + ret = errno; + + fprintf(stderr, _("%s: read failed: %s\n"), + progname, strerror(ret)); + return -ret; + } + if (ret != len) { + fprintf(stderr, _("%s: error - read only %zd of %zd bytes\n"), + progname, ret, len); + return -EIO; + } + + return 0; +} + +static int +calculate_buf_csums( + struct xfs_rtgroup *rtg, + void *data_buf, + xfs_rgblock_t rgbno, + unsigned int *nr) +{ + struct xfs_mount *mp = rtg_mount(rtg); + unsigned int boff = xfs_rgb_to_rtcsumoff(mp, rgbno); + struct xfs_trans_res tres = M_RES(mp)->tr_csum; + xfs_daddr_t csum_daddr; + void *csum_buf; + struct xfs_buf *csum_bp; + struct xfs_trans *tp; + int error; + unsigned int i = 0; + + *nr = min(*nr, xfs_rtcsum_len_to_extlen(mp, mp->m_rtcsum_bsize - boff)); + + error = xfs_rtcsum_bmap(rtg, rgbno, &csum_daddr); + if (error) { + do_error( +_("Cannot bmap new csum file for RG %u, error %d\n"), + rtg_rgno(rtg), error); + return error; + } + + tres.tr_logres = xfs_calc_csum_reservation(mp, + xfs_extlen_to_rtcsum_len(mp, *nr)); + ASSERT(tres.tr_logres <= M_RES(mp)->tr_csum.tr_logres); + + error = libxfs_trans_alloc(mp, &tres, 0, 0, 0, &tp); + if (error) { + do_error( +_("Transaction allocation for csum repair failed: %d\n"), + error); + return error; + } + + error = libxfs_buf_read(mp->m_ddev_targp, csum_daddr, + BTOBB(mp->m_rtcsum_bsize), 0, &csum_bp, + &xfs_rtcsum_buf_ops); + if (error) { + do_error( +_("Cannot get buffer for new csum file for RG %u, error %d\n"), + rtg_rgno(rtg), error); + return error; + } + + libxfs_trans_bjoin(tp, csum_bp); + xfs_trans_buf_set_type(tp, csum_bp, XFS_BLFT_RTCSUM_BUF); + + csum_buf = csum_bp->b_addr + boff; + for (i = 0; i < *nr; i++) { + union xfs_csum csum; + + xfs_csum_seed(mp, &csum); + xfs_csum_gen(mp, data_buf + XFS_FSB_TO_B(mp, i), + mp->m_sb.sb_blocksize, &csum); + if (i == 0) + printf("calculated checksum 0x%x\n", csum.crc32c); + xfs_csum_finalize(mp, csum_buf, &csum); + csum_buf += (1u << mp->m_rtcsum_shift); + } + + libxfs_trans_log_buf(tp, csum_bp, boff, + boff + xfs_extlen_to_rtcsum_len(mp, *nr) - 1); + return -libxfs_trans_commit(tp); +} + +void +calculate_rtgroup_csums( + struct xfs_rtgroup *rtg) +{ + struct xfs_mount *mp = rtg_mount(rtg); + xfs_rgblock_t rgbno; + int error; + void *data_buf; + + error = posix_memalign(&data_buf, sysconf(_SC_PAGESIZE), + XFS_FSB_TO_B(mp, DATA_BLOCKS_PER_BUF)); + if (error) { + do_warn( +_("Failed to allocate memory for csum rebuild\n")); + return; + } + for (rgbno = 0; rgbno < rtg_blocks(rtg); rgbno += DATA_BLOCKS_PER_BUF) { + unsigned int nr_blocks, done = 0; + + nr_blocks = min(DATA_BLOCKS_PER_BUF, rtg_blocks(rtg) - rgbno); + error = read_data(mp->m_rtdev_targp, + xfs_gbno_to_daddr(rtg_group(rtg), rgbno), + XFS_FSB_TO_BB(mp, nr_blocks), + data_buf); + if (error) { + do_warn( +_("Reading data at RG %u/%u failed, error %d"), + rtg_rgno(rtg), rgbno, error); + continue; + } + + do { + unsigned int n = nr_blocks - done; + + calculate_buf_csums(rtg, + data_buf + XFS_FSB_TO_B(mp, done), + rgbno + done, &n); + done += n; + } while (done < nr_blocks); + } + + kfree(data_buf); +} -- 2.53.0