Linux Btrfs filesystem development
 help / color / mirror / Atom feed
* btrfs: spurious fs-verity "FILE CORRUPTED" with large data folios (v7.2+)
@ 2026-10-10  8:38 Tito Duarte
  2026-10-10  9:02 ` Qu Wenruo
  0 siblings, 1 reply; 2+ messages in thread
From: Tito Duarte @ 2026-10-10  8:38 UTC (permalink / raw)
  To: linux-btrfs; +Cc: fsverity

Hi,

Since v7.2, reading fs-verity files on btrfs intermittently logs fs-verity
verification failures although the data is intact:

  fs-verity (loop0p2, inode 2980): FILE CORRUPTED! pos=1355776, level=-1,
    want_hash=sha512:8c662f97..., real_hash=sha512:2d23913d3759ef01...

The read is retried and userspace gets correct data, btrfs logs no csum
error, and a file without fs-verity on the same filesystem always reads
back correctly. Both SHA-256 and SHA-512 files are affected.

Cause
-----

end_folio_read() -> btrfs_verify_folio() calls fsverity_verify_folio(),
which verifies the whole folio; start/len are only used for the
uptodate check (fs/btrfs/extent_io.c in v7.2, unchanged in master as of
2026-10-07). With large data folios, a folio can span two extents and is
then filled by two bios. When the first bio completes, the whole folio is
hashed, including blocks the second bio has not read yet.

Large data folios came with cc38d178ff33 (v6.17-rc1, behind
CONFIG_BTRFS_EXPERIMENTAL) and became the default with 9bce95edb1b4
("btrfs: move large data folios out of experimental features", v7.2).
Neither commit is wrong in itself; they expose the whole-folio verification.
In v6.12 a data folio is one page, so whole-folio and per-block
verification are the same thing.

Evidence
--------

1. Natural occurrence, arm64 VM (Apple Virtualization, virtio-blk,
   4K pages), reproducer below, 30 mount/read cycles each:

     6.12.111 (Debian trixie)                     0 reports
     7.2.6    (Debian trixie-backports)          10
     7.2.8    (Debian testing)                   10, twice
     7.3-rc3  (Ubuntu mainline build)            10, twice
     7.2.8    (Debian) + the change below         0, three runs (90 cycles)

   On this machine real_hash was always the hash of a 4 KiB zero block
   and the failing block was always the first block of an extent (10 of
   10 checked). That fits the first bio completing first: the first
   unread block is then the start of the next extent in the folio.

2. Forced reordering, x86_64 QEMU/TCG: delaying every second read bio of
   a verity inode by 20 ms, 10 cycles each:

     9bce95edb1b4^                                  0 reports
     9bce95edb1b4                                 100
     cc38d178ff33^  (BTRFS_EXPERIMENTAL=y)          0
     cc38d178ff33   (BTRFS_EXPERIMENTAL=y)         91
     v7.2                                          99
     v7.2 + the change below                        0

   In all failures (160 logged) the folio had blocks that were not
   uptodate yet, and the failing block lay inside the folio but outside
   the range whose bio had just completed. Here real_hash was not the
   zero-block hash: the unread blocks held stale page contents.

3. Distribution kernel: openSUSE Tumbleweed kernel-default 7.2.8-1.1,
   unmodified, with dm-delay reordering reads: 13 and 12 reports in 10
   cycles. Without the delay x86_64 hits it rarely here (0 to 1 in 120
   cycles under TCG), which I attribute to TCG's timing, not to the
   architecture.

No checksum mismatches and no btrfs errors in any run.

Suggested change
----------------

Verify only the range whose read completed. Each completion then hashes
its own blocks, independent of completion order, and the folio stays
locked and not uptodate until the last range is done.

diff --git a/fs/btrfs/extent_io.c b/fs/btrfs/extent_io.c
index f032f08..6b6cf0e 100644
--- a/fs/btrfs/extent_io.c
+++ b/fs/btrfs/extent_io.c
@@ -482,7 +482,13 @@ static bool btrfs_verify_folio(struct fsverity_info *vi, struct folio *folio,
 
 	if (!vi || btrfs_folio_test_uptodate(fs_info, folio, start, len))
 		return true;
-	return fsverity_verify_folio(vi, folio);
+	/*
+	 * Only verify the range whose read just completed.  With large folios
+	 * the other blocks of the folio may belong to a different extent and
+	 * bio that has not completed yet, so verifying the whole folio here
+	 * would hash blocks that have not been read.
+	 */
+	return fsverity_verify_blocks(vi, folio, len, offset_in_folio(folio, start));
 }
 
 static void end_folio_read(struct fsverity_info *vi, struct folio *folio,

Not covered by my testing:
- Subpage: fsverity_verify_blocks() needs offset and length aligned to
  the Merkle tree block size. With a Merkle block larger than the sector
  size (e.g. 64K pages, 4K sectors, 64K Merkle blocks) a single-sector
  range is not aligned; that case would need verification deferred until
  the whole Merkle block is read.
- Compressed extents and holes (same end_folio_read() path, untested).

If the approach is right I can send it as a proper patch with a Fixes:
tag, or leave it to whoever prefers to handle the subpage case.

Reproducer (root, scratch machine; needs btrfs-progs, fsverity-utils)
---------------------------------------------------------------------

-----8<-----
#!/bin/bash
set -euo pipefail
CYCLES=${1:-30}
T=$(mktemp -d /var/tmp/btrfs-verity-repro.XXXX)
cleanup() { umount "$T/mnt" 2>/dev/null || true; rm -rf "$T"; }
trap cleanup EXIT
mkdir "$T/mnt"; truncate -s 2G "$T/img"; mkfs.btrfs -q "$T/img"; mount -o loop "$T/img" "$T/mnt"
head -c 40M /dev/urandom > "$T/src"
off=0; i=0
while [ $off -lt 10240 ]; do
  n=$(( 97 + (i * 137) % 331 ))          # 97..427 blocks: extent starts not folio-aligned
  for f in a b c; do
    dd if="$T/src" of="$T/mnt/$f" bs=4096 skip=$off seek=$off count=$n conv=notrunc status=none
    sync
  done
  off=$((off + n)); i=$((i + 1))
done
fsverity enable "$T/mnt/a" --hash-alg=sha512
fsverity enable "$T/mnt/b" --hash-alg=sha256   # c stays without fs-verity
sync
want=$(sha256sum < "$T/src" | cut -d' ' -f1)
reports=0; mismatches=0
for c in $(seq 1 "$CYCLES"); do
  umount "$T/mnt"; mount -o loop,ro "$T/img" "$T/mnt"
  dmesg -C; echo 3 > /proc/sys/vm/drop_caches
  for f in a b c; do
    [ "$(sha256sum < "$T/mnt/$f" | cut -d' ' -f1)" = "$want" ] || mismatches=$((mismatches + 1))
  done
  reports=$((reports + $(dmesg | grep -c 'FILE CORRUPTED' || true)))
done
echo "$(uname -r): $CYCLES cycles: $reports fs-verity reports, $mismatches checksum mismatches"
-----8<-----

Disclosure: I used an AI coding assistant (Claude) for parts of the
analysis, the test harness and this write-up. The numbers above are from
real runs on my machines; logs are available on request.

Thanks,
Tito Duarte

^ permalink raw reply related	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-10  9:02 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-10  8:38 btrfs: spurious fs-verity "FILE CORRUPTED" with large data folios (v7.2+) Tito Duarte
2026-10-10  9:02 ` Qu Wenruo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox