Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Qu Wenruo <quwenruo.btrfs@gmx.com>
To: c.f.a@posteo.de, linux-btrfs@vger.kernel.org
Subject: Re: [BUG] qgroup: btrfs_remove_qgroup() can leave an in-memory qgroup, without on-disk items, every later rescan gets cancelled
Date: Thu, 1 Oct 2026 07:19:48 +0930	[thread overview]
Message-ID: <9a46521c-8bb0-403e-b587-d8088020a0ea@gmx.com> (raw)
In-Reply-To: <14846642-85d3-4afb-ad99-1b74b22f5f90@posteo.de>



在 2026/9/30 18:52, c.f.a@posteo.de 写道:
> Hi,
> 
[...]
> 
> Analysis
> --------
> btrfs_remove_qgroup() (qgroup.c:1806) does, in this order:
> 
>    1839  del_qgroup_item()             - deletes the INFO and LIMIT items
>    1843  loop over qgroup->groups:
>    1846    __del_qgroup_relation()
>    1848    if (ret) goto out;
>    1889  del_qgroup_rb()               - only reached on success
>    1897  btrfs_sysfs_del_one_qgroup()
> 
> __del_qgroup_relation() (qgroup.c:1628) returns the result of
> quick_update_accounting() (qgroup.c:1674) when the relation was found.
> quick_update_accounting() (qgroup.c:1540) starts with ret = 1 and only
> clears it when the member's excl == rfer (1549); otherwise it sets
> BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT and returns 1. So 1 means "needs a
> rescan", not failure, but btrfs_remove_qgroup() treats it as an error and
> bails out after the items have already been deleted. The transaction is
> not aborted (btrfs_ioctl_qgroup_create() just ends it, ioctl.c:3737ff), so
> the item deletion is committed while the struct btrfs_qgroup stays in the
> rbtree and in sysfs. Its relation to the parent is already gone from the
> rbtree (del_relation_rb() at 1673), which is why a later destroy succeeds
> in removing it but reports -ENOENT from __del_qgroup_relation().

Thanks for the report, I'll fix it soon by not treating >0 as an error.

And also changes those older direct INCONSISTENT bit setting to the 
newer helpers.

Thanks,
Qu

> 
> The same happens if __del_qgroup_relation() returns -ENOENT because
> neither relation item exists and the relation was not found.
> 
> Once the stale qgroup gets dirtied (a rescan zeroes and dirties all
> qgroups in qgroup_rescan_zero_tracking(), qgroup.c:4028),
> btrfs_run_qgroups() fails in update_qgroup_info_item() and
> update_qgroup_limit_item() with -ENOENT (qgroup.c:3144-3151) and calls
> qgroup_mark_inconsistent(), which sets
> BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN (qgroup.c:391-393), so
> rescan_should_stop() (qgroup.c:3839) stops the rescan worker. Because the
> qgroups are already inconsistent at that point, qgroup_mark_inconsistent()
> does not print its reason, which makes this hard to diagnose: the only
> visible message is "qgroup scan paused".
> 
> How it gets triggered
> ---------------------
> Any removal of a level-0 qgroup that is a member of a higher-level qgroup
> while the member's rfer != excl. One path that needs no user action:
> btrfs_drop_snapshot() calls btrfs_qgroup_cleanup_dropped_subvolume()
> (extent-tree.c:6498), which commits and then calls btrfs_remove_qgroup().
> If the qgroups are inconsistent at that time, the dropped subvolume's
> numbers are stale and non-zero, quick_update_accounting() returns 1, and
> the caller only warns on ret < 0, so the stale qgroup appears silently.
> 
> With snapper this is easy to hit: creating a snapshot into a group
> qgroup marks the qgroups inconsistent as soon as the group shares any
> extent with a subvolume outside it ("qgroup inherit needs a rescan"), and
> snapper's cleanup then deletes old snapshots before a rescan has run. On
> my machine the first occurrence was exactly that sequence (snapshot at
> 14:00 marked the qgroups inconsistent, snapper-cleanup at 14:35 deleted
> snapshots, every rescan afterwards was cancelled).
> 
> Rescans are being cancelled the same way again since this morning, after
> snapper-cleanup deleted four snapshots at 06:59 (followed by its usual
> 'btrfs qgroup clear-stale'). As far as I can tell the qgroups were
> consistent at that time, so this would be a different trigger. I have
> not traced this occurrence yet; sysfs shows three level-0 qgroups with
> rfer = excl = 0 as candidates.
> 
> Reproducer
> ----------
> The script below (loop device, needs root) follows the snapper pattern:
> create a snapshot with '-i 1/100' from a subvolume outside 1/100, which
> goes through qgroup_mark_inconsistent() ("qgroup inherit needs a rescan")
> and thereby stops accounting, then delete that snapshot, wait for the
> drop and run a rescan.
> 
> Note that making the qgroups inconsistent via 'qgroup assign --no-rescan'
> instead does not reproduce it: quick_update_accounting() only sets
> BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT, accounting continues, the dropped
> snapshot's numbers reach zero and the removal succeeds.
> 
> Output on 7.2.7 (trimmed):
> 
>    == setup: subvolume a, snapshot s created into parent qgroup 1/100
>    snapshot s = 0/257, inconsistent = 1
>    BTRFS warning (device loop2): qgroup marked inconsistent, qgroup 
> inherit needs a rescan
>    Qgroupid    Referenced    Exclusive Parent   Child   Path
>    0/256         64.02MiB     16.00KiB -        -       a
>    0/257         64.02MiB     16.00KiB 1/100    -       s
>    1/100            0.00B        0.00B -        0/257   <0 member qgroups>
> 
>    == delete s and wait until it is dropped
>    Subvolume id 257 is gone (1/1)
>    0/257 in sysfs (in memory):  yes
>    0/257 in quota tree (on disk): no
> 
>    == destroy 0/257 by hand, rescan again
>    ERROR: unable to destroy quota group: No such file or directory
>    0/257 in sysfs: no
> 
> So after the drop, 0/257 is half-removed exactly as on my real
> filesystem. On this tiny filesystem the rescan itself still finishes
> (presumably because it is done within a single commit); on the real one 
> (~700 GiB,
> ~80 subvolumes) it is cancelled at the first transaction commit every
> time, as described above.
> 
> Possible fixes
> --------------
> - In btrfs_remove_qgroup(), do not treat a positive return of
>    __del_qgroup_relation() as an error (the qgroups are already flagged
>    inconsistent), and tolerate -ENOENT for a relation that no longer
>    exists; or
> - remove the relations before deleting the INFO/LIMIT items, so that a
>    failure leaves the qgroup intact instead of half-deleted.
> 
> Workaround
> ----------
> Destroy the qgroup by hand ('btrfs qgroup destroy 0/<id>', ignoring
> ENOENT) or remount; the stale qgroups can be found by comparing
> /sys/fs/btrfs/<uuid>/qgroups/0_* with 'btrfs subvolume list -a'.
> 
> Thanks,
> Carl Ahring
> 
> 
> 
> ---8<--- btrfs-qgroup-repro.sh ---8<---
> #!/usr/bin/env bash
> # Reproducer for a half-removed qgroup: deleting a snapshot whose level-0
> # qgroup is a member of a parent qgroup while qgroup accounting is off
> # (inconsistent after a snapshot inherit that needs a rescan).
> # Works on a throwaway loop device, touches nothing else. Needs root.
> set -eu
> 
> img=$(mktemp /tmp/qg-repro.XXXXXX.img)
> mnt=$(mktemp -d /tmp/qg-repro.XXXXXX)
> truncate -s 2G "$img"
> dev=$(losetup -f --show "$img")
> cleanup() { umount "$mnt" 2>/dev/null || true; losetup -d "$dev"; rm -f 
> "$img"; rmdir "$mnt"; }
> trap cleanup EXIT
> 
> mkfs.btrfs -q -f "$dev"
> mount "$dev" "$mnt"
> uuid=$(findmnt -no UUID "$mnt")
> q=/sys/fs/btrfs/$uuid/qgroups
> step() { printf '\n== %s\n' "$1"; }
> 
> step "setup: subvolume a, snapshot s created into parent qgroup 1/100"
> btrfs quota enable "$mnt"
> btrfs subvolume create "$mnt/a" >/dev/null
> dd if=/dev/urandom of="$mnt/a/f" bs=1M count=64 status=none
> sync
> btrfs quota rescan -w "$mnt" >/dev/null
> btrfs qgroup create 1/100 "$mnt"
> # What snapper does with QGROUP=1/100. a is not in 1/100, so the kernel
> # cannot inherit quickly and calls qgroup_mark_inconsistent() ("qgroup
> # inherit needs a rescan"), which also sets NO_ACCOUNTING: from now on the
> # numbers of s stay frozen at rfer = 64 MiB, excl = 16 KiB.
> btrfs subvolume snapshot -i 1/100 "$mnt/a" "$mnt/s" >/dev/null
> sync
> sid=$(btrfs inspect-internal rootid "$mnt/s")
> echo "snapshot s = 0/$sid, inconsistent = $(cat "$q/inconsistent")"
> dmesg | grep -i 'btrfs' | tail -1
> btrfs qgroup show -pc "$mnt"
> 
> step "delete s and wait until it is dropped"
> btrfs subvolume delete "$mnt/s" >/dev/null
> btrfs subvolume sync "$mnt"
> sleep 2   # btrfs_qgroup_cleanup_dropped_subvolume() runs right after 
> the drop
> sync
> echo "0/$sid in sysfs (in memory):  $([ -e "$q/0_$sid" ] && echo yes || 
> echo no)"
> echo "0/$sid in quota tree (on disk): $(btrfs qgroup show "$mnt" 2>/dev/ 
> null | awk '{print $1}' | grep -qx "0/$sid" && echo yes || echo no)"
> 
> step "rescan"
> start=$(date +%s)
> btrfs quota rescan -w "$mnt" || true
> echo "returned after $(( $(date +%s) - start ))s, inconsistent = $(cat 
> "$q/inconsistent")"
> dmesg | grep -i 'btrfs' | tail -4
> 
> step "destroy 0/$sid by hand, rescan again"
> btrfs qgroup destroy "0/$sid" "$mnt" || true
> echo "0/$sid in sysfs: $([ -e "$q/0_$sid" ] && echo yes || echo no)"
> btrfs quota rescan -w "$mnt" || true
> echo "inconsistent = $(cat "$q/inconsistent")"
> dmesg | grep -i 'btrfs' | tail -2


      reply	other threads:[~2026-09-30 21:49 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  9:22 [BUG] qgroup: btrfs_remove_qgroup() can leave an in-memory qgroup, without on-disk items, every later rescan gets cancelled c.f.a
2026-09-30 21:49 ` Qu Wenruo [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=9a46521c-8bb0-403e-b587-d8088020a0ea@gmx.com \
    --to=quwenruo.btrfs@gmx.com \
    --cc=c.f.a@posteo.de \
    --cc=linux-btrfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox