All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Dongjiang Zhu" <zhudongjiang@fygo.io>
To: <linux-btrfs@vger.kernel.org>
Subject: [RFC PATCH 0/5] btrfs: tighten qgroup rescan lifecycle handling
Date: Wed, 16 Sep 2026 11:14:48 +0800	[thread overview]
Message-ID: <cover.1789524388.git.zhudongjiang@fygo.io> (raw)

While fixing the confusing error returned for a qgroup rescan request
in simple quota mode, we found further problems in the qgroup rescan
lifecycle.

The following sequence can make a filesystem hang during mount:

  1. Enable full qgroups and wait for the initial rescan to finish.
  2. Start another qgroup rescan.
  3. Disable qgroups while the rescan worker is running.
  4. Enable simple quotas with "quota enable -s".
  5. Unmount and mount the filesystem again.

The rescan worker stops after quotas are disabled, but preserves RESCAN
as if the scan were only paused. Quota disable waits for the worker to
stop, but the stale bit remains. A subsequent simple quota enable
inherits it and writes ON|SIMPLE_MODE|SCANNING to the new status item.

On the next mount, qgroup_rescan_init() rejects this state. However,
its return value is assigned after the existing error cleanup branch
has already been skipped. The qgroup configuration and its sysfs
kobjects are left behind, so mount failure cleanup waits for a kobject
reference that cannot be released.

Fixing the mount error cleanup makes this state fail cleanly instead
of hanging. Preventing the stale state requires fixing rescan flag
cleanup, but clearing RESCAN alone is not sufficient. Two approaches
were considered:

  A. Clear RESCAN in the rescan worker and setup error paths. This fixes
     the running-worker case above, but leaves a race before worker
     queuing. Quota disable can proceed while rescan setup is still in
     progress, and a subsequent simple quota enable can commit a status
     item containing ON|SIMPLE_MODE|SCANNING. If the old rescan ioctl
     later returns -ENOTCONN and clears RESCAN, it only changes the
     in-memory flag; the committed status item is not updated.

  B. Clear RESCAN directly in quota disable. This prevents the new
     simple quota configuration from inheriting the bit, but does not
     stop the old rescan setup. After quota disable and simple quota
     enable, the old ioctl can still run qgroup_rescan_zero_tracking()
     against the newly created qgroup configuration. This was observed
     in the zero-tracking experiment.

Both approaches leave the same synchronization gap. Rescan
initialization, transaction commit, zero tracking, and worker queuing
are separate steps. Quota disable waits for a running worker, but does
not wait for setup that has not yet marked the worker as running.
Neither flag-clearing approach prevents an old setup from crossing
the quota disable/enable boundary.

Several fixes have addressed individual consequences of this window:

commit 331cd9461412 ("btrfs: fix race between quota enable and quota rescan ioctl")
commit e12496677503 ("btrfs: qgroup: fix race between quota disable and quota rescan ioctl")
commit b7adbf9ada35 ("btrfs: fix race between quota rescan and disable leading to NULL pointer deref")

These fixes protect individual objects and failure paths, but do not
serialize the complete rescan setup with quota mode changes.

This series combines the necessary flag cleanup with serialization of
the complete userspace rescan setup. Take subvol_sem for read across
initialization, transaction commit, zero tracking, and worker queuing;
quota enable and disable already take it for write. This prevents
quota configuration changes during setup. The lock is released after
setup and is not held for the duration of the asynchronous scan.

The series also fixes related result-reporting and remount issues:

Patch 1: Clean up qgroup state and sysfs entries when mount-time rescan initialization fails.
Patch 2: Fix cancellation reporting and emit the final result even when no status-update transaction is available.
Patch 3: Serialize rescan setup with quota enable and disable.
Patch 4: Clear stale RESCAN state in worker and setup failure paths.
Patch 5: Fix rescan stop and resume handling during remount.

Dongjiang Zhu (5):
  btrfs: qgroup: clean up config after rescan resume failure
  btrfs: qgroup: fix rescan result reporting
  btrfs: qgroup: serialize rescan setup with quota changes
  btrfs: qgroup: clear RESCAN in overlooked cleanup paths
  btrfs: qgroup: fix rescan handling during remount

 fs/btrfs/disk-io.c |  3 +--
 fs/btrfs/ioctl.c   |  3 +++
 fs/btrfs/qgroup.c  | 67 ++++++++++++++++++++--------------------------
 fs/btrfs/super.c   |  1 +
 4 files changed, 34 insertions(+), 40 deletions(-)

-- 
2.39.5

             reply	other threads:[~2026-09-16  3:15 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  3:14 Dongjiang Zhu [this message]
2026-09-16  3:14 ` [RFC PATCH 1/5] btrfs: qgroup: clean up config after rescan resume failure Dongjiang Zhu
2026-09-16  3:14 ` [RFC PATCH 2/5] btrfs: qgroup: fix rescan result reporting Dongjiang Zhu
2026-09-16  3:14 ` [RFC PATCH 3/5] btrfs: qgroup: serialize rescan setup with quota changes Dongjiang Zhu
2026-09-16  3:14 ` [RFC PATCH 4/5] btrfs: qgroup: clear RESCAN in overlooked cleanup paths Dongjiang Zhu
2026-09-16  3:14 ` [RFC PATCH 5/5] btrfs: qgroup: fix rescan handling during remount Dongjiang Zhu
2026-09-16  3:21   ` Dongjiang Zhu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cover.1789524388.git.zhudongjiang@fygo.io \
    --to=zhudongjiang@fygo.io \
    --cc=linux-btrfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.