Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Boris Burkov <boris@bur.io>
To: Qu Wenruo <wqu@suse.com>
Cc: linux-btrfs@vger.kernel.org, kernel-team@fb.com
Subject: Re: [PATCH] btrfs: keep unused block groups queued when a pass fails
Date: Wed, 16 Sep 2026 16:04:47 -0700	[thread overview]
Message-ID: <20260916230447.GA1706558@zen.localdomain> (raw)
In-Reply-To: <04e6331d-6b7a-4a2f-9eca-8aa7f0e4042a@suse.com>

On Thu, Sep 17, 2026 at 08:21:37AM +0930, Qu Wenruo wrote:
> 
> 
> 在 2026/9/17 07:47, Boris Burkov 写道:
> > Once any block_group sets ret!=0 in the main loop of
> > btrfs_delete_unused_bgs(), the check
> >    if (ret || btrfs_mixed_space_info(space_info)) {
> >            btrfs_put_block_group(block_group);
> >            continue;
> >    }
> > skips the rest of the unused bgs while unlinking them from
> > fs_info->unused_bgs. There is no "level triggered" re-queueing of empty
> > block groups onto fs_info->unused_bgs so it is possible to leak quite a
> > bit of space this way and unless we happen to get a balance or
> > re-use/re-empty one of these bgs, they are leaked for good, which can
> > lead to a spurious enospc later.
> > 
> > While I have observed such leaked blocked groups that are empty but not
> > on the unused_bgs list on production systems, I have not observed that
> > it is definitely due to this issue. I also reproduced this behavior by
> > injecting an ENOSPC error from btrfs_start_trans_remove_block_group
> > which can also fail with ENOMEM, so this feels like a legitimate
> > injection point.
> 
> Do have happen to know which error caused this non-zero @ret?

I do not have a trace or any other error log in dmesg to conclude where
it came from. I just have boxes with a lot of empty bgs not linked on
the unused list and inferred this mechanism.

> 
> I did a quick glance into the loop, it looks like it's not that easy to get
> a non-zero @ret:
> 
> - inc_block_group_ro() failure
>   @ret is reset to 0, so not this path.
> 
> - btrfs_zone_finish()
>   I guess meta is not deploying zoned btrfs in production.
> 
> - btrfs_star_trans_remove_block_group()
>   This can return -ENOSPC, especially considering we have just marked
>   one bg read-only, thus even stealing from global rsv, we may still
>   fail with ENOSPC here.

I suspect ENOSPC or ENOMEM from this one, personally. Agreed with the
rest of your points on the other possible error sites.

> 
>   Although I'd say, that means the inc_block_group_ro() checks are not
>   doing the correct reserved space checking, and that may be the real
>   problem.
> 
> - btrfs_remove_chunk()
>   If it failed, the trans is already aborted.
> 
> > 
> > To fix it, instead of checking ret in the loop, just break out of the
> > loop when ret != 0. Also, link the bg to the retry list at the
> > individual failure sites so that the failing bg is not leaked.
> > 
> > Assisted-by: LLM (reproducer/error injection)
> > Signed-off-by: Boris Burkov <boris@bur.io>
> 
> Otherwise the handling looks correct to me, doing the proper handling on
> error, other than delaying it to the next iteration.
> 
> Reviewed-by: Qu Wenruo <wqu@suse.com>
> 
> Thanks,
> Qu
> > ---
> >   fs/btrfs/block-group.c | 7 ++++++-
> >   1 file changed, 6 insertions(+), 1 deletion(-)
> > 
> > diff --git a/fs/btrfs/block-group.c b/fs/btrfs/block-group.c
> > index ee182369254c..2eb09c9901c9 100644
> > --- a/fs/btrfs/block-group.c
> > +++ b/fs/btrfs/block-group.c
> > @@ -1612,7 +1612,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info)
> >   		space_info = block_group->space_info;
> > -		if (ret || btrfs_mixed_space_info(space_info)) {
> > +		if (btrfs_mixed_space_info(space_info)) {
> >   			btrfs_put_block_group(block_group);
> >   			continue;
> >   		}
> > @@ -1727,6 +1727,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info)
> >   		ret = inc_block_group_ro(block_group, false);
> >   		up_write(&space_info->groups_sem);
> >   		if (ret < 0) {
> > +			btrfs_link_bg_list(block_group, &retry_list);
> >   			ret = 0;
> >   			goto next;
> >   		}
> > @@ -1749,6 +1750,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info)
> >   						     block_group->start);
> >   		if (IS_ERR(trans)) {
> >   			btrfs_dec_block_group_ro(block_group);
> > +			btrfs_link_bg_list(block_group, &retry_list);
> >   			ret = PTR_ERR(trans);
> >   			goto next;
> >   		}
> > @@ -1759,6 +1761,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info)
> >   		 */
> >   		if (!clean_pinned_extents(trans, block_group)) {
> >   			btrfs_dec_block_group_ro(block_group);
> > +			btrfs_link_bg_list(block_group, &retry_list);
> >   			goto end_trans;
> >   		}
> > @@ -1845,6 +1848,8 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info)
> >   next:
> >   		btrfs_put_block_group(block_group);
> >   		spin_lock(&fs_info->unused_bgs_lock);
> > +		if (ret)
> > +			break;
> >   	}
> >   	list_splice_tail(&retry_list, &fs_info->unused_bgs);
> >   	spin_unlock(&fs_info->unused_bgs_lock);
> 

      reply	other threads:[~2026-09-16 23:04 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 22:17 [PATCH] btrfs: keep unused block groups queued when a pass fails Boris Burkov
2026-09-16 22:51 ` Qu Wenruo
2026-09-16 23:04   ` Boris Burkov [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260916230447.GA1706558@zen.localdomain \
    --to=boris@bur.io \
    --cc=kernel-team@fb.com \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=wqu@suse.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox