From: Qu Wenruo <wqu@suse.com>
To: Zhang Boyang <zhangboyang.id@gmail.com>,
Qu Wenruo <quwenruo.btrfs@gmx.com>,
linux-btrfs@vger.kernel.org
Cc: David Sterba <dsterba@suse.com>, Filipe Manana <fdmanana@kernel.org>
Subject: Re: [BUG] two raid consistency bugs
Date: Wed, 15 Jul 2026 18:06:21 +0930 [thread overview]
Message-ID: <8248f43d-d6ab-47e4-9ed9-65bb5c281f4b@suse.com> (raw)
In-Reply-To: <8e02253e-2bd2-4ca1-8def-f91afaac4cd6@gmail.com>
在 2026/7/15 17:08, Zhang Boyang 写道:
> Hi,
>
> On 2026/7/15 14:10, Qu Wenruo wrote:
>>
>>
>> 在 2026/7/15 15:16, Zhang Boyang 写道:
>>> Hi,
>>>
>>> On 2026/7/15 05:40, Qu Wenruo wrote:
>>>>> At first power failure during transaction N, metadata trees of
>>>>> generation N are written to disk A, but super is not committed.
>>>>> Nothing
>>>>> is written to disk B.
>>>>>
>>>>> At second power failure during a different transaction N,
>>>>
>>>> If it's a different transaction, why it will still have the same
>>>> transid N?
>>>>
>>>
>>> Because the power failure occurred just before writing superblock
>>> (transid N). The superblock on disk still has transid N-1.
>>>
>>> After reboot, looking at superblock which transid is N-1, btrfs has
>>> no idea of transid N existed previously, so it uses transid N for new
>>> transcation.
>>
>> Then there should be no problem at least at the next mount after the
>> power loss.
>>
>>>
>>>>> nothing is
>>>>> written to disk A, but metadata trees and super is committed to
>>>>> disk B.
>>>>>
>>>>> This creates a ambiguous generation N in two disks. Currently btrfs
>>>>> can't detect this, and can lead to severe damages.
>>
>> At the next mount, btrfs should detect device B has the latest super
>> block, and use that as the super block to mount.
>>
>> Since metadata are all written to device B, even disk A may have some
>> stale tree blocks with transid N, stale tree blocks still need to meet
>> other conditions like root owner, level, first key checks.
>>
>> I won't say that's impossible, and won't say we shouldn't do anything
>> to address it, but this is a variant of the split brain problems
>> mentioned in the past.
>>
>
> I think this is a very special variant of split brain problem. This bug
> can occur even if no degraded mounts are involved. All devices are
> presented to btrfs at every mounts. Personally I don't like to call this
> bug as a split brain problem because I think split brain should only
> related to degraded mounts.
>
>> I strongly recommend to find out that thread and check if any of the
>> ideas are explored before and if they have their limits.
>>
>
> I did read some of these threads. If split brain problems are solved,
> this bug can be solved as well. But I think fixing this particular bug
> is also beneficial.
I do not agree.
You're introducing more and more code just to handle some very specific
corner cases.
E.g. for your particular case, you will need a very specific write
situation. You're introducing a feature that is very hard to hit under
most situations, but we will always bear the burden.
I am not even sure if you'll still contribute in the next 5 years, thus
I won't bet my 5 cents on that this feature will be properly maintained.
To me, if you really bother this particular situation for whatever
reason, just introduce a special harden mount option, that any
barrier/super block write failure will mark the fs error.
That will be a much safer bet than any of your proposal.
next prev parent reply other threads:[~2026-07-15 8:36 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-14 16:10 [BUG] two raid consistency bugs Zhang Boyang
2026-07-14 16:10 ` [PATCH 1/2] fstests: btrfs/348: test ambiguous generation handling on raid1 profile Zhang Boyang
2026-07-14 16:10 ` [PATCH 2/2] fstests: btrfs/349: test if latest tree-log is choosen at mount time " Zhang Boyang
2026-07-14 21:40 ` [BUG] two raid consistency bugs Qu Wenruo
2026-07-15 5:46 ` Zhang Boyang
2026-07-15 6:10 ` Qu Wenruo
2026-07-15 7:38 ` Zhang Boyang
2026-07-15 8:36 ` Qu Wenruo [this message]
2026-07-15 9:25 ` Zhang Boyang
2026-07-15 9:44 ` Qu Wenruo
2026-07-15 10:07 ` Zhang Boyang
2026-07-15 10:11 ` Qu Wenruo
2026-07-15 10:41 ` Zhang Boyang
2026-07-15 13:11 ` Alan Huang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=8248f43d-d6ab-47e4-9ed9-65bb5c281f4b@suse.com \
--to=wqu@suse.com \
--cc=dsterba@suse.com \
--cc=fdmanana@kernel.org \
--cc=linux-btrfs@vger.kernel.org \
--cc=quwenruo.btrfs@gmx.com \
--cc=zhangboyang.id@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox