Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Qu Wenruo <wqu@suse.com>
To: Zhang Boyang <zhangboyang.id@gmail.com>,
	Qu Wenruo <quwenruo.btrfs@gmx.com>,
	linux-btrfs@vger.kernel.org
Cc: David Sterba <dsterba@suse.com>, Filipe Manana <fdmanana@kernel.org>
Subject: Re: [BUG] two raid consistency bugs
Date: Wed, 15 Jul 2026 18:06:21 +0930	[thread overview]
Message-ID: <8248f43d-d6ab-47e4-9ed9-65bb5c281f4b@suse.com> (raw)
In-Reply-To: <8e02253e-2bd2-4ca1-8def-f91afaac4cd6@gmail.com>



在 2026/7/15 17:08, Zhang Boyang 写道:
> Hi,
> 
> On 2026/7/15 14:10, Qu Wenruo wrote:
>>
>>
>> 在 2026/7/15 15:16, Zhang Boyang 写道:
>>> Hi,
>>>
>>> On 2026/7/15 05:40, Qu Wenruo wrote:
>>>>> At first power failure during transaction N, metadata trees of
>>>>> generation N are written to disk A, but super is not committed. 
>>>>> Nothing
>>>>> is written to disk B.
>>>>>
>>>>> At second power failure during a different transaction N,
>>>>
>>>> If it's a different transaction, why it will still have the same 
>>>> transid N?
>>>>
>>>
>>> Because the power failure occurred just before writing superblock 
>>> (transid N). The superblock on disk still has transid N-1.
>>>
>>> After reboot, looking at superblock which transid is N-1, btrfs has 
>>> no idea of transid N existed previously, so it uses transid N for new 
>>> transcation.
>>
>> Then there should be no problem at least at the next mount after the 
>> power loss.
>>
>>>
>>>>> nothing is
>>>>> written to disk A, but metadata trees and super is committed to 
>>>>> disk B.
>>>>>
>>>>> This creates a ambiguous generation N in two disks. Currently btrfs
>>>>> can't detect this, and can lead to severe damages.
>>
>> At the next mount, btrfs should detect device B has the latest super 
>> block, and use that as the super block to mount.
>>
>> Since metadata are all written to device B, even disk A may have some 
>> stale tree blocks with transid N, stale tree blocks still need to meet 
>> other conditions like root owner, level, first key checks.
>>
>> I won't say that's impossible, and won't say we shouldn't do anything 
>> to address it, but this is a variant of the split brain problems 
>> mentioned in the past.
>>
> 
> I think this is a very special variant of split brain problem. This bug 
> can occur even if no degraded mounts are involved. All devices are 
> presented to btrfs at every mounts. Personally I don't like to call this 
> bug as a split brain problem because I think split brain should only 
> related to degraded mounts.
> 
>> I strongly recommend to find out that thread and check if any of the 
>> ideas are explored before and if they have their limits.
>>
> 
> I did read some of these threads. If split brain problems are solved, 
> this bug can be solved as well. But I think fixing this particular bug 
> is also beneficial.

I do not agree.

You're introducing more and more code just to handle some very specific 
corner cases.

E.g. for your particular case, you will need a very specific write 
situation. You're introducing a feature that is very hard to hit under 
most situations, but we will always bear the burden.

I am not even sure if you'll still contribute in the next 5 years, thus 
I won't bet my 5 cents on that this feature will be properly maintained.


To me, if you really bother this particular situation for whatever 
reason, just introduce a special harden mount option, that any 
barrier/super block write failure will mark the fs error.

That will be a much safer bet than any of your proposal.

  reply	other threads:[~2026-07-15  8:36 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-14 16:10 [BUG] two raid consistency bugs Zhang Boyang
2026-07-14 16:10 ` [PATCH 1/2] fstests: btrfs/348: test ambiguous generation handling on raid1 profile Zhang Boyang
2026-07-14 16:10 ` [PATCH 2/2] fstests: btrfs/349: test if latest tree-log is choosen at mount time " Zhang Boyang
2026-07-14 21:40 ` [BUG] two raid consistency bugs Qu Wenruo
2026-07-15  5:46   ` Zhang Boyang
2026-07-15  6:10     ` Qu Wenruo
2026-07-15  7:38       ` Zhang Boyang
2026-07-15  8:36         ` Qu Wenruo [this message]
2026-07-15  9:25           ` Zhang Boyang
2026-07-15  9:44             ` Qu Wenruo
2026-07-15 10:07               ` Zhang Boyang
2026-07-15 10:11                 ` Qu Wenruo
2026-07-15 10:41                   ` Zhang Boyang
2026-07-15 13:11                   ` Alan Huang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=8248f43d-d6ab-47e4-9ed9-65bb5c281f4b@suse.com \
    --to=wqu@suse.com \
    --cc=dsterba@suse.com \
    --cc=fdmanana@kernel.org \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=quwenruo.btrfs@gmx.com \
    --cc=zhangboyang.id@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox