Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Emil Heimpel <broetchenrackete@gmail.com>
To: Qu Wenruo <quwenruo.btrfs@gmx.com>
Cc: linux-btrfs@vger.kernel.org
Subject: Re: Errors after successful disk replace
Date: Tue, 19 Oct 2021 12:38:43 +0000 (UTC)	[thread overview]
Message-ID: <0e218606-5e07-4d7d-b81f-519895af2dfa@gmail.com> (raw)
In-Reply-To: <5d8df74b-6094-de7d-1b47-885126cf4bc6@gmx.com>

So it finished after 2 minutes?

[Tue Oct 19 14:13:51 2021] BTRFS info (device sde1): continuing dev_replace from <missing disk> (devid 1) to target /dev/sde1 @74%
[Tue Oct 19 14:15:39 2021] BTRFS info (device sde1): dev_replace from <missing disk> (devid 1) to /dev/sde1 finished


Now I at least have an expected filesystem show:

Label: 'BlueButter'  uuid: 7e3378e6-da46-4a60-b9b8-1bcc306986e3
        Total devices 6 FS bytes used 20.96TiB
        devid    1 size 7.28TiB used 5.46TiB path /dev/sde1
        devid    2 size 7.28TiB used 5.46TiB path /dev/sdb1
        devid    3 size 2.73TiB used 2.73TiB path /dev/sdg1
        devid    4 size 2.73TiB used 2.73TiB path /dev/sdd1
        devid    5 size 7.28TiB used 4.81TiB path /dev/sdf1
        devid    6 size 7.28TiB used 5.33TiB path /dev/sdc1

And a nondegraded remount worked too.

Thanks,
Emil

Oct 19, 2021 14:20:21 Qu Wenruo <quwenruo.btrfs@gmx.com>:

> 
> 
> On 2021/10/19 20:16, Emil Heimpel wrote:
>> Color me suprised:
>> 
>> 
>> [74713.072745] BTRFS info (device sde1): flagging fs with big metadata feature
>> [74713.072755] BTRFS info (device sde1): allowing degraded mounts
>> [74713.072758] BTRFS info (device sde1): using free space tree
>> [74713.072760] BTRFS info (device sde1): has skinny extents
>> [74713.104297] BTRFS warning (device sde1): devid 1 uuid 51645efd-bf95-458d-b5ae-b31623533abb is missing
>> [74714.675001] BTRFS info (device sde1): bdev (efault) errs: wr 52950, rd 8161, flush 0, corrupt 1221, gen 0
>> [74714.675015] BTRFS info (device sde1): bdev /dev/sdb1 errs: wr 0, rd 0, flush 0, corrupt 228, gen 0
>> [74714.675025] BTRFS info (device sde1): bdev /dev/sdc1 errs: wr 0, rd 0, flush 0, corrupt 140, gen 0
>> [74751.033383] BTRFS info (device sde1): continuing dev_replace from <missing disk> (devid 1) to target /dev/sde1 @74%
>> [bluemond@BlueQ ~]$ sudo btrfs replace status  -1 /mnt/btrfsrepair/
>> 74.9% done, 0 write errs, 0 uncorr. read errs
>> 
>> I guess I just wait?
> 
> Yep, wait and stay alert, better to also keep an eye on the dmesg.
> 
> But this also means, previous replace didn't really finish, which may
> mean the replace ioctl is not reporting the proper status, and can be a
> possible bug.
> 
> Thanks,
> Qu
> 
>> 
>> Oct 19, 2021 13:37:09 Qu Wenruo <quwenruo.btrfs@gmx.com>:
>> 
>>> 
>>> 
>>> On 2021/10/19 18:49, Emil Heimpel wrote:
>>>> 
>>>> Oct 19, 2021 07:35:54 Qu Wenruo <quwenruo.btrfs@gmx.com>:
>>>> 
>>>>> 
>>>>> 
>>>>> On 2021/10/19 11:54, Emil Heimpel wrote:
>>>>>> …
>>>>> 
>>>>> Any dmesg of that time?
>>>>> 
>>>> 
>>>> Nothing after the replace finished:
>>>> 
>>>> 1634463961.245751 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663044222976 for dev (efault)
>>>> 1634463961.255819 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663045795840 for dev (efault)
>>>> 1634463961.275815 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663046582272 for dev (efault)
>>>> 1634463961.275922 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663047368704 for dev (efault)
>>>> 1634463961.339074 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663048155136 for dev (efault)
>>>> 1634463961.339248 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663048941568 for dev (efault)
>>> 
>>> *failed*...
>>> 
>>>> 1634475910.611261 BlueQ kernel: sd 9:0:2:0: attempting task abort!scmd(0x0000000046fead3f), outstanding for 7120 ms & timeout 7000 ms
>>>> 1634475910.615126 BlueQ kernel: sd 9:0:2:0: [sdd] tag#840 CDB: ATA command pass through(16) 85 08 2e 00 00 00 01 00 00 00 00 00 00 00 ec 00
>>>> 1634475910.615429 BlueQ kernel: scsi target9:0:2: handle(0x000b), sas_address(0x4433221105000000), phy(5)
>>>> 1634475910.615691 BlueQ kernel: scsi target9:0:2: enclosure logical id(0x590b11c022f3fb00), slot(6)
>>> 
>>> And ATA commands failure.
>>> 
>>> I don't believe the replace finished without problem, and the involved
>>> device is /dev/sdd.
>>> 
>>>> 1634475910.787911 BlueQ kernel: sd 9:0:2:0: task abort: SUCCESS scmd(0x0000000046fead3f)
>>>> 1634475910.807083 BlueQ kernel: sd 9:0:2:0: Power-on or device reset occurred
>>>> 1634475949.877998 BlueQ kernel: sd 9:0:2:0: Power-on or device reset occurred
>>>> 1634525944.213931 BlueQ kernel: perf: interrupt took too long (3138 > 3137), lowering kernel.perf_event_max_sample_rate to 63600
>>>> 1634533791.168760 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 22996545634304 for dev (efault)
>>>> 1634552685.203559 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 23816815706112 for dev (efault)
>>> 
>>> You won't want to see this message at all.
>>> 
>>> This means, you're running RAID56, as btrfs has write-hole problem,
>>> which will degrade the robust of RAID56 byte by byte for each unclean
>>> shutdown.
>>> 
>>> I guess the write hole problem has already make the repair failed for
>>> the replace.
>>> 
>>> Thus after a successful mount, scrub and manually file checking is
>>> almost a must.
>>> 
>>>> 1634558977.979621 BlueQ kernel: BTRFS info (device sdb1): dev_replace from <missing disk> (devid 1) to /dev/sde1 finished
>>>> 1634560793.132731 BlueQ kernel: zram0: detected capacity change from 32610864 to 0
>>>> 1634560793.169379 BlueQ kernel: zram: Removed device: zram0
>>>> 1634560883.549481 BlueQ kernel: watchdog: watchdog0: watchdog did not stop!
>>>> 1634560883.556038 BlueQ systemd-shutdown[1]: Syncing filesystems and block devices.
>>>> 1634560883.572840 BlueQ systemd-shutdown[1]: Sending SIGTERM to remaining processes...
>>>> 
>>>> 
>>>> 
>>>> 
>>>>>> …
>>>>> 
>>>>> And dmesg for the failed mount?
>>>>> 
>>>> 
>>>> Oops, I must have missed that it failed because of missing devid 1 too...
>>>> 
>>>> 1634562944.145383 BlueQ kernel: BTRFS info (device sde1): flagging fs with big metadata feature
>>>> 1634562944.145529 BlueQ kernel: BTRFS info (device sde1): force zstd compression, level 2
>>>> 1634562944.145650 BlueQ kernel: BTRFS info (device sde1): using free space tree
>>>> 1634562944.145697 BlueQ kernel: BTRFS info (device sde1): has skinny extents
>>>> 1634562944.148709 BlueQ kernel: BTRFS error (device sde1): devid 1 uuid 51645efd-bf95-458d-b5ae-b31623533abb is missing
>>>> 1634562944.148764 BlueQ kernel: BTRFS error (device sde1): failed to read chunk tree: -2
>>>> 1634562944.185369 BlueQ kernel: BTRFS error (device sde1): open_ctree failed
>>> 
>>> This doesn't sound correct.
>>> 
>>> If a device is properly replaced, it should have the same devid number.
>>> 
>>> I guess you have tried to add a new device before, and then tried to
>>> replace the missing device, right?
>>> 
>>> 
>>> Anyway, have you tried to mount it degraded and then remove the missing
>>> device?
>>> 
>>> Since you're using RAID56, I guess degrade mount should work.
>>> 
>>> Thanks,
>>> Qu
>>> 
>>>> 
>>>>> Thanks,
>>>>> Qu
>>>>>> …
>>>> 

  reply	other threads:[~2021-10-19 12:38 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2021-10-19  3:54 Errors after successful disk replace Emil Heimpel
2021-10-19  5:35 ` Qu Wenruo
2021-10-19 10:49   ` Emil Heimpel
2021-10-19 11:37     ` Qu Wenruo
2021-10-19 12:10       ` Emil Heimpel
2021-10-19 12:16       ` Emil Heimpel
2021-10-19 12:20         ` Qu Wenruo
2021-10-19 12:38           ` Emil Heimpel [this message]
2021-10-19 12:46             ` Qu Wenruo
2021-10-26 12:16               ` Emil Heimpel
2021-10-26 12:17                 ` Qu Wenruo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0e218606-5e07-4d7d-b81f-519895af2dfa@gmail.com \
    --to=broetchenrackete@gmail.com \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=quwenruo.btrfs@gmx.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox