All of lore.kernel.org
 help / color / mirror / Atom feed
From: Emil Heimpel <broetchenrackete@gmail.com>
To: Qu Wenruo <quwenruo.btrfs@gmx.com>
Cc: linux-btrfs@vger.kernel.org
Subject: Re: Errors after successful disk replace
Date: Tue, 19 Oct 2021 12:38:43 +0000 (UTC)	[thread overview]
Message-ID: <0e218606-5e07-4d7d-b81f-519895af2dfa@gmail.com> (raw)
In-Reply-To: <5d8df74b-6094-de7d-1b47-885126cf4bc6@gmx.com>

So it finished after 2 minutes?

[Tue Oct 19 14:13:51 2021] BTRFS info (device sde1): continuing dev_replace from <missing disk> (devid 1) to target /dev/sde1 @74%
[Tue Oct 19 14:15:39 2021] BTRFS info (device sde1): dev_replace from <missing disk> (devid 1) to /dev/sde1 finished


Now I at least have an expected filesystem show:

Label: 'BlueButter'  uuid: 7e3378e6-da46-4a60-b9b8-1bcc306986e3
        Total devices 6 FS bytes used 20.96TiB
        devid    1 size 7.28TiB used 5.46TiB path /dev/sde1
        devid    2 size 7.28TiB used 5.46TiB path /dev/sdb1
        devid    3 size 2.73TiB used 2.73TiB path /dev/sdg1
        devid    4 size 2.73TiB used 2.73TiB path /dev/sdd1
        devid    5 size 7.28TiB used 4.81TiB path /dev/sdf1
        devid    6 size 7.28TiB used 5.33TiB path /dev/sdc1

And a nondegraded remount worked too.

Thanks,
Emil

Oct 19, 2021 14:20:21 Qu Wenruo <quwenruo.btrfs@gmx.com>:

> 
> 
> On 2021/10/19 20:16, Emil Heimpel wrote:
>> Color me suprised:
>> 
>> 
>> [74713.072745] BTRFS info (device sde1): flagging fs with big metadata feature
>> [74713.072755] BTRFS info (device sde1): allowing degraded mounts
>> [74713.072758] BTRFS info (device sde1): using free space tree
>> [74713.072760] BTRFS info (device sde1): has skinny extents
>> [74713.104297] BTRFS warning (device sde1): devid 1 uuid 51645efd-bf95-458d-b5ae-b31623533abb is missing
>> [74714.675001] BTRFS info (device sde1): bdev (efault) errs: wr 52950, rd 8161, flush 0, corrupt 1221, gen 0
>> [74714.675015] BTRFS info (device sde1): bdev /dev/sdb1 errs: wr 0, rd 0, flush 0, corrupt 228, gen 0
>> [74714.675025] BTRFS info (device sde1): bdev /dev/sdc1 errs: wr 0, rd 0, flush 0, corrupt 140, gen 0
>> [74751.033383] BTRFS info (device sde1): continuing dev_replace from <missing disk> (devid 1) to target /dev/sde1 @74%
>> [bluemond@BlueQ ~]$ sudo btrfs replace status  -1 /mnt/btrfsrepair/
>> 74.9% done, 0 write errs, 0 uncorr. read errs
>> 
>> I guess I just wait?
> 
> Yep, wait and stay alert, better to also keep an eye on the dmesg.
> 
> But this also means, previous replace didn't really finish, which may
> mean the replace ioctl is not reporting the proper status, and can be a
> possible bug.
> 
> Thanks,
> Qu
> 
>> 
>> Oct 19, 2021 13:37:09 Qu Wenruo <quwenruo.btrfs@gmx.com>:
>> 
>>> 
>>> 
>>> On 2021/10/19 18:49, Emil Heimpel wrote:
>>>> 
>>>> Oct 19, 2021 07:35:54 Qu Wenruo <quwenruo.btrfs@gmx.com>:
>>>> 
>>>>> 
>>>>> 
>>>>> On 2021/10/19 11:54, Emil Heimpel wrote:
>>>>>> …
>>>>> 
>>>>> Any dmesg of that time?
>>>>> 
>>>> 
>>>> Nothing after the replace finished:
>>>> 
>>>> 1634463961.245751 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663044222976 for dev (efault)
>>>> 1634463961.255819 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663045795840 for dev (efault)
>>>> 1634463961.275815 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663046582272 for dev (efault)
>>>> 1634463961.275922 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663047368704 for dev (efault)
>>>> 1634463961.339074 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663048155136 for dev (efault)
>>>> 1634463961.339248 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 17663048941568 for dev (efault)
>>> 
>>> *failed*...
>>> 
>>>> 1634475910.611261 BlueQ kernel: sd 9:0:2:0: attempting task abort!scmd(0x0000000046fead3f), outstanding for 7120 ms & timeout 7000 ms
>>>> 1634475910.615126 BlueQ kernel: sd 9:0:2:0: [sdd] tag#840 CDB: ATA command pass through(16) 85 08 2e 00 00 00 01 00 00 00 00 00 00 00 ec 00
>>>> 1634475910.615429 BlueQ kernel: scsi target9:0:2: handle(0x000b), sas_address(0x4433221105000000), phy(5)
>>>> 1634475910.615691 BlueQ kernel: scsi target9:0:2: enclosure logical id(0x590b11c022f3fb00), slot(6)
>>> 
>>> And ATA commands failure.
>>> 
>>> I don't believe the replace finished without problem, and the involved
>>> device is /dev/sdd.
>>> 
>>>> 1634475910.787911 BlueQ kernel: sd 9:0:2:0: task abort: SUCCESS scmd(0x0000000046fead3f)
>>>> 1634475910.807083 BlueQ kernel: sd 9:0:2:0: Power-on or device reset occurred
>>>> 1634475949.877998 BlueQ kernel: sd 9:0:2:0: Power-on or device reset occurred
>>>> 1634525944.213931 BlueQ kernel: perf: interrupt took too long (3138 > 3137), lowering kernel.perf_event_max_sample_rate to 63600
>>>> 1634533791.168760 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 22996545634304 for dev (efault)
>>>> 1634552685.203559 BlueQ kernel: BTRFS error (device sdb1): failed to rebuild valid logical 23816815706112 for dev (efault)
>>> 
>>> You won't want to see this message at all.
>>> 
>>> This means, you're running RAID56, as btrfs has write-hole problem,
>>> which will degrade the robust of RAID56 byte by byte for each unclean
>>> shutdown.
>>> 
>>> I guess the write hole problem has already make the repair failed for
>>> the replace.
>>> 
>>> Thus after a successful mount, scrub and manually file checking is
>>> almost a must.
>>> 
>>>> 1634558977.979621 BlueQ kernel: BTRFS info (device sdb1): dev_replace from <missing disk> (devid 1) to /dev/sde1 finished
>>>> 1634560793.132731 BlueQ kernel: zram0: detected capacity change from 32610864 to 0
>>>> 1634560793.169379 BlueQ kernel: zram: Removed device: zram0
>>>> 1634560883.549481 BlueQ kernel: watchdog: watchdog0: watchdog did not stop!
>>>> 1634560883.556038 BlueQ systemd-shutdown[1]: Syncing filesystems and block devices.
>>>> 1634560883.572840 BlueQ systemd-shutdown[1]: Sending SIGTERM to remaining processes...
>>>> 
>>>> 
>>>> 
>>>> 
>>>>>> …
>>>>> 
>>>>> And dmesg for the failed mount?
>>>>> 
>>>> 
>>>> Oops, I must have missed that it failed because of missing devid 1 too...
>>>> 
>>>> 1634562944.145383 BlueQ kernel: BTRFS info (device sde1): flagging fs with big metadata feature
>>>> 1634562944.145529 BlueQ kernel: BTRFS info (device sde1): force zstd compression, level 2
>>>> 1634562944.145650 BlueQ kernel: BTRFS info (device sde1): using free space tree
>>>> 1634562944.145697 BlueQ kernel: BTRFS info (device sde1): has skinny extents
>>>> 1634562944.148709 BlueQ kernel: BTRFS error (device sde1): devid 1 uuid 51645efd-bf95-458d-b5ae-b31623533abb is missing
>>>> 1634562944.148764 BlueQ kernel: BTRFS error (device sde1): failed to read chunk tree: -2
>>>> 1634562944.185369 BlueQ kernel: BTRFS error (device sde1): open_ctree failed
>>> 
>>> This doesn't sound correct.
>>> 
>>> If a device is properly replaced, it should have the same devid number.
>>> 
>>> I guess you have tried to add a new device before, and then tried to
>>> replace the missing device, right?
>>> 
>>> 
>>> Anyway, have you tried to mount it degraded and then remove the missing
>>> device?
>>> 
>>> Since you're using RAID56, I guess degrade mount should work.
>>> 
>>> Thanks,
>>> Qu
>>> 
>>>> 
>>>>> Thanks,
>>>>> Qu
>>>>>> …
>>>> 

  reply	other threads:[~2021-10-19 12:38 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2021-10-19  3:54 Errors after successful disk replace Emil Heimpel
2021-10-19  5:35 ` Qu Wenruo
2021-10-19 10:49   ` Emil Heimpel
2021-10-19 11:37     ` Qu Wenruo
2021-10-19 12:10       ` Emil Heimpel
2021-10-19 12:16       ` Emil Heimpel
2021-10-19 12:20         ` Qu Wenruo
2021-10-19 12:38           ` Emil Heimpel [this message]
2021-10-19 12:46             ` Qu Wenruo
2021-10-26 12:16               ` Emil Heimpel
2021-10-26 12:17                 ` Qu Wenruo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0e218606-5e07-4d7d-b81f-519895af2dfa@gmail.com \
    --to=broetchenrackete@gmail.com \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=quwenruo.btrfs@gmx.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.