From: Russell Coker <russell@coker.com.au>
To: linux-btrfs@vger.kernel.org
Subject: Re: btrfs dev sta not updating
Date: Tue, 23 Jun 2020 18:00:05 +1000 [thread overview]
Message-ID: <5752066.y3qnG1rHMR@liv> (raw)
In-Reply-To: <08121825-9c05-87c4-4015-f6f508193fa8@suse.com>
On Tuesday, 23 June 2020 4:03:37 PM AEST Nikolay Borisov wrote:
> > I have a USB stick that's corrupted, I get the above kernel messages when
> > I
> > try to copy files from it. But according to btrfs dev sta it has had 0
> > read and 0 corruption errors.
> >
> > root@xev:/mnt/tmp# btrfs dev sta .
> > [/dev/sdc1].write_io_errs 0
> > [/dev/sdc1].read_io_errs 0
> > [/dev/sdc1].flush_io_errs 0
> > [/dev/sdc1].corruption_errs 0
> > [/dev/sdc1].generation_errs 0
> > root@xev:/mnt/tmp# uname -a
> > Linux xev 5.6.0-2-amd64 #1 SMP Debian 5.6.14-1 (2020-05-23) x86_64
> > GNU/Linux
> The read/write io err counters are updated when even repair bio have
> failed. So in your case you had some checksum errors, but btrfs managed
> to repair them by reading from a different mirror. In this case those
> aren't really counted as io errs since in the end you did get the
> correct data.
In this case I'm getting application IO errors and lost data, so if the error
count is designed to not count recovered errors then it's still not doing the
right thing.
# md5sum *
md5sum: 'Rise of the Machines S1 Ep6 - Mega Digger-qcOpMtIWsrgN.mp4': Input/
output error
md5sum: 'Rise of the Machines S1 Ep7 - Ultimate Dragster-Ke9hMplfQAWF.mp4':
Input/output error
md5sum: 'Rise of the Machines S1 Ep8 - Aircraft Carrier-Qxht6qMEwMKr.mp4':
Input/output error
^C
# btrfs dev sta .
[/dev/sdc1].write_io_errs 0
[/dev/sdc1].read_io_errs 0
[/dev/sdc1].flush_io_errs 0
[/dev/sdc1].corruption_errs 0
[/dev/sdc1].generation_errs 0
# tail /var/log/kern.log
Jun 23 17:48:40 xev kernel: [417603.547748] BTRFS warning (device sdc1): csum
failed root 5 ino 275 off 59580416 csum 0x8941f998 expected csum 0xb5b581fc
mirror 1
Jun 23 17:48:40 xev kernel: [417603.609861] BTRFS warning (device sdc1): csum
failed root 5 ino 275 off 60628992 csum 0x8941f998 expected csum 0x4b6c9883
mirror 1
Jun 23 17:48:40 xev kernel: [417603.672251] BTRFS warning (device sdc1): csum
failed root 5 ino 275 off 61677568 csum 0x8941f998 expected csum 0x89f5fb68
mirror 1
# uname -a
Linux xev 5.6.0-2-amd64 #1 SMP Debian 5.6.14-1 (2020-05-23) x86_64 GNU/Linux
On Tuesday, 23 June 2020 4:17:55 PM AEST waxhead wrote:
> I don't think this is what most people expect.
> A simple way to solve this could be to put the non-fatal errors in
> parentheses if this can be done easily.
I think that most people would expect a "device stats" command to just give
stats of the device and not refer to what happens at the higher level. If a
device is giving corruption or read errors then "device stats" should tell
that.
On Tuesday, 23 June 2020 5:11:05 PM AEST Nikolay Borisov wrote:
> read_io_errs. But this leads to a different can of worms - if a user
> sees read_io_errs should they be worried because potentially some data
> is stale or not (give we won't be distinguishing between unrepairable vs
> transient ones).
If a user sees errors reported their degree of worry should be based on the
degree to which they use RAID and have decent backups. If you have RAID-1 and
only 1 device has errors then you are OK. If you have 2 devices with errors
then you have a problem.
Below is an example of a zpool having errors that were corrected. The DEVICE
had an unrecoverable error, but the RAID-Z pool recovered it from other
devices. It states that "Applications are unaffected" so the user knows the
degree of worry that should be involved.
# zpool status
pool: pet630
state: ONLINE
status: One or more devices has experienced an unrecoverable error. An
attempt was made to correct the error. Applications are unaffected.
action: Determine if the device needs to be replaced, and clear the errors
using 'zpool clear' or replace the device with 'zpool replace'.
see: http://zfsonlinux.org/msg/ZFS-8000-9P
scan: scrub repaired 380K in 156h39m with 0 errors on Sat Jun 20 13:03:26
2020
config:
NAME STATE READ WRITE CKSUM
pet630 ONLINE 0 0 0
raidz1-0 ONLINE 0 0 0
sdf ONLINE 0 0 0
sdq ONLINE 0 0 0
sdd ONLINE 0 0 0
sdh ONLINE 0 0 0
sdi ONLINE 41 0 1
--
My Main Blog http://etbe.coker.com.au/
My Documents Blog http://doc.coker.com.au/
next prev parent reply other threads:[~2020-06-23 8:00 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2020-06-23 2:09 btrfs dev sta not updating Russell Coker
2020-06-23 6:03 ` Nikolay Borisov
2020-06-23 6:17 ` waxhead
2020-06-23 7:11 ` Nikolay Borisov
2020-06-23 8:00 ` Russell Coker [this message]
2020-06-23 8:17 ` Nikolay Borisov
2020-06-23 9:48 ` Russell Coker
2020-06-23 11:13 ` Nikolay Borisov
2020-06-23 11:21 ` Russell Coker
2020-06-24 11:39 ` Zygo Blaxell
2020-06-24 13:04 ` Nikolay Borisov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=5752066.y3qnG1rHMR@liv \
--to=russell@coker.com.au \
--cc=linux-btrfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox