CEPH filesystem development
 help / color / mirror / Atom feed
From: Guido Winkelmann <guido-ceph@thisisnotatest.de>
To: Josh Durgin <josh.durgin@inktank.com>
Cc: "ceph-devel@vger.kernel.org" <ceph-devel@vger.kernel.org>
Subject: Re: Random data corruption in VM, possibly caused by rbd
Date: Thu, 07 Jun 2012 23:36:23 +0200	[thread overview]
Message-ID: <1432839.r57HJoU1Hp@tolkien> (raw)
In-Reply-To: <4FD10575.7010300@inktank.com>

On Thursday 07 June 2012 12:48:05 Josh Durgin wrote:
> On 06/07/2012 11:04 AM, Guido Winkelmann wrote:
> > Hi,
> > 
> > I'm using Ceph with RBD to provide network-transparent disk images for
> > KVM-
> > based virtual servers. The last two days, I've been hunting some weird
> > elusive bug where data in the virtual machines would be corrupted in
> > weird ways. It usually manifests in files having some random data -
> > usually zeroes - at the start before the actual contents that should be
> > in there start.
> 
> I definitely want to figure out what's going on with this.
> A few questions:
> 
> Are you using rbd caching? If so, what settings?

I'm not using rbd caching, and I wasn't planning on even trying before I have 
a much better understanding of how it affects VM migration.
 
> In either case, does the corruption still occur if you
> switch caching on/off? There are different I/O paths here,
> and this might tell us if the problem is on the client side.
> 
> Another thing to try is turning off sparse reads on the osd by setting
> filestore fiemap threshold = 0

Okay, I will try these things tomorrow.
 
[...]
> > The ceph cluster uses btrf for the osd's data dirs. The journal is on a
> > tmpfs. (This is not a production setup - luckily.)
> > The virtual machine is using ext4 as its filesystem.
> > There were no obvious other problems with either the ceph cluster or the
> > KVM host machines.
> 
> Were there any nodes with osds restarted during the test runs? I wonder
> if it's a problem with losing the tmpfs journal.

No, from the point when the rbd volume was created, all nodes were online all 
the time. No nodes were added or removed.
 
> As Oliver suggested, switching the osd data dir filesystem might help
> too.

Again, I'll try that tomorrow. BTW, I could use some advice on how to go about 
that. Right I would stop one osd process (not the whole machine), reformat and 
remount its btrfs devices as XFS, delete the journal, restart the osd, wait 
until the cluster is healthy again, repeat for all the osds in the cluster. Is 
that sufficient?

Oh, one other thing I just thought of:
The rbd volume in question was created as a copy, using the rbd cp command, 
from a template volume. I cannot recall seeing any corruption while using the 
original volume (which was created using rbd import). Maybe the bug only bites 
volumes that have been created as copies of other volumes? I'll have to do 
more tests along those lines as well...

Regards,
	Guido


  reply	other threads:[~2012-06-07 21:36 UTC|newest]

Thread overview: 30+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2012-06-07 18:04 Random data corruption in VM, possibly caused by rbd Guido Winkelmann
2012-06-07 18:18 ` Stefan Priebe
2012-06-07 18:37   ` Guido Winkelmann
2012-06-07 19:54     ` Andrey Korolyov
2012-06-07 21:03       ` Guido Winkelmann
2012-06-07 21:53     ` Marcus Sorensen
2012-06-07 22:12       ` Guido Winkelmann
2012-06-07 18:40 ` Oliver Francke
2012-06-07 19:48 ` Josh Durgin
2012-06-07 21:36   ` Guido Winkelmann [this message]
2012-06-07 22:13     ` Tommi Virtanen
2012-06-08 12:55   ` Guido Winkelmann
2012-06-08 13:08     ` Guido Winkelmann
2012-06-08 13:36     ` Oliver Francke
2012-06-08 13:55       ` Sage Weil
2012-06-08 14:50         ` Josh Durgin
2012-06-08 15:39           ` Oliver Francke
2012-06-08 17:15           ` Guido Winkelmann
2012-06-10  3:04             ` Sage Weil
2012-06-10  3:07               ` Sage Weil
2012-06-11 14:15               ` Guido Winkelmann
2012-06-11 15:50         ` Guido Winkelmann
2012-06-11 16:30           ` Sage Weil
2012-06-11 17:07             ` Guido Winkelmann
2012-06-11 17:12               ` Sage Weil
2012-06-11 17:29               ` Josh Durgin
2012-06-12 12:31             ` Guido Winkelmann
2012-06-15 12:14               ` Stefan Majer
2012-06-15 15:38                 ` Josh Durgin
2012-06-15 18:50                   ` Josh Durgin

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1432839.r57HJoU1Hp@tolkien \
    --to=guido-ceph@thisisnotatest.de \
    --cc=ceph-devel@vger.kernel.org \
    --cc=josh.durgin@inktank.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox