CEPH filesystem development
 help / color / mirror / Atom feed
* un-even data filled on OSDs
@ 2016-06-07  7:32 M Ranga Swami Reddy
       [not found] ` <CANA9Uk4NH7H4kQjEzeNp26bvar5DRWvnNdPff-wCJDRfb1be_A-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  0 siblings, 1 reply; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-07  7:32 UTC (permalink / raw)
  To: ceph-devel; +Cc: ceph-users

Hello,
I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
with >85% of data and few OSDs filled with ~60%-70% of data.

Any reason why the unevenly OSDs filling happned? do I need to any
tweaks on configuration to fix the above? Please advise.

PS: Ceph version is - 0.80.7

Thanks
Swami

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found] ` <CANA9Uk4NH7H4kQjEzeNp26bvar5DRWvnNdPff-wCJDRfb1be_A-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2016-06-07 12:30   ` Sage Weil
  2016-06-07 12:40     ` M Ranga Swami Reddy
  0 siblings, 1 reply; 19+ messages in thread
From: Sage Weil @ 2016-06-07 12:30 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: ceph-devel, ceph-users

On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> Hello,
> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
> with >85% of data and few OSDs filled with ~60%-70% of data.
> 
> Any reason why the unevenly OSDs filling happned? do I need to any
> tweaks on configuration to fix the above? Please advise.
> 
> PS: Ceph version is - 0.80.7

Jewel and the latest hammer point release have an improved 
reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry 
run) to correct this.

sage

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-07 12:30   ` Sage Weil
@ 2016-06-07 12:40     ` M Ranga Swami Reddy
       [not found]       ` <CANA9Uk6YXFpAyp+jUgq7Cc9nRwR4Ha=UO-0nPrEAxy8A_GxGNA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  0 siblings, 1 reply; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-07 12:40 UTC (permalink / raw)
  To: Sage Weil; +Cc: ceph-devel, ceph-users

Hi Sage,
>Jewel and the latest hammer point release have an improved
>reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> run) to correct this.

Thank you....But not planning to upgrade the cluster soon.
So, in this case - are there any tunable options will help? like
"crush tunable optimal" or so?
OR any other configuration options change will help?


Thanks
Swami


On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>> Hello,
>> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>> with >85% of data and few OSDs filled with ~60%-70% of data.
>>
>> Any reason why the unevenly OSDs filling happned? do I need to any
>> tweaks on configuration to fix the above? Please advise.
>>
>> PS: Ceph version is - 0.80.7
>
> Jewel and the latest hammer point release have an improved
> reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> run) to correct this.
>
> sage
>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]       ` <CANA9Uk6YXFpAyp+jUgq7Cc9nRwR4Ha=UO-0nPrEAxy8A_GxGNA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2016-06-07 13:07         ` Sage Weil
  2016-06-07 13:11           ` M Ranga Swami Reddy
  0 siblings, 1 reply; 19+ messages in thread
From: Sage Weil @ 2016-06-07 13:07 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: ceph-devel, ceph-users

On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> Hi Sage,
> >Jewel and the latest hammer point release have an improved
> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> > run) to correct this.
> 
> Thank you....But not planning to upgrade the cluster soon.
> So, in this case - are there any tunable options will help? like
> "crush tunable optimal" or so?
> OR any other configuration options change will help?

Firefly also has reweight-by-utilization... it's just a bit less friendly 
than the newer versions.  CRUSH tunables don't generally help here unless 
you have lots of OSDs that are down+out.

Note that firefly is no longer supported.

sage


> 
> 
> Thanks
> Swami
> 
> 
> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> >> Hello,
> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
> >> with >85% of data and few OSDs filled with ~60%-70% of data.
> >>
> >> Any reason why the unevenly OSDs filling happned? do I need to any
> >> tweaks on configuration to fix the above? Please advise.
> >>
> >> PS: Ceph version is - 0.80.7
> >
> > Jewel and the latest hammer point release have an improved
> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> > run) to correct this.
> >
> > sage
> >
> --
> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> 
> 

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-07 13:07         ` Sage Weil
@ 2016-06-07 13:11           ` M Ranga Swami Reddy
  2016-06-07 13:21             ` Sage Weil
  2016-06-08  1:05             ` Blair Bethwaite
  0 siblings, 2 replies; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-07 13:11 UTC (permalink / raw)
  To: Sage Weil; +Cc: ceph-devel, ceph-users

OK, understood...
To fix the nearfull warn, I am reducing the weight of a specific OSD,
which filled >85%..
Is this work-around advisable?

Thanks
Swami

On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>> Hi Sage,
>> >Jewel and the latest hammer point release have an improved
>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>> > run) to correct this.
>>
>> Thank you....But not planning to upgrade the cluster soon.
>> So, in this case - are there any tunable options will help? like
>> "crush tunable optimal" or so?
>> OR any other configuration options change will help?
>
> Firefly also has reweight-by-utilization... it's just a bit less friendly
> than the newer versions.  CRUSH tunables don't generally help here unless
> you have lots of OSDs that are down+out.
>
> Note that firefly is no longer supported.
>
> sage
>
>
>>
>>
>> Thanks
>> Swami
>>
>>
>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>> >> Hello,
>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>> >>
>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>> >> tweaks on configuration to fix the above? Please advise.
>> >>
>> >> PS: Ceph version is - 0.80.7
>> >
>> > Jewel and the latest hammer point release have an improved
>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>> > run) to correct this.
>> >
>> > sage
>> >
>> --
>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>> the body of a message to majordomo@vger.kernel.org
>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>
>>

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-07 13:11           ` M Ranga Swami Reddy
@ 2016-06-07 13:21             ` Sage Weil
       [not found]               ` <alpine.DEB.2.11.1606070921430.6221-Wo5lQnKln9t9PHm/lf2LFUEOCMrvLtNR@public.gmane.org>
  2016-06-08  1:05             ` Blair Bethwaite
  1 sibling, 1 reply; 19+ messages in thread
From: Sage Weil @ 2016-06-07 13:21 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: ceph-devel, ceph-users

On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> OK, understood...
> To fix the nearfull warn, I am reducing the weight of a specific OSD,
> which filled >85%..
> Is this work-around advisable?

Sure.  This is what reweight-by-utilization does for you, but 
automatically.

sage

> 
> Thanks
> Swami
> 
> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> >> Hi Sage,
> >> >Jewel and the latest hammer point release have an improved
> >> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> >> > run) to correct this.
> >>
> >> Thank you....But not planning to upgrade the cluster soon.
> >> So, in this case - are there any tunable options will help? like
> >> "crush tunable optimal" or so?
> >> OR any other configuration options change will help?
> >
> > Firefly also has reweight-by-utilization... it's just a bit less friendly
> > than the newer versions.  CRUSH tunables don't generally help here unless
> > you have lots of OSDs that are down+out.
> >
> > Note that firefly is no longer supported.
> >
> > sage
> >
> >
> >>
> >>
> >> Thanks
> >> Swami
> >>
> >>
> >> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
> >> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> >> >> Hello,
> >> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
> >> >> with >85% of data and few OSDs filled with ~60%-70% of data.
> >> >>
> >> >> Any reason why the unevenly OSDs filling happned? do I need to any
> >> >> tweaks on configuration to fix the above? Please advise.
> >> >>
> >> >> PS: Ceph version is - 0.80.7
> >> >
> >> > Jewel and the latest hammer point release have an improved
> >> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> >> > run) to correct this.
> >> >
> >> > sage
> >> >
> >> --
> >> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
> >> the body of a message to majordomo@vger.kernel.org
> >> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> >>
> >>
> 
> 

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]               ` <alpine.DEB.2.11.1606070921430.6221-Wo5lQnKln9t9PHm/lf2LFUEOCMrvLtNR@public.gmane.org>
@ 2016-06-07 13:52                 ` Corentin Bonneton
       [not found]                   ` <1F4849EF-DC23-468C-9008-7F1CCEE50F36-Ir4F46R2WEs@public.gmane.org>
  2016-06-07 15:11                 ` Markus Blank-Burian
  1 sibling, 1 reply; 19+ messages in thread
From: Corentin Bonneton @ 2016-06-07 13:52 UTC (permalink / raw)
  To: Sage Weil; +Cc: ceph-devel, ceph-users


[-- Attachment #1.1: Type: text/plain, Size: 2670 bytes --]

Hello,
You how much your PG pools since he first saw you have left too big.

--
Cordialement,
Corentin BONNETON


> Le 7 juin 2016 à 15:21, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> a écrit :
> 
> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>> OK, understood...
>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>> which filled >85%..
>> Is this work-around advisable?
> 
> Sure.  This is what reweight-by-utilization does for you, but 
> automatically.
> 
> sage
> 
>> 
>> Thanks
>> Swami
>> 
>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>> Hi Sage,
>>>>> Jewel and the latest hammer point release have an improved
>>>>> reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>> run) to correct this.
>>>> 
>>>> Thank you....But not planning to upgrade the cluster soon.
>>>> So, in this case - are there any tunable options will help? like
>>>> "crush tunable optimal" or so?
>>>> OR any other configuration options change will help?
>>> 
>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>> you have lots of OSDs that are down+out.
>>> 
>>> Note that firefly is no longer supported.
>>> 
>>> sage
>>> 
>>> 
>>>> 
>>>> 
>>>> Thanks
>>>> Swami
>>>> 
>>>> 
>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>> Hello,
>>>>>> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>>> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>>> 
>>>>>> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>>> tweaks on configuration to fix the above? Please advise.
>>>>>> 
>>>>>> PS: Ceph version is - 0.80.7
>>>>> 
>>>>> Jewel and the latest hammer point release have an improved
>>>>> reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>> run) to correct this.
>>>>> 
>>>>> sage
>>>>> 
>>>> --
>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>> 
>>>> 
>> 
>> 
> _______________________________________________
> ceph-users mailing list
> ceph-users-idqoXFIVOFJgJs9I8MT0rw@public.gmane.org
> http://lists.ceph.com/listinfo.cgi/ceph-users-ceph.com


[-- Attachment #1.2: Type: text/html, Size: 5523 bytes --]

[-- Attachment #2: Type: text/plain, Size: 178 bytes --]

_______________________________________________
ceph-users mailing list
ceph-users-idqoXFIVOFJgJs9I8MT0rw@public.gmane.org
http://lists.ceph.com/listinfo.cgi/ceph-users-ceph.com

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]               ` <alpine.DEB.2.11.1606070921430.6221-Wo5lQnKln9t9PHm/lf2LFUEOCMrvLtNR@public.gmane.org>
  2016-06-07 13:52                 ` Corentin Bonneton
@ 2016-06-07 15:11                 ` Markus Blank-Burian
  1 sibling, 0 replies; 19+ messages in thread
From: Markus Blank-Burian @ 2016-06-07 15:11 UTC (permalink / raw)
  To: 'Sage Weil', 'M Ranga Swami Reddy'
  Cc: 'ceph-devel', 'ceph-users'

Hello Sage,

are there any development plans to improve PG distribution to OSDs? 

We use CephFS on Infernalis and objects are distributed very well across the
PGs. But the automatic PG distribution creates large fluctuations, even in the
simplest case for same-sized OSDs and a flat hierarchy using 3x replication.
Testing with crushtool, we would need an unreasonable high number of PG/OSD to
have a nearly-flat distribution.
Reweighting the PGs with "ceph osd reweight" helped for our cache-pool since it
has a flat hierarchy (one OSD/host). The EC pool on the other hand has a
hierarchy with around 2-4 OSDs (1-4TB / OSD) per host and redundancy is
distributed across hosts (6-8TB / host). I assume, there would be no way around
tuning the host weights in the crush map in this case?

Markus

-----Original Message-----
From: ceph-users [mailto:ceph-users-bounces-idqoXFIVOFJgJs9I8MT0rw@public.gmane.org] On Behalf Of Sage
Weil
Sent: Dienstag, 7. Juni 2016 15:22
To: M Ranga Swami Reddy <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
Cc: ceph-devel <ceph-devel-u79uwXL29TY76Z2rM5mHXA@public.gmane.org>; ceph-users
<ceph-users-idqoXFIVOFJgJs9I8MT0rw@public.gmane.org>
Subject: Re: [ceph-users] un-even data filled on OSDs

On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> OK, understood...
> To fix the nearfull warn, I am reducing the weight of a specific OSD, 
> which filled >85%..
> Is this work-around advisable?

Sure.  This is what reweight-by-utilization does for you, but automatically.

sage

> 
> Thanks
> Swami
> 
> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> >> Hi Sage,
> >> >Jewel and the latest hammer point release have an improved 
> >> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... 
> >> >to dry
> >> > run) to correct this.
> >>
> >> Thank you....But not planning to upgrade the cluster soon.
> >> So, in this case - are there any tunable options will help? like 
> >> "crush tunable optimal" or so?
> >> OR any other configuration options change will help?
> >
> > Firefly also has reweight-by-utilization... it's just a bit less 
> > friendly than the newer versions.  CRUSH tunables don't generally 
> > help here unless you have lots of OSDs that are down+out.
> >
> > Note that firefly is no longer supported.
> >
> > sage
> >
> >
> >>
> >>
> >> Thanks
> >> Swami
> >>
> >>
> >> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
> >> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
> >> >> Hello,
> >> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs 
> >> >> filled with >85% of data and few OSDs filled with ~60%-70% of data.
> >> >>
> >> >> Any reason why the unevenly OSDs filling happned? do I need to 
> >> >> any tweaks on configuration to fix the above? Please advise.
> >> >>
> >> >> PS: Ceph version is - 0.80.7
> >> >
> >> > Jewel and the latest hammer point release have an improved 
> >> > reweight-by-utilization (ceph osd test-reweight-by-utilization 
> >> > ... to dry
> >> > run) to correct this.
> >> >
> >> > sage
> >> >
> >> --
> >> To unsubscribe from this list: send the line "unsubscribe 
> >> ceph-devel" in the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org 
> >> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> >>
> >>
> 
> 
_______________________________________________
ceph-users mailing list
ceph-users-idqoXFIVOFJgJs9I8MT0rw@public.gmane.org
http://lists.ceph.com/listinfo.cgi/ceph-users-ceph.com

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]                   ` <1F4849EF-DC23-468C-9008-7F1CCEE50F36-Ir4F46R2WEs@public.gmane.org>
@ 2016-06-07 16:45                     ` M Ranga Swami Reddy
  0 siblings, 0 replies; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-07 16:45 UTC (permalink / raw)
  To: Corentin Bonneton; +Cc: ceph-devel, ceph-users

In my cluster:
 351 OSDs with same size and 8192 pgs per pool. And 60% RAW space used.

Thanks
Swami


On Tue, Jun 7, 2016 at 7:22 PM, Corentin Bonneton <list@titin.fr> wrote:
> Hello,
> You how much your PG pools since he first saw you have left too big.
>
> --
> Cordialement,
> Corentin BONNETON
>
>
> Le 7 juin 2016 à 15:21, Sage Weil <sage@newdream.net> a écrit :
>
> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>
> OK, understood...
> To fix the nearfull warn, I am reducing the weight of a specific OSD,
> which filled >85%..
> Is this work-around advisable?
>
>
> Sure.  This is what reweight-by-utilization does for you, but
> automatically.
>
> sage
>
>
> Thanks
> Swami
>
> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>
> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>
> Hi Sage,
>
> Jewel and the latest hammer point release have an improved
> reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> run) to correct this.
>
>
> Thank you....But not planning to upgrade the cluster soon.
> So, in this case - are there any tunable options will help? like
> "crush tunable optimal" or so?
> OR any other configuration options change will help?
>
>
> Firefly also has reweight-by-utilization... it's just a bit less friendly
> than the newer versions.  CRUSH tunables don't generally help here unless
> you have lots of OSDs that are down+out.
>
> Note that firefly is no longer supported.
>
> sage
>
>
>
>
> Thanks
> Swami
>
>
> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>
> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>
> Hello,
> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
> with >85% of data and few OSDs filled with ~60%-70% of data.
>
> Any reason why the unevenly OSDs filling happned? do I need to any
> tweaks on configuration to fix the above? Please advise.
>
> PS: Ceph version is - 0.80.7
>
>
> Jewel and the latest hammer point release have an improved
> reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
> run) to correct this.
>
> sage
>
> --
> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>
>
>
>
> _______________________________________________
> ceph-users mailing list
> ceph-users@lists.ceph.com
> http://lists.ceph.com/listinfo.cgi/ceph-users-ceph.com
>
>
_______________________________________________
ceph-users mailing list
ceph-users@lists.ceph.com
http://lists.ceph.com/listinfo.cgi/ceph-users-ceph.com

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-07 13:11           ` M Ranga Swami Reddy
  2016-06-07 13:21             ` Sage Weil
@ 2016-06-08  1:05             ` Blair Bethwaite
  2016-06-08  5:04               ` M Ranga Swami Reddy
  1 sibling, 1 reply; 19+ messages in thread
From: Blair Bethwaite @ 2016-06-08  1:05 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: Sage Weil, ceph-devel, ceph-users

Swami,

Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
that'll work with Firefly and allow you to only tune down weight of a
specific number of overfull OSDs.

Cheers,

On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
> OK, understood...
> To fix the nearfull warn, I am reducing the weight of a specific OSD,
> which filled >85%..
> Is this work-around advisable?
>
> Thanks
> Swami
>
> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>> Hi Sage,
>>> >Jewel and the latest hammer point release have an improved
>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>> > run) to correct this.
>>>
>>> Thank you....But not planning to upgrade the cluster soon.
>>> So, in this case - are there any tunable options will help? like
>>> "crush tunable optimal" or so?
>>> OR any other configuration options change will help?
>>
>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>> than the newer versions.  CRUSH tunables don't generally help here unless
>> you have lots of OSDs that are down+out.
>>
>> Note that firefly is no longer supported.
>>
>> sage
>>
>>
>>>
>>>
>>> Thanks
>>> Swami
>>>
>>>
>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>> >> Hello,
>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>> >>
>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>> >> tweaks on configuration to fix the above? Please advise.
>>> >>
>>> >> PS: Ceph version is - 0.80.7
>>> >
>>> > Jewel and the latest hammer point release have an improved
>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>> > run) to correct this.
>>> >
>>> > sage
>>> >
>>> --
>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>> the body of a message to majordomo@vger.kernel.org
>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>
>>>
> --
> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html



-- 
Cheers,
~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-08  1:05             ` Blair Bethwaite
@ 2016-06-08  5:04               ` M Ranga Swami Reddy
  2016-06-08  5:08                 ` Blair Bethwaite
  2016-06-09 10:20                 ` M Ranga Swami Reddy
  0 siblings, 2 replies; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-08  5:04 UTC (permalink / raw)
  To: Blair Bethwaite; +Cc: Sage Weil, ceph-devel, ceph-users

Blair - Thanks for the script...Btw, is this script has option for dry run?

Thanks
Swami

On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
<blair.bethwaite@gmail.com> wrote:
> Swami,
>
> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
> that'll work with Firefly and allow you to only tune down weight of a
> specific number of overfull OSDs.
>
> Cheers,
>
> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>> OK, understood...
>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>> which filled >85%..
>> Is this work-around advisable?
>>
>> Thanks
>> Swami
>>
>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>> Hi Sage,
>>>> >Jewel and the latest hammer point release have an improved
>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>> > run) to correct this.
>>>>
>>>> Thank you....But not planning to upgrade the cluster soon.
>>>> So, in this case - are there any tunable options will help? like
>>>> "crush tunable optimal" or so?
>>>> OR any other configuration options change will help?
>>>
>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>> you have lots of OSDs that are down+out.
>>>
>>> Note that firefly is no longer supported.
>>>
>>> sage
>>>
>>>
>>>>
>>>>
>>>> Thanks
>>>> Swami
>>>>
>>>>
>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>> >> Hello,
>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>> >>
>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>> >> tweaks on configuration to fix the above? Please advise.
>>>> >>
>>>> >> PS: Ceph version is - 0.80.7
>>>> >
>>>> > Jewel and the latest hammer point release have an improved
>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>> > run) to correct this.
>>>> >
>>>> > sage
>>>> >
>>>> --
>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>> the body of a message to majordomo@vger.kernel.org
>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>
>>>>
>> --
>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>> the body of a message to majordomo@vger.kernel.org
>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>
>
>
> --
> Cheers,
> ~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-08  5:04               ` M Ranga Swami Reddy
@ 2016-06-08  5:08                 ` Blair Bethwaite
       [not found]                   ` <CA+z5DsxrovFX6GGmcfWAhNm69fb3s4-M6wMSR1u5vzdV8LTSSg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  2016-06-09 10:20                 ` M Ranga Swami Reddy
  1 sibling, 1 reply; 19+ messages in thread
From: Blair Bethwaite @ 2016-06-08  5:08 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: Sage Weil, ceph-devel, ceph-users

It runs by default in dry-run mode, which IMHO opinion should be the
default for operations like this. IIRC you add "-d -r" to make it
actually apply the re-weighting.

Cheers,

On 8 June 2016 at 15:04, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
> Blair - Thanks for the script...Btw, is this script has option for dry run?
>
> Thanks
> Swami
>
> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
> <blair.bethwaite@gmail.com> wrote:
>> Swami,
>>
>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>> that'll work with Firefly and allow you to only tune down weight of a
>> specific number of overfull OSDs.
>>
>> Cheers,
>>
>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>>> OK, understood...
>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>> which filled >85%..
>>> Is this work-around advisable?
>>>
>>> Thanks
>>> Swami
>>>
>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>> Hi Sage,
>>>>> >Jewel and the latest hammer point release have an improved
>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>> > run) to correct this.
>>>>>
>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>> So, in this case - are there any tunable options will help? like
>>>>> "crush tunable optimal" or so?
>>>>> OR any other configuration options change will help?
>>>>
>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>> you have lots of OSDs that are down+out.
>>>>
>>>> Note that firefly is no longer supported.
>>>>
>>>> sage
>>>>
>>>>
>>>>>
>>>>>
>>>>> Thanks
>>>>> Swami
>>>>>
>>>>>
>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>> >> Hello,
>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>> >>
>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>> >>
>>>>> >> PS: Ceph version is - 0.80.7
>>>>> >
>>>>> > Jewel and the latest hammer point release have an improved
>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>> > run) to correct this.
>>>>> >
>>>>> > sage
>>>>> >
>>>>> --
>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>> the body of a message to majordomo@vger.kernel.org
>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>
>>>>>
>>> --
>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>> the body of a message to majordomo@vger.kernel.org
>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>
>>
>>
>> --
>> Cheers,
>> ~Blairo



-- 
Cheers,
~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]                   ` <CA+z5DsxrovFX6GGmcfWAhNm69fb3s4-M6wMSR1u5vzdV8LTSSg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2016-06-08  5:09                     ` M Ranga Swami Reddy
  0 siblings, 0 replies; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-08  5:09 UTC (permalink / raw)
  To: Blair Bethwaite; +Cc: ceph-devel, ceph-users

Thats great.. Will try this..

Thanks
Swami

On Wed, Jun 8, 2016 at 10:38 AM, Blair Bethwaite
<blair.bethwaite-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
> It runs by default in dry-run mode, which IMHO opinion should be the
> default for operations like this. IIRC you add "-d -r" to make it
> actually apply the re-weighting.
>
> Cheers,
>
> On 8 June 2016 at 15:04, M Ranga Swami Reddy <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>> Blair - Thanks for the script...Btw, is this script has option for dry run?
>>
>> Thanks
>> Swami
>>
>> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
>> <blair.bethwaite-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>>> Swami,
>>>
>>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>>> that'll work with Firefly and allow you to only tune down weight of a
>>> specific number of overfull OSDs.
>>>
>>> Cheers,
>>>
>>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>>>> OK, understood...
>>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>>> which filled >85%..
>>>> Is this work-around advisable?
>>>>
>>>> Thanks
>>>> Swami
>>>>
>>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>> Hi Sage,
>>>>>> >Jewel and the latest hammer point release have an improved
>>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>> > run) to correct this.
>>>>>>
>>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>>> So, in this case - are there any tunable options will help? like
>>>>>> "crush tunable optimal" or so?
>>>>>> OR any other configuration options change will help?
>>>>>
>>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>>> you have lots of OSDs that are down+out.
>>>>>
>>>>> Note that firefly is no longer supported.
>>>>>
>>>>> sage
>>>>>
>>>>>
>>>>>>
>>>>>>
>>>>>> Thanks
>>>>>> Swami
>>>>>>
>>>>>>
>>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>> >> Hello,
>>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>>> >>
>>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>>> >>
>>>>>> >> PS: Ceph version is - 0.80.7
>>>>>> >
>>>>>> > Jewel and the latest hammer point release have an improved
>>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>> > run) to correct this.
>>>>>> >
>>>>>> > sage
>>>>>> >
>>>>>> --
>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>>
>>>>>>
>>>> --
>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>
>>>
>>>
>>> --
>>> Cheers,
>>> ~Blairo
>
>
>
> --
> Cheers,
> ~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-08  5:04               ` M Ranga Swami Reddy
  2016-06-08  5:08                 ` Blair Bethwaite
@ 2016-06-09 10:20                 ` M Ranga Swami Reddy
       [not found]                   ` <CANA9Uk48oQ7sjd614LzYy=7HxYfmO8qZxoQCC62wYeOdsBPFeg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  1 sibling, 1 reply; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-09 10:20 UTC (permalink / raw)
  To: Blair Bethwaite; +Cc: ceph-devel, ceph-users

Hi Blari,
I ran the script and results are below:
==
./crush-reweight-by-utilization.py
average_util: 0.587024, overload_util: 0.704429, underload_util: 0.587024.
reweighted:
43 (0.852690 >= 0.704429) [1.000000 -> 0.950000]
238 (0.845154 >= 0.704429) [1.000000 -> 0.950000]
104 (0.827908 >= 0.704429) [1.000000 -> 0.950000]
173 (0.817063 >= 0.704429) [1.000000 -> 0.950000]
==

is the above scripts says to reweight 43 -> 0.95?

Thanks
Swami

On Wed, Jun 8, 2016 at 10:34 AM, M Ranga Swami Reddy
<swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
> Blair - Thanks for the script...Btw, is this script has option for dry run?
>
> Thanks
> Swami
>
> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
> <blair.bethwaite-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>> Swami,
>>
>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>> that'll work with Firefly and allow you to only tune down weight of a
>> specific number of overfull OSDs.
>>
>> Cheers,
>>
>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>>> OK, understood...
>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>> which filled >85%..
>>> Is this work-around advisable?
>>>
>>> Thanks
>>> Swami
>>>
>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>> Hi Sage,
>>>>> >Jewel and the latest hammer point release have an improved
>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>> > run) to correct this.
>>>>>
>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>> So, in this case - are there any tunable options will help? like
>>>>> "crush tunable optimal" or so?
>>>>> OR any other configuration options change will help?
>>>>
>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>> you have lots of OSDs that are down+out.
>>>>
>>>> Note that firefly is no longer supported.
>>>>
>>>> sage
>>>>
>>>>
>>>>>
>>>>>
>>>>> Thanks
>>>>> Swami
>>>>>
>>>>>
>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>> >> Hello,
>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>> >>
>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>> >>
>>>>> >> PS: Ceph version is - 0.80.7
>>>>> >
>>>>> > Jewel and the latest hammer point release have an improved
>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>> > run) to correct this.
>>>>> >
>>>>> > sage
>>>>> >
>>>>> --
>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>
>>>>>
>>> --
>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>
>>
>>
>> --
>> Cheers,
>> ~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]                   ` <CANA9Uk48oQ7sjd614LzYy=7HxYfmO8qZxoQCC62wYeOdsBPFeg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2016-06-09 12:20                     ` Blair Bethwaite
  2016-06-10  2:08                       ` M Ranga Swami Reddy
  0 siblings, 1 reply; 19+ messages in thread
From: Blair Bethwaite @ 2016-06-09 12:20 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: ceph-devel, ceph-users

Swami,

Run it with the help option for more context:
"./crush-reweight-by-utilization.py --help". In your example below
it's reporting to you what changes it would make to your OSD reweight
values based on the default option settings (because you didn't
specify any options). To make the script actually apply those weight
changes you need the "-d -r" or "--doit --really" flags.

If you want to get an idea of the impact that the weight changes will
have before actually starting to move data then I suggest setting
norecover and nobackfill (ceph osd set ...) on your cluster before
making the weight changes, you can then examine "ceph -s" output
(looking at "objects misplaced" to determine the scale of recovery
required. Unset the flags once ready to start or back-out the reweight
settings if you change your mind. You'll also want to lower these
recovery and backfill tunables to reduce impact to client I/O (and if
possible do not do this reweight change during peak I/O hours):
ceph tell osd.* injectargs '--osd-max-backfills 1'
ceph tell osd.* injectargs '--osd-max-recovery-threads 1'
ceph tell osd.* injectargs '--osd-recovery-op-priority 1'
ceph tell osd.* injectargs '--osd-client-op-priority 63'
ceph tell osd.* injectargs '--osd-recovery-max-active 1'

Cheers,

On 9 June 2016 at 20:20, M Ranga Swami Reddy <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
> Hi Blari,
> I ran the script and results are below:
> ==
> ./crush-reweight-by-utilization.py
> average_util: 0.587024, overload_util: 0.704429, underload_util: 0.587024.
> reweighted:
> 43 (0.852690 >= 0.704429) [1.000000 -> 0.950000]
> 238 (0.845154 >= 0.704429) [1.000000 -> 0.950000]
> 104 (0.827908 >= 0.704429) [1.000000 -> 0.950000]
> 173 (0.817063 >= 0.704429) [1.000000 -> 0.950000]
> ==
>
> is the above scripts says to reweight 43 -> 0.95?
>
> Thanks
> Swami
>
> On Wed, Jun 8, 2016 at 10:34 AM, M Ranga Swami Reddy
> <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>> Blair - Thanks for the script...Btw, is this script has option for dry run?
>>
>> Thanks
>> Swami
>>
>> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
>> <blair.bethwaite-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>>> Swami,
>>>
>>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>>> that'll work with Firefly and allow you to only tune down weight of a
>>> specific number of overfull OSDs.
>>>
>>> Cheers,
>>>
>>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>>>> OK, understood...
>>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>>> which filled >85%..
>>>> Is this work-around advisable?
>>>>
>>>> Thanks
>>>> Swami
>>>>
>>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>> Hi Sage,
>>>>>> >Jewel and the latest hammer point release have an improved
>>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>> > run) to correct this.
>>>>>>
>>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>>> So, in this case - are there any tunable options will help? like
>>>>>> "crush tunable optimal" or so?
>>>>>> OR any other configuration options change will help?
>>>>>
>>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>>> you have lots of OSDs that are down+out.
>>>>>
>>>>> Note that firefly is no longer supported.
>>>>>
>>>>> sage
>>>>>
>>>>>
>>>>>>
>>>>>>
>>>>>> Thanks
>>>>>> Swami
>>>>>>
>>>>>>
>>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage-BnTBU8nroG7k1uMJSBkQmQ@public.gmane.org> wrote:
>>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>> >> Hello,
>>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>>> >>
>>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>>> >>
>>>>>> >> PS: Ceph version is - 0.80.7
>>>>>> >
>>>>>> > Jewel and the latest hammer point release have an improved
>>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>> > run) to correct this.
>>>>>> >
>>>>>> > sage
>>>>>> >
>>>>>> --
>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>>
>>>>>>
>>>> --
>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>> the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>
>>>
>>>
>>> --
>>> Cheers,
>>> ~Blairo



-- 
Cheers,
~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-09 12:20                     ` Blair Bethwaite
@ 2016-06-10  2:08                       ` M Ranga Swami Reddy
  2016-06-10  2:10                         ` Blair Bethwaite
       [not found]                         ` <CANA9Uk4Bddii2vBgExF3tM9hs2CByefep8mpUZpHKFCyzEFyXg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  0 siblings, 2 replies; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-10  2:08 UTC (permalink / raw)
  To: Blair Bethwaite; +Cc: Sage Weil, ceph-devel, ceph-users

Blair - Thanks for the details. I used to set the low priority for
recovery during the rebalance/recovery activity.
Even though I set the recovery_priority as 5 (instead of 1) and
client-op_priority set as 63, some of my customers complained that
their VMs are not reachable for a few mins/secs during the reblancing
task. Not sure, these low priority configurations are doing the job as
its.

Thanks
Swami

On Thu, Jun 9, 2016 at 5:50 PM, Blair Bethwaite
<blair.bethwaite@gmail.com> wrote:
> Swami,
>
> Run it with the help option for more context:
> "./crush-reweight-by-utilization.py --help". In your example below
> it's reporting to you what changes it would make to your OSD reweight
> values based on the default option settings (because you didn't
> specify any options). To make the script actually apply those weight
> changes you need the "-d -r" or "--doit --really" flags.
>
> If you want to get an idea of the impact that the weight changes will
> have before actually starting to move data then I suggest setting
> norecover and nobackfill (ceph osd set ...) on your cluster before
> making the weight changes, you can then examine "ceph -s" output
> (looking at "objects misplaced" to determine the scale of recovery
> required. Unset the flags once ready to start or back-out the reweight
> settings if you change your mind. You'll also want to lower these
> recovery and backfill tunables to reduce impact to client I/O (and if
> possible do not do this reweight change during peak I/O hours):
> ceph tell osd.* injectargs '--osd-max-backfills 1'
> ceph tell osd.* injectargs '--osd-max-recovery-threads 1'
> ceph tell osd.* injectargs '--osd-recovery-op-priority 1'
> ceph tell osd.* injectargs '--osd-client-op-priority 63'
> ceph tell osd.* injectargs '--osd-recovery-max-active 1'
>
> Cheers,
>
> On 9 June 2016 at 20:20, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>> Hi Blari,
>> I ran the script and results are below:
>> ==
>> ./crush-reweight-by-utilization.py
>> average_util: 0.587024, overload_util: 0.704429, underload_util: 0.587024.
>> reweighted:
>> 43 (0.852690 >= 0.704429) [1.000000 -> 0.950000]
>> 238 (0.845154 >= 0.704429) [1.000000 -> 0.950000]
>> 104 (0.827908 >= 0.704429) [1.000000 -> 0.950000]
>> 173 (0.817063 >= 0.704429) [1.000000 -> 0.950000]
>> ==
>>
>> is the above scripts says to reweight 43 -> 0.95?
>>
>> Thanks
>> Swami
>>
>> On Wed, Jun 8, 2016 at 10:34 AM, M Ranga Swami Reddy
>> <swamireddy@gmail.com> wrote:
>>> Blair - Thanks for the script...Btw, is this script has option for dry run?
>>>
>>> Thanks
>>> Swami
>>>
>>> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
>>> <blair.bethwaite@gmail.com> wrote:
>>>> Swami,
>>>>
>>>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>>>> that'll work with Firefly and allow you to only tune down weight of a
>>>> specific number of overfull OSDs.
>>>>
>>>> Cheers,
>>>>
>>>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>>>>> OK, understood...
>>>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>>>> which filled >85%..
>>>>> Is this work-around advisable?
>>>>>
>>>>> Thanks
>>>>> Swami
>>>>>
>>>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>>>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>>> Hi Sage,
>>>>>>> >Jewel and the latest hammer point release have an improved
>>>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>>> > run) to correct this.
>>>>>>>
>>>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>>>> So, in this case - are there any tunable options will help? like
>>>>>>> "crush tunable optimal" or so?
>>>>>>> OR any other configuration options change will help?
>>>>>>
>>>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>>>> you have lots of OSDs that are down+out.
>>>>>>
>>>>>> Note that firefly is no longer supported.
>>>>>>
>>>>>> sage
>>>>>>
>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> Thanks
>>>>>>> Swami
>>>>>>>
>>>>>>>
>>>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>>>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>>> >> Hello,
>>>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>>>> >>
>>>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>>>> >>
>>>>>>> >> PS: Ceph version is - 0.80.7
>>>>>>> >
>>>>>>> > Jewel and the latest hammer point release have an improved
>>>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>>> > run) to correct this.
>>>>>>> >
>>>>>>> > sage
>>>>>>> >
>>>>>>> --
>>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>>> the body of a message to majordomo@vger.kernel.org
>>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>>>
>>>>>>>
>>>>> --
>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>> the body of a message to majordomo@vger.kernel.org
>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>
>>>>
>>>>
>>>> --
>>>> Cheers,
>>>> ~Blairo
>
>
>
> --
> Cheers,
> ~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-10  2:08                       ` M Ranga Swami Reddy
@ 2016-06-10  2:10                         ` Blair Bethwaite
  2016-06-10  8:18                           ` M Ranga Swami Reddy
       [not found]                         ` <CANA9Uk4Bddii2vBgExF3tM9hs2CByefep8mpUZpHKFCyzEFyXg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  1 sibling, 1 reply; 19+ messages in thread
From: Blair Bethwaite @ 2016-06-10  2:10 UTC (permalink / raw)
  To: M Ranga Swami Reddy; +Cc: Sage Weil, ceph-devel, ceph-users

Hi Swami,

That's a known issue, which I believe is much improved in Jewel thanks
to a priority queue added somewhere in the OSD op path (I think). If I
were you I'd be planning to get off Firefly and upgrade.

Cheers,

On 10 June 2016 at 12:08, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
> Blair - Thanks for the details. I used to set the low priority for
> recovery during the rebalance/recovery activity.
> Even though I set the recovery_priority as 5 (instead of 1) and
> client-op_priority set as 63, some of my customers complained that
> their VMs are not reachable for a few mins/secs during the reblancing
> task. Not sure, these low priority configurations are doing the job as
> its.
>
> Thanks
> Swami
>
> On Thu, Jun 9, 2016 at 5:50 PM, Blair Bethwaite
> <blair.bethwaite@gmail.com> wrote:
>> Swami,
>>
>> Run it with the help option for more context:
>> "./crush-reweight-by-utilization.py --help". In your example below
>> it's reporting to you what changes it would make to your OSD reweight
>> values based on the default option settings (because you didn't
>> specify any options). To make the script actually apply those weight
>> changes you need the "-d -r" or "--doit --really" flags.
>>
>> If you want to get an idea of the impact that the weight changes will
>> have before actually starting to move data then I suggest setting
>> norecover and nobackfill (ceph osd set ...) on your cluster before
>> making the weight changes, you can then examine "ceph -s" output
>> (looking at "objects misplaced" to determine the scale of recovery
>> required. Unset the flags once ready to start or back-out the reweight
>> settings if you change your mind. You'll also want to lower these
>> recovery and backfill tunables to reduce impact to client I/O (and if
>> possible do not do this reweight change during peak I/O hours):
>> ceph tell osd.* injectargs '--osd-max-backfills 1'
>> ceph tell osd.* injectargs '--osd-max-recovery-threads 1'
>> ceph tell osd.* injectargs '--osd-recovery-op-priority 1'
>> ceph tell osd.* injectargs '--osd-client-op-priority 63'
>> ceph tell osd.* injectargs '--osd-recovery-max-active 1'
>>
>> Cheers,
>>
>> On 9 June 2016 at 20:20, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>>> Hi Blari,
>>> I ran the script and results are below:
>>> ==
>>> ./crush-reweight-by-utilization.py
>>> average_util: 0.587024, overload_util: 0.704429, underload_util: 0.587024.
>>> reweighted:
>>> 43 (0.852690 >= 0.704429) [1.000000 -> 0.950000]
>>> 238 (0.845154 >= 0.704429) [1.000000 -> 0.950000]
>>> 104 (0.827908 >= 0.704429) [1.000000 -> 0.950000]
>>> 173 (0.817063 >= 0.704429) [1.000000 -> 0.950000]
>>> ==
>>>
>>> is the above scripts says to reweight 43 -> 0.95?
>>>
>>> Thanks
>>> Swami
>>>
>>> On Wed, Jun 8, 2016 at 10:34 AM, M Ranga Swami Reddy
>>> <swamireddy@gmail.com> wrote:
>>>> Blair - Thanks for the script...Btw, is this script has option for dry run?
>>>>
>>>> Thanks
>>>> Swami
>>>>
>>>> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
>>>> <blair.bethwaite@gmail.com> wrote:
>>>>> Swami,
>>>>>
>>>>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>>>>> that'll work with Firefly and allow you to only tune down weight of a
>>>>> specific number of overfull OSDs.
>>>>>
>>>>> Cheers,
>>>>>
>>>>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>>>>>> OK, understood...
>>>>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>>>>> which filled >85%..
>>>>>> Is this work-around advisable?
>>>>>>
>>>>>> Thanks
>>>>>> Swami
>>>>>>
>>>>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>>>>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>>>> Hi Sage,
>>>>>>>> >Jewel and the latest hammer point release have an improved
>>>>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>>>> > run) to correct this.
>>>>>>>>
>>>>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>>>>> So, in this case - are there any tunable options will help? like
>>>>>>>> "crush tunable optimal" or so?
>>>>>>>> OR any other configuration options change will help?
>>>>>>>
>>>>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>>>>> you have lots of OSDs that are down+out.
>>>>>>>
>>>>>>> Note that firefly is no longer supported.
>>>>>>>
>>>>>>> sage
>>>>>>>
>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>> Thanks
>>>>>>>> Swami
>>>>>>>>
>>>>>>>>
>>>>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>>>>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>>>> >> Hello,
>>>>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>>>>> >>
>>>>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>>>>> >>
>>>>>>>> >> PS: Ceph version is - 0.80.7
>>>>>>>> >
>>>>>>>> > Jewel and the latest hammer point release have an improved
>>>>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>>>> > run) to correct this.
>>>>>>>> >
>>>>>>>> > sage
>>>>>>>> >
>>>>>>>> --
>>>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>>>> the body of a message to majordomo@vger.kernel.org
>>>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>>>>
>>>>>>>>
>>>>>> --
>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>> the body of a message to majordomo@vger.kernel.org
>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>
>>>>>
>>>>>
>>>>> --
>>>>> Cheers,
>>>>> ~Blairo
>>
>>
>>
>> --
>> Cheers,
>> ~Blairo



-- 
Cheers,
~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
  2016-06-10  2:10                         ` Blair Bethwaite
@ 2016-06-10  8:18                           ` M Ranga Swami Reddy
  0 siblings, 0 replies; 19+ messages in thread
From: M Ranga Swami Reddy @ 2016-06-10  8:18 UTC (permalink / raw)
  To: Blair Bethwaite; +Cc: Sage Weil, ceph-devel, ceph-users

Thanks Blair. Yes, will plan to upgrade my cluster.

Thanks
Swami

On Fri, Jun 10, 2016 at 7:40 AM, Blair Bethwaite
<blair.bethwaite@gmail.com> wrote:
> Hi Swami,
>
> That's a known issue, which I believe is much improved in Jewel thanks
> to a priority queue added somewhere in the OSD op path (I think). If I
> were you I'd be planning to get off Firefly and upgrade.
>
> Cheers,
>
> On 10 June 2016 at 12:08, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>> Blair - Thanks for the details. I used to set the low priority for
>> recovery during the rebalance/recovery activity.
>> Even though I set the recovery_priority as 5 (instead of 1) and
>> client-op_priority set as 63, some of my customers complained that
>> their VMs are not reachable for a few mins/secs during the reblancing
>> task. Not sure, these low priority configurations are doing the job as
>> its.
>>
>> Thanks
>> Swami
>>
>> On Thu, Jun 9, 2016 at 5:50 PM, Blair Bethwaite
>> <blair.bethwaite@gmail.com> wrote:
>>> Swami,
>>>
>>> Run it with the help option for more context:
>>> "./crush-reweight-by-utilization.py --help". In your example below
>>> it's reporting to you what changes it would make to your OSD reweight
>>> values based on the default option settings (because you didn't
>>> specify any options). To make the script actually apply those weight
>>> changes you need the "-d -r" or "--doit --really" flags.
>>>
>>> If you want to get an idea of the impact that the weight changes will
>>> have before actually starting to move data then I suggest setting
>>> norecover and nobackfill (ceph osd set ...) on your cluster before
>>> making the weight changes, you can then examine "ceph -s" output
>>> (looking at "objects misplaced" to determine the scale of recovery
>>> required. Unset the flags once ready to start or back-out the reweight
>>> settings if you change your mind. You'll also want to lower these
>>> recovery and backfill tunables to reduce impact to client I/O (and if
>>> possible do not do this reweight change during peak I/O hours):
>>> ceph tell osd.* injectargs '--osd-max-backfills 1'
>>> ceph tell osd.* injectargs '--osd-max-recovery-threads 1'
>>> ceph tell osd.* injectargs '--osd-recovery-op-priority 1'
>>> ceph tell osd.* injectargs '--osd-client-op-priority 63'
>>> ceph tell osd.* injectargs '--osd-recovery-max-active 1'
>>>
>>> Cheers,
>>>
>>> On 9 June 2016 at 20:20, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>>>> Hi Blari,
>>>> I ran the script and results are below:
>>>> ==
>>>> ./crush-reweight-by-utilization.py
>>>> average_util: 0.587024, overload_util: 0.704429, underload_util: 0.587024.
>>>> reweighted:
>>>> 43 (0.852690 >= 0.704429) [1.000000 -> 0.950000]
>>>> 238 (0.845154 >= 0.704429) [1.000000 -> 0.950000]
>>>> 104 (0.827908 >= 0.704429) [1.000000 -> 0.950000]
>>>> 173 (0.817063 >= 0.704429) [1.000000 -> 0.950000]
>>>> ==
>>>>
>>>> is the above scripts says to reweight 43 -> 0.95?
>>>>
>>>> Thanks
>>>> Swami
>>>>
>>>> On Wed, Jun 8, 2016 at 10:34 AM, M Ranga Swami Reddy
>>>> <swamireddy@gmail.com> wrote:
>>>>> Blair - Thanks for the script...Btw, is this script has option for dry run?
>>>>>
>>>>> Thanks
>>>>> Swami
>>>>>
>>>>> On Wed, Jun 8, 2016 at 6:35 AM, Blair Bethwaite
>>>>> <blair.bethwaite@gmail.com> wrote:
>>>>>> Swami,
>>>>>>
>>>>>> Try https://github.com/cernceph/ceph-scripts/blob/master/tools/crush-reweight-by-utilization.py,
>>>>>> that'll work with Firefly and allow you to only tune down weight of a
>>>>>> specific number of overfull OSDs.
>>>>>>
>>>>>> Cheers,
>>>>>>
>>>>>> On 7 June 2016 at 23:11, M Ranga Swami Reddy <swamireddy@gmail.com> wrote:
>>>>>>> OK, understood...
>>>>>>> To fix the nearfull warn, I am reducing the weight of a specific OSD,
>>>>>>> which filled >85%..
>>>>>>> Is this work-around advisable?
>>>>>>>
>>>>>>> Thanks
>>>>>>> Swami
>>>>>>>
>>>>>>> On Tue, Jun 7, 2016 at 6:37 PM, Sage Weil <sage@newdream.net> wrote:
>>>>>>>> On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>>>>> Hi Sage,
>>>>>>>>> >Jewel and the latest hammer point release have an improved
>>>>>>>>> >reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>>>>> > run) to correct this.
>>>>>>>>>
>>>>>>>>> Thank you....But not planning to upgrade the cluster soon.
>>>>>>>>> So, in this case - are there any tunable options will help? like
>>>>>>>>> "crush tunable optimal" or so?
>>>>>>>>> OR any other configuration options change will help?
>>>>>>>>
>>>>>>>> Firefly also has reweight-by-utilization... it's just a bit less friendly
>>>>>>>> than the newer versions.  CRUSH tunables don't generally help here unless
>>>>>>>> you have lots of OSDs that are down+out.
>>>>>>>>
>>>>>>>> Note that firefly is no longer supported.
>>>>>>>>
>>>>>>>> sage
>>>>>>>>
>>>>>>>>
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> Thanks
>>>>>>>>> Swami
>>>>>>>>>
>>>>>>>>>
>>>>>>>>> On Tue, Jun 7, 2016 at 6:00 PM, Sage Weil <sage@newdream.net> wrote:
>>>>>>>>> > On Tue, 7 Jun 2016, M Ranga Swami Reddy wrote:
>>>>>>>>> >> Hello,
>>>>>>>>> >> I have aorund 100 OSDs in my ceph cluster. In this a few OSDs filled
>>>>>>>>> >> with >85% of data and few OSDs filled with ~60%-70% of data.
>>>>>>>>> >>
>>>>>>>>> >> Any reason why the unevenly OSDs filling happned? do I need to any
>>>>>>>>> >> tweaks on configuration to fix the above? Please advise.
>>>>>>>>> >>
>>>>>>>>> >> PS: Ceph version is - 0.80.7
>>>>>>>>> >
>>>>>>>>> > Jewel and the latest hammer point release have an improved
>>>>>>>>> > reweight-by-utilization (ceph osd test-reweight-by-utilization ... to dry
>>>>>>>>> > run) to correct this.
>>>>>>>>> >
>>>>>>>>> > sage
>>>>>>>>> >
>>>>>>>>> --
>>>>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>>>>> the body of a message to majordomo@vger.kernel.org
>>>>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>>>>>
>>>>>>>>>
>>>>>>> --
>>>>>>> To unsubscribe from this list: send the line "unsubscribe ceph-devel" in
>>>>>>> the body of a message to majordomo@vger.kernel.org
>>>>>>> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>>>>>>
>>>>>>
>>>>>>
>>>>>> --
>>>>>> Cheers,
>>>>>> ~Blairo
>>>
>>>
>>>
>>> --
>>> Cheers,
>>> ~Blairo
>
>
>
> --
> Cheers,
> ~Blairo

^ permalink raw reply	[flat|nested] 19+ messages in thread

* Re: un-even data filled on OSDs
       [not found]                         ` <CANA9Uk4Bddii2vBgExF3tM9hs2CByefep8mpUZpHKFCyzEFyXg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2016-06-10 10:29                           ` Max A. Krasilnikov
  0 siblings, 0 replies; 19+ messages in thread
From: Max A. Krasilnikov @ 2016-06-10 10:29 UTC (permalink / raw)
  Cc: ceph-devel, ceph-users

Hello! 

On Fri, Jun 10, 2016 at 07:38:10AM +0530, swamireddy wrote:

> Blair - Thanks for the details. I used to set the low priority for
> recovery during the rebalance/recovery activity.
> Even though I set the recovery_priority as 5 (instead of 1) and
> client-op_priority set as 63, some of my customers complained that
> their VMs are not reachable for a few mins/secs during the reblancing
> task. Not sure, these low priority configurations are doing the job as
> its.

It is true up to Hammer at least. I have no possibility to test it on jevel
setup due to my company policy (part of my cluster already is jevel, but I can
not continue upgrade due to direct directive).

> Thanks
> Swami

-- 
WBR, Max A. Krasilnikov

^ permalink raw reply	[flat|nested] 19+ messages in thread

end of thread, other threads:[~2016-06-10 10:29 UTC | newest]

Thread overview: 19+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2016-06-07  7:32 un-even data filled on OSDs M Ranga Swami Reddy
     [not found] ` <CANA9Uk4NH7H4kQjEzeNp26bvar5DRWvnNdPff-wCJDRfb1be_A-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2016-06-07 12:30   ` Sage Weil
2016-06-07 12:40     ` M Ranga Swami Reddy
     [not found]       ` <CANA9Uk6YXFpAyp+jUgq7Cc9nRwR4Ha=UO-0nPrEAxy8A_GxGNA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2016-06-07 13:07         ` Sage Weil
2016-06-07 13:11           ` M Ranga Swami Reddy
2016-06-07 13:21             ` Sage Weil
     [not found]               ` <alpine.DEB.2.11.1606070921430.6221-Wo5lQnKln9t9PHm/lf2LFUEOCMrvLtNR@public.gmane.org>
2016-06-07 13:52                 ` Corentin Bonneton
     [not found]                   ` <1F4849EF-DC23-468C-9008-7F1CCEE50F36-Ir4F46R2WEs@public.gmane.org>
2016-06-07 16:45                     ` M Ranga Swami Reddy
2016-06-07 15:11                 ` Markus Blank-Burian
2016-06-08  1:05             ` Blair Bethwaite
2016-06-08  5:04               ` M Ranga Swami Reddy
2016-06-08  5:08                 ` Blair Bethwaite
     [not found]                   ` <CA+z5DsxrovFX6GGmcfWAhNm69fb3s4-M6wMSR1u5vzdV8LTSSg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2016-06-08  5:09                     ` M Ranga Swami Reddy
2016-06-09 10:20                 ` M Ranga Swami Reddy
     [not found]                   ` <CANA9Uk48oQ7sjd614LzYy=7HxYfmO8qZxoQCC62wYeOdsBPFeg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2016-06-09 12:20                     ` Blair Bethwaite
2016-06-10  2:08                       ` M Ranga Swami Reddy
2016-06-10  2:10                         ` Blair Bethwaite
2016-06-10  8:18                           ` M Ranga Swami Reddy
     [not found]                         ` <CANA9Uk4Bddii2vBgExF3tM9hs2CByefep8mpUZpHKFCyzEFyXg-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2016-06-10 10:29                           ` Max A. Krasilnikov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox