The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* Re: + mm-oom-killer-kills-more-than-needed.patch added to -mm tree
       [not found] <200809222238.m8MMcgAp001695@imap1.linux-foundation.org>
@ 2008-09-23 12:40 ` Oleg Nesterov
  2008-09-23 17:25   ` Chad Zanonie
  0 siblings, 1 reply; 4+ messages in thread
From: Oleg Nesterov @ 2008-09-23 12:40 UTC (permalink / raw)
  To: akpm; +Cc: linux-kernel, chad.zanonie, npiggin, rientjes

On 09/22, Andrew Morton wrote:
>
> ------------------------------------------------------
> Subject: mm: oom-killer kills more than needed
> From: "Chad Zanonie" <chad.zanonie@gmail.com>
>
> Possibility exists for an exiting application to be in between marking its
> mm NULL and calling mmput when out_of_memory is invoked. 
> select_bad_process() will continue past this process as opposed to
> returning -1UL due to its mm being NULL.  This causes the oom killer in
> certain scenarios to not only kill the memory culprit, but also kill the
> runner up.
>
> EXIT_DEAD seems to be the only flag that guarantees that mmput() has
> finished.

I don't think this is right.

Let's suppose we have a single zombie. Now select_bad_process() always
returns -1 ? IOW, doesn't this means that, say,

	$ perl -e 'fork && sleep'

disables oom-kill completely and forever?

Hmm. But please see below. This doesn't happen because the usage
of EXIT_DEAD is not right.

> Checking for PF_KTHREAD should replace p->mm regardless.

Yes, almost every check for ->mm in oom_kill.c is not right.

> Adding EXIT_DEAD to the check seems to prevent unnecessary kills in local
> testing.

This is strange, could you re-test? Because

> @@ -216,7 +216,7 @@ static struct task_struct *select_bad_pr
>  		 * skip kernel threads and tasks which have already released
>  		 * their mm.
>  		 */
> -		if (!p->mm)
> +		if (p->flags & PF_KTHREAD || p->flags & EXIT_DEAD)
                                                ^^^^^^^^^^^^^^^^^

this is not possible. EXIT_DEAD lives in ->exit_state, not in ->flags.

Oleg.


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: + mm-oom-killer-kills-more-than-needed.patch added to -mm tree
  2008-09-23 12:40 ` + mm-oom-killer-kills-more-than-needed.patch added to -mm tree Oleg Nesterov
@ 2008-09-23 17:25   ` Chad Zanonie
  2008-09-23 17:37     ` Chad Zanonie
  2008-09-24 16:58     ` Oleg Nesterov
  0 siblings, 2 replies; 4+ messages in thread
From: Chad Zanonie @ 2008-09-23 17:25 UTC (permalink / raw)
  To: Oleg Nesterov; +Cc: akpm, linux-kernel, npiggin, rientjes

On Tue, Sep 23, 2008 at 5:40 AM, Oleg Nesterov <oleg@tv-sign.ru> wrote:
> On 09/22, Andrew Morton wrote:
>>
>> ------------------------------------------------------
>> Subject: mm: oom-killer kills more than needed
>> From: "Chad Zanonie" <chad.zanonie@gmail.com>
>>
>> Possibility exists for an exiting application to be in between marking its
>> mm NULL and calling mmput when out_of_memory is invoked.
>> select_bad_process() will continue past this process as opposed to
>> returning -1UL due to its mm being NULL.  This causes the oom killer in
>> certain scenarios to not only kill the memory culprit, but also kill the
>> runner up.
>>
>> EXIT_DEAD seems to be the only flag that guarantees that mmput() has
>> finished.
>
> I don't think this is right.
>
> Let's suppose we have a single zombie. Now select_bad_process() always
> returns -1 ? IOW, doesn't this means that, say,
>
>        $ perl -e 'fork && sleep'
>
> disables oom-kill completely and forever?

Ugh, looks like my description was slightly incorrect. I don't mean
for it to return -1 upon noticing a null mm. I mean for
select_bad_process to not skip (continue) over processes that haven't
provably finished unmapping their memory (which EXIT_DEAD infers).

>
> Hmm. But please see below. This doesn't happen because the usage
> of EXIT_DEAD is not right.
>
>> Checking for PF_KTHREAD should replace p->mm regardless.
>
> Yes, almost every check for ->mm in oom_kill.c is not right.
>
>> Adding EXIT_DEAD to the check seems to prevent unnecessary kills in local
>> testing.
>
> This is strange, could you re-test? Because
>
>> @@ -216,7 +216,7 @@ static struct task_struct *select_bad_pr
>>                * skip kernel threads and tasks which have already released
>>                * their mm.
>>                */
>> -             if (!p->mm)
>> +             if (p->flags & PF_KTHREAD || p->flags & EXIT_DEAD)
>                                                ^^^^^^^^^^^^^^^^^
>
> this is not possible. EXIT_DEAD lives in ->exit_state, not in ->flags.

Good catch. I had actually done the testing in an older kernel that
used PF_DEAD and while noticing the change to only EXIT_DEAD forgot to
use the new appropriate flag in the patch.

>
> Oleg.
>
>

I'm new to this endeavor. Should I propose a new patch, or, will
things be fixed from this omission? I still support the EXIT_DEAD
inclusion, as it'll allow the OOM killer to make the similar progress
that !p->mm provided.

Chad

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: + mm-oom-killer-kills-more-than-needed.patch added to -mm tree
  2008-09-23 17:25   ` Chad Zanonie
@ 2008-09-23 17:37     ` Chad Zanonie
  2008-09-24 16:58     ` Oleg Nesterov
  1 sibling, 0 replies; 4+ messages in thread
From: Chad Zanonie @ 2008-09-23 17:37 UTC (permalink / raw)
  To: Oleg Nesterov; +Cc: akpm, linux-kernel, npiggin, rientjes

Before I propagate this blunder anymore, I've found the root of my mistake.

I really mean TASK_DEAD, not EXIT_DEAD.

(p->state & TASK_DEAD)

-Chad

On Tue, Sep 23, 2008 at 10:25 AM, Chad Zanonie <chad.zanonie@gmail.com> wrote:
> On Tue, Sep 23, 2008 at 5:40 AM, Oleg Nesterov <oleg@tv-sign.ru> wrote:
>> On 09/22, Andrew Morton wrote:
>>>
>>> ------------------------------------------------------
>>> Subject: mm: oom-killer kills more than needed
>>> From: "Chad Zanonie" <chad.zanonie@gmail.com>
>>>
>>> Possibility exists for an exiting application to be in between marking its
>>> mm NULL and calling mmput when out_of_memory is invoked.
>>> select_bad_process() will continue past this process as opposed to
>>> returning -1UL due to its mm being NULL.  This causes the oom killer in
>>> certain scenarios to not only kill the memory culprit, but also kill the
>>> runner up.
>>>
>>> EXIT_DEAD seems to be the only flag that guarantees that mmput() has
>>> finished.
>>
>> I don't think this is right.
>>
>> Let's suppose we have a single zombie. Now select_bad_process() always
>> returns -1 ? IOW, doesn't this means that, say,
>>
>>        $ perl -e 'fork && sleep'
>>
>> disables oom-kill completely and forever?
>
> Ugh, looks like my description was slightly incorrect. I don't mean
> for it to return -1 upon noticing a null mm. I mean for
> select_bad_process to not skip (continue) over processes that haven't
> provably finished unmapping their memory (which EXIT_DEAD infers).
>
>>
>> Hmm. But please see below. This doesn't happen because the usage
>> of EXIT_DEAD is not right.
>>
>>> Checking for PF_KTHREAD should replace p->mm regardless.
>>
>> Yes, almost every check for ->mm in oom_kill.c is not right.
>>
>>> Adding EXIT_DEAD to the check seems to prevent unnecessary kills in local
>>> testing.
>>
>> This is strange, could you re-test? Because
>>
>>> @@ -216,7 +216,7 @@ static struct task_struct *select_bad_pr
>>>                * skip kernel threads and tasks which have already released
>>>                * their mm.
>>>                */
>>> -             if (!p->mm)
>>> +             if (p->flags & PF_KTHREAD || p->flags & EXIT_DEAD)
>>                                                ^^^^^^^^^^^^^^^^^
>>
>> this is not possible. EXIT_DEAD lives in ->exit_state, not in ->flags.
>
> Good catch. I had actually done the testing in an older kernel that
> used PF_DEAD and while noticing the change to only EXIT_DEAD forgot to
> use the new appropriate flag in the patch.
>
>>
>> Oleg.
>>
>>
>
> I'm new to this endeavor. Should I propose a new patch, or, will
> things be fixed from this omission? I still support the EXIT_DEAD
> inclusion, as it'll allow the OOM killer to make the similar progress
> that !p->mm provided.
>
> Chad
>

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: + mm-oom-killer-kills-more-than-needed.patch added to -mm tree
  2008-09-23 17:25   ` Chad Zanonie
  2008-09-23 17:37     ` Chad Zanonie
@ 2008-09-24 16:58     ` Oleg Nesterov
  1 sibling, 0 replies; 4+ messages in thread
From: Oleg Nesterov @ 2008-09-24 16:58 UTC (permalink / raw)
  To: Chad Zanonie; +Cc: akpm, linux-kernel, npiggin, rientjes

On 09/23, Chad Zanonie wrote:
>
> On Tue, Sep 23, 2008 at 5:40 AM, Oleg Nesterov <oleg@tv-sign.ru> wrote:
> > On 09/22, Andrew Morton wrote:
> >>
> >> ------------------------------------------------------
> >> Subject: mm: oom-killer kills more than needed
> >> From: "Chad Zanonie" <chad.zanonie@gmail.com>
> >>
> >> Possibility exists for an exiting application to be in between marking its
> >> mm NULL and calling mmput when out_of_memory is invoked.
> >> select_bad_process() will continue past this process as opposed to
> >> returning -1UL due to its mm being NULL.  This causes the oom killer in
> >> certain scenarios to not only kill the memory culprit, but also kill the
> >> runner up.
> >>
> >> EXIT_DEAD seems to be the only flag that guarantees that mmput() has
> >> finished.
> >
> > I don't think this is right.
> >
> > Let's suppose we have a single zombie. Now select_bad_process() always
> > returns -1 ? IOW, doesn't this means that, say,
> >
> >        $ perl -e 'fork && sleep'
> >
> > disables oom-kill completely and forever?
>
> Ugh, looks like my description was slightly incorrect. I don't mean
> for it to return -1 upon noticing a null mm. I mean for
> select_bad_process to not skip (continue) over processes that haven't
> provably finished unmapping their memory

Yes I see. But if select_bad_process() sees the PF_EXITING task it returns
-1 (unless the task == current). So, with this change select_bad_process()
will return -1 more often, because it doesn't skip the tasks without ->mm.
And of course, !mm implies PF_EXITING.

I don't claim this is wrong. Unless we use EXIT_DEAD as the patch did, in
that case the 'fork && sleep' above really disables oom-kill.

I must admit I don't understand why this change is good but this does not
matter, I don't understand oom-kill anyway (but I think it has numerous
bugs ;).

> Before I propagate this blunder anymore, I've found the root of my mistake.
>
> I really mean TASK_DEAD, not EXIT_DEAD.
>
> (p->state & TASK_DEAD)

Yes, this should work. But I think this "defers" the decision too far.

You can check "p->exit_state != 0". But still this is a bit strange,
and needs a comment. For example, you can check p->files == NULL with
the same effect to verify that the task has already passed
exit_mm()->mmput().

The task can spend a lot of time before it sets TASK_DEAD or ->exit_state,
and again, during this time oom-kill will be "disabled". Contrary,
the window between "tsk->mm = NULL;" and mmput() in exit_mm() is very
small. Well, unless CONFIG_MM_OWNER.


In short, I can't judge this patch, but could you please improve the
changelog?

Oleg.


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2008-09-24 16:52 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <200809222238.m8MMcgAp001695@imap1.linux-foundation.org>
2008-09-23 12:40 ` + mm-oom-killer-kills-more-than-needed.patch added to -mm tree Oleg Nesterov
2008-09-23 17:25   ` Chad Zanonie
2008-09-23 17:37     ` Chad Zanonie
2008-09-24 16:58     ` Oleg Nesterov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox