From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1755389AbbITNJS (ORCPT ); Sun, 20 Sep 2015 09:09:18 -0400 Received: from mx1.redhat.com ([209.132.183.28]:35765 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1755263AbbITNJR (ORCPT ); Sun, 20 Sep 2015 09:09:17 -0400 Date: Sun, 20 Sep 2015 15:06:16 +0200 From: Oleg Nesterov To: Michal Hocko Cc: Linus Torvalds , Kyle Walker , Christoph Lameter , Andrew Morton , David Rientjes , Johannes Weiner , Vladimir Davydov , linux-mm , Linux Kernel Mailing List , Stanislav Kozina , Tetsuo Handa Subject: Re: can't oom-kill zap the victim's memory? Message-ID: <20150920130616.GB2104@redhat.com> References: <1442512783-14719-1-git-send-email-kwalker@redhat.com> <20150919150316.GB31952@redhat.com> <20150920093332.GA20562@dhcp22.suse.cz> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20150920093332.GA20562@dhcp22.suse.cz> User-Agent: Mutt/1.5.18 (2008-05-17) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 09/20, Michal Hocko wrote: > > On Sat 19-09-15 15:24:02, Linus Torvalds wrote: > > On Sat, Sep 19, 2015 at 8:03 AM, Oleg Nesterov wrote: > > > + > > > +static void oom_unmap_func(struct work_struct *work) > > > +{ > > > + struct mm_struct *mm = xchg(&oom_unmap_mm, NULL); > > > + > > > + if (!atomic_inc_not_zero(&mm->mm_users)) > > > + return; > > > + > > > + // If this is not safe we can do use_mm() + unuse_mm() > > > + down_read(&mm->mmap_sem); > > > > I don't think this is safe. > > > > What makes you sure that we might not deadlock on the mmap_sem here? > > For all we know, the process that is going out of memory is in the > > middle of a mmap(), and already holds the mmap_sem for writing. No? > > > > So at the very least that needs to be a trylock, I think. > > Agreed. Why? See my reply to Linus's email. Just in case, yes sure the unconditonal down_read() is suboptimal, but this is minor compared to other problems we need to solve. > > And I'm not > > sure zap_page_range() is ok with the mmap_sem only held for reading. > > Normally our rule is that you can *populate* the page tables > > concurrently, but you can't tear the down > > Actually mmap_sem for reading should be sufficient because we do not > alter the layout. Both MADV_DONTNEED and MADV_FREE require read mmap_sem > for example. Yes, but see the ->vm_flags check in madvise_dontneed(). Oleg.