All of lore.kernel.org
 help / color / mirror / Atom feed
From: Peter Xu <peterx@redhat.com>
To: Lorenzo Stoakes <lstoakes@gmail.com>
Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org,
	Matthew Wilcox <willy@infradead.org>,
	Andrea Arcangeli <aarcange@redhat.com>,
	John Hubbard <jhubbard@nvidia.com>,
	Mike Rapoport <rppt@kernel.org>,
	David Hildenbrand <david@redhat.com>,
	Vlastimil Babka <vbabka@suse.cz>,
	"Kirill A . Shutemov" <kirill@shutemov.name>,
	Andrew Morton <akpm@linux-foundation.org>,
	Mike Kravetz <mike.kravetz@oracle.com>,
	James Houghton <jthoughton@google.com>,
	Hugh Dickins <hughd@google.com>
Subject: Re: [PATCH 6/7] mm/gup: Accelerate thp gup even for "pages != NULL"
Date: Mon, 19 Jun 2023 15:37:30 -0400	[thread overview]
Message-ID: <ZJCuepgy3+66S03G@x1n> (raw)
In-Reply-To: <d8c76484-1030-44a3-b148-7e69fa84243a@lucifer.local>

On Sat, Jun 17, 2023 at 09:27:22PM +0100, Lorenzo Stoakes wrote:
> On Tue, Jun 13, 2023 at 05:53:45PM -0400, Peter Xu wrote:
> > The acceleration of THP was done with ctx.page_mask, however it'll be
> > ignored if **pages is non-NULL.
> >
> > The old optimization was introduced in 2013 in 240aadeedc4a ("mm:
> > accelerate mm_populate() treatment of THP pages").  It didn't explain why
> > we can't optimize the **pages non-NULL case.  It's possible that at that
> > time the major goal was for mm_populate() which should be enough back then.
> >
> > Optimize thp for all cases, by properly looping over each subpage, doing
> > cache flushes, and boost refcounts / pincounts where needed in one go.
> >
> > This can be verified using gup_test below:
> >
> >   # chrt -f 1 ./gup_test -m 512 -t -L -n 1024 -r 10
> >
> > Before:    13992.50 ( +-8.75%)
> > After:       378.50 (+-69.62%)
> >
> > Signed-off-by: Peter Xu <peterx@redhat.com>
> > ---
> >  mm/gup.c | 36 +++++++++++++++++++++++++++++-------
> >  1 file changed, 29 insertions(+), 7 deletions(-)
> >
> > diff --git a/mm/gup.c b/mm/gup.c
> > index a2d1b3c4b104..cdabc8ea783b 100644
> > --- a/mm/gup.c
> > +++ b/mm/gup.c
> > @@ -1210,16 +1210,38 @@ static long __get_user_pages(struct mm_struct *mm,
> >  			goto out;
> >  		}
> >  next_page:
> > -		if (pages) {
> > -			pages[i] = page;
> > -			flush_anon_page(vma, page, start);
> > -			flush_dcache_page(page);
> > -			ctx.page_mask = 0;
> > -		}
> > -
> >  		page_increm = 1 + (~(start >> PAGE_SHIFT) & ctx.page_mask);
> >  		if (page_increm > nr_pages)
> >  			page_increm = nr_pages;
> > +
> > +		if (pages) {
> > +			struct page *subpage;
> > +			unsigned int j;
> > +
> > +			/*
> > +			 * This must be a large folio (and doesn't need to
> > +			 * be the whole folio; it can be part of it), do
> > +			 * the refcount work for all the subpages too.
> > +			 * Since we already hold refcount on the head page,
> > +			 * it should never fail.
> > +			 *
> > +			 * NOTE: here the page may not be the head page
> > +			 * e.g. when start addr is not thp-size aligned.
> > +			 */
> > +			if (page_increm > 1)
> > +				WARN_ON_ONCE(
> > +				    try_grab_folio(compound_head(page),
> > +						   page_increm - 1,
> > +						   foll_flags) == NULL);
> 
> I'm not sure this should be warning but otherwise ignoring this returning
> NULL?  This feels like a case that could come up in realtiy,
> e.g. folio_ref_try_add_rcu() fails, or !folio_is_longterm_pinnable().

Note that we hold already at least 1 refcount on the folio (also mentioned
in the comment above this chunk of code), so both folio_ref_try_add_rcu()
and folio_is_longterm_pinnable() should already have been called on the
same folio and passed.  If it will fail it should have already, afaict.

I still don't see how that would trigger if the refcount won't overflow.

Here what I can do is still guard this try_grab_folio() and fail the GUP if
for any reason it failed.  Perhaps then it means I'll also keep that one
untouched in hugetlb_follow_page_mask() too.  But I suppose keeping the
WARN_ON_ONCE() seems still proper.

Thanks,

-- 
Peter Xu



  reply	other threads:[~2023-06-19 19:37 UTC|newest]

Thread overview: 39+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2023-06-13 21:53 [PATCH 0/7] mm/gup: Unify hugetlb, speed up thp Peter Xu
2023-06-13 21:53 ` [PATCH 1/7] mm/hugetlb: Handle FOLL_DUMP well in follow_page_mask() Peter Xu
2023-06-14 23:24   ` Mike Kravetz
2023-06-16  8:08   ` David Hildenbrand
2023-06-13 21:53 ` [PATCH 2/7] mm/hugetlb: Fix hugetlb_follow_page_mask() on permission checks Peter Xu
2023-06-14 15:31   ` David Hildenbrand
2023-06-14 15:46     ` Peter Xu
2023-06-14 15:57       ` David Hildenbrand
2023-06-15  0:11       ` Mike Kravetz
2023-06-13 21:53 ` [PATCH 3/7] mm/hugetlb: Add page_mask for hugetlb_follow_page_mask() Peter Xu
2023-06-15  0:17   ` Mike Kravetz
2023-06-16  8:11   ` David Hildenbrand
2023-06-19 21:43   ` Peter Xu
2023-06-20  7:01     ` David Hildenbrand
2023-06-20 14:40       ` Peter Xu
2023-06-13 21:53 ` [PATCH 4/7] mm/hugetlb: Prepare hugetlb_follow_page_mask() for FOLL_PIN Peter Xu
2023-06-14 14:57   ` David Hildenbrand
2023-06-14 15:11     ` Peter Xu
2023-06-14 15:17       ` David Hildenbrand
2023-06-14 15:31         ` Peter Xu
2023-06-14 15:47           ` David Hildenbrand
2023-06-14 15:51             ` Peter Xu
2023-06-15  0:25               ` Mike Kravetz
2023-06-15 19:42                 ` Peter Xu
2023-06-13 21:53 ` [PATCH 5/7] mm/gup: Cleanup next_page handling Peter Xu
2023-06-17 19:48   ` Lorenzo Stoakes
2023-06-17 20:00     ` Lorenzo Stoakes
2023-06-19 19:18       ` Peter Xu
2023-06-13 21:53 ` [PATCH 6/7] mm/gup: Accelerate thp gup even for "pages != NULL" Peter Xu
2023-06-14 14:58   ` Matthew Wilcox
2023-06-14 15:19     ` Peter Xu
2023-06-14 15:35       ` Peter Xu
2023-06-17 20:27   ` Lorenzo Stoakes
2023-06-19 19:37     ` Peter Xu [this message]
2023-06-19 20:24       ` Peter Xu
2023-06-13 21:53 ` [PATCH 7/7] mm/gup: Retire follow_hugetlb_page() Peter Xu
2023-06-14 14:37   ` Jason Gunthorpe
2023-06-17 20:40   ` Lorenzo Stoakes
2023-06-19 19:41     ` Peter Xu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ZJCuepgy3+66S03G@x1n \
    --to=peterx@redhat.com \
    --cc=aarcange@redhat.com \
    --cc=akpm@linux-foundation.org \
    --cc=david@redhat.com \
    --cc=hughd@google.com \
    --cc=jhubbard@nvidia.com \
    --cc=jthoughton@google.com \
    --cc=kirill@shutemov.name \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=lstoakes@gmail.com \
    --cc=mike.kravetz@oracle.com \
    --cc=rppt@kernel.org \
    --cc=vbabka@suse.cz \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.