From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 77481CD1299 for ; Wed, 10 Apr 2024 15:29:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender: Content-Transfer-Encoding:Content-Type:List-Subscribe:List-Help:List-Post: List-Archive:List-Unsubscribe:List-Id:In-Reply-To:MIME-Version:References: Message-ID:Subject:Cc:To:From:Date:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=hf9T91rZV3bTwwnsTYLbmI/lBQ33YOKSb5CXjP5CCHU=; b=RMpD1WgW3xXlAR Wxn1uqbnCNFn6u2atpk6+2BI3MwX4cWa/0nWcUAyCOYCFAneViDA7KtaIXt6r/yt7xWzekDz4SInm dEJwPX1+tPXtnAibOjUG0Jk0X1D7+b/jW3HTGKCdXU+QpC8jSo1Cc1clIoNsduoRHa+AZMOG7/Iyg BbrNl+/4Mdod4azX+MnwkSaHkJ7nkZc5WfnzqMItRHHXy/Qld+40w9/r1uc+1NlGfPnCsPtVuDB4e qm6at/8XSIiX74U2PkEY8p+VVZw/DM8aawYRY8yMHt9FUoWqhHKwljIo3GQhUIAnbuMu5u2aDz0en utJ25BCR9OfZ8xw/FZhg==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.97.1 #2 (Red Hat Linux)) id 1ruZsi-00000007mHv-1ELX; Wed, 10 Apr 2024 15:29:00 +0000 Received: from us-smtp-delivery-124.mimecast.com ([170.10.129.124]) by bombadil.infradead.org with esmtps (Exim 4.97.1 #2 (Red Hat Linux)) id 1ruZsc-00000007mDn-0WZt for linux-riscv@lists.infradead.org; Wed, 10 Apr 2024 15:28:57 +0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1712762933; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=eaActkU+0iuCIKN1UtzhnOPW3PrvAQBvA6rHVjQkAnbIXvivOMfb3Z8Vy+yZai2JefaURB cLCWrquvEUihEyo2WTKsG+e/XBAiAUjorWnTV3zH4KHap3aiuYRPnhOGfUc7rBujhLdN60 4yWHXiuAu7M+cQDWXhtZwSk9Kc12/80= Received: from mail-qt1-f198.google.com (mail-qt1-f198.google.com [209.85.160.198]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-491-IFaPWbhTPQOJnBHnrd9JKQ-1; Wed, 10 Apr 2024 11:28:49 -0400 X-MC-Unique: IFaPWbhTPQOJnBHnrd9JKQ-1 Received: by mail-qt1-f198.google.com with SMTP id d75a77b69052e-434ed2e412fso6707341cf.0 for ; Wed, 10 Apr 2024 08:28:43 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1712762923; x=1713367723; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=oxqZdf6wWG6AX6A/Igr3E6gVajd9NfNgIpfRLgB+dQPwKwC/dk7wXWvAednwjf2jiv 1dhOLcX7efGXt+khmwuCFqxQzU2v/HXKRmQbHBROaSaHicQbih8euI8g8Fv7OzozhoOd Ehkp41W0dkJ0dPWQDcBTVjdePa4BQuxu0HQxTZI8GOFB4g458ZTHrjNWn32AAJx81xYR sGjG9O5vjPMo2uVrzdg41Sq8/Y3EpkpxIZUUiwajNwUITIyaPbqckaqrEwp2w6ZkNtD1 gBZP3bS4vgdEsIEUYhweExdGxj4sB/lQ23pSCmbeZyPljUF2/xMWDUlMEv1Q/wNE+Lc7 iUbw== X-Forwarded-Encrypted: i=1; AJvYcCWGit7mkN6PW1/ydH+l9zLpCA6nOq6A/zkvbbKZZrUg8buVLTYTMMXiUl6C0xAp8nIYiuxoDPbvhaGXJSpDV4HntTYp57VadFWn0EsDlJ4K X-Gm-Message-State: AOJu0YxTNaLn8By1dOYxi/34JRTKI/nNJ8BhI9/x7EaVoj4fRKvSSBmO pXsC1JIHE725llAJOKNLiF7Kjz1iF7BzHuaGSC2eRtwpEJo3jZ3D33nd5tXEN+2p6ybeJiT1bT/ LyRvufuAvrnnZiQiNY8XqFoYlEGdEWtEXQw47ml0XRSBigI+sSPQubQWInru1mkb87A== X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168084qvb.2.1712762923245; Wed, 10 Apr 2024 08:28:43 -0700 (PDT) X-Google-Smtp-Source: AGHT+IGZs5MCkL0IJwntJvXwMvElTS5kl/yk1ExPO6tDqr2NK62DfB6pdbW0p2/DjTaMvJz9lDcKPw== X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168047qvb.2.1712762922544; Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Received: from x1n (pool-99-254-121-117.cpe.net.cable.rogers.com. [99.254.121.117]) by smtp.gmail.com with ESMTPSA id ek1-20020ad45981000000b00699032e555bsm5208543qvb.127.2024.04.10.08.28.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Date: Wed, 10 Apr 2024 11:28:39 -0400 From: Peter Xu To: Jason Gunthorpe Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, Michael Ellerman , Christophe Leroy , Matthew Wilcox , Rik van Riel , Lorenzo Stoakes , Axel Rasmussen , Yang Shi , John Hubbard , linux-arm-kernel@lists.infradead.org, "Kirill A . Shutemov" , Andrew Jones , Vlastimil Babka , Mike Rapoport , Andrew Morton , Muchun Song , Christoph Hellwig , linux-riscv@lists.infradead.org, James Houghton , David Hildenbrand , Andrea Arcangeli , "Aneesh Kumar K . V" , Mike Kravetz Subject: Re: [PATCH v3 00/12] mm/gup: Unify hugetlb, part 2 Message-ID: References: <20240321220802.679544-1-peterx@redhat.com> <20240322161000.GJ159172@nvidia.com> <20240326140252.GH6245@nvidia.com> <20240405181633.GH5383@nvidia.com> <20240409234355.GJ5383@nvidia.com> MIME-Version: 1.0 In-Reply-To: <20240409234355.GJ5383@nvidia.com> X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com Content-Disposition: inline X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20240410_082854_264163_7D4F3884 X-CRM114-Status: GOOD ( 37.94 ) X-BeenThere: linux-riscv@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: "linux-riscv" Errors-To: linux-riscv-bounces+linux-riscv=archiver.kernel.org@lists.infradead.org On Tue, Apr 09, 2024 at 08:43:55PM -0300, Jason Gunthorpe wrote: > On Fri, Apr 05, 2024 at 05:42:44PM -0400, Peter Xu wrote: > > In short, hugetlb mappings shouldn't be special comparing to other huge pXd > > and large folio (cont-pXd) mappings for most of the walkers in my mind, if > > not all. I need to look at all the walkers and there can be some tricky > > ones, but I believe that applies in general. It's actually similar to what > > I did with slow gup here. > > I think that is the big question, I also haven't done the research to > know the answer. > > At this point focusing on moving what is reasonable to the pXX_* API > makes sense to me. Then reviewing what remains and making some > decision. > > > Like this series, for cont-pXd we'll need multiple walks comparing to > > before (when with hugetlb_entry()), but for that part I'll provide some > > performance tests too, and we also have a fallback plan, which is to detect > > cont-pXd existance, which will also work for large folios. > > I think we can optimize this pretty easy. > > > > I think if you do the easy places for pXX conversion you will have a > > > good idea about what is needed for the hard places. > > > > Here IMHO we don't need to understand "what is the size of this hugetlb > > vma" > > Yeh, I never really understood why hugetlb was linked to the VMA.. The > page table is self describing, obviously. Attaching to vma still makes sense to me, where we should definitely avoid a mixture of hugetlb and !hugetlb pages in a single vma - hugetlb pages are allocated, managed, ... totally differently. And since hugetlb is designed as file-based (which also makes sense to me, at least for now), it's also natural that it's vma-attached. > > > or "which level of pgtable does this hugetlb vma pages locate", > > Ditto > > > because we may not need that, e.g., when we only want to collect some smaps > > statistics. "whether it's hugetlb" may matter, though. E.g. in the mm > > walker we see a huge pmd, it can be a thp, it can be a hugetlb (when > > hugetlb_entry removed), we may need extra check later to put things into > > the right bucket, but for the walker itself it doesn't necessarily need > > hugetlb_entry(). > > Right, places may still need to know it is part of a huge VMA because we > have special stuff linked to that. > > > > But then again we come back to power and its big list of page sizes > > > and variety :( Looks like some there have huge sizes at the pgd level > > > at least. > > > > Yeah this is something I want to be super clear, because I may miss > > something: we don't have real pgd pages, right? Powerpc doesn't even > > define p4d_leaf(), AFAICT. > > AFAICT it is because it hides it all in hugepd. IMHO one thing we can benefit from such hugepd rework is, if we can squash all the hugepds like what Christophe does, then we push it one more layer down, and we have a good chance all things should just work. So again my Power brain is close to zero, but now I'm referring to what Christophe shared in the other thread: https://github.com/linuxppc/wiki/wiki/Huge-pages Together with: https://lore.kernel.org/r/288f26f487648d21fd9590e40b390934eaa5d24a.1711377230.git.christophe.leroy@csgroup.eu Where it has: --- a/arch/powerpc/platforms/Kconfig.cputype +++ b/arch/powerpc/platforms/Kconfig.cputype @@ -98,6 +98,7 @@ config PPC_BOOK3S_64 select ARCH_ENABLE_HUGEPAGE_MIGRATION if HUGETLB_PAGE && MIGRATION select ARCH_ENABLE_SPLIT_PMD_PTLOCK select ARCH_ENABLE_THP_MIGRATION if TRANSPARENT_HUGEPAGE + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS select ARCH_SUPPORTS_NUMA_BALANCING select HAVE_MOVE_PMD @@ -290,6 +291,7 @@ config PPC_BOOK3S config PPC_E500 select FSL_EMB_PERFMON bool + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS if PHYS_64BIT || PPC64 select PPC_SMP_MUXED_IPI select PPC_DOORBELL So I think it means we have three PowerPC systems that supports hugepd right now (besides the 8xx which Christophe is trying to drop support there), besides 8xx we still have book3s_64 and E500. Let's check one by one: - book3s_64 - hash - 64K: p4d is not used, largest pgsize pgd 16G @pud level. It means after squashing it'll be a bunch of cont-pmd, all good. - 4K: p4d also not used, largest pgsize pgd 128G, after squashed it'll be cont-pud. all good. - radix - 64K: largest 1G @pud, then cont-pmd after squashed. all good. - 4K: largest 1G @pud, then cont-pmd, all good. - e500 & 8xx - both of them use 2-level pgtables (pgd + pte), after squashed hugepd @pgd level they become cont-pte. all good. I think the trick here is there'll be no pgd leaves after hugepd squashing to lower levels, then since PowerPC seems to never have p4d, then all things fall into pud or lower. We seem to be all good there? > > If the goal is to purge hugepd then some of the options might turn out > to convert hugepd into huge p4d/pgd, as I understand it. It would be > nice to have certainty on this at least. Right. I hope the pmd/pud plan I proposed above can already work too with such ambicious goal too. But review very welcomed from either you or Christophe. PS: I think I'll also have a closer look at Christophe's series this week or next. > > We have effectively three APIs to parse a single page table and > currently none of the APIs can return 100% of the data for power. Thanks, -- Peter Xu _______________________________________________ linux-riscv mailing list linux-riscv@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-riscv From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from lists.ozlabs.org (lists.ozlabs.org [112.213.38.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 02403CD128A for ; Wed, 10 Apr 2024 15:29:33 +0000 (UTC) Authentication-Results: lists.ozlabs.org; dkim=fail reason="signature verification failed" (1024-bit key; unprotected) header.d=redhat.com header.i=@redhat.com header.a=rsa-sha256 header.s=mimecast20190719 header.b=H46YB+to; dkim=fail reason="signature verification failed" (1024-bit key) header.d=redhat.com header.i=@redhat.com header.a=rsa-sha256 header.s=mimecast20190719 header.b=H46YB+to; dkim-atps=neutral Received: from boromir.ozlabs.org (localhost [IPv6:::1]) by lists.ozlabs.org (Postfix) with ESMTP id 4VF6FS4hKvz3dTr for ; Thu, 11 Apr 2024 01:29:32 +1000 (AEST) Authentication-Results: lists.ozlabs.org; dkim=pass (1024-bit key; unprotected) header.d=redhat.com header.i=@redhat.com header.a=rsa-sha256 header.s=mimecast20190719 header.b=H46YB+to; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.a=rsa-sha256 header.s=mimecast20190719 header.b=H46YB+to; dkim-atps=neutral Authentication-Results: lists.ozlabs.org; spf=pass (sender SPF authorized) smtp.mailfrom=redhat.com (client-ip=170.10.129.124; helo=us-smtp-delivery-124.mimecast.com; envelope-from=peterx@redhat.com; receiver=lists.ozlabs.org) Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by lists.ozlabs.org (Postfix) with ESMTPS id 4VF6Df6VCFz3bs0 for ; Thu, 11 Apr 2024 01:28:50 +1000 (AEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1712762926; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=H46YB+to2jmBuWQ8iXD81PAplnTxQhy4xP2ps2K6cIakExX1lAYR1QbCjrKkfdm14OYJuz zZ+wwV4coKm3fH43pOC9FYmSDcqWZUkOPVDSoiDve7YLw0A1KVNzHwh1mHsAUMy6cQpjaJ aFSgw0SDCJKHqubrgyjsd7QJS7qbtcI= DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1712762926; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=H46YB+to2jmBuWQ8iXD81PAplnTxQhy4xP2ps2K6cIakExX1lAYR1QbCjrKkfdm14OYJuz zZ+wwV4coKm3fH43pOC9FYmSDcqWZUkOPVDSoiDve7YLw0A1KVNzHwh1mHsAUMy6cQpjaJ aFSgw0SDCJKHqubrgyjsd7QJS7qbtcI= Received: from mail-qt1-f198.google.com (mail-qt1-f198.google.com [209.85.160.198]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-413-39DLB4ttPpeIKwi0AN_y5Q-1; Wed, 10 Apr 2024 11:28:44 -0400 X-MC-Unique: 39DLB4ttPpeIKwi0AN_y5Q-1 Received: by mail-qt1-f198.google.com with SMTP id d75a77b69052e-4311d908f3cso29912151cf.1 for ; Wed, 10 Apr 2024 08:28:44 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1712762923; x=1713367723; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=nNYXaELdTuICuTy0uYyN1nq1A5Fbi8Fs8yHm0MA8l9MRhGuQTdi+OoRoWy2qOAAcn1 gvJpeRPun4e0RmhL/29chUSchOkezjiMwQNK/8sELThHxHSYAJY19sXHhq2HT0z1EG0u bKR3DkQ6ha9HiUzQKN8GnkVJE+a6KYBSIH2ihHjKES/EoC5LI+SnWvQDZD8JASMQbEx8 WW+dk7dT8y73Rz0JkPdLDkN02NHHj6XLOnxXsYT93mWqAhOyRe6smr0H3Cd7h3La0QvY Ex3QqLMK89FV3jUbCFphpL4Q3MKQ+/EMshKCub1qchQiCHm5tN5G0bfx92IIVV569GLs Aj1Q== X-Forwarded-Encrypted: i=1; AJvYcCVzB1Ti5p2lb5rM17YHMy05htDDCnAhvLR4U50HWPyBuFI1XDUtaqlOOiUwqMFXriWt6t66J5lL3n6+8Vf2SEhs0KieFKX2+idtbWzNRQ== X-Gm-Message-State: AOJu0YwxWDQ2ajuIt2NVDFGOVEYL6Z9Y+IJHGeZ38vkms0EdqfqdSWGj hXLmG06jAxUzSHNtGM62D8sPVO1wrnGUJlsytHxFbF4sfA29luM6VrJ9bO0BPpw2pl5vVR2KsPe J/5FT0BrwZyJqNZZDQpK65lPDrKpvQmhheuYNlrR7ASmBC3AZr+DOaX95FB4c9fk= X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168086qvb.2.1712762923247; Wed, 10 Apr 2024 08:28:43 -0700 (PDT) X-Google-Smtp-Source: AGHT+IGZs5MCkL0IJwntJvXwMvElTS5kl/yk1ExPO6tDqr2NK62DfB6pdbW0p2/DjTaMvJz9lDcKPw== X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168047qvb.2.1712762922544; Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Received: from x1n (pool-99-254-121-117.cpe.net.cable.rogers.com. [99.254.121.117]) by smtp.gmail.com with ESMTPSA id ek1-20020ad45981000000b00699032e555bsm5208543qvb.127.2024.04.10.08.28.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Date: Wed, 10 Apr 2024 11:28:39 -0400 From: Peter Xu To: Jason Gunthorpe Subject: Re: [PATCH v3 00/12] mm/gup: Unify hugetlb, part 2 Message-ID: References: <20240321220802.679544-1-peterx@redhat.com> <20240322161000.GJ159172@nvidia.com> <20240326140252.GH6245@nvidia.com> <20240405181633.GH5383@nvidia.com> <20240409234355.GJ5383@nvidia.com> MIME-Version: 1.0 In-Reply-To: <20240409234355.GJ5383@nvidia.com> X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=utf-8 Content-Disposition: inline X-BeenThere: linuxppc-dev@lists.ozlabs.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: Linux on PowerPC Developers Mail List List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Cc: James Houghton , David Hildenbrand , Yang Shi , Andrew Jones , linux-mm@kvack.org, linux-riscv@lists.infradead.org, Andrea Arcangeli , "Aneesh Kumar K . V" , Matthew Wilcox , Christoph Hellwig , linux-arm-kernel@lists.infradead.org, Axel Rasmussen , Rik van Riel , John Hubbard , "Kirill A . Shutemov" , Vlastimil Babka , Lorenzo Stoakes , Muchun Song , linux-kernel@vger.kernel.org, Andrew Morton , linuxppc-dev@lists.ozlabs.org, Mike Rapoport , Mike Kravetz Errors-To: linuxppc-dev-bounces+linuxppc-dev=archiver.kernel.org@lists.ozlabs.org Sender: "Linuxppc-dev" On Tue, Apr 09, 2024 at 08:43:55PM -0300, Jason Gunthorpe wrote: > On Fri, Apr 05, 2024 at 05:42:44PM -0400, Peter Xu wrote: > > In short, hugetlb mappings shouldn't be special comparing to other huge pXd > > and large folio (cont-pXd) mappings for most of the walkers in my mind, if > > not all. I need to look at all the walkers and there can be some tricky > > ones, but I believe that applies in general. It's actually similar to what > > I did with slow gup here. > > I think that is the big question, I also haven't done the research to > know the answer. > > At this point focusing on moving what is reasonable to the pXX_* API > makes sense to me. Then reviewing what remains and making some > decision. > > > Like this series, for cont-pXd we'll need multiple walks comparing to > > before (when with hugetlb_entry()), but for that part I'll provide some > > performance tests too, and we also have a fallback plan, which is to detect > > cont-pXd existance, which will also work for large folios. > > I think we can optimize this pretty easy. > > > > I think if you do the easy places for pXX conversion you will have a > > > good idea about what is needed for the hard places. > > > > Here IMHO we don't need to understand "what is the size of this hugetlb > > vma" > > Yeh, I never really understood why hugetlb was linked to the VMA.. The > page table is self describing, obviously. Attaching to vma still makes sense to me, where we should definitely avoid a mixture of hugetlb and !hugetlb pages in a single vma - hugetlb pages are allocated, managed, ... totally differently. And since hugetlb is designed as file-based (which also makes sense to me, at least for now), it's also natural that it's vma-attached. > > > or "which level of pgtable does this hugetlb vma pages locate", > > Ditto > > > because we may not need that, e.g., when we only want to collect some smaps > > statistics. "whether it's hugetlb" may matter, though. E.g. in the mm > > walker we see a huge pmd, it can be a thp, it can be a hugetlb (when > > hugetlb_entry removed), we may need extra check later to put things into > > the right bucket, but for the walker itself it doesn't necessarily need > > hugetlb_entry(). > > Right, places may still need to know it is part of a huge VMA because we > have special stuff linked to that. > > > > But then again we come back to power and its big list of page sizes > > > and variety :( Looks like some there have huge sizes at the pgd level > > > at least. > > > > Yeah this is something I want to be super clear, because I may miss > > something: we don't have real pgd pages, right? Powerpc doesn't even > > define p4d_leaf(), AFAICT. > > AFAICT it is because it hides it all in hugepd. IMHO one thing we can benefit from such hugepd rework is, if we can squash all the hugepds like what Christophe does, then we push it one more layer down, and we have a good chance all things should just work. So again my Power brain is close to zero, but now I'm referring to what Christophe shared in the other thread: https://github.com/linuxppc/wiki/wiki/Huge-pages Together with: https://lore.kernel.org/r/288f26f487648d21fd9590e40b390934eaa5d24a.1711377230.git.christophe.leroy@csgroup.eu Where it has: --- a/arch/powerpc/platforms/Kconfig.cputype +++ b/arch/powerpc/platforms/Kconfig.cputype @@ -98,6 +98,7 @@ config PPC_BOOK3S_64 select ARCH_ENABLE_HUGEPAGE_MIGRATION if HUGETLB_PAGE && MIGRATION select ARCH_ENABLE_SPLIT_PMD_PTLOCK select ARCH_ENABLE_THP_MIGRATION if TRANSPARENT_HUGEPAGE + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS select ARCH_SUPPORTS_NUMA_BALANCING select HAVE_MOVE_PMD @@ -290,6 +291,7 @@ config PPC_BOOK3S config PPC_E500 select FSL_EMB_PERFMON bool + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS if PHYS_64BIT || PPC64 select PPC_SMP_MUXED_IPI select PPC_DOORBELL So I think it means we have three PowerPC systems that supports hugepd right now (besides the 8xx which Christophe is trying to drop support there), besides 8xx we still have book3s_64 and E500. Let's check one by one: - book3s_64 - hash - 64K: p4d is not used, largest pgsize pgd 16G @pud level. It means after squashing it'll be a bunch of cont-pmd, all good. - 4K: p4d also not used, largest pgsize pgd 128G, after squashed it'll be cont-pud. all good. - radix - 64K: largest 1G @pud, then cont-pmd after squashed. all good. - 4K: largest 1G @pud, then cont-pmd, all good. - e500 & 8xx - both of them use 2-level pgtables (pgd + pte), after squashed hugepd @pgd level they become cont-pte. all good. I think the trick here is there'll be no pgd leaves after hugepd squashing to lower levels, then since PowerPC seems to never have p4d, then all things fall into pud or lower. We seem to be all good there? > > If the goal is to purge hugepd then some of the options might turn out > to convert hugepd into huge p4d/pgd, as I understand it. It would be > nice to have certainty on this at least. Right. I hope the pmd/pud plan I proposed above can already work too with such ambicious goal too. But review very welcomed from either you or Christophe. PS: I think I'll also have a closer look at Christophe's series this week or next. > > We have effectively three APIs to parse a single page table and > currently none of the APIs can return 100% of the data for power. Thanks, -- Peter Xu From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 9E727CD128A for ; Wed, 10 Apr 2024 15:29:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender: Content-Transfer-Encoding:Content-Type:List-Subscribe:List-Help:List-Post: List-Archive:List-Unsubscribe:List-Id:In-Reply-To:MIME-Version:References: Message-ID:Subject:Cc:To:From:Date:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=zEEaFVgXAhftKhibMI8/uHWQPRDITsAST5TuK0wdeC8=; b=1sCbd8czgt8MIS +eXAXKzgaEwOUFSgPKoTEOnIv7PmwSJrs4PTPylxIR7313mgGKDoCBZpOMxf0SdgGLThrvYBag/Y/ JcXNhEfuGosLrNJjrEqUSVsPOCKxLZ7yp1+xL3VhLZ+J+p91LD0UCK/PTERsII8VWZQ/XPpdwhfRd /UELO9BaQenMHAHmJbBLanzFSc83d+UycBn1HteuUW9ahLWEgmHAto5vDCF3azdz4n67ho6nGWwQa JNTOh0xq2UmJr7kKgIj9mwbRK/Jkh56tH19DhtOuvL7EfwFMmH8E4ls9Li5+g+lYYCMCsZXGXtC2L 50J37mTodMKANoxTDgMw==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.97.1 #2 (Red Hat Linux)) id 1ruZsh-00000007mHF-0aGs; Wed, 10 Apr 2024 15:28:59 +0000 Received: from us-smtp-delivery-124.mimecast.com ([170.10.133.124]) by bombadil.infradead.org with esmtps (Exim 4.97.1 #2 (Red Hat Linux)) id 1ruZsb-00000007mCJ-2j3w for linux-arm-kernel@lists.infradead.org; Wed, 10 Apr 2024 15:28:56 +0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1712762925; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=fKfXx6daHhl5PrK4E9aFrZnlTgfulS/1gTixVI5lppIUn4QLmTYyjG7U4WgxMei2HJbxBT 7SVDXcWl22mWqiKidMiB7KG1dnUUDPhhg9qCiyxMp65QfNZ8hvoqNH50JXwt6dWgt5YVGi M08HmR8HVH0LI9O+xUtY7i/e5LucBTw= Received: from mail-qt1-f200.google.com (mail-qt1-f200.google.com [209.85.160.200]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-402-Ex2iEFlkPjKktPWI30M6MQ-1; Wed, 10 Apr 2024 11:28:44 -0400 X-MC-Unique: Ex2iEFlkPjKktPWI30M6MQ-1 Received: by mail-qt1-f200.google.com with SMTP id d75a77b69052e-434ed2e412fso6707381cf.0 for ; Wed, 10 Apr 2024 08:28:44 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1712762923; x=1713367723; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=AlsgRp2gTTN/9HZOjS78kxXHOpswAVVXdp3OGPfkIQBgY91AnlwmC/hSCUzSiCrie4 Lafd86PiXwVI1z0IcCSZ08KV1Vzqj0usUXXWPYSPJIF35QB+ILGh85DsWyRB8/9LXRdm +CJoi5IMQsi2ZS16VCDLJgBIWbo1wNXIU91MCw/ZVjpoNVxnIbW1RzFJvkhllhF+zeGZ dIZu2qNHw1sA6ZqtxwOlWtGLVYwE4xDFoa37F95Kf4v7iZS1dxXIP65ZhjG39l0NVpK+ HEK5uy7X+sbtBbY6FVwpIblK+0U+lL/3L0YKrke/RmLXzca9Mo4sjGBVmqu5fQgDBidk YImw== X-Forwarded-Encrypted: i=1; AJvYcCU0mVx4Ibs5Ckkdkddhh9w9V47R6SQCCZBCOQr7OT8rmcNHcphiMOah1ch2eN5kpv1x2aOU7qpZv+vYMznL25VZCVwPcLyyWlT023E1xK9a1+YzWBU= X-Gm-Message-State: AOJu0YxjJMpH3DDKwoDGgsuL69ySrFJEokUbD1ugSDLu9xMnngmjNqgY 6AvBRT7xfRxUnLVOFuPJGhYDI2d3n0aPHva0buya9E8C+k4GFan38annn6eSYS21Vol2pohS6qZ DP1AQNoO5vT4q4m3/P+SKys05Aexb4D2idmH9fWyTzsuZ2bSl/Vz+wGiZmMoy0mEuxLFAz1i9 X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168097qvb.2.1712762923256; Wed, 10 Apr 2024 08:28:43 -0700 (PDT) X-Google-Smtp-Source: AGHT+IGZs5MCkL0IJwntJvXwMvElTS5kl/yk1ExPO6tDqr2NK62DfB6pdbW0p2/DjTaMvJz9lDcKPw== X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168047qvb.2.1712762922544; Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Received: from x1n (pool-99-254-121-117.cpe.net.cable.rogers.com. [99.254.121.117]) by smtp.gmail.com with ESMTPSA id ek1-20020ad45981000000b00699032e555bsm5208543qvb.127.2024.04.10.08.28.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Date: Wed, 10 Apr 2024 11:28:39 -0400 From: Peter Xu To: Jason Gunthorpe Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, Michael Ellerman , Christophe Leroy , Matthew Wilcox , Rik van Riel , Lorenzo Stoakes , Axel Rasmussen , Yang Shi , John Hubbard , linux-arm-kernel@lists.infradead.org, "Kirill A . Shutemov" , Andrew Jones , Vlastimil Babka , Mike Rapoport , Andrew Morton , Muchun Song , Christoph Hellwig , linux-riscv@lists.infradead.org, James Houghton , David Hildenbrand , Andrea Arcangeli , "Aneesh Kumar K . V" , Mike Kravetz Subject: Re: [PATCH v3 00/12] mm/gup: Unify hugetlb, part 2 Message-ID: References: <20240321220802.679544-1-peterx@redhat.com> <20240322161000.GJ159172@nvidia.com> <20240326140252.GH6245@nvidia.com> <20240405181633.GH5383@nvidia.com> <20240409234355.GJ5383@nvidia.com> MIME-Version: 1.0 In-Reply-To: <20240409234355.GJ5383@nvidia.com> X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com Content-Disposition: inline X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20240410_082853_779175_2D35B51F X-CRM114-Status: GOOD ( 39.60 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Tue, Apr 09, 2024 at 08:43:55PM -0300, Jason Gunthorpe wrote: > On Fri, Apr 05, 2024 at 05:42:44PM -0400, Peter Xu wrote: > > In short, hugetlb mappings shouldn't be special comparing to other huge pXd > > and large folio (cont-pXd) mappings for most of the walkers in my mind, if > > not all. I need to look at all the walkers and there can be some tricky > > ones, but I believe that applies in general. It's actually similar to what > > I did with slow gup here. > > I think that is the big question, I also haven't done the research to > know the answer. > > At this point focusing on moving what is reasonable to the pXX_* API > makes sense to me. Then reviewing what remains and making some > decision. > > > Like this series, for cont-pXd we'll need multiple walks comparing to > > before (when with hugetlb_entry()), but for that part I'll provide some > > performance tests too, and we also have a fallback plan, which is to detect > > cont-pXd existance, which will also work for large folios. > > I think we can optimize this pretty easy. > > > > I think if you do the easy places for pXX conversion you will have a > > > good idea about what is needed for the hard places. > > > > Here IMHO we don't need to understand "what is the size of this hugetlb > > vma" > > Yeh, I never really understood why hugetlb was linked to the VMA.. The > page table is self describing, obviously. Attaching to vma still makes sense to me, where we should definitely avoid a mixture of hugetlb and !hugetlb pages in a single vma - hugetlb pages are allocated, managed, ... totally differently. And since hugetlb is designed as file-based (which also makes sense to me, at least for now), it's also natural that it's vma-attached. > > > or "which level of pgtable does this hugetlb vma pages locate", > > Ditto > > > because we may not need that, e.g., when we only want to collect some smaps > > statistics. "whether it's hugetlb" may matter, though. E.g. in the mm > > walker we see a huge pmd, it can be a thp, it can be a hugetlb (when > > hugetlb_entry removed), we may need extra check later to put things into > > the right bucket, but for the walker itself it doesn't necessarily need > > hugetlb_entry(). > > Right, places may still need to know it is part of a huge VMA because we > have special stuff linked to that. > > > > But then again we come back to power and its big list of page sizes > > > and variety :( Looks like some there have huge sizes at the pgd level > > > at least. > > > > Yeah this is something I want to be super clear, because I may miss > > something: we don't have real pgd pages, right? Powerpc doesn't even > > define p4d_leaf(), AFAICT. > > AFAICT it is because it hides it all in hugepd. IMHO one thing we can benefit from such hugepd rework is, if we can squash all the hugepds like what Christophe does, then we push it one more layer down, and we have a good chance all things should just work. So again my Power brain is close to zero, but now I'm referring to what Christophe shared in the other thread: https://github.com/linuxppc/wiki/wiki/Huge-pages Together with: https://lore.kernel.org/r/288f26f487648d21fd9590e40b390934eaa5d24a.1711377230.git.christophe.leroy@csgroup.eu Where it has: --- a/arch/powerpc/platforms/Kconfig.cputype +++ b/arch/powerpc/platforms/Kconfig.cputype @@ -98,6 +98,7 @@ config PPC_BOOK3S_64 select ARCH_ENABLE_HUGEPAGE_MIGRATION if HUGETLB_PAGE && MIGRATION select ARCH_ENABLE_SPLIT_PMD_PTLOCK select ARCH_ENABLE_THP_MIGRATION if TRANSPARENT_HUGEPAGE + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS select ARCH_SUPPORTS_NUMA_BALANCING select HAVE_MOVE_PMD @@ -290,6 +291,7 @@ config PPC_BOOK3S config PPC_E500 select FSL_EMB_PERFMON bool + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS if PHYS_64BIT || PPC64 select PPC_SMP_MUXED_IPI select PPC_DOORBELL So I think it means we have three PowerPC systems that supports hugepd right now (besides the 8xx which Christophe is trying to drop support there), besides 8xx we still have book3s_64 and E500. Let's check one by one: - book3s_64 - hash - 64K: p4d is not used, largest pgsize pgd 16G @pud level. It means after squashing it'll be a bunch of cont-pmd, all good. - 4K: p4d also not used, largest pgsize pgd 128G, after squashed it'll be cont-pud. all good. - radix - 64K: largest 1G @pud, then cont-pmd after squashed. all good. - 4K: largest 1G @pud, then cont-pmd, all good. - e500 & 8xx - both of them use 2-level pgtables (pgd + pte), after squashed hugepd @pgd level they become cont-pte. all good. I think the trick here is there'll be no pgd leaves after hugepd squashing to lower levels, then since PowerPC seems to never have p4d, then all things fall into pud or lower. We seem to be all good there? > > If the goal is to purge hugepd then some of the options might turn out > to convert hugepd into huge p4d/pgd, as I understand it. It would be > nice to have certainty on this at least. Right. I hope the pmd/pud plan I proposed above can already work too with such ambicious goal too. But review very welcomed from either you or Christophe. PS: I think I'll also have a closer look at Christophe's series this week or next. > > We have effectively three APIs to parse a single page table and > currently none of the APIs can return 100% of the data for power. Thanks, -- Peter Xu _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1EB42CD128A for ; Wed, 10 Apr 2024 15:28:49 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id A67B16B00BA; Wed, 10 Apr 2024 11:28:48 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id A17406B00BB; Wed, 10 Apr 2024 11:28:48 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 8B9246B00BC; Wed, 10 Apr 2024 11:28:48 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0016.hostedemail.com [216.40.44.16]) by kanga.kvack.org (Postfix) with ESMTP id 694D76B00BA for ; Wed, 10 Apr 2024 11:28:48 -0400 (EDT) Received: from smtpin11.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay07.hostedemail.com (Postfix) with ESMTP id 247671607B2 for ; Wed, 10 Apr 2024 15:28:48 +0000 (UTC) X-FDA: 81994004736.11.AB803A6 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) by imf08.hostedemail.com (Postfix) with ESMTP id 1ABFF16000D for ; Wed, 10 Apr 2024 15:28:45 +0000 (UTC) Authentication-Results: imf08.hostedemail.com; dkim=pass header.d=redhat.com header.s=mimecast20190719 header.b=fKfXx6da; dmarc=pass (policy=none) header.from=redhat.com; spf=pass (imf08.hostedemail.com: domain of peterx@redhat.com designates 170.10.133.124 as permitted sender) smtp.mailfrom=peterx@redhat.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1712762926; a=rsa-sha256; cv=none; b=7QkJqrZTDtrVhU65ZT78ujDr2qEqGJv9iThOG64ic2cJVCiSBZ2BsYaSCkqDGL5i072RXP lBEfJpbeu2oLUBnJNqdTodMSR/w4GzMEXUAASRRFfpAn46WP/5uA2JnuuOPy9D0D30gcvY quENEp9xSNX2pZvCRtqoIIUcsCPk5NA= ARC-Authentication-Results: i=1; imf08.hostedemail.com; dkim=pass header.d=redhat.com header.s=mimecast20190719 header.b=fKfXx6da; dmarc=pass (policy=none) header.from=redhat.com; spf=pass (imf08.hostedemail.com: domain of peterx@redhat.com designates 170.10.133.124 as permitted sender) smtp.mailfrom=peterx@redhat.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1712762926; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=zP5oBHnc7zi7KisH2rHeuW55nB75sgw0Yy9RsHsscJPy2YkSRVQFicoxVlmvJ0vb7UmyhV xqSSZvZpxJB5ah1dCs9fTqC/i3zKI8JW/luSuC8HP8vxhMzGh3oWt1x7H3dP/7gALz+ywT EgGYGuygumCfwKBf7JsORbQaVhGMGBw= DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1712762925; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=fKfXx6daHhl5PrK4E9aFrZnlTgfulS/1gTixVI5lppIUn4QLmTYyjG7U4WgxMei2HJbxBT 7SVDXcWl22mWqiKidMiB7KG1dnUUDPhhg9qCiyxMp65QfNZ8hvoqNH50JXwt6dWgt5YVGi M08HmR8HVH0LI9O+xUtY7i/e5LucBTw= Received: from mail-qt1-f198.google.com (mail-qt1-f198.google.com [209.85.160.198]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-402-pPkIg4aBOWugZc5Klm60Hg-1; Wed, 10 Apr 2024 11:28:44 -0400 X-MC-Unique: pPkIg4aBOWugZc5Klm60Hg-1 Received: by mail-qt1-f198.google.com with SMTP id d75a77b69052e-434ed2e412fso6707401cf.0 for ; Wed, 10 Apr 2024 08:28:44 -0700 (PDT) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1712762923; x=1713367723; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=Ci1baLlwDbOm6Wup8/nHleUUTbOV/R7CG6mxApYjz0M=; b=RZqN7fntjH9MiGoIAi010BWMvsRZBjTLvrkadIM+Rsg6YZK/CBPidvW+p7cskIz32m GGwqYG/dW7sPzSZ6eCxIOsRgIWrEGzYkwosEfzKS6yfdb/ahifNxVRqORfyIOEFkR3Fk l5LrVAN3scsiBVMpVCDbuPJ0p8jQGO+rTfVtsF6ujic0mejxstbl9pGvPJvSKqjA+dGD LPPzthWMhHH8QwoTE9WY0k22CgGUWPb9D9xHKaQ9X1RhTsfsZRKiCqMpbvNGwD666bHs xmzUPCxgyqmY24vtUVbviWTjd/Vr9e5OXKG8jNUOMnrCVTDZXLVtM1Mmfzjh8GpRdSnT 2SJA== X-Gm-Message-State: AOJu0YzJnxG0/K3zZNDECV/iO5RIM6F+hx4Cb1pMCGVbz7X+jUMOdEUY lZ+I83L4iXC17eG+2MA8pZmvrkZ8RIrbcGRx1SEDNFO2GnuYHMUv0u4us6FY61JpRJweLCTS3gH H2/alfmt4P15iKfwFPQ4Fy3uOE5OAtM2RffY1LIjrX0pXvt6w X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168075qvb.2.1712762923235; Wed, 10 Apr 2024 08:28:43 -0700 (PDT) X-Google-Smtp-Source: AGHT+IGZs5MCkL0IJwntJvXwMvElTS5kl/yk1ExPO6tDqr2NK62DfB6pdbW0p2/DjTaMvJz9lDcKPw== X-Received: by 2002:a05:6214:5090:b0:69b:ce6:271b with SMTP id kk16-20020a056214509000b0069b0ce6271bmr3168047qvb.2.1712762922544; Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Received: from x1n (pool-99-254-121-117.cpe.net.cable.rogers.com. [99.254.121.117]) by smtp.gmail.com with ESMTPSA id ek1-20020ad45981000000b00699032e555bsm5208543qvb.127.2024.04.10.08.28.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 10 Apr 2024 08:28:42 -0700 (PDT) Date: Wed, 10 Apr 2024 11:28:39 -0400 From: Peter Xu To: Jason Gunthorpe Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, Michael Ellerman , Christophe Leroy , Matthew Wilcox , Rik van Riel , Lorenzo Stoakes , Axel Rasmussen , Yang Shi , John Hubbard , linux-arm-kernel@lists.infradead.org, "Kirill A . Shutemov" , Andrew Jones , Vlastimil Babka , Mike Rapoport , Andrew Morton , Muchun Song , Christoph Hellwig , linux-riscv@lists.infradead.org, James Houghton , David Hildenbrand , Andrea Arcangeli , "Aneesh Kumar K . V" , Mike Kravetz Subject: Re: [PATCH v3 00/12] mm/gup: Unify hugetlb, part 2 Message-ID: References: <20240321220802.679544-1-peterx@redhat.com> <20240322161000.GJ159172@nvidia.com> <20240326140252.GH6245@nvidia.com> <20240405181633.GH5383@nvidia.com> <20240409234355.GJ5383@nvidia.com> MIME-Version: 1.0 In-Reply-To: <20240409234355.GJ5383@nvidia.com> X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com Content-Type: text/plain; charset=utf-8 Content-Disposition: inline X-Rspam-User: X-Rspamd-Server: rspam06 X-Rspamd-Queue-Id: 1ABFF16000D X-Stat-Signature: 7nddpgkys8g5nqm5yfdxt6c6qswzamt4 X-HE-Tag: 1712762925-645730 X-HE-Meta: U2FsdGVkX1+FGIanG9EETi5Wm8lImSex2kPXqsUY25mXMGUlO9toEUlYyjite1k20n1Npjz8MCDPxY2bw5MyE0IGD9ln3F9Vq//o6/7oGED2TCDu3J1p2mok6Sf7+yKUFyPULHIQNKi3RGLiPFzuCMOFeHCu+8kdd6Bt22xFkt7JTfcUGBOmdSPWR8d5r3SDSOXRx/F+3FfA5gCy8m4DZ3UmqsqmC6HXa/+OkU4U9kG0EC3Km4CpLmCCr1q7JfxxdH0MVqQJPtleD+tjyTsxl5JaMVu4VVXbyxDqDA6HNBs+HyT470VeLUHq3QQVCCvdHgfEWdC28MBBr3PXmj/pc37a6kXvuR/OzLoULYwJ+rTKt4dw6GdMBcAyi88UjESQk+XTNDCY9kRJrJFXfFTLyDLxtBPesueZK9z/xakS7gRRnZLy+qRG7BODrI0u+S7FFKFGvVCYkV07KU6iD9OuC06ppiMrBstBedgtz4j+s4I9z9PVWpKxdCI0CZ7g1RS4aDbyxp6+CtGcAyO7o7rUFcyRNYSWXQCNqJdrrbyv0BYGRjRQVmMfyhBmpqi2zBizcog8J096ToADrp6mSDkU1d6csetjb3fWbz+FgtWzkxYGwYlP7ZHHOqM+BbNywtaQ+IyesN5RN+9F9LatG8KCfsrBa5BzN/XvCfbxrtg/kMCxpLcz74FMtACKKiGJIelsAIv0o5IEzaLApFcTI5f60ZcPXCBDkeg/csFRkomL41oQ5Gsv4fbVXpUv2uleooborx4GI1+yDjdnSeTs+r/OsD2WFi5QTjg+JIEB1wtWjZ6p8NjxPFVfC4ulslt8K6VfCLUyaNU1mzotJQSzwOrcyBdRc6F5hT5zhWqNd4jqbb6bO2aR7gBfDu77b3coR4IikaCDYcCKtjYF/9P6WZZMCr3CJVWgbQbu2c7PT1RLe49S5PBIC4KKvTfEecuh15b/OQNt7LXpIi5coZ+SQkp IVfZmGxp I01Jz4FYcdEj6XyN77lR+8qmcRGng8XAMR0AOqQP4/gjb67a5Sb2nigjePZWpYDzenunh7M38qj77pwz6iFkQAVrnmR674fCzJxNM6xbVM5CGiLZMV2/3ZRKPkIWDfLnbfBxWb7Nf4EE9dXpVgYse7iJImP1G0soYJ5F809QTp9cMxn7b5j4GCTJWgGJ2c+kpXrEF9u8zIO4eOy9cCDs6DX3kapGi78MPViTvZoMadTC6CnA70DklfViWpvYkULTq+gem5yROzRtroZVkTfdlAgRuJSqdVj7jOjCiak2O32GLCaJvlNPGy6Gy0u2LsN355HnMJ7QKbTbKk+0oBamEFgbJ0tbRU/y/DMj4YSzJqdYgsOeJSJm81KY32mKLl7YdgkfhK7Z/vD20CdGJ7IiMOgYHnMSVWL93e/FAVW8HVKLj0kwii9kCb0MTU52d/IdOi9oShtiupuH/JHmtgRNk6bOXoXoGk2QhoSG4vHPPHrIDsTATV3kFPkI9V76haAz7Lkvb X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: On Tue, Apr 09, 2024 at 08:43:55PM -0300, Jason Gunthorpe wrote: > On Fri, Apr 05, 2024 at 05:42:44PM -0400, Peter Xu wrote: > > In short, hugetlb mappings shouldn't be special comparing to other huge pXd > > and large folio (cont-pXd) mappings for most of the walkers in my mind, if > > not all. I need to look at all the walkers and there can be some tricky > > ones, but I believe that applies in general. It's actually similar to what > > I did with slow gup here. > > I think that is the big question, I also haven't done the research to > know the answer. > > At this point focusing on moving what is reasonable to the pXX_* API > makes sense to me. Then reviewing what remains and making some > decision. > > > Like this series, for cont-pXd we'll need multiple walks comparing to > > before (when with hugetlb_entry()), but for that part I'll provide some > > performance tests too, and we also have a fallback plan, which is to detect > > cont-pXd existance, which will also work for large folios. > > I think we can optimize this pretty easy. > > > > I think if you do the easy places for pXX conversion you will have a > > > good idea about what is needed for the hard places. > > > > Here IMHO we don't need to understand "what is the size of this hugetlb > > vma" > > Yeh, I never really understood why hugetlb was linked to the VMA.. The > page table is self describing, obviously. Attaching to vma still makes sense to me, where we should definitely avoid a mixture of hugetlb and !hugetlb pages in a single vma - hugetlb pages are allocated, managed, ... totally differently. And since hugetlb is designed as file-based (which also makes sense to me, at least for now), it's also natural that it's vma-attached. > > > or "which level of pgtable does this hugetlb vma pages locate", > > Ditto > > > because we may not need that, e.g., when we only want to collect some smaps > > statistics. "whether it's hugetlb" may matter, though. E.g. in the mm > > walker we see a huge pmd, it can be a thp, it can be a hugetlb (when > > hugetlb_entry removed), we may need extra check later to put things into > > the right bucket, but for the walker itself it doesn't necessarily need > > hugetlb_entry(). > > Right, places may still need to know it is part of a huge VMA because we > have special stuff linked to that. > > > > But then again we come back to power and its big list of page sizes > > > and variety :( Looks like some there have huge sizes at the pgd level > > > at least. > > > > Yeah this is something I want to be super clear, because I may miss > > something: we don't have real pgd pages, right? Powerpc doesn't even > > define p4d_leaf(), AFAICT. > > AFAICT it is because it hides it all in hugepd. IMHO one thing we can benefit from such hugepd rework is, if we can squash all the hugepds like what Christophe does, then we push it one more layer down, and we have a good chance all things should just work. So again my Power brain is close to zero, but now I'm referring to what Christophe shared in the other thread: https://github.com/linuxppc/wiki/wiki/Huge-pages Together with: https://lore.kernel.org/r/288f26f487648d21fd9590e40b390934eaa5d24a.1711377230.git.christophe.leroy@csgroup.eu Where it has: --- a/arch/powerpc/platforms/Kconfig.cputype +++ b/arch/powerpc/platforms/Kconfig.cputype @@ -98,6 +98,7 @@ config PPC_BOOK3S_64 select ARCH_ENABLE_HUGEPAGE_MIGRATION if HUGETLB_PAGE && MIGRATION select ARCH_ENABLE_SPLIT_PMD_PTLOCK select ARCH_ENABLE_THP_MIGRATION if TRANSPARENT_HUGEPAGE + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS select ARCH_SUPPORTS_NUMA_BALANCING select HAVE_MOVE_PMD @@ -290,6 +291,7 @@ config PPC_BOOK3S config PPC_E500 select FSL_EMB_PERFMON bool + select ARCH_HAS_HUGEPD if HUGETLB_PAGE select ARCH_SUPPORTS_HUGETLBFS if PHYS_64BIT || PPC64 select PPC_SMP_MUXED_IPI select PPC_DOORBELL So I think it means we have three PowerPC systems that supports hugepd right now (besides the 8xx which Christophe is trying to drop support there), besides 8xx we still have book3s_64 and E500. Let's check one by one: - book3s_64 - hash - 64K: p4d is not used, largest pgsize pgd 16G @pud level. It means after squashing it'll be a bunch of cont-pmd, all good. - 4K: p4d also not used, largest pgsize pgd 128G, after squashed it'll be cont-pud. all good. - radix - 64K: largest 1G @pud, then cont-pmd after squashed. all good. - 4K: largest 1G @pud, then cont-pmd, all good. - e500 & 8xx - both of them use 2-level pgtables (pgd + pte), after squashed hugepd @pgd level they become cont-pte. all good. I think the trick here is there'll be no pgd leaves after hugepd squashing to lower levels, then since PowerPC seems to never have p4d, then all things fall into pud or lower. We seem to be all good there? > > If the goal is to purge hugepd then some of the options might turn out > to convert hugepd into huge p4d/pgd, as I understand it. It would be > nice to have certainty on this at least. Right. I hope the pmd/pud plan I proposed above can already work too with such ambicious goal too. But review very welcomed from either you or Christophe. PS: I think I'll also have a closer look at Christophe's series this week or next. > > We have effectively three APIs to parse a single page table and > currently none of the APIs can return 100% of the data for power. Thanks, -- Peter Xu