From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 6787121147A for ; Fri, 10 Jan 2025 14:54:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736520889; cv=none; b=S9SsYn4J/ktI6XAA6Pum2deuK6+zvXKThwcXdtMHfWptSKbeRnQ+a/NF9UjbZeESlyU6MjXE8YtDZuoWwaNAMX9gObBIK9jfJuAmVuG9GkPmEBDn+vZRUzeGvw/72T3ip2Baz4udhLYxxA3KnRiCBjxjmWyIm/kzjqdUqts1HOs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736520889; c=relaxed/simple; bh=bDKg8luLOYMLT2wWpcKKaWCR2izb0RwPH8HgQuthG2c=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=LE6dbyqRBkm8lVFn02kPpx1nKK7hmJpzApMbBjB0ZD+W7VXs9EU5RnFCt0kcliZo6rgEMXbBX1ctUoOp9B1SOYjtjNIIZ6OSXsGxRxL26guq45Smi0u/sLBuk2WIhtRfYQdRWfdXCTsVHIXkwxYI0ykRAkOB5STRXjIFFQfwH8Q= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id EC9AF1477; Fri, 10 Jan 2025 06:55:14 -0800 (PST) Received: from [10.50.66.95] (PW040MKD.arm.com [10.50.66.95]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id A379D3F59E; Fri, 10 Jan 2025 06:54:30 -0800 (PST) Message-ID: <27ae4d80-38cd-4d6b-a49c-dad3f0ffbde3@arm.com> Date: Fri, 10 Jan 2025 20:24:25 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [RFC 08/11] khugepaged: introduce khugepaged_scan_bitmap for mTHP support To: Nico Pache , linux-kernel@vger.kernel.org, linux-mm@kvack.org Cc: ryan.roberts@arm.com, anshuman.khandual@arm.com, catalin.marinas@arm.com, cl@gentwo.org, vbabka@suse.cz, mhocko@suse.com, apopple@nvidia.com, dave.hansen@linux.intel.com, will@kernel.org, baohua@kernel.org, jack@suse.cz, srivatsa@csail.mit.edu, haowenchao22@gmail.com, hughd@google.com, aneesh.kumar@kernel.org, yang@os.amperecomputing.com, peterx@redhat.com, ioworker0@gmail.com, wangkefeng.wang@huawei.com, ziy@nvidia.com, jglisse@google.com, surenb@google.com, vishal.moola@gmail.com, zokeefe@google.com, zhengqi.arch@bytedance.com, jhubbard@nvidia.com, 21cnbao@gmail.com, willy@infradead.org, kirill.shutemov@linux.intel.com, david@redhat.com, aarcange@redhat.com, raquini@redhat.com, sunnanyong@huawei.com, usamaarif642@gmail.com, audra@redhat.com, akpm@linux-foundation.org References: <20250108233128.14484-1-npache@redhat.com> <20250108233128.14484-9-npache@redhat.com> Content-Language: en-US From: Dev Jain In-Reply-To: <20250108233128.14484-9-npache@redhat.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 09/01/25 5:01 am, Nico Pache wrote: > khugepaged scans PMD ranges for potential collapse to a hugepage. To add > mTHP support we use this scan to instead record chunks of fully utilized > sections of the PMD. > > create a bitmap to represent a PMD in order MTHP_MIN_ORDER chunks. > by default we will set this to order 3. The reasoning is that for 4K 512 > PMD size this results in a 64 bit bitmap which has some optimizations. > For other arches like ARM64 64K, we can set a larger order if needed. > > khugepaged_scan_bitmap uses a stack struct to recursively scan a bitmap > that represents chunks of fully utilized regions. We can then determine > what mTHP size fits best and in the following patch, we set this bitmap > while scanning the PMD. > > max_ptes_none is used as a scale to determine how "full" an order must > be before being considered for collapse. > > Signed-off-by: Nico Pache > --- > include/linux/khugepaged.h | 4 +- > mm/khugepaged.c | 129 +++++++++++++++++++++++++++++++++++-- > 2 files changed, 126 insertions(+), 7 deletions(-) > [--snip--] > > +// Recursive function to consume the bitmap > +static int khugepaged_scan_bitmap(struct mm_struct *mm, unsigned long address, > + int referenced, int unmapped, struct collapse_control *cc, > + bool *mmap_locked, unsigned long enabled_orders) > +{ > + u8 order, offset; > + int num_chunks; > + int bits_set, max_percent, threshold_bits; > + int next_order, mid_offset; > + int top = -1; > + int collapsed = 0; > + int ret; > + struct scan_bit_state state; > + > + cc->mthp_bitmap_stack[++top] = (struct scan_bit_state) > + { HPAGE_PMD_ORDER - MIN_MTHP_ORDER, 0 }; > + > + while (top >= 0) { > + state = cc->mthp_bitmap_stack[top--]; > + order = state.order; > + offset = state.offset; > + num_chunks = 1 << order; > + // Skip mTHP orders that are not enabled > + if (!(enabled_orders >> (order + MIN_MTHP_ORDER)) & 1) > + goto next; > + > + // copy the relavant section to a new bitmap > + bitmap_shift_right(cc->mthp_bitmap_temp, cc->mthp_bitmap, offset, > + MTHP_BITMAP_SIZE); > + > + bits_set = bitmap_weight(cc->mthp_bitmap_temp, num_chunks); > + > + // Check if the region is "almost full" based on the threshold > + max_percent = ((HPAGE_PMD_NR - khugepaged_max_ptes_none - 1) * 100) > + / (HPAGE_PMD_NR - 1); > + threshold_bits = (max_percent * num_chunks) / 100; > + > + if (bits_set >= threshold_bits) { > + ret = collapse_huge_page(mm, address, referenced, unmapped, cc, > + mmap_locked, order + MIN_MTHP_ORDER, offset * MIN_MTHP_NR); > + if (ret == SCAN_SUCCEED) > + collapsed += (1 << (order + MIN_MTHP_ORDER)); > + continue; > + } We are going to the lower order when it is not in the allowed mask of orders, or when we are below the threshold. What to do when these conditions do not happen, and the reason for collapse failure is collapse_huge_page()? For example, if you start with a PMD order scan, and collapse_huge_page() fails, then you hit "continue", and then exit the loop because there is nothing else in the stack, so we exit without trying mTHPs. > + > +next: > + if (order > 0) { > + next_order = order - 1; > + mid_offset = offset + (num_chunks / 2); > + cc->mthp_bitmap_stack[++top] = (struct scan_bit_state) > + { next_order, mid_offset }; > + cc->mthp_bitmap_stack[++top] = (struct scan_bit_state) > + { next_order, offset }; > + } > + } > + return collapsed; > +} > +