From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 631ED3242DF for ; Mon, 27 Apr 2026 15:05:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777302314; cv=none; b=bjhnLK8FxMHXLX7ltwXf0aYL6p9SR6B3Z2sNY8DJiqyRXaLXvghG5RBONCd0gQIsQGUR8atzWpBbg2eLKYnOokKF8x0LLHSkL0A/kTrDEsU5HcvfTarhWl7EcpusAhaj55bxD38AfrvKktyAOhWz2mtt69JavFQaJSig0Gr5H2Y= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1777302314; c=relaxed/simple; bh=LF4eMdK0ID4NJDAUmuKQ8VsmPWmO7w6FTTlpJhcDycs=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=qHdFtqSy+Btum4sGomGgIJ/9ZwZqZ1Um8xvmK86eH6vTUiBBPP8hQojmOwlar8cHxw9Xx+QbBlYO4fgMDY/JG7OcNBlD04YS9TVEKIclM1UccQ1EtHCRg9AEJvqLnZcx5eA67+Hs0YbukAa2qQs+dz6gs7pFaigiJc2KnM4Be48= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b=ZvPqTD/8; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b="ZvPqTD/8" Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 0699B1FC7; Mon, 27 Apr 2026 08:05:05 -0700 (PDT) Received: from [10.164.148.37] (MacBook-Pro.blr.arm.com [10.164.148.37]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPSA id 952753F7B4; Mon, 27 Apr 2026 08:05:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1777302310; bh=LF4eMdK0ID4NJDAUmuKQ8VsmPWmO7w6FTTlpJhcDycs=; h=Date:Subject:To:Cc:References:From:In-Reply-To:From; b=ZvPqTD/8uz+vouJEXQWe5V4WnPgvNPoD0zRcg/+h/FtE7s8IlA6PhWIUySQ6Ot+HJ 7KCalBDGW0j/7MxTtrtqV9TDUTDg/99gV52ADTb5nu8Ml/zvWGSKnJ6WAsgr+vS0FR BRD/orHCUaiFgQi9kEUiKiBo8YlSGbTrzSKMfScs= Message-ID: Date: Mon, 27 Apr 2026 20:34:57 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [RFC PATCH 0/8] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory To: "Barry Song (Xiaomi)" , linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org, catalin.marinas@arm.com, will@kernel.org, akpm@linux-foundation.org, urezki@gmail.com Cc: linux-kernel@vger.kernel.org, anshuman.khandual@arm.com, ryan.roberts@arm.com, ajd@linux.ibm.com, rppt@kernel.org, david@kernel.org, Xueyuan.chen21@gmail.com References: <20260408025115.27368-1-baohua@kernel.org> Content-Language: en-US From: Dev Jain In-Reply-To: <20260408025115.27368-1-baohua@kernel.org> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On 08/04/26 8:21 am, Barry Song (Xiaomi) wrote: > This patchset accelerates ioremap, vmalloc, and vmap when the memory > is physically fully or partially contiguous. Two techniques are used: > > 1. Avoid page table zigzag when setting PTEs/PMDs for multiple memory > segments > 2. Use batched mappings wherever possible in both vmalloc and ARM64 > layers > > Patches 1–2 extend ARM64 vmalloc CONT-PTE mapping to support multiple > CONT-PTE regions instead of just one. > > Patches 3–4 extend vmap_small_pages_range_noflush() to support page > shifts other than PAGE_SHIFT. This allows mapping multiple memory > segments for vmalloc() without zigzagging page tables. > > Patches 5–8 add huge vmap support for contiguous pages. This not only > improves performance but also enables PMD or CONT-PTE mapping for the > vmapped area, reducing TLB pressure. > > Many thanks to Xueyuan Chen for his substantial testing efforts > on RK3588 boards. > > On the RK3588 8-core ARM64 SoC, with tasks pinned to CPU2 and > the performance CPUfreq policy enabled, Xueyuan’s tests report: > > * ioremap(1 MB): 1.2× faster > * vmalloc(1 MB) mapping time (excluding allocation) with > VM_ALLOW_HUGE_VMAP: 1.5× faster > * vmap(): 5.6× faster when memory includes some order-8 pages, > with no regression observed for order-0 pages > > Barry Song (Xiaomi) (8): > arm64/hugetlb: Extend batching of multiple CONT_PTE in a single PTE > setup > arm64/vmalloc: Allow arch_vmap_pte_range_map_size to batch multiple > CONT_PTE > mm/vmalloc: Extend vmap_small_pages_range_noflush() to support larger > page_shift sizes > mm/vmalloc: Eliminate page table zigzag for huge vmalloc mappings > mm/vmalloc: map contiguous pages in batches for vmap() if possible > mm/vmalloc: align vm_area so vmap() can batch mappings > mm/vmalloc: Coalesce same page_shift mappings in vmap to avoid pgtable > zigzag > mm/vmalloc: Stop scanning for compound pages after encountering small > pages in vmap > > arch/arm64/include/asm/vmalloc.h | 6 +- > arch/arm64/mm/hugetlbpage.c | 10 ++ > mm/vmalloc.c | 178 +++++++++++++++++++++++++------ > 3 files changed, 161 insertions(+), 33 deletions(-) > Hi Barry, have you got the chance to work on v2?