From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 090CF14F70 for ; Thu, 25 Jun 2026 02:57:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782356227; cv=none; b=AR+cIsbMw/M/V1H3j2vgCMTFsnyeIb/ur97E3lvvIhG8UsP07xtY+vvp39n6u1pR7YIMKbFiARtzEdlKjkb9z5ABEOWGSqFLqnQ2lxVFaiYN416YT+A3fpP2Si/Oe7pJEYBULd3WE7tmh1Y+04EBhvW3tDR5cml7Q7ryyBU0gN8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782356227; c=relaxed/simple; bh=QLBddajJX+4op3KljiFnb0Y6pLOfwf7j5/+SuIU+Fx8=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=ckdF1TUkoRW6HZ/dxvsNwc/9LyeFTODeIr4YDQAeL8DV4IbLHMIsC9VDXfPrlBjz7FrgsZdiMVmETtMEGlC0mC8VTLcCdfjWZak3T7OuvNrVIuqy3VwHu7yet+bgfjbW5SbRMq1L24IM30tAHsuZfPKxN0G8XlYa7FMauVIuudU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=wUfWXj5a; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="wUfWXj5a" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0B1481F000E9; Thu, 25 Jun 2026 02:57:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1782356225; bh=1VUqMECiZlY9thhbQvmy35ab0ws4JsaGHxxluesab0Q=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=wUfWXj5aUF9eoVta8SgZEnnNduNGTOKD5CBkcZw8lAQZ5nDQRfynZOikmNZair0Ja xXlXcSLqbYwRVFC0GBj35O/iJAfNJHkfNGgjOUcxYCbjg/mRCa74VQbQf+iHrTDgCb jyvQry5GV5UdWA2ROhhIbcXZti87XuVpMIrNdw/8= Date: Wed, 24 Jun 2026 19:57:04 -0700 From: Andrew Morton To: Wen Jiang Cc: linux-mm@kvack.org, linux-arm-kernel@lists.infradead.org, catalin.marinas@arm.com, will@kernel.org, urezki@gmail.com, baohua@kernel.org, Xueyuan.chen21@gmail.com, dev.jain@arm.com, rppt@kernel.org, david@kernel.org, ryan.roberts@arm.com, anshuman.khandual@arm.com, ajd@linux.ibm.com, linux-kernel@vger.kernel.org, jiangwen6@xiaomi.com, shanghaoqiang@xiaomi.com Subject: Re: [PATCH v4 0/6] mm/vmalloc: Speed up ioremap, vmalloc and vmap with contiguous memory Message-Id: <20260624195704.5c29c0353163babb721585ca@linux-foundation.org> In-Reply-To: <20260618084726.1070022-1-jiangwen6@xiaomi.com> References: <20260618084726.1070022-1-jiangwen6@xiaomi.com> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Thu, 18 Jun 2026 16:47:20 +0800 Wen Jiang wrote: > This patchset accelerates ioremap, vmalloc, and vmap when the memory > is physically fully or partially contiguous. Two techniques are used: Thanks. > 1. Avoid page table rewalk when setting PTEs/PMDs for multiple memory > segments > 2. Use batched mappings wherever possible in both vmalloc and ARM64 > layers > > Besides accelerating the mapping path, this also enables large > mappings (PMD and cont-PTE) for vmap, which are currently not > supported. > > Patches 1-2 extend ARM64 vmalloc CONT-PTE mapping to support multiple > CONT-PTE regions instead of just one. > > Patch 3 extracts a common helper vmap_set_ptes() that consolidates PTE > mapping logic between the ioremap and vmalloc/vmap paths, handling both > CONT_PTE and regular PTE mappings. This prepares for the next patch. > > Patch 4 extends the page table walk path to support page shifts other > than PAGE_SHIFT and eliminates the page table rewalk for huge vmalloc > mappings. The function is renamed from vmap_small_pages_range_noflush() > to vmap_pages_range_noflush_walk(). > > Patches 5-6 add huge vmap support for contiguous pages, including > support for non-compound pages with pfn alignment verification. > > On the RK3588 8-core ARM64 SoC, with tasks pinned to a little core and > the performance CPUfreq policy enabled, benchmark results: > > * ioremap(1 MB): 1.35x faster (3407 ns -> 2526 ns) > * vmalloc(1 MB) mapping time (excluding allocation) with > VM_ALLOW_HUGE_VMAP: 1.42x faster (5.00 us -> 3.53us) > * vmap(100MB) with order-8 pages: 8.3x faster (1235 us -> 149 us) Nice. > Many thanks to Xueyuan Chen for his testing efforts on RK3588 boards. Indeed. I see Dev had a good look at v3 - hopefully he (and Ulad) (and more ARM folks) have time to go through this. Is there any effect on anything other than arm64? I'm wondering how much testing these changes will really get in mm.git and linux-next. How is our selftests coverage of these changes? Is there some existing selftest which will exercise these new features? You diligently went through the Sashiko report against v3 (thanks). Please pass an eye across its v4 report, see if something new popped up? https://sashiko.dev/#/patchset/20260618084726.1070022-1-jiangwen6@xiaomi.com