From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A75442737E8 for ; Tue, 19 Aug 2025 14:42:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1755614568; cv=none; b=W4MKmPcrlBQ/05H01Mz/fmVfdd42c2mC2CbwOV0a/Y/MQn+TVXrlfQwlsZ6w4GcV8yYfqFvR38JM0+SrmUvLbjaDGVCG8KMt29xeLDVA/UpN3RS1/tNB+CT+MkppN4rO0ypKa/0BkU13O6NxIQXQkxFLoYSY5eBUL/xnZF0v2tg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1755614568; c=relaxed/simple; bh=AOzXG7x8oyUCXfGZusv/OBJNoH8IM4j4ujSL9LGLkSs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=thIsRv1/n1ZJET2P+c3ickavN1TI8MUs0uJnYWgZaiD14UGZe+P0gAhbm4TxaXgbnz/tkhDBpstd0BEMRlZTrb795oHbJP7HnUCp5Au8qoVJgxV1Y1+P/qxKfMon/73a/hYCFZT+7ZGB2a5y6VTdU8q/9OSc9q2d+yqcSHVNODU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SxUsFBe1; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SxUsFBe1" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A09E9C4CEF1; Tue, 19 Aug 2025 14:42:45 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1755614567; bh=AOzXG7x8oyUCXfGZusv/OBJNoH8IM4j4ujSL9LGLkSs=; h=From:To:Cc:Subject:Date:In-Reply-To:References:From; b=SxUsFBe15wMaYRqrWwKx1jdRWZDil8tnv1mnpoR8xhIPoELFGXfVIaO53MXNYQVQq zq212d10K56XJH3A20d0/9c8LPUo4YrMcWYBl2qYDjRRrJdpDyCKq8CoMQ2wpO3GA5 D3/zleY2XMmamzj/DhaoFVI0k2Vs/eUN9wfMUVwbb36o+jMDj+1p6CzY/Xoq4rlksB 733G6gt8r28Unb7tRyDwgEs6/i03i8JCFR0mNFXxegERcpeuI3fWm6WGuVnVHcw00y ZCiyCsEdK/0EOv21MEuoCljRcezUA6zGGofwc8ytUbmvGZxTlCb1N8L3iiheHzj0zE vbUWX9rMjP1zQ== From: Sasha Levin To: stable@vger.kernel.org Cc: Anshuman Khandual , David Hildenbrand , Dev Jain , Alexander Gordeev , Catalin Marinas , Will Deacon , Ryan Roberts , Paul Walmsley , Palmer Dabbelt , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Christian Borntraeger , Sven Schnelle , Andrew Morton , Sasha Levin Subject: [PATCH 6.6.y] mm/ptdump: take the memory hotplug lock inside ptdump_walk_pgd() Date: Tue, 19 Aug 2025 10:42:43 -0400 Message-ID: <20250819144243.519849-1-sashal@kernel.org> X-Mailer: git-send-email 2.50.1 In-Reply-To: <2025081858-sabbath-blunt-7735@gregkh> References: <2025081858-sabbath-blunt-7735@gregkh> Precedence: bulk X-Mailing-List: stable@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: Anshuman Khandual [ Upstream commit 59305202c67fea50378dcad0cc199dbc13a0e99a ] Memory hot remove unmaps and tears down various kernel page table regions as required. The ptdump code can race with concurrent modifications of the kernel page tables. When leaf entries are modified concurrently, the dump code may log stale or inconsistent information for a VA range, but this is otherwise not harmful. But when intermediate levels of kernel page table are freed, the dump code will continue to use memory that has been freed and potentially reallocated for another purpose. In such cases, the ptdump code may dereference bogus addresses, leading to a number of potential problems. To avoid the above mentioned race condition, platforms such as arm64, riscv and s390 take memory hotplug lock, while dumping kernel page table via the sysfs interface /sys/kernel/debug/kernel_page_tables. Similar race condition exists while checking for pages that might have been marked W+X via /sys/kernel/debug/kernel_page_tables/check_wx_pages which in turn calls ptdump_check_wx(). Instead of solving this race condition again, let's just move the memory hotplug lock inside generic ptdump_check_wx() which will benefit both the scenarios. Drop get_online_mems() and put_online_mems() combination from all existing platform ptdump code paths. Link: https://lkml.kernel.org/r/20250620052427.2092093-1-anshuman.khandual@arm.com Fixes: bbd6ec605c0f ("arm64/mm: Enable memory hot remove") Signed-off-by: Anshuman Khandual Acked-by: David Hildenbrand Reviewed-by: Dev Jain Acked-by: Alexander Gordeev [s390] Cc: Catalin Marinas Cc: Will Deacon Cc: Ryan Roberts Cc: Paul Walmsley Cc: Palmer Dabbelt Cc: Alexander Gordeev Cc: Gerald Schaefer Cc: Heiko Carstens Cc: Vasily Gorbik Cc: Christian Borntraeger Cc: Sven Schnelle Cc: Signed-off-by: Andrew Morton Signed-off-by: Sasha Levin --- arch/arm64/mm/ptdump_debugfs.c | 3 --- arch/s390/mm/dump_pagetables.c | 2 -- mm/ptdump.c | 2 ++ 3 files changed, 2 insertions(+), 5 deletions(-) diff --git a/arch/arm64/mm/ptdump_debugfs.c b/arch/arm64/mm/ptdump_debugfs.c index 68bf1a125502..1e308328c079 100644 --- a/arch/arm64/mm/ptdump_debugfs.c +++ b/arch/arm64/mm/ptdump_debugfs.c @@ -1,6 +1,5 @@ // SPDX-License-Identifier: GPL-2.0 #include -#include #include #include @@ -9,9 +8,7 @@ static int ptdump_show(struct seq_file *m, void *v) { struct ptdump_info *info = m->private; - get_online_mems(); ptdump_walk(m, info); - put_online_mems(); return 0; } DEFINE_SHOW_ATTRIBUTE(ptdump); diff --git a/arch/s390/mm/dump_pagetables.c b/arch/s390/mm/dump_pagetables.c index b51666967aa1..4721ada81a02 100644 --- a/arch/s390/mm/dump_pagetables.c +++ b/arch/s390/mm/dump_pagetables.c @@ -249,11 +249,9 @@ static int ptdump_show(struct seq_file *m, void *v) .marker = address_markers, }; - get_online_mems(); mutex_lock(&cpa_mutex); ptdump_walk_pgd(&st.ptdump, &init_mm, NULL); mutex_unlock(&cpa_mutex); - put_online_mems(); return 0; } DEFINE_SHOW_ATTRIBUTE(ptdump); diff --git a/mm/ptdump.c b/mm/ptdump.c index 03c1bdae4a43..e46df2c24d85 100644 --- a/mm/ptdump.c +++ b/mm/ptdump.c @@ -152,6 +152,7 @@ void ptdump_walk_pgd(struct ptdump_state *st, struct mm_struct *mm, pgd_t *pgd) { const struct ptdump_range *range = st->range; + get_online_mems(); mmap_write_lock(mm); while (range->start != range->end) { walk_page_range_novma(mm, range->start, range->end, @@ -159,6 +160,7 @@ void ptdump_walk_pgd(struct ptdump_state *st, struct mm_struct *mm, pgd_t *pgd) range++; } mmap_write_unlock(mm); + put_online_mems(); /* Flush out the last page */ st->note_page(st, 0, -1, 0); -- 2.50.1