From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 04C2C389104; Wed, 30 Sep 2026 17:01:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790787666; cv=none; b=rVJJRjG1KSkhmJzJcyjt416PNVOCcNrneoB9xVMkqNS+qX2TmX1+s27SxQAd1Nw5u+/LgndDNkTuqrm/T6Gd61WkXvrIm7TPX8NjiVMT2VaYHu/sQrO9cuucb/3FX5wsy2Wd+JY4AWT7dtZLhCL0mMv/GymNQ/rVrQgecf+Luk4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790787666; c=relaxed/simple; bh=h9x/mxPTAiW2PZEbCc6WKqIt0vh/1yNoTxCtPvnbKkU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gMGgAiI+G+SxcTn1qC/BRBYLg430YJKDrzq4918n5hA3m00mqb1rcV11WJ4Egx7GHBwjek2K6x4joJDNTuTHLSqsAbfMMkcJ6eDPFYoKg3ctRTo4ZV9XfmbCz84Vi0PmJumkS0ENX8000m0MmECxC9aexom4hzhozec8YL4uXBA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b=zE87LZKR; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b="zE87LZKR" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DE3F41F000FF; Wed, 30 Sep 2026 17:01:03 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linuxfoundation.org; s=korg; t=1790787664; bh=pWw2znjjNKnTOVDYukRpTm6eixM07EFXCOHKEx1yzug=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=zE87LZKRlFopgV1OetATmRN1qthZvuXY1GjK29mn1q10rK+cCjK/WNTU0xDYM0pJV IMayUYd9EFsza3G2WcBGQejWNSonQ/oVprLM0yrAx/SBboTQWByC0B2X1+GxIZjVaY H3BoWfKynP6cP2f2UJPG9ErMJ2eflVRXtD2wikNE= From: Greg Kroah-Hartman To: stable@vger.kernel.org Cc: Greg Kroah-Hartman , patches@lists.linux.dev, Jinjiang Tu , Andrew Morton , Lance Yang , "Lorenzo Stoakes (ARM)" , "David Hildenbrand (Arm)" , "Vlastimil Babka (SUSE)" , Minchan Kim , Harry Yoo , Hiroyouki Kamezawa , Jann Horn , Kefeng Wang , Larry Woodman , "Liam R. Howlett" , Nanyong Sun , Rik van Riel Subject: [PATCH 7.2 311/457] mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish Date: Wed, 30 Sep 2026 17:26:56 +0200 Message-ID: <20260930152352.738293412@linuxfoundation.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260930152346.024115587@linuxfoundation.org> References: <20260930152346.024115587@linuxfoundation.org> User-Agent: quilt/0.69 X-stable: review X-Patchwork-Hint: ignore Precedence: bulk X-Mailing-List: patches@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit 7.2-stable review patch. If anyone has any objections, please let me know. ------------------ From: Jinjiang Tu commit b6ac0b3f6013c168f22cad97e79967accacb08e1 upstream. On arm64 server, we find that a task trying to grab the anon_vma lock triggers hungtask. INFO: task main:2354726 blocked for more than 120 seconds. Tainted: G E 5.10.0-0021.aarch64 #1 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. task:main state:D stack: 0 pid:2354726 ppid:2350673 flags:0x00000a01 Call trace: __switch_to+0x7c/0xbc __schedule+0x3b4/0x8a0 schedule+0x50/0xe0 rwsem_down_write_slowpath+0x3cc/0x6cc down_write+0x60/0x260 __anon_vma_prepare+0x6c/0x210 do_anonymous_page+0x258/0x660 handle_pte_fault+0x188/0x214 __handle_mm_fault+0x1b0/0x380 handle_mm_fault+0xf4/0x284 do_page_fault+0x19c/0x494 do_translation_fault+0xcc/0xf8 do_mem_abort+0x48/0xac el0_da+0x44/0x80 el0_sync_handler+0x88/0xb4 el0_sync+0x160/0x180 After analyzing the vmcore, we found the anon_vma->root->rwsem.count is -1. There is another anon_vma whose anon_vma->root->rwsem.count is 1, the anon_vma->root->rwsem.owner shows the lock is held, but the stack of the task shows the task doesn't hold the anon_vma lock. After adding more debugging info, we found __anon_vma_prepare() reuses anon_vma and triggers the UAF of anon_vma->root due to missing memory barrier, leading to locking and unlocking two different anon_vma->root, thus leading to an anon_vma will never be unlocked, and another anon_vma couldn't be locked anymore. This race requires two adjacent VMAs that are not merged but are anon_vma-compatible (e.g., they differ in VMA_ACCESS_FLAGS that can be changed by mprotect()). Two threads fault on each VMA concurrently, both calling __anon_vma_prepare() with only mmap_lock held for reading. THREAD A THREAD B __anon_vma_prepare __anon_vma_prepare find_mergeable_anon_vma() -> NULL anon_vma = anon_vma_alloc(); anon_vma->root = anon_vma; // the two stores may be reordered vma->anon_vma = anon_vma; // finds A's anon_vma anon_vma = find_mergeable_anon_vma(vma); anon_vma_lock_write(anon_vma); // may still see the old root down_write(&anon_vma->root->rwsem); anon_vma_unlock_write(anon_vma); // see the new root, never unlock old up_write(&anon_vma->root->rwsem); thread A triggers page fault and calls __anon_vma_prepare() to prepare anon_vma for the faulting vma. __anon_vma_prepare() allocates and initializes a new anon_vma, and then publishes it to the vma with a plain store. anon_vma_prepare() only requires the mmap_lock to be held for reading, so two threads can fault on adjacent VMAs at the same time. While thread A publishes a new anon_vma, thread B could find the anon_vma via find_mergeable_anon_vma() and then locks anon_vma->root->rwsem. The store to anon_vma->root in anon_vma_alloc() and the store to vma->anon_vma can be reordered. The anon_vma_lock_write() and spin_lock() only provide acquire semantics, which do not prevent prior stores from being reordered after them. The release semantics of the corresponding spin_unlock() and anon_vma_unlock_write() come too late, the store to vma->anon_vma is already published before they take effect. As a result, thread B can observe the following order: vma->anon_vma = anon_vma; anon_vma->root = anon_vma; The anon_vma slab is SLAB_TYPESAFE_BY_RCU, so a newly allocated anon_vma may reuse memory from a previously freed one. The constructor (anon_vma_ctor) does not reset anon_vma->root, and __put_anon_vma() doesn't clear it either, so the old root value persists until anon_vma_alloc() overwrites it. If that store isn't visible, thread B reads a root that points to the old anon_vma and locks it. As a result, thread B can call anon_vma_lock_write() with the old root, and call anon_vma_unlock_write() with the new root, leading to an anon_vma will never be unlocked, and another anon_vma couldn't be locked anymore (its count is dropped from 0 to -1 due to wrong unlock). To fix it, change the plain store `vma->anon_vma = anon_vma` to store release, so that the fields of anon_vma are visible before anon_vma is published to vma->anon_vma. At read side, the load of anon_vma and anon_vma->root have address dependency. According to Documentation/memory-barriers.txt and some investigations, only Alpha needs address-dependency barriers and it has been handled by READ_ONCE() in reusable_anon_vma(). We reproduced this issue in v5.10 with KSM enabled. The kernel doesn't merge commit cf7e7a3503df ("mm: prevent KSM from breaking VMA merging for new VMAs"), so there are many adjacent VMAs that aren't merged but are compatible for anon_vma. Without this fix, our production environment could reproduce this issue about 2-5 times each month. After adding a smp_mb() before anon_vma_lock_write(anon_vma) in __anon_vma_prepare(), which is different to this patch, this issue hasn't been reproduced for one month. Link: https://lore.kernel.org/20260908122924.554373-1-tujinjiang@huawei.com Fixes: 5c341ee1dfc8 ("mm: track the root (oldest) anon_vma") Signed-off-by: Jinjiang Tu Signed-off-by: Andrew Morton Reviewed-by: Lance Yang Reviewed-by: Lorenzo Stoakes (ARM) Acked-by: David Hildenbrand (Arm) Acked-by: Vlastimil Babka (SUSE) Cc: Minchan Kim Cc: Harry Yoo Cc: Hiroyouki Kamezawa Cc: Jann Horn Cc: Jinjiang Tu Cc: Kefeng Wang Cc: Larry Woodman Cc: Liam R. Howlett Cc: Nanyong Sun Cc: Rik van Riel Cc: Signed-off-by: Greg Kroah-Hartman --- mm/rmap.c | 6 +++++- mm/vma.c | 8 ++++++++ 2 files changed, 13 insertions(+), 1 deletion(-) --- a/mm/rmap.c +++ b/mm/rmap.c @@ -209,7 +209,11 @@ int __anon_vma_prepare(struct vm_area_st /* page_table_lock to protect against threads */ spin_lock(&mm->page_table_lock); if (likely(!vma->anon_vma)) { - vma->anon_vma = anon_vma; + /* + * Make anon_vma fields visible before anon_vma is published. + * Paired with an address dependency in reusable_anon_vma(). + */ + smp_store_release(&vma->anon_vma, anon_vma); anon_vma_chain_assign(vma, avc, anon_vma); anon_vma_interval_tree_insert(avc, &anon_vma->rb_root); anon_vma->num_active_vmas++; --- a/mm/vma.c +++ b/mm/vma.c @@ -1995,6 +1995,13 @@ static int anon_vma_compatible(struct vm * acceptable for merging, so we can do all of this optimistically. But * we do that READ_ONCE() to make sure that we never re-load the pointer. * + * The READ_ONCE() establishes an address dependency between anon_vma and + * any access to its fields, which pairs with the assignment to + * vma->anon_vma performed with release semantics in __anon_vma_prepare(). + * + * This is especially important as anon_vma's are SLAB_TYPESAFE_BY_RCU so + * accessing an uninitialised anon_vma's fields may result in a UAF. + * * IOW: that the "list_is_singular()" test on the anon_vma_chain only * matters for the 'stable anon_vma' case (ie the thing we want to avoid * is to return an anon_vma that is "complex" due to having gone through @@ -2009,6 +2016,7 @@ static struct anon_vma *reusable_anon_vm struct vm_area_struct *b) { if (anon_vma_compatible(a, b)) { + /* Paired with a memory barrier in __anon_vma_prepare(). */ struct anon_vma *anon_vma = READ_ONCE(old->anon_vma); if (anon_vma && list_is_singular(&old->anon_vma_chain))