From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from out-173.mta0.migadu.com (out-173.mta0.migadu.com [91.218.175.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8613D3F44E8 for ; Thu, 28 May 2026 13:30:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.173 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779975030; cv=none; b=hSF6HpBNTFi1djrf2+nMA9qZ9Bxu5UAf67OPRRAXj81COV3ZwUDO91euWtxI78CbCqGXqxW3q8oYGj9F7caR58jh9ee77XhOjIoOMMSCZEWP6mH/3200sLfZeVXtxuk7weyXajm8bSWqGmZDryjVoV7FYSVyis3c+m7Y0ZT2NYA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779975030; c=relaxed/simple; bh=60OxiMo42dilbQTJSpdna78BcBAGrO3oDLbTHFznUBM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BCdiWJ6MwTwX46K7KocKcABRj1aPVnD14WXZ5JzOGC5w9nSoMaDnt+K4kluQ21IqBpxP3EAk366doGh0gzyGoGXNwA5QZdLFhwQQaHzXgXBd8g1A7eyKJdyWPiXfTcpBvI8fZgcOf7Yq9ZCB7Nf21c497OG0sL1tiYtKWMkPAec= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=HIyd9o0V; arc=none smtp.client-ip=91.218.175.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="HIyd9o0V" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1779975026; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=xtP6vXc+wJ1AKv++CgvfrsjrOTISB3WM3uSjHYG3gWE=; b=HIyd9o0V6brTfZATywJxIFk831aiMRHc+XNfJ8OFdKNEZu/JOtvt1X6UUEetvJSxDR4MQ4 W5oNzudzyvuB7V8LEJlYz/BxhuUaw/VIQKTPj53ulQkUDg5Wju+1H954dpOWAOhjzMNKXk +BAkjgcMQrRSNLrH7PJmtp3E0gzimbU= From: Kaitao Cheng To: dennis@kernel.org, tj@kernel.org, cl@gentwo.org, akpm@linux-foundation.org Cc: mhocko@suse.com, vbabka@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, muchun.song@linux.dev, Kaitao Cheng Subject: [PATCH 2/2] mm/percpu: Avoid pcpu_alloc_mutex recursion from reclaim Date: Thu, 28 May 2026 21:29:17 +0800 Message-ID: <20260528132917.81123-3-kaitao.cheng@linux.dev> In-Reply-To: <20260528132917.81123-1-kaitao.cheng@linux.dev> References: <20260528132917.81123-1-kaitao.cheng@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Migadu-Flow: FLOW_OUT From: Kaitao Cheng pcpu_alloc_noprof() takes pcpu_alloc_mutex for sleepable allocations so that it can create chunks and populate backing pages. If reclaim is entered while that mutex is already held, and reclaim reaches a path which allocates percpu memory, the nested allocation can try to take pcpu_alloc_mutex again. That creates a reclaim recursion dependency: pcpu_alloc_noprof(GFP_KERNEL) mutex_lock(&pcpu_alloc_mutex) reclaim pcpu_alloc_noprof(GFP_NOIO/GFP_NOFS) mutex_lock(&pcpu_alloc_mutex) Avoid this by treating percpu allocations from reclaim context as atomic. Such allocations may still be served from already available and populated areas, but they must not enter the mutex-protected slow path or create new chunks. If no space is available, fail the allocation and let the normal balance work handle replenishment outside reclaim. Update the function comment to describe that reclaim context allocations are atomic regardless of whether the supplied GFP mask would otherwise allow blocking. This patch is a preventive fix. There may not currently be any path that calls pcpu_alloc_noprof(GFP_NOIO/GFP_NOFS) from direct reclaim context. Fixes: 9a5b183941b5 ("mm, percpu: do not consider sleepable allocations atomic") Signed-off-by: Kaitao Cheng --- mm/percpu.c | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/mm/percpu.c b/mm/percpu.c index 1bb38467390b..9c30e5897813 100644 --- a/mm/percpu.c +++ b/mm/percpu.c @@ -1803,9 +1803,9 @@ static void pcpu_memalloc_scope_restore(gfp_t gfp, unsigned int flags) * @gfp: allocation flags * * Allocate percpu area of @size bytes aligned at @align. If @gfp doesn't - * contain %GFP_KERNEL, the allocation is atomic. If @gfp has __GFP_NOWARN - * then no warning will be triggered on invalid or failed allocation - * requests. + * allow blocking, or if allocation is requested from reclaim context, the + * allocation is atomic. If @gfp has __GFP_NOWARN then no warning will be + * triggered on invalid or failed allocation requests. * * RETURNS: * Percpu pointer to the allocated area on success, NULL on failure. @@ -1828,7 +1828,12 @@ void __percpu *pcpu_alloc_noprof(size_t size, size_t align, bool reserved, gfp = current_gfp_context(gfp); /* whitelisted flags that can be passed to the backing allocators */ pcpu_gfp = gfp & (GFP_KERNEL | __GFP_NORETRY | __GFP_NOWARN); - is_atomic = !gfpflags_allow_blocking(gfp); + /* + * Reclaim can be entered while pcpu_alloc_mutex is already held by + * another percpu allocation. Avoid recursing back into the mutex from + * reclaim; best-effort allocations from already populated areas are OK. + */ + is_atomic = !gfpflags_allow_blocking(gfp) || current->reclaim_state; do_warn = !(gfp & __GFP_NOWARN); /* -- 2.50.1 (Apple Git-155)