From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-00069f02.pphosted.com (mx0b-00069f02.pphosted.com [205.220.177.32]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AE0A23BB674 for ; Mon, 17 Aug 2026 08:42:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=205.220.177.32 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786956176; cv=none; b=F2/HY0axu5MdJEQuUUz2j1mlRHfKpFHfBsDg8VdHhoCKB+cnWwA8lbK6YtzlEgl8rxxbBw5VH406qWDXaa8kkXGd9WMN+bNo6uZPV47qhd1WW8JUVpDmRM/9gdwd82R5dj6Ry7bcY1wqkYSab3p/YQest/23P/LTl5gio3Q/Rmo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786956176; c=relaxed/simple; bh=a2KJIZsBCp7HG0WuJvlyZO0bvJfD+OjLpmEYzQLs178=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=uKs2NMWnpy5L5/f82hbsjIM6M0Bydq4q5Wf/OhhZVCatqAanyd5hGfuKE/Q6ka6SaN4cDXD9XYGBU5by36Z26hEdi6H59ZOdkDxYALuGytu/H88cU+X7+Fs3HnecH++Aug99pF1nR2GNuo/yatosdzw1dRKIyZIPFOyVOIE++1s= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oracle.com; spf=pass smtp.mailfrom=oracle.com; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b=Oq44cy7b; arc=none smtp.client-ip=205.220.177.32 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oracle.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=oracle.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="Oq44cy7b" Received: from pps.filterd (m0246631.ppops.net [127.0.0.1]) by mx0b-00069f02.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67GMr7NW3275841; Mon, 17 Aug 2026 08:42:24 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s= corp-2025-04-25; bh=jlXb8C140Mg9tU3S5wDgtAO0on14dmgA56I5Kpd/yuw=; b= Oq44cy7bsUJTI17IlUNeMJJDxkDn/IRvOLK3piI1Ad5gUG4kevPQ08Objb2n0G2p L8tBERKi/NTN2JYvjaVcBmz8XD4qhzgIBLcfV1x04IBcYrtHMiXdDutpE6/dOtWc 6CWmeCxbkFvC2KoFhH1VB6cv382tskCH1NYFcA0oskWiTBoa7sLqP9e88+D17zLu /95Db8kEcEVy7JRvJwZEM9Xc/XqK21AsgXQ3+OT8kU61dr/kMGot18XU/PW0ShFt t+gVGgvPzAi5Wh1Xd9jDKBTyUG4qQWe7pKcz0nF1Rr7yvzc+0RSsK4sTJh/81oSc MiZFAXtXMz5w60SWFb/blw== Received: from iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com (iadpaimrmta02.appoci.oracle.com [147.154.18.20]) by mx0b-00069f02.pphosted.com (PPS) with ESMTPS id 4g2ew5215f-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Mon, 17 Aug 2026 08:42:23 +0000 (GMT) Received: from pps.filterd (iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com [127.0.0.1]) by iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com (8.18.1.7/8.18.1.7) with ESMTP id 67H8ebc2035032; Mon, 17 Aug 2026 08:42:23 GMT Received: from pps.reinject (localhost [127.0.0.1]) by iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com (PPS) with ESMTPS id 4g2ejp4a66-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Mon, 17 Aug 2026 08:42:23 +0000 (GMT) Received: from iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com (iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com [127.0.0.1]) by pps.reinject (8.18.1.12/8.18.1.12) with ESMTP id 67H8gKbd040948; Mon, 17 Aug 2026 08:42:22 GMT Received: from lab61.no.oracle.com (lab61.no.oracle.com [10.172.144.82]) by iadpaimrmta02.imrmtpd1.prodappiadaev1.oraclevcn.com (PPS) with ESMTP id 4g2ejp4a4p-2; Mon, 17 Aug 2026 08:42:22 +0000 (GMT) From: =?UTF-8?q?H=C3=A5kon=20Bugge?= To: linux-kernel@vger.kernel.org, Peter Zijlstra , Ingo Molnar , Will Deacon , Boqun Feng , Waiman Long , Chris Wilson Cc: John Stultz , Bradley Morgan , Tejun Heo , =?UTF-8?q?H=C3=A5kon=20Bugge?= , Ingo Molnar Subject: [PATCH v4 2/2] test-ww_mutex: Fix deadlock in test_cycle_work Date: Mon, 17 Aug 2026 10:42:16 +0200 Message-ID: <20260817084218.318326-2-haakon.bugge@oracle.com> X-Mailer: git-send-email 2.43.5 In-Reply-To: <20260817084218.318326-1-haakon.bugge@oracle.com> References: <20260817084218.318326-1-haakon.bugge@oracle.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-16_06,2026-08-12_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 mlxlogscore=999 lowpriorityscore=0 malwarescore=0 bulkscore=0 suspectscore=0 adultscore=0 phishscore=0 mlxscore=0 spamscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.19.0-2606160000 definitions=main-2608170064 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODE3MDA2NCBTYWx0ZWRfXyQelpYEFzR0n F8nrGrLwBMWKEGapsZlr3mtINLIjBqhdT/+Oe89bnCRSku+X5UURlxAvUvghatDp9hSUKt/VyYY JOGHZP6kcrcmy4XpHJPbJgN2dEN/5yWA2/aIrY+hBMkn6n90A3NqQzY64j7R5LiT9eQgq360dno XCyTWpzvxdlX/7KEgjbQH54K40lecWODXV51P0jgJeWChIHGslupCA2+ureDuOTjpEc8mVg5MRJ ekWW20TQSjkg10/tpjWlznZuDfC4ldlMc0DoYH0/fGIVGXQ/5CdPRh6unz8VYa5vDyu44iXmUGt ix0jn5gbWUFDaGB75KAbi8G2iMos/km6c1JhYYXj4RHmQ7jIvC28rKnkO6WugEVU31AolDn50Mk X3GpkwYZzBuSLz7vKe2rLxeimw+t3Wb3gfAU+LDzqnIx7TkqZZxTFMimvxPsrQgC+jXZ3ChD7kQ SLamks1drAM/pVFDC85kv9OSHMiZyGKIPFmOp2OI= X-Proofpoint-GUID: nkYBmxiWcoR6Cl9RYnhPzG5Z34gxVxxJ X-Authority-Analysis: v=2.4 cv=cYbiaHDM c=1 sm=1 tr=0 ts=6a82c970 b=1 cx=c_pps a=e1sVV491RgrpLwSTMOnk8w==:117 a=e1sVV491RgrpLwSTMOnk8w==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=M51BFTxLslgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=jiCTI4zE5U7BLdzWsZGv:22 a=o5oIOnhZENCTenyL_yNV:22 a=yPCof4ZbAAAA:8 a=jcFIXsoQAAAA:8 a=WAdc_XZaImYrliwInUAA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 a=M1-Q_cRM6PFdmZSamPpU:22 a=5yU3S35YU4bGjq-dph-N:22 a=Bho9c0fBagfJEIQBS7DQ:22 cc=ntf awl=host:13509 X-Proofpoint-Spam-Info: AW1haW4tMjYwODE3MDA2NCBTYWx0ZWRfX8fniOfkP4j7k 1VvU9+eGLJQMj3J86jOvvWb8uxwqOoLQ5XQ8Dg2aDUVJyV0NMVD7cgyXwqW8BPCJsPiVkbeXpsG EHQg0hXqUZRb6xMQgbOTrXV67QhI5Jz3pgmVa2m9CjQTOdH2ta3L X-Proofpoint-ORIG-GUID: nkYBmxiWcoR6Cl9RYnhPzG5Z34gxVxxJ When running with N online CPUs, where N is fairly large, let's say 512, a deadlock may happen in test_cycle_work() when running with N + 1 kernel threads. The reason is that the min_active is too small in order to let N + 1 worker threads run concurrently. The following is my analyzes. The test sets up a circular dependency with ww_mutexes and completions: work[1] completes signal[0] work[2] completes signal[1] ... work[511] completes signal[510] work[512] would complete signal[511] work[0] completes signal[512] Therefore: work[0..510] -> blocked in ww_mutex_lock(b_mutex) work[511] -> blocked in wait_for_completion(signal[511]) work[512] -> inactive; test_cycle_work() has not started Because the last worker thread has not started, worker 511 hangs forever in wait_for_completion(). This bug produces the following splats (slightly edited for better brevity). This for the worker threads hung in ww_mutex_lock(): INFO: task kworker/u2066:1:13658 blocked for more than 123 seconds. Workqueue: test-ww_mutex test_cycle_work [test_ww_mutex] Call Trace: __schedule+0x28e/0x670 schedule+0x27/0xa0 schedule_preempt_disabled+0x15/0x30 __ww_mutex_lock.constprop.0+0x841/0xe00 test_cycle_work+0x9d/0x150 [test_ww_mutex] process_one_work+0x196/0x370 worker_thread+0x1af/0x320 kthread+0xe3/0x120 ret_from_fork+0x19e/0x260 ret_from_fork_asm+0x1a/0x30 And this for the single worker thread 511, hung in wait_for_completion(): Workqueue: test-ww_mutex test_cycle_work [test_ww_mutex] Call Trace: __schedule+0x28e/0x670 schedule+0x27/0xa0 schedule_timeout+0xac/0xf0 __wait_for_common+0x97/0x1b0 test_cycle_work+0x80/0x160 [test_ww_mutex] process_one_work+0x196/0x370 worker_thread+0x1af/0x320 kthread+0xe3/0x120 ret_from_fork+0x19e/0x260 ret_from_fork_asm+0x1a/0x30 We fix this by adjusting the wq's min_active parameter. When num_online_cpus() has been sampled in run_tests(), we adjust the {min,max}_active values of the wq, to make sure both are set to ncpus + 1 in order to avoid the above deadlock. Note that this fix is invariant to num_online_cpus() changing after it has been sampled, because the RC here is the number of runnable worker threads vs. threads created, not per se the number of online CPUs. Also, in order to avoid exceeding WQ_MAX_ACTIVE, we create cycle_ncpus and clamp it. We do not want to change ncpus for the other tests, not affected by this bug. Fixes: d1b42b800e5d ("locking/ww_mutex: Add kselftests for resolving ww_mutex cyclic deadlocks") Signed-off-by: HÃ¥kon Bugge Reviewed-by: Bradley Morgan --- v3 -> v4: * No changes v2 -> v3: * No changes v1 -> v2: * Added Bradley's r-b * Reworded comment in run_tests() --- kernel/locking/test-ww_mutex.c | 13 ++++++++++++- 1 file changed, 12 insertions(+), 1 deletion(-) diff --git a/kernel/locking/test-ww_mutex.c b/kernel/locking/test-ww_mutex.c index 47e016a4f4fea..e11f6b869b4d4 100644 --- a/kernel/locking/test-ww_mutex.c +++ b/kernel/locking/test-ww_mutex.c @@ -676,6 +676,7 @@ static int stress(struct ww_class *class, int nlocks, int nthreads, unsigned int static int run_tests(struct ww_class *class) { int ncpus = num_online_cpus(); + int cycle_ncpus = min_t(int, ncpus, WQ_MAX_ACTIVE - 1); int ret, i; ret = test_mutex(class); @@ -696,7 +697,17 @@ static int run_tests(struct ww_class *class) return ret; } - ret = test_cycle(class, ncpus); + /* + * test_cycle_work() has a linear dependency which requires + * all kernel threads to be run-able at once. With N CPUs and + * N + 1 worker threads, deadlock may happen. Hence, adjust + * min_active. Raise max first, min_active is clamped to it. + * Cap N so that N + 1 doesn't exceed WQ_MAX_ACTIVE. + */ + workqueue_set_max_active(wq, cycle_ncpus + 1); + workqueue_set_min_active(wq, cycle_ncpus + 1); + + ret = test_cycle(class, cycle_ncpus); if (ret) return ret; -- 2.43.5