From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 29D034477F8 for ; Tue, 4 Aug 2026 10:41:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785840072; cv=none; b=bWtX9PWtuLkTRszINYo6g/pztRTRgIRri6izKhReWt6zONIQNxTw2I4igzCYOpfCaddgPPh0NyPyFiGImY3nJO7b+C/bA2LMFSW2Qg9u0DzrRx6m//Qqsobu3tBPfaGu3PcM6hmp1wBJNsTm47Qrva/qVcHCIvhzxMK8yG4C7AU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785840072; c=relaxed/simple; bh=jRVCFGH8MFmVsO5IRWzmet0dlZgIjoFtKEQpFGXSpys=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=P87F67tIgajjSPQilgHID6hDGMr7C3KndPYoNRgKDubd2E3qg3Tc5yoh3rdMkBo4QMe9F4rwqHjEXFZ1VvMxTFuutp1Fk/+maauW1a4xt1+vhtPgv5+Vit96xAjN4GIBXW4lFf1exyvQy7yAWs3fn+F7SZsrhag/SUyGv/T5woQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=mQl/X1b+; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="mQl/X1b+" Received: from pps.filterd (m0360083.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6748HXrS295917; Tue, 4 Aug 2026 10:40:44 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=XZTsOV r5fRA7ElFV8/LqC7tsT9j9fLay7cNKi0RN0JQ=; b=mQl/X1b+vaEV2ytVtlovX9 gVqDtG1bt/jfQKurshotPhNrTFHBZUf+XD+bRqubLtJVaarFlI1QAvxbdPG4exN1 kc1YgwBAFmwsMYal12Fn7DrTzoDDd1ftlDHDKCGGPRo3CFf7yM9gBzisMUdLOfAC 9umvIRAwsa1bpFLAWnByZ/8uOjyqv6AOy1Xcq3yNhXYlgAavv99GXLPptAfQ27qb 3zAEy5csCowsukEqofoxInKmhoORs8h+eGuroZLRMAOfMTXG9b643ufVIR1XwdEO 2zsIUVUHbieNzm+lROVIJk8+GubCVWPhfEdbpmDxE7+US7ib37yhJtGL2DQdU6GQ == Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8a3wcum-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Tue, 04 Aug 2026 10:40:43 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 674AQI0P021144; Tue, 4 Aug 2026 10:40:42 GMT Received: from smtprelay06.fra02v.mail.ibm.com ([9.218.2.230]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fsu4qhpve-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Tue, 04 Aug 2026 10:40:42 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (smtpav02.fra02v.mail.ibm.com [10.20.54.101]) by smtprelay06.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 674Aee2r32702726 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Tue, 4 Aug 2026 10:40:41 GMT Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id CF0522004B; Tue, 4 Aug 2026 10:40:40 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 9D9F220043; Tue, 4 Aug 2026 10:40:36 +0000 (GMT) Received: from [9.39.27.209] (unknown [9.39.27.209]) by smtpav02.fra02v.mail.ibm.com (Postfix) with ESMTP; Tue, 4 Aug 2026 10:40:36 +0000 (GMT) Message-ID: <4d531a51-4ec9-4df0-98e4-ec1f314891ba@linux.ibm.com> Date: Tue, 4 Aug 2026 16:10:35 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v4] sched/fair: Preserve wake-affine CPU for non-SMT reciprocal sync wakeups To: K Prateek Nayak , "Shubhang Kaushik (Ampere)" , Peter Zijlstra , Vincent Guittot , Ingo Molnar , Mel Gorman Cc: "Christoph Lameter (Ampere)" , Shubhang Kaushik , linux-kernel@vger.kernel.org, Juri Lelli , Dietmar Eggemann , Steven Rostedt , Ben Segall , Valentin Schneider , Christian Loehle , Madadi Vineeth Reddy References: <20260803-b4-sched-sync-wakeup-v4-1-52333b0cfb79@gentwo.org> <98bbe2d7-2401-4713-b49e-12f5dfc36471@linux.ibm.com> <97bc48b8-f00d-4c75-97ac-6feadf7d3372@amd.com> Content-Language: en-US From: Shrikanth Hegde In-Reply-To: <97bc48b8-f00d-4c75-97ac-6feadf7d3372@amd.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Authority-Analysis: v=2.4 cv=E6P9Y6dl c=1 sm=1 tr=0 ts=6a71c1ac cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=iQ6ETzBq9ecOQQE5vZCe:22 a=VwQbUJbxAAAA:8 a=yPCof4ZbAAAA:8 a=Enq7bIfXd3ymGqJVygkA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 X-Proofpoint-ORIG-GUID: bKDvYLee8Jvtj7mQP7ir_k6BObVmbmuU X-Proofpoint-GUID: B85vEtcJNZnHRF_YVz-sOLkOlYT2nsai X-Proofpoint-Spam-Info: AW1haW4tMjYwODA0MDA4MSBTYWx0ZWRfX/vYiAhvejOqe QZNGQmufNR0MCfQqezNflPS+h460VRHoxCSC8EP6LR/BkxuOH6myr0+w6uLfe1KXvcLWuI6Dc08 OBe8oYbm6vCs4JCJFg4ApCnky/ZIpCg= X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA0MDA4MSBTYWx0ZWRfXw/VzthQxWRFQ d95b/j+a5eQ39Jw8g3PLiPc0NN2QbcX4rH9betXf8NOA0EvRqEV3g3jU/Y6Nllk/iOtMjxR9dNk X8oVslZL2HFOYfiTK/BM4WPyKA0j0eBdr6ehafj0uTsSAeac1hNdQxgnueK7zzxI+NW6hAWop9z iu8AMs/Bsj/9Qd2FSDkDHDEokGsG/Uy8+ZfnFXTio/87GSEhXMoUrNkcrRuiLrchI5SHXCjlPAc USrbMCQQ/zH2zdcPznZgpASjbCi/8ovjasXTsTcFmew//sbDjb/4Sgxchsjz85PwsGKvDC+edPy ORtIO9xTF0G7SE9Yuq0uuPQAFNN2wCNI13f9lr11pXGU/hnZDyass4q3bTCMd/S5n5TBUMLCYMk z0MljrZS6Q5pN4M/8aYC7FFEYMP8eqAB5hJVqDgfAJNHU6aW5rMy9cwEkxEX72rq7tumLJ5s8tz ssesH8fMVhoSHEoRB5A== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-04_02,2026-08-03_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 bulkscore=0 clxscore=1015 lowpriorityscore=0 priorityscore=1501 suspectscore=0 adultscore=0 spamscore=0 malwarescore=0 impostorscore=0 phishscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608040081 Hi Prateek. On 8/4/26 2:12 PM, K Prateek Nayak wrote: > Hello Shrikanth, > > On 8/4/2026 10:10 AM, Shrikanth Hegde wrote: >> As I said in v3, before we add bells/whistles to sync path, i want >> to know what is expected of sync behavior today. >> And that should be documented in Documentation/scheduler/ >> >> Be it, >> - current way of hint only and scheduler can still choose an idle core/idle cpu etc. >> - Should it be enforcing it to waker cpu if waker cpu has only one task. >> - Whatever the policy maybe. >> >> Current api usage is tricky to use and effect is visible in real life workloads. >> The case I mentioned in v3 of networking code using sync api leads to strange >> results due to sync mechanism. >> - It depends whether waker/wakee are running on same node. >> - Result of wake_wide. >> In other end, user sees inconsistent latency/throughput. >> >> We can keep on adding minor changes to sync api path, >> but one benchmark will benefit and one will suffer. >> Having the behavior documented is a good start. >> >> Peter, Ingo, Vincent, Mel, Prateek, >> What do you guys think? > > Currently it is very arbitrary and WF_SYNC may, or may not, indicate a > true voluntary blocking behavior. For example, anon_pipe_read() uses a > wake_up_interruptible_sync_poll() to wake up writers once reader has > drained the pipe but if you think about it, why would the reader block > soon after just having the data it needed? Doesn't "perf bench sched pipe" also use anon_pipe_read/write? - 27.49% 0.32% sched-pipe [kernel.kallsyms] [k] ksys_read - 27.17% ksys_read - 26.57% vfs_read - 22.33% anon_pipe_read > > Here are the results on my Zen4 system from running perf bench > sched messaging (threads + pipes) at varying worker counts with > all wake_up_interruptible_sync_poll converted to > wake_up_interruptible_poll: > > Test: tip no_sync > 1-groups: 3.79 (0.00 pct) 3.35 (11.60 pct) > 2-groups: 3.85 (0.00 pct) 3.41 (11.42 pct) > 4-groups: 4.02 (0.00 pct) 3.32 (17.41 pct) > 8-groups: 4.33 (0.00 pct) 4.38 (-1.15 pct) > 16-groups: 6.09 (0.00 pct) 6.12 (-0.49 pct) > --- > > So seems like WF_SYNC hint on this machine with perf bench sched > messaging (thread + pipes) pattern is actually holding it back. > Lemme check processes ... > > Test: tip no_sync > 1-groups: 3.48 (0.00 pct) 3.08 (11.49 pct) > 2-groups: 3.80 (0.00 pct) 3.07 (19.21 pct) > 4-groups: 3.91 (0.00 pct) 3.09 (20.97 pct) > 8-groups: 4.13 (0.00 pct) 4.10 (0.72 pct) > 16-groups: 5.81 (0.00 pct) 5.74 (1.20 pct) > > Similar stuff. At some point it was pretty bad for Zen3 but > situation might have changed since ¯\_(ツ)_/¯ I'll let you > know once I have a machine. > > But ... If I have true 1:1 waiting on pipe as in the case of > "perf bench sched pipe -l 1000000" I go from ~2.5usecs/op on > average to ~4.2usecs/op which is close to a 50% increase in the > benchmark time so that WF_SYNC hint can also help if all we have > is looping over a read waiting for one page worth of write. > > The way I look at WF_SYNC nowadays is that it indicates a local LLC > wakeup is beneficial. wake_wide() doesn't even consider WF_SYNc and > simply uses wake-wakee flips and then only at want_affine() do we > actually check the sync hint. > Yes, it is difficult to say when would sync actually kick in. Also, if we say hint, onus now falls on scheduler to optimize all call sites. Clearly comments around __wake_up_sync_key are outdated. > Most benefit come from wake_affine_idle() for the 1 task case where > target is set to current and select_idle_sibling() uses that as the > target from there on. > > It could purely be a coincidence that it benefits at all - most of > these microbenchmark we have always hit the same two syscall (mostly > read() and write() in turns) and a most of benefit for those comes > from kernel instructions being primed in cache. > > I remember a while back, removing the effect of WF_SYNC on the > networking side hampered a lot of performance - especially for > localhost communications. This one specifically > https://lore.kernel.org/lkml/20220711224704.1672831-1-libo.chen@oracle.com/ > > Let me see if things have miraculously changed there too but for > TCP sockets in real world, with blocking for ACKs, I think > WF_SYNC still makes sense there but I feel most of the benefits > from sync are a second order effect. > But today it calls sync even for non-blocking.