From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 636F3221FB4 for ; Sat, 1 Aug 2026 03:56:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785556588; cv=none; b=s5OTmIMDEPt1fgs6uiXHFXlShMILqZ8ZmNt/tRsi8YuwrUVLRguSVxTVFL8+sQOMyorsPO1aLF19Q50/K0eBsZkHS1TuNOr0AOWaZcCEerKxggUyWWElUjyvFivDGOM3Qb+xEj3HI3LCcaS+AkgpkNWYCe/mZnqdZBorJmf7jhQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785556588; c=relaxed/simple; bh=J5XDgB2bDXEM1RnvKtxS1WBL2arl2oKzMHXQcDmtkB0=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=rUBNAUNm2t/iJPu7kEby4xYbKTDW0iCTb3P6vHZalxYWoYcMzN1tHkGFbcg3qhAZf9bmlMUj4AUYYP9HF+TGDqjIByUdNQCFgX1JZYoDMM2as2RO5w7tKF43L8RQ6Hf5OeoyX64th2nefqUpcKbbswd5Exfxagm1m2sgAEvzdw8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=dy46NG7B; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="dy46NG7B" Received: from pps.filterd (m0353729.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6712vcXo1783679; Sat, 1 Aug 2026 03:55:56 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:message-id:mime-version :subject:to; s=pp1; bh=QNtGJBKM3HhOe0FPzbtRLq7CKBAHna54jM2a+/Lom nY=; b=dy46NG7BKFbsCKcphpkx7fTzq/WYqUVw128G4kTW9k9qh+bUPdbwt8Euv 0VmqQ15P4GjGrCz0rLbfK2sKrrNxkQ5Xp2PopQ9IT3pLiaKjKTU5U8IQNPIIV592 ecKnyRy9CsJZXDrs7XBOlHzjSrmW1+j4G8LX8VT/IldHpgNJFGLa0wFuWw2BfNYG HXeFxvc5TAVC//Ewa/s+42woWDfax9t+n1pKPQHFNQVjMjQwGAbOZJQNW4zyE6N3 zwI+MbDVSO36OjEsY8JsjLdKDMNt1kjsXGsnBGWuI6Rt2IpbJxS1w+jtHfPcJvnH TrEvBgLw8E4U8kTywHBLs67GW4uMw== Received: from ppma23.wdc07v.mail.ibm.com (5d.69.3da9.ip4.static.sl-reverse.com [169.61.105.93]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs8fq83ns-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Sat, 01 Aug 2026 03:55:55 +0000 (GMT) Received: from pps.filterd (ppma23.wdc07v.mail.ibm.com [127.0.0.1]) by ppma23.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6713fJNJ027927; Sat, 1 Aug 2026 03:55:54 GMT Received: from smtprelay07.fra02v.mail.ibm.com ([9.218.2.229]) by ppma23.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fn8yhtmdd-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Sat, 01 Aug 2026 03:55:54 +0000 (GMT) Received: from smtpav04.fra02v.mail.ibm.com (smtpav04.fra02v.mail.ibm.com [10.20.54.103]) by smtprelay07.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6713tqo635848686 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Sat, 1 Aug 2026 03:55:52 GMT Received: from smtpav04.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 8AD9F20043; Sat, 1 Aug 2026 03:55:52 +0000 (GMT) Received: from smtpav04.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 789A120040; Sat, 1 Aug 2026 03:55:49 +0000 (GMT) Received: from li-fdfde5cc-27d0-11b2-a85c-e224154bf6d4.ibm.com.com (unknown [9.43.87.78]) by smtpav04.fra02v.mail.ibm.com (Postfix) with ESMTP; Sat, 1 Aug 2026 03:55:49 +0000 (GMT) From: Madadi Vineeth Reddy To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak Cc: linux-kernel@vger.kernel.org, sh@gentwo.org, Madadi Vineeth Reddy Subject: [PATCH] sched/fair: Let sync wakeups target the waker's core Date: Sat, 1 Aug 2026 09:25:32 +0530 Message-ID: <20260801035532.260625-1-vineethr@linux.ibm.com> X-Mailer: git-send-email 2.54.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: Hw7JjpfyDM8JScBEqADueFUIans4fKYR X-Proofpoint-ORIG-GUID: kNGJAVTY1T0Y-aPz4fPJbetA5VJ5mRFm X-Proofpoint-Spam-Info: AW1haW4tMjYwODAxMDAyMiBTYWx0ZWRfX+sajZ4qfH6Qp Nv16n79ar/JKcSnPAk76JL+5cq14p3aGe1z6BG4M7vsVcHZLGCWit8S4tuabkuQ5RdAyN6yD8QE riC5PR0JiOitMOFF/RRFrrxD4Vu7lTY= X-Authority-Analysis: v=2.4 cv=K8cS2SWI c=1 sm=1 tr=0 ts=6a6d6e4c cx=c_pps a=3Bg1Hr4SwmMryq2xdFQyZA==:117 a=3Bg1Hr4SwmMryq2xdFQyZA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=uAbxVGIbfxUO_5tXvNgY:22 a=VwQbUJbxAAAA:8 a=PuvxfXWCAAAA:8 a=zd2uoN0lAAAA:8 a=VnNF1IyMAAAA:8 a=XddnbLw1eWHioJh3XxsA:9 a=uAr15Ul7AJ1q7o2wzYQp:22 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODAxMDAyMiBTYWx0ZWRfXwYe9L8B2An4V UEP8VSHQRzB/xFjGl1ZyXar/FdgR7TJuUjX+GETYwsHF+wW/KNE3Wf3Y9vgZx7sdTd6xSEYNbrq MdQyCeINGRXcPNJJ3CV2dmd2iIOd2STCkHYgKOJ9VYtMgUyVbVFDW+P2vszY0A43Ycp6tQPSbz6 rSiSBigPqJX08lBN4D/NDHZ9oDsAfYibxwL6XKDAkD+cCs4oLb1+e9ZgIcQWJ3biLhc/aE8DHn3 ZOo87sLEFbjUMdsibDg4GkCVrOghRi7OTEVYjWj+PJnRUzoY6zgaZOyUhRkEfW64cFsHrPq2OLJ tT6UJ7L7kPmHn6qCK709T2sa0gfVahFibOZD7CRCIMP1/XjGqeLiEVLoMPPH7YXDq8tvwz5vSh+ Q3+1TQeWzvzyq/HPiz9Je40RsaGXUlBdfr7e7PceaQvbjbP3RAVi27jyQbK25XjiXnSDoKqNr37 c4NFFko6/V/PNMz++ag== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-07-31_07,2026-07-30_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 spamscore=0 impostorscore=0 bulkscore=0 priorityscore=1501 lowpriorityscore=0 malwarescore=0 phishscore=0 suspectscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608010022 WF_SYNC tells the scheduler the waker is about to block, and wake_affine_idle() already acts on it: when the waker's runqueue holds a single runnable task it returns the waker's CPU. select_idle_sibling() then discards that decision, because available_idle_cpu() is false for a CPU that is still running the waker, and the scan continues elsewhere in the LLC. When the wakee's previous CPU is idle and shares cache with the target, SIS returns it early and nothing is lost. Once prev_cpu is busy, though, the scan has no way to distinguish "this core is busy" from "this core is running only the task that is about to sleep", so the wakee is placed on a cold CPU even though the waker's core is about to have capacity and already holds the data in cache. Pass the waker's CPU down to select_idle_core() in that case and let it count as idle. The waker's core then remains an idle-core candidate and the wakee is placed on one of its sibling threads. The change is a no-op when the waker's core has no idle sibling, when the waker's runqueue holds more than one runnable task, when the wakee's previous CPU is already a valid target, and on non-SMT systems. Tested on POWER11, SMT8, 160 CPUs / 20 cores. 5 runs per case, all figures normalised to baseline. POWER11: [producer_consumer] (time/access, median, lower is better) ================================================== case load base base+this patch time/access -l 5 1.00 0.89 (+11.11%) time/access -l 10 1.00 0.92 ( +7.69%) time/access -l 20 1.00 1.00 ( +0.00%) time/access -l 100 1.00 1.00 ( +0.00%) A strict two-task handoff. As the per-iteration work grows, the wakeup path becomes a smaller fraction of each iteration and placement matters less: 11% at -l 5, nothing from -l 20 upwards. [hackbench] (mean completion time, lower is better; 150000 loops) ================================================================= case load baseline base+this patch sd% process-pipe 1-group 1.00 0.94 ( +6.39%) 6.2 thread-pipe 1-group 1.00 0.92 ( +7.83%) 8.4 process-pipe 2-group 1.00 0.94 ( +6.18%) 7.1 thread-pipe 2-group 1.00 0.97 ( +3.12%) 7.9 process-pipe 4-group 1.00 0.98 ( +1.69%) 7.2 thread-pipe 4-group 1.00 1.09 ( -8.80%) 11.9 process-pipe 8-group 1.00 1.02 ( -2.32%) 3.8 thread-pipe 8-group 1.00 0.99 ( +0.56%) 2.6 process-socket 2-group 1.00 0.95 ( +5.20%) 6.1 thread-socket 2-group 1.00 1.02 ( -2.13%) 5.8 Each group is 40 tasks, so 1 group is 25% of the 160 CPUs and 8 groups is 200%. The patch needs an idle core in the LLC, so its impact shrinks as the utilization increase. Mostly numbers are positive and within run to run variation. [schbench] (mean p99 wakeup latency, lower is better) ===================================================== case load baseline base+this patch sd% p99-latency 8-wkr 1.00 1.06 ( -6.06%) 8.3 p99-latency 40-wkr 1.00 0.95 ( +5.26%) 7.8 p99-latency 80-wkr 1.00 1.08 ( -7.58%) 9.9 p99-latency 240-wkr 1.00 1.00 ( -0.05%) 0.6 [schbench] (mean current rps, higher is better) =============================================== case load baseline base+this patch sd% rps 8-wkr 1.00 1.00 ( -0.17%) 0.4 rps 40-wkr 1.00 1.01 ( +0.87%) 4.9 rps 80-wkr 1.00 0.94 ( -5.55%) 7.9 rps 240-wkr 1.00 1.00 ( -0.28%) 0.3 schbench numbers are within run to run variation and hence not much impacted with this patch. Signed-off-by: Madadi Vineeth Reddy --- Shubhang Kaushik returns the waker CPU directly from the wake-affine branch in select_task_rq_fair(); v3 restricts that to !sched_smt_active(), since on SMT it stacks the pair onto one hardware thread while siblings sit idle: https://lore.kernel.org/lkml/20260727-b4-sched-sync-wakeup-v3-1-90cf481dbd85@gentwo.org/ Prateek proposed moving the decision into select_idle_sibling() under if (!has_idle_core), preferring an idle sibling of prev and then stacking on the waker: https://lore.kernel.org/lkml/f3d5530f-3811-42af-8c34-c40cf314deed@amd.com/ This patch covers the disjoint case: on a sync wakeup the waker's core already holds the data, so where an idle core exists in the LLC this redirects the wakee onto a sibling of the waker's core rather than a cold one, and where none exists it is a no-op. I mentioned this approach on the v2 thread of Shubhang's patch: https://lore.kernel.org/all/60a584c5-25ac-4077-a725-a2f9ee74318d@linux.ibm.com/ --- kernel/sched/fair.c | 49 +++++++++++++++++++++++++++++++++++---------- 1 file changed, 38 insertions(+), 11 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index df8c9c2c7918..bb75b228c817 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1301,7 +1301,7 @@ static bool update_deadline(struct cfs_rq *cfs_rq, struct sched_entity *se) #include "pelt.h" -static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu); +static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu, int sync_cpu); static unsigned long task_h_load(struct task_struct *p); static unsigned long capacity_of(int cpu); @@ -8604,13 +8604,24 @@ void __update_idle_core(struct rq *rq) * sd_balance_shared->has_idle_cores and enabled through update_idle_core() * above. */ -static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpus, int *idle_cpu) +static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpus, + int *idle_cpu, int sync_cpu) { bool idle = true; int cpu; for_each_cpu(cpu, cpu_smt_mask(core)) { - if (!available_idle_cpu(cpu)) { + bool sync_waker = (cpu == sync_cpu); + + /* + * @sync_cpu, if set, is running a waker that is about to + * block with nothing else runnable behind it. Treat it as + * idle so this core stays an idle-core candidate: placing + * the wakee on a sibling keeps the cache sharing that + * stacking on the waker's rq would get, without serialising + * the wakee behind the waker's remaining work. + */ + if (!available_idle_cpu(cpu) && !sync_waker) { idle = false; if (*idle_cpu == -1) { if (choose_sched_idle_rq(cpu_rq(cpu), p) && @@ -8622,7 +8633,12 @@ static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpu } break; } - if (*idle_cpu == -1 && cpumask_test_cpu(cpu, cpus)) + + /* + * The waker is not idle yet, so it must not be offered as + * the fallback target if this core turns out to be busy. + */ + if (!sync_waker && *idle_cpu == -1 && cpumask_test_cpu(cpu, cpus)) *idle_cpu = cpu; } @@ -8661,7 +8677,8 @@ static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int t * comparing the average scan cost (tracked in sd->avg_scan_cost) against the * average idle time for this rq (as found in rq->avg_idle). */ -static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool has_idle_core, int target) +static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool has_idle_core, + int target, int sync_cpu) { struct cpumask *cpus = this_cpu_cpumask_var_ptr(select_rq_mask); int i, cpu, idle_cpu = -1, nr = INT_MAX; @@ -8694,7 +8711,7 @@ static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool continue; if (has_idle_core) { - i = select_idle_core(p, cpu, cpus, &idle_cpu); + i = select_idle_core(p, cpu, cpus, &idle_cpu, sync_cpu); if ((unsigned int)i < nr_cpumask_bits) return i; } else { @@ -8711,7 +8728,7 @@ static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool for_each_cpu_wrap(cpu, cpus, target + 1) { if (has_idle_core) { - i = select_idle_core(p, cpu, cpus, &idle_cpu); + i = select_idle_core(p, cpu, cpus, &idle_cpu, sync_cpu); if ((unsigned int)i < nr_cpumask_bits) return i; @@ -8928,7 +8945,7 @@ static inline bool asym_fits_cpu(unsigned long util, /* * Try and locate an idle core/thread in the LLC cache domain. */ -static int select_idle_sibling(struct task_struct *p, int prev, int target) +static int select_idle_sibling(struct task_struct *p, int prev, int target, int sync_cpu) { bool has_idle_core = false; struct sched_domain *sd; @@ -9037,7 +9054,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) } } - i = select_idle_cpu(p, sd, has_idle_core, target); + i = select_idle_cpu(p, sd, has_idle_core, target, sync_cpu); if ((unsigned)i < nr_cpumask_bits) return i; @@ -9733,8 +9750,18 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag); /* Fast path */ - if (wake_flags & WF_TTWU) - return select_idle_sibling(p, prev_cpu, new_cpu); + if (wake_flags & WF_TTWU) { + int sync_cpu = -1; + + if (want_affine && sync && new_cpu == cpu) { + struct rq *rq = cpu_rq(cpu); + + if ((rq->nr_running - cfs_h_nr_delayed(rq)) == 1) + sync_cpu = cpu; + } + + return select_idle_sibling(p, prev_cpu, new_cpu, sync_cpu); + } return new_cpu; } -- 2.43.0