From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 573483AD501 for ; Tue, 4 Aug 2026 12:14:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845645; cv=none; b=BxjxVl0KDGjNztJa6FeudMAZs2yIWfx91m5M4EROevFz4xd7/k78enpacNKReVgwu9Rkjm+psriAC3vZw9lvHKdwaK2edYLVuyWJKCcvncv7mh0X6RfcwTTtYj6YmcovScai6qRD8E9CiTzjzBdnzCJld/iB+YzwE+eIq2fvw24= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845645; c=relaxed/simple; bh=2ieBMczgxKiReA72i7aosV4dp/MllKVQ9dIB5xIxmSI=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=i8nk4pKGMe9VZazg37T9xugBk+yuh1j1mVQoMtNl9YeC1IrOF45c+S2NM5aX+pZSKaq38rRYFb5FPSNekfsjtdOlhbMah4TxD0mh2X01ivYHOz/3gzspu82Y1+P6TgOTCx5b0Z5wPeuC2r58SnCcLOtYWe0YGh/UvAAExI7aLtc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=E9gyPDW8; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="E9gyPDW8" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 6748IVLT236337; Tue, 4 Aug 2026 12:13:38 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=j2Vnhd oNBXfpAgPB6nH2eeAcokcx1ijlo1O39dNv2z4=; b=E9gyPDW8mL8L5Wfo9Qxpm5 RfoEkBHGAeEP5H9ML4zH0k5s9zyhgvRQNEBZbkQxOahsQOEXqLHk3wLZ9rQjPe5r hqKt5M42V3lJC9sXSRACh2KfMyWEYCNwJnGkL/P3CVQuHGChpNSuwDbnUUYyA1r0 DgNewbjLi7TPyCLAzZVbqD5trTT7vFtTGiUwHCZgTKGoMvmOOsADMmX9Pk3BtS2a QGRdazucCFziydsr6zv3auQMydx9R3EyOrO/pvl407pYzBMcvHQvtQhBPKeWgW1a hfigQECkUltSvz6ADe5Y32OugewE5zjM4fE5NeblveCsFztjjOinjd6cMoCni+BA == Received: from ppma23.wdc07v.mail.ibm.com (5d.69.3da9.ip4.static.sl-reverse.com [169.61.105.93]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fs67hnk63-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Tue, 04 Aug 2026 12:13:37 +0000 (GMT) Received: from pps.filterd (ppma23.wdc07v.mail.ibm.com [127.0.0.1]) by ppma23.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 674CBK65032146; Tue, 4 Aug 2026 12:13:37 GMT Received: from smtprelay02.wdc07v.mail.ibm.com ([172.16.1.69]) by ppma23.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4fsvmh9sd6-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Tue, 04 Aug 2026 12:13:37 +0000 (GMT) Received: from smtpav03.dal12v.mail.ibm.com (smtpav03.dal12v.mail.ibm.com [10.241.53.102]) by smtprelay02.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 674CDalT23724712 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Tue, 4 Aug 2026 12:13:36 GMT Received: from smtpav03.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 76FA25805A; Tue, 4 Aug 2026 12:13:36 +0000 (GMT) Received: from smtpav03.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 3C0F85803F; Tue, 4 Aug 2026 12:13:32 +0000 (GMT) Received: from [9.43.77.145] (unknown [9.43.77.145]) by smtpav03.dal12v.mail.ibm.com (Postfix) with ESMTP; Tue, 4 Aug 2026 12:13:31 +0000 (GMT) Message-ID: <40536fd1-cdc6-4fac-a78f-1ed0df1fcce6@linux.ibm.com> Date: Tue, 4 Aug 2026 17:43:30 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] sched/fair: Let sync wakeups target the waker's core To: K Prateek Nayak Cc: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , linux-kernel@vger.kernel.org, sh@gentwo.org, Madadi Vineeth Reddy References: <20260801035532.260625-1-vineethr@linux.ibm.com> <492e6bb3-504d-486a-ad9d-226e9d7235a3@amd.com> Content-Language: en-US From: Madadi Vineeth Reddy In-Reply-To: <492e6bb3-504d-486a-ad9d-226e9d7235a3@amd.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-Spam-Info: AW1haW4tMjYwODA0MDA5NCBTYWx0ZWRfXzWSOZT3znB2M Bv1Hfun97WxNQffPgLRSmEWGw1t3eqcntXy15YirsGfQcJ94dlQcGZO2GLsiH87iS8uDfk2cuM5 r9Th1IxhNOQ59SUNlZ6Zg6cIg0Lx44M= X-Authority-Analysis: v=2.4 cv=I7VVgtgg c=1 sm=1 tr=0 ts=6a71d772 cx=c_pps a=3Bg1Hr4SwmMryq2xdFQyZA==:117 a=3Bg1Hr4SwmMryq2xdFQyZA==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=gNxIMHBx9n27ugbkDLAA:9 a=QEXdDO2ut3YA:10 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA0MDA5NCBTYWx0ZWRfXxTlOU71pzSBh LBUWXYs7w5B1k6/Q14WFEU9b5ypcWOaeZzAN1tCThI0B9IasQnHJ8/ow/fGmTmE5zeArTPnNKij GXZEwc+9Kek/bcmBrnnKZp/hORwL90iouiqnvyLpeJqJoUvLUpYnsCNuAf09hBGtxvVUKOJ0Gz5 bdjAhkHA7UNzop8aP/pblIkU7idvAQBT9YBfe6ZxQWdRyrCqTO1nw1ywWyf4QN20/qxHvH33HyC jdkDd//mJ1xlERKok7GdLoPVx4MfQi9P3CbYwAcT+7gZ1U9iez87MQwPi8ZeTP00Mj0eqXIsOsA a02jMC47HlImgWZWM/cDBhiUruvCNqb+PGCpIM8QeIt9fDuL6/R9QfisfiqeVLaND9t71g7zOGi dZTWzn9WxMDEH9cF8NPVu0F3gjn7Ao9OpUKw/H/PYxTsGyTIQYj7z7f5UN074WS/7zROGKieDUe UiFIU6wvqbkejFp/2ag== X-Proofpoint-ORIG-GUID: MSPSBKLFIIbn5rquJpcDUqgIZlnPMbnN X-Proofpoint-GUID: efgMUBFUO-r0c_IqVCIOo7EzQErWZ4HU X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-04_02,2026-08-03_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 bulkscore=0 impostorscore=0 clxscore=1015 priorityscore=1501 suspectscore=0 malwarescore=0 adultscore=0 lowpriorityscore=0 phishscore=0 spamscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608040094 Hello Prateek, On 04/08/26 10:19, K Prateek Nayak wrote: > Hello Vineeth, > > On 8/1/2026 9:25 AM, Madadi Vineeth Reddy wrote: >> -static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpus, int *idle_cpu) >> +static int select_idle_core(struct task_struct *p, int core, struct cpumask *cpus, >> + int *idle_cpu, int sync_cpu) >> { >> bool idle = true; >> int cpu; >> >> for_each_cpu(cpu, cpu_smt_mask(core)) { >> - if (!available_idle_cpu(cpu)) { >> + bool sync_waker = (cpu == sync_cpu); >> + >> + /* >> + * @sync_cpu, if set, is running a waker that is about to >> + * block with nothing else runnable behind it. Treat it as >> + * idle so this core stays an idle-core candidate: placing >> + * the wakee on a sibling keeps the cache sharing that >> + * stacking on the waker's rq would get, without serialising >> + * the wakee behind the waker's remaining work. >> + */ >> + if (!available_idle_cpu(cpu) && !sync_waker) { > > If I'm not wrong, all you want to make is the sync_waker appear idle and > then see if you can then consider that core as idle core or not right? > Correct. > Why can't this be done in select_idle_sibling() extending that early > check for (!has_idle_core && cpus_share_cache(prev, target)) condition > and then initializing "idle_cpu" in select_idle_cpu() accordingly? > > Something along the lines of: > > (Only build tested) > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > index df8c9c2c7918..dd62bceb3838 100644 > --- a/kernel/sched/fair.c > +++ b/kernel/sched/fair.c > @@ -1301,7 +1301,6 @@ static bool update_deadline(struct cfs_rq *cfs_rq, struct sched_entity *se) > > #include "pelt.h" > > -static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu); > static unsigned long task_h_load(struct task_struct *p); > static unsigned long capacity_of(int cpu); > > @@ -8661,10 +8660,11 @@ static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int t > * comparing the average scan cost (tracked in sd->avg_scan_cost) against the > * average idle time for this rq (as found in rq->avg_idle). > */ > -static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool has_idle_core, int target) > +static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, bool has_idle_core, > + int target, int idle_cpu) > { > struct cpumask *cpus = this_cpu_cpumask_var_ptr(select_rq_mask); > - int i, cpu, idle_cpu = -1, nr = INT_MAX; > + int i, cpu, nr = INT_MAX; > > if (sched_feat(SIS_UTIL) && sd->shared) { > /* > @@ -8928,7 +8928,7 @@ static inline bool asym_fits_cpu(unsigned long util, > /* > * Try and locate an idle core/thread in the LLC cache domain. > */ > -static int select_idle_sibling(struct task_struct *p, int prev, int target) > +static int select_idle_sibling(struct task_struct *p, int prev, int target, int sync_cpu) > { > bool has_idle_core = false; > struct sched_domain *sd; > @@ -9028,16 +9028,19 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) > return target; > > if (sched_smt_active()) { > + int cpu = ((unsigned)sync_cpu < nr_cpumask_bits) ? sync_cpu : prev; > + > has_idle_core = test_idle_cores(target); > > - if (!has_idle_core && cpus_share_cache(prev, target)) { > - i = select_idle_smt(p, sd, prev); > - if ((unsigned int)i < nr_cpumask_bits) > + if (sync_cpu == target || (!has_idle_core && cpus_share_cache(prev, target))) { > + i = select_idle_smt(p, sd, cpu); > + > + if (!has_idle_core && ((unsigned int)i < nr_cpumask_bits)) > return i; > } > } > > - i = select_idle_cpu(p, sd, has_idle_core, target); > + i = select_idle_cpu(p, sd, has_idle_core, target, i); `i` which is passed could be garbage value if we don't enter sched_smt_active block. > if ((unsigned)i < nr_cpumask_bits) > return i; > This is different from what I wanted to achieve in a couple of ways. - Calling `select_idle_smt()` in sync case, would only give an idle CPU in that core but doesn't test if that core is idle. That would be lost. - You return `i` only when `!has_idle_core`, but I wanted to return waker core given that rest of the siblings in that waker core are idle even though there are other idle cores present in the LLC. I agree that this could be done in `select_idle_sibling` but with a helper function. Something like below (build tested) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index df8c9c2c7918..448af3c4b183 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1301,7 +1301,7 @@ static bool update_deadline(struct cfs_rq *cfs_rq, struct sched_entity *se) #include "pelt.h" -static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu); +static int select_idle_sibling(struct task_struct *p, int prev_cpu, int cpu, bool sync_core); static unsigned long task_h_load(struct task_struct *p); static unsigned long capacity_of(int cpu); @@ -8656,6 +8656,27 @@ static int select_idle_smt(struct task_struct *p, struct sched_domain *sd, int t return -1; } +static int select_idle_sync_core(struct task_struct *p, struct sched_domain *sd, + int target) +{ + int cpu, idle_sibling = -1; + + for_each_cpu(cpu, cpu_smt_mask(target)) { + if (cpu == target) + continue; + + if (!available_idle_cpu(cpu)) + return -1; + + if (idle_sibling == -1 && + cpumask_test_cpu(cpu, sched_domain_span(sd)) && + cpumask_test_cpu(cpu, p->cpus_ptr)) + idle_sibling = cpu; + } + + return idle_sibling; +} + /* * Scan the LLC domain for idle CPUs; this is dynamically regulated by * comparing the average scan cost (tracked in sd->avg_scan_cost) against the @@ -8928,7 +8949,7 @@ static inline bool asym_fits_cpu(unsigned long util, /* * Try and locate an idle core/thread in the LLC cache domain. */ -static int select_idle_sibling(struct task_struct *p, int prev, int target) +static int select_idle_sibling(struct task_struct *p, int prev, int target, bool sync_core) { bool has_idle_core = false; struct sched_domain *sd; @@ -9035,6 +9056,12 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target) if ((unsigned int)i < nr_cpumask_bits) return i; } + + if (sync_core) { + i = select_idle_sync_core(p, sd, target); + if ((unsigned int)i < nr_cpumask_bits) + return i; + } } i = select_idle_cpu(p, sd, has_idle_core, target); @@ -9733,8 +9760,16 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag); /* Fast path */ - if (wake_flags & WF_TTWU) - return select_idle_sibling(p, prev_cpu, new_cpu); + if (wake_flags & WF_TTWU) { + bool sync_core = false; + if (want_affine && sync && new_cpu == cpu) { + struct rq *rq = cpu_rq(cpu); + + sync_core = (rq->nr_running - cfs_h_nr_delayed(rq)) == 1; + } + + return select_idle_sibling(p, prev_cpu, new_cpu, sync_core); + } return new_cpu; } This should also solve the issue raised by Zhan Xusheng and first target waker core given that it is idle by giving exception to waker cpu. Thoughts? Thanks, Vineeth > @@ -9733,8 +9736,18 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags) > return sched_balance_find_dst_cpu(sd, p, cpu, prev_cpu, sd_flag); > > /* Fast path */ > - if (wake_flags & WF_TTWU) > - return select_idle_sibling(p, prev_cpu, new_cpu); > + if (wake_flags & WF_TTWU) { > + int sync_cpu = -1; > + > + if (want_affine && sync && new_cpu == cpu) { > + struct rq *rq = cpu_rq(cpu); > + > + if ((rq->nr_running - cfs_h_nr_delayed(rq)) == 1) > + sync_cpu = cpu; > + } > + > + return select_idle_sibling(p, prev_cpu, new_cpu, sync_cpu); > + } > > return new_cpu; > } > --- > > You can probably infer sync hint by checking > "target == smp_preocessor_id()" too in select_idle_sibling() instead of > passing it on. > >> idle = false; >> if (*idle_cpu == -1) { >> if (choose_sched_idle_rq(cpu_rq(cpu), p) &&