From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AA9F0351C20 for ; Thu, 18 Jun 2026 04:18:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781756304; cv=none; b=eBCHfUbeCzqkxrn6/TNZxTASY+UtkIOTZW/jB8IRVW43D29en+rsH+K86i2rpmed7LhRoSXdrS2xkAwkPq0bRqtsE00LrimCx+U91Vtb7eluqf6eXZZNoqNF42kAjx1iyFwKL/E9iVVUgKoQ4zpR+K4TE9ZgvRCM5dsgCu8vx0k= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781756304; c=relaxed/simple; bh=KIFE1mR3DpqZTyAb1y3uxc/4N9grW6bAMWESkxEcrf0=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=Sau+WjsahB5s/ZoPFw+VCvbWL/2d3D7iYlbvQVs5MdmoTJmRvaVzGSzfV/3+R5mPcoKbu5Ju0mQqMptPqFyUdJfZyJNEUpY8LECOWTd0EXXlgbOlgDt8dTg1elingK0wn8GtgcQTWKb4WBJ6SDtoWMU2JOYXO6DVO1F+rpFVV5I= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=XvkH8rL9; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="XvkH8rL9" Received: from pps.filterd (m0360072.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 65HHnBaN979279; Thu, 18 Jun 2026 04:18:02 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=Vh/SDQ WdQMdMA7Z2DVSkcFywgYdA3VnPYxcI0UL6pew=; b=XvkH8rL9fjoNEgb1o39CYx Ma8MmXzkHfjG9jbiBUcLVgG99badql55yfr+UdlvUy1z+l0VTDN0sfIJdCzLEWvh 7fJI1um6/bqGfPN30lcALSzUJ4XinrM5Rp2oHqS2VVyDNfTMfsPdn2jwcKlkMkxo W7cXxBRsmVXJcoLQwgO7aFVyOblcLpaujSmHHvJl3oIPKYhhkr0Dj4N9CApQzvb0 KedjJ6MyR5YgcCosFu+2rZkySJLXBDdQlDUm1efRuHoUBP9mo7Cri2VOZZpfYfr/ J6IOEcWdCqO+giQTtl4XbSo7KHGWH3hnvaxT4QF5v0C6cfJD7+uOOAmP96YVZxkg == Received: from ppma22.wdc07v.mail.ibm.com (5c.69.3da9.ip4.static.sl-reverse.com [169.61.105.92]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4eueqx65sp-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Thu, 18 Jun 2026 04:18:01 +0000 (GMT) Received: from pps.filterd (ppma22.wdc07v.mail.ibm.com [127.0.0.1]) by ppma22.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 65I44ePk003430; Thu, 18 Jun 2026 04:18:00 GMT Received: from smtprelay02.fra02v.mail.ibm.com ([9.218.2.226]) by ppma22.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4ev1729wc0-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Thu, 18 Jun 2026 04:18:00 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (smtpav02.fra02v.mail.ibm.com [10.20.54.101]) by smtprelay02.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 65I4HudS41681200 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 18 Jun 2026 04:17:56 GMT Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 6C53220043; Thu, 18 Jun 2026 04:17:56 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 3E02A20040; Thu, 18 Jun 2026 04:17:50 +0000 (GMT) Received: from [9.123.5.233] (unknown [9.123.5.233]) by smtpav02.fra02v.mail.ibm.com (Postfix) with ESMTP; Thu, 18 Jun 2026 04:17:50 +0000 (GMT) Message-ID: Date: Thu, 18 Jun 2026 09:47:49 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v4 06/20] sched/core: allow only preferred CPUs in is_cpu_allowed To: Yury Norov Cc: linux-kernel@vger.kernel.org, mingo@kernel.org, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, kprateek.nayak@amd.com, iii@linux.ibm.com, tglx@kernel.org, gregkh@linuxfoundation.org, pbonzini@redhat.com, seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com, rostedt@goodmis.org, dietmar.eggemann@arm.com, mgorman@suse.de, bsegall@google.com, maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com, chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org, arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com, tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org, rafael@kernel.org References: <20260617174139.155540-1-sshegde@linux.ibm.com> <20260617174139.155540-7-sshegde@linux.ibm.com> From: Shrikanth Hegde Content-Language: en-US In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: UpiAmozfczinym3Zew_ITQXKUw3e8cXr X-Authority-Analysis: v=2.4 cv=Le0MLDfi c=1 sm=1 tr=0 ts=6a337179 cx=c_pps a=5BHTudwdYE3Te8bg5FgnPg==:117 a=5BHTudwdYE3Te8bg5FgnPg==:17 a=IkcTkHD0fZMA:10 a=FelO9ux0wxsA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=RzCfie-kr_QcCd8fBx8p:22 a=VnNF1IyMAAAA:8 a=6LVG20tIGdCBSfm8B2IA:9 a=QEXdDO2ut3YA:10 X-Proofpoint-Spam-Info: AW1haW4tMjYwNjE4MDAzMCBTYWx0ZWRfXz/+vg0dKMVdM mkO7U7TUapI7AIQW3v0lk2+x3XxYsmM3AKfpN72IU/QLh3Gh7N3rCIfTSW5sBrkLothWw4fE1Kj 2o/SMHfH8n/2qRAcrnJ3ykh/oQ30Lx0= X-Proofpoint-ORIG-GUID: 7ad-Wrq0Sv6E4eFG2NiAMNR4eWstaTOx X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNjE4MDAzMCBTYWx0ZWRfXyNcsBT5HJgkL H3JonmX/ek45BnvGCheFiTXWLjXGtON38wQt/A8B0/BjUlWKfxZ6GiSveerRpeYEYdEi/dcY9/f Mwj1jOZP3WIVWjStJmuWvbGWshpbUVaBruylMVTU/WT9ZC1u64Dbo12BytfYNmlRbsm/OvTL2FC Nf1CpJbkhPvhxswFyPltBgTmO9uRfx1KJboa/uXn6CWMyr60NEQpHMWq8+7b7AdddLmMnvBqQAT jej+92DFSlaTHvPVoIyj7mn3PiuIotRmOOrHrMz4/fZaTbXnJDzuzOj2WUoaLIE2q0dUoqQYFI/ SEB+EnLe2l1EodyCOUyiHHpFNu9Jt2XO5d1L7AAaFgqiWSaBIk8fB/aoBFOtV3HjEa2QJKOgC5l ApyFDvo4k0sf/q/8ftNFH3oGdvNnmKliS94BId/8aLcBHXyBsoKUXbCOWrj5+DrSZUvGDIa46Yo DkropqYWFjtIH0pu3pg== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.125,FMLib:17.12.100.49 definitions=2026-06-17_02,2026-06-17_03,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 clxscore=1015 priorityscore=1501 malwarescore=0 spamscore=0 suspectscore=0 impostorscore=0 phishscore=0 lowpriorityscore=0 bulkscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2606180030 On 6/18/26 9:02 AM, Yury Norov wrote: > On Wed, Jun 17, 2026 at 11:11:25PM +0530, Shrikanth Hegde wrote: >> When possible, choose a preferred CPUs to pick. >> >> Push task mechanism uses stopper thread which going to call >> select_fallback_rq and use this mechanism to pick only a preferred CPU. >> >> When task is affined only to non-preferred CPUs it should continue to >> run there. Detect that by checking if cpus_ptr and cpu_preferred_mask >> intersect or not. >> >> Since is_cpu_allowed can be called directly or repeatedly in >> select_fallback_rq, encode the info in task_struct->has_preferred_cpu_state >> if the path is via select_fallback_rq or not. >> This helps to avoid N**2 complexity for the rare cases. >> >> Signed-off-by: Shrikanth Hegde >> --- >> v3->v4: >> - Missing case of PF_KTHREAD is avoided. >> - Add a new field in task_struct which encodes intersection of >> tasks affinity and preferred CPUs and path its coming from. >> >> include/linux/sched.h | 1 + >> kernel/sched/core.c | 34 ++++++++++++++++++++++++++++++++-- >> kernel/sched/sched.h | 18 ++++++++++++++++++ >> 3 files changed, 51 insertions(+), 2 deletions(-) >> >> diff --git a/include/linux/sched.h b/include/linux/sched.h >> index fc6ecb3869dd..2d0b1a6d50ac 100644 >> --- a/include/linux/sched.h >> +++ b/include/linux/sched.h >> @@ -1657,6 +1657,7 @@ struct task_struct { >> #ifdef CONFIG_UNWIND_USER >> struct unwind_task_info unwind_info; >> #endif >> + int has_preferred_cpu_state; > > Shouldn't this be protected with the config? Since preferred is defined always, i don;t see a reason to add it again here. > >> >> /* CPU-specific state of this task: */ >> struct thread_struct thread; >> diff --git a/kernel/sched/core.c b/kernel/sched/core.c >> index 9e16946c9d62..714816cfa975 100644 >> --- a/kernel/sched/core.c >> +++ b/kernel/sched/core.c >> @@ -2500,6 +2500,8 @@ static inline bool rq_has_pinned_tasks(struct rq *rq) >> */ >> static inline bool is_cpu_allowed(struct task_struct *p, int cpu) >> { >> + bool task_check_preferred_cpu = false; > > Initialization is not needed. ok > >> + >> /* When not in the task's cpumask, no point in looking further. */ >> if (!task_allowed_on_cpu(p, cpu)) >> return false; >> @@ -2508,9 +2510,22 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu) >> if (is_migration_disabled(p)) >> return cpu_online(cpu); >> >> + /* >> + * This is essential to maintain user affinities when preferred >> + * CPUs change. A task pinned on non-preferred CPU should continue >> + * to run there, since this is non-user triggered. >> + * >> + * If CPU is non-preferred and task can run on other CPUs which are >> + * currently preferred, then choose those other CPUs instead >> + */ >> + task_check_preferred_cpu = !cpu_preferred(cpu) && task_has_preferred_cpus(p); >> + >> /* Non kernel threads are not allowed during either online or offline. */ >> - if (!(p->flags & PF_KTHREAD)) >> + if (!(p->flags & PF_KTHREAD)) { >> + if (task_check_preferred_cpu) >> + return false; >> return cpu_active(cpu); >> + } >> >> /* KTHREAD_IS_PER_CPU is always allowed. */ >> if (kthread_is_per_cpu(p)) >> @@ -2520,6 +2535,10 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu) >> if (cpu_dying(cpu)) >> return false; >> >> + /* Try on preferred CPU first if possible*/ >> + if (task_check_preferred_cpu) >> + return false; >> + >> /* But are allowed during online. */ >> return cpu_online(cpu); >> } >> @@ -3549,6 +3568,14 @@ static int select_fallback_rq(int cpu, struct task_struct *p) >> enum { cpuset, possible, fail } state = cpuset; >> int dest_cpu; >> >> + /* >> + * Cache value whether task's affinity spans preferred CPUs. > > Because it's cached, it should go inside is_cpu_allowed(), I think. > >> + * This helps to avoid repeating the same for each CPU >> + * later in the loop. Encode call to is_cpu_allowed coming >> + * via select_fallback_rq. >> + */ >> + p->has_preferred_cpu_state = task_has_preferred_cpus(p) << 8 | 0x1; > > This looks weird. Your intention is to store three states: not cached, has > preferred CPUs and has not preferred CPUs, > > Why don't you create an enum for it? Or a couple of flags? I think what prateek suggested in other thread looks same. I will give that a try. > >> + >> /* >> * If the node that the CPU is on has been offlined, cpu_to_node() >> * will return -1. There is no CPU on the node, and we should >> @@ -3560,7 +3587,7 @@ static int select_fallback_rq(int cpu, struct task_struct *p) >> /* Look for allowed, online CPU in same node. */ >> for_each_cpu(dest_cpu, nodemask) { >> if (is_cpu_allowed(p, dest_cpu)) >> - return dest_cpu; >> + goto clear_and_return; >> } >> } >> >> @@ -3604,6 +3631,8 @@ static int select_fallback_rq(int cpu, struct task_struct *p) >> } >> } >> >> +clear_and_return: >> + p->has_preferred_cpu_state = 0; > It is reset to indicate that any subsequent direct calls to is_cpu_allowed can't use the old cached value of select_fallback_rq. So events could be, - cpu marked as non preferred - select_fallback_rq (sets the p->has_preferred_cpu_state) Lets say CPU(300-450) are marked as non-preferred and Task affinity is (200-350) - task moved out. Now either task's affinity changed or preferred_mask has changed. while CPU(400) maybe still marked as non-preferred but CPU(340) is marked as preferred. - Subsequent call to is_cpu_allowed (CPU=340) can't assume the old value. > What for resetting it here? I think it should be zeroed only on update > of preferred cpumask. In other words, to properly implement caching, > you need to have a global counter incremented on each > cpu_preferred_mask update, and in task_has_preferred_cpus() you do: > > { > if (p->preferred_cpu_updates == atomic_read(preferred_cpumask_updates)) > return p->has_preferred_cpus; > > p->preferred_cpu_updates = atomic_read(preferred_cpumask_updates); > p->has_preferred_cpus = cpumask_intersects(...); > } > > Do you have any numbers that justify this caching? The best practice > is to put performance optimizations at the end of the series and > provide some sort of benchmark supporting it. > This was to avoid N**2 aspect that was there in select_fallback_rq. Its more of the functional aspect which i mentioned above which this needs to take care as well. >> return dest_cpu; >> } >> >> @@ -4612,6 +4641,7 @@ static void __sched_fork(u64 clone_flags, struct task_struct *p) >> init_numa_balancing(clone_flags, p); >> p->wake_entry.u_flags = CSD_TYPE_TTWU; >> p->migration_pending = NULL; >> + p->has_preferred_cpu_state = 0; >> init_sched_mm(p); >> } >> >> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h >> index c7c2dea65edd..38fd84b0b8f8 100644 >> --- a/kernel/sched/sched.h >> +++ b/kernel/sched/sched.h >> @@ -4213,4 +4213,22 @@ DEFINE_CLASS_IS_UNCONDITIONAL(sched_change) >> >> #include "ext.h" >> >> +/* >> + * has_preferred_cpu_state is encoding two bits of information. >> + * First Byte is to encode where the call to is_cpu_allowed coming from. >> + * Second Byte is to encode the intersection of task affinity >> + * and cpu_preferred_mask. >> + * >> + * If 1st Byte is set, call to is_cpu_allowed coming from select_fallback_rq. >> + * That helps to avoid repeated calculation keeping time complexity same. >> + */ >> +static inline bool task_has_preferred_cpus(struct task_struct *p) > > This function should be void because you change the task state. > It doesn't alter p->has_preferred_cpu_state. No? >> +{ >> + int cached_value = p->has_preferred_cpu_state; >> + >> + if (cached_value & 0x1) >> + return p->has_preferred_cpu_state >> 8; >> + else >> + return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask); >> +} >> #endif /* _KERNEL_SCHED_SCHED_H */ >> -- >> 2.47.3