From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7F24B3FDBEA; Thu, 3 Sep 2026 06:34:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788417278; cv=none; b=l4Frm4Z/m+lf7eBpoQsWN1/8rBJ4SKrHJvJ9dM7WNwA9H3a94uUtxuq7uY9yR57YsnAnKvdEK6LgMrxqB66ZIaJl8+zMBDocITQLKX+vuzfI2LvMHZOphYbRxQmymuR+vKLYjyvwpa/4UqMKv7mXA8coXLutlUenIYNyadQ2BF8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788417278; c=relaxed/simple; bh=FFIqL1vOC52cOedJ/mLnt9LwToWuxsqC0oOXPpdNguw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=teIDc1+34Os60Q39xb14qQeAP3oybYrPWeTKJSEHI2itL2yy13b/7VvroL2FcTu3w9f96q4ipUkA+tcL6X3d/nt9WGh8B+iksFM/vn1U2CuUW+g4T+dsBlsyEXwAksllOWnzxwdB5cAmy4VhiiW1CoRcpyR/ByIwyaekPRd0IY8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=Cujb5GZA; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="Cujb5GZA" Received: from pps.filterd (m0360072.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68361W022239205; Thu, 3 Sep 2026 06:34:12 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=XRkH9dkKGaFlD5hLq brlIRfogZ6fqIq/djjbF/YKrTg=; b=Cujb5GZALG02cO7yQjVDD8+8InvDVPDaD gmVRFPsvodLS4LII8xi10mP9cGuWFMqbCqqbbi2ub69bO+TzJp7rfrLxkTXFWL9D a5WqIqvUidzMJoRl3poyYrvr+fWRMfygW/349/XG26cup+jJGvRq5hA+pdQVa4am P4j1GBt3SBANAuLBwCm/ASjkl8wDxNJPornhG2p9LDayJpHR0gKO9izKA6cArH6i yMsINUd73/cFn4FLVre4VtKsxohPQDXhNFaz8xgAdUvANQ8BQ121+ITys43BCISi /lX4sR7sDVVYVp7N10X6YD/SlnaTcLQoFah7eyTHZPLM5KUs8U6og== Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gbq2tjutb-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Thu, 03 Sep 2026 06:34:11 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 6836QL9u020580; Thu, 3 Sep 2026 06:34:10 GMT Received: from smtprelay02.fra02v.mail.ibm.com ([9.218.2.226]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4gcbygp76f-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Thu, 03 Sep 2026 06:34:10 +0000 (GMT) Received: from smtpav05.fra02v.mail.ibm.com (smtpav05.fra02v.mail.ibm.com [10.20.54.104]) by smtprelay02.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 6836Y50I38928654 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Thu, 3 Sep 2026 06:34:05 GMT Received: from smtpav05.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 1299920043; Thu, 3 Sep 2026 06:34:05 +0000 (GMT) Received: from smtpav05.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id BC1ED20040; Thu, 3 Sep 2026 06:33:56 +0000 (GMT) Received: from li-7bb28a4c-2dab-11b2-a85c-887b5c60d769.ibm.com.com (unknown [9.39.27.193]) by smtpav05.fra02v.mail.ibm.com (Postfix) with ESMTP; Thu, 3 Sep 2026 06:33:56 +0000 (GMT) From: Shrikanth Hegde To: linux-kernel@vger.kernel.org, mingo@kernel.org, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, yury.norov@gmail.com, kprateek.nayak@amd.com, iii@linux.ibm.com, corbet@lwn.net, meted@linux.ibm.com, ynorov@nvidia.com Cc: sshegde@linux.ibm.com, tglx@kernel.org, gregkh@linuxfoundation.org, pbonzini@redhat.com, seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com, rostedt@goodmis.org, dietmar.eggemann@arm.com, maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com, chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org, arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com, tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org, rafael@kernel.org, rdunlap@infradead.org, kernellwp@gmail.com, linux-doc@vger.kernel.org, jgross@suse.com, virtualization@lists.linux.dev, sunlightlinux@gmail.com Subject: [PATCH v12 08/13] sched/core: Push current task from non preferred CPU Date: Thu, 3 Sep 2026 12:02:35 +0530 Message-ID: <20260903063240.268775-9-sshegde@linux.ibm.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260903063240.268775-1-sshegde@linux.ibm.com> References: <20260903063240.268775-1-sshegde@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Authority-Analysis: v=2.4 cv=bc1bluPB c=1 sm=1 tr=0 ts=6a9914e4 cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=RzCfie-kr_QcCd8fBx8p:22 a=VnNF1IyMAAAA:8 a=rPhwLmmYJdnLsqmuV48A:9 X-Proofpoint-Spam-Info: AW1haW4tMjYwOTAzMDA1NyBTYWx0ZWRfX97SmqsMXlZcC i2JBzum/vfbgIAZio0UFmPv/rnXeathvZ+zQFjs4+z6s6dcpSHqCJ/dQWFopyny8IT2b/0uyAr2 DGFPMpkRg9EF7fh9HeN2CmqUXszkZ20= X-Proofpoint-ORIG-GUID: Y9NIcWc22nZjlIX4fZ9GtK4OmzppeqmM X-Proofpoint-GUID: pwXKtR2Q5Kp7zlUrb-sx_HLIqfMcdu8Z X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTAzMDA1NyBTYWx0ZWRfX9WAEfX+LbSRB Z/qOSxFTrrTxsEzpYlf05DJJIrTW+hZdzc9ub3JNBfBMnfpojPlWavQSxlACnhfruFc5HKpfDDM ezviasbRHD47mZQ4n7FtZYLIWx6TgD2NUnIuD0fBY0ePxpb1l2TFLMMCb5oI668UHgP1Bt8isqP U/CyWEVK57Frz3JIqRpenJ4YGsJUA0vzXqaPAa8UOh4XqQoFbFPFaR1rlhEHaRygkLBKbI1Qhab d6h4CloSoUceA1ghBQdWNx84G1+Ik1adgL0mOpUcLjVSaPZsebZK5SFyM317vY7eapXlYPN8pdm vhxgJtn15Vcy41KGHWUrVa9QjCBt60AmcDOUzszIbqDQOWHp26BYTmr7Lk7+o1HswpKUl2rk5VS h7V6LVHZPzJGEij5ykm1sw5nAwu7MmoSldxyvVsGsX5UPWlfeK47tqYZ7OcYhGQZSspEKCFMaFt XP/kAKZzFNofK/jORUA== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-03_02,2026-09-02_04,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 bulkscore=0 suspectscore=0 adultscore=0 malwarescore=0 spamscore=0 lowpriorityscore=0 phishscore=0 clxscore=1015 priorityscore=1501 impostorscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2609030057 Actively push out the current running task on a non-preferred CPU. Since the task is currently running, a stopper thread must be queued to push the task out. However, if the task is pinned only to non-preferred CPUs, it will continue running there. This helps to maintain userspace affinities, unlike CPU hotplug or isolated cpusets. Though the code is similar to __balance_push_cpu_stop and quite close to push_cpu_stop, it is kept separate as it provides a cleaner implementation specifically for CONFIG_PREFERRED_CPU. Add the push_task_work_done flag to protect the work buffer. For now, only the currently running task is pushed out. This keeps the code simpler. In the future, an optimization may be added to move all queued tasks on the runqueue. This works only for the FAIR scheduling class. Signed-off-by: Shrikanth Hegde --- kernel/sched/core.c | 81 ++++++++++++++++++++++++++++++++++++++++++++ kernel/sched/sched.h | 8 +++++ 2 files changed, 89 insertions(+) diff --git a/kernel/sched/core.c b/kernel/sched/core.c index b4ef2e92d786..35e7eedad104 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -5808,6 +5808,9 @@ void sched_tick(void) unsigned long hw_pressure; u64 resched_latency; + if (!cpu_preferred(cpu)) + sched_push_current_non_preferred_cpu(rq); + if (housekeeping_cpu(cpu, HK_TYPE_KERNEL_NOISE)) arch_scale_freq_tick(); @@ -11202,3 +11205,81 @@ void sched_change_end(struct sched_change_ctx *ctx) p->sched_class->prio_changed(rq, p, ctx->prio); } } + +#ifdef CONFIG_PREFERRED_CPU +static DEFINE_PER_CPU(struct cpu_stop_work, npc_push_task_work); + +static int sched_non_preferred_cpu_push_stop(void *arg) +{ + struct task_struct *p = arg; + struct rq *rq = this_rq(); + struct rq_flags rf; + int cpu; + + if (cpu_preferred(rq->cpu)) { + scoped_guard(rq_lock_irqsave, rq) + rq->push_task_work_done = false; + put_task_struct(p); + return 0; + } + + raw_spin_lock_irq(&p->pi_lock); + + /* This could take rq lock. So call it before rq lock is taken */ + cpu = select_fallback_rq(rq->cpu, p); + rq_lock(rq, &rf); + rq->push_task_work_done = false; + update_rq_clock(rq); + + context_unsafe_alias(rq); + + if (task_rq(p) == rq && task_on_rq_queued(p) && + !is_migration_disabled(p)) + rq = __migrate_task(rq, &rf, p, cpu); + + rq_unlock(rq, &rf); + raw_spin_unlock_irq(&p->pi_lock); + put_task_struct(p); + + return 0; +} + +/* + * Push the current task running on non-preferred CPU(npc). + * Using this non preferred CPU will lead to more contention + * in the host. So it is better not to use this CPU. + * + * Since task is running, call a stopper to push the task out. This is + * similar to how task moves during hotplug. In select_fallback_rq a + * preferred CPU will be chosen and henceforth task shouldn't come back to + * this CPU again. + * + * Works for FAIR class only. + * + * If task is affined only on non-preferred CPUs, no point in moving it out. + */ +void sched_push_current_non_preferred_cpu(struct rq *rq) +{ + struct task_struct *push_task = rq->curr; + + scoped_guard(rq_lock, rq) { + /* Push the task if its explicit affinity allows */ + if (!task_can_sched_on_preferred(rq->cpu, push_task)) + return; + + /* There is already a stopper thread. Don't race with it. */ + if (rq->push_task_work_done) + return; + + if (is_migration_disabled(push_task)) + return; + + rq->push_task_work_done = true; + } + + /* sched_tick runs with interrupts disabled. */ + get_task_struct(push_task); + stop_one_cpu_nowait(rq->cpu, sched_non_preferred_cpu_push_stop, + push_task, this_cpu_ptr(&npc_push_task_work)); +} +#endif diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 6c3ad70e58b8..678e44134acf 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -1298,6 +1298,8 @@ struct rq { struct list_head cfs_tasks; + bool push_task_work_done; + struct sched_avg avg_rt; struct sched_avg avg_dl; #ifdef CONFIG_HAVE_SCHED_AVG_IRQ @@ -4280,4 +4282,10 @@ DEFINE_CLASS_IS_UNCONDITIONAL(sched_change) #include "ext/ext.h" +#ifdef CONFIG_PREFERRED_CPU +void sched_push_current_non_preferred_cpu(struct rq *rq); +#else /* !CONFIG_PREFERRED_CPU */ +static inline void sched_push_current_non_preferred_cpu(struct rq *rq) { } +#endif + #endif /* _KERNEL_SCHED_SCHED_H */ -- 2.52.0