From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC7C436F900 for ; Mon, 10 Aug 2026 10:38:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786358339; cv=none; b=RSsFmn1nhnoy4Idd9y//dNy9juVojuj1pQGZP1PH2Qs04GqZ6xL4qvLIBHSefdgRcXj7myIEii1XsKrbwTSCVrHacElzaD0gfUuK/HtDI1rl7e+cLqeml4UCd4NrF9q0AOXouqlhpw+ha/Y3+zjz/4IRsG6i4WoX1evmENbhu9k= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786358339; c=relaxed/simple; bh=NxhFv5lyzFLf2vSzmUy37aBN2gSWkr0pikoNjTpW8ZU=; h=Message-ID:Date:MIME-Version:Subject:From:To:Cc:References: In-Reply-To:Content-Type; b=NmW/c09GkwXljcevoGs2nLiyXlfeMqqFAt6MHIbOe9LQXstykle0HNu58DpZWkiDpKpAah7UDiWa/dUhz9YOGG6LNMFvVAxdyAABzxxhCJrKFxOn1DphmvwWy6TmJTeYxzAgrrz4twMWWV6z5fmXS48wMDeRtALnBz1DVfbhj2A= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=mv37nLvv; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="mv37nLvv" Received: from pps.filterd (m0360083.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67A9VkYV1209411; Mon, 10 Aug 2026 10:38:31 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=cd2+8w JoAsrwnFPXV1bVbo8tBMhRyJTEkzQRSYC6rzU=; b=mv37nLvvaCSTBS9iOdSkpu OxTAm2mslHjxw8ulFo+GC8J818WwgiqXMyYAXBEZCo8WdIvfFkBvhMLYtJSUkF7+ v8xuWDDwfU8Xb7D4bDnwcZSRj2nx3zjgv6AaxPcEnlsfOMlhx7Z3FT2Y70XIEndc //zxBjlju3eC6q6BraUuvoUk5ee65LV0ltRzY1gwXK+xpfC6gRHV6uqgY7P5AVwX k2x3da9EcavFSvlf+wJLIU1c9k5GWdnF3Mi4WyGmVrp4aqQ+N+c+bJR1YPjukf5J Gf/DW/i7jQhX+DfvwpVOfPXWaJvt73mfMIPuUGhwCxN9XVIe3+3SPR3geplJTGZw == Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fwvq97esb-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Mon, 10 Aug 2026 10:38:30 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 67AAQYVO011885; Mon, 10 Aug 2026 10:38:29 GMT Received: from smtprelay01.fra02v.mail.ibm.com ([9.218.2.227]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fxh0g48js-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Mon, 10 Aug 2026 10:38:29 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (smtpav02.fra02v.mail.ibm.com [10.20.54.101]) by smtprelay01.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 67AAcR2Z34603470 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Mon, 10 Aug 2026 10:38:27 GMT Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 815532004B; Mon, 10 Aug 2026 10:38:27 +0000 (GMT) Received: from smtpav02.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 1A06620043; Mon, 10 Aug 2026 10:38:24 +0000 (GMT) Received: from [9.39.23.237] (unknown [9.39.23.237]) by smtpav02.fra02v.mail.ibm.com (Postfix) with ESMTP; Mon, 10 Aug 2026 10:38:23 +0000 (GMT) Message-ID: <3fb4a198-7221-48cf-9481-1728e8ca12ad@linux.ibm.com> Date: Mon, 10 Aug 2026 16:08:23 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v5] sched/fair: Prefer fully idle cores for NOHZ balancing From: Shrikanth Hegde To: Andrea Righi , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot Cc: Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Christian Loehle , Phil Auld , Mete Durlu , linux-kernel@vger.kernel.org References: <20260806134355.592145-1-arighi@nvidia.com> Content-Language: en-US In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-GUID: 8WEhhkqYSyLrdvFxw2I7Pjw6HgOJEz0X X-Authority-Analysis: v=2.4 cv=PbDPQChd c=1 sm=1 tr=0 ts=6a79aa26 cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=iQ6ETzBq9ecOQQE5vZCe:22 a=VwQbUJbxAAAA:8 a=zd2uoN0lAAAA:8 a=VnNF1IyMAAAA:8 a=KKAkSRfTAAAA:8 a=Ikd4Dj_1AAAA:8 a=mKuV-UcRH07PTvzR3PAA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 a=cvBusfyB2V15izCimMoJ:22 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODEwMDA5MSBTYWx0ZWRfX9eE/AxGYxp3k i/ofKAdA9Nihw6saho91uZ9xlLfIhvkez5irdcfkMZmLLira5VOplYTCF85U2AoNfkPZq8ewt+T ZFHOCeRU9EeZ0roj9WY+ZnfZf/KL/n2dPFs0oLUWvciu7FuGxiiGvYxcIk7EJsuF9cvQNseE3z0 WSEeKiV3rJAEWOJ5G36o8YXTLcemSeu+acvm+CzXlBC77NFtossGia1/aqC7G5B5RS0mtbX2OV9 7WQdlVAVqgvSHY6E+ew8jcB0LC0hSqHCGZpQ4PwOy9lv6h/RoVk/BabYj6Vdee4n62v6Kv+4yOH w8Q/qEt2ncxwHX91ugxFdC+J4/znVjb0LQMb42v0EWlSshNr7bVGK9A5WosCDOqP+NYmSqKjgTA cZ47cl39SxY1ayvdAyKioxMGqV2ZAfnW/zmf101e4gNNz434WWkeExvCFQaeqScT5GWf4c3v2nU sGEQqTHa8Bi5aZShg5g== X-Proofpoint-ORIG-GUID: 4QM7Knt9r5k_bOkh8b9rnVA2oZskuLFT X-Proofpoint-Spam-Info: AW1haW4tMjYwODEwMDA5MSBTYWx0ZWRfX5+WaplJ8VUY6 S7iqkJiCvPrbl06TsxKdrvripv2U08fWQP/5QemTHvAjr2An00Y7SkdgdzKBd6AKw8fgVt0vrbX vR5h70ZWl3VKtTFdRJr05L9Y008FQdk= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-10_02,2026-08-07_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 bulkscore=0 impostorscore=0 malwarescore=0 adultscore=0 clxscore=1015 priorityscore=1501 suspectscore=0 phishscore=0 lowpriorityscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608100091 Hi Andrea. On 8/10/26 3:11 PM, Shrikanth Hegde wrote: > > > On 8/6/26 7:13 PM, Andrea Righi wrote: >> find_new_ilb() selects the first idle housekeeping CPU without >> considering whether another thread is running on the same physical core. >> On an SMT system, the idle load balancer can therefore activate both >> siblings even when another housekeeping CPU has an entirely idle core. >> >> On most SMT systems, this is not problematic because the idle load >> balancer is a short-lived activity and the transient wakeup of a sibling >> has negligible performance impact. >> >> However, this can be particularly costly on NVIDIA Olympus cores used in >> Vera. Briefly activating an otherwise idle sibling can reduce the >> performance available to the other sibling and this effect does not >> necessarily end once the activated sibling becomes idle: after the ILB >> finishes and its CPU enters WFI, full single-thread performance is >> restored only after the sibling has remained idle for a qualification >> interval (10 Ki cycles on the tested Vera system). Repeated short >> sibling wakeups can therefore sustain the interference even with little >> actual overlap. >> >> Prevent this by preferring an idle housekeeping CPU whose entire SMT >> core is idle. Retain the first idle CPU as a fallback when no fully idle >> core is available, so NOHZ balancing continues to make forward progress. >> Once a partially busy core has been examined, skip its remaining SMT >> siblings to avoid repeating the core-idle check on wide SMT systems. >> >> Tests performed using an ad hoc GEMM benchmark running one CPU-intensive >> task per SMT core within its CPU affinity mask improved from >> approximately 6.2 TFLOP/s to 9.4 TFLOP/s. >> >> Note that this preference may wake a fully idle physical core instead of >> using an idle sibling of an active core, potentially increasing ILB >> wakeup latency or energy consumption on some architectures. It may also >> scan additional CPUs before selecting the one to run the ILB. The >> selection falls back to the first idle CPU when no fully idle SMT core >> is available. Non-SMT systems continue to select the first idle >> housekeeping CPU. >> >> Tested-by: K Prateek Nayak >> Reviewed-by: K Prateek Nayak >> Reviewed-by: Mete Durlu >> Reviewed-by: Vincent Guittot >> Reviewed-by: Shrikanth Hegde >> Signed-off-by: Andrea Righi >> --- >> Changes in v5: >>   - Collect Tested-by and Reviewed-by tags >>   - Reorder local variable declarations (Prateek Nayak) >>   - Link to v4: https://lore.kernel.org/all/20260804151324.918020-1- >> arighi@nvidia.com/ >> >> Changes in v4: >>   - Remove redundant this_cpu check (Prateek Nayak, Vincent Guittot) >>   - Link to v3: https://lore.kernel.org/all/20260731191957.3199642-1- >> arighi@nvidia.com/ >> >> Changes in v3: >>   - After finding an idle fallback, skip all siblings when a busy CPU is >>     encountered, avoiding per-CPU traversal of known-busy cores (Mete >> Durlu) >>   - Link to v2: https://lore.kernel.org/all/20260729163225.1987068-1- >> arighi@nvidia.com/ >> >> Changes in v2: >>   - Avoid repeated is_core_idle() checks on wide SMT systems by pruning >>     the remaining siblings of a partially busy core (Prateek Nayak) >>   - Link to v1: https://lore.kernel.org/r/20260728214442.1648483-1- >> arighi@nvidia.com/ >> >>   kernel/sched/fair.c | 55 ++++++++++++++++++++++++++++++++++++--------- >>   1 file changed, 44 insertions(+), 11 deletions(-) >> >> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >> index 37001c63452e5..89bec68622db9 100644 >> --- a/kernel/sched/fair.c >> +++ b/kernel/sched/fair.c >> @@ -13964,29 +13964,62 @@ static inline int on_null_domain(struct rq *rq) >>    */ >>   static inline int find_new_ilb(void) >>   { >> -    int this_cpu = smp_processor_id(); >> -    const struct cpumask *hk_mask; >> -    int ilb_cpu; >> +    int ilb_cpu, fallback = -1; >> +    struct cpumask *ilb_cpus; >> + >> +    lockdep_assert_irqs_disabled(); >> + >> +    /* >> +     * Reuse the per-CPU select_rq_mask, which is protected from >> concurrent >> +     * use on this CPU by having interrupts disabled. >> +     */ >> +    ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask); >> +    cpumask_and(ilb_cpus, nohz.idle_cpus_mask, >> +            housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); >> + >> +    for_each_cpu(ilb_cpu, ilb_cpus) { >> +        if (!idle_cpu(ilb_cpu)) { >> +            /* >> +             * Once an idle fallback exists, a busy CPU proves that >> +             * this core cannot be fully idle. Skip its siblings. >> +             */ >> +            if (sched_smt_active() && fallback >= 0) >> +                cpumask_andnot(ilb_cpus, ilb_cpus, >> cpu_smt_mask(ilb_cpu)); >> +            continue; >> +        } >> -    hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); >> +        /* >> +         * Running the idle load balancer on an idle sibling of a busy >> +         * SMT core can reduce the capacity available to its sibling. >> Prefer >> +         * a CPU whose entire core is idle, but retain the first idle >> CPU as >> +         * a fallback so idle balancing can still make progress when >> no fully >> +         * idle core exists. >> +         */ >> +        if (sched_smt_active() && !is_core_idle(ilb_cpu)) { > > nit: > > Now, when making the changes for paravirt series, > i remembered sched_smt_active() is not needed anymore. > since cpu_smt_mask in that case just points to just that cpu. I guess it is good to have since it skips the cpumask_and which might save few cycles. Also, I see peter has pulled in v4. So you can ignore this comment. Sorry for the noise. > >> +            if (fallback < 0) >> +                fallback = ilb_cpu; >> -    for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) { >> -        if (ilb_cpu == this_cpu) >> +            /* >> +             * The core is not idle, so there is no need to check >> +             * any of its other SMT siblings. >> +             */ >> +            cpumask_andnot(ilb_cpus, ilb_cpus, >> +                       cpu_smt_mask(ilb_cpu)); >>               continue; >> +        } >> -        if (idle_cpu(ilb_cpu)) >> -            return ilb_cpu; >> +        return ilb_cpu; >>       } >> -    return -1; >> +    return fallback; >>   } >>   /* >>    * Kick a CPU to do the NOHZ balancing, if it is time for it, via a >> cross-CPU >>    * SMP function call (IPI). >>    * >> - * We pick the first idle CPU in the HK_TYPE_KERNEL_NOISE >> housekeeping set >> - * (if there is one). >> + * Prefer a CPU on a fully idle core in the HK_TYPE_KERNEL_NOISE >> housekeeping >> + * set. Fall back to the first idle CPU when no fully idle core exists. >>    */ >>   static void kick_ilb(unsigned int flags) >>   { >