From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from SJ2PR03CU001.outbound.protection.outlook.com (mail-westusazon11012048.outbound.protection.outlook.com [52.101.43.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3585B45C705 for ; Tue, 4 Aug 2026 14:59:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.43.48 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785855558; cv=fail; b=TsqVh0vrRsraL7R6PTOvRWuKedN6KX9lF+yu7GAU3UeJTdtUxfnLY+IfncoeXkGyB1MY6UkTJxIzWCFUIvax5F7fFSQOplC1s+UEDE5N5uWNxQPfj9Tg/01XOCy8ZyMjzeXBmSlRNzfnHmBr3aywqoxt3Hlekton9xbGXmeG0mk= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785855558; c=relaxed/simple; bh=i7il+RktHzAFnDHZZXXKK5APOuIHrGMi1a+P+fbsLMw=; h=Date:From:To:Cc:Subject:Message-ID:References:Content-Type: Content-Disposition:In-Reply-To:MIME-Version; b=E/0rUH87Ij40qqIdXBFaheQjbxpdX7uVS86bscNmFfpjKMKLpdWx6hL2NfcGN0NnZQX+ZEH4njsc5BZEGzDBCxtpl9HWMpQw+6qWsBSPpems2DvN6usR5cRj9I7mKW7qkKUaEFN6SPG6zGXiBGwVAg2SooSmlVH28GNawlO+NoQ= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=el1pUovw; arc=fail smtp.client-ip=52.101.43.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="el1pUovw" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=grjbx/d4usLIKEAuULYUSclAuHBuQiR0JnLefrcRXT2EqfhR7SiX0BfUEQLSUMiZlxYN/Vm2BQRwAx9vIjfZ0eQeAud9IVb85QFosDixoAtogXlTFc6pFrfCuJA97Infb4Mm70VEKMPNPuvfCzEP1eLG/qWQUzmDxItHMo5lfC0sKiGwcYZzen7DHNrGM521B66pH3gQu6J2Mq9owJwbcwH/flajR6uJgWrJ2IWLpOxQKLOoZHj6ov/ti0LHYU7CkqMNfhLkM9ZrZE/kvrVTDOsB/4DHsUrftRMEX9CA9KKMwclk1AoKUghNKmrin49N2tf99nXDJblTleq5z73zCA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=ikMFKm5hauPqmnD+rcRk2OV6//oi2mLWovsqXQsLrNY=; b=oP58C3xP14MVca4UkLy0DbPHVJTfd6bdeOe3/lNZvq1go0Nq6E42lFq+fTjYJcHc9zodGVjQ23hD1F1HohVyqEMpD0E8dvMvjVoSnUUKUBpowamN3B0ftdF+gjbCm/CpncTyOoLmZGuPsxde795gxhTFCTgDR/jSIb8wdPfJ89D4AfXFtXCi5II2GZVLLawtP469lHKsTKZ+y9zMj4oEp8Bec7JKTswcF+mYIBqrW5YB1delTPy2w9prYrqjQShOjpk816xhvDgjRWliFQmKhZZnsJl4eBXgvBmHecXchd29Btj7G5OMz3DhpZAHtPOxCwrCKnj5QNUcx0NNiGpN/A== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=nvidia.com; dmarc=pass action=none header.from=nvidia.com; dkim=pass header.d=nvidia.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=ikMFKm5hauPqmnD+rcRk2OV6//oi2mLWovsqXQsLrNY=; b=el1pUovwAT62mjn5eE2D/iEBkbwgpriBCHOXeBRqFDdT3SciffQ0mvI6N+Y63lpD1YoG/83sCih5Wy23E6h65lSHD8Zc7BO/y8/+Lq9vw5qihs5RlDoGVJlRQ6PYZgUR2TYt7sRGS6X0/1njbypX8cqZ67udr6cizRBXRy3w9679NC0R6mQSd0xmTjRz4YtoMWDRNKJWGJg3kISz9GgxUHmlJb8mWqV4NEqUd56m+SS36AufrGdq+D5Sy0Usy/N/ENwAZyPCoWkeYHffSCz+Y8/KPN51ivhi3X/gbiZKmI++LZ1fPlnQqXFhJ6JtB2mXdQe44gpsBs3RheWRPZ4/Wg== Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=nvidia.com; Received: from DM6PR12MB4827.namprd12.prod.outlook.com (2603:10b6:5:1d6::14) by SJ2PR12MB8652.namprd12.prod.outlook.com (2603:10b6:a03:53a::18) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.270.17; Tue, 4 Aug 2026 14:59:05 +0000 Received: from DM6PR12MB4827.namprd12.prod.outlook.com ([fe80::6261:3040:864b:159c]) by DM6PR12MB4827.namprd12.prod.outlook.com ([fe80::6261:3040:864b:159c%5]) with mapi id 15.21.0270.016; Tue, 4 Aug 2026 14:59:05 +0000 Date: Tue, 4 Aug 2026 16:58:52 +0200 From: Andrea Righi To: Mete Durlu Cc: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Christian Loehle , Shrikanth Hegde , Phil Auld , linux-kernel@vger.kernel.org Subject: Re: [PATCH v3] sched/fair: Prefer fully idle cores for NOHZ balancing Message-ID: References: <20260731191957.3199642-1-arighi@nvidia.com> <533a1615-e0f8-4c7d-b212-b24760f198d3@linux.ibm.com> Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <533a1615-e0f8-4c7d-b212-b24760f198d3@linux.ibm.com> X-ClientProxiedBy: MI2PEPF00000B79.ITAP293.PROD.OUTLOOK.COM (2603:10a6:298:1::40d) To DM6PR12MB4827.namprd12.prod.outlook.com (2603:10b6:5:1d6::14) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DM6PR12MB4827:EE_|SJ2PR12MB8652:EE_ X-MS-Office365-Filtering-Correlation-Id: 4d86b416-146c-409d-91bd-08def238ed66 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|7416014|1800799024|376014|23010399003|366016|6133799003|22082099003|18002099003|4143699003|11063799006|56012099006|10067099003; X-Microsoft-Antispam-Message-Info: nSed5/T/p6kr0xHvMsMBfqq53K/IqQOjeE90iri6fJs3+sBx6c2DYMiBmDDBcX51Vjovy0gYgw/+o7WH3ceMqbwhY+ya4TRA64NYAAj1Yr/eBQTUuQDl2fF3T04lP9r4/OfYF8h8/XyTdt0MUdq33UKJmnV8pS1yvZiI7jYTWuYWw1BIlaXKFUc2jTRCGlCtl8vw1zmWSvds5Yg1PDChOEYoLSxLWCKZPF8Lxbv/Ja4A0bgFSKkS6Xdi1FODFaaTDl+K4Rfceem1Qka461EGLTnAKBbArIQ8YdnSW5TTzO2LfLLkDrJbY8BhvdbWDo4pp199RJLVDZt1cZKsKgrRnazjCcHkBwehUdlFhdlNqXxcN3Gk4u0xnrTU3iJm+oUKJ4ku2e9wWUfRuRK5PO0q2qnqPS3dOx0olj7373ymClp2FhAHn3sdW1vZ+bfBLmxsWSV0ulNMmAkcdsJKMROwLNVO+2PKbmM0Mob600fAqADAsoyzczD9o/PdGQAlZrALY2OKXre4rnhu3CyF5IvGCavizP+bFLdkUVFS57VQrFknYEHZ3CMLgDzsCdEgbEFCuf5QUuIbNkEFhOLb3T4A9LO3+nNvp/torg6XbruEwD4BsRUwx5PT16F31vCpR+5SRHjhv7GYEwvIVgQP+8EkqMmIkX5DF8tBGgwWvSQUIDk= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:DM6PR12MB4827.namprd12.prod.outlook.com;PTR:;CAT:NONE;SFS:(13230040)(7416014)(1800799024)(376014)(23010399003)(366016)(6133799003)(22082099003)(18002099003)(4143699003)(11063799006)(56012099006)(10067099003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?ikF68FasPIMgWnGQlmvh3U3swt3YfkctcUBCh/ROHXAH0shnzR4qWTB3j64W?= =?us-ascii?Q?QxzyKm2ZvPncdTcBsBAjo/E5FXCRwwUVsbCoFEnYZmlZgpfvmK2mt7UZZCaD?= =?us-ascii?Q?T520+/LZQhE3wxF7a0969b/Ursq2z4VtDnvplr2WYaGnUThy0GYY1Q3TfO8Y?= =?us-ascii?Q?7dBhRijUhne9FSCQ7K9A8fy2ec/q73IBT6EUaCr0ai+phos3G9qqlSDDYYzv?= =?us-ascii?Q?6oMCWgfXd+Zfns1+C3UN3caaMA57h+So+SSQKdRSlgek0bbWGg/sNY/lgBTH?= =?us-ascii?Q?+YWMrtpi/faP1MURSIneUKFZwBXg4cyz0HoQqhQYEPcA9RcTxucrsfD83os6?= =?us-ascii?Q?0L1VrVzM8IU/DqW7CMe3ZFBhhEimu3x3fKeSgmiJPt8XSu9xtQ9UFLo80ocS?= =?us-ascii?Q?Splz8EDpvBc9qZ/xa2AfWenAC0C6WSF9QPOJ+DQITROWEcBUTyEQiAJ+RwL4?= =?us-ascii?Q?AAOnMRMZNwsimip75xpJlEBhI0SxqeOMXwoiVqyHuTwt7qO2OfzZJVMtCREW?= =?us-ascii?Q?2EjmyOgAgTaLmJx1AjS8Z91Ao1TvIUrBm89iJktYzVtS8kxbI7/H4yMfDW9q?= =?us-ascii?Q?NkXPbClrYlC2BdDriuUq9tOYRRjWgmZRruDD0rHnYK0iuCEXX7LW9/ByvbNf?= =?us-ascii?Q?+9EEyidSvaSWXosZVsA9zbBSFcgN0bL5KZmMJzgh59d24kht2QtGu7p5jcyu?= =?us-ascii?Q?JE9WWgi9uW9OSzHwK8ryBQhPce3vIReCEd4KQFMffiLljJqa2H7ofcFzzrq7?= =?us-ascii?Q?rw5Ov/xD9/w+XYWQ+JZ0eHBcrgt6a3dxbf3t4qIiO2oH5lTSNanD564Ohxec?= =?us-ascii?Q?7MZAljqjcvKVjqAz2bpQsqDgW9uIxoqfoLsmf4doA+wHkWzkWv/HzbARimWC?= =?us-ascii?Q?dvulGCJ3h80wjD68GGXHswnMpZ6Hn++J2zSuqiMD3n/drfxfPrHjLaR/pAS2?= =?us-ascii?Q?nMX1AsBWVoyZx8GKevbFkebUeSB0yXo6LmkwHE7YbA73bRZecJMpNgRkg29O?= =?us-ascii?Q?Uqs6Ni9uOwIBwLoY6/Gt4bXkUlplxAA1WZRm5lu+y84yeBUYT87mv+3iiXL0?= =?us-ascii?Q?QHe47Z8yC9MpIS/ie7Jkh/CVVDqRk754Zn/O4JiKiH2RGAAGEYFXzPHl+K1n?= =?us-ascii?Q?dftCyOYilz0QCPhasy5AhFX293+GwGtKwAX61Ktj+j6tqnxYTiULDyVAKod1?= =?us-ascii?Q?eS9qRBRt1uE51YAipRDklQrZmIu5rel0plmgP/DDF7q1diHvwL0+CFM3YpYt?= =?us-ascii?Q?N++zCObKL8fGp19PMlivFQi+Z9KzcpIbOENU2XC/jn/t3thHsWYjKW9wxF2w?= =?us-ascii?Q?fbvtyESf8oBPdjmDjEkpSaAhW+uhcmG/1m3DCzmfQUIK3SIMNfPmnqCm+nKR?= =?us-ascii?Q?AVe4XBa7xgM1VxpxbegiWSPfChLL1tMFNC6iRh9bQjfFQ08KcooieTVIn9KT?= =?us-ascii?Q?ExZUx3SwXTuP0Uom2zDgT93Bk7+RxYFwK8PS2Fe+/30leoC5NMANoSopFgti?= =?us-ascii?Q?2Q7ADGmFlj4k/rl0+pxrDMmbFpb77FGuRaNL6YWITJMXcRaST4qfcg1YxAFp?= =?us-ascii?Q?BXVrhzaSfKhYp/Rea/xwlIK31O+E0vpVAq7QTgFYTq+EQmSCQlm8MP7yuI++?= =?us-ascii?Q?Q/CT+WavI00QkXpTqNQjA7vWHJelk0TaOYwFlhyz01ZBJqlOE146cWccBHFO?= =?us-ascii?Q?uE0idN2MAS5yxdcB0/rB/pEFKPd5hd8A8wilRM5zRJTS5O0/u3HJoIvX5zrL?= =?us-ascii?Q?92dFnJUOYA=3D=3D?= X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-Network-Message-Id: 4d86b416-146c-409d-91bd-08def238ed66 X-MS-Exchange-CrossTenant-AuthSource: DM6PR12MB4827.namprd12.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 04 Aug 2026 14:59:05.2177 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: 6j/9TKwaI7R0h4SEsOHAWFkHLZ+qRhMDCDkfI2Kpye8HK2vX8aRPUMXDzxzu97Kt5pmqe6Z4SQdJQkvaNPxDLQ== X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ2PR12MB8652 Hi Mete, On Tue, Aug 04, 2026 at 02:36:48PM +0200, Mete Durlu wrote: > Hi, > > > find_new_ilb() selects the first idle housekeeping CPU without > > considering whether another thread is running on the same physical core. > > On an SMT system, the idle load balancer can therefore activate both > > siblings even when another housekeeping CPU has an entirely idle core. > > > > On most SMT systems, this is not problematic because the idle load > > balancer is a short-lived activity and the transient wakeup of a sibling > > has negligible performance impact. > > > > However, this can be particularly costly on NVIDIA Olympus cores used in > > Vera. Briefly activating an otherwise idle sibling can reduce the > > performance available to the other sibling and this effect does not > > necessarily end once the activated sibling becomes idle: after the ILB > > finishes and its CPU enters WFI, full single-thread performance is > > restored only after the sibling has remained idle for a qualification > > interval (10 Ki cycles on the tested Vera system). Repeated short > > sibling wakeups can therefore sustain the interference even with little > > actual overlap. > > > > Prevent this by preferring an idle housekeeping CPU whose entire SMT > > core is idle. Retain the first idle CPU as a fallback when no fully idle > > core is available, so NOHZ balancing continues to make forward progress. > > Once a partially busy core has been examined, skip its remaining SMT > > siblings to avoid repeating the core-idle check on wide SMT systems. > > > > Tests performed using an ad hoc GEMM benchmark running one CPU-intensive > > task per SMT core within its CPU affinity mask improved from > > approximately 6.2 TFLOP/s to 9.4 TFLOP/s. > > Although what you describe above with siblings suffering interference > does not really fit to s390, I'd like to hear more about what sort > of GEMM (general matrix multiplication) tests you did. > > I tested this patch with a couple of different tools > - perf bench sched pipe > - hackbench > - uperf > - cyclictest > - stress-ng (3d-matrix and cyclic) > > Didn't come across any meaningful difference in any of them on multiple > runs each. So I was curious about the exact sort of benchmark you > mention here. Thanks for testing on s390! The original benchmark I used is based on an internal NVPL container that I can't share publicly. However, I tried with the public OpenBLAS SGEMM benchmark and I can see exactly the same behavior and effects: https://github.com/OpenMathLib/OpenBLAS My test machine has the following topology (arm64): CPUs: 352 Sockets: 2 Cores per socket: 88 Threads per core: 2 NUMA node 0 CPUs: 0-87,176-263 NUMA node 1 CPUs: 88-175,264-351 The SMT sibling pairs on NUMA node 0 are (0,176), (1,177), ..., (87,263). I only used NUMA node 0, to prevent adding potential NUMA side effects. I used OpenBLAS v0.3.33, built using GCC 13.3.0 with the ARMv8 SVE kernels and OpenMP threading: $ make -j176 \ TARGET=ARMV8SVE \ USE_OPENMP=1 \ NUM_THREADS=176 \ NOFORTRAN=1 $ make -C benchmark sgemm.goto \ TARGET=ARMV8SVE \ USE_OPENMP=1 \ NUM_THREADS=176 \ NOFORTRAN=1 The unpatched kernel was tip/master, while the patched kernel used the same base with only this change applied. A 10-loop runs produced: unpatched: 5.088 TFLOP/s patched: 7.143 TFLOP/s Two additional 5-loop runs produced: unpatched: 5.338, 5.246 TFLOP/s patched: 7.082, 7.144 TFLOP/s The median across all the measurements increased from 5.246 TFLOP/s to 7.143 TFLOP/s, an improvement of approximately 36.2%. The absolute throughput is lower than my initial NVPL benchmark, as expected from the different GEMM implementations, but the relative behavior seems to be consistent. > > One minor nit for the diff below; > > > > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c > > index 37001c63452e5..574b6b3ee922a 100644 > > --- a/kernel/sched/fair.c > > +++ b/kernel/sched/fair.c > > @@ -13965,28 +13965,66 @@ static inline int on_null_domain(struct rq *rq) > > static inline int find_new_ilb(void) > > { > > int this_cpu = smp_processor_id(); > > - const struct cpumask *hk_mask; > > - int ilb_cpu; > > + struct cpumask *ilb_cpus; > > + int ilb_cpu, fallback = -1; > > + > > + lockdep_assert_irqs_disabled(); > > - hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); > > + /* > > + * Reuse the per-CPU select_rq_mask, which is protected from concurrent > > + * use on this CPU by having interrupts disabled. > > + */ > > + ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask); > > + cpumask_and(ilb_cpus, nohz.idle_cpus_mask, > > + housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); > > - for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) { > > + for_each_cpu(ilb_cpu, ilb_cpus) { > > if (ilb_cpu == this_cpu) > > continue; > > - if (idle_cpu(ilb_cpu)) > > - return ilb_cpu; > > + if (!idle_cpu(ilb_cpu)) { > > + /* > > + * Once an idle fallback exists, a busy CPU proves that > > + * this core cannot be fully idle. Skip its siblings. > > + */ > > + if (sched_smt_active() && fallback >= 0) > > + cpumask_andnot(ilb_cpus, ilb_cpus, > > + cpu_smt_mask(ilb_cpu)); > > nit; > With line break this if block is now taking multiple lines and deserves > its own curly braces. The cpumask_andnot() invocation fits within the line-length limit, so I'll move it on the same line. > > With or without the nit, feel free to add my r-b to v4, I doubt removal > of the "this_cpu" check will change anything as it is a dud. > > Reviewed By: Mete Durlu > Thanks, -Andrea