From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pg1-f176.google.com (mail-pg1-f176.google.com [209.85.215.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A4DAA3A7F6D for ; Tue, 28 Jul 2026 11:27:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.176 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785238081; cv=none; b=mvWy6H1bWyZjkmRpVy5luOAHvf7R7iMs+zBTgPSUk2GWegiGmJLHDdZHZuKdoQIPw5YSgVGv5bLcXi6H8DdkK0BWmnAY9Guljp1IJsRfGZ+nB5JOFWTNmu4h9BXmnf/WMRcZxUIKa912OHTNw1ryKkwaY3i5fncvDEFTlpEdfQ8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785238081; c=relaxed/simple; bh=TGC9VfFo0ZJZ2pXBQ4kmPfIwzNAivnl6LaE9/OB6VCE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:To:Cc; b=ISzXWcdLj3kckSOufZgqTWYSDxw0NwpMLerO3CCXnluWztRnxBnwg6XYzYCmjt5S7sscL47hwRbdLYzLpx5AQeomtF6+zvbzcUWCc8yVEpTb93m3A5kSaYlICFLmvqKYQ5WOdWbGf13ZmQ8HmU7pYJWFDzYWxhCH13ja/QzgKic= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=fjZi7d1w; arc=none smtp.client-ip=209.85.215.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="fjZi7d1w" Received: by mail-pg1-f176.google.com with SMTP id 41be03b00d2f7-c999f162c9aso2714284a12.3 for ; Tue, 28 Jul 2026 04:27:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785238079; x=1785842879; darn=vger.kernel.org; h=cc:to:message-id:content-transfer-encoding:content-type :mime-version:subject:date:from:from:to:cc:subject:date:message-id :reply-to:content-type; bh=CwTWUmVicSqt3aLit4DqwSibGK3y5Q0JEJNw6JJq6fQ=; b=fjZi7d1wtIUwSmF3R9KtUf5ngsdd94TPfSWoc4lQSifWnCcawGmoF/9xeePNw1fUpn TAUYCKMxAcdAeOb6sOcuDQQly+0gPxlQgWr7UotCjcwqBNP/QmmzOiX+irp0+dErlC72 63pffHHxjMBFPeHOymaRmvdOhruLJB8aCblGhsRcPm9rrDaEMmwB5yTTPS2uavInY6VA 3RbkgxkGfw0COLa6KyJW/wUqY0PjtYP/tFXBh4zh6YHB/x4LjI8LhqOrlx5ZS+Tql56+ EtlxqegjazhY6c75QwcJDGLtejueEVVMVevYSPn/ixjcljDT9aPfCmTssnzFzzh6BfTr /QIQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785238079; x=1785842879; h=cc:to:message-id:content-transfer-encoding:content-type :mime-version:subject:date:from:x-gm-gg:x-gm-message-state:from:to :cc:subject:date:message-id:reply-to:content-type; bh=CwTWUmVicSqt3aLit4DqwSibGK3y5Q0JEJNw6JJq6fQ=; b=ohF/FDiOLmOLQps6GVCv7RATu+tCmhHqqnistfrlitQZa4QupxP0nnrSawm91AQd1q l8vPaAvgIWh3/sfOf2hLb6woX9hXpso002VkdxqK8WMnjq/HtTYkq+T3TBaxDtTJOshL wZkCjxe0ugGeHjUOEqXogXh3uc46cDFQ2o1SsHIlJp0nwuIZUubT3vwOhbcYeFoQNePd 3ugFcuPWtkadsirigPlX+Jrs4R3xx/ui1+KbhJCHxCG/f8h8g1QiGTgxSoCVGKT8XyN4 2jLWkVwchsC8tSmkWEqOzgX1CgoWhj9NkbOBtOqDr0MbjAME1jRiqizDWuGJ4dmDtvY6 2c4w== X-Gm-Message-State: AOJu0YyEl/EKbPKF5KonPpPM4fw8WssZUJh/jat5tvslCfmZAnKacomr n1qgcQn7SQFiflHgNg42a5zhWnBf8sMV3b6m1i0eJTWvvKorbHyPuV6E X-Gm-Gg: AR+sD109L1c6pWs9wY88lNtp9ZMGTtpFP6ri517qp49phUiEoYnR3jUl3Cf2037hyHH 9/tBMUiDkOuCeWILisXA+AMCAK45eK5bATov5KqWcDzK92Ou5elZUOBTJKpNpJaVr13JxW9DDF6 vbCLfZHDbm5az9XKI95ESoa8su5MzwT+EErrYqhffAIf9AfVOmBd9W7xoZVfu5rQ5p+6WXeMdnz uWh8fCe53ImeC8iE6O7hFhiGhtxmq2VbWmy3zW1DVf3ybuo5Rt/wdSCPiHtRHAY4060FWp6AbMf 2Sxk+nZ2uVGyUJtPcXRhyU2hjsYfbJDPrXzbnrC+U/6TjwGtPObBKxBz8HrZHl2/iztK694T19A xIzmDq03wLPfeb5HG/iht8dVe12LLRK4nCVh89wFNGE0kGq+wix7GY7dWm6Z5CYQfM7pud8+O1Z 6aI7y1gM55sw== X-Received: by 2002:a05:6a20:939e:b0:3c3:8222:b8c9 with SMTP id adf61e73a8af0-3c8ba647551mr2882428637.67.1785238078981; Tue, 28 Jul 2026 04:27:58 -0700 (PDT) Received: from [127.0.1.1] ([23.254.208.9]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc593b57sm62564803eec.25.2026.07.28.04.27.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 28 Jul 2026 04:27:58 -0700 (PDT) From: Jing Wu Date: Tue, 28 Jul 2026 19:27:54 +0800 Subject: [PATCH] x86/aperfmperf: Refresh stale sample via IPI for busy NOHZ_FULL CPUs Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260728-bug-isolatecpu-cpufreq-v1-1-e95d34db8bcd@gmail.com> X-B4-Tracking: v=1; b=H4sIADmSaGoC/x2MQQqAIBAAvxJ7bqEkS/pKdFBbayGstCII/550m MMcZl6IFJgi9MULgW6OvPksdVmAXbSfCXnKDqISbdUJheaakeO26pPsfmHGBTpQtY2WzmglTQc 53gM5fv7xMKb0AXZyw/9oAAAA To: Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , "Peter Zijlstra (Intel)" , "Paul E. McKenney" , "Rafael J. Wysocki" Cc: linux-kernel@vger.kernel.org, Qiliang Yuan , Jian Zhang , Jing Wu X-Mailer: b4 0.13.0 An isolated CPU covered by nohz_full stops its periodic tick once it has only one runnable task, since sched_can_stop_tick() only checks scheduling-class fairness and has no notion of cpufreq reporting needs. arch_scale_freq_tick() runs only from scheduler_tick(), so cpu_samples for that CPU is never refreshed again. arch_freq_get_on_cpu() then permanently hits its staleness check and falls back to cpufreq_quick_get(), which returns whatever policy->cur was left at (typically the P-state floor). This happens even though HWP hardware keeps running the CPU at full turbo autonomously, as confirmed by turbostat and by directly reading APERF/MPERF. Reproduce on an isolated, nohz_full CPU with intel_pstate/HWP by loading it and watching scaling_cur_freq stay pinned at the floor: taskset -c $CPU stress --cpu 1 & for i in $(seq 10); do cat /sys/devices/system/cpu/cpu$CPU/cpufreq/scaling_cur_freq sleep 0.5 done turbostat --cpu $CPU --interval 1 --num_iterations 5 scaling_cur_freq stays at the floor for the whole run, while turbostat's Bzy_MHz confirms the CPU is actually at full turbo. Refresh the stale sample with one on-demand arch_scale_freq_tick() via IPI before falling back, but only when the target CPU is online and not idle. APERF/MPERF both stop advancing during idle (C1+), so a delta computed over an arbitrarily long stale window still yields a correct busy-time frequency average. Fixes: 7d84c1ebf9dd ("x86/aperfmperf: Replace aperfmperf_get_khz()") Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Co-developed-by: Jian Zhang Signed-off-by: Jian Zhang Signed-off-by: Jing Wu --- arch/x86/kernel/cpu/aperfmperf.c | 34 +++++++++++++++++++++++++++++++++- 1 file changed, 33 insertions(+), 1 deletion(-) diff --git a/arch/x86/kernel/cpu/aperfmperf.c b/arch/x86/kernel/cpu/aperfmperf.c index 7ffc78d5ebf21..e51544db9f9bb 100644 --- a/arch/x86/kernel/cpu/aperfmperf.c +++ b/arch/x86/kernel/cpu/aperfmperf.c @@ -12,6 +12,7 @@ #include #include #include +#include #include #include #include @@ -503,16 +504,41 @@ void arch_scale_freq_tick(void) */ #define MAX_SAMPLE_AGE ((unsigned long)HZ / 50) +static void aperfmperf_snapshot_cpu_ipi(void *info) +{ + arch_scale_freq_tick(); +} + +/* + * A NOHZ_FULL CPU with a single runnable task (e.g. an isolated CPU running + * a pinned PMD/busy-poll workload) stops its periodic tick, so nothing ever + * calls arch_scale_freq_tick() for it again and cpu_samples goes stale + * forever, not just for one MAX_SAMPLE_AGE window. Force one on-demand + * sample via IPI so a deliberate frequency read doesn't report the P-state + * floor from before isolation took effect. Skip idle CPUs: their frequency + * genuinely doesn't matter and there is no point poking them with an IPI. + */ +static bool aperfmperf_refresh_stale_sample(int cpu) +{ + if (!cpu_online(cpu) || idle_cpu(cpu)) + return false; + + smp_call_function_single(cpu, aperfmperf_snapshot_cpu_ipi, NULL, 1); + return true; +} + int arch_freq_get_on_cpu(int cpu) { struct aperfmperf *s = per_cpu_ptr(&cpu_samples, cpu); unsigned int seq, freq; unsigned long last; + bool refreshed = false; u64 acnt, mcnt; if (!cpu_feature_enabled(X86_FEATURE_APERFMPERF)) goto fallback; +again: do { seq = raw_read_seqcount_begin(&s->seq); last = s->last_update; @@ -524,8 +550,14 @@ int arch_freq_get_on_cpu(int cpu) * Bail on invalid count and when the last update was too long ago, * which covers idle and NOHZ full CPUs. */ - if (!mcnt || (jiffies - last) > MAX_SAMPLE_AGE) + if (!mcnt || (jiffies - last) > MAX_SAMPLE_AGE) { + if (!refreshed) { + refreshed = true; + if (aperfmperf_refresh_stale_sample(cpu)) + goto again; + } goto fallback; + } return div64_u64((cpu_khz * acnt), mcnt); --- base-commit: 502d801f0ab03e4f32f9a33d203154ce84887921 change-id: 20260728-bug-isolatecpu-cpufreq-864a5fba85b7 Best regards, -- Jing Wu