From: srinivas pandruvada <srinivas.pandruvada@linux.intel.com>
To: "Rafael J. Wysocki" <rafael@kernel.org>,
Jing Wu <realwujing@gmail.com>,
linux-pm@vger.kernel.org
Cc: Viresh Kumar <viresh.kumar@linaro.org>,
Doug Smythies <dsmythies@telus.net>,
"Rafael J. Wysocki" <rafael.j.wysocki@intel.com>,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v1] cpufreq: intel_pstate: Adjust policy->cur in active mode to policy
Date: Fri, 31 Jul 2026 10:20:03 -0700 [thread overview]
Message-ID: <306864970e3cbb3650617427aa83eb8277175377.camel@linux.intel.com> (raw)
In-Reply-To: <5144014.31r3eYUQgx@rafael.j.wysocki>
On Wed, 2026-07-29 at 20:40 +0200, Rafael J. Wysocki wrote:
> On Wednesday, July 29, 2026 10:59:24 AM CEST Jing Wu wrote:
> > When cpu->policy is CPUFREQ_POLICY_PERFORMANCE,
> > intel_pstate_set_policy()
> > pins the CPU to a fixed pstate (max(min_pstate, max_perf_ratio))
> > and
> > programs it directly, precisely because, per the existing comment,
> > "NOHZ_FULL CPUs need this as the governor callback may not be
> > invoked
> > on them". Two lines later it still unconditionally clobbers policy-
> > >cur
> > down to policy->min, discarding the pinned value it just computed
> > and
> > applied.
> >
> > arch_freq_get_on_cpu() falls back to cpufreq_quick_get(), i.e.
> > policy->cur, whenever its APERF/MPERF sample goes stale. A CPU
> > whose
> > tick keeps running refreshes that sample constantly and rarely hits
> > the fallback, but an isolated CPU covered by nohz_full with a
> > single
> > runnable task never gets another tick, so it permanently reports
> > the
> > floor through this fallback - even though it is genuinely pinned
> > to,
> > and running at, the frequency computed just above.
> >
> > Set policy->cur to the exact pinned frequency (pstate * scaling) in
> > the CPUFREQ_POLICY_PERFORMANCE branch instead, and only fall back
> > to
> > policy->min for the general case, where the frequency genuinely
> > isn't
> > known without a fresh sample.
> >
> > Fixes: d51847acb018 ("cpufreq: intel_pstate: set stale CPU
> > frequency to minimum")
> > Co-developed-by: Qiliang Yuan <yuanql9@chinatelecom.cn>
> > Signed-off-by: Qiliang Yuan <yuanql9@chinatelecom.cn>
> > Signed-off-by: Jing Wu <realwujing@gmail.com>
> > ---
> > drivers/cpufreq/intel_pstate.c | 20 +++++++++++++++-----
> > 1 file changed, 15 insertions(+), 5 deletions(-)
> >
> > diff --git a/drivers/cpufreq/intel_pstate.c
> > b/drivers/cpufreq/intel_pstate.c
> > index 5a0eeb84d3821..b2c60c4931dcd 100644
> > --- a/drivers/cpufreq/intel_pstate.c
> > +++ b/drivers/cpufreq/intel_pstate.c
> > @@ -2908,8 +2908,23 @@ static int intel_pstate_set_policy(struct
> > cpufreq_policy *policy)
> > */
> > intel_pstate_clear_update_util_hook(policy->cpu);
> > intel_pstate_set_pstate(cpu, pstate);
> > +
> > + /*
> > + * Report the exact pinned frequency instead of
> > the floor:
> > + * the CPU is pinned to pstate here and nothing
> > else changes
> > + * it, unlike the general case below.
> > + */
> > + policy->cur = pstate * cpu->pstate.scaling;
> > } else {
> > intel_pstate_set_update_util_hook(policy->cpu);
> > +
> > + /*
> > + * Keep policy->cur within limits here: outside of
> > the pinned
> > + * CPUFREQ_POLICY_PERFORMANCE case above, it is
> > never updated
> > + * by the intel_pstate driver, but it is used as a
> > stale
> > + * frequency value.
> > + */
> > + policy->cur = policy->min;
> > }
> >
> > if (hwp_active) {
> > @@ -2922,11 +2937,6 @@ static int intel_pstate_set_policy(struct
> > cpufreq_policy *policy)
> > intel_pstate_clear_update_util_hook(policy
> > ->cpu);
> > intel_pstate_hwp_set(policy->cpu);
> > }
> > - /*
> > - * policy->cur is never updated with the intel_pstate
> > driver, but it
> > - * is used as a stale frequency value. So, keep it within
> > limits.
> > - */
> > - policy->cur = policy->min;
> >
> > mutex_unlock(&intel_pstate_limits_lock);
> >
> >
> > ---
>
> Good idea overall, but it takes a bit more to do this. In
> particular, the HWP
> case needs some more care.
>
> Also, I don't think that this really is a fix. The code works as
> intended,
> although what it does is sometimes confusing.
>
> Below is my version of this change (on top of linux-next), please let
> me know
> if it works for you.
>
> Thanks!
>
> ---
> From: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
> Subject: [PATCH v1] cpufreq: intel_pstate: Adjust policy->cur in
> active mode to policy
>
> Since arch_freq_get_on_cpu() on x86 falls back to
> cpufreq_quick_get(),
> which effectively causes policy->cur to be returned when intel_pstate
> is used, adjust intel_pstate_set_policy() to set policy->cur to
> reflect
> the P-state that is actually going to be requested in the
> "performance"
> policy case instead of setting it to policy->min (which is confusing
> because it causes scaling_cur_freq to show the minimum frequency
> while
> the CPU is likely running at the maximum one).
>
> For this purpose, rearrange intel_pstate_set_policy() to handle the
> HWP
> case separately, to avoid calling intel_pstate_set_pstate()
> pointlessly
> with HWP enabled, and use the observation that with HWP enabled in
> the
> active mode, the utilization update hook is only needed when HWP
> boost
> is used and the policy is not "performance".
>
> Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Acked-by: Srinivas Pandruvada <srinivas.pandruvada@linux.intel.com>
> ---
> drivers/cpufreq/intel_pstate.c | 38 +++++++++++++++++++++++-------
> --------
> 1 file changed, 23 insertions(+), 15 deletions(-)
>
> --- a/drivers/cpufreq/intel_pstate.c
> +++ b/drivers/cpufreq/intel_pstate.c
> @@ -2870,6 +2870,7 @@ static void intel_pstate_set_pstate(stru
>
> static int intel_pstate_set_policy(struct cpufreq_policy *policy)
> {
> + unsigned int freq = policy->min;
> struct cpudata *cpu;
>
> if (!policy->cpuinfo.max_freq)
> @@ -2885,7 +2886,23 @@ static int intel_pstate_set_policy(struc
>
> intel_pstate_update_perf_limits(cpu, policy->min, policy-
> >max);
>
> - if (cpu->policy == CPUFREQ_POLICY_PERFORMANCE) {
> + if (hwp_active) {
> + /*
> + * The active mode only requires an update util hook
> if HWP
> + * boost is used and the policy is not
> "performance".
> + */
> + if (hwp_boost && cpu->policy !=
> CPUFREQ_POLICY_PERFORMANCE) {
> + intel_pstate_set_update_util_hook(policy-
> >cpu);
> + } else {
> + intel_pstate_clear_update_util_hook(policy-
> >cpu);
> + if (cpu->policy ==
> CPUFREQ_POLICY_PERFORMANCE) {
> + freq = cpu->max_perf_ratio * cpu-
> >pstate.scaling;
> + if (cpu->pstate.scaling != cpu-
> >pstate.perf_ctl_scaling)
> + freq = rounddown(freq, cpu-
> >pstate.perf_ctl_scaling);
> + }
> + }
> + intel_pstate_hwp_set(policy->cpu);
> + } else if (cpu->policy == CPUFREQ_POLICY_PERFORMANCE) {
> int pstate = max(cpu->pstate.min_pstate, cpu-
> >max_perf_ratio);
>
> /*
> @@ -2894,25 +2910,17 @@ static int intel_pstate_set_policy(struc
> */
> intel_pstate_clear_update_util_hook(policy->cpu);
> intel_pstate_set_pstate(cpu, pstate);
> + freq = pstate * cpu->pstate.scaling;
> } else {
> intel_pstate_set_update_util_hook(policy->cpu);
> }
> -
> - if (hwp_active) {
> - /*
> - * When hwp_boost was active before and dynamically
> it
> - * was turned off, in that case we need to clear the
> - * update util hook.
> - */
> - if (!hwp_boost)
> - intel_pstate_clear_update_util_hook(policy-
> >cpu);
> - intel_pstate_hwp_set(policy->cpu);
> - }
> /*
> - * policy->cur is never updated with the intel_pstate
> driver, but it
> - * is used as a stale frequency value. So, keep it within
> limits.
> + * policy->cur is never updated in the intel_pstate driver,
> but it is
> + * used as a stale frequency value, so set it to reflect the
> actual
> + * requested P-state in the "performance" policy case and to
> the min
> + * otherwise.
> */
> - policy->cur = policy->min;
> + policy->cur = freq;
>
> mutex_unlock(&intel_pstate_limits_lock);
>
>
>
>
next prev parent reply other threads:[~2026-07-31 17:20 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-29 8:59 [PATCH] cpufreq: intel_pstate: Sync policy->cur to the pinned pstate Jing Wu
2026-07-29 18:40 ` [PATCH v1] cpufreq: intel_pstate: Adjust policy->cur in active mode to policy Rafael J. Wysocki
2026-07-29 23:29 ` Doug Smythies
2026-07-30 15:01 ` Doug Smythies
2026-07-31 17:05 ` Rafael J. Wysocki (Intel)
2026-07-31 17:20 ` srinivas pandruvada [this message]
2026-08-03 18:15 ` [PATCH v1] cpufreq: intel_pstate: Avoid using DESIRED_PERF when DEC is enabled Rafael J. Wysocki
2026-08-04 17:58 ` srinivas pandruvada
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=306864970e3cbb3650617427aa83eb8277175377.camel@linux.intel.com \
--to=srinivas.pandruvada@linux.intel.com \
--cc=dsmythies@telus.net \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pm@vger.kernel.org \
--cc=rafael.j.wysocki@intel.com \
--cc=rafael@kernel.org \
--cc=realwujing@gmail.com \
--cc=viresh.kumar@linaro.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.