Linux Power Management development
 help / color / mirror / Atom feed
* [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint
@ 2026-07-21  9:23 Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 1/5] pm: qos: add notifiers for CPU latency and wakeup latency QoS Yaxiong Tian
                   ` (4 more replies)
  0 siblings, 5 replies; 8+ messages in thread
From: Yaxiong Tian @ 2026-07-21  9:23 UTC (permalink / raw)
  To: rafael, daniel.lezcano, christian.loehle, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest, Yaxiong Tian

cpuidle_governor_latency_req() is evaluated on every idle-state
selection.  It aggregates the per-CPU resume latency with the global
CPU latency and wakeup latency QoS limits by repeatedly calling
get_cpu_device() and pm_qos_read_value() cpu_latency_qos_limit()
cpu_wakeup_latency_qos_limit().

Use ftrace's function_graph, we can see:
parent:          do_idle
parent_total_ns: 36671010505
parent_count:    5994

SYMBOL                                                      TIME_NS    %ROOT  %PARENT      COUNT
--------------------------------------------------------------------------------------------------
do_idle                                                 36671010505  100.00%  100.00%       5994
  cpuidle_idle_call                                     35566528731   96.99%   96.99%       8570
    call_cpuidle                                        35476606844   96.74%   99.75%       8561
      cpuidle_enter                                     35472932468   96.73%   99.99%       8526
    cpuidle_select                                         52097031    0.14%    0.15%       8580
      menu_select                                          49555181    0.14%   95.12%       8580
        tick_nohz_get_sleep_length                         28843887    0.08%   58.21%       8570
        cpuidle_governor_latency_req                        9567488    0.03%   19.31%       8580
        tick_nohz_tick_stopped                              2057031    0.01%    4.15%      15695
    cpuidle_reflect                                        11427201    0.03%    0.03%       8561
      menu_reflect                                          6579506    0.02%   57.58%       8526
        tick_nohz_idle_got_tick                             2231559    0.01%   33.92%       8526
      __sysvec_apic_timer_interrupt                          105520    0.00%    0.92%          3
    tick_nohz_idle_stop_tick                                8279641    0.02%    0.02%       1475
    ---- skip

The majority of the time spent in cpuidle_enter for CPUs entering
idle state has already been charged to the idle path. Among the
remaining contributors, cpuidle_governor_latency_req() accounts 
for a non-negligible portion of the overall latency.

Under the menu governor this shows up hot: ftrace data shows,
 cpuidle_governor_latency_req() accounts for about 19.9% of
menu_select() time (~1.9 us/call).  After caching the aggregated
value per CPU and invalidating via QoS notifiers, that share drops to
about 4.2% (~0.3 us/call), roughly a 6x reduction on this path.

The ftrace data before and after the optimization is shown below:
1) original
parent:          menu_select
parent_total_ns: 160492937
parent_count:    16718

SYMBOL                                                      TIME_NS    %ROOT  %PARENT      COUNT
--------------------------------------------------------------------------------------------------
menu_select                                               160492937  100.00%  100.00%      16718
  tick_nohz_get_sleep_length                              100262940   62.47%   62.47%      16698
    tick_nohz_next_event                                   67891299   42.30%   67.71%      16689
      rcu_needs_cpu                                         2825649    1.76%    4.16%      16689
      timekeeping_max_deferment                             2380377    1.48%    3.51%      15296
    hrtimer_next_event_without                             17865162   11.13%   17.82%      15296
      hrtimer_bases_next_event_without                      2707631    1.69%   15.16%      15296
      _raw_spin_lock_irqsave                                2461132    1.53%   13.78%      15296
        native_queued_spin_lock_slowpath                        177    0.00%    0.01%          1
      _raw_spin_unlock_irqrestore                           2364647    1.47%   13.24%      15296
    can_stop_idle_tick                                      2906072    1.81%    2.90%      16698
  cpuidle_governor_latency_req                             31988150   19.93%   19.93%      16718
    get_cpu_device                                          4122502    2.57%   12.89%      16718
    pm_qos_read_value                                       3427318    2.14%   10.71%      16718
    cpu_latency_qos_limit                                   3005804    1.87%    9.40%      16718
    cpu_wakeup_latency_qos_limit                            3005475    1.87%    9.40%      16718
  tick_nohz_tick_stopped                                    4551981    2.84%    2.84%      29496

%ROOT = share of menu_select; %PARENT = share of immediate caller (inclusive)

2) post-optimized
parent:          menu_select
parent_total_ns: 55428604
parent_count:    7626

SYMBOL                                                      TIME_NS    %ROOT  %PARENT      COUNT
--------------------------------------------------------------------------------------------------
menu_select                                                55428604  100.00%  100.00%       7626
  tick_nohz_get_sleep_length                               37464607   67.59%   67.59%       7544
    tick_nohz_next_event                                   24076913   43.44%   64.27%       7522
      get_next_timer_interrupt                             16332381   29.47%   67.83%       5854
      rcu_needs_cpu                                         1489633    2.69%    6.19%       7522
      timekeeping_max_deferment                              870851    1.57%    3.62%       5586
    hrtimer_next_event_without                              6140153   11.08%   16.39%       5586
      hrtimer_bases_next_event_without                       979000    1.77%   15.94%       5586
      _raw_spin_lock_irqsave                                 808786    1.46%   13.17%       5586
      _raw_spin_unlock_irqrestore                            785518    1.42%   12.79%       5586
    can_stop_idle_tick                                      1583137    2.86%    4.23%       7544
  cpuidle_governor_latency_req                              2321119    4.19%    4.19%       7626
  tick_nohz_tick_stopped                                    1863015    3.36%    3.36%      12751

%ROOT = share of menu_select; %PARENT = share of immediate caller (inclusive)

A self-test case is introduced in patch 5 to validate that the existing
functionality remains intact. This can be handled as a standalone task.

This series:
  - adds notifier APIs for the global CPU/wakeup latency QoS
  - lets cpuidle subscribe and maintain a per-CPU generation
  - invalidates only the affected CPU on per-CPU resume latency
    changes
  - caches the aggregated constraint in cpuidle_governor_latency_req()
  - adds a kselftest covering the three QoS input paths

Yaxiong Tian (5):
  pm: qos: add notifiers for CPU latency and wakeup latency QoS
  cpuidle: subscribe to global latency QoS notifiers
  cpuidle: invalidate latency gen on per-CPU resume QoS changes
  cpuidle: cache aggregated governor latency QoS constraint
  selftests/cpuidle: add latency_req QoS idle-state selection test

 drivers/cpuidle/cpuidle.c                     |  15 +-
 drivers/cpuidle/cpuidle.h                     |   2 +
 drivers/cpuidle/governor.c                    | 141 +++++++-
 include/linux/pm_qos.h                        |  23 ++
 kernel/power/qos.c                            |  50 +++
 tools/testing/selftests/Makefile              |   1 +
 tools/testing/selftests/cpuidle/Makefile      |   6 +
 tools/testing/selftests/cpuidle/config        |   3 +
 .../cpuidle/cpuidle_latency_req_qos.py        | 331 ++++++++++++++++++
 tools/testing/selftests/cpuidle/settings      |   2 +
 10 files changed, 566 insertions(+), 8 deletions(-)
 create mode 100644 tools/testing/selftests/cpuidle/Makefile
 create mode 100644 tools/testing/selftests/cpuidle/config
 create mode 100755 tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
 create mode 100644 tools/testing/selftests/cpuidle/settings

-- 
2.43.0

^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH 1/5] pm: qos: add notifiers for CPU latency and wakeup latency QoS
  2026-07-21  9:23 [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint Yaxiong Tian
@ 2026-07-21  9:25 ` Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 2/5] cpuidle: subscribe to global latency QoS notifiers Yaxiong Tian
                   ` (3 subsequent siblings)
  4 siblings, 0 replies; 8+ messages in thread
From: Yaxiong Tian @ 2026-07-21  9:25 UTC (permalink / raw)
  To: rafael, daniel.lezcano, christian.loehle, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest, Yaxiong Tian

Expose notifier registration for the system-wide CPU latency and
wakeup latency QoS constraints so listeners can react when the
aggregate target value changes.  pm_qos_update_target() already
invokes constraint notifiers when the effective value changes.

Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
---
 include/linux/pm_qos.h | 23 +++++++++++++++++++
 kernel/power/qos.c     | 50 ++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 73 insertions(+)

diff --git a/include/linux/pm_qos.h b/include/linux/pm_qos.h
index 6cea4455f867..3759b453c6d6 100644
--- a/include/linux/pm_qos.h
+++ b/include/linux/pm_qos.h
@@ -149,6 +149,8 @@ bool cpu_latency_qos_request_active(struct pm_qos_request *req);
 void cpu_latency_qos_add_request(struct pm_qos_request *req, s32 value);
 void cpu_latency_qos_update_request(struct pm_qos_request *req, s32 new_value);
 void cpu_latency_qos_remove_request(struct pm_qos_request *req);
+int cpu_latency_qos_add_notifier(struct notifier_block *notifier);
+int cpu_latency_qos_remove_notifier(struct notifier_block *notifier);
 #else
 static inline s32 cpu_latency_qos_limit(void) { return INT_MAX; }
 static inline bool cpu_latency_qos_request_active(struct pm_qos_request *req)
@@ -160,15 +162,36 @@ static inline void cpu_latency_qos_add_request(struct pm_qos_request *req,
 static inline void cpu_latency_qos_update_request(struct pm_qos_request *req,
 						  s32 new_value) {}
 static inline void cpu_latency_qos_remove_request(struct pm_qos_request *req) {}
+static inline int cpu_latency_qos_add_notifier(struct notifier_block *notifier)
+{
+	return 0;
+}
+static inline int
+cpu_latency_qos_remove_notifier(struct notifier_block *notifier)
+{
+	return 0;
+}
 #endif
 
 #ifdef CONFIG_PM_QOS_CPU_SYSTEM_WAKEUP
 s32 cpu_wakeup_latency_qos_limit(void);
+int cpu_wakeup_latency_qos_add_notifier(struct notifier_block *notifier);
+int cpu_wakeup_latency_qos_remove_notifier(struct notifier_block *notifier);
 #else
 static inline s32 cpu_wakeup_latency_qos_limit(void)
 {
 	return PM_QOS_RESUME_LATENCY_NO_CONSTRAINT;
 }
+static inline int
+cpu_wakeup_latency_qos_add_notifier(struct notifier_block *notifier)
+{
+	return 0;
+}
+static inline int
+cpu_wakeup_latency_qos_remove_notifier(struct notifier_block *notifier)
+{
+	return 0;
+}
 #endif
 
 #ifdef CONFIG_PM
diff --git a/kernel/power/qos.c b/kernel/power/qos.c
index 1944dbeb0d4c..392ab3dbb923 100644
--- a/kernel/power/qos.c
+++ b/kernel/power/qos.c
@@ -212,12 +212,15 @@ bool pm_qos_update_flags(struct pm_qos_flags *pqf,
 #ifdef CONFIG_CPU_IDLE
 /* Definitions related to the CPU latency QoS. */
 
+static BLOCKING_NOTIFIER_HEAD(cpu_latency_qos_notifiers);
+
 static struct pm_qos_constraints cpu_latency_constraints = {
 	.list = PLIST_HEAD_INIT(cpu_latency_constraints.list),
 	.target_value = PM_QOS_CPU_LATENCY_DEFAULT_VALUE,
 	.default_value = PM_QOS_CPU_LATENCY_DEFAULT_VALUE,
 	.no_constraint_value = PM_QOS_CPU_LATENCY_DEFAULT_VALUE,
 	.type = PM_QOS_MIN,
+	.notifiers = &cpu_latency_qos_notifiers,
 };
 
 static inline bool cpu_latency_qos_value_invalid(s32 value)
@@ -335,6 +338,28 @@ void cpu_latency_qos_remove_request(struct pm_qos_request *req)
 }
 EXPORT_SYMBOL_GPL(cpu_latency_qos_remove_request);
 
+/**
+ * cpu_latency_qos_add_notifier - Add CPU latency QoS change notifier.
+ * @notifier: Notifier block managed by the caller.
+ */
+int cpu_latency_qos_add_notifier(struct notifier_block *notifier)
+{
+	return blocking_notifier_chain_register(&cpu_latency_qos_notifiers,
+						notifier);
+}
+EXPORT_SYMBOL_GPL(cpu_latency_qos_add_notifier);
+
+/**
+ * cpu_latency_qos_remove_notifier - Remove CPU latency QoS change notifier.
+ * @notifier: Notifier block previously registered.
+ */
+int cpu_latency_qos_remove_notifier(struct notifier_block *notifier)
+{
+	return blocking_notifier_chain_unregister(&cpu_latency_qos_notifiers,
+						  notifier);
+}
+EXPORT_SYMBOL_GPL(cpu_latency_qos_remove_notifier);
+
 /* User space interface to the CPU latency QoS via misc device. */
 
 static int cpu_latency_qos_open(struct inode *inode, struct file *filp)
@@ -417,12 +442,15 @@ static struct miscdevice cpu_latency_qos_miscdev = {
 
 #ifdef CONFIG_PM_QOS_CPU_SYSTEM_WAKEUP
 /* The CPU system wakeup latency QoS. */
+static BLOCKING_NOTIFIER_HEAD(cpu_wakeup_latency_qos_notifiers);
+
 static struct pm_qos_constraints cpu_wakeup_latency_constraints = {
 	.list = PLIST_HEAD_INIT(cpu_wakeup_latency_constraints.list),
 	.target_value = PM_QOS_RESUME_LATENCY_NO_CONSTRAINT,
 	.default_value = PM_QOS_RESUME_LATENCY_NO_CONSTRAINT,
 	.no_constraint_value = PM_QOS_RESUME_LATENCY_NO_CONSTRAINT,
 	.type = PM_QOS_MIN,
+	.notifiers = &cpu_wakeup_latency_qos_notifiers,
 };
 
 /**
@@ -436,6 +464,28 @@ s32 cpu_wakeup_latency_qos_limit(void)
 	return pm_qos_read_value(&cpu_wakeup_latency_constraints);
 }
 
+/**
+ * cpu_wakeup_latency_qos_add_notifier - Add wakeup latency QoS notifier.
+ * @notifier: Notifier block managed by the caller.
+ */
+int cpu_wakeup_latency_qos_add_notifier(struct notifier_block *notifier)
+{
+	return blocking_notifier_chain_register(&cpu_wakeup_latency_qos_notifiers,
+						notifier);
+}
+EXPORT_SYMBOL_GPL(cpu_wakeup_latency_qos_add_notifier);
+
+/**
+ * cpu_wakeup_latency_qos_remove_notifier - Remove wakeup latency QoS notifier.
+ * @notifier: Notifier block previously registered.
+ */
+int cpu_wakeup_latency_qos_remove_notifier(struct notifier_block *notifier)
+{
+	return blocking_notifier_chain_unregister(&cpu_wakeup_latency_qos_notifiers,
+						  notifier);
+}
+EXPORT_SYMBOL_GPL(cpu_wakeup_latency_qos_remove_notifier);
+
 static int cpu_wakeup_latency_qos_open(struct inode *inode, struct file *filp)
 {
 	struct pm_qos_request *req;
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* [PATCH 2/5] cpuidle: subscribe to global latency QoS notifiers
  2026-07-21  9:23 [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 1/5] pm: qos: add notifiers for CPU latency and wakeup latency QoS Yaxiong Tian
@ 2026-07-21  9:25 ` Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 3/5] cpuidle: invalidate latency gen on per-CPU resume QoS changes Yaxiong Tian
                   ` (2 subsequent siblings)
  4 siblings, 0 replies; 8+ messages in thread
From: Yaxiong Tian @ 2026-07-21  9:25 UTC (permalink / raw)
  To: rafael, daniel.lezcano, christian.loehle, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest, Yaxiong Tian

Register for CPU latency and wakeup latency QoS notifications and
bump a per-CPU generation counter when either aggregate constraint
changes.  This prepares for caching the governor latency requirement
without introducing a direct qos-to-cpuidle call dependency.

Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
---
 drivers/cpuidle/governor.c | 58 ++++++++++++++++++++++++++++++++++++++
 1 file changed, 58 insertions(+)

diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
index 5d0e7f78c6c5..bc4a70d30c34 100644
--- a/drivers/cpuidle/governor.c
+++ b/drivers/cpuidle/governor.c
@@ -8,10 +8,13 @@
  * This code is licenced under the GPL.
  */
 
+#include <linux/atomic.h>
 #include <linux/cpu.h>
 #include <linux/cpuidle.h>
+#include <linux/init.h>
 #include <linux/mutex.h>
 #include <linux/module.h>
+#include <linux/notifier.h>
 #include <linux/pm_qos.h>
 
 #include "cpuidle.h"
@@ -22,6 +25,61 @@ LIST_HEAD(cpuidle_governors);
 struct cpuidle_governor *cpuidle_curr_governor;
 struct cpuidle_governor *cpuidle_prev_governor;
 
+/*
+ * Per-CPU generation bumped to invalidate that CPU's cached latency
+ * constraint.  Consumers of the generation are added in later changes.
+ */
+static DEFINE_PER_CPU(atomic_t, latency_req_gen);
+
+static void cpuidle_latency_req_invalidate_cpu(unsigned int cpu)
+{
+	atomic_inc(per_cpu_ptr(&latency_req_gen, cpu));
+}
+
+static void cpuidle_latency_req_invalidate_all(void)
+{
+	unsigned int cpu;
+
+	for_each_possible_cpu(cpu)
+		cpuidle_latency_req_invalidate_cpu(cpu);
+}
+
+static int cpuidle_global_qos_notify(struct notifier_block *nb,
+				     unsigned long action, void *data)
+{
+	cpuidle_latency_req_invalidate_all();
+	return NOTIFY_OK;
+}
+
+static struct notifier_block cpuidle_latency_qos_nb = {
+	.notifier_call = cpuidle_global_qos_notify,
+};
+
+#ifdef CONFIG_PM_QOS_CPU_SYSTEM_WAKEUP
+static struct notifier_block cpuidle_wakeup_qos_nb = {
+	.notifier_call = cpuidle_global_qos_notify,
+};
+#endif
+
+static int __init cpuidle_latency_req_init(void)
+{
+	int ret;
+
+	ret = cpu_latency_qos_add_notifier(&cpuidle_latency_qos_nb);
+	if (ret)
+		return ret;
+
+#ifdef CONFIG_PM_QOS_CPU_SYSTEM_WAKEUP
+	ret = cpu_wakeup_latency_qos_add_notifier(&cpuidle_wakeup_qos_nb);
+	if (ret) {
+		cpu_latency_qos_remove_notifier(&cpuidle_latency_qos_nb);
+		return ret;
+	}
+#endif
+	return 0;
+}
+core_initcall(cpuidle_latency_req_init);
+
 /**
  * cpuidle_find_governor - finds a governor of the specified name
  * @str: the name
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* [PATCH 3/5] cpuidle: invalidate latency gen on per-CPU resume QoS changes
  2026-07-21  9:23 [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 1/5] pm: qos: add notifiers for CPU latency and wakeup latency QoS Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 2/5] cpuidle: subscribe to global latency QoS notifiers Yaxiong Tian
@ 2026-07-21  9:25 ` Yaxiong Tian
  2026-07-21  9:25 ` [PATCH 4/5] cpuidle: cache aggregated governor latency QoS constraint Yaxiong Tian
  2026-07-21  9:26 ` [PATCH 5/5] selftests/cpuidle: add latency_req QoS idle-state selection test Yaxiong Tian
  4 siblings, 0 replies; 8+ messages in thread
From: Yaxiong Tian @ 2026-07-21  9:25 UTC (permalink / raw)
  To: rafael, daniel.lezcano, christian.loehle, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest, Yaxiong Tian

Register a DEV_PM_QOS_RESUME_LATENCY notifier for each CPU device and
bump only that CPU's latency_req generation when its resume latency
constraint changes.

Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
---
 drivers/cpuidle/cpuidle.c  | 15 +++++++++---
 drivers/cpuidle/cpuidle.h  |  2 ++
 drivers/cpuidle/governor.c | 48 +++++++++++++++++++++++++++++++++++++-
 3 files changed, 61 insertions(+), 4 deletions(-)

diff --git a/drivers/cpuidle/cpuidle.c b/drivers/cpuidle/cpuidle.c
index 2d2f40a2cb81..9495f904fe85 100644
--- a/drivers/cpuidle/cpuidle.c
+++ b/drivers/cpuidle/cpuidle.c
@@ -613,6 +613,8 @@ static void __cpuidle_unregister_device(struct cpuidle_device *dev)
 {
 	struct cpuidle_driver *drv = cpuidle_get_cpu_driver(dev);
 
+	cpuidle_latency_req_notifier_unregister(dev->cpu);
+
 	list_del(&dev->device_list);
 	per_cpu(cpuidle_devices, dev->cpu) = NULL;
 	module_put(drv->owner);
@@ -661,10 +663,17 @@ static int __cpuidle_register_device(struct cpuidle_device *dev)
 
 	ret = cpuidle_coupled_register_device(dev);
 	if (ret)
-		__cpuidle_unregister_device(dev);
-	else
-		dev->registered = 1;
+		goto unreg;
+
+	ret = cpuidle_latency_req_notifier_register(cpu);
+	if (ret)
+		goto unreg;
 
+	dev->registered = 1;
+	return 0;
+
+unreg:
+	__cpuidle_unregister_device(dev);
 	return ret;
 }
 
diff --git a/drivers/cpuidle/cpuidle.h b/drivers/cpuidle/cpuidle.h
index 52701d9588f1..151fa9ebe483 100644
--- a/drivers/cpuidle/cpuidle.h
+++ b/drivers/cpuidle/cpuidle.h
@@ -25,6 +25,8 @@ extern void cpuidle_uninstall_idle_handler(void);
 /* governors */
 extern struct cpuidle_governor *cpuidle_find_governor(const char *str);
 extern int cpuidle_switch_governor(struct cpuidle_governor *gov);
+int cpuidle_latency_req_notifier_register(unsigned int cpu);
+void cpuidle_latency_req_notifier_unregister(unsigned int cpu);
 
 /* sysfs */
 
diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
index bc4a70d30c34..d286ccf19a69 100644
--- a/drivers/cpuidle/governor.c
+++ b/drivers/cpuidle/governor.c
@@ -27,10 +27,18 @@ struct cpuidle_governor *cpuidle_prev_governor;
 
 /*
  * Per-CPU generation bumped to invalidate that CPU's cached latency
- * constraint.  Consumers of the generation are added in later changes.
+ * constraint.  Global QoS changes invalidate every CPU; per-CPU resume
+ * latency changes invalidate only the affected CPU.
  */
 static DEFINE_PER_CPU(atomic_t, latency_req_gen);
 
+struct cpuidle_cpu_qos_nb {
+	struct notifier_block nb;
+	unsigned int cpu;
+};
+
+static DEFINE_PER_CPU(struct cpuidle_cpu_qos_nb, cpuidle_cpu_qos_nb);
+
 static void cpuidle_latency_req_invalidate_cpu(unsigned int cpu)
 {
 	atomic_inc(per_cpu_ptr(&latency_req_gen, cpu));
@@ -61,6 +69,44 @@ static struct notifier_block cpuidle_wakeup_qos_nb = {
 };
 #endif
 
+static int cpuidle_cpu_qos_notify(struct notifier_block *nb,
+				 unsigned long action, void *data)
+{
+	struct cpuidle_cpu_qos_nb *qos_nb =
+		container_of(nb, struct cpuidle_cpu_qos_nb, nb);
+
+	cpuidle_latency_req_invalidate_cpu(qos_nb->cpu);
+	return NOTIFY_OK;
+}
+
+int cpuidle_latency_req_notifier_register(unsigned int cpu)
+{
+	struct device *device = get_cpu_device(cpu);
+	struct cpuidle_cpu_qos_nb *qos_nb =
+		per_cpu_ptr(&cpuidle_cpu_qos_nb, cpu);
+
+	if (!device)
+		return -ENODEV;
+
+	qos_nb->cpu = cpu;
+	qos_nb->nb.notifier_call = cpuidle_cpu_qos_notify;
+	return dev_pm_qos_add_notifier(device, &qos_nb->nb,
+				       DEV_PM_QOS_RESUME_LATENCY);
+}
+
+void cpuidle_latency_req_notifier_unregister(unsigned int cpu)
+{
+	struct device *device = get_cpu_device(cpu);
+	struct cpuidle_cpu_qos_nb *qos_nb =
+		per_cpu_ptr(&cpuidle_cpu_qos_nb, cpu);
+
+	if (!device)
+		return;
+
+	dev_pm_qos_remove_notifier(device, &qos_nb->nb,
+				   DEV_PM_QOS_RESUME_LATENCY);
+}
+
 static int __init cpuidle_latency_req_init(void)
 {
 	int ret;
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* [PATCH 4/5] cpuidle: cache aggregated governor latency QoS constraint
  2026-07-21  9:23 [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint Yaxiong Tian
                   ` (2 preceding siblings ...)
  2026-07-21  9:25 ` [PATCH 3/5] cpuidle: invalidate latency gen on per-CPU resume QoS changes Yaxiong Tian
@ 2026-07-21  9:25 ` Yaxiong Tian
  2026-07-23 12:44   ` Christian Loehle
  2026-07-21  9:26 ` [PATCH 5/5] selftests/cpuidle: add latency_req QoS idle-state selection test Yaxiong Tian
  4 siblings, 1 reply; 8+ messages in thread
From: Yaxiong Tian @ 2026-07-21  9:25 UTC (permalink / raw)
  To: rafael, daniel.lezcano, christian.loehle, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest, Yaxiong Tian

cpuidle_governor_latency_req() runs on every idle-state selection and
aggregates per-CPU resume latency with the global CPU latency and
wakeup latency QoS limits.  That path repeatedly walks get_cpu_device()
and pm_qos_read_value(), which shows up hot under menu_select().

Cache the aggregated constraint per CPU and refresh it only when the
corresponding per-CPU generation changes.  The generation is already
bumped by the global and per-CPU resume latency QoS notifiers added
earlier in this series.

On a menu governor profile, cpuidle_governor_latency_req() drops from
about 19.9% of menu_select time to about 4.2%, roughly a 6x reduction
in cost per call (~1.9 us down to ~0.3 us).

Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
---
 drivers/cpuidle/governor.c | 37 ++++++++++++++++++++++++++++++++-----
 1 file changed, 32 insertions(+), 5 deletions(-)

diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
index d286ccf19a69..67909470c14b 100644
--- a/drivers/cpuidle/governor.c
+++ b/drivers/cpuidle/governor.c
@@ -32,11 +32,17 @@ struct cpuidle_governor *cpuidle_prev_governor;
  */
 static DEFINE_PER_CPU(atomic_t, latency_req_gen);
 
+struct cpuidle_latency_req_cache {
+	unsigned int gen;
+	s64 latency_ns;
+};
+
 struct cpuidle_cpu_qos_nb {
 	struct notifier_block nb;
 	unsigned int cpu;
 };
 
+static DEFINE_PER_CPU(struct cpuidle_latency_req_cache, latency_req_cache);
 static DEFINE_PER_CPU(struct cpuidle_cpu_qos_nb, cpuidle_cpu_qos_nb);
 
 static void cpuidle_latency_req_invalidate_cpu(unsigned int cpu)
@@ -212,10 +218,22 @@ int cpuidle_register_governor(struct cpuidle_governor *gov)
  */
 s64 cpuidle_governor_latency_req(unsigned int cpu)
 {
-	struct device *device = get_cpu_device(cpu);
-	int device_req = dev_pm_qos_raw_resume_latency(device);
-	int global_req = cpu_latency_qos_limit();
-	int global_wake_req = cpu_wakeup_latency_qos_limit();
+	struct cpuidle_latency_req_cache *cache;
+	unsigned int gen;
+	struct device *device;
+	int device_req, global_req, global_wake_req;
+	s64 latency_ns;
+
+	cache = per_cpu_ptr(&latency_req_cache, cpu);
+	gen = atomic_read(per_cpu_ptr(&latency_req_gen, cpu));
+
+	if (likely(READ_ONCE(cache->gen) == gen))
+		return READ_ONCE(cache->latency_ns);
+
+	device = get_cpu_device(cpu);
+	device_req = dev_pm_qos_raw_resume_latency(device);
+	global_req = cpu_latency_qos_limit();
+	global_wake_req = cpu_wakeup_latency_qos_limit();
 
 	if (global_req > global_wake_req)
 		global_req = global_wake_req;
@@ -223,5 +241,14 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
 	if (device_req > global_req)
 		device_req = global_req;
 
-	return (s64)device_req * NSEC_PER_USEC;
+	latency_ns = (s64)device_req * NSEC_PER_USEC;
+
+	WRITE_ONCE(cache->latency_ns, latency_ns);
+	/*
+	 * Store gen last so a concurrent invalidate cannot leave a stale
+	 * latency_ns marked as current.
+	 */
+	WRITE_ONCE(cache->gen, gen);
+
+	return latency_ns;
 }
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* [PATCH 5/5] selftests/cpuidle: add latency_req QoS idle-state selection test
  2026-07-21  9:23 [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint Yaxiong Tian
                   ` (3 preceding siblings ...)
  2026-07-21  9:25 ` [PATCH 4/5] cpuidle: cache aggregated governor latency QoS constraint Yaxiong Tian
@ 2026-07-21  9:26 ` Yaxiong Tian
  2026-07-23 12:51   ` Christian Loehle
  4 siblings, 1 reply; 8+ messages in thread
From: Yaxiong Tian @ 2026-07-21  9:26 UTC (permalink / raw)
  To: rafael, daniel.lezcano, christian.loehle, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest, Yaxiong Tian

Verify that CPU latency, wakeup latency, and per-CPU resume latency
QoS ceilings prevent governors from selecting idle states whose exit
latency exceeds the constraint, by comparing cpuidle state usage
deltas under each QoS path.

USE:
sudo make -C tools/testing/selftests TARGETS=cpuidle run_tests

Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
---
 tools/testing/selftests/Makefile              |   1 +
 tools/testing/selftests/cpuidle/Makefile      |   6 +
 tools/testing/selftests/cpuidle/config        |   3 +
 .../cpuidle/cpuidle_latency_req_qos.py        | 331 ++++++++++++++++++
 tools/testing/selftests/cpuidle/settings      |   2 +
 5 files changed, 343 insertions(+)
 create mode 100644 tools/testing/selftests/cpuidle/Makefile
 create mode 100644 tools/testing/selftests/cpuidle/config
 create mode 100755 tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
 create mode 100644 tools/testing/selftests/cpuidle/settings

diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Makefile
index 8189f333814c..a8fb620bebd1 100644
--- a/tools/testing/selftests/Makefile
+++ b/tools/testing/selftests/Makefile
@@ -12,6 +12,7 @@ TARGETS += clone3
 TARGETS += connector
 TARGETS += core
 TARGETS += cpufreq
+TARGETS += cpuidle
 TARGETS += cpu-hotplug
 TARGETS += damon
 TARGETS += devices/error_logs
diff --git a/tools/testing/selftests/cpuidle/Makefile b/tools/testing/selftests/cpuidle/Makefile
new file mode 100644
index 000000000000..f960d72c3a65
--- /dev/null
+++ b/tools/testing/selftests/cpuidle/Makefile
@@ -0,0 +1,6 @@
+# SPDX-License-Identifier: GPL-2.0
+all:
+
+TEST_PROGS := cpuidle_latency_req_qos.py
+
+include ../lib.mk
diff --git a/tools/testing/selftests/cpuidle/config b/tools/testing/selftests/cpuidle/config
new file mode 100644
index 000000000000..86e8f87d4e62
--- /dev/null
+++ b/tools/testing/selftests/cpuidle/config
@@ -0,0 +1,3 @@
+CONFIG_CPU_IDLE=y
+CONFIG_CPU_IDLE_GOV_MENU=y
+CONFIG_PM_QOS_CPU_SYSTEM_WAKEUP=y
diff --git a/tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py b/tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
new file mode 100755
index 000000000000..875b623da530
--- /dev/null
+++ b/tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
@@ -0,0 +1,331 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+"""
+cpuidle: verify latency_req QoS ceilings restrict idle-state selection.
+
+Constrains exit latency via each of the three QoS inputs and checks that
+cpuidle states whose exit latency exceeds the ceiling do not gain usage:
+
+  1) /dev/cpu_dma_latency
+  2) /dev/cpu_wakeup_latency
+  3) /sys/devices/system/cpu/cpuN/power/pm_qos_resume_latency_us
+"""
+
+from __future__ import annotations
+
+import glob
+import os
+import struct
+import sys
+import time
+from dataclasses import dataclass
+from typing import Dict, List, Optional, Tuple
+
+# Source tree: tools/testing/selftests/cpuidle/../kselftest
+# Install tree: kselftest_install/cpuidle/../kselftest
+sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)),
+				"..", "kselftest"))
+import ksft
+
+
+CPUIDLE_BASE = "/sys/devices/system/cpu"
+DMA_LAT_DEV = "/dev/cpu_dma_latency"
+WAKEUP_LAT_DEV = "/dev/cpu_wakeup_latency"
+
+# Keep windows short so the whole collection fits under settings timeout.
+IDLE_SEC = 2.0
+CPU = 0
+
+
+@dataclass
+class IdleState:
+	index: int
+	name: str
+	latency_us: int
+	residency_us: int
+	usage: int
+	time_us: int
+
+
+def read_states(cpu: int) -> Dict[int, IdleState]:
+	base = f"{CPUIDLE_BASE}/cpu{cpu}/cpuidle"
+	states: Dict[int, IdleState] = {}
+	paths = sorted(
+		glob.glob(f"{base}/state*"),
+		key=lambda p: int(os.path.basename(p).replace("state", "")),
+	)
+	for path in paths:
+		idx = int(os.path.basename(path).replace("state", ""))
+		states[idx] = IdleState(
+			index=idx,
+			name=open(f"{path}/name").read().strip(),
+			latency_us=int(open(f"{path}/latency").read()),
+			residency_us=int(open(f"{path}/residency").read()),
+			usage=int(open(f"{path}/usage").read()),
+			time_us=int(open(f"{path}/time").read()),
+		)
+	return states
+
+
+def usage_delta(before: Dict[int, IdleState],
+		after: Dict[int, IdleState]) -> Dict[int, int]:
+	return {i: after[i].usage - before[i].usage for i in before}
+
+
+def pick_ceilings(states: Dict[int, IdleState]) -> List[int]:
+	nonzero = sorted({s.latency_us for s in states.values() if s.latency_us > 0})
+	out: List[int] = []
+	if nonzero:
+		out.append(nonzero[0])
+	if len(nonzero) >= 2:
+		mid = (nonzero[0] + nonzero[1]) // 2
+		out.append(mid if mid > nonzero[0] else max(1, nonzero[1] - 1))
+	if len(nonzero) >= 3:
+		out.append((nonzero[1] + nonzero[2]) // 2)
+	# Dedup, keep order
+	seen = set()
+	uniq = []
+	for c in out:
+		if c not in seen and c >= 1:
+			seen.add(c)
+			uniq.append(c)
+	return uniq or [1]
+
+
+def idle_on_cpu(seconds: float) -> None:
+	end = time.monotonic() + seconds
+	burn_end = time.monotonic() + min(0.1, seconds / 10)
+	while time.monotonic() < burn_end:
+		pass
+	while time.monotonic() < end:
+		time.sleep(0.05)
+
+
+def allowed_states(states: Dict[int, IdleState], ceiling: int) -> List[int]:
+	return [i for i, s in states.items() if s.latency_us <= ceiling]
+
+
+def forbidden_states(states: Dict[int, IdleState], ceiling: int) -> List[int]:
+	return [i for i, s in states.items() if s.latency_us > ceiling]
+
+
+def fmt_state_list(states: Dict[int, IdleState], idxs: List[int]) -> str:
+	if not idxs:
+		return "(none)"
+	return ", ".join(
+		f"state{i}:{states[i].name}(lat={states[i].latency_us})"
+		for i in idxs
+	)
+
+
+def forbidden_violations(states: Dict[int, IdleState], udelta: Dict[int, int],
+			 ceiling: int) -> List[str]:
+	"""States that must not be entered but still gained usage."""
+	bad = []
+	for i in forbidden_states(states, ceiling):
+		if udelta[i] > 0:
+			s = states[i]
+			bad.append(
+				f"state{i}({s.name},lat={s.latency_us}) usage+={udelta[i]}"
+			)
+	return bad
+
+
+def print_usage_table(states: Dict[int, IdleState], udelta: Dict[int, int],
+		      ceiling: int) -> None:
+	ksft.print_msg(
+		f"{'idx':>3} {'name':<12} {'lat':>6} {'d_usage':>8} {'expect':>8}"
+	)
+	for i in sorted(states):
+		s = states[i]
+		expect = "allow" if s.latency_us <= ceiling else "forbid"
+		ksft.print_msg(
+			f"{i:3d} {s.name:<12} {s.latency_us:6d} {udelta[i]:8d} {expect:>8}"
+		)
+
+
+class DmaLatencyGuard:
+	def __init__(self, latency_us: int):
+		self.latency_us = latency_us
+		self.fd = -1
+
+	def __enter__(self):
+		self.fd = os.open(DMA_LAT_DEV, os.O_RDWR)
+		os.write(self.fd, struct.pack("i", int(self.latency_us)))
+		return self
+
+	def __exit__(self, *args):
+		if self.fd >= 0:
+			os.close(self.fd)
+			self.fd = -1
+
+
+class WakeupLatencyGuard:
+	def __init__(self, latency_us: int):
+		self.latency_us = latency_us
+		self.fd = -1
+
+	def __enter__(self):
+		self.fd = os.open(WAKEUP_LAT_DEV, os.O_RDWR)
+		os.write(self.fd, struct.pack("i", int(self.latency_us)))
+		return self
+
+	def __exit__(self, *args):
+		if self.fd >= 0:
+			os.close(self.fd)
+			self.fd = -1
+
+
+class ResumeLatencyGuard:
+	def __init__(self, cpu: int, latency_us: int):
+		self.path = f"{CPUIDLE_BASE}/cpu{cpu}/power/pm_qos_resume_latency_us"
+		self.latency_us = latency_us
+		self.prev: Optional[str] = None
+
+	def __enter__(self):
+		self.prev = open(self.path).read().strip()
+		with open(self.path, "w") as f:
+			f.write(f"{int(self.latency_us)}\n")
+		return self
+
+	def __exit__(self, *args):
+		restore = "0" if self.prev in ("0", "n/a", None) else self.prev
+		with open(self.path, "w") as f:
+			f.write(f"{restore}\n")
+
+
+def run_case(desc: str, states: Dict[int, IdleState], ceiling: int,
+	     guard) -> None:
+	allow = allowed_states(states, ceiling)
+	forbid = forbidden_states(states, ceiling)
+
+	ksft.print_msg(f"=== {desc} ===")
+	ksft.print_msg(f"ceiling={ceiling}us")
+	ksft.print_msg(f"allowed:   {fmt_state_list(states, allow)}")
+	ksft.print_msg(f"forbidden: {fmt_state_list(states, forbid)}")
+
+	with guard:
+		time.sleep(0.05)
+		before = read_states(CPU)
+		idle_on_cpu(IDLE_SEC)
+		after = read_states(CPU)
+
+	ud = usage_delta(before, after)
+	total = sum(ud.values())
+	violations = forbidden_violations(states, ud, ceiling)
+
+	print_usage_table(states, ud, ceiling)
+	ksft.print_msg(
+		f"total_usage+={total} "
+		f"violations={violations if violations else 'none'}"
+	)
+
+	if total <= 0:
+		ksft.test_result_fail(f"{desc}: too little idle activity ({total})")
+		return
+	if violations:
+		ksft.test_result_fail(f"{desc}: {'; '.join(violations)}")
+		return
+	ksft.test_result_pass(desc)
+
+
+def build_plan(states: Dict[int, IdleState],
+	       ceilings: List[int]) -> List[Tuple[str, int, object]]:
+	"""Return list of (description, ceiling, context-manager factory args)."""
+	cases: List[Tuple[str, int, object]] = []
+
+	have_dma = os.path.exists(DMA_LAT_DEV)
+	have_wakeup = os.path.exists(WAKEUP_LAT_DEV)
+	resume_path = f"{CPUIDLE_BASE}/cpu{CPU}/power/pm_qos_resume_latency_us"
+	have_resume = os.path.exists(resume_path)
+
+	for ceiling in ceilings:
+		if have_dma:
+			cases.append(
+				(f"cpu_dma_latency ceiling={ceiling}", ceiling,
+				 ("dma", ceiling))
+			)
+		else:
+			cases.append(
+				(f"cpu_dma_latency ceiling={ceiling}", ceiling,
+				 ("skip", "missing /dev/cpu_dma_latency"))
+			)
+
+		if have_wakeup:
+			cases.append(
+				(f"cpu_wakeup_latency ceiling={ceiling}", ceiling,
+				 ("wakeup", ceiling))
+			)
+		else:
+			cases.append(
+				(f"cpu_wakeup_latency ceiling={ceiling}", ceiling,
+				 ("skip", "missing /dev/cpu_wakeup_latency"))
+			)
+
+		if have_resume:
+			cases.append(
+				(f"pm_qos_resume_latency_us ceiling={ceiling}",
+				 ceiling, ("resume", ceiling))
+			)
+		else:
+			cases.append(
+				(f"pm_qos_resume_latency_us ceiling={ceiling}",
+				 ceiling, ("skip", f"missing {resume_path}"))
+			)
+
+	return cases
+
+
+def main() -> None:
+	ksft.print_header()
+
+	if os.geteuid() != 0:
+		ksft.set_plan(1)
+		ksft.test_result_skip("must run as root")
+		ksft.finished()
+
+	cpuidle_dir = f"{CPUIDLE_BASE}/cpu{CPU}/cpuidle"
+	if not os.path.isdir(cpuidle_dir):
+		ksft.set_plan(1)
+		ksft.test_result_skip(f"no cpuidle sysfs at {cpuidle_dir}")
+		ksft.finished()
+
+	states = read_states(CPU)
+	if not states:
+		ksft.set_plan(1)
+		ksft.test_result_skip("no cpuidle states")
+		ksft.finished()
+
+	gov_path = f"{CPUIDLE_BASE}/cpuidle/current_governor"
+	gov = open(gov_path).read().strip() if os.path.exists(gov_path) else "?"
+	ksft.print_msg(f"governor={gov} cpu={CPU}")
+	for i in sorted(states):
+		s = states[i]
+		ksft.print_msg(
+			f"state{i}: {s.name} latency={s.latency_us}us "
+			f"residency={s.residency_us}us"
+		)
+
+	ceilings = pick_ceilings(states)
+	ksft.print_msg(f"ceilings_us={ceilings}")
+	cases = build_plan(states, ceilings)
+	ksft.set_plan(len(cases))
+
+	for desc, ceiling, kind in cases:
+		tag, arg = kind[0], kind[1]
+		if tag == "skip":
+			ksft.test_result_skip(f"{desc}: {arg}")
+			continue
+		if tag == "dma":
+			guard = DmaLatencyGuard(arg)
+		elif tag == "wakeup":
+			guard = WakeupLatencyGuard(arg)
+		else:
+			guard = ResumeLatencyGuard(CPU, arg)
+		run_case(desc, states, ceiling, guard)
+
+	ksft.finished()
+
+
+if __name__ == "__main__":
+	main()
diff --git a/tools/testing/selftests/cpuidle/settings b/tools/testing/selftests/cpuidle/settings
new file mode 100644
index 000000000000..5b445e716562
--- /dev/null
+++ b/tools/testing/selftests/cpuidle/settings
@@ -0,0 +1,2 @@
+# Multiple QoS paths x several latency ceilings x idle windows.
+timeout=180
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* Re: [PATCH 4/5] cpuidle: cache aggregated governor latency QoS constraint
  2026-07-21  9:25 ` [PATCH 4/5] cpuidle: cache aggregated governor latency QoS constraint Yaxiong Tian
@ 2026-07-23 12:44   ` Christian Loehle
  0 siblings, 0 replies; 8+ messages in thread
From: Christian Loehle @ 2026-07-23 12:44 UTC (permalink / raw)
  To: Yaxiong Tian, rafael, daniel.lezcano, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest

On 7/21/26 10:25, Yaxiong Tian wrote:
> cpuidle_governor_latency_req() runs on every idle-state selection and
> aggregates per-CPU resume latency with the global CPU latency and
> wakeup latency QoS limits.  That path repeatedly walks get_cpu_device()
> and pm_qos_read_value(), which shows up hot under menu_select().
> 
> Cache the aggregated constraint per CPU and refresh it only when the
> corresponding per-CPU generation changes.  The generation is already
> bumped by the global and per-CPU resume latency QoS notifiers added
> earlier in this series.
> 
> On a menu governor profile, cpuidle_governor_latency_req() drops from
> about 19.9% of menu_select time to about 4.2%, roughly a 6x reduction
> in cost per call (~1.9 us down to ~0.3 us).
> 
> Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
> ---
>  drivers/cpuidle/governor.c | 37 ++++++++++++++++++++++++++++++++-----
>  1 file changed, 32 insertions(+), 5 deletions(-)
> 
> diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
> index d286ccf19a69..67909470c14b 100644
> --- a/drivers/cpuidle/governor.c
> +++ b/drivers/cpuidle/governor.c
> @@ -32,11 +32,17 @@ struct cpuidle_governor *cpuidle_prev_governor;
>   */
>  static DEFINE_PER_CPU(atomic_t, latency_req_gen);
>  
> +struct cpuidle_latency_req_cache {
> +	unsigned int gen;
> +	s64 latency_ns;
> +};
> +
>  struct cpuidle_cpu_qos_nb {
>  	struct notifier_block nb;
>  	unsigned int cpu;
>  };
>  
> +static DEFINE_PER_CPU(struct cpuidle_latency_req_cache, latency_req_cache);

So this will initialize with latency_req_cache.gen = 0 and latency.req_cache.latency_ns = 0...
Come to think of it maybe a no-req run is also good as the first and last selftest case.

>  static DEFINE_PER_CPU(struct cpuidle_cpu_qos_nb, cpuidle_cpu_qos_nb);
>  
>  static void cpuidle_latency_req_invalidate_cpu(unsigned int cpu)
> @@ -212,10 +218,22 @@ int cpuidle_register_governor(struct cpuidle_governor *gov)
>   */
>  s64 cpuidle_governor_latency_req(unsigned int cpu)
>  {
> -	struct device *device = get_cpu_device(cpu);
> -	int device_req = dev_pm_qos_raw_resume_latency(device);
> -	int global_req = cpu_latency_qos_limit();
> -	int global_wake_req = cpu_wakeup_latency_qos_limit();
> +	struct cpuidle_latency_req_cache *cache;
> +	unsigned int gen;
> +	struct device *device;
> +	int device_req, global_req, global_wake_req;
> +	s64 latency_ns;
> +
> +	cache = per_cpu_ptr(&latency_req_cache, cpu);
> +	gen = atomic_read(per_cpu_ptr(&latency_req_gen, cpu));
> +
> +	if (likely(READ_ONCE(cache->gen) == gen))
> +		return READ_ONCE(cache->latency_ns);
> +
> +	device = get_cpu_device(cpu);
> +	device_req = dev_pm_qos_raw_resume_latency(device);
> +	global_req = cpu_latency_qos_limit();
> +	global_wake_req = cpu_wakeup_latency_qos_limit();
>  
>  	if (global_req > global_wake_req)
>  		global_req = global_wake_req;
> @@ -223,5 +241,14 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
>  	if (device_req > global_req)
>  		device_req = global_req;
>  
> -	return (s64)device_req * NSEC_PER_USEC;
> +	latency_ns = (s64)device_req * NSEC_PER_USEC;
> +
> +	WRITE_ONCE(cache->latency_ns, latency_ns);
> +	/*
> +	 * Store gen last so a concurrent invalidate cannot leave a stale
> +	 * latency_ns marked as current.
> +	 */

This doesn't guarantee the ordering.

> +	WRITE_ONCE(cache->gen, gen);
> +
> +	return latency_ns;
>  }


^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 5/5] selftests/cpuidle: add latency_req QoS idle-state selection test
  2026-07-21  9:26 ` [PATCH 5/5] selftests/cpuidle: add latency_req QoS idle-state selection test Yaxiong Tian
@ 2026-07-23 12:51   ` Christian Loehle
  0 siblings, 0 replies; 8+ messages in thread
From: Christian Loehle @ 2026-07-23 12:51 UTC (permalink / raw)
  To: Yaxiong Tian, rafael, daniel.lezcano, lenb, pavel, shuah
  Cc: linux-kernel, linux-pm, linux-kselftest

On 7/21/26 10:26, Yaxiong Tian wrote:
> Verify that CPU latency, wakeup latency, and per-CPU resume latency
> QoS ceilings prevent governors from selecting idle states whose exit
> latency exceeds the constraint, by comparing cpuidle state usage
> deltas under each QoS path.
> 
> USE:
> sudo make -C tools/testing/selftests TARGETS=cpuidle run_tests

FWIW since you're mostly there anyway, can you extend this to cover cpuidle
state disable too? Ideally also interaction between disable and latency_req.
I think that would be quite useful.

> 
> Signed-off-by: Yaxiong Tian <tianyaxiong@kylinos.cn>
> ---
>  tools/testing/selftests/Makefile              |   1 +
>  tools/testing/selftests/cpuidle/Makefile      |   6 +
>  tools/testing/selftests/cpuidle/config        |   3 +
>  .../cpuidle/cpuidle_latency_req_qos.py        | 331 ++++++++++++++++++
>  tools/testing/selftests/cpuidle/settings      |   2 +
>  5 files changed, 343 insertions(+)
>  create mode 100644 tools/testing/selftests/cpuidle/Makefile
>  create mode 100644 tools/testing/selftests/cpuidle/config
>  create mode 100755 tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
>  create mode 100644 tools/testing/selftests/cpuidle/settings
> 
> diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Makefile
> index 8189f333814c..a8fb620bebd1 100644
> --- a/tools/testing/selftests/Makefile
> +++ b/tools/testing/selftests/Makefile
> @@ -12,6 +12,7 @@ TARGETS += clone3
>  TARGETS += connector
>  TARGETS += core
>  TARGETS += cpufreq
> +TARGETS += cpuidle
>  TARGETS += cpu-hotplug
>  TARGETS += damon
>  TARGETS += devices/error_logs
> diff --git a/tools/testing/selftests/cpuidle/Makefile b/tools/testing/selftests/cpuidle/Makefile
> new file mode 100644
> index 000000000000..f960d72c3a65
> --- /dev/null
> +++ b/tools/testing/selftests/cpuidle/Makefile
> @@ -0,0 +1,6 @@
> +# SPDX-License-Identifier: GPL-2.0
> +all:
> +
> +TEST_PROGS := cpuidle_latency_req_qos.py
> +
> +include ../lib.mk
> diff --git a/tools/testing/selftests/cpuidle/config b/tools/testing/selftests/cpuidle/config
> new file mode 100644
> index 000000000000..86e8f87d4e62
> --- /dev/null
> +++ b/tools/testing/selftests/cpuidle/config
> @@ -0,0 +1,3 @@
> +CONFIG_CPU_IDLE=y
> +CONFIG_CPU_IDLE_GOV_MENU=y
> +CONFIG_PM_QOS_CPU_SYSTEM_WAKEUP=y
> diff --git a/tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py b/tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
> new file mode 100755
> index 000000000000..875b623da530
> --- /dev/null
> +++ b/tools/testing/selftests/cpuidle/cpuidle_latency_req_qos.py
> @@ -0,0 +1,331 @@
> +#!/usr/bin/env python3
> +# SPDX-License-Identifier: GPL-2.0
> +"""
> +cpuidle: verify latency_req QoS ceilings restrict idle-state selection.
> +
> +Constrains exit latency via each of the three QoS inputs and checks that
> +cpuidle states whose exit latency exceeds the ceiling do not gain usage:
> +
> +  1) /dev/cpu_dma_latency
> +  2) /dev/cpu_wakeup_latency
> +  3) /sys/devices/system/cpu/cpuN/power/pm_qos_resume_latency_us
> +"""
> +
> +from __future__ import annotations
> +
> +import glob
> +import os
> +import struct
> +import sys
> +import time
> +from dataclasses import dataclass
> +from typing import Dict, List, Optional, Tuple
> +
> +# Source tree: tools/testing/selftests/cpuidle/../kselftest
> +# Install tree: kselftest_install/cpuidle/../kselftest
> +sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)),
> +				"..", "kselftest"))
> +import ksft
> +
> +
> +CPUIDLE_BASE = "/sys/devices/system/cpu"
> +DMA_LAT_DEV = "/dev/cpu_dma_latency"
> +WAKEUP_LAT_DEV = "/dev/cpu_wakeup_latency"
> +
> +# Keep windows short so the whole collection fits under settings timeout.
> +IDLE_SEC = 2.0
> +CPU = 0
> +
> +
> +@dataclass
> +class IdleState:
> +	index: int
> +	name: str
> +	latency_us: int
> +	residency_us: int
> +	usage: int
> +	time_us: int
> +
> +
> +def read_states(cpu: int) -> Dict[int, IdleState]:
> +	base = f"{CPUIDLE_BASE}/cpu{cpu}/cpuidle"
> +	states: Dict[int, IdleState] = {}
> +	paths = sorted(
> +		glob.glob(f"{base}/state*"),
> +		key=lambda p: int(os.path.basename(p).replace("state", "")),
> +	)
> +	for path in paths:
> +		idx = int(os.path.basename(path).replace("state", ""))
> +		states[idx] = IdleState(
> +			index=idx,
> +			name=open(f"{path}/name").read().strip(),
> +			latency_us=int(open(f"{path}/latency").read()),
> +			residency_us=int(open(f"{path}/residency").read()),
> +			usage=int(open(f"{path}/usage").read()),
> +			time_us=int(open(f"{path}/time").read()),
> +		)
> +	return states
> +
> +
> +def usage_delta(before: Dict[int, IdleState],
> +		after: Dict[int, IdleState]) -> Dict[int, int]:
> +	return {i: after[i].usage - before[i].usage for i in before}
> +
> +
> +def pick_ceilings(states: Dict[int, IdleState]) -> List[int]:
> +	nonzero = sorted({s.latency_us for s in states.values() if s.latency_us > 0})
> +	out: List[int] = []
> +	if nonzero:
> +		out.append(nonzero[0])
> +	if len(nonzero) >= 2:
> +		mid = (nonzero[0] + nonzero[1]) // 2
> +		out.append(mid if mid > nonzero[0] else max(1, nonzero[1] - 1))
> +	if len(nonzero) >= 3:
> +		out.append((nonzero[1] + nonzero[2]) // 2)
> +	# Dedup, keep order
> +	seen = set()
> +	uniq = []
> +	for c in out:
> +		if c not in seen and c >= 1:
> +			seen.add(c)
> +			uniq.append(c)
> +	return uniq or [1]
> +
> +
> +def idle_on_cpu(seconds: float) -> None:
> +	end = time.monotonic() + seconds
> +	burn_end = time.monotonic() + min(0.1, seconds / 10)
> +	while time.monotonic() < burn_end:
> +		pass
> +	while time.monotonic() < end:
> +		time.sleep(0.05)
> +
> +
> +def allowed_states(states: Dict[int, IdleState], ceiling: int) -> List[int]:
> +	return [i for i, s in states.items() if s.latency_us <= ceiling]
> +
> +
> +def forbidden_states(states: Dict[int, IdleState], ceiling: int) -> List[int]:
> +	return [i for i, s in states.items() if s.latency_us > ceiling]
> +
> +
> +def fmt_state_list(states: Dict[int, IdleState], idxs: List[int]) -> str:
> +	if not idxs:
> +		return "(none)"
> +	return ", ".join(
> +		f"state{i}:{states[i].name}(lat={states[i].latency_us})"
> +		for i in idxs
> +	)
> +
> +
> +def forbidden_violations(states: Dict[int, IdleState], udelta: Dict[int, int],
> +			 ceiling: int) -> List[str]:
> +	"""States that must not be entered but still gained usage."""
> +	bad = []
> +	for i in forbidden_states(states, ceiling):
> +		if udelta[i] > 0:
> +			s = states[i]
> +			bad.append(
> +				f"state{i}({s.name},lat={s.latency_us}) usage+={udelta[i]}"
> +			)
> +	return bad
> +
> +
> +def print_usage_table(states: Dict[int, IdleState], udelta: Dict[int, int],
> +		      ceiling: int) -> None:
> +	ksft.print_msg(
> +		f"{'idx':>3} {'name':<12} {'lat':>6} {'d_usage':>8} {'expect':>8}"
> +	)
> +	for i in sorted(states):
> +		s = states[i]
> +		expect = "allow" if s.latency_us <= ceiling else "forbid"
> +		ksft.print_msg(
> +			f"{i:3d} {s.name:<12} {s.latency_us:6d} {udelta[i]:8d} {expect:>8}"
> +		)
> +
> +
> +class DmaLatencyGuard:
> +	def __init__(self, latency_us: int):
> +		self.latency_us = latency_us
> +		self.fd = -1
> +
> +	def __enter__(self):
> +		self.fd = os.open(DMA_LAT_DEV, os.O_RDWR)
> +		os.write(self.fd, struct.pack("i", int(self.latency_us)))
> +		return self
> +
> +	def __exit__(self, *args):
> +		if self.fd >= 0:
> +			os.close(self.fd)
> +			self.fd = -1
> +
> +
> +class WakeupLatencyGuard:
> +	def __init__(self, latency_us: int):
> +		self.latency_us = latency_us
> +		self.fd = -1
> +
> +	def __enter__(self):
> +		self.fd = os.open(WAKEUP_LAT_DEV, os.O_RDWR)
> +		os.write(self.fd, struct.pack("i", int(self.latency_us)))
> +		return self
> +
> +	def __exit__(self, *args):
> +		if self.fd >= 0:
> +			os.close(self.fd)
> +			self.fd = -1
> +
> +
> +class ResumeLatencyGuard:
> +	def __init__(self, cpu: int, latency_us: int):
> +		self.path = f"{CPUIDLE_BASE}/cpu{cpu}/power/pm_qos_resume_latency_us"
> +		self.latency_us = latency_us
> +		self.prev: Optional[str] = None
> +
> +	def __enter__(self):
> +		self.prev = open(self.path).read().strip()
> +		with open(self.path, "w") as f:
> +			f.write(f"{int(self.latency_us)}\n")
> +		return self
> +
> +	def __exit__(self, *args):
> +		restore = "0" if self.prev in ("0", "n/a", None) else self.prev

See pm_qos_resume_latency_us_store(), n/a and 0 are distinct.

> +		with open(self.path, "w") as f:
> +			f.write(f"{restore}\n")
> +
> +
> +def run_case(desc: str, states: Dict[int, IdleState], ceiling: int,
> +	     guard) -> None:
> +	allow = allowed_states(states, ceiling)
> +	forbid = forbidden_states(states, ceiling)
> +
> +	ksft.print_msg(f"=== {desc} ===")
> +	ksft.print_msg(f"ceiling={ceiling}us")
> +	ksft.print_msg(f"allowed:   {fmt_state_list(states, allow)}")
> +	ksft.print_msg(f"forbidden: {fmt_state_list(states, forbid)}")
> +
> +	with guard:
> +		time.sleep(0.05)
> +		before = read_states(CPU)
> +		idle_on_cpu(IDLE_SEC)
> +		after = read_states(CPU)
> +
> +	ud = usage_delta(before, after)
> +	total = sum(ud.values())
> +	violations = forbidden_violations(states, ud, ceiling)
> +
> +	print_usage_table(states, ud, ceiling)
> +	ksft.print_msg(
> +		f"total_usage+={total} "
> +		f"violations={violations if violations else 'none'}"
> +	)
> +
> +	if total <= 0:
> +		ksft.test_result_fail(f"{desc}: too little idle activity ({total})")
> +		return
> +	if violations:
> +		ksft.test_result_fail(f"{desc}: {'; '.join(violations)}")
> +		return
> +	ksft.test_result_pass(desc)
> +
> +
> +def build_plan(states: Dict[int, IdleState],
> +	       ceilings: List[int]) -> List[Tuple[str, int, object]]:
> +	"""Return list of (description, ceiling, context-manager factory args)."""
> +	cases: List[Tuple[str, int, object]] = []
> +
> +	have_dma = os.path.exists(DMA_LAT_DEV)
> +	have_wakeup = os.path.exists(WAKEUP_LAT_DEV)
> +	resume_path = f"{CPUIDLE_BASE}/cpu{CPU}/power/pm_qos_resume_latency_us"
> +	have_resume = os.path.exists(resume_path)
> +
> +	for ceiling in ceilings:
> +		if have_dma:
> +			cases.append(
> +				(f"cpu_dma_latency ceiling={ceiling}", ceiling,
> +				 ("dma", ceiling))
> +			)
> +		else:
> +			cases.append(
> +				(f"cpu_dma_latency ceiling={ceiling}", ceiling,
> +				 ("skip", "missing /dev/cpu_dma_latency"))
> +			)
> +
> +		if have_wakeup:
> +			cases.append(
> +				(f"cpu_wakeup_latency ceiling={ceiling}", ceiling,
> +				 ("wakeup", ceiling))
> +			)
> +		else:
> +			cases.append(
> +				(f"cpu_wakeup_latency ceiling={ceiling}", ceiling,
> +				 ("skip", "missing /dev/cpu_wakeup_latency"))
> +			)
> +
> +		if have_resume:
> +			cases.append(
> +				(f"pm_qos_resume_latency_us ceiling={ceiling}",
> +				 ceiling, ("resume", ceiling))
> +			)
> +		else:
> +			cases.append(
> +				(f"pm_qos_resume_latency_us ceiling={ceiling}",
> +				 ceiling, ("skip", f"missing {resume_path}"))
> +			)
> +
> +	return cases
> +
> +
> +def main() -> None:
> +	ksft.print_header()
> +
> +	if os.geteuid() != 0:
> +		ksft.set_plan(1)
> +		ksft.test_result_skip("must run as root")
> +		ksft.finished()
> +
> +	cpuidle_dir = f"{CPUIDLE_BASE}/cpu{CPU}/cpuidle"
> +	if not os.path.isdir(cpuidle_dir):
> +		ksft.set_plan(1)
> +		ksft.test_result_skip(f"no cpuidle sysfs at {cpuidle_dir}")
> +		ksft.finished()
> +
> +	states = read_states(CPU)
> +	if not states:
> +		ksft.set_plan(1)
> +		ksft.test_result_skip("no cpuidle states")
> +		ksft.finished()
> +
> +	gov_path = f"{CPUIDLE_BASE}/cpuidle/current_governor"
> +	gov = open(gov_path).read().strip() if os.path.exists(gov_path) else "?"
> +	ksft.print_msg(f"governor={gov} cpu={CPU}")
> +	for i in sorted(states):
> +		s = states[i]
> +		ksft.print_msg(
> +			f"state{i}: {s.name} latency={s.latency_us}us "
> +			f"residency={s.residency_us}us"
> +		)
> +
> +	ceilings = pick_ceilings(states)
> +	ksft.print_msg(f"ceilings_us={ceilings}")
> +	cases = build_plan(states, ceilings)
> +	ksft.set_plan(len(cases))
> +
> +	for desc, ceiling, kind in cases:
> +		tag, arg = kind[0], kind[1]
> +		if tag == "skip":
> +			ksft.test_result_skip(f"{desc}: {arg}")
> +			continue
> +		if tag == "dma":
> +			guard = DmaLatencyGuard(arg)
> +		elif tag == "wakeup":
> +			guard = WakeupLatencyGuard(arg)
> +		else:
> +			guard = ResumeLatencyGuard(CPU, arg)
> +		run_case(desc, states, ceiling, guard)
> +
> +	ksft.finished()
> +
> +
> +if __name__ == "__main__":
> +	main()
> diff --git a/tools/testing/selftests/cpuidle/settings b/tools/testing/selftests/cpuidle/settings
> new file mode 100644
> index 000000000000..5b445e716562
> --- /dev/null
> +++ b/tools/testing/selftests/cpuidle/settings
> @@ -0,0 +1,2 @@
> +# Multiple QoS paths x several latency ceilings x idle windows.
> +timeout=180


^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-07-23 12:52 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-21  9:23 [PATCH 0/5] cpuidle: speed up do_idle() by caching the governor latency QoS constraint Yaxiong Tian
2026-07-21  9:25 ` [PATCH 1/5] pm: qos: add notifiers for CPU latency and wakeup latency QoS Yaxiong Tian
2026-07-21  9:25 ` [PATCH 2/5] cpuidle: subscribe to global latency QoS notifiers Yaxiong Tian
2026-07-21  9:25 ` [PATCH 3/5] cpuidle: invalidate latency gen on per-CPU resume QoS changes Yaxiong Tian
2026-07-21  9:25 ` [PATCH 4/5] cpuidle: cache aggregated governor latency QoS constraint Yaxiong Tian
2026-07-23 12:44   ` Christian Loehle
2026-07-21  9:26 ` [PATCH 5/5] selftests/cpuidle: add latency_req QoS idle-state selection test Yaxiong Tian
2026-07-23 12:51   ` Christian Loehle

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox