* [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock
@ 2024-12-02 4:38 Changwoo Min
2024-12-02 4:38 ` [PATCH v2 1/5] sched_ext: Implement scx_rq_clock_update/stale() Changwoo Min
` (4 more replies)
0 siblings, 5 replies; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 4:38 UTC (permalink / raw)
To: tj, void, mingo, peterz; +Cc: changwoo, kernel-dev, linux-kernel
Many BPF schedulers (such as scx_central, scx_lavd, scx_rusty, scx_bpfland,
scx_flash) frequently call bpf_ktime_get_ns() for tracking tasks' runtime
properties. If supported, bpf_ktime_get_ns() eventually reads a hardware
timestamp counter (TSC). However, reading a hardware TSC is not
performant in some hardware platforms, degrading IPC.
This patchset addresses the performance problem of reading hardware TSC
by leveraging the rq clock in the scheduler core, introducing a
scx_bpf_clock_get_ns() function for BPF schedulers. Whenever the rq clock
is fresh and valid, scx_bpf_clock_get_ns() provides the rq clock, which is
already updated by the scheduler core (update_rq_clock), so it can reduce
reading the hardware TSC.
When the rq lock is released (rq_unpin_lock) or a long-running
operations are done by the BPF scheduler (ops.running, ops.update_idle),
the rq clock is invalidated, so a subsequent scx_bpf_clock_get_ns() call
gets the fresh sched_clock for the caller.
In addition, scx_bpf_clock_get_ns() guarantees the clock is
monotonically non-decreasing for the same CPU, so the clock cannot go
backward in the same CPU.
Using scx_bpf_clock_get_ns() reduces the number of reading hardware TSC
by 40-70% (65% for scx_lavd, 58% for scx_bpfland, and 43% for scx_rusty)
for the following benchmark.
perf bench -f simple sched messaging -t -g 20 -l 6000
The patchset begins by managing the status of rq clock in the scheduler
core, then implementing scx_bpf_clock_get_ns(), and finally applying it
to the BPF schedulers.
ChangeLog v1 -> v2:
- Rename SCX_RQ_CLK_UPDATED to SCX_RQ_CLK_VALID to denote the validity
of an rq clock clearly.
- Rearrange the clock and flags fields in struct scx_rq to make sure
they are in the same cacheline to minimize the cache misses
- Add an additional explanation to the commit message in the 2/5 patch
describing when the rq clock will be reused with an example.
- Fix typos
- Rebase the code to the tip of Tejun's sched_ext repo
Changwoo Min (5):
sched_ext: Implement scx_rq_clock_update/stale()
sched_ext: Manage the validity of scx_rq_clock
sched_ext: Implement scx_bpf_clock_get_ns()
sched_ext: Add scx_bpf_clock_get_ns() for BPF scheduler
sched_ext: Replace bpf_ktime_get_ns() to scx_bpf_clock_get_ns()
kernel/sched/core.c | 6 +-
kernel/sched/ext.c | 74 ++++++++++++++++++++++++
kernel/sched/sched.h | 24 +++++++-
tools/sched_ext/include/scx/common.bpf.h | 1 +
tools/sched_ext/include/scx/compat.bpf.h | 5 ++
tools/sched_ext/scx_central.bpf.c | 4 +-
tools/sched_ext/scx_flatcg.bpf.c | 2 +-
7 files changed, 110 insertions(+), 6 deletions(-)
--
2.47.1
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v2 1/5] sched_ext: Implement scx_rq_clock_update/stale()
2024-12-02 4:38 [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock Changwoo Min
@ 2024-12-02 4:38 ` Changwoo Min
2024-12-02 4:38 ` [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock Changwoo Min
` (3 subsequent siblings)
4 siblings, 0 replies; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 4:38 UTC (permalink / raw)
To: tj, void, mingo, peterz; +Cc: changwoo, kernel-dev, linux-kernel
scx_rq_clock_update() and scx_rq_clock_stale() manage the status of an
rq clock. scx_rq_clock_update() keeps the rq clock in memory and its
status valid. scx_rq_clock_stale() invalidates the current rq clock
not to use the cached rq clock.
Signed-off-by: Changwoo Min <changwoo@igalia.com>
---
kernel/sched/sched.h | 22 +++++++++++++++++++++-
1 file changed, 21 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 76f5f53a645f..09d5a64f78ce 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -754,6 +754,7 @@ enum scx_rq_flags {
SCX_RQ_BAL_PENDING = 1 << 2, /* balance hasn't run yet */
SCX_RQ_BAL_KEEP = 1 << 3, /* balance decided to keep current */
SCX_RQ_BYPASSING = 1 << 4,
+ SCX_RQ_CLK_VALID = 1 << 5, /* RQ clock is fresh and valid */
SCX_RQ_IN_WAKEUP = 1 << 16,
SCX_RQ_IN_BALANCE = 1 << 17,
@@ -766,8 +767,9 @@ struct scx_rq {
unsigned long ops_qseq;
u64 extra_enq_flags; /* see move_task_to_local_dsq() */
u32 nr_running;
- u32 flags;
u32 cpuperf_target; /* [0, SCHED_CAPACITY_SCALE] */
+ u64 clock; /* cached per-rq clock -- see scx_bpf_clock_get_ns() */
+ u32 flags;
bool cpu_released;
cpumask_var_t cpus_to_kick;
cpumask_var_t cpus_to_kick_if_idle;
@@ -1353,6 +1355,24 @@ static inline void rq_set_donor(struct rq *rq, struct task_struct *t)
/* Do nothing */
}
+#ifdef CONFIG_SCHED_CLASS_EXT
+static inline void scx_rq_clock_update(struct rq *rq, u64 clock)
+{
+ rq->scx.clock = clock;
+ rq->scx.flags |= SCX_RQ_CLK_VALID;
+}
+
+static inline void scx_rq_clock_stale(struct rq *rq)
+{
+ rq->scx.flags &= ~SCX_RQ_CLK_VALID;
+}
+
+#else
+static inline void scx_rq_clock_update(struct rq *rq, u64 clock) {}
+static inline void scx_rq_clock_stale(struct rq *rq) {}
+
+#endif /* CONFIG_SCHED_CLASS_EXT */
+
#ifdef CONFIG_SCHED_CORE
static inline struct cpumask *sched_group_span(struct sched_group *sg);
--
2.47.1
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock
2024-12-02 4:38 [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock Changwoo Min
2024-12-02 4:38 ` [PATCH v2 1/5] sched_ext: Implement scx_rq_clock_update/stale() Changwoo Min
@ 2024-12-02 4:38 ` Changwoo Min
2024-12-02 9:58 ` Peter Zijlstra
2024-12-02 4:38 ` [PATCH v2 3/5] sched_ext: Implement scx_bpf_clock_get_ns() Changwoo Min
` (2 subsequent siblings)
4 siblings, 1 reply; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 4:38 UTC (permalink / raw)
To: tj, void, mingo, peterz; +Cc: changwoo, kernel-dev, linux-kernel
An rq clock becomes valid when it is updated using update_rq_clock()
and invalidated when the rq is unlocked using rq_unpin_lock(). Also,
after long-running operations -- ops.running() and ops.update_idle() --
in a BPF scheduler, the sched_ext core invalidates the rq clock.
Let's suppose the following timeline in the scheduler core:
T1. rq_lock(rq)
T2. update_rq_clock(rq)
T3. a sched_ext BPF operation
T4. rq_unlock(rq)
T5. a sched_ext BPF operation
T6. rq_lock(rq)
T7. update_rq_clock(rq)
For [T2, T4), we consider that rq clock is valid (SCX_RQ_CLK_VALID is
set), so scx_bpf_clock_get_ns() calls during [T2, T4) (including T3)
will return the rq clock updated at T2. For duration [T4, T7),
when a BPF scheduler can still call scx_bpf_clock_get_ns() (T5), we
consider the rq clock is invalid (SCX_RQ_CLK_VALID is unset at T4). So
when calling scx_bpf_clock_get_ns() at T5, we will return a fresh clock
value by calling sched_clock() internally.
One example of calling scx_bpf_clock_get_ns(), when the rq clock is
invalid (like T5), is in scx_central [1]. The scx_central scheduler uses
a BPF timer for preemptive scheduling. In every msec, the timer callback
checks if the currently running tasks exceed their timeslice. At the
beginning of the BPF timer callback (central_timerfn in scx_central.bpf.c),
scx_central gets the current time. When the BPF timer callback runs, the rq
clock could be invalid, the same as T5. In this case, scx_bpf_clock_get_ns()
returns a fresh clock value rather than returning the old one (T2).
[1] https://github.com/sched-ext/scx/blob/main/scheds/c/scx_central.bpf.c
Signed-off-by: Changwoo Min <changwoo@igalia.com>
---
kernel/sched/core.c | 6 +++++-
kernel/sched/ext.c | 3 +++
kernel/sched/sched.h | 2 +-
3 files changed, 9 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 95e40895a519..ab8015c8cab4 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -789,6 +789,7 @@ static void update_rq_clock_task(struct rq *rq, s64 delta)
void update_rq_clock(struct rq *rq)
{
s64 delta;
+ u64 clock;
lockdep_assert_rq_held(rq);
@@ -800,11 +801,14 @@ void update_rq_clock(struct rq *rq)
SCHED_WARN_ON(rq->clock_update_flags & RQCF_UPDATED);
rq->clock_update_flags |= RQCF_UPDATED;
#endif
+ clock = sched_clock_cpu(cpu_of(rq));
+ scx_rq_clock_update(rq, clock);
- delta = sched_clock_cpu(cpu_of(rq)) - rq->clock;
+ delta = clock - rq->clock;
if (delta < 0)
return;
rq->clock += delta;
+
update_rq_clock_task(rq, delta);
}
diff --git a/kernel/sched/ext.c b/kernel/sched/ext.c
index 7fff1d045477..ac279a657d50 100644
--- a/kernel/sched/ext.c
+++ b/kernel/sched/ext.c
@@ -2928,6 +2928,8 @@ static void set_next_task_scx(struct rq *rq, struct task_struct *p, bool first)
if (SCX_HAS_OP(running) && (p->scx.flags & SCX_TASK_QUEUED))
SCX_CALL_OP_TASK(SCX_KF_REST, running, p);
+ scx_rq_clock_stale(rq);
+
clr_task_runnable(p, true);
/*
@@ -3590,6 +3592,7 @@ void __scx_update_idle(struct rq *rq, bool idle)
{
int cpu = cpu_of(rq);
+ scx_rq_clock_stale(rq);
if (SCX_HAS_OP(update_idle) && !scx_rq_bypassing(rq)) {
SCX_CALL_OP(SCX_KF_REST, update_idle, cpu_of(rq), idle);
if (!static_branch_unlikely(&scx_builtin_idle_enabled))
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 09d5a64f78ce..1c1345274b02 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1766,7 +1766,7 @@ static inline void rq_unpin_lock(struct rq *rq, struct rq_flags *rf)
if (rq->clock_update_flags > RQCF_ACT_SKIP)
rf->clock_update_flags = RQCF_UPDATED;
#endif
-
+ scx_rq_clock_stale(rq);
lockdep_unpin_lock(__rq_lockp(rq), rf->cookie);
}
--
2.47.1
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v2 3/5] sched_ext: Implement scx_bpf_clock_get_ns()
2024-12-02 4:38 [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock Changwoo Min
2024-12-02 4:38 ` [PATCH v2 1/5] sched_ext: Implement scx_rq_clock_update/stale() Changwoo Min
2024-12-02 4:38 ` [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock Changwoo Min
@ 2024-12-02 4:38 ` Changwoo Min
2024-12-02 4:38 ` [PATCH v2 4/5] sched_ext: Add scx_bpf_clock_get_ns() for BPF scheduler Changwoo Min
2024-12-02 4:38 ` [PATCH v2 5/5] sched_ext: Replace bpf_ktime_get_ns() to scx_bpf_clock_get_ns() Changwoo Min
4 siblings, 0 replies; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 4:38 UTC (permalink / raw)
To: tj, void, mingo, peterz; +Cc: changwoo, kernel-dev, linux-kernel
Returns a high-performance monotonically non-decreasing clock for the
current CPU. The clock returned is in nanoseconds.
It provides the following properties:
1) High performance: Many BPF schedulers call bpf_ktime_get_ns()
frequently to account for execution time and track tasks' runtime
properties. Unfortunately, in some hardware platforms, bpf_ktime_get_ns()
-- which eventually reads a hardware timestamp counter -- is neither
performant nor scalable. scx_bpf_clock_get_ns() aims to provide a
high-performance clock by using the rq clock in the scheduler core
whenever possible.
2) High enough resolution for the BPF scheduler use cases: In most BPF
scheduler use cases, the required clock resolution is lower than the
most accurate hardware clock (e.g., rdtsc in x86). scx_bpf_clock_get_ns()
basically uses the rq clock in the scheduler core whenever it is valid.
It considers that the rq clock is valid from the time the rq clock is
updated (update_rq_clock) until the rq is unlocked (rq_unpin_lock).
In addition, it invalidates the rq clock after long operations --
ops.running() and ops.update_idle() -- in a BPF scheduler.
3) Monotonically non-decreasing clock for the same CPU:
scx_bpf_clock_get_ns() guarantees the clock never goes backward when
comparing them in the same CPU. On the other hand, when comparing clocks
in different CPUs, there is no such guarantee -- the clock can go backward.
It provides a monotonically *non-decreasing* clock so that it would provide
the same clock values in two different scx_bpf_clock_get_ns() calls in the
same CPU during the same period of when the rq clock is valid.
Signed-off-by: Changwoo Min <changwoo@igalia.com>
---
kernel/sched/ext.c | 71 ++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 71 insertions(+)
diff --git a/kernel/sched/ext.c b/kernel/sched/ext.c
index ac279a657d50..9e2656e21593 100644
--- a/kernel/sched/ext.c
+++ b/kernel/sched/ext.c
@@ -7546,6 +7546,76 @@ __bpf_kfunc struct cgroup *scx_bpf_task_cgroup(struct task_struct *p)
}
#endif
+/**
+ * scx_bpf_clock_get_ns - Returns a high-performance monotonically
+ * non-decreasing clock for the current CPU. The clock returned is in
+ * nanoseconds.
+ *
+ * It provides the following properties:
+ *
+ * 1) High performance: Many BPF schedulers call bpf_ktime_get_ns() frequently
+ * to account for execution time and track tasks' runtime properties.
+ * Unfortunately, in some hardware platforms, bpf_ktime_get_ns() -- which
+ * eventually reads a hardware timestamp counter -- is neither performant nor
+ * scalable. scx_bpf_clock_get_ns() aims to provide a high-performance clock
+ * by using the rq clock in the scheduler core whenever possible.
+ *
+ * 2) High enough resolution for the BPF scheduler use cases: In most BPF
+ * scheduler use cases, the required clock resolution is lower than the most
+ * accurate hardware clock (e.g., rdtsc in x86). scx_bpf_clock_get_ns()
+ * basically uses the rq clock in the scheduler core whenever it is valid.
+ * It considers that the rq clock is valid from the time the rq clock is
+ * updated (update_rq_clock) until the rq is unlocked (rq_unpin_lock).
+ * In addition, it invalidates the rq clock after long operations --
+ * ops.running() and ops.update_idle().
+ *
+ * 3) Monotonically non-decreasing clock for the same CPU:
+ * scx_bpf_clock_get_ns() guarantees the clock never goes backward when
+ * comparing them in the same CPU. On the other hand, when comparing clocks
+ * in different CPUs, there is no such guarantee -- the clock can go backward.
+ * It provides a monotonically *non-decreasing* clock so that it would provide
+ * the same clock values in two different scx_bpf_clock_get_ns() calls in the
+ * same CPU during the same period of when the rq clock is valid.
+ */
+__bpf_kfunc u64 scx_bpf_clock_get_ns(void)
+{
+ static DEFINE_PER_CPU(u64, prev_clk);
+ struct rq *rq = this_rq();
+ u64 pr_clk, cr_clk;
+
+ preempt_disable();
+ pr_clk = __this_cpu_read(prev_clk);
+
+ /*
+ * If the rq clock is invalid, start a new rq clock period
+ * with a fresh sched_clock().
+ */
+ if (!(rq->scx.flags & SCX_RQ_CLK_VALID)) {
+ cr_clk = sched_clock();
+ scx_rq_clock_update(rq, cr_clk);
+ }
+ /*
+ * If the rq clock is valid, use the cached rq clock
+ * whenever the clock does not go backward.
+ */
+ else {
+ cr_clk = rq->scx.clock;
+ /*
+ * If the clock goes backward, start a new rq clock period
+ * with a fresh sched_clock().
+ */
+ if (pr_clk > cr_clk) {
+ cr_clk = sched_clock();
+ scx_rq_clock_update(rq, cr_clk);
+ }
+ }
+
+ __this_cpu_write(prev_clk, cr_clk);
+ preempt_enable();
+
+ return cr_clk;
+}
+
__bpf_kfunc_end_defs();
BTF_KFUNCS_START(scx_kfunc_ids_any)
@@ -7577,6 +7647,7 @@ BTF_ID_FLAGS(func, scx_bpf_cpu_rq)
#ifdef CONFIG_CGROUP_SCHED
BTF_ID_FLAGS(func, scx_bpf_task_cgroup, KF_RCU | KF_ACQUIRE)
#endif
+BTF_ID_FLAGS(func, scx_bpf_clock_get_ns)
BTF_KFUNCS_END(scx_kfunc_ids_any)
static const struct btf_kfunc_id_set scx_kfunc_set_any = {
--
2.47.1
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v2 4/5] sched_ext: Add scx_bpf_clock_get_ns() for BPF scheduler
2024-12-02 4:38 [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock Changwoo Min
` (2 preceding siblings ...)
2024-12-02 4:38 ` [PATCH v2 3/5] sched_ext: Implement scx_bpf_clock_get_ns() Changwoo Min
@ 2024-12-02 4:38 ` Changwoo Min
2024-12-02 4:38 ` [PATCH v2 5/5] sched_ext: Replace bpf_ktime_get_ns() to scx_bpf_clock_get_ns() Changwoo Min
4 siblings, 0 replies; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 4:38 UTC (permalink / raw)
To: tj, void, mingo, peterz; +Cc: changwoo, kernel-dev, linux-kernel
scx_bpf_clock_get_ns() is added to the header files so the BPF
scheduler can use it.
Signed-off-by: Changwoo Min <changwoo@igalia.com>
---
tools/sched_ext/include/scx/common.bpf.h | 1 +
tools/sched_ext/include/scx/compat.bpf.h | 5 +++++
2 files changed, 6 insertions(+)
diff --git a/tools/sched_ext/include/scx/common.bpf.h b/tools/sched_ext/include/scx/common.bpf.h
index 2f36b7b6418d..230c7f2e8ad6 100644
--- a/tools/sched_ext/include/scx/common.bpf.h
+++ b/tools/sched_ext/include/scx/common.bpf.h
@@ -72,6 +72,7 @@ bool scx_bpf_task_running(const struct task_struct *p) __ksym;
s32 scx_bpf_task_cpu(const struct task_struct *p) __ksym;
struct rq *scx_bpf_cpu_rq(s32 cpu) __ksym;
struct cgroup *scx_bpf_task_cgroup(struct task_struct *p) __ksym __weak;
+u64 scx_bpf_clock_get_ns(void) __ksym __weak;
/*
* Use the following as @it__iter when calling scx_bpf_dsq_move[_vtime]() from
diff --git a/tools/sched_ext/include/scx/compat.bpf.h b/tools/sched_ext/include/scx/compat.bpf.h
index d56520100a26..d295c59e3f05 100644
--- a/tools/sched_ext/include/scx/compat.bpf.h
+++ b/tools/sched_ext/include/scx/compat.bpf.h
@@ -125,6 +125,11 @@ bool scx_bpf_dispatch_vtime_from_dsq___compat(struct bpf_iter_scx_dsq *it__iter,
false; \
})
+#define scx_bpf_clock_get_ns() \
+ (bpf_ksym_exists(scx_bpf_clock_get_ns) ? \
+ scx_bpf_clock_get_ns() : \
+ bpf_ktime_get_ns())
+
/*
* Define sched_ext_ops. This may be expanded to define multiple variants for
* backward compatibility. See compat.h::SCX_OPS_LOAD/ATTACH().
--
2.47.1
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v2 5/5] sched_ext: Replace bpf_ktime_get_ns() to scx_bpf_clock_get_ns()
2024-12-02 4:38 [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock Changwoo Min
` (3 preceding siblings ...)
2024-12-02 4:38 ` [PATCH v2 4/5] sched_ext: Add scx_bpf_clock_get_ns() for BPF scheduler Changwoo Min
@ 2024-12-02 4:38 ` Changwoo Min
4 siblings, 0 replies; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 4:38 UTC (permalink / raw)
To: tj, void, mingo, peterz; +Cc: changwoo, kernel-dev, linux-kernel
In the BPF schedulers that use bpf_ktime_get_ns() -- scx_central and
scx_flatcg, replace bpf_ktime_get_ns() calls to scx_bpf_clock_get_ns().
Signed-off-by: Changwoo Min <changwoo@igalia.com>
---
tools/sched_ext/scx_central.bpf.c | 4 ++--
tools/sched_ext/scx_flatcg.bpf.c | 2 +-
2 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/tools/sched_ext/scx_central.bpf.c b/tools/sched_ext/scx_central.bpf.c
index e6fad6211f6c..cb7428b6a198 100644
--- a/tools/sched_ext/scx_central.bpf.c
+++ b/tools/sched_ext/scx_central.bpf.c
@@ -245,7 +245,7 @@ void BPF_STRUCT_OPS(central_running, struct task_struct *p)
s32 cpu = scx_bpf_task_cpu(p);
u64 *started_at = ARRAY_ELEM_PTR(cpu_started_at, cpu, nr_cpu_ids);
if (started_at)
- *started_at = bpf_ktime_get_ns() ?: 1; /* 0 indicates idle */
+ *started_at = scx_bpf_clock_get_ns() ?: 1; /* 0 indicates idle */
}
void BPF_STRUCT_OPS(central_stopping, struct task_struct *p, bool runnable)
@@ -258,7 +258,7 @@ void BPF_STRUCT_OPS(central_stopping, struct task_struct *p, bool runnable)
static int central_timerfn(void *map, int *key, struct bpf_timer *timer)
{
- u64 now = bpf_ktime_get_ns();
+ u64 now = scx_bpf_clock_get_ns();
u64 nr_to_kick = nr_queued;
s32 i, curr_cpu;
diff --git a/tools/sched_ext/scx_flatcg.bpf.c b/tools/sched_ext/scx_flatcg.bpf.c
index 4e3afcd260bf..15351bf2f053 100644
--- a/tools/sched_ext/scx_flatcg.bpf.c
+++ b/tools/sched_ext/scx_flatcg.bpf.c
@@ -734,7 +734,7 @@ void BPF_STRUCT_OPS(fcg_dispatch, s32 cpu, struct task_struct *prev)
struct fcg_cpu_ctx *cpuc;
struct fcg_cgrp_ctx *cgc;
struct cgroup *cgrp;
- u64 now = bpf_ktime_get_ns();
+ u64 now = scx_bpf_clock_get_ns();
bool picked_next = false;
cpuc = find_cpu_ctx();
--
2.47.1
^ permalink raw reply related [flat|nested] 8+ messages in thread
* Re: [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock
2024-12-02 4:38 ` [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock Changwoo Min
@ 2024-12-02 9:58 ` Peter Zijlstra
2024-12-02 14:50 ` Changwoo Min
0 siblings, 1 reply; 8+ messages in thread
From: Peter Zijlstra @ 2024-12-02 9:58 UTC (permalink / raw)
To: Changwoo Min; +Cc: tj, void, mingo, changwoo, kernel-dev, linux-kernel
On Mon, Dec 02, 2024 at 01:38:46PM +0900, Changwoo Min wrote:
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 95e40895a519..ab8015c8cab4 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -789,6 +789,7 @@ static void update_rq_clock_task(struct rq *rq, s64 delta)
> void update_rq_clock(struct rq *rq)
> {
> s64 delta;
> + u64 clock;
>
> lockdep_assert_rq_held(rq);
>
> @@ -800,11 +801,14 @@ void update_rq_clock(struct rq *rq)
> SCHED_WARN_ON(rq->clock_update_flags & RQCF_UPDATED);
> rq->clock_update_flags |= RQCF_UPDATED;
> #endif
> + clock = sched_clock_cpu(cpu_of(rq));
> + scx_rq_clock_update(rq, clock);
This adds a write to a second cacheline for *everyone* always.
Please don't do that.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock
2024-12-02 9:58 ` Peter Zijlstra
@ 2024-12-02 14:50 ` Changwoo Min
0 siblings, 0 replies; 8+ messages in thread
From: Changwoo Min @ 2024-12-02 14:50 UTC (permalink / raw)
To: Peter Zijlstra; +Cc: tj, void, mingo, kernel-dev, linux-kernel
Hello Peter,
On 24. 12. 2. 18:58, Peter Zijlstra wrote:
> On Mon, Dec 02, 2024 at 01:38:46PM +0900, Changwoo Min wrote:
>> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
>> index 95e40895a519..ab8015c8cab4 100644
>> --- a/kernel/sched/core.c
>> +++ b/kernel/sched/core.c
>> @@ -789,6 +789,7 @@ static void update_rq_clock_task(struct rq *rq, s64 delta)
>> void update_rq_clock(struct rq *rq)
>> {
>> s64 delta;
>> + u64 clock;
>>
>> lockdep_assert_rq_held(rq);
>>
>> @@ -800,11 +801,14 @@ void update_rq_clock(struct rq *rq)
>> SCHED_WARN_ON(rq->clock_update_flags & RQCF_UPDATED);
>> rq->clock_update_flags |= RQCF_UPDATED;
>> #endif
>> + clock = sched_clock_cpu(cpu_of(rq));
>> + scx_rq_clock_update(rq, clock);
>
> This adds a write to a second cacheline for *everyone* always.
>
> Please don't do that.
>
Thank you for the suggestion! I will update the flags only when
the scx is enabled by checking it inside the
scx_rq_clock_update() and scx_rq_clock_stale() functions.
Regards,
Changwoo Min
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2024-12-02 14:51 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2024-12-02 4:38 [PATCH v2 0/5] sched_ext: Support high-performance monotonically non-decreasing clock Changwoo Min
2024-12-02 4:38 ` [PATCH v2 1/5] sched_ext: Implement scx_rq_clock_update/stale() Changwoo Min
2024-12-02 4:38 ` [PATCH v2 2/5] sched_ext: Manage the validity of scx_rq_clock Changwoo Min
2024-12-02 9:58 ` Peter Zijlstra
2024-12-02 14:50 ` Changwoo Min
2024-12-02 4:38 ` [PATCH v2 3/5] sched_ext: Implement scx_bpf_clock_get_ns() Changwoo Min
2024-12-02 4:38 ` [PATCH v2 4/5] sched_ext: Add scx_bpf_clock_get_ns() for BPF scheduler Changwoo Min
2024-12-02 4:38 ` [PATCH v2 5/5] sched_ext: Replace bpf_ktime_get_ns() to scx_bpf_clock_get_ns() Changwoo Min
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox