From: Ankur Arora <ankur.a.arora@oracle.com>
To: linux-kernel@vger.kernel.org, linux-arch@vger.kernel.org,
linux-arm-kernel@lists.infradead.org, linux-pm@vger.kernel.org,
bpf@vger.kernel.org
Cc: arnd@arndb.de, catalin.marinas@arm.com, will@kernel.org,
peterz@infradead.org, akpm@linux-foundation.org,
mark.rutland@arm.com, harisokn@amazon.com, cl@gentwo.org,
ast@kernel.org, rafael@kernel.org, daniel.lezcano@linaro.org,
memxor@gmail.com, zhenglifeng1@huawei.com,
xueshuai@linux.alibaba.com, rdunlap@infradead.org,
david.laight.linux@gmail.com, broonie@kernel.org,
joao.m.martins@oracle.com, boris.ostrovsky@oracle.com,
konrad.wilk@oracle.com, ashok.bhat@arm.com,
Ankur Arora <ankur.a.arora@oracle.com>
Subject: [PATCH v15 01/16] asm-generic: barrier: Add smp_cond_load_relaxed_timeout()
Date: Mon, 31 Aug 2026 13:22:36 -0700 [thread overview]
Message-ID: <20260831202251.305046-2-ankur.a.arora@oracle.com> (raw)
In-Reply-To: <20260831202251.305046-1-ankur.a.arora@oracle.com>
Add smp_cond_load_relaxed_timeout(), which extends
smp_cond_load_relaxed() to allow waiting for a duration.
The interface loops around waiting for the condition variable to change
while peridically doing a time-check. It uses cpu_poll_relax() to slow
down the busy-wait, which, unless overridden by the architecture code,
amounts to a cpu_relax().
There are two ways for the time-check to fail: the timeout case or,
@time_expr_ns returning an invalid value (negative or zero). The second
failure mode allows for clocks attached to the clock-domain of
@cond_expr -- clocks which might cease to operate meaningfully once
some state internal to @cond_expr has changed -- to fail.
Evaluation of @time_expr_ns: in the fastpath we want to keep the
performance close to smp_cond_load_relaxed(). So defer evaluation
of the potentially costly @time_expr_ns to the slowpath.
This also means that there will always be some hardware dependent
duration that has passed in cpu_poll_relax() iterations at the time
of first evaluation. Additionally cpu_poll_relax() is not guaranteed
to return at timeout boundary. In sum, expect timeout overshoot when
we exit due to expiration of the timeout.
The number of spin iterations before time-check, SMP_TIMEOUT_POLL_COUNT
is chosen to be 200 by default. With a cpu_poll_relax() iteration
taking ~20-30 cycles (measured on a variety of x86 platforms), we
expect a time-check every ~4000-6000 cycles.
If a architecture provides a waiting implementation for cpu_poll_relax()
(and signifies that by defining CPU_POLL_RELAX_WAITS) we define
SMP_TIMEOUT_POLL_COUNT to 1.
Lastly, config option ARCH_HAS_CPU_RELAX indicates availability of a
cpu_poll_relax() that is cheaper than polling. Long timeout values
might not make sense for architectures not having ARCH_HAS_CPU_RELAX.
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Will Deacon <will@kernel.org>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: linux-arch@vger.kernel.org
Signed-off-by: Ankur Arora <ankur.a.arora@oracle.com>
---
Notes:
- deferred time-check is now limited to a single usec. This allows the
fastpath to not pay the time check penalty while ensuring that
architecturs with precise waits (like WFET) don't end up with a
huge overshoot.
- smp_cond_load_relaxed_timeout() now checks that timeout_ns fits in
s64. (This is kept as a wrapper around __smp_cond_load_relaxed_timeout()
to avoid cluttering the core logic.)
- SMP_TIMEOUT_POLL_COUNT is explicitly defined to 1 if the
architecture provides a waiting implementation of cpu_poll_relax()
(defines CPU_POLL_RELAX_WAITS)
- this was the only sane behaviour from arch code but now the code
explicitly encodes it.
- __smp_cond_load_relaxed_timeout now uses prefixed names to avoid
any possibility of collision.
- minor language changes in the commit message; comment block around
cpu_poll_relax() and CPU_POLL_RELAX_WAITS
Catalin: the changes are a significant number of lines of code.
Though IMO only minor updates to the logic. But I didn't want to stretch
your R-by too far. Could you take another look?
include/asm-generic/barrier.h | 106 ++++++++++++++++++++++++++++++++++
1 file changed, 106 insertions(+)
diff --git a/include/asm-generic/barrier.h b/include/asm-generic/barrier.h
index b99cb57dfccc..4437d27c46b9 100644
--- a/include/asm-generic/barrier.h
+++ b/include/asm-generic/barrier.h
@@ -273,6 +273,112 @@ do { \
})
#endif
+/*
+ * Number of times we iterate in the loop before doing the time check.
+ */
+#ifndef SMP_TIMEOUT_POLL_COUNT
+#ifdef CPU_POLL_RELAX_WAITS
+#define SMP_TIMEOUT_POLL_COUNT 1 /* Wait mode. No need to poll. */
+#else
+/*
+ * Assume that cpu_poll_relax() provides a small blip in the pipeline.
+ * Combine a reasonable number of cpu_poll_relax() instances before
+ * doing anything substantial like a time-check.
+ * Note that this assumes that the rest of the loop (largely evaluation
+ * of the loop condition is relatively cheap.)
+ */
+#define SMP_TIMEOUT_POLL_COUNT 200
+#endif
+#endif
+
+/*
+ * cpu_poll_relax() stitches up two kinds of primitives: ones that provide
+ * a momentary blip in the pipeline (ex. cpu_relax() on x86), or ones that
+ * support waiting for @ptr value to change, coupled with a precise (or not)
+ * timeout.
+ *
+ * We keep both together because the objective is to minimize expensive
+ * operations while polling on @ptr waiting for it to change. Either
+ * version allows for that.
+ * The arguments (@ptr, @val, @timeout_ns) are only needed for waiting
+ * implementations.
+ *
+ * Note that platforms with a suitable cpu_poll_relax() implementation are
+ * expected to define ARCH_HAS_CPU_RELAX.
+ */
+#ifndef cpu_poll_relax
+#define cpu_poll_relax(ptr, val, timeout_ns) cpu_relax()
+#endif
+
+/**
+ * smp_cond_load_relaxed_timeout() - (Spin) wait for cond with no ordering
+ * guarantees until a timeout expires.
+ * @ptr: pointer to the variable to wait on.
+ * @cond_expr: boolean expression to wait for.
+ * @time_expr_ns: expression that evaluates to monotonic time (in ns) or,
+ * on failure, returns zero or a negative value.
+ * @timeout_ns: timeout value in ns
+ * Both of the above are expected to be compatible with s64; the signed
+ * value is used to handle the failure case in @time_expr_ns.
+ *
+ * Equivalent to using READ_ONCE() on the condition variable.
+ *
+ * Callers that expect to wait for prolonged durations might want
+ * to take into account the availability of ARCH_HAS_CPU_RELAX.
+ *
+ * Note that @ptr is expected to point to a memory address. Using this
+ * interface with MMIO will be slower (since SMP_TIMEOUT_POLL_COUNT is
+ * tuned for memory) and might also break in interesting architecture
+ * dependent ways.
+ */
+#ifndef smp_cond_load_relaxed_timeout
+#define __smp_cond_load_relaxed_timeout(ptr, cond_expr, \
+ time_expr_ns, timeout_ns) \
+({ \
+ typeof(ptr) __PTR = (ptr); \
+ __unqual_scalar_typeof(*(ptr)) VAL; \
+ u32 __scl_count = 0, __scl_spin = SMP_TIMEOUT_POLL_COUNT; \
+ s64 __scl_timeout = NSEC_PER_USEC; \
+ s64 __scl_time_now, __scl_time_end = 0; \
+ \
+ for (;;) { \
+ VAL = READ_ONCE(*__PTR); \
+ if (cond_expr) \
+ break; \
+ cpu_poll_relax(__PTR, VAL, (u64)__scl_timeout); \
+ if (++__scl_count < __scl_spin) \
+ continue; \
+ __scl_time_now = (s64)(time_expr_ns); \
+ if (unlikely(__scl_time_end == 0)) { \
+ __scl_timeout = (s64)(timeout_ns); \
+ __scl_time_end = __scl_time_now + __scl_timeout;\
+ } \
+ __scl_timeout = __scl_time_end - __scl_time_now; \
+ if (__scl_time_now <= 0 || __scl_timeout <= 0) { \
+ VAL = READ_ONCE(*__PTR); \
+ break; \
+ } \
+ __scl_count = 0; \
+ } \
+ (typeof(*(ptr)))VAL; \
+})
+
+#define smp_cond_load_relaxed_timeout(ptr, cond_expr, \
+ time_expr_ns, timeout_ns) \
+({ \
+ __unqual_scalar_typeof(*(ptr)) VAL; \
+ s64 __scl_timeout_ns = (s64)(timeout_ns); \
+ \
+ if (__scl_timeout_ns < 0) \
+ VAL = READ_ONCE(*(ptr)); \
+ else \
+ VAL = __smp_cond_load_relaxed_timeout(ptr, cond_expr, \
+ time_expr_ns, \
+ __scl_timeout_ns);\
+ (typeof(*(ptr)))VAL; \
+})
+#endif
+
/*
* pmem_wmb() ensures that all stores for which the modification
* are written to persistent storage by preceding instructions have
--
2.43.7
next prev parent reply other threads:[~2026-08-31 20:24 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 20:22 [PATCH v15 00/16] barrier: Add smp_cond_load_{relaxed,acquire}_timeout() Ankur Arora
2026-08-31 20:22 ` Ankur Arora [this message]
2026-08-31 20:22 ` [PATCH v15 02/16] arm64: barrier: Support smp_cond_load_relaxed_timeout() Ankur Arora
2026-08-31 20:22 ` [PATCH v15 03/16] arm64/delay: move, fixup usecs_to_cycles() Ankur Arora
2026-08-31 20:22 ` [PATCH v15 04/16] arm64: support WFET in smp_cond_load_relaxed_timeout() Ankur Arora
2026-08-31 21:16 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 05/16] arm64: rqspinlock: Remove private copy of smp_cond_load_acquire_timewait() Ankur Arora
2026-08-31 20:22 ` [PATCH v15 06/16] asm-generic: barrier: Add smp_cond_load_acquire_timeout() Ankur Arora
2026-08-31 21:17 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 07/16] atomic: Add atomic_cond_read_*_timeout() Ankur Arora
2026-08-31 21:16 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 08/16] locking/atomic: scripts: build atomic_long_cond_read_*_timeout() Ankur Arora
2026-08-31 20:22 ` [PATCH v15 09/16] bpf/rqspinlock: switch check_timeout() to a clock interface Ankur Arora
2026-08-31 21:16 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 10/16] bpf/rqspinlock: Use smp_cond_load_acquire_timeout() Ankur Arora
2026-08-31 21:31 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 11/16] sched: add need-resched timed wait interface Ankur Arora
2026-08-31 20:22 ` [PATCH v15 12/16] cpuidle/poll_state: Wait for need-resched via tif_need_resched_relaxed_wait() Ankur Arora
2026-08-31 20:22 ` [PATCH v15 13/16] arm64/delay: enable testing smp_cond_load_relaxed_timeout() Ankur Arora
2026-08-31 21:16 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 14/16] barrier: add tests for smp_cond_load_*_timeout() Ankur Arora
2026-08-31 21:17 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 15/16] barrier: timeout validity checks for smp_cond_load_relaxed_timeout() Ankur Arora
2026-08-31 21:16 ` bot+bpf-ci
2026-08-31 20:22 ` [PATCH v15 16/16] barrier: timeout validity checks for smp_cond_load_acquire_timeout() Ankur Arora
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260831202251.305046-2-ankur.a.arora@oracle.com \
--to=ankur.a.arora@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=arnd@arndb.de \
--cc=ashok.bhat@arm.com \
--cc=ast@kernel.org \
--cc=boris.ostrovsky@oracle.com \
--cc=bpf@vger.kernel.org \
--cc=broonie@kernel.org \
--cc=catalin.marinas@arm.com \
--cc=cl@gentwo.org \
--cc=daniel.lezcano@linaro.org \
--cc=david.laight.linux@gmail.com \
--cc=harisokn@amazon.com \
--cc=joao.m.martins@oracle.com \
--cc=konrad.wilk@oracle.com \
--cc=linux-arch@vger.kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pm@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=memxor@gmail.com \
--cc=peterz@infradead.org \
--cc=rafael@kernel.org \
--cc=rdunlap@infradead.org \
--cc=will@kernel.org \
--cc=xueshuai@linux.alibaba.com \
--cc=zhenglifeng1@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox