* [PATCH 1/3] perf/core: publish the aux_event link with release semantics
2026-09-22 1:13 [PATCH 0/3] perf/core: order three publications against their readers Jaidev Shastri via B4 Relay
@ 2026-09-22 1:13 ` Jaidev Shastri via B4 Relay
2026-09-22 1:24 ` sashiko-bot
2026-09-22 1:13 ` [PATCH 2/3] perf/core: install the guest callbacks before the guest_state gate Jaidev Shastri via B4 Relay
2026-09-22 1:13 ` [PATCH 3/3] perf/core: publish perf_event_cache with release semantics Jaidev Shastri via B4 Relay
2 siblings, 1 reply; 7+ messages in thread
From: Jaidev Shastri via B4 Relay @ 2026-09-22 1:13 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
Ian Rogers, Adrian Hunter, James Clark
Cc: linux-perf-users, linux-kernel, Jaidev Shastri
From: Jaidev Shastri <jaidevshastri@vt.edu>
perf_get_aux_event() links an aux_output event to its group leader with
a plain store to event->aux_event once the leader has been validated.
perf_aux_output_begin() and the AUX sample path read the link with plain
loads, from the PMU interrupt on the CPU the event is scheduled on.
Store the link with smp_store_release() and read it with
smp_load_acquire(), so that a reader that sees the link also sees the
state of the leader it points at.
Found with MBCheck, a static herd7-based memory consistency checker.
Signed-off-by: Jaidev Shastri <jaidevshastri@vt.edu>
---
kernel/events/core.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
diff --git a/kernel/events/core.c b/kernel/events/core.c
index db7b76d6b..7cce3fc7c 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -2332,7 +2332,8 @@ static int perf_get_aux_event(struct perf_event *event,
* group in torn down, the aux_output events loose their
* link to the aux_event and can't schedule any more.
*/
- event->aux_event = group_leader;
+ /* Pairs with the smp_load_acquire() in the AUX output paths. */
+ smp_store_release(&event->aux_event, group_leader);
return 1;
}
@@ -7974,7 +7975,8 @@ static unsigned long perf_prepare_sample_aux(struct perf_event *event,
struct perf_sample_data *data,
size_t size)
{
- struct perf_event *sampler = event->aux_event;
+ /* Pairs with the smp_store_release() in perf_get_aux_event(). */
+ struct perf_event *sampler = smp_load_acquire(&event->aux_event);
struct perf_buffer *rb;
data->aux_size = 0;
@@ -8046,7 +8048,8 @@ static void perf_aux_sample_output(struct perf_event *event,
struct perf_output_handle *handle,
struct perf_sample_data *data)
{
- struct perf_event *sampler = event->aux_event;
+ /* Pairs with the smp_store_release() in perf_get_aux_event(). */
+ struct perf_event *sampler = smp_load_acquire(&event->aux_event);
struct perf_buffer *rb;
unsigned long pad;
long size;
--
2.43.0
^ permalink raw reply related [flat|nested] 7+ messages in thread* [PATCH 2/3] perf/core: install the guest callbacks before the guest_state gate
2026-09-22 1:13 [PATCH 0/3] perf/core: order three publications against their readers Jaidev Shastri via B4 Relay
2026-09-22 1:13 ` [PATCH 1/3] perf/core: publish the aux_event link with release semantics Jaidev Shastri via B4 Relay
@ 2026-09-22 1:13 ` Jaidev Shastri via B4 Relay
2026-09-22 1:25 ` sashiko-bot
2026-09-22 1:13 ` [PATCH 3/3] perf/core: publish perf_event_cache with release semantics Jaidev Shastri via B4 Relay
2 siblings, 1 reply; 7+ messages in thread
From: Jaidev Shastri via B4 Relay @ 2026-09-22 1:13 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
Ian Rogers, Adrian Hunter, James Clark
Cc: linux-perf-users, linux-kernel, Jaidev Shastri
From: Jaidev Shastri <jaidevshastri@vt.edu>
perf_register_guest_info_callbacks() updates the __perf_guest_state
static call first and __perf_guest_get_ip and the interrupt handlers
afterwards. perf_instruction_pointer() and perf_misc_flags() test
perf_guest_state() and then call perf_guest_get_ip().
Without CONFIG_HAVE_STATIC_CALL, which arm64 selects only under CFI, a
static call is a plain load of key->func. A PMI on another CPU can then
see the new state callback while get_ip still resolves to the RET0
default, and report an instruction pointer of zero for a guest sample.
Install the callees first and the gate last, separated by smp_wmb(). The
reader keeps its control dependency.
Found with MBCheck, a static herd7-based memory consistency checker.
Signed-off-by: Jaidev Shastri <jaidevshastri@vt.edu>
---
kernel/events/core.c | 13 ++++++++++++-
1 file changed, 12 insertions(+), 1 deletion(-)
diff --git a/kernel/events/core.c b/kernel/events/core.c
index 7cce3fc7c..3b5c59263 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -7760,7 +7760,14 @@ void perf_register_guest_info_callbacks(struct perf_guest_info_callbacks *cbs)
return;
rcu_assign_pointer(perf_guest_cbs, cbs);
- static_call_update(__perf_guest_state, cbs->state);
+
+ /*
+ * perf_guest_state() is the gate the PMI paths test before they call
+ * perf_guest_get_ip() and the interrupt handlers. On architectures
+ * without inline static calls the updates are plain pointer stores,
+ * so install the callees first and the gate last, behind a write
+ * barrier.
+ */
static_call_update(__perf_guest_get_ip, cbs->get_ip);
/* Implementing ->handle_intel_pt_intr is optional. */
@@ -7771,6 +7778,10 @@ void perf_register_guest_info_callbacks(struct perf_guest_info_callbacks *cbs)
if (cbs->handle_mediated_pmi)
static_call_update(__perf_guest_handle_mediated_pmi,
cbs->handle_mediated_pmi);
+
+ /* Order the callee updates above before the gate below. */
+ smp_wmb();
+ static_call_update(__perf_guest_state, cbs->state);
}
EXPORT_SYMBOL_GPL(perf_register_guest_info_callbacks);
--
2.43.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH 2/3] perf/core: install the guest callbacks before the guest_state gate
2026-09-22 1:13 ` [PATCH 2/3] perf/core: install the guest callbacks before the guest_state gate Jaidev Shastri via B4 Relay
@ 2026-09-22 1:25 ` sashiko-bot
0 siblings, 0 replies; 7+ messages in thread
From: sashiko-bot @ 2026-09-22 1:25 UTC (permalink / raw)
To: Jaidev Shastri; +Cc: linux-perf-users
> From: Jaidev Shastri <jaidevshastri@vt.edu>
>
> perf_register_guest_info_callbacks() updates the __perf_guest_state
> static call first and __perf_guest_get_ip and the interrupt handlers
> afterwards. perf_instruction_pointer() and perf_misc_flags() test
> perf_guest_state() and then call perf_guest_get_ip().
>
> Without CONFIG_HAVE_STATIC_CALL, which arm64 selects only under CFI, a
> static call is a plain load of key->func. A PMI on another CPU can then
> see the new state callback while get_ip still resolves to the RET0
> default, and report an instruction pointer of zero for a guest sample.
>
> Install the callees first and the gate last, separated by smp_wmb(). The
> reader keeps its control dependency.
>
> Found with MBCheck, a static herd7-based memory consistency checker.
>
> Signed-off-by: Jaidev Shastri <jaidevshastri@vt.edu>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921-mb-perf-core-v1-0-ac56a719c02b@vt.edu?part=2
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 3/3] perf/core: publish perf_event_cache with release semantics
2026-09-22 1:13 [PATCH 0/3] perf/core: order three publications against their readers Jaidev Shastri via B4 Relay
2026-09-22 1:13 ` [PATCH 1/3] perf/core: publish the aux_event link with release semantics Jaidev Shastri via B4 Relay
2026-09-22 1:13 ` [PATCH 2/3] perf/core: install the guest callbacks before the guest_state gate Jaidev Shastri via B4 Relay
@ 2026-09-22 1:13 ` Jaidev Shastri via B4 Relay
2026-09-22 1:24 ` sashiko-bot
2 siblings, 1 reply; 7+ messages in thread
From: Jaidev Shastri via B4 Relay @ 2026-09-22 1:13 UTC (permalink / raw)
To: Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo,
Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa,
Ian Rogers, Adrian Hunter, James Clark
Cc: linux-perf-users, linux-kernel, Jaidev Shastri
From: Jaidev Shastri <jaidevshastri@vt.edu>
perf_event_init() creates perf_event_cache with a plain store, after the
pmu_idr has been set up. free_event_rcu() reads the pointer with a plain
load from an RCU callback on any CPU.
Store it with smp_store_release() and read it with smp_load_acquire().
Found with MBCheck, a static herd7-based memory consistency checker.
Signed-off-by: Jaidev Shastri <jaidevshastri@vt.edu>
---
kernel/events/core.c | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/kernel/events/core.c b/kernel/events/core.c
index 3b5c59263..74cfb13fe 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -5313,7 +5313,8 @@ static void free_event_rcu(struct rcu_head *head)
put_pid_ns(event->ns);
perf_event_free_filter(event);
kfree(event->addr_filter_ranges);
- kmem_cache_free(perf_event_cache, event);
+ /* Pairs with the smp_store_release() in perf_event_init(). */
+ kmem_cache_free(smp_load_acquire(&perf_event_cache), event);
}
static void ring_buffer_attach(struct perf_event *event,
@@ -15404,7 +15405,8 @@ void __init perf_event_init(void)
ret = init_hw_breakpoint();
WARN(ret, "hw_breakpoint initialization failed with: %d", ret);
- perf_event_cache = KMEM_CACHE(perf_event, SLAB_PANIC);
+ /* Pairs with the smp_load_acquire() in free_event_rcu(). */
+ smp_store_release(&perf_event_cache, KMEM_CACHE(perf_event, SLAB_PANIC));
/*
* Build time assertion that we keep the data_head at the intended
--
2.43.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH 3/3] perf/core: publish perf_event_cache with release semantics
2026-09-22 1:13 ` [PATCH 3/3] perf/core: publish perf_event_cache with release semantics Jaidev Shastri via B4 Relay
@ 2026-09-22 1:24 ` sashiko-bot
0 siblings, 0 replies; 7+ messages in thread
From: sashiko-bot @ 2026-09-22 1:24 UTC (permalink / raw)
To: Jaidev Shastri; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] perf: unnecessary and inconsistent memory barriers for boot-initialized variable
--
commit 502c74ee5072b4297e79a78cf708d77e31d6b0e5
Author: Jaidev Shastri <jaidevshastri@vt.edu>
Subject: perf/core: publish perf_event_cache with release semantics
This patch replaces plain reads and writes of perf_event_cache with
smp_store_release() and smp_load_acquire() to address a static checker
warning about memory consistency.
> diff --git a/kernel/events/core.c b/kernel/events/core.c
> index 3b5c592630e17..74cfb13fec7c0 100644
> --- a/kernel/events/core.c
> +++ b/kernel/events/core.c
> @@ -5313,7 +5313,8 @@ static void free_event_rcu(struct rcu_head *head)
> put_pid_ns(event->ns);
> perf_event_free_filter(event);
> kfree(event->addr_filter_ranges);
> - kmem_cache_free(perf_event_cache, event);
> + /* Pairs with the smp_store_release() in perf_event_init(). */
> + kmem_cache_free(smp_load_acquire(&perf_event_cache), event);
[Severity: Low]
Is this acquire barrier necessary? Since perf_event_cache is initialized
synchronously during early kernel boot, it should already be causally
guaranteed to be fully visible across all CPUs by the time any tasks are
spawned to interact with it.
Additionally, if a memory barrier were required to prevent data races, is
it problematic that the allocation path still uses a plain load?
kernel/events/core.c:perf_event_alloc() {
...
event = kmem_cache_alloc_node(perf_event_cache, GFP_KERNEL | __GFP_ZERO, node);
...
}
[ ... ]
> @@ -15404,7 +15405,8 @@ void __init perf_event_init(void)
> ret = init_hw_breakpoint();
> WARN(ret, "hw_breakpoint initialization failed with: %d", ret);
>
> - perf_event_cache = KMEM_CACHE(perf_event, SLAB_PANIC);
> + /* Pairs with the smp_load_acquire() in free_event_rcu(). */
> + smp_store_release(&perf_event_cache, KMEM_CACHE(perf_event, SLAB_PANIC));
[Severity: Low]
Since perf_event_init() is an __init function that runs synchronously during
early system boot before tasks could potentially allocate or free perf events,
does this require a release barrier?
Instead of adding synchronization overhead to silence the static checker,
would it be more appropriate to annotate the declaration of perf_event_cache
with __ro_after_init? This documents the initialization phase ordering and
safely silences static checkers while avoiding unnecessary code complexity.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921-mb-perf-core-v1-0-ac56a719c02b@vt.edu?part=3
^ permalink raw reply [flat|nested] 7+ messages in thread