Linux Trace Kernel
 help / color / mirror / Atom feed
* [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields
@ 2026-08-10  6:32 Michael Wu
  2026-08-10 14:45 ` Steven Rostedt
  0 siblings, 1 reply; 6+ messages in thread
From: Michael Wu @ 2026-08-10  6:32 UTC (permalink / raw)
  To: Steven Rostedt, Masami Hiramatsu
  Cc: Mathieu Desnoyers, linux-kernel, linux-trace-kernel

The following sequence may leads race between event_define_fields()
and update_event_fields():
  CPU0 (module A, pri=1 notifier)        CPU1 (module B, pri=0 notifier)
  ===============================        ===============================
  event_define_fields(call_A)            trace_event_update_all()
    for each f:                            list_for_each_entry(...,
      list_add(&f->link,                                &ftrace_events)
               &class->fields)               -> finds call_A
        f->link.next = next;        (2)
                                           update_event_fields(call_A)
        WRITE_ONCE(class->fields->next,
                   &f->link);       (4)
                                             list_for_each_entry(field,
                                               &class->fields, link)
                                             -> field = class->fields->next
                                                  = &f->link
                                                  = f (offset 0)
                                             -> arm64 weak ordering:
                                                (4) visible before (2)
                                                field->link.next == 0
                                             -> next iteration:
                                                field = (void *)0 = NULL
                                             -> crash at NULL->type (0x18)

This produces the following panic:
   Unable to handle kernel access ... at virtual address 0000000000000018
   pc : update_event_fields+0xf8/0x368
   Call trace:
    update_event_fields+0xf8/0x368
    trace_event_update_all+0x7c/0x2b4
    trace_module_notify+0x4c/0x1dc
    notifier_call_chain+0x84/0x168
    blocking_notifier_call_chain_robust+0x64/0xd4
    load_module+0x10c8/0x123c
    __arm64_sys_finit_module+0x230/0x31c

Fix by taking event_mutex in trace_event_update_all() before
trace_event_sem.

Fixes: b3bc8547d3be ("tracing: Have TRACE_DEFINE_ENUM affect trace event types as well")
Cc: stable@vger.kernel.org
Signed-off-by: Michael Wu <michael@allwinnertech.com>
---
 kernel/trace/trace_events.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/kernel/trace/trace_events.c b/kernel/trace/trace_events.c
index 956692856fa8..9632788da5af 100644
--- a/kernel/trace/trace_events.c
+++ b/kernel/trace/trace_events.c
@@ -3566,6 +3566,7 @@ void trace_event_update_all(struct trace_eval_map **map, int len)
 	int last_i;
 	int i;
 
+	mutex_lock(&event_mutex);
 	down_write(&trace_event_sem);
 	list_for_each_entry_safe(call, p, &ftrace_events, list) {
 		/* events are usually grouped together with systems */
@@ -3604,6 +3605,7 @@ void trace_event_update_all(struct trace_eval_map **map, int len)
 		cond_resched();
 	}
 	up_write(&trace_event_sem);
+	mutex_unlock(&event_mutex);
 }
 
 static bool event_in_systems(struct trace_event_call *call,
-- 
2.29.0

^ permalink raw reply related	[flat|nested] 6+ messages in thread

* Re: [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields
  2026-08-10  6:32 [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields Michael Wu
@ 2026-08-10 14:45 ` Steven Rostedt
  2026-08-11  6:00   ` Michael Wu
  0 siblings, 1 reply; 6+ messages in thread
From: Steven Rostedt @ 2026-08-10 14:45 UTC (permalink / raw)
  To: Michael Wu
  Cc: Masami Hiramatsu, Mathieu Desnoyers, linux-kernel,
	linux-trace-kernel

On Mon, 10 Aug 2026 14:32:30 +0800
Michael Wu <michael@allwinnertech.com> wrote:

> The following sequence may leads race between event_define_fields()
> and update_event_fields():

>   CPU0 (module A, pri=1 notifier)        CPU1 (module B, pri=0 notifier)

What does the above mean? Are you loading two modules at the same time?
What does "pri=X notifier" mean? What function calls are these coming from?

-- Steve

>   ===============================        ===============================
>   event_define_fields(call_A)            trace_event_update_all()
>     for each f:                            list_for_each_entry(...,
>       list_add(&f->link,                                &ftrace_events)
>                &class->fields)               -> finds call_A
>         f->link.next = next;        (2)
>                                            update_event_fields(call_A)
>         WRITE_ONCE(class->fields->next,
>                    &f->link);       (4)
>                                              list_for_each_entry(field,
>                                                &class->fields, link)
>                                              -> field = class->fields->next  
>                                                   = &f->link
>                                                   = f (offset 0)
>                                              -> arm64 weak ordering:  
>                                                 (4) visible before (2)
>                                                 field->link.next == 0
>                                              -> next iteration:  
>                                                 field = (void *)0 = NULL
>                                              -> crash at NULL->type (0x18)  
> 
> This produces the following panic:
>    Unable to handle kernel access ... at virtual address 0000000000000018
>    pc : update_event_fields+0xf8/0x368
>    Call trace:
>     update_event_fields+0xf8/0x368
>     trace_event_update_all+0x7c/0x2b4
>     trace_module_notify+0x4c/0x1dc
>     notifier_call_chain+0x84/0x168
>     blocking_notifier_call_chain_robust+0x64/0xd4
>     load_module+0x10c8/0x123c
>     __arm64_sys_finit_module+0x230/0x31c
> 
> Fix by taking event_mutex in trace_event_update_all() before
> trace_event_sem.
> 
> Fixes: b3bc8547d3be ("tracing: Have TRACE_DEFINE_ENUM affect trace event types as well")
> Cc: stable@vger.kernel.org
> Signed-off-by: Michael Wu <michael@allwinnertech.com>

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields
  2026-08-10 14:45 ` Steven Rostedt
@ 2026-08-11  6:00   ` Michael Wu
  2026-08-11 13:00     ` Steven Rostedt
  0 siblings, 1 reply; 6+ messages in thread
From: Michael Wu @ 2026-08-11  6:00 UTC (permalink / raw)
  To: Steven Rostedt
  Cc: Masami Hiramatsu, Mathieu Desnoyers, linux-kernel,
	linux-trace-kernel

> What does the above mean? Are you loading two modules at the same time?
Two modules (A and B) are loaded simultaneously on different CPUs. On the arm64,
 when CPU0's trace_module_notify [pri=1] and CPU1's trace_module_notify [pri=0] 
simultaneously perform operations on call_A, because they are in different cache lines,
CPU1 may observe WRITE_ONCE(head->next, &f->link) in step (4) before f->link.next=next in step (2).
 At this time, CPU1 reads an uninitialized f->link.next and performs an operation that causes to crash.

> What does "pri=X notifier" mean? What function calls are these coming from?
`pri=X notifier` represents `trace_events.c:trace_module_notify [pri=1]` and `trace.c:trace_module_notify [pri=0]`, respectively.
 CPU0 (loads module A)                      CPU1 (loads module B)
  ===============================            ===============================
  load_module(A)                             load_module(B)
    blocking_notifier_call_chain_robust        blocking_notifier_call_chain_robust
         notifier_call_chain                     notifier_call_chain
          nb = trace_events.c:                     nb = trace.c:
          trace_module_notify [pri=1]              trace_module_notify [pri=0]
            mutex_lock(&event_mutex)                 trace_event_update_all() 
            trace_module_add_events(A)               down_write(&trace_event_sem) 
              __register_event(call_A)                
              __add_event_to_tracers(call_A)           
                event_define_fields(call_A)            
                  for each f:                          
                    f = kmem_cache_alloc()             
                    list_add(&f->link,               	
                             &class->fields)           
                      f->link.next=next; (2)       	   
                      WRITE_ONCE(head->next,        
                                &f->link); (4)       update_event_fields(call_A)   											           
            mutex_unlock(&event_mutex)	               list_for_each_entry(field, 
                                                         &class->fields, link) 
                                                         field = class->fields->next 
                                                               = &f->link 
                                                       	       = f (offset 0) 
                                                      up_write(&trace_event_sem)  
On 8/10/2026 10:45 PM, Steven Rostedt wrote:
> On Mon, 10 Aug 2026 14:32:30 +0800
> Michael Wu <michael@allwinnertech.com> wrote:
> 
>> The following sequence may leads race between event_define_fields()
>> and update_event_fields():
> 
>>   CPU0 (module A, pri=1 notifier)        CPU1 (module B, pri=0 notifier)
> 
> What does the above mean? Are you loading two modules at the same time?
> What does "pri=X notifier" mean? What function calls are these coming from?
> 
> -- Steve
> 
>>   ===============================        ===============================
>>   event_define_fields(call_A)            trace_event_update_all()
>>     for each f:                            list_for_each_entry(...,
>>       list_add(&f->link,                                &ftrace_events)
>>                &class->fields)               -> finds call_A
>>         f->link.next = next;        (2)
>>                                            update_event_fields(call_A)
>>         WRITE_ONCE(class->fields->next,
>>                    &f->link);       (4)
>>                                              list_for_each_entry(field,
>>                                                &class->fields, link)
>>                                              -> field = class->fields->next  
>>                                                   = &f->link
>>                                                   = f (offset 0)
>>                                              -> arm64 weak ordering:  
>>                                                 (4) visible before (2)
>>                                                 field->link.next == 0
>>                                              -> next iteration:  
>>                                                 field = (void *)0 = NULL
>>                                              -> crash at NULL->type (0x18)  
>>
>> This produces the following panic:
>>    Unable to handle kernel access ... at virtual address 0000000000000018
>>    pc : update_event_fields+0xf8/0x368
>>    Call trace:
>>     update_event_fields+0xf8/0x368
>>     trace_event_update_all+0x7c/0x2b4
>>     trace_module_notify+0x4c/0x1dc
>>     notifier_call_chain+0x84/0x168
>>     blocking_notifier_call_chain_robust+0x64/0xd4
>>     load_module+0x10c8/0x123c
>>     __arm64_sys_finit_module+0x230/0x31c
>>
>> Fix by taking event_mutex in trace_event_update_all() before
>> trace_event_sem.
>>
>> Fixes: b3bc8547d3be ("tracing: Have TRACE_DEFINE_ENUM affect trace event types as well")
>> Cc: stable@vger.kernel.org
>> Signed-off-by: Michael Wu <michael@allwinnertech.com>


-- 
Regards,
Michael Wu

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields
  2026-08-11  6:00   ` Michael Wu
@ 2026-08-11 13:00     ` Steven Rostedt
  2026-08-12  2:14       ` Michael Wu
  0 siblings, 1 reply; 6+ messages in thread
From: Steven Rostedt @ 2026-08-11 13:00 UTC (permalink / raw)
  To: Michael Wu
  Cc: Masami Hiramatsu, Mathieu Desnoyers, linux-kernel,
	linux-trace-kernel

On Tue, 11 Aug 2026 14:00:05 +0800
Michael Wu <michael@allwinnertech.com> wrote:

> > What does the above mean? Are you loading two modules at the same time?  
> Two modules (A and B) are loaded simultaneously on different CPUs. On the arm64,
>  when CPU0's trace_module_notify [pri=1] and CPU1's trace_module_notify [pri=0] 
> simultaneously perform operations on call_A, because they are in different cache lines,
> CPU1 may observe WRITE_ONCE(head->next, &f->link) in step (4) before f->link.next=next in step (2).
>  At this time, CPU1 reads an uninitialized f->link.next and performs an operation that causes to crash.

This is still way too verbose. Is this AI written? If so, AI is *not* your friend.


> 
> > What does "pri=X notifier" mean? What function calls are these coming from?  
> `pri=X notifier` represents `trace_events.c:trace_module_notify [pri=1]` and `trace.c:trace_module_notify [pri=0]`, respectively.

Why are the priorities of the notifiers important here?

I honestly didn't know one was allowed to load two modules at the same time
and thought that it the module logic would prevent that. But if that's not
the case, then yeah, we need protection.


>  CPU0 (loads module A)                      CPU1 (loads module B)
>   ===============================            ===============================
>   load_module(A)                             load_module(B)
>     blocking_notifier_call_chain_robust        blocking_notifier_call_chain_robust
>          notifier_call_chain                     notifier_call_chain
>           nb = trace_events.c:                     nb = trace.c:
>           trace_module_notify [pri=1]              trace_module_notify [pri=0]
>             mutex_lock(&event_mutex)                 trace_event_update_all() 
>             trace_module_add_events(A)               down_write(&trace_event_sem) 
>               __register_event(call_A)                
>               __add_event_to_tracers(call_A)           
>                 event_define_fields(call_A)            
>                   for each f:                          
>                     f = kmem_cache_alloc()             
>                     list_add(&f->link,               	
>                              &class->fields)           
>                       f->link.next=next; (2)       	   
>                       WRITE_ONCE(head->next,        
>                                 &f->link); (4)       update_event_fields(call_A)   											           
>             mutex_unlock(&event_mutex)	               list_for_each_entry(field, 
>                                                          &class->fields, link) 
>                                                          field = class->fields->next 
>                                                                = &f->link 
>                                                        	       = f (offset 0) 
>                                                       up_write(&trace_event_sem)  

Basically this can be summed up to being:

 CPU0 (loads module A)                      CPU1 (loads module B)
 ===============================            ===============================
 load_module(A)                             load_module(B)
   notifier_call_chain                        notifier_call_chain
     trace_module_notify                        trace_module_notify
       mutex_lock(&event_mutex)                   trace_event_update_all() 
         trace_module_add_events(A)                 down_write(&trace_event_sem) 
            __register_event(call_A)                
              __add_event_to_tracers(call_A)           
                event_define_fields(call_A)            
                  for each f:                         list_for_each_entry(field, 
                    list_add(&f->link,                                    &class->fields, link) 
                             &class->fields)            field = class->fields->next;

Where you can see that one is being read while the other is being written
to. You do not need to go into details of the cache visibility here because
this is an obvious race condition. All that information just distracts from
the real issue that is being fixed.

Less is more when it comes to describing a bug.

I'll rewrite you change log and take the patch.

Thanks,

-- Steve

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields
  2026-08-11 13:00     ` Steven Rostedt
@ 2026-08-12  2:14       ` Michael Wu
  2026-08-12 13:51         ` Steven Rostedt
  0 siblings, 1 reply; 6+ messages in thread
From: Michael Wu @ 2026-08-12  2:14 UTC (permalink / raw)
  To: Steven Rostedt
  Cc: Masami Hiramatsu, Mathieu Desnoyers, linux-kernel,
	linux-trace-kernel

On 8/11/2026 9:00 PM, Steven Rostedt wrote:
> This is still way too verbose. Is this AI written? If so, AI is *not* your friend.
This commit was written by me, not by AI. I apologize for wasting your time. Thank you for rewriting the change.
> Why are the priorities of the notifiers important here?
I probably wanted to describe the entire process, so I've gone on to say too much.
> Less is more when it comes to describing a bug.
Got it.

> On Tue, 11 Aug 2026 14:00:05 +0800
> Michael Wu <michael@allwinnertech.com> wrote:
> 
>>> What does the above mean? Are you loading two modules at the same time?  
>> Two modules (A and B) are loaded simultaneously on different CPUs. On the arm64,
>>  when CPU0's trace_module_notify [pri=1] and CPU1's trace_module_notify [pri=0] 
>> simultaneously perform operations on call_A, because they are in different cache lines,
>> CPU1 may observe WRITE_ONCE(head->next, &f->link) in step (4) before f->link.next=next in step (2).
>>  At this time, CPU1 reads an uninitialized f->link.next and performs an operation that causes to crash.
> 
> This is still way too verbose. Is this AI written? If so, AI is *not* your friend.
> 
> 
>>
>>> What does "pri=X notifier" mean? What function calls are these coming from?  
>> `pri=X notifier` represents `trace_events.c:trace_module_notify [pri=1]` and `trace.c:trace_module_notify [pri=0]`, respectively.
> 
> Why are the priorities of the notifiers important here?
> 
> I honestly didn't know one was allowed to load two modules at the same time
> and thought that it the module logic would prevent that. But if that's not
> the case, then yeah, we need protection.
> 
> 
>>  CPU0 (loads module A)                      CPU1 (loads module B)
>>   ===============================            ===============================
>>   load_module(A)                             load_module(B)
>>     blocking_notifier_call_chain_robust        blocking_notifier_call_chain_robust
>>          notifier_call_chain                     notifier_call_chain
>>           nb = trace_events.c:                     nb = trace.c:
>>           trace_module_notify [pri=1]              trace_module_notify [pri=0]
>>             mutex_lock(&event_mutex)                 trace_event_update_all() 
>>             trace_module_add_events(A)               down_write(&trace_event_sem) 
>>               __register_event(call_A)                
>>               __add_event_to_tracers(call_A)           
>>                 event_define_fields(call_A)            
>>                   for each f:                          
>>                     f = kmem_cache_alloc()             
>>                     list_add(&f->link,               	
>>                              &class->fields)           
>>                       f->link.next=next; (2)       	   
>>                       WRITE_ONCE(head->next,        
>>                                 &f->link); (4)       update_event_fields(call_A)   											           
>>             mutex_unlock(&event_mutex)	               list_for_each_entry(field, 
>>                                                          &class->fields, link) 
>>                                                          field = class->fields->next 
>>                                                                = &f->link 
>>                                                        	       = f (offset 0) 
>>                                                       up_write(&trace_event_sem)  
> 
> Basically this can be summed up to being:
> 
>  CPU0 (loads module A)                      CPU1 (loads module B)
>  ===============================            ===============================
>  load_module(A)                             load_module(B)
>    notifier_call_chain                        notifier_call_chain
>      trace_module_notify                        trace_module_notify
>        mutex_lock(&event_mutex)                   trace_event_update_all() 
>          trace_module_add_events(A)                 down_write(&trace_event_sem) 
>             __register_event(call_A)                
>               __add_event_to_tracers(call_A)           
>                 event_define_fields(call_A)            
>                   for each f:                         list_for_each_entry(field, 
>                     list_add(&f->link,                                    &class->fields, link) 
>                              &class->fields)            field = class->fields->next;
> 
> Where you can see that one is being read while the other is being written
> to. You do not need to go into details of the cache visibility here because
> this is an obvious race condition. All that information just distracts from
> the real issue that is being fixed.
> 
> Less is more when it comes to describing a bug.
> 
> I'll rewrite you change log and take the patch.
> 
> Thanks,
> 
> -- Steve


-- 
Regards,
Michael Wu

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields
  2026-08-12  2:14       ` Michael Wu
@ 2026-08-12 13:51         ` Steven Rostedt
  0 siblings, 0 replies; 6+ messages in thread
From: Steven Rostedt @ 2026-08-12 13:51 UTC (permalink / raw)
  To: Michael Wu
  Cc: Masami Hiramatsu, Mathieu Desnoyers, linux-kernel,
	linux-trace-kernel

On Wed, 12 Aug 2026 10:14:24 +0800
Michael Wu <michael@allwinnertech.com> wrote:

> On 8/11/2026 9:00 PM, Steven Rostedt wrote:
> > This is still way too verbose. Is this AI written? If so, AI is *not* your friend.  

> This commit was written by me, not by AI. I apologize for wasting your time. Thank you for rewriting the change.

Heh, AI tends to be very verbose, which is why I was thinking it was AI.

> > Why are the priorities of the notifiers important here?  

> I probably wanted to describe the entire process, so I've gone on to say too much.

Yeah, adding more than what is needed distracts from the issue.

> > Less is more when it comes to describing a bug.  

> Got it.

Great. And thanks for the fix.

-- Steve

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-08-12 13:50 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-10  6:32 [PATCH v5] tracing: Fix race between update_event_fields and, event_define_fields Michael Wu
2026-08-10 14:45 ` Steven Rostedt
2026-08-11  6:00   ` Michael Wu
2026-08-11 13:00     ` Steven Rostedt
2026-08-12  2:14       ` Michael Wu
2026-08-12 13:51         ` Steven Rostedt

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox