Linux Trace Kernel
 help / color / mirror / Atom feed
From: Vincent Donnefort <vdonnefort@google.com>
To: Steven Rostedt <rostedt@goodmis.org>
Cc: mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org,
	mathieu.desnoyers@efficios.com, kernel-team@android.com,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v5 06/10] tracing: Fix subbuf resize races with trace_pipe_raw readers
Date: Fri, 14 Aug 2026 09:02:55 +0100	[thread overview]
Message-ID: <an7Lr4wMOwg_-vgX@google.com> (raw)
In-Reply-To: <20260813213153.4ac1e1ff@gandalf.local.home>

On Thu, Aug 13, 2026 at 09:31:53PM -0400, Steven Rostedt wrote:
> On Thu, 13 Aug 2026 14:11:48 +0100
> Vincent Donnefort <vdonnefort@google.com> wrote:
> 
> > diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c
> > index a00ab8a9cbd0..83292d90599e 100644
> > --- a/kernel/trace/ring_buffer.c
> > +++ b/kernel/trace/ring_buffer.c
> > @@ -6976,34 +6976,52 @@ EXPORT_SYMBOL_GPL(ring_buffer_swap_cpu);
> >   * ring_buffer_alloc_read_page - allocate a page to read from buffer
> >   * @buffer: the buffer to allocate for.
> >   * @cpu: the cpu buffer to allocate.
> > + * @prev: The previous page to be repurposed (can be NULL).
> >   *
> > - * This function is used in conjunction with ring_buffer_read_page.
> > + * This function is used in conjunction with ring_buffer_read_page().
> >   * When reading a full page from the ring buffer, these functions
> >   * can be used to speed up the process. The calling function should
> >   * allocate a few pages first with this function. Then when it
> >   * needs to get pages from the ring buffer, it passes the result
> > - * of this function into ring_buffer_read_page, which will swap
> > + * of this function into ring_buffer_read_page(), which will swap
> >   * the page that was allocated, with the read page of the buffer.
> >   *
> > + * If @prev is provided, and it has a different order than the current
> > + * subbuffer order, its payload will be freed and re-allocated. If it
> > + * already matches the order, it is simply returned.
> > + *
> >   * Returns:
> >   *  The page allocated, or ERR_PTR
> >   */
> > -struct buffer_data_read_page *
> > -ring_buffer_alloc_read_page(struct trace_buffer *buffer, int cpu)
> > +struct buffer_data_read_page *ring_buffer_alloc_read_page(struct trace_buffer *buffer, int cpu,
> > +							  struct buffer_data_read_page *prev)
> 
> I think we should do this differently. I don't like the "prev" argument.
> Instead, let's pass by address.
> 
> int ring_buffer_alloc_read_page(struct trace_buffer *buffer, int cpu, struct buffer_data_read_page **rpage)
> 
>

ack.

> 
> 
> > diff --git a/kernel/trace/trace.c b/kernel/trace/trace.c
> > index 395238b2b715..f9399f391ac6 100644
> > --- a/kernel/trace/trace.c
> > +++ b/kernel/trace/trace.c
> > @@ -7080,8 +7080,8 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> >  {
> >  	struct ftrace_buffer_info *info = filp->private_data;
> >  	struct trace_iterator *iter = &info->iter;
> > -	void *trace_data;
> > -	int page_size;
> > +	void *trace_data, *prev_spare;
> > +	unsigned int spare_size;
> >  	ssize_t ret = 0;
> >  	ssize_t size;
> >  
> > @@ -7091,36 +7091,30 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> >  	if (iter->snapshot && tracer_uses_snapshot(iter->tr->current_trace))
> >  		return -EBUSY;
> >  
> > -	page_size = ring_buffer_subbuf_size_get(iter->array_buffer->buffer);
> > +again:
> 
> 
> > +	prev_spare = info->spare;
> > +	if (prev_spare) {
> > +		spare_size = ring_buffer_read_page_size(info->spare);
> >  
> > -	/* Make sure the spare matches the current sub buffer size */
> > -	if (info->spare) {
> > -		if (page_size != info->spare_size) {
> > -			ring_buffer_free_read_page(iter->array_buffer->buffer,
> > -						   info->spare_cpu, info->spare);
> > -			info->spare = NULL;
> > -		}
> > +		/* Do we have previous read data to read? */
> > +		if (info->read < spare_size)
> > +			goto read;
> >  	}
> >  
> > -	if (!info->spare) {
> > -		info->spare = ring_buffer_alloc_read_page(iter->array_buffer->buffer,
> > -							  iter->cpu_file);
> > -		if (IS_ERR(info->spare)) {
> > -			ret = PTR_ERR(info->spare);
> > -			info->spare = NULL;
> > -		} else {
> > -			info->spare_cpu = iter->cpu_file;
> > -			info->spare_size = page_size;
> > -		}
> > -	}
> > -	if (!info->spare)
> > +	/* Make sure the read page order is aligned with the current buffer subbuf order */
> > +	info->spare = ring_buffer_alloc_read_page(iter->array_buffer->buffer, iter->cpu_file,
> > +						  prev_spare);
> > +	if (IS_ERR(info->spare)) {
> > +		ret = PTR_ERR(info->spare);
> > +		info->spare = NULL;
> > +		ring_buffer_free_read_page(iter->array_buffer->buffer, info->spare_cpu, prev_spare);
> >  		return ret;
> > +	}
> 
> 
> instead of the above:
> 
> 	ret = ring_buffer_alloc_read_page(iter->array_buffer->buffer, iter->cpu_file,
> 					  &info->space);
> 
> Where the above could do (under lock):
> 
> 	if (*rpage) {
> 		if (*rpage)->order == buffer->subbuf_order)
> 			return 0;
> 		ring_buffer_free_read_page(*rpage);
> 		*rpage = NULL;
> 	}
> 
> 	*rpage = all the reader page;
> 
> That is, lets completely remove the responsibility of the user having to
> keep track of the buffer order here.
> 
> >  
> > -	/* Do we have previous read data to read? */
> > -	if (info->read < page_size)
> > -		goto read;
> > +	spare_size = ring_buffer_read_page_size(info->spare);
> > +	info->read = spare_size;
> > +	info->spare_cpu = iter->cpu_file;
> >  
> > - again:
> >  	trace_access_lock(iter->cpu_file);
> >  	ret = ring_buffer_read_page(iter->array_buffer->buffer,
> >  				    info->spare,
> 
> And if we could make ring_buffer_read_page() return -EAGAIN if the spare is
> not the proper order. And only for that case.

Sounds good.

-- 
Vincent

> 
> 
> > @@ -7129,6 +7123,10 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> >  	trace_access_unlock(iter->cpu_file);
> >  
> >  	if (ret < 0) {
> > +		/* Did we race with ring_buffer_subbuf_order_set ? */
> > +		if (spare_size != ring_buffer_subbuf_size_get(iter->array_buffer->buffer))
> > +			goto again;
> 
> Then here we can just have:
> 
> 		if (ret == -EAGAIN)
> 			goto again;
> 
> -- Steve
> 
> 
> > +
> >  		if (trace_empty(iter) && !iter->closed) {
> >  			if (update_last_data_if_empty(iter->tr))
> >  				return 0;
> > @@ -7142,12 +7140,14 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> >  
> >  			goto again;
> >  		}
> > +
> >  		return 0;
> >  	}
> >  
> >  	info->read = 0;
> > +
> >   read:
> > -	size = page_size - info->read;
> > +	size = spare_size - info->read;
> >  	if (size > count)
> >  		size = count;
> >  	trace_data = ring_buffer_read_page_data(info->spare);


  reply	other threads:[~2026-08-14  8:03 UTC|newest]

Thread overview: 29+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-13 13:11 [PATCH v5 00/10] ring-buffer: Fixes for subbuf resizing and persistent buffers Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 01/10] ring-buffer: Free cpu_buffer::free_page with subbuf_order Vincent Donnefort
2026-08-13 13:54   ` sashiko-bot
2026-08-13 15:49     ` Masami Hiramatsu
2026-08-13 16:02       ` Vincent Donnefort
2026-08-13 16:03         ` Vincent Donnefort
2026-08-13 16:18   ` Masami Hiramatsu
2026-08-13 13:11 ` [PATCH v5 02/10] ring-buffer: Hold cpu_buffer::lock when resizing a subbuf Vincent Donnefort
2026-08-13 14:01   ` sashiko-bot
2026-08-13 13:11 ` [PATCH v5 03/10] ring-buffer: Make cpu_buffer::free_page a buffer_data_read_page Vincent Donnefort
2026-08-13 13:56   ` sashiko-bot
2026-08-14  1:04     ` Steven Rostedt
2026-08-13 13:11 ` [PATCH v5 04/10] ring-buffer: Fix subbuf resize race with ring buffer readers Vincent Donnefort
2026-08-13 13:51   ` sashiko-bot
2026-08-14  1:12     ` Steven Rostedt
2026-08-13 13:11 ` [PATCH v5 05/10] ring-buffer: Fix subbuf resize race with ring_buffer_alloc_read_page() Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 06/10] tracing: Fix subbuf resize races with trace_pipe_raw readers Vincent Donnefort
2026-08-13 13:53   ` sashiko-bot
2026-08-14  1:31   ` Steven Rostedt
2026-08-14  8:02     ` Vincent Donnefort [this message]
2026-08-13 13:11 ` [PATCH v5 07/10] ring-buffer: Dynamically calculate max_data_size Vincent Donnefort
2026-08-14  1:34   ` Steven Rostedt
2026-08-14  7:59     ` Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 08/10] ring-buffer: Remove trace_buffer::cpus Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 09/10] ring-buffer: Remove ring_buffer_per_cpu::mapped Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 10/10] ring-buffer: Make nr_pages unsigned int Vincent Donnefort
2026-08-13 13:55   ` sashiko-bot
2026-08-14  1:41     ` Steven Rostedt
2026-08-14  8:07       ` Vincent Donnefort

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=an7Lr4wMOwg_-vgX@google.com \
    --to=vdonnefort@google.com \
    --cc=kernel-team@android.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=rostedt@goodmis.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox