From: Vincent Donnefort <vdonnefort@google.com>
To: Steven Rostedt <rostedt@goodmis.org>
Cc: mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org,
mathieu.desnoyers@efficios.com, kernel-team@android.com,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v5 06/10] tracing: Fix subbuf resize races with trace_pipe_raw readers
Date: Fri, 14 Aug 2026 09:02:55 +0100 [thread overview]
Message-ID: <an7Lr4wMOwg_-vgX@google.com> (raw)
In-Reply-To: <20260813213153.4ac1e1ff@gandalf.local.home>
On Thu, Aug 13, 2026 at 09:31:53PM -0400, Steven Rostedt wrote:
> On Thu, 13 Aug 2026 14:11:48 +0100
> Vincent Donnefort <vdonnefort@google.com> wrote:
>
> > diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c
> > index a00ab8a9cbd0..83292d90599e 100644
> > --- a/kernel/trace/ring_buffer.c
> > +++ b/kernel/trace/ring_buffer.c
> > @@ -6976,34 +6976,52 @@ EXPORT_SYMBOL_GPL(ring_buffer_swap_cpu);
> > * ring_buffer_alloc_read_page - allocate a page to read from buffer
> > * @buffer: the buffer to allocate for.
> > * @cpu: the cpu buffer to allocate.
> > + * @prev: The previous page to be repurposed (can be NULL).
> > *
> > - * This function is used in conjunction with ring_buffer_read_page.
> > + * This function is used in conjunction with ring_buffer_read_page().
> > * When reading a full page from the ring buffer, these functions
> > * can be used to speed up the process. The calling function should
> > * allocate a few pages first with this function. Then when it
> > * needs to get pages from the ring buffer, it passes the result
> > - * of this function into ring_buffer_read_page, which will swap
> > + * of this function into ring_buffer_read_page(), which will swap
> > * the page that was allocated, with the read page of the buffer.
> > *
> > + * If @prev is provided, and it has a different order than the current
> > + * subbuffer order, its payload will be freed and re-allocated. If it
> > + * already matches the order, it is simply returned.
> > + *
> > * Returns:
> > * The page allocated, or ERR_PTR
> > */
> > -struct buffer_data_read_page *
> > -ring_buffer_alloc_read_page(struct trace_buffer *buffer, int cpu)
> > +struct buffer_data_read_page *ring_buffer_alloc_read_page(struct trace_buffer *buffer, int cpu,
> > + struct buffer_data_read_page *prev)
>
> I think we should do this differently. I don't like the "prev" argument.
> Instead, let's pass by address.
>
> int ring_buffer_alloc_read_page(struct trace_buffer *buffer, int cpu, struct buffer_data_read_page **rpage)
>
>
ack.
>
>
> > diff --git a/kernel/trace/trace.c b/kernel/trace/trace.c
> > index 395238b2b715..f9399f391ac6 100644
> > --- a/kernel/trace/trace.c
> > +++ b/kernel/trace/trace.c
> > @@ -7080,8 +7080,8 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> > {
> > struct ftrace_buffer_info *info = filp->private_data;
> > struct trace_iterator *iter = &info->iter;
> > - void *trace_data;
> > - int page_size;
> > + void *trace_data, *prev_spare;
> > + unsigned int spare_size;
> > ssize_t ret = 0;
> > ssize_t size;
> >
> > @@ -7091,36 +7091,30 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> > if (iter->snapshot && tracer_uses_snapshot(iter->tr->current_trace))
> > return -EBUSY;
> >
> > - page_size = ring_buffer_subbuf_size_get(iter->array_buffer->buffer);
> > +again:
>
>
> > + prev_spare = info->spare;
> > + if (prev_spare) {
> > + spare_size = ring_buffer_read_page_size(info->spare);
> >
> > - /* Make sure the spare matches the current sub buffer size */
> > - if (info->spare) {
> > - if (page_size != info->spare_size) {
> > - ring_buffer_free_read_page(iter->array_buffer->buffer,
> > - info->spare_cpu, info->spare);
> > - info->spare = NULL;
> > - }
> > + /* Do we have previous read data to read? */
> > + if (info->read < spare_size)
> > + goto read;
> > }
> >
> > - if (!info->spare) {
> > - info->spare = ring_buffer_alloc_read_page(iter->array_buffer->buffer,
> > - iter->cpu_file);
> > - if (IS_ERR(info->spare)) {
> > - ret = PTR_ERR(info->spare);
> > - info->spare = NULL;
> > - } else {
> > - info->spare_cpu = iter->cpu_file;
> > - info->spare_size = page_size;
> > - }
> > - }
> > - if (!info->spare)
> > + /* Make sure the read page order is aligned with the current buffer subbuf order */
> > + info->spare = ring_buffer_alloc_read_page(iter->array_buffer->buffer, iter->cpu_file,
> > + prev_spare);
> > + if (IS_ERR(info->spare)) {
> > + ret = PTR_ERR(info->spare);
> > + info->spare = NULL;
> > + ring_buffer_free_read_page(iter->array_buffer->buffer, info->spare_cpu, prev_spare);
> > return ret;
> > + }
>
>
> instead of the above:
>
> ret = ring_buffer_alloc_read_page(iter->array_buffer->buffer, iter->cpu_file,
> &info->space);
>
> Where the above could do (under lock):
>
> if (*rpage) {
> if (*rpage)->order == buffer->subbuf_order)
> return 0;
> ring_buffer_free_read_page(*rpage);
> *rpage = NULL;
> }
>
> *rpage = all the reader page;
>
> That is, lets completely remove the responsibility of the user having to
> keep track of the buffer order here.
>
> >
> > - /* Do we have previous read data to read? */
> > - if (info->read < page_size)
> > - goto read;
> > + spare_size = ring_buffer_read_page_size(info->spare);
> > + info->read = spare_size;
> > + info->spare_cpu = iter->cpu_file;
> >
> > - again:
> > trace_access_lock(iter->cpu_file);
> > ret = ring_buffer_read_page(iter->array_buffer->buffer,
> > info->spare,
>
> And if we could make ring_buffer_read_page() return -EAGAIN if the spare is
> not the proper order. And only for that case.
Sounds good.
--
Vincent
>
>
> > @@ -7129,6 +7123,10 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> > trace_access_unlock(iter->cpu_file);
> >
> > if (ret < 0) {
> > + /* Did we race with ring_buffer_subbuf_order_set ? */
> > + if (spare_size != ring_buffer_subbuf_size_get(iter->array_buffer->buffer))
> > + goto again;
>
> Then here we can just have:
>
> if (ret == -EAGAIN)
> goto again;
>
> -- Steve
>
>
> > +
> > if (trace_empty(iter) && !iter->closed) {
> > if (update_last_data_if_empty(iter->tr))
> > return 0;
> > @@ -7142,12 +7140,14 @@ ssize_t tracing_buffers_read(struct file *filp, char __user *ubuf,
> >
> > goto again;
> > }
> > +
> > return 0;
> > }
> >
> > info->read = 0;
> > +
> > read:
> > - size = page_size - info->read;
> > + size = spare_size - info->read;
> > if (size > count)
> > size = count;
> > trace_data = ring_buffer_read_page_data(info->spare);
next prev parent reply other threads:[~2026-08-14 8:03 UTC|newest]
Thread overview: 32+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-13 13:11 [PATCH v5 00/10] ring-buffer: Fixes for subbuf resizing and persistent buffers Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 01/10] ring-buffer: Free cpu_buffer::free_page with subbuf_order Vincent Donnefort
2026-08-13 13:54 ` sashiko-bot
2026-08-13 15:49 ` Masami Hiramatsu
2026-08-13 16:02 ` Vincent Donnefort
2026-08-13 16:03 ` Vincent Donnefort
2026-08-13 16:18 ` Masami Hiramatsu
2026-08-13 13:11 ` [PATCH v5 02/10] ring-buffer: Hold cpu_buffer::lock when resizing a subbuf Vincent Donnefort
2026-08-13 14:01 ` sashiko-bot
2026-08-13 13:11 ` [PATCH v5 03/10] ring-buffer: Make cpu_buffer::free_page a buffer_data_read_page Vincent Donnefort
2026-08-13 13:56 ` sashiko-bot
2026-08-14 1:04 ` Steven Rostedt
2026-08-13 13:11 ` [PATCH v5 04/10] ring-buffer: Fix subbuf resize race with ring buffer readers Vincent Donnefort
2026-08-13 13:51 ` sashiko-bot
2026-08-14 1:12 ` Steven Rostedt
2026-08-13 13:11 ` [PATCH v5 05/10] ring-buffer: Fix subbuf resize race with ring_buffer_alloc_read_page() Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 06/10] tracing: Fix subbuf resize races with trace_pipe_raw readers Vincent Donnefort
2026-08-13 13:53 ` sashiko-bot
2026-08-14 1:31 ` Steven Rostedt
2026-08-14 8:02 ` Vincent Donnefort [this message]
2026-08-13 13:11 ` [PATCH v5 07/10] ring-buffer: Dynamically calculate max_data_size Vincent Donnefort
2026-08-14 1:34 ` Steven Rostedt
2026-08-14 7:59 ` Vincent Donnefort
2026-08-14 12:46 ` Steven Rostedt
2026-08-14 12:50 ` Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 08/10] ring-buffer: Remove trace_buffer::cpus Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 09/10] ring-buffer: Remove ring_buffer_per_cpu::mapped Vincent Donnefort
2026-08-13 13:11 ` [PATCH v5 10/10] ring-buffer: Make nr_pages unsigned int Vincent Donnefort
2026-08-13 13:55 ` sashiko-bot
2026-08-14 1:41 ` Steven Rostedt
2026-08-14 8:07 ` Vincent Donnefort
2026-08-14 12:37 ` Steven Rostedt
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=an7Lr4wMOwg_-vgX@google.com \
--to=vdonnefort@google.com \
--cc=kernel-team@android.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=rostedt@goodmis.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.