BPF List
 help / color / mirror / Atom feed
From: "Alexei Starovoitov" <alexei.starovoitov@gmail.com>
To: "Emil Tsalapatis" <emil@etsalapatis.com>, <bpf@vger.kernel.org>
Cc: <andrii@kernel.org>, <eddyz87@gmail.com>, <memxor@gmail.com>,
	<daniel@iogearbox.net>,
	"Mykola Lysenko" <nickolay.lysenko@gmail.com>
Subject: Re: [PATCH bpf-next v3 3/6] bpf: Fix arena race between page free and alloc leading to incoherency
Date: Wed, 23 Sep 2026 22:42:02 +0000	[thread overview]
Message-ID: <DLN2450HFL9V.1JUQ5FJMYCMM0@gmail.com> (raw)
In-Reply-To: <20260923191125.5311-4-emil@etsalapatis.com>

On Wed, Sep 23, 2026 at 07:11 PM Emil Tsalapatis <emil@etsalapatis.com> wrote:
> Solve this ABA problem by preventing range reallocation until
> TLB invalidation/unmapping is complete. First, mark the range
> freed but unavailable. Afterwards, drop the spinlock and
> flush the kernel TLB and zap user page tables. Then pick up
> the lock again and mark the ranges as available once again,
> completing the free operation.

The range is allocated in the range tree before the free.
can we keep it allocated until flush_tlb_kernel_range() and
zap_pages() are done and call range_tree_set() only after that ?
bpf_arena_reserve_pages() creates the same 'allocated without pages'
state. arena_vm_fault() shouldn't populate such range, I think.
Then there is no need for the 3rd state in the range tree.

[...]
> +		unavail_node = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
> +		if (IS_ERR(unavail_node)) {
> +			ret = PTR_ERR(unavail_node);
> +			/* Kick off another attempt at the end of this call. */
> +			if (ret == -EAGAIN) {
> +				llist_add(pos, &arena->free_spans);
> +				retry = true;
> +				continue;
> +			}
> +
> +			/*
> +			 * An -ENOMEM failure is the same failure mode as in
> +			 * the defer: path of arena_free_pages(). Do not treat
> +			 * the leak as a bug.
> +			 */
> +			if (ret != -ENOMEM)
> +				WARN_ON_ONCE(ret);
> +
> +			kfree_nolock(s);
> +			continue;
> +		}

This makes it worse.
Unavailable ranges don't merge, so range_tree_set_unavail() always
allocates a node. When kmalloc_nolock() fails the span is dropped
before the ptes are cleared and the pages stay mapped until map free.
Today the worker unmaps and frees the pages and range_tree_set()
failure costs only the address range.

pw-bot: cr

  reply	other threads:[~2026-09-23 22:42 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-23 19:11 [PATCH bpf-next v3 0/6] bpf: Fix arena memory incoherence Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 1/6] bpf: Update is_range_tree_set to work for consecutive ranges Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 2/6] bpf: Track availability information for ranges in range tree Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 3/6] bpf: Fix arena race between page free and alloc leading to incoherency Emil Tsalapatis
2026-09-23 22:42   ` Alexei Starovoitov [this message]
2026-09-24 19:08     ` Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 4/6] bpf: Add explicit state machine for arena free spans Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 5/6] bpf: Atomically update PTE and range tree in arena VM fault handler Emil Tsalapatis
2026-09-23 19:28   ` sashiko-bot
2026-09-23 20:11   ` bot+bpf-ci
2026-09-23 22:42   ` Alexei Starovoitov
2026-09-23 19:11 ` [PATCH bpf-next v3 6/6] selftests/bpf: Add arena allocation race tests Emil Tsalapatis
2026-09-23 20:12   ` bot+bpf-ci

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLN2450HFL9V.1JUQ5FJMYCMM0@gmail.com \
    --to=alexei.starovoitov@gmail.com \
    --cc=andrii@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=emil@etsalapatis.com \
    --cc=memxor@gmail.com \
    --cc=nickolay.lysenko@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox