From: "Alexei Starovoitov" <alexei.starovoitov@gmail.com>
To: "Emil Tsalapatis" <emil@etsalapatis.com>, <bpf@vger.kernel.org>
Cc: <andrii@kernel.org>, <eddyz87@gmail.com>, <memxor@gmail.com>,
<daniel@iogearbox.net>,
"Mykola Lysenko" <nickolay.lysenko@gmail.com>
Subject: Re: [PATCH bpf-next v3 3/6] bpf: Fix arena race between page free and alloc leading to incoherency
Date: Wed, 23 Sep 2026 22:42:02 +0000 [thread overview]
Message-ID: <DLN2450HFL9V.1JUQ5FJMYCMM0@gmail.com> (raw)
In-Reply-To: <20260923191125.5311-4-emil@etsalapatis.com>
On Wed, Sep 23, 2026 at 07:11 PM Emil Tsalapatis <emil@etsalapatis.com> wrote:
> Solve this ABA problem by preventing range reallocation until
> TLB invalidation/unmapping is complete. First, mark the range
> freed but unavailable. Afterwards, drop the spinlock and
> flush the kernel TLB and zap user page tables. Then pick up
> the lock again and mark the ranges as available once again,
> completing the free operation.
The range is allocated in the range tree before the free.
can we keep it allocated until flush_tlb_kernel_range() and
zap_pages() are done and call range_tree_set() only after that ?
bpf_arena_reserve_pages() creates the same 'allocated without pages'
state. arena_vm_fault() shouldn't populate such range, I think.
Then there is no need for the 3rd state in the range tree.
[...]
> + unavail_node = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
> + if (IS_ERR(unavail_node)) {
> + ret = PTR_ERR(unavail_node);
> + /* Kick off another attempt at the end of this call. */
> + if (ret == -EAGAIN) {
> + llist_add(pos, &arena->free_spans);
> + retry = true;
> + continue;
> + }
> +
> + /*
> + * An -ENOMEM failure is the same failure mode as in
> + * the defer: path of arena_free_pages(). Do not treat
> + * the leak as a bug.
> + */
> + if (ret != -ENOMEM)
> + WARN_ON_ONCE(ret);
> +
> + kfree_nolock(s);
> + continue;
> + }
This makes it worse.
Unavailable ranges don't merge, so range_tree_set_unavail() always
allocates a node. When kmalloc_nolock() fails the span is dropped
before the ptes are cleared and the pages stay mapped until map free.
Today the worker unmaps and frees the pages and range_tree_set()
failure costs only the address range.
pw-bot: cr
next prev parent reply other threads:[~2026-09-23 22:42 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 19:11 [PATCH bpf-next v3 0/6] bpf: Fix arena memory incoherence Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 1/6] bpf: Update is_range_tree_set to work for consecutive ranges Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 2/6] bpf: Track availability information for ranges in range tree Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 3/6] bpf: Fix arena race between page free and alloc leading to incoherency Emil Tsalapatis
2026-09-23 22:42 ` Alexei Starovoitov [this message]
2026-09-24 19:08 ` Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 4/6] bpf: Add explicit state machine for arena free spans Emil Tsalapatis
2026-09-23 19:11 ` [PATCH bpf-next v3 5/6] bpf: Atomically update PTE and range tree in arena VM fault handler Emil Tsalapatis
2026-09-23 19:28 ` sashiko-bot
2026-09-23 20:11 ` bot+bpf-ci
2026-09-23 22:42 ` Alexei Starovoitov
2026-09-23 19:11 ` [PATCH bpf-next v3 6/6] selftests/bpf: Add arena allocation race tests Emil Tsalapatis
2026-09-23 20:12 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLN2450HFL9V.1JUQ5FJMYCMM0@gmail.com \
--to=alexei.starovoitov@gmail.com \
--cc=andrii@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=memxor@gmail.com \
--cc=nickolay.lysenko@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox