BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next 0/2] bpf: Fix task work round ownership race
@ 2026-09-18  8:00 Yun Lu
  2026-09-18  8:00 ` [PATCH bpf-next 1/2] bpf: Fix task work round ownership during cancellation Yun Lu
                   ` (2 more replies)
  0 siblings, 3 replies; 9+ messages in thread
From: Yun Lu @ 2026-09-18  8:00 UTC (permalink / raw)
  To: ast, daniel, andrii, eddyz87, memxor, martin.lau, song,
	yonghong.song, jolsa, emil, ihor.solodrai
  Cc: bpf

From: Yun Lu <luyun@kylinos.cn>

This series fixes a race in the bpf_task_work scheduling kfuncs and
adds a regression test for it.

A task_work callback scheduled for a task running on another CPU can
complete before bpf_task_work_irq() regains control after
task_work_add().  The callback's cleanup releases ctx->task, so a
concurrent map value deletion that publishes the FREED state makes the
resumed irq_work handler call task_work_cancel(NULL, &ctx->work),
dereferencing task_struct::task_works through a NULL task (RIP at
task_work_cancel+0xd, RDI == 0, CR2 at the task_works offset).  The
same cancellation is reached when the deletion wins before the
callback runs and the callback's bailout path drops the last ctx
refcount, clearing ctx->task.

Patch 1 gives each scheduling round an ownership count: the irq_work
scheduler, the published callback and an asynchronous canceller hold a
reference while using this round's task/prog/work, and only the last
user releases them, before the ctx can be reused from STANDBY.  The
count is zero based with the STANDBY -> PENDING transition acting as
the zero-to-one gate, so a finished round cannot be revived: a
canceller either pins a still-active round (ctx->task guaranteed
valid) or declines to cancel.  Separate irq_work objects for
scheduling, cancellation and deferred destruction avoid reinitializing
an irq_work that may still be queued or running.  FREED remains
terminal and the scheduling path stays atomic-only, preserving NMI
safety.

The race window between task_work_add() and the SCHEDULING ->
SCHEDULED cmpxchg is only a few instructions wide, so the fix was
verified with a deterministic reproduction that widens exactly this
window with a debug delay, while scheduling cross-CPU and deleting the
map value concurrently: without the fix the kernel panics reliably;
with it, all interleavings complete.  A 900-round stress test rotating
through ctx reuse, deletion after the callback and deletion right
after scheduling also runs clean, with no leaks reported by kmemleak.

Patch 2 adds a selftest that keeps steady pressure on the
interleaving: it schedules cross-CPU, deletes the map value at
different points of a round, uses READY/DONE handshakes so scheduling
errors cannot be missed, and tags every callback with a generation so
a late callback cannot satisfy a later round's assertions.

Yun Lu (2):
  bpf: Fix task work round ownership during cancellation
  selftests/bpf: Add task work round ownership race test

 kernel/bpf/helpers.c                                  | 123 ++++++++++--
 .../selftests/bpf/prog_tests/test_task_work.c         | 283 +++++++++++++++++++++
 .../selftests/bpf/progs/task_work_race.c              | 125 ++++++++
 3 files changed, 503 insertions(+), 28 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-21 10:23 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18  8:00 [PATCH bpf-next 0/2] bpf: Fix task work round ownership race Yun Lu
2026-09-18  8:00 ` [PATCH bpf-next 1/2] bpf: Fix task work round ownership during cancellation Yun Lu
2026-09-18 17:54   ` Alexei Starovoitov
2026-09-21 10:22     ` luyun
2026-09-18  8:00 ` [PATCH bpf-next 2/2] selftests/bpf: Add task work round ownership race test Yun Lu
2026-09-18  8:13   ` sashiko-bot
2026-09-18  9:12   ` bot+bpf-ci
2026-09-18 16:50 ` [PATCH bpf-next 0/2] bpf: Fix task work round ownership race Mykyta Yatsenko
2026-09-18 17:17   ` Mykyta Yatsenko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox