BPF List
 help / color / mirror / Atom feed
From: luyun <luyun_611@163.com>
To: Mykyta Yatsenko <mykyta.yatsenko5@gmail.com>,
	bot+bpf-ci@kernel.org, ast@kernel.org, daniel@iogearbox.net,
	andrii@kernel.org, eddyz87@gmail.com, memxor@gmail.com,
	martin.lau@linux.dev, song@kernel.org, yonghong.song@linux.dev,
	jolsa@kernel.org, emil@etsalapatis.com, ihor.solodrai@linux.dev
Cc: yatsenko@meta.com, bpf@vger.kernel.org, martin.lau@kernel.org,
	mason@kernel.org
Subject: Re: [PATCH bpf-next v4 3/3] selftests/bpf: Add task work scheduling race test
Date: Wed, 30 Sep 2026 16:37:57 +0800	[thread overview]
Message-ID: <42361d59-bd08-49cd-9896-9594e7521ee8@163.com> (raw)
In-Reply-To: <e443d46b-1062-484a-a20e-20c0d85afa48@gmail.com>


在 2026/9/29 20:00, Mykyta Yatsenko 写道:
>
> On 9/29/26 9:12 AM, bot+bpf-ci@kernel.org wrote:
>>> selftests/bpf: Add task work scheduling race test
>>>
>>> Exercise task-work scheduling and callback execution on different CPUs
>>> while userspace reuses or deletes the associated map value.  Run 1500
>>> rounds rotating through context reuse, deletion after the callback has
>>> recorded its generation, and deletion racing with scheduling.
>>>
>>> Publish the scheduling result and its generation in one packed value, so
>>> userspace does not rely on ordering between separate map slots.  Record
>>> callbacks in separate per-generation slots so an earlier callback cannot
>>> satisfy a later round's completion check.
>>>
>>> Retry transient -EBUSY results while the context is unavailable for
>>> reuse.  In the deletion-race case, allow -ENOENT if the element is
>>> removed before a retry finds it.
>>>
>>> Choose two CPUs from the process affinity mask and skip if fewer than
>>> two are available.  Bound the thread-start wait, report thread creation
>>> and affinity failures, and stop the worker threads with an atomic
>>> handshake.
>>>
>>> Signed-off-by: Yun Lu <luyun@kylinos.cn>
>> This isn't a bug, but could the changelog say which of the preceding fixes
>> this test is meant to cover, and which interleaving each variant targets,
>> instead of listing implementation details like the bounded thread-start
>> wait and the atomic stop handshake?
>>
>>> diff --git a/tools/testing/selftests/bpf/prog_tests/test_task_work.c b/tools/testing/selftests/bpf/prog_tests/test_task_work.c
>>> --- a/tools/testing/selftests/bpf/prog_tests/test_task_work.c
>>> +++ b/tools/testing/selftests/bpf/prog_tests/test_task_work.c
>>> @@ -5,10 +5,14 @@
>>>   #include <stdio.h>
>>>   #include "task_work.skel.h"
>>>   #include "task_work_fail.skel.h"
>>> +#include "task_work_race.skel.h"
>>>   #include <linux/bpf.h>
>>>   #include <linux/perf_event.h>
>>>   #include <sys/syscall.h>
>>>   #include <time.h>
>>> +#include <pthread.h>
>>> +#include <sched.h>
>>> +#include <unistd.h>
>> [ ... ]
>>
>>> +static void *task_work_race_trigger(void *arg)
>>> +{
>>> +	struct task_work_race_ctx *ctx = arg;
>>> +	int err;
>>> +
>>> +	err = task_work_race_pin_cpu(ctx->trigger_cpu);
>>> +	if (err)
>>> +		__atomic_store_n(&ctx->setup_err, err, __ATOMIC_RELEASE);
>>> +	__atomic_store_n(&ctx->trigger_tid, syscall(__NR_gettid),
>>> +			 __ATOMIC_RELEASE);
>>> +	while (!__atomic_load_n(&ctx->stop, __ATOMIC_ACQUIRE))
>>> +		getppid();
>>> +	return NULL;
>>> +}
>>> +
>>> +static void *task_work_race_target(void *arg)
>>> +{
>>> +	struct task_work_race_ctx *ctx = arg;
>>> +	int err;
>>> +
>>> +	err = task_work_race_pin_cpu(ctx->target_cpu);
>>> +	if (err)
>>> +		__atomic_store_n(&ctx->setup_err, err, __ATOMIC_RELEASE);
>>> +	__atomic_store_n(&ctx->target_tid, syscall(__NR_gettid),
>>> +			 __ATOMIC_RELEASE);
>>> +	while (!__atomic_load_n(&ctx->stop, __ATOMIC_ACQUIRE))
>>> +		getppid();
>>> +	return NULL;
>>> +}
>> This isn't a bug, but could the trigger and target threads share one
>> function with a small per-thread argument (cpu and tid pointer)? And could
>> the repeated 2000000 / 100 poll bound become a named constant, like
>> TASK_WORK_RACE_ROUNDS?
>>
> I think this is worth addressing.
Hi, Mykyta

Thanks for your reviewing.
I carefully considered the 3 suggestions from the bot+bpf-ci,
and believe they are all necessary. I will incorporate them into
the v5 revision.

I will send out the v5 revision later (selftest changes only).

---

Thanks,

Yun Lu

>> [ ... ]
>>
>>> +		err = task_work_race_check_done(skel, seq);
>>> +		/*
>>> +		 * Variant 2 deletes the element right after READY, so losing
>>> +		 * the race against the scheduling kfunc is a valid outcome:
>>> +		 * the schedule itself fails with -EBUSY (ctx already FREED),
>>> +		 * and the retry on the next tracepoint invocation fails with
>>> +		 * -ENOENT because the element is gone.  Either way the round
>>> +		 * exercised the deletion paths; a later variant 0 round
>>> +		 * recreates the element.
>>> +		 */
>>> +		if (variant == 2 && (err == -EBUSY || err == -ENOENT))
>>> +			err = 0;
>> This isn't a bug, but can err actually be -EBUSY here, given that
>> race_sched_work() returns without publishing RESULT on -EBUSY? If not,
>> could the condition just check -ENOENT?
>>
>> [ ... ]
>>
>>> diff --git a/tools/testing/selftests/bpf/progs/task_work_race.c b/tools/testing/selftests/bpf/progs/task_work_race.c
>>> --- /dev/null
>>> +++ b/tools/testing/selftests/bpf/progs/task_work_race.c
>> [ ... ]
>>
>> ---
>> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
>> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>>
>> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36538172291


      reply	other threads:[~2026-09-30  8:38 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29  7:27 [PATCH bpf-next v4 0/3] bpf: Fix task work scheduling races Yun Lu
2026-09-29  7:27 ` [PATCH bpf-next v4 1/3] bpf: Fix NULL task dereference in task work cancellation Yun Lu
2026-09-29  7:27 ` [PATCH bpf-next v4 2/3] bpf: Prevent task work context reuse while irq_work is busy Yun Lu
2026-09-29 12:02   ` Mykyta Yatsenko
2026-09-29  7:28 ` [PATCH bpf-next v4 3/3] selftests/bpf: Add task work scheduling race test Yun Lu
2026-09-29  8:12   ` bot+bpf-ci
2026-09-29 12:00     ` Mykyta Yatsenko
2026-09-30  8:37       ` luyun [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=42361d59-bd08-49cd-9896-9594e7521ee8@163.com \
    --to=luyun_611@163.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bot+bpf-ci@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=emil@etsalapatis.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=jolsa@kernel.org \
    --cc=martin.lau@kernel.org \
    --cc=martin.lau@linux.dev \
    --cc=mason@kernel.org \
    --cc=memxor@gmail.com \
    --cc=mykyta.yatsenko5@gmail.com \
    --cc=song@kernel.org \
    --cc=yatsenko@meta.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox