All of lore.kernel.org
 help / color / mirror / Atom feed
From: Tejun Heo <tj@kernel.org>
To: Andrea Righi <arighi@nvidia.com>
Cc: David Vernet <void@manifault.com>,
	Changwoo Min <changwoo@igalia.com>,
	John Stultz <jstultz@google.com>, Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Christian Loehle <christian.loehle@arm.com>,
	David Dai <david.dai@linux.dev>, Koba Ko <kobak@nvidia.com>,
	Aiqun Yu <aiqun.yu@oss.qualcomm.com>,
	Shuah Khan <shuah@kernel.org>,
	sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [PATCH 10/15] sched_ext: Handle proxy-exec races in remote DSQ transfers
Date: Mon, 3 Aug 2026 11:36:39 -1000	[thread overview]
Message-ID: <anEJ5xj2bhXg5GSf@slm.duckdns.org> (raw)
In-Reply-To: <20260728154425.1549660-11-arighi@nvidia.com>

Hello,

On Tue, Jul 28, 2026 at 05:43:28PM +0200, Andrea Righi wrote:
...
> @@ -135,6 +135,9 @@ enum scx_ent_flags {
> @@ -145,6 +148,8 @@ enum scx_ent_flags {
>  	SCX_TASK_REENQ_IMMED	= 2 << SCX_TASK_REENQ_REASON_SHIFT,
>  	SCX_TASK_REENQ_PREEMPTED = 3 << SCX_TASK_REENQ_REASON_SHIFT,
>  	SCX_TASK_REENQ_CAP	= 4 << SCX_TASK_REENQ_REASON_SHIFT,
> +	SCX_TASK_REENQ_MIGRATION_DISABLED = 5 << SCX_TASK_REENQ_REASON_SHIFT,
> +	SCX_TASK_REENQ_PROXY	= 6 << SCX_TASK_REENQ_REASON_SHIFT,

Given that MIGRATION_DISABLED can only happen with proxy execution, it may
be better to name it accordingly. Maybe that's too long. Do we have to
distinguish between PROXY and MIGRATION_DISABLED? We can count both as
PROXY, no?

> +/*
> + * Proxy execution can change @p's execution and migration-disabled state
> + * without touching its DSQ entry or clearing holding_cpu. Check those states
> + * with @p's rq locked. Without proxy execution, the holding_cpu handshake is
> + * sufficient and this must not affect the existing migration path.
> + */
> +static u32 task_move_reject_reason(struct task_struct *p)
> +{
> +	struct rq *src_rq = task_rq(p);
> +
> +	lockdep_assert_rq_held(src_rq);
> +
> +	if (!sched_proxy_exec())
> +		return SCX_TASK_REENQ_NONE;
> +
> +	/* @p may be rq->curr under another task's scheduling context. */
> +	if (task_on_cpu(src_rq, p))
> +		return SCX_TASK_REENQ_PROXY;
> +
> +	/*
> +	 * Reject only BPF-directed migration. proxy_migrate_task() may still
> +	 * move a blocked donor's scheduling context to its lock owner's CPU.
> +	 */

I have a hard time understanding this comment. Can you explain a bit more?

> +	if (is_migration_disabled(p))
> +		return SCX_TASK_REENQ_MIGRATION_DISABLED;
> +
> +	/* Don't move an active scheduling context off its source rq. */
> +	if (task_current_donor(src_rq, p))
> +		return SCX_TASK_REENQ_PROXY;
> +
> +	return SCX_TASK_REENQ_NONE;
> +}
> +
> +/*
> + * Park a task whose remote transfer raced with proxy execution. Reenqueueing
> + * from the source rq makes the task's owning scheduler choose its placement
> + * again and preserves sub-scheduler containment.
> + */
> +static void scx_reject_task(struct scx_sched *sch, struct rq *rq,
> +			    struct task_struct *p, u64 enq_flags, u32 reason)
> +{
> +	lockdep_assert_rq_held(rq);
> +	WARN_ON_ONCE(reason != SCX_TASK_REENQ_MIGRATION_DISABLED &&
> +		     reason != SCX_TASK_REENQ_PROXY);
> +	WARN_ON_ONCE(p->scx.reject_reason);
> +
> +	p->scx.holding_cpu = -1;
> +	p->scx.reject_reason = reason;
> +	p->scx.flags &= ~SCX_TASK_IMMED;
> +	enq_flags &= ~(SCX_ENQ_IMMED | SCX_ENQ_PREEMPT);

SCX_ENQ_HEAD likely needs clearing too. After applying the rescue patchset,
scx_resolve_local_dsq() has:

	/*
	 * Diverting to rescue or reject, neither of which honors IMMED, PREEMPT
	 * or HEAD - a diversion has no priority and IMMED is not allowed on
	 * non-local DSQs. Strip the enq and task flags along with the slice.
	 */
	*enq_flags &= ~(SCX_ENQ_IMMED | SCX_ENQ_PREEMPT | SCX_ENQ_HEAD |
			SCX_ENQ_APPLY_SLICE | SCX_ENQ_SLICE_DFL);
	p->scx.flags &= ~SCX_TASK_IMMED;

We should probably factor that out and use that whenever we're diverting.

> @@ -2533,6 +2617,15 @@ static struct rq *move_task_between_dsqs(struct scx_sched *sch,
>  
>  	if (dst_dsq->id == SCX_DSQ_LOCAL) {
>  		dst_rq = container_of(dst_dsq, struct rq, scx.local_dsq);
> +		reject_reason = src_rq != dst_rq ?
> +			task_move_reject_reason(p) : SCX_TASK_REENQ_NONE;
> +		if (unlikely(reject_reason)) {
> +			dispatch_dequeue_locked(p, src_dsq);
> +			raw_spin_unlock(&src_dsq->lock);
> +			scx_reject_task(sch, src_rq, p, enq_flags,
> +					reject_reason);
> +			return src_rq;
> +		}

I wonder whether it'd be cleaner if we just mark the task for rejection and
then make scx_resolve_local_dsq() resolve that to reject dsq instead of
directly inserting from each site. I guess the problem is that the rejection
has to be on the source rq, not the destination one. I hope there's a neater
way to do this.

>  /*
> - * Drain @rq->scx.reject_dsq and reenqueue each task so that its owning BPF
> - * scheduler chooses placement again.
> + * Drain ready tasks from @rq->scx.reject_dsq and reenqueue them so that their
> + * owning BPF schedulers choose placement again. Proxy-active tasks remain
> + * parked until proxy resolution schedules another drain after switch-out.
>   *
>   * A task can be re-rejected repeatedly. Reenqueues are bounded per task by
>   * SCX_REENQ_MAX_REPEAT in scx_do_enqueue_task(), which ejects the owning
> - * scheduler. The private list below prevents a task from being revisited in
> - * the same round.
> + * scheduler.
>   */
>  static void scx_reenq_reject(struct rq *rq)
>  {
>  	LIST_HEAD(tasks);
>  	struct task_struct *p, *n;
> +	bool proxy_pending = false;
>  
>  	lockdep_assert_rq_held(rq);
>  
> -	if (list_empty(&rq->scx.reject_dsq.list))
> +	if (list_empty(&rq->scx.reject_dsq.list)) {
> +		rq->scx.flags &= ~SCX_RQ_PROXY_REENQ;
>  		return;
> +	}
>  
>  	/*
> -	 * Move tasks to a private list so a task re-rejected by
> +	 * Move ready tasks to a private list so a task re-rejected by
>  	 * scx_do_enqueue_task() below isn't revisited this round.
>  	 */
>  	list_for_each_entry_safe(p, n, &rq->scx.reject_dsq.list, scx.dsq_list.node) {
>  		u32 reason = p->scx.reject_reason;
>  
>  		/* migration_pending tasks should have bypassed to local DSQ */
> -		if (WARN_ON_ONCE(p->migration_pending))
> -			continue;
> +		WARN_ON_ONCE(p->migration_pending);

Why this change? Also, wouldn't this now be allowed to happen? During task
move, if it hits migration_pending due to proxy exec, the task would be put
on reject DSQ, right? Reenq can trigger from other sources before
migration_pending is cleared and then would see the migration_pending set.
Shouldn't it just continue?

>  		if (WARN_ON_ONCE(!reason))
>  			continue;
>  
> +		if (reason == SCX_TASK_REENQ_PROXY &&
> +		    (task_on_cpu(rq, p) || task_current_donor(rq, p))) {
> +			proxy_pending = true;
> +			continue;
> +		}

ie. Shouldn't migration_pending() be one of the || conditions above? And
then if it's still migraiton_pending, that should trigger a warning?

Thanks.

-- 
tejun

  reply	other threads:[~2026-08-03 21:36 UTC|newest]

Thread overview: 34+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28 15:43 [PATCHSET v10 sched_ext/for-7.3] sched: Make proxy execution compatible with sched_ext Andrea Righi
2026-07-28 15:43 ` [PATCH 01/15] sched/core: Avoid false migration warning for proxy donors Andrea Righi
2026-07-28 15:43 ` [PATCH 02/15] sched: Make NOHZ CFS bandwidth checks follow proxy donor Andrea Righi
2026-07-28 15:43 ` [PATCH 03/15] sched: Add helper to block retained proxy donors Andrea Righi
2026-07-28 15:43 ` [PATCH 04/15] sched: Skip class callbacks with SCHED_FLAG_KEEP_PARAMS Andrea Righi
2026-07-28 16:03   ` sashiko-bot
2026-07-28 15:43 ` [PATCH 05/15] sched: Add prepare_switch() class callback Andrea Righi
2026-07-28 15:43 ` [PATCH 06/15] sched: Add sched_ext hooks for proxy execution Andrea Righi
2026-07-28 15:43 ` [PATCH 07/15] sched_ext: Block proxy donors across scheduler transitions Andrea Righi
2026-08-03 20:36   ` Tejun Heo
2026-08-05  7:02     ` Andrea Righi
2026-08-06  9:07       ` Andrea Righi
2026-07-28 15:43 ` [PATCH 08/15] sched_ext: Fix ops.running/stopping() pairing for proxy-exec donors Andrea Righi
2026-07-28 15:43 ` [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path Andrea Righi
2026-07-28 16:00   ` sashiko-bot
2026-08-03 20:35   ` Tejun Heo
2026-08-03 20:38     ` Tejun Heo
2026-08-05  8:50       ` Andrea Righi
2026-07-28 15:43 ` [PATCH 10/15] sched_ext: Handle proxy-exec races in remote DSQ transfers Andrea Righi
2026-08-03 21:36   ` Tejun Heo [this message]
2026-08-05 16:44     ` Andrea Righi
2026-07-28 15:43 ` [PATCH 11/15] sched_ext: Split curr|donor references properly Andrea Righi
2026-07-28 16:25   ` sashiko-bot
2026-07-28 15:43 ` [PATCH 12/15] sched_ext: Delegate proxy donor admission to BPF schedulers Andrea Righi
2026-08-03 22:18   ` Tejun Heo
2026-08-06  6:10     ` Andrea Righi
2026-07-28 15:43 ` [PATCH 13/15] sched_ext: Add selftest for blocked donor admission Andrea Righi
2026-07-28 15:43 ` [PATCH 14/15] sched_ext: scx_qmap: Add proxy execution support Andrea Righi
2026-08-03 22:22   ` Tejun Heo
2026-08-06  7:28     ` Andrea Righi
2026-07-28 15:43 ` [PATCH 15/15] sched: Allow enabling proxy exec with sched_ext Andrea Righi
2026-08-03 22:19   ` Tejun Heo
  -- strict thread matches above, loose matches on Subject: below --
2026-08-10 15:13 [PATCHSET v11 sched_ext/for-7.3] sched: Make proxy execution compatible " Andrea Righi
2026-08-10 15:13 ` [PATCH 10/15] sched_ext: Handle proxy-exec races in remote DSQ transfers Andrea Righi
2026-08-10 16:04   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anEJ5xj2bhXg5GSf@slm.duckdns.org \
    --to=tj@kernel.org \
    --cc=aiqun.yu@oss.qualcomm.com \
    --cc=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=christian.loehle@arm.com \
    --cc=david.dai@linux.dev \
    --cc=dietmar.eggemann@arm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kobak@nvidia.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.