All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Andrea Righi" <arighi@nvidia.com>
Cc: sched-ext@lists.linux.dev
Subject: Re: [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path
Date: Mon, 10 Aug 2026 15:55:34 +0000	[thread overview]
Message-ID: <20260810155534.D27801F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260810151523.86994-10-arighi@nvidia.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Re-rejecting a task triggers false-positive `WARN_ON_ONCE`s and permanently loses the rejection reason due to conflicting mask management.
--

--- Patch [9]: [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path ---
commit 0787bee4f7073d43925067d6fff35bd1a30ed061
Author: Andrea Righi <arighi@nvidia.com>

sched_ext: Generalize the reject DSQ reenqueue path

This commit generalizes the reject DSQ re-enqueue path by making it
unconditional and carrying the re-enqueue reason directly in p->scx.flags.

> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c

[ ... ]

> @@ -4598,23 +4600,22 @@ static void scx_reenq_reject(struct rq *rq)
>  
>  	lockdep_assert_rq_held(rq);
>  
> -	if (!scx_has_subs() || list_empty(&rq->scx.reject_dsq.list))
> +	if (list_empty(&rq->scx.reject_dsq.list))
>  		return;
>  
>  	/*
> -	 * Move to a private list so a task re-rejected by the
> +	 * Move tasks to a private list so a task re-rejected by
>  	 * scx_do_enqueue_task() below isn't revisited this round.
>  	 */
>  	list_for_each_entry_safe(p, n, &rq->scx.reject_dsq.list, scx.dsq_list.node) {
> +		u32 reason = p->scx.flags & SCX_TASK_REENQ_REASON_MASK;
> +
>  		/* migration_pending tasks should have bypassed to local DSQ */
> -		if (WARN_ON_ONCE(p->migration_pending))
> -			continue;
> +		WARN_ON_ONCE(p->migration_pending);
> +		WARN_ON_ONCE(!reason);

[Severity: High]
Will this WARN_ON_ONCE(!reason) trigger on the next tick if a task is
synchronously re-rejected during scx_do_enqueue_task()?

When a task is pushed back to the reject_dsq during the enqueue attempt,
its rejection reason mask seems to get cleared immediately after by the code
below, leaving the reason as 0 on the next drain.

>  
>  		scx_dispatch_dequeue(rq, p);
> -
> -		if (WARN_ON_ONCE(p->scx.flags & SCX_TASK_REENQ_REASON_MASK))
> -			p->scx.flags &= ~SCX_TASK_REENQ_REASON_MASK;
> -		p->scx.flags |= SCX_TASK_REENQ_CAP;
> +		p->scx.flags |= reason;
>  
>  		list_add_tail(&p->scx.dsq_list.node, &tasks);
>  	}
> @@ -4627,9 +4628,6 @@ static void scx_reenq_reject(struct rq *rq)
>  		p->scx.flags &= ~SCX_TASK_REENQ_REASON_MASK;

[Severity: High]
Does unconditionally clearing the mask here corrupt the state of newly
re-rejected tasks that were just put back on the reject_dsq?

If scx_do_enqueue_task() re-rejects the task, it gets pushed back via
scx_resolve_local_dsq().

Clearing the mask here wipes out the rejection reason while the task sits
on the reject_dsq. Without a valid rejection reason, the BPF scheduler
cannot correct the placement (e.g., missing capabilities).

Could this lead to repeated rejections until the task hits
SCX_REENQ_MAX_REPEAT, triggering __scx_exit() and crashing the BPF scheduler?

A similar issue seems to exist in reenq_local(), where the mask is also
unconditionally cleared after scx_do_enqueue_task() returns.

>  	}
>  }
> -#else
> -static void scx_reenq_reject(struct rq *rq) {}
> -#endif

[ ... ]

> diff --git a/kernel/sched/ext/sub.c b/kernel/sched/ext/sub.c
> --- a/kernel/sched/ext/sub.c
> +++ b/kernel/sched/ext/sub.c
> @@ -752,6 +752,9 @@ struct scx_dispatch_q *scx_resolve_local_dsq(struct scx_sched *sch, struct rq *r
>  
>  	p->scx.reenq_reason_caps = missing;
>  	p->scx.reenq_reason_cid = cid;
> +	if (WARN_ON_ONCE(p->scx.flags & SCX_TASK_REENQ_REASON_MASK))
> +		p->scx.flags &= ~SCX_TASK_REENQ_REASON_MASK;
> +	p->scx.flags |= SCX_TASK_REENQ_CAP;

[Severity: High]
Will this WARN_ON_ONCE fire incorrectly on valid re-rejections?

scx_reenq_reject() deliberately preserves the reason mask prior to the
enqueue attempt. If a task gets synchronously re-rejected here, it will
still have its mask set, which triggers this warning.

>  
>  	return &rq->scx.reject_dsq;
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260810151523.86994-1-arighi@nvidia.com?part=9

  reply	other threads:[~2026-08-10 15:55 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-10 15:13 [PATCHSET v11 sched_ext/for-7.3] sched: Make proxy execution compatible with sched_ext Andrea Righi
2026-08-10 15:13 ` [PATCH 01/15] sched: Make NOHZ CFS bandwidth checks follow proxy donor Andrea Righi
2026-08-10 15:35   ` sashiko-bot
2026-08-10 15:13 ` [PATCH 02/15] sched/core: Avoid false migration warning for proxy donors Andrea Righi
2026-08-10 15:13 ` [PATCH 03/15] sched: Pass next class to sched_change_begin() Andrea Righi
2026-08-10 15:13 ` [PATCH 04/15] sched: Add helper to block retained proxy donors Andrea Righi
2026-08-10 15:13 ` [PATCH 05/15] sched: Add sched_ext hooks for proxy execution Andrea Righi
2026-08-10 15:13 ` [PATCH 06/15] sched_ext: Block proxy donors across scheduler transitions Andrea Righi
2026-08-10 15:13 ` [PATCH 07/15] sched_ext: Fix ops.running/stopping() pairing for proxy-exec donors Andrea Righi
2026-08-10 15:13 ` [PATCH 08/15] sched_ext: Move reject DSQ draining into core Andrea Righi
2026-08-10 15:13 ` [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path Andrea Righi
2026-08-10 15:55   ` sashiko-bot [this message]
2026-08-10 15:13 ` [PATCH 10/15] sched_ext: Handle proxy-exec races in remote DSQ transfers Andrea Righi
2026-08-10 16:04   ` sashiko-bot
2026-08-10 15:13 ` [PATCH 11/15] sched_ext: Split curr|donor references properly Andrea Righi
2026-08-10 16:06   ` sashiko-bot
2026-08-10 15:13 ` [PATCH 12/15] sched_ext: Delegate proxy donor admission to BPF schedulers Andrea Righi
2026-08-10 15:13 ` [PATCH 13/15] sched_ext: Add selftest for blocked donor admission Andrea Righi
2026-08-10 15:14 ` [PATCH 14/15] sched_ext: scx_qmap: Add proxy execution support Andrea Righi
2026-08-10 15:14 ` [PATCH 15/15] sched: Allow enabling proxy exec with sched_ext Andrea Righi
  -- strict thread matches above, loose matches on Subject: below --
2026-07-28 15:43 [PATCHSET v10 sched_ext/for-7.3] sched: Make proxy execution compatible " Andrea Righi
2026-07-28 15:43 ` [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path Andrea Righi
2026-07-28 16:00   ` sashiko-bot
2026-08-03 20:35   ` Tejun Heo
2026-08-03 20:38     ` Tejun Heo
2026-08-05  8:50       ` Andrea Righi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260810155534.D27801F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=arighi@nvidia.com \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=sched-ext@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.