All of lore.kernel.org
 help / color / mirror / Atom feed
From: Kuba Piecuch <jpiecuch@google.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
	Andrea Righi <arighi@nvidia.com>,
	 Changwoo Min <changwoo@igalia.com>
Cc: Kuba Piecuch <jpiecuch@google.com>,
	Emil Tsalapatis <emil@etsalapatis.com>,
	sched-ext@lists.linux.dev,  linux-kernel@vger.kernel.org
Subject: [PATCHSET v2 sched_ext/for-7.3-fixes] sched_ext: Fix missing ops.dequeue() on remote local DSQ moves
Date: Wed, 30 Sep 2026 11:47:21 +0000	[thread overview]
Message-ID: <20260930114725.331370-1-jpiecuch@google.com> (raw)

Hi,

When a task in the BPF scheduler's custody is moved to another CPU's
local DSQ, ops.dequeue() is deferred until the task is picked and then
reported with SCX_DEQ_CORE_SCHED_EXEC, instead of being called without
flags when the task is inserted into the destination DSQ. Patch 1 fixes
this by reverting the enqueue side of b75aaea24c9f ("sched_ext: Properly
mark SCX-internal migrations via sticky_cpu"). Patch 2 adds a selftest.

For 7.1.y, patch 1 also needs 18d62044cda7 ("sched_ext: Preserve rq
tracking across local DSQ dispatch"), as noted in its stable tags.

This is based on sched_ext/for-7.3-fixes (d35a535d3e3e). Merging it into
sched_ext/for-next gives a trivial conflict in the selftests Makefile,
where enq_blocked was added next to dequeue_remote; keep both.

Testing was done on x86_64 in virtme-ng with 4 vCPUs in an SMT topology
(2 cores x 2 threads) and CONFIG_SCHED_CORE=y. Without patch 1,
dequeue_remote fails in every run (30/30), e.g.:

  sched_ext: dequeue_remote: dequeue_remote.bpf.c:141: 15 (rcu_preempt): late ops.dequeue() with SCX_DEQ_CORE_SCHED_EXEC (enq_cpu=3 cpu=2 seq=1)
     ...
     ops_dequeue+0x114/0x170
     set_next_task_scx+0x104/0x1e0
     __pick_next_task+0xc7/0x180
     __schedule+0x154/0x1870

and the full sched_ext selftest suite reports 31 passed, 1 skipped
(nohz_tick), 1 failed (dequeue_remote).

With patch 1, dequeue_remote passes in every run (30/30), with ~140k
custody enqueues per run, ~90k of them followed by the task running on
another CPU. The full suite reports 32 passed, 1 skipped (nohz_tick),
0 failed. dequeue_remote also passes 10/10 with the runner and its
workers sharing a core cookie, where legitimate core-sched picks out of
custody do happen.

v2:
 - Reordered to put the fix first and folded the Makefile entry into the
   test patch (Tejun, Andrea).
 - Fix: clear p->scx.sticky_cpu right after reading it, making the fix a
   revert of the enqueue side of b75aaea24c9f (Andrea). Reworded the
   comment and shortened the description (Tejun). Noted 18d62044cda7 as
   a 7.1.y prerequisite (Tejun, Andrea). Added Andrea's Reviewed-by.
 - Test: check p->core_cookie on SCX_DEQ_CORE_SCHED_EXEC instead of
   skipping when core scheduling is in use (Tejun).
 - Test: count CPUs with sched_getaffinity() (Tejun), pop past stale queue
   entries in ops.dispatch(), reset and print all counters per scenario
   (Tejun), destroy the struct_ops link and reap workers on error paths
   (Sashiko).
 - Test: deduplicated the descriptions and lowercased single-line
   comments (Tejun).

v1: https://lore.kernel.org/r/20260929161730.185271-1-jpiecuch@google.com

Thanks,
Kuba

Assisted-by: Claude:claude-opus-5.5

Kuba Piecuch (2):
  sched_ext: Call ops.dequeue() when a task arrives on a remote local
    DSQ
  selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ
    moves

 kernel/sched/ext/ext.c                        |  10 +-
 tools/testing/selftests/sched_ext/Makefile    |   1 +
 .../selftests/sched_ext/dequeue_remote.bpf.c  | 265 ++++++++++++++++++
 .../selftests/sched_ext/dequeue_remote.c      | 202 +++++++++++++
 4 files changed, 475 insertions(+), 3 deletions(-)
 create mode 100644 tools/testing/selftests/sched_ext/dequeue_remote.bpf.c
 create mode 100644 tools/testing/selftests/sched_ext/dequeue_remote.c


base-commit: d35a535d3e3e71f92a051d94e415f43612c4edfc
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


             reply	other threads:[~2026-09-30 11:47 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 11:47 Kuba Piecuch [this message]
2026-09-30 11:47 ` [PATCH v2 1/2] sched_ext: Call ops.dequeue() when a task arrives on a remote local DSQ Kuba Piecuch
2026-09-30 11:47 ` [PATCH v2 2/2] selftests/sched_ext: Add a test for ops.dequeue() on remote local DSQ moves Kuba Piecuch
2026-09-30 13:53   ` Andrea Righi
2026-09-30 14:27     ` Kuba Piecuch

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930114725.331370-1-jpiecuch@google.com \
    --to=jpiecuch@google.com \
    --cc=arighi@nvidia.com \
    --cc=changwoo@igalia.com \
    --cc=emil@etsalapatis.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.