From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 650BA30D40F for ; Sun, 23 Aug 2026 14:28:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=216.40.44.15 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787495297; cv=none; b=UMjv4zgadwKTvTamI9rrYdNq2yV35ABU0KLdBWfa3LSF/7G7uCsi76n689A+BCobRGO9mjM56X2vinp0NR5DG8ejPDLtLuAy7rqSH79GLPVQEvdL75tSqHEmf+8f1/2Lu0J1v8kSzdGAyvn5usCOULqjAqMQGzB0wiWNxWdUNeM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787495297; c=relaxed/simple; bh=BZQhJbqAapRYIxjdbW/Sm/6sEoHWmfHiHwxdGA+SQ0s=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=BG4Hunhza3cfC9rLMDEzyNhu7uvdviz6OpRX7DTiAQMK4ABsnXfXB8hKdtxp4t2IwIaYEY57tWYekpd8aPCepjSS/VnanaA1T+dZGp54qoSP0Nu8dQ14GZ3lnpPWsV2HMQppHLes6BiyO71dWVwaSIRrUPICf+Re3V/bJiMuPXA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=goodmis.org; spf=pass smtp.mailfrom=goodmis.org; dkim=pass (1024-bit key) header.d=goodmis.org header.i=@goodmis.org header.b=pTRTipk/; arc=none smtp.client-ip=216.40.44.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=goodmis.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=goodmis.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=goodmis.org header.i=@goodmis.org header.b="pTRTipk/" Received: from omf06.hostedemail.com (lb01a-stub [10.200.18.249]) by unirelay04.hostedemail.com (Postfix) with ESMTP id 214B41A0340; Sun, 23 Aug 2026 14:28:14 +0000 (UTC) Received: from [HIDDEN] (Authenticated sender: rostedt@goodmis.org) by omf06.hostedemail.com (Postfix) with ESMTPA id 86C452000F; Sun, 23 Aug 2026 14:28:12 +0000 (UTC) Date: Sun, 23 Aug 2026 10:28:11 -0400 From: Steven Rostedt To: Con Kolivas Cc: linux-kernel Subject: Re: [ANNOUNCE] linux-7.2-ck1, MuQSS CPU scheduler for linux-7.2 Message-ID: <20260823102811.14c3f543@fedora> In-Reply-To: References: X-Mailer: Claws Mail 4.4.0 (GTK 3.24.52; x86_64-redhat-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit X-Stat-Signature: xkmq3z48imizqxkry1z4zcjf8yiwypqr X-Rspamd-Server: rspamout04 X-Rspamd-Queue-Id: 86C452000F X-Session-Marker: 726F737465647440676F6F646D69732E6F7267 X-Session-ID: U2FsdGVkX18bNivES5WJmaMS0ctXCr2BDFr+8jqOE7s= DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=goodmis.org; h=date:from:to:cc:subject:message-id:in-reply-to:references:mime-version:content-type:content-transfer-encoding; s=dkim1; bh=NnLGNppk7e4p9e2cGHIGGz+jnOIuTfiCbNqmtSoEXUM=; b=pTRTipk/YtgT/r7Phg7fkR5iB01GEu5pWtosPkeStkt+ozP7HNQz4LJElH5nFgf3q6PjUuTPufQK9aASerfJBUW06+6uNYGfxs/xpY66d0qWSDwK+OCAjnH2nddlnNTxRtV8GEA2OZaYk+53bHNpxjdj6U3K8LWstttaDutJmV0= X-HE-Tag: 1787495292-395261 X-HE-Meta: U2FsdGVkX19BMkq55ll5IzJORQzqqSlDA7cllQuDPsU2AjBgNm/1LMMOWmBpeDWaeGD34RzO7jgWw//jXPpWNhwLKZW1LmZoi3MDex64QQXVt0Mjt7T+FhlOlbHEbCJ3paoKufbyww5YzgNz4AHbZcbAN2bLEnTs1KkQWz7rfnE3aqZfgbcqh4775leW6om1VdpQaQzwLmg+dHoYR9eTSHyBZXkXA16oMB1dg4snWp4orDKPGfZx5hhM5mvQf1Ho9+Ho8K0p927W4qwsbkVCsIOdlnl8fybmol0n41H27xiWr+xNUc1gHchv2b2VIQS72jkVYEV+Xg8NleJO4znQZ1EUmddbGmXnMFGxvMqymZ5aRfCBPqr6QMnViqwtA7+kuRowk5jxJKTWSyQe/DflsAabuG+464tclm9vuO80vHsV9VcfKE8jCwVlGabDiEPXqiZ0SWL9E333Or/E+oQdJw== On Sun, 23 Aug 2026 12:08:19 +1000 Con Kolivas wrote: > Hi again Steve et. al > > > > [1] https://lwn.net/Articles/1020596/ > > > [2] https://docs.google.com/presentation/d/1dtm0AiiTI30gTFeKj95vmSyirYmk5_QiR_lh17l_Moo/edit?usp=sharing > > I'm now caught up with your presentation as presented in the youtube > video linked in the lwn article. Thanks, very informative and > thoughtful. I was unable to access the google doc but have requested > read access - though I believe it was all presented on the video. > > Your governor idea is not as dissimilar to plugsched as may appear on > the surface. Plugsched built in all the schedulers into the kernel and > allowed you to boot the scheduler of your choice at boot time; it was > not to just build one scheduler into the kernel. Making it switch on > the fly was a pipe-dream goal but since it got shot down in > spectacular fashion I did not pursue it further. Its code is also so > outdated that literally nothing is of relevance in the current kernel > tree. Yeah, the main reason I call it a "governor" and not "plugable" is because I wanted to stress that it's not for any scheduler. It is for an environment. Similar to the power management governors. Saying "plugable" gives a "plug and play" feel that I wanted to avoid. Ingo's complaint back then was that we would have hundreds of schedulers. Now with sched_ext, that's exactly what we have. Thus, the governor was an idea to bring back collaboration between folks that work in the same environment. > > If you do pursue the governor idea there are a few things worth noting > about how high up and broad the hooks need to be. > One overhead problem with plugsched was it added a layer of > indirection to every single scheduler function call that was shared > between different schedulers. The cost of this may be considered > either trivially irrelevant or not remotely worth it depending on your > viewpoint. A the time I wrote plugsched, Itanic[sic] was still an > active architecture and the indirection was considered a huge > downside. I'm not sure you are aware that the Linux kernel today has something called a "static_call"[1]. It's a direct function call that can be switched at runtime with run-time code modification. It is something that I was planning on using. > There are four broad aspects to achieving low latency with muqss which > all need to be adopted to reproduce its behaviour, in order of > decreasing importance. > 1. Policy - the simple ordering aspect based on deadline, timeslice > interval etc. based on a shared monotonically increasing nanosecond > time counter. > 2. Shared access to a global queue - BFS did this by having only one > queue. MuQSS was created as a way to address scalability concerns by > reintroducing separate runqueues. It became clear very quickly that > policy alone did not reproduce the behaviour of BFS and that's where > the idea for having shared runqueues came about. The more the > runqueues were shared, the closer the latency approximated BFS'. The > default configuration chooses MC - Multicore. For virtually all > desktops and mobile devices that means they all end up with one > runqueue anyway. It is pre-configurable in kconfig, but also boot-time > selectable. > 3. Busy and idle load balancing. In MuQSS' case the busy balancing > happens by proxy through the next task selection, but idle balancing > is handled separately. Mainline handles both of these separately from > policy. For my idea, I would have the governors only affect SCHED_OTHER and SCHED_IDLE. The RT and DL schedulers would not be affected. But having the governor take over the idle and other tasks, it would have more control of what it could do. > 4. Highres timer based scheduling to effect the nanosecond timers. > This is to disentangle the scheduler's latency dependency on the > chosen jiffy Hz which ties all other subsystem components to that > resolution and/or overhead. I believe mainline is going in this direction too. > > None of these are insurmountable endpoints with enough LLM tokens, but > I suspect there will at least be one/some indirection somewhere in the > implementation. Again, indirection is solved by the static calls. When spectre and meltdown mitigations were introduced into the kernel, indirect calls took a hugh hit. Tracepoints used them quite extensively and it slowed down hackbench with the sched_switch tracepoint active by 10%! I started working on a way to replace indirect calls with a direct call that could be modified. Then Josh Poimbeouf took over and Linus didn't like the implementation.Finally, Peter Zijlstra got it into the kernel. They are expensive to change, but for things that do not change often (like enabling a tracepoint, picking a KVM implementation, or picking a specific scheduler governor) it is really useful. The hackbench slowdown disappeared and tracepoints are actually even faster than what they were with the indirect calls. -- Steve [1] https://lwn.net/Articles/815908/