linux-kernel.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: Hao Jia <jiahao.kernel@gmail.com>
To: Chris Friesen <chris.friesen@windriver.com>,
	LKML <linux-kernel@vger.kernel.org>,
	hanguangjiang@lixiang.com
Cc: osandov@fb.com, Peter Zijlstra <peterz@infradead.org>
Subject: Re: sched: observed instability under stress in 6.12 and mainline
Date: Mon, 8 Sep 2025 10:22:31 +0800	[thread overview]
Message-ID: <33cac97d-d73c-98e7-3a8b-852460cb1909@gmail.com> (raw)
In-Reply-To: <7cd74213-5654-aac0-54d0-4f4b1a7f0fef@gmail.com>



On 2025/9/8 09:51, Hao Jia wrote:
> 
> 
> On 2025/9/5 00:33, Chris Friesen wrote:
>> Hi,
>>
>> I'd like to draw the attention of the scheduler maintainers to a 
>> number of kernel bugzilla reports submitted by a colleague a couple of 
>> weeks ago:
>>
>> 6.12.18:
>> https://bugzilla.kernel.org/show_bug.cgi?id=220447
>> https://bugzilla.kernel.org/show_bug.cgi?id=220448
>>
>> v6.16-rt3
>> https://bugzilla.kernel.org/show_bug.cgi?id=220450
>> https://bugzilla.kernel.org/show_bug.cgi?id=220449
>>
>> There seems to be something wrong with either the logic or the 
>> locking. In one case this resulted in a NULL pointer dereference in 
>> pick_next_entity().  In another case it resulted in 
>> BUG_ON(!rq->nr_running) in dequeue_top_rt_rq() and 
>> SCHED_WARN_ON(!se->on_rq) in update_entity_lag().
>>
>> My colleague suggests that the NULL pointer dereference may be due to 
>> pick_eevdf() returning NULL in pick_next_entity().
>>
>> I did some digging and found that 
>> https://gitlab.com/linux-kernel/stable/-/commit/86b37810 would not 
>> have been included in 6.12.18, but the equivalent fix should have been 
>> in the 6.16 load.
>>
>> We haven't yet bottomed out the root cause.
>>
>> Any suggestions or assistance would be appreciated.
>>
>> Thanks,
>> Chris
>>
>>
> 
> Maybe this patch can be useful for your problem.
> https://lore.kernel.org/all/tencent_3177343A3163451463643E434C61911B4208@qq.com/
> 
> If I understand correctly, we may dequeue_entity twice in 
> rt_mutex_setprio()/__sched_setscheduler(). cfs_bandwidth may break the 
> state of p->on_rq and se->on_rq.
> 

Perhaps the "Defer throttle when task exits to user" patch set from the 
sched/core branch can also fix this bug.

https://lore.kernel.org/all/20250829081120.806-4-ziqianlu@bytedance.com/

Thanks,
Hao

  reply	other threads:[~2025-09-08  2:22 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-09-04 16:33 sched: observed instability under stress in 6.12 and mainline Chris Friesen
2025-09-08  1:51 ` Hao Jia
2025-09-08  2:22   ` Hao Jia [this message]
2025-10-13  3:03   ` Jiping Ma
2025-10-13  5:54     ` Hao Jia

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=33cac97d-d73c-98e7-3a8b-852460cb1909@gmail.com \
    --to=jiahao.kernel@gmail.com \
    --cc=chris.friesen@windriver.com \
    --cc=hanguangjiang@lixiang.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=osandov@fb.com \
    --cc=peterz@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).