From: Lance Yang <lance.yang@linux.dev>
To: lirongqing <lirongqing@baidu.com>, Petr Mladek <pmladek@suse.com>
Cc: wireguard@lists.zx2c4.com, linux-arm-kernel@lists.infradead.org,
"Liam R . Howlett" <Liam.Howlett@oracle.com>,
linux-doc@vger.kernel.org, David Hildenbrand <david@redhat.com>,
Randy Dunlap <rdunlap@infradead.org>,
Stanislav Fomichev <sdf@fomichev.me>,
linux-aspeed@lists.ozlabs.org,
Andrew Jeffery <andrew@codeconstruct.com.au>,
Joel Stanley <joel@jms.id.au>,
Russell King <linux@armlinux.org.uk>,
Lorenzo Stoakes <lorenzo.stoakes@oracle.com>,
Shuah Khan <shuah@kernel.org>,
Steven Rostedt <rostedt@goodmis.org>,
Jonathan Corbet <corbet@lwn.net>,
Joel Granados <joel.granados@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
Phil Auld <pauld@redhat.com>,
linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org,
Masami Hiramatsu <mhiramat@kernel.org>,
Jakub Kicinski <kuba@kernel.org>,
Pawan Gupta <pawan.kumar.gupta@linux.intel.com>,
Simon Horman <horms@kernel.org>,
Anshuman Khandual <anshuman.khandual@arm.com>,
Florian Westphal <fw@strlen.de>,
netdev@vger.kernel.org, Kees Cook <kees@kernel.org>,
Arnd Bergmann <arnd@arndb.de>,
"Paul E . McKenney" <paulmck@kernel.org>,
Feng Tang <feng.tang@linux.alibaba.com>,
"Jason A . Donenfeld" <Jason@zx2c4.com>
Subject: Re: [PATCH][v3] hung_task: Panic after fixed number of hung tasks
Date: Tue, 14 Oct 2025 18:59:07 +0800 [thread overview]
Message-ID: <3acdcd15-7e52-4a9a-9492-a434ed609dcc@linux.dev> (raw)
In-Reply-To: <aO4boXFaIb0_Wiif@pathway.suse.cz>
On 2025/10/14 17:45, Petr Mladek wrote:
> On Tue 2025-10-14 13:23:58, Lance Yang wrote:
>> Thanks for the patch!
>>
>> I noticed the implementation panics only when N tasks are detected
>> within a single scan, because total_hung_task is reset for each
>> check_hung_uninterruptible_tasks() run.
>
> Great catch!
>
> Does it make sense?
> Is is the intended behavior, please?
>
>> So some suggestions to align the documentation with the code's
>> behavior below :)
>
>> On 2025/10/12 19:50, lirongqing wrote:
>>> From: Li RongQing <lirongqing@baidu.com>
>>>
>>> Currently, when 'hung_task_panic' is enabled, the kernel panics
>>> immediately upon detecting the first hung task. However, some hung
>>> tasks are transient and the system can recover, while others are
>>> persistent and may accumulate progressively.
>
> My understanding is that this patch wanted to do:
>
> + report even temporary stalls
> + panic only when the stall was much longer and likely persistent
>
> Which might make some sense. But the code does something else.
Cool. Sounds good to me!
>
>>> --- a/kernel/hung_task.c
>>> +++ b/kernel/hung_task.c
>>> @@ -229,9 +232,11 @@ static void check_hung_task(struct task_struct *t, unsigned long timeout)
>>> */
>>> sysctl_hung_task_detect_count++;
>>> + total_hung_task = sysctl_hung_task_detect_count - prev_detect_count;
>>> trace_sched_process_hang(t);
>>> - if (sysctl_hung_task_panic) {
>>> + if (sysctl_hung_task_panic &&
>>> + (total_hung_task >= sysctl_hung_task_panic)) {
>>> console_verbose();
>>> hung_task_show_lock = true;
>>> hung_task_call_panic = true;
>
> I would expect that this patch added another counter, similar to
> sysctl_hung_task_detect_count. It would be incremented only
> once per check when a hung task was detected. And it would
> be cleared (reset) when no hung task was found.
Much cleaner. We could add an internal counter for that, yeah. No need
to expose it to userspace ;)
Petr's suggestion seems to align better with the goal of panicking on
persistent hangs, IMHO. Panic after N consecutive checks with hung tasks.
@RongQing does that work for you?
next prev parent reply other threads:[~2025-10-14 21:59 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-10-12 11:50 [PATCH][v3] hung_task: Panic after fixed number of hung tasks lirongqing
2025-10-12 13:26 ` [PATCH v3] " Markus Elfring
2025-10-13 2:14 ` [外部邮件] " Li,Rongqing
2025-10-14 1:37 ` [PATCH][v3] " Randy Dunlap
2025-10-14 2:03 ` [外部邮件] " Li,Rongqing
2025-10-14 5:23 ` Lance Yang
2025-10-14 5:33 ` [外部邮件] " Li,Rongqing
2025-10-14 9:45 ` Petr Mladek
2025-10-14 10:49 ` [????] " Li,Rongqing
2025-10-14 13:09 ` Petr Mladek
2025-10-15 2:04 ` [????] " Li,Rongqing
2025-10-14 10:59 ` Lance Yang [this message]
2025-10-14 11:18 ` [外部邮件] " Li,Rongqing
2025-10-14 11:40 ` Lance Yang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=3acdcd15-7e52-4a9a-9492-a434ed609dcc@linux.dev \
--to=lance.yang@linux.dev \
--cc=Jason@zx2c4.com \
--cc=Liam.Howlett@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=andrew@codeconstruct.com.au \
--cc=anshuman.khandual@arm.com \
--cc=arnd@arndb.de \
--cc=corbet@lwn.net \
--cc=david@redhat.com \
--cc=feng.tang@linux.alibaba.com \
--cc=fw@strlen.de \
--cc=horms@kernel.org \
--cc=joel.granados@kernel.org \
--cc=joel@jms.id.au \
--cc=kees@kernel.org \
--cc=kuba@kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-aspeed@lists.ozlabs.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux@armlinux.org.uk \
--cc=lirongqing@baidu.com \
--cc=lorenzo.stoakes@oracle.com \
--cc=mhiramat@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pauld@redhat.com \
--cc=paulmck@kernel.org \
--cc=pawan.kumar.gupta@linux.intel.com \
--cc=pmladek@suse.com \
--cc=rdunlap@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sdf@fomichev.me \
--cc=shuah@kernel.org \
--cc=wireguard@lists.zx2c4.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.