From: Petr Mladek <pmladek@suse.com>
To: Lance Yang <lance.yang@linux.dev>
Cc: Aaron Tomlin <atomlin@atomlin.com>,
sean@ashe.io, linux-kernel@vger.kernel.org, mhiramat@kernel.org,
akpm@linux-foundation.org, gregkh@linuxfoundation.org
Subject: Re: [PATCH v3 2/2] hung_task: Enable runtime reset of hung_task_detect_count
Date: Wed, 17 Dec 2025 14:09:14 +0100 [thread overview]
Message-ID: <aUKresftPnbndSBo@pathway.suse.cz> (raw)
In-Reply-To: <79e00527-0ada-4c43-915e-1f8f6589d051@linux.dev>
On Wed 2025-12-17 15:31:25, Lance Yang wrote:
> On 2025/12/16 11:00, Aaron Tomlin wrote:
> > Introduce support for writing to /proc/sys/kernel/hung_task_detect_count.
> >
> > Writing any value to this file atomically resets the counter of detected
> > hung tasks to zero. This grants system administrators the ability to clear
> > the cumulative diagnostic history after resolving an incident, simplifying
> > monitoring without requiring a system restart.
>
> > --- a/Documentation/admin-guide/sysctl/kernel.rst
> > +++ b/Documentation/admin-guide/sysctl/kernel.rst
> > @@ -418,7 +418,7 @@ hung_task_detect_count
> > ======================
> > Indicates the total number of tasks that have been detected as hung since
> > -the system boot.
> > +the system boot. The counter can be reset to zero when written to.
> > This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
> > diff --git a/kernel/hung_task.c b/kernel/hung_task.c
> > index 5902573200c0..01ce46a107b0 100644
> > --- a/kernel/hung_task.c
> > +++ b/kernel/hung_task.c
> > @@ -375,6 +375,31 @@ static long hung_timeout_jiffies(unsigned long last_checked,
> > }
> > #ifdef CONFIG_SYSCTL
> > +
> > +/**
> > + * proc_dohung_task_detect_count - proc handler for hung_task_detect_count
> > + * @table: Pointer to the struct ctl_table definition for this proc entry
> > + * @write: Flag indicating the operation
> > + * @buffer: User space buffer for data transfer
> > + * @lenp: Pointer to the length of the data being transferred
> > + * @ppos: Pointer to the current file offset
> > + *
> > + * This handler is used for reading the current hung task detection count
> > + * and for resetting it to zero when a write operation is performed.
> > + * Returns 0 on success or a negative error code on failure.
> > + */
> > +static int proc_dohung_task_detect_count(const struct ctl_table *table, int write,
> > + void *buffer, size_t *lenp, loff_t *ppos)
> > +{
> > + if (!write)
> > + return proc_doulongvec_minmax(table, write, buffer, lenp, ppos);
> > +
> > + WRITE_ONCE(sysctl_hung_task_detect_count, 0);
>
> The reset uses WRITE_ONCE() but the increment in check_hung_task()
> is plain ++, it's just a stat cunter so fine though ;)
I was about to wave this away as well. But I am afraid that
it is even more complicated.
The counter is used to decide how many hung tasks were found in
by a single check, see:
static void check_hung_uninterruptible_tasks(unsigned long timeout)
{
[...]
unsigned long prev_detect_count = sysctl_hung_task_detect_count;
[...]
for_each_process_thread(g, t) {
[...]
check_hung_task(t, timeout, prev_detect_count);
}
[...]
if (!(sysctl_hung_task_detect_count - prev_detect_count))
return;
if (need_warning || hung_task_call_panic) {
si_mask |= SYS_INFO_LOCKS;
if (sysctl_hung_task_all_cpu_backtrace)
si_mask |= SYS_INFO_ALL_BT;
}
sys_info(si_mask);
if (hung_task_call_panic)
panic("hung_task: blocked tasks");
}
Any race might cause false positives and even panic() !!!
And the race window is rather big (checking all processes).
This race can't be prevented by an atomic read/write.
IMHO, the only solution would be to add some locking,
e.g. mutex and take it around the entire
check_hung_uninterruptible_tasks().
On one hand, it might work. The code is called from
a task context...
But adding a lock into a lockup-detector code triggers
some warning bells in my head.
Honestly, I am not sure if it is worth it. I understand
the motivation for the reset. But IMHO, even more important
is to make sure that the watchdog works as expected.
Best Regards,
Petr
next prev parent reply other threads:[~2025-12-17 13:09 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-12-16 3:00 [PATCH v3 0/2] hung_task: Provide runtime reset interface for hung task detector Aaron Tomlin
2025-12-16 3:00 ` [PATCH v3 1/2] hung_task: Introduce helper for hung task warning Aaron Tomlin
2025-12-17 9:39 ` Petr Mladek
2025-12-21 21:52 ` Aaron Tomlin
2025-12-16 3:00 ` [PATCH v3 2/2] hung_task: Enable runtime reset of hung_task_detect_count Aaron Tomlin
2025-12-17 7:31 ` Lance Yang
2025-12-17 13:09 ` Petr Mladek [this message]
2025-12-17 13:36 ` Lance Yang
2025-12-19 3:09 ` Aaron Tomlin
2025-12-19 12:03 ` Petr Mladek
2025-12-19 14:15 ` Lance Yang
2025-12-21 21:00 ` Aaron Tomlin
2025-12-17 12:48 ` Petr Mladek
2025-12-17 13:21 ` Lance Yang
2025-12-18 9:08 ` Joel Granados
2025-12-19 14:18 ` Lance Yang
2025-12-21 21:26 ` Aaron Tomlin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aUKresftPnbndSBo@pathway.suse.cz \
--to=pmladek@suse.com \
--cc=akpm@linux-foundation.org \
--cc=atomlin@atomlin.com \
--cc=gregkh@linuxfoundation.org \
--cc=lance.yang@linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=mhiramat@kernel.org \
--cc=sean@ashe.io \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.