From: Frederic Weisbecker <frederic@kernel.org>
To: Peter Zijlstra <peterz@infradead.org>
Cc: LKML <linux-kernel@vger.kernel.org>,
Alexey Dobriyan <adobriyan@gmail.com>,
Wei Li <liwei391@huawei.com>,
Mirsad Goran Todorovac <mirsad.todorovac@alu.unizg.hr>,
Thomas Gleixner <tglx@linutronix.de>,
Yu Liao <liaoyu15@huawei.com>, Hillf Danton <hdanton@sina.com>,
Ingo Molnar <mingo@kernel.org>
Subject: Re: [PATCH 4/6] timers/nohz: Add a comment about broken iowait counter update race
Date: Fri, 10 Feb 2023 17:10:03 +0100 [thread overview]
Message-ID: <Y+ZsWx4gnx4Cak7D@lothringen> (raw)
In-Reply-To: <Y+ZXLy0P0Sggbrxc@hirez.programming.kicks-ass.net>
On Fri, Feb 10, 2023 at 03:39:43PM +0100, Peter Zijlstra wrote:
> On Fri, Feb 10, 2023 at 03:09:15PM +0100, Frederic Weisbecker wrote:
> > The per-cpu iowait task counter is incremented locally upon sleeping.
> > But since the task can be woken to (and by) another CPU, the counter may
> > then be decremented remotely. This is the source of a race involving
> > readers VS writer of idle/iowait sleeptime.
> >
> > The following scenario shows an example where a /proc/stat reader
> > observes a pending sleep time as IO whereas that pending sleep time
> > later eventually gets accounted as non-IO.
> >
> > CPU 0 CPU 1 CPU 2
> > ----- ----- ------
> > //io_schedule() TASK A
> > current->in_iowait = 1
> > rq(0)->nr_iowait++
> > //switch to idle
> > // READ /proc/stat
> > // See nr_iowait_cpu(0) == 1
> > return ts->iowait_sleeptime +
> > ktime_sub(ktime_get(), ts->idle_entrytime)
> >
> > //try_to_wake_up(TASK A)
> > rq(0)->nr_iowait--
> > //idle exit
> > // See nr_iowait_cpu(0) == 0
> > ts->idle_sleeptime += ktime_sub(ktime_get(), ts->idle_entrytime)
> >
> > As a result subsequent reads on /proc/stat may expose backward progress.
> >
> > This is unfortunately hardly fixable. Just add a comment about that
> > condition.
>
> It is far worse than that, the whole concept of per-cpu iowait is
> absurd. Also see the comment near nr_iowait().
Alas I know :-(
next prev parent reply other threads:[~2023-02-10 16:10 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-02-10 14:09 [PATCH 0/6] timers/nohz: Fixes and cleanups Frederic Weisbecker
2023-02-10 14:09 ` [PATCH 1/6] timers/nohz: Restructure and reshuffle struct tick_sched Frederic Weisbecker
2023-02-10 14:09 ` [PATCH 2/6] timers/nohz: Only ever update sleeptime from idle exit Frederic Weisbecker
2023-02-10 14:09 ` [PATCH 3/6] timers/nohz: Protect idle/iowait sleep time under seqcount Frederic Weisbecker
2023-02-10 14:09 ` [PATCH 4/6] timers/nohz: Add a comment about broken iowait counter update race Frederic Weisbecker
2023-02-10 14:39 ` Peter Zijlstra
2023-02-10 16:10 ` Frederic Weisbecker [this message]
2023-02-10 14:09 ` [PATCH 5/6] timers/nohz: Remove middle-function __tick_nohz_idle_stop_tick() Frederic Weisbecker
2023-02-10 14:09 ` [PATCH 6/6] MAINTAINERS: Remove stale email address Frederic Weisbecker
2023-02-10 14:41 ` [PATCH 0/6] timers/nohz: Fixes and cleanups Peter Zijlstra
2023-02-13 18:36 ` Thomas Gleixner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=Y+ZsWx4gnx4Cak7D@lothringen \
--to=frederic@kernel.org \
--cc=adobriyan@gmail.com \
--cc=hdanton@sina.com \
--cc=liaoyu15@huawei.com \
--cc=linux-kernel@vger.kernel.org \
--cc=liwei391@huawei.com \
--cc=mingo@kernel.org \
--cc=mirsad.todorovac@alu.unizg.hr \
--cc=peterz@infradead.org \
--cc=tglx@linutronix.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox