From: Peter Zijlstra <peterz@infradead.org>
To: venkatesh.pallipadi@intel.com
Cc: Gautham R Shenoy <ego@in.ibm.com>,
Vaidyanathan Srinivasan <svaidy@linux.vnet.ibm.com>,
Ingo Molnar <mingo@elte.hu>, Thomas Gleixner <tglx@linutronix.de>,
Arjan van de Ven <arjan@infradead.org>,
linux-kernel@vger.kernel.org,
Suresh Siddha <suresh.b.siddha@intel.com>
Subject: Re: [patch 2/2] sched: Scale the nohz_tracker logic by making it per NUMA node
Date: Mon, 21 Dec 2009 14:11:46 +0100 [thread overview]
Message-ID: <1261401106.4314.137.camel@laptop> (raw)
In-Reply-To: <20091211013056.450920000@intel.com>
On Thu, 2009-12-10 at 17:27 -0800, venkatesh.pallipadi@intel.com wrote:
> plain text document attachment
> (0002-sched-Scale-the-nohz_tracker-logic-by-making-it-per.patch)
> Having one idle CPU doing the rebalancing for all the idle CPUs in
> nohz mode does not scale well with increasing number of cores and
> sockets. Make the nohz_tracker per NUMA node. This results in multiple
> idle load balancing happening at NUMA node level and idle load balancer
> only does the rebalance domain among all the other nohz CPUs in that
> NUMA node.
>
> This addresses the below problem with the current nohz ilb logic
> * The lone balancer may end up spending a lot of time doing the
> * balancing on
> behalf of nohz CPUs, especially with increasing number of sockets and
> cores in the platform.
Right, so I think the whole NODE idea here is wrong, it all seems to
work out properly if you simply pick one sched domain larger than the
one that contains all of the current socket and contains an idle unit.
Except that the sched domain stuff is not properly aware of bigger
topology things atm.
The sched domain tree should not view node as the largest structure and
we should remove that current random node split crap we have.
Instead the sched domains should continue to express the topology, like
nodes within 1 hop, nodes within 2 hops, etc.
Then this nohz idle balancing should pick the socket level (which might
be larger than the node level), and walks up the domain tree, until we
reach a level where it has a whole idle group.
This means that we'll always span at least 2 sockets, which means we'll
gracefully deal with the overload scenario.
prev parent reply other threads:[~2009-12-21 13:12 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2009-12-11 1:27 [patch 0/2] sched: Change nohz ilb logic from pull to push model venkatesh.pallipadi
2009-12-11 1:27 ` [patch 1/2] sched: Change the " venkatesh.pallipadi
2009-12-14 22:18 ` Peter Zijlstra
2009-12-21 12:13 ` Peter Zijlstra
2009-12-21 13:00 ` Peter Zijlstra
2009-12-23 0:15 ` Pallipadi, Venkatesh
2009-12-11 1:27 ` [patch 2/2] sched: Scale the nohz_tracker logic by making it per NUMA node venkatesh.pallipadi
2009-12-14 22:21 ` Peter Zijlstra
2009-12-14 22:32 ` Pallipadi, Venkatesh
2009-12-14 22:58 ` Peter Zijlstra
2009-12-15 1:00 ` Pallipadi, Venkatesh
2009-12-15 10:21 ` Peter Zijlstra
2009-12-21 13:11 ` Peter Zijlstra [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1261401106.4314.137.camel@laptop \
--to=peterz@infradead.org \
--cc=arjan@infradead.org \
--cc=ego@in.ibm.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@elte.hu \
--cc=suresh.b.siddha@intel.com \
--cc=svaidy@linux.vnet.ibm.com \
--cc=tglx@linutronix.de \
--cc=venkatesh.pallipadi@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.