From: Peter Zijlstra <peterz@infradead.org>
To: David Stevens <stevensd@google.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>, Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H . Peter Anvin" <hpa@zytor.com>,
Andrew Morton <akpm@linux-foundation.org>,
Dave Chinner <david@fromorbit.com>, Qi Zheng <qi.zheng@linux.dev>,
Roman Gushchin <roman.gushchin@linux.dev>,
Muchun Song <muchun.song@linux.dev>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Uladzislau Rezki <urezki@gmail.com>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R . Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Kees Cook <kees@kernel.org>,
Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
Clark Williams <clrkwllms@kernel.org>,
suleiman@google.com, linux-kernel@vger.kernel.org,
linux-arm-kernel@lists.infradead.org, linux-mm@kvack.org,
linux-rt-devel@lists.linux.dev
Subject: Re: [RFC 06/10] Reclaim memory from blocked kernel stacks
Date: Sat, 29 Aug 2026 10:39:29 +0200 [thread overview]
Message-ID: <20260829083929.GZ776954@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <CAOiLmNHWGSwAqH815X3JWMrQCc2nh-fRNsNZ0hjQGv8JvFndag@mail.gmail.com>
On Fri, Aug 28, 2026 at 05:18:05PM -0700, David Stevens wrote:
> On Fri, Aug 28, 2026 at 5:04 AM Peter Zijlstra <peterz@infradead.org> wrote:
> >
> > On Thu, Aug 27, 2026 at 04:29:44PM -0700, David Stevens wrote:
> > > @@ -4320,8 +4319,18 @@ int try_to_wake_up(struct task_struct *p, unsigned int state, int wake_flags)
> > > * A similar smp_rmb() lives in __task_needs_rq_lock().
> > > */
> > > smp_rmb();
> > > - if (READ_ONCE(p->on_rq) && ttwu_runnable(p, wake_flags))
> > > + if (READ_ONCE(p->on_rq) && ttwu_runnable(p, wake_flags)) {
> > > + trace_sched_waking(p);
> > > + break;
> > > + }
> > > +
> > > + if (!ensure_stack_is_present(p, &need_deferred_repopulate)) {
> > > + WRITE_ONCE(p->__state, TASK_STACK_RECLAIM);
> > > + do_deferred_repopulate_wake = need_deferred_repopulate;
> > > break;
> > > + }
> > > +
> > > + trace_sched_waking(p);
> >
> > Absolutely not; ensure_stack_is_present() must not call
> > repopulate_stack() while holding ->pi_lock. Not happening.
>
> The optimistic fast path for repopulate_stack() could be modified to
> try pulling from a pre-allocated pool of zero'ed pages. That would
> reduce the function to a couple of memcg_kmem_charge_page() calls and
Afaict memcg_kmem_charge_page() ends up in a local_lock, which is a
spinlock, so that cannot be.
Most, if not everything, in mm/ is build around being preemptible and
thus not suitable for use under raw_spinlock_t.
> then vmap_pages_range() to repopulate the stack's page tables. That
vmap_page_range() can end up in the allocator, which I suppose is ruled
out by the vmap having been populated before, but it still has a
might_sleep() that will scream AFAICT.
> wouldn't require touching any locks except a raw_spinlock protecting
> the pre-allocated pool (or just make it per_cpu). In terms of cost,
> this would involve a couple of atomic operations for the page pool
> lock and the memcg charging plus non-atomic operations on 5-10 other
> cache lines.
>
> Is that within the scope of what can be done under the pi_lock? If
> that's still not happening, I can see how things look if we always
> defer wakeup to a workqueue.
As long as it really is all atomics it should be fine. If there is a
lock, it must be raw_spinlock_t, but ideally no new locks nested under
pi_lock.
next prev parent reply other threads:[~2026-08-29 8:39 UTC|newest]
Thread overview: 40+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 23:29 [RFC 00/10] Reclaimable kernel stacks David Stevens
2026-08-27 23:29 ` [RFC 01/10] Add !MEMCG memcg_list_lru_alloc implementation David Stevens
2026-08-27 23:29 ` [RFC 02/10] mm/vmalloc: Skip vmallocinfo NUMA stats for VM_SPARSE David Stevens
2026-08-27 23:29 ` [RFC 03/10] fork: refactor vmap stack alloc/free into helpers David Stevens
2026-08-27 23:29 ` [RFC 04/10] mm: vmalloc: support creating aligned vm areas David Stevens
2026-08-27 23:29 ` [RFC 05/10] fork: allocate reclaimable stacks with VM_SPARSE David Stevens
2026-08-27 23:29 ` [RFC 06/10] Reclaim memory from blocked kernel stacks David Stevens
2026-08-27 23:53 ` sashiko-bot
2026-08-28 11:54 ` Peter Zijlstra
2026-08-28 12:01 ` Peter Zijlstra
2026-08-28 12:04 ` Peter Zijlstra
2026-08-29 0:18 ` David Stevens
2026-08-29 8:39 ` Peter Zijlstra [this message]
2026-08-29 8:43 ` Peter Zijlstra
2026-08-28 12:41 ` Peter Zijlstra
2026-08-28 12:57 ` Peter Zijlstra
2026-08-28 23:33 ` David Stevens
2026-08-28 13:36 ` Sebastian Andrzej Siewior
2026-08-28 13:59 ` Peter Zijlstra
2026-08-28 14:25 ` Peter Zijlstra
2026-08-28 15:58 ` Sebastian Andrzej Siewior
2026-08-28 15:10 ` Sebastian Andrzej Siewior
2026-08-28 19:08 ` Steven Rostedt
2026-08-28 19:13 ` Steven Rostedt
2026-08-28 19:17 ` Steven Rostedt
2026-08-28 20:50 ` David Stevens
2026-08-29 8:49 ` Peter Zijlstra
2026-08-28 21:17 ` David Stevens
2026-08-27 23:29 ` [RFC 07/10] Reclaim stacks via a shrinker David Stevens
2026-08-27 23:29 ` [RFC 08/10] Set PF_RECLAIMABLE_STACK in various places David Stevens
2026-08-27 23:43 ` sashiko-bot
2026-08-28 6:33 ` K Prateek Nayak
2026-08-27 23:29 ` [RFC 09/10] x86: Enable reclaimable stacks David Stevens
2026-08-27 23:29 ` [RFC 10/10] arm64: " David Stevens
2026-08-28 12:47 ` [RFC 00/10] Reclaimable kernel stacks Peter Zijlstra
2026-08-28 14:33 ` Steven Rostedt
2026-08-28 14:35 ` Peter Zijlstra
2026-08-28 14:45 ` Peter Zijlstra
2026-08-28 16:10 ` Steven Rostedt
2026-08-28 17:58 ` David Stevens
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260829083929.GZ776954@noisy.programming.kicks-ass.net \
--to=peterz@infradead.org \
--cc=akpm@linux-foundation.org \
--cc=bigeasy@linutronix.de \
--cc=bp@alien8.de \
--cc=bsegall@google.com \
--cc=catalin.marinas@arm.com \
--cc=clrkwllms@kernel.org \
--cc=dave.hansen@linux.intel.com \
--cc=david@fromorbit.com \
--cc=david@kernel.org \
--cc=dietmar.eggemann@arm.com \
--cc=hpa@zytor.com \
--cc=juri.lelli@redhat.com \
--cc=kees@kernel.org \
--cc=kprateek.nayak@amd.com \
--cc=liam@infradead.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-rt-devel@lists.linux.dev \
--cc=ljs@kernel.org \
--cc=mgorman@suse.de \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=muchun.song@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=stevensd@google.com \
--cc=suleiman@google.com \
--cc=surenb@google.com \
--cc=tglx@kernel.org \
--cc=urezki@gmail.com \
--cc=vbabka@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=will@kernel.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox