From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B17743B0AED for ; Mon, 20 Jul 2026 14:55:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784559344; cv=none; b=E1HURuZK9pOSCg+t0ccaLx2SCKLY6cmvazALpopN8S9Tqg3liwCYtCzxGyv9pI5R3JAOuPY5LQy1K3VtlmAZXWV70xnbWFnlZHPaCLKAsEUdOiPC1/gDJRwOrHYL344qNBNGGfi2zFBrPP5muyo1kBTLJXdtEC7jjnhVqMllqeU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784559344; c=relaxed/simple; bh=nMkihhtdVnIMv0IX212u0tTZ3JriE6BXLZw2Tneet+I=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=X+YpV8rTrVQNKpXB/WyqAOrWpc+5w7H88Lk+S64i0MKvkuaf0sg/hn6Sh8Mh7QKYTUjlwRTu+BOK3Y2/+41gA81D1Biz2GMr5s8eSteOh9yD149L2OfIhPvqrk09JiJ5z/rGEVjMwI+iaRtcpg6mX0Uv0jmlRjmL3yH/WgF70g8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=FS0yNZ1Q; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=9ZHqkb2a; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="FS0yNZ1Q"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="9ZHqkb2a" Date: Mon, 20 Jul 2026 16:55:39 +0200 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1784559340; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=afXqtlbn3lqINCIWtGAhI4ma+pzsU1WiINToBwwzt7c=; b=FS0yNZ1QE2HaKz8MK9U8MhgU3LY1Vz0KP51u70sVJ7Mu49ob6ZSQUNzB02W6Ga9qBDmHdC CMydtIASWpnDhkWGthUwTfGe9ZUHGQnhXiyn5yYpZJ20kBcV2xYhfF7irUD3FGGZVe3FsJ Pt1FeN48DmCGdkrEyYfKE1yQN+S1kjoJ0uFwDvWY8iNwbndplu7CbaG4WuRTBvFjHuvl0X iaJfg7xfWvVUZSLT3I4yUqMEP4m51wKscW6KsktqyuhBsas8IrWp3cyE3a46lybdnG4H/J Md88PBWHatgngQJ2JxZSq37XMMi+MexRb1B5ojHb2QKVPhyWmWJqepW4yIen9Q== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1784559340; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=afXqtlbn3lqINCIWtGAhI4ma+pzsU1WiINToBwwzt7c=; b=9ZHqkb2ap6IW8RBF9EowDQDj5AXMq8Kvw+NV+75pmnHFjLCFjw3SVMfUZniMkk2cXGNH/1 u5Kg6smF38iU2SAw== From: Sebastian Andrzej Siewior To: Yao Kai Cc: linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, dvhart@infradead.org, dave@stgolabs.net, andrealmeid@igalia.com, liuyongqiang13@huawei.com Subject: Re: [PATCH 2/2] futex/requeue: Prevent rcuwait use-after-free during requeue PI Message-ID: <20260720145539.nNAar9Bm@linutronix.de> References: <20260717084922.4153317-1-yaokai34@huawei.com> <20260717084922.4153317-3-yaokai34@huawei.com> <20260717093829.K07Gk1tS@linutronix.de> <235279a2-e6e4-4b6a-b76b-8a972a3e2af8@huawei.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline In-Reply-To: <235279a2-e6e4-4b6a-b76b-8a972a3e2af8@huawei.com> On 2026-07-20 10:50:55 [+0800], Yao Kai wrote: > > > On 7/17/2026 5:38 PM, Sebastian Andrzej Siewior wrote: > > On 2026-07-17 16:49:22 [+0800], Yao Kai wrote: > > > On PREEMPT_RT, FUTEX_CMP_REQUEUE_PI can trigger a KASAN report: > > > > > > BUG: KASAN: slab-out-of-bounds in _raw_spin_lock_irqsave+0x76/0xe0 > > > Call Trace: > > > _raw_spin_lock_irqsave+0x76/0xe0 > > > try_to_wake_up+0xab/0x1540 > > > rcuwait_wake_up+0x39/0x60 > > > futex_requeue+0x18c3/0x1e10 > > > > > > The futex_q used by futex_wait_requeue_pi() is allocated on the waiter's > > > stack. An early wakeup can race with a PI requeue as follows: > > > > > > waiter requeue task > > > ------ ------------ > > > futex_wait_requeue_pi() > > > futex_do_wait() > > > schedule() > > > > > > * timeout/signal wakes waiter * > > > > > > futex_requeue_pi_wakeup_sync() > > > IN_PROGRESS -> WAIT > > > rcuwait_wait_event() > > > requeue_pi_wake_futex() > > > task = READ_ONCE(q->task) > > > futex_requeue_pi_complete() > > > WAIT -> LOCKED > > > return LOCKED > > > return > > > // q lifetime ends > > > rcuwait_wake_up() > > > > > > futex_requeue_pi_complete() publishes LOCKED before calling > > > rcuwait_wake_up(). Once the waiter observes LOCKED, it can return from > > > futex_wait_requeue_pi() and let q go out of scope before rcuwait_wake_up() > > > reads q->requeue_wait.task and passes the stale pointer to > > > try_to_wake_up(). > > > > > > Skip rcuwait_wake_up() for Q_REQUEUE_PI_LOCKED. requeue_pi_wake_futex() > > > already saves q->task before publishing LOCKED and wakes the saved task > > > afterward. > > > > > > Fixes: 07d91ef510fb1 ("futex: Prevent requeue_pi() lock nesting issue on RT") > > > Cc: stable@vger.kernel.org > > > Signed-off-by: Yao Kai > > > --- > > > kernel/futex/requeue.c | 9 +++++++-- > > > 1 file changed, 7 insertions(+), 2 deletions(-) > > > > > > diff --git a/kernel/futex/requeue.c b/kernel/futex/requeue.c > > > index abc652b5b2dd..59e587775d9b 100644 > > > --- a/kernel/futex/requeue.c > > > +++ b/kernel/futex/requeue.c > > > @@ -155,8 +155,13 @@ static inline void futex_requeue_pi_complete(struct futex_q *q, int locked) > > > } while (!atomic_try_cmpxchg(&q->requeue_state, &old, new)); > > > #ifdef CONFIG_PREEMPT_RT > > > - /* If the waiter interleaved with the requeue let it know */ > > > - if (unlikely(old == Q_REQUEUE_PI_WAIT)) > > > + /* > > > + * If the waiter interleaved with the requeue, let it know. For LOCKED, > > > + * q may already be invalid, so requeue_pi_wake_futex() wakes the saved > > > + * task instead. > > > + */ > > > + if (unlikely(old == Q_REQUEUE_PI_WAIT) && > > > + new != Q_REQUEUE_PI_LOCKED) > > > rcuwait_wake_up(&q->requeue_wait); > > > > Your whole assumption is based on the requeue_state in > > futex_requeue_pi_wakeup_sync() changes from Q_REQUEUE_PI_IN_PROGRESS to > > Q_REQUEUE_PI_WAIT and the rcuwait_wait_event() does not wait because the > > condition becomes true before that happens. So the rcuwait_wake_up() > > could access q.requeue_wait which is allocated on behalf of the waiter > > which is gone. Certainly possible. But if we skip the wait in thise case > > we probably miss the 99% cases where the waiter did wait, no? > > > > This looks similar to commit b549113738e8c ("futex: Prevent > > use-after-free during requeue-PI"). > > > > > #endif > > > } > > > > Sebastian > > No wakeup is missed. Q_REQUEUE_PI_LOCKED is only published by > requeue_pi_wake_futex(), which saves q->task before publishing the state > and calls wake_up_state(task, TASK_NORMAL) afterwards. The other > futex_requeue_pi_complete() callers produce DONE or an error state, and > the rcuwait wakeup is retained for those paths. T1 T2 futex_requeue_pi_wakeup_sync() old = Q_REQUEUE_PI_IN_PROGRESS new = Q_REQUEUE_PI_WAIT cmpxchg() if (old == Q_REQUEUE_PI_IN_PROGRESS) futex_requeue_pi_complete old = Q_REQUEUE_PI_WAIT new = Q_REQUEUE_PI_LOCKED cmpxchg() rcuwait_wait_event() rcuwait_wake_up(&q->requeue_wait); So this your case then T2 updates the state before T1 enters sleep because the condition was true before that. So you patch makes sense. However given the more common case: T1 T2 futex_requeue_pi_wakeup_sync() old = Q_REQUEUE_PI_IN_PROGRESS new = Q_REQUEUE_PI_WAIT cmpxchg() futex_requeue_pi_complete old = Q_REQUEUE_PI_WAIT new = Q_REQUEUE_PI_LOCKED cmpxchg() if (old == Q_REQUEUE_PI_IN_PROGRESS) rcuwait_wait_event() rcuwait_wake_up(&q->requeue_wait); You will miss to wake T1 if T2 skips the wake, as suggested. Or do I miss something? > Thanks, > Yao Sebastian