From: Alice Ryhl <aliceryhl@google.com>
To: Dmitry Ilvokhin <d@ilvokhin.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Peter Zijlstra <peterz@infradead.org>,
"Paul E. McKenney" <paulmck@kernel.org>,
Boqun Feng <boqun@kernel.org>,
Dmitry Vyukov <dvyukov@google.com>,
Thomas Gleixner <tglx@kernel.org>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Randy Dunlap <rdunlap@infradead.org>,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] rseq: defer time slice extension yield for sys_futex_wake
Date: Tue, 1 Sep 2026 20:27:44 +0000 [thread overview]
Message-ID: <apc1QGjH0ObFx2rw@google.com> (raw)
In-Reply-To: <apa71yN7lYR3FKeI@shell.ilvokhin.com>
On Tue, Sep 01, 2026 at 11:49:43AM +0000, Dmitry Ilvokhin wrote:
> On Mon, Aug 31, 2026 at 12:57:26PM +0000, Alice Ryhl wrote:
> > When a task is granted an rseq scheduler time slice extension, it is
> > expected to finish its critical section and relinquish the CPU via
> > rseq_slice_yield(2). If the task issues any other system call while a
> > grant is active, rseq_syscall_enter_work() forces an immediate
> > reschedule on syscall entry via cond_resched(). This may cause
> > significant latency penalty for userspace lock implementations that use
> > rseq time slice extensions when unlocking the futex.
> >
> > In a userspace mutex unlock sequence:
> > 1. The lock is released in userspace.
> > 2. If there are waiters, the unlocking thread calls sys_futex_wake()
> > to wake a sleeping waiter.
>
> Just out of curiosity, is there any publicly available implementation of
> a mutex that combines rseq time-slice extensions and futexes?
>
> I'm interested in learning more about how these two mechanisms interact
> with each other.
This came up while I was looking into implementing one for use in Tokio.
There, we have quite a few cases where we have a doubly linked list
protected by a mutex, and I really really want to avoid preemption while
that lock is held. Some of those locks are known problems for contention
in Tokio.
But that implementation is still WIP.
Some pseudocode:
struct rseq_futex {
// 0 = unlocked, 1 = locked, 2 = contended
int futex;
};
void mutex_lock(struct rseq_futex *mutex) {
for (;;) {
__rseq_abi.slice_ctrl.request = 1;
int expected = 0;
// Change state from 0 -> 1 to lock.
if (compare_exchange(&mutex->futex, &expected, 1))
return;
// lock is taken, use slow-path
bool was_granted = __rseq_abi.slice_ctrl.granted;
__rseq_abi.slice_ctrl.request = 0;
// If the state is not already 2, then change it so that
// mutex_unlock() knows to wake us up.
if (expected == 2 || compare_exchange(&mutex->futex, &expected, 2)) {
// State is 2, we can sleep until that changes.
futex_wait(&mutex->futex, 2);
} else if (was_granted) {
rseq_slice_yield();
}
}
}
void mutex_unlock(struct rseq_futex *mutex) {
// Unlock the mutex and return whether it was contended.
int prev = atomic_swap(&mutex->futex, 0);
// End the time slice extension for the critical region
bool was_granted = __rseq_abi.slice_ctrl.granted;
__rseq_abi.slice_ctrl.request = 0;
if (prev == 2) {
futex_wake(&mutex->futex);
} else if (was_granted) {
rseq_slice_yield();
}
}
Though now that I think more about it, perhaps unlock should look like
this, to avoid a preemption point just immediately before futex_wake().
void mutex_unlock(struct rseq_futex *mutex) {
// Unlock the mutex and return whether it was contended.
int prev = atomic_exchange(&mutex->futex, 0);
// Wake up contended waiters
if (prev == 2)
futex_wake(&mutex->futex);
// End the time slice extension for the critical region
bool was_granted = __rseq_abi.slice_ctrl.granted;
__rseq_abi.slice_ctrl.request = 0;
if (was_granted)
rseq_slice_yield();
}
next prev parent reply other threads:[~2026-09-01 20:27 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 12:57 [PATCH] rseq: defer time slice extension yield for sys_futex_wake Alice Ryhl
2026-08-31 13:45 ` Mathieu Desnoyers
2026-08-31 14:12 ` Alice Ryhl
2026-09-01 11:49 ` Dmitry Ilvokhin
2026-09-01 20:27 ` Alice Ryhl [this message]
2026-09-04 5:21 ` Thomas Gleixner
2026-09-04 13:53 ` [PATCH] rseq: defer time slice extension yield for sys_futex_wakey Dmitry Ilvokhin
2026-09-05 20:01 ` Thomas Gleixner
2026-09-07 13:21 ` Alice Ryhl
2026-09-07 16:09 ` Thomas Gleixner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apc1QGjH0ObFx2rw@google.com \
--to=aliceryhl@google.com \
--cc=boqun@kernel.org \
--cc=corbet@lwn.net \
--cc=d@ilvokhin.com \
--cc=dvyukov@google.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=paulmck@kernel.org \
--cc=peterz@infradead.org \
--cc=rdunlap@infradead.org \
--cc=skhan@linuxfoundation.org \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.