From: Jinjie Ruan <ruanjinjie@huawei.com>
To: <viro@zeniv.linux.org.uk>, <brauner@kernel.org>, <jack@suse.cz>,
<bcrl@kvack.org>, <tytso@mit.edu>, <adilger.kernel@dilger.ca>,
<libaokun@linux.alibaba.com>, <ojaswin@linux.ibm.com>,
<ritesh.list@gmail.com>, <yi.zhang@huawei.com>,
<pmladek@suse.com>, <rostedt@goodmis.org>,
<andriy.shevchenko@linux.intel.com>, <linux@rasmusvillemoes.dk>,
<senozhatsky@chromium.org>, <akpm@linux-foundation.org>,
<davem@davemloft.net>, <edumazet@google.com>, <kuba@kernel.org>,
<pabeni@redhat.com>, <horms@kernel.org>, <socketcan@hartkopp.net>,
<mkl@pengutronix.de>, <kuniyu@google.com>, <willemb@google.com>,
<jhs@mojatatu.com>, <jiri@resnulli.us>, <kees@kernel.org>,
<cyphar@cyphar.com>, <tglx@kernel.org>, <liuhangbin@gmail.com>,
<sdf@fomichev.me>, <nb@tipi-net.de>,
<linux-fsdevel@vger.kernel.org>, <linux-aio@kvack.org>,
<linux-kernel@vger.kernel.org>, <linux-ext4@vger.kernel.org>,
<netdev@vger.kernel.org>, <linux-can@vger.kernel.org>
Cc: <ruanjinjie@huawei.com>
Subject: [PATCH 00/11] Convert barrier pairs to acquire/release for better performance
Date: Tue, 25 Aug 2026 17:54:11 +0800 [thread overview]
Message-ID: <20260825095422.3166067-1-ruanjinjie@huawei.com> (raw)
Hi,
This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
smp_store_release()/smp_load_acquire() across various subsystems.
Background
==========
Many architectures support load acquire and store release instructions
which can replace explicit memory barriers and save cycles. As noted
in the ARM architecture reference [1]:
"Weaker ordering requirements that are imposed by Load-Acquire and
Store-Release instructions allow for micro-architectural
optimizations, which could reduce some of the performance impacts
that are otherwise imposed by an explicit memory barrier.
If the ordering requirement is satisfied using either a Load-Acquire
or Store-Release, then it would be preferable to use these
instructions instead of a DMB."
On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
barriers. Replacing the read barrier with smp_load_acquire() reduces
this to 8 cycles on an Ampere Altra.
We also observed significant barrier overhead while profiling Unxibench
syscall test on arm64: a single getuid() call is ~8ns slower than on
a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
the use of LDAR, eliminating the measurable overhead.
This motivated a broader search for existing barrier pairs that can
be converted to the lighter acquire/release semantics.
Changes
=======
Each patch in this series targets a specific barrier pair where the
publish/subscribe pattern is already present:
- Writers populate data, then publish a flag/count/pointer via
smp_store_release()
- Readers load the flag/count/pointer via smp_load_acquire(), then
consume the data
This preserves the existing memory ordering guarantees while allowing
architectures with native acquire/release instructions (e.g. arm64's
STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
On architectures without native support, the generated code is
generally no worse than the explicit barrier pair.
The conversions are mechanical and no functional change is intended.
[1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
[2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
Jinjie Ruan (11):
user_namespace: Use acquire/release for nr_extents synchronization
lib/vsprintf: Use acquire/release for ptr_key publication
fs: aio: Use acquire/release for ring->tail publication
fs: Use acquire/release for fdtable resize synchronization
pidfs: Use test_bit_acquire() for attr flag tests
super: Use acquire for SB_BORN check in super_cache_count()
ext4: Convert group-count barrier protocol to acquire/release
soreuseport: publish num_socks with acquire/release
net: sched: act_gact: use acquire/release for tcfg_ptype
8021q: publish vlan_devices_arrays entries with acquire/release
can: isotp: publish tx.state with smp_store_release()
fs/aio.c | 12 +++++-------
fs/ext4/ext4.h | 10 +++-------
fs/ext4/mballoc.c | 6 ++----
fs/ext4/resize.c | 19 +++++++++++--------
fs/file.c | 10 ++++------
fs/pidfs.c | 6 ++----
fs/super.c | 7 +++----
kernel/user_namespace.c | 24 +++++++++++++-----------
lib/vsprintf.c | 11 ++++-------
net/8021q/vlan.c | 6 ++----
net/8021q/vlan.h | 8 +++-----
net/can/isotp.c | 4 ++--
net/core/sock_reuseport.c | 20 ++++++++------------
net/sched/act_gact.c | 12 ++++--------
14 files changed, 66 insertions(+), 89 deletions(-)
--
2.34.1
next reply other threads:[~2026-08-25 9:53 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-25 9:54 Jinjie Ruan [this message]
2026-08-25 9:54 ` [PATCH 01/11] user_namespace: Use acquire/release for nr_extents synchronization Jinjie Ruan
2026-08-25 9:54 ` [PATCH 02/11] lib/vsprintf: Use acquire/release for ptr_key publication Jinjie Ruan
2026-08-25 9:54 ` [PATCH 03/11] fs: aio: Use acquire/release for ring->tail publication Jinjie Ruan
2026-08-25 9:54 ` [PATCH 04/11] fs: Use acquire/release for fdtable resize synchronization Jinjie Ruan
2026-08-25 9:54 ` [PATCH 05/11] pidfs: Use test_bit_acquire() for attr flag tests Jinjie Ruan
2026-08-25 9:54 ` [PATCH 06/11] super: Use acquire for SB_BORN check in super_cache_count() Jinjie Ruan
2026-08-25 9:54 ` [PATCH 07/11] ext4: Convert group-count barrier protocol to acquire/release Jinjie Ruan
2026-08-25 9:54 ` [PATCH 08/11] soreuseport: publish num_socks with acquire/release Jinjie Ruan
2026-08-25 9:54 ` [PATCH 09/11] net: sched: act_gact: use acquire/release for tcfg_ptype Jinjie Ruan
2026-08-25 9:54 ` [PATCH 10/11] 8021q: publish vlan_devices_arrays entries with acquire/release Jinjie Ruan
2026-08-25 9:54 ` [PATCH 11/11] can: isotp: publish tx.state with smp_store_release() Jinjie Ruan
2026-08-25 11:46 ` Oliver Hartkopp
2026-08-26 3:33 ` Jinjie Ruan
2026-08-26 8:04 ` Oliver Hartkopp
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825095422.3166067-1-ruanjinjie@huawei.com \
--to=ruanjinjie@huawei.com \
--cc=adilger.kernel@dilger.ca \
--cc=akpm@linux-foundation.org \
--cc=andriy.shevchenko@linux.intel.com \
--cc=bcrl@kvack.org \
--cc=brauner@kernel.org \
--cc=cyphar@cyphar.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=jack@suse.cz \
--cc=jhs@mojatatu.com \
--cc=jiri@resnulli.us \
--cc=kees@kernel.org \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=libaokun@linux.alibaba.com \
--cc=linux-aio@kvack.org \
--cc=linux-can@vger.kernel.org \
--cc=linux-ext4@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux@rasmusvillemoes.dk \
--cc=liuhangbin@gmail.com \
--cc=mkl@pengutronix.de \
--cc=nb@tipi-net.de \
--cc=netdev@vger.kernel.org \
--cc=ojaswin@linux.ibm.com \
--cc=pabeni@redhat.com \
--cc=pmladek@suse.com \
--cc=ritesh.list@gmail.com \
--cc=rostedt@goodmis.org \
--cc=sdf@fomichev.me \
--cc=senozhatsky@chromium.org \
--cc=socketcan@hartkopp.net \
--cc=tglx@kernel.org \
--cc=tytso@mit.edu \
--cc=viro@zeniv.linux.org.uk \
--cc=willemb@google.com \
--cc=yi.zhang@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).