* Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
[not found] <20260902074805.398540-1-ruanjinjie@huawei.com>
@ 2026-09-04 12:28 ` Petr Mladek
0 siblings, 0 replies; only message in thread
From: Petr Mladek @ 2026-09-04 12:28 UTC (permalink / raw)
To: Jinjie Ruan
Cc: bcrl, viro, brauner, jack, tytso, adilger.kernel, libaokun,
ojaswin, ritesh.list, yi.zhang, sforshee, akpm, rostedt,
andriy.shevchenko, linux, senozhatsky, kees, tglx, linux-fsdevel,
linux-aio, linux-kernel, linux-ext4, linux-riscv
Adding Risc-V list into Cc.
On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:
> Hi,
>
> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> smp_store_release()/smp_load_acquire() across various subsystems.
>
> Background
> ==========
>
> Many architectures support load acquire and store release instructions
> which can replace explicit memory barriers and save cycles. As noted
> in the ARM architecture reference [1]:
>
> "Weaker ordering requirements that are imposed by Load-Acquire and
> Store-Release instructions allow for micro-architectural
> optimizations, which could reduce some of the performance impacts
> that are otherwise imposed by an explicit memory barrier.
>
> If the ordering requirement is satisfied using either a Load-Acquire
> or Store-Release, then it would be preferable to use these
> instructions instead of a DMB."
>
> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
> barriers. Replacing the read barrier with smp_load_acquire() reduces
> this to 8 cycles on an Ampere Altra.
I wonder if this is true on all other architectures:
+ It seems that Arm gets the gain because the instruction
does both load/store + barrier. It helps even when
the barrier is full.
+ Some other architectures need two instructions. One for the
load/store and the other for the barrier. But the barrier
is weaker, it synchronizes just reads or just writes.
For example, I see the following in riscv/include/asm/barrier.h:
<paste riscv/include/asm/barrier.h>
#define smp_mb() RISCV_FENCE(rw, rw)
#define smp_rmb() RISCV_FENCE(r, r)
#define smp_wmb() RISCV_FENCE(w, w)
#define smp_store_release(p, v) \
do { \
RISCV_FENCE(rw, w); \
WRITE_ONCE(*p, v); \
} while (0)
#define smp_load_acquire(p) \
({ \
typeof(*p) ___p1 = READ_ONCE(*p); \
RISCV_FENCE(r, rw); \
___p1; \
})
</paste riscv/include/asm/barrier.h>
I wonder whether:
+ RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
+ RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)
so it might cause performance regression there...
Best Regards,
Petr
> We also observed significant barrier overhead while profiling Unxibench
> syscall test on arm64: a single getuid() call is ~8ns slower than on
> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
> the use of LDAR, eliminating the measurable overhead.
>
> This motivated a broader search for existing barrier pairs that can
> be converted to the lighter acquire/release semantics.
>
> Changes
> =======
>
> Each patch in this series targets a specific barrier pair where the
> publish/subscribe pattern is already present:
>
> - Writers populate data, then publish a flag/count/pointer via
> smp_store_release()
>
> - Readers load the flag/count/pointer via smp_load_acquire(), then
> consume the data
>
> This preserves the existing memory ordering guarantees while allowing
> architectures with native acquire/release instructions (e.g. arm64's
> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
> On architectures without native support, the generated code is
> generally no worse than the explicit barrier pair.
>
> The conversions are mechanical and no functional change is intended.
>
> Testing (Kunpeng HIP09 arm64 server)
> ====================================
>
> 1. UNIXBENCH syscall
> Baseline: 715.27
> Patched: 718.83
> Improvement: +0.50%
>
> 2. fs/aio (fio + null_blk, 4 jobs):
> Baseline: 1441k IOPS, 86.46us
> Patched: 1452k IOPS, 85.80us
> Improvement: ~0.8%
>
> Both improvements are consistent across runs and align with the
> expected savings from replacing DMB with LDAR/STLR on arm64.
>
> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>
> Changes in v3:
> - Add Reviewed-by.
> - Split out network patch set as Kuniyuki suggested.
> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
>
> Changes in v2:
> - Fix pre-existing issue for ext4 and 8021q [3].
> - Fix missing copy_mnt_idmap() udapte [3].
> - Drop nacked isotp patch.
> - Add test data.
> - Add Reviewed-by and update fs patch as Jan suggested.
>
> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>
> Jinjie Ruan (8):
> user_namespace: Use acquire/release for nr_extents synchronization
> lib/vsprintf: Use acquire/release for ptr_key publication
> fs: aio: Use acquire/release for ring->tail publication
> fs: Use acquire/release for fdtable resize synchronization
> pidfs: Use test_bit_acquire() for attr flag tests
> super: Use acquire for SB_BORN check in super_cache_count()
> ext4: Fix out-of-bounds read in ext4_get_group_info()
> ext4: Convert group-count barrier protocol to acquire/release
>
> fs/aio.c | 10 ++++------
> fs/ext4/balloc.c | 2 +-
> fs/ext4/ext4.h | 10 +++-------
> fs/ext4/mballoc.c | 6 ++----
> fs/ext4/resize.c | 19 +++++++++++--------
> fs/file.c | 10 ++++------
> fs/mnt_idmapping.c | 5 ++---
> fs/pidfs.c | 6 ++----
> fs/super.c | 6 ++----
> kernel/user_namespace.c | 24 +++++++++++++-----------
> lib/vsprintf.c | 11 ++++-------
> 11 files changed, 48 insertions(+), 61 deletions(-)
>
> --
> 2.34.1
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] only message in thread
only message in thread, other threads:[~2026-09-04 12:29 UTC | newest]
Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <20260902074805.398540-1-ruanjinjie@huawei.com>
2026-09-04 12:28 ` [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance Petr Mladek
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox