Linux EXT4 FS development
 help / color / mirror / Atom feed
From: Petr Mladek <pmladek@suse.com>
To: Jinjie Ruan <ruanjinjie@huawei.com>
Cc: bcrl@kvack.org, viro@zeniv.linux.org.uk, brauner@kernel.org,
	jack@suse.cz, tytso@mit.edu, adilger.kernel@dilger.ca,
	libaokun@linux.alibaba.com, ojaswin@linux.ibm.com,
	ritesh.list@gmail.com, yi.zhang@huawei.com, sforshee@kernel.org,
	akpm@linux-foundation.org, rostedt@goodmis.org,
	andriy.shevchenko@linux.intel.com, linux@rasmusvillemoes.dk,
	senozhatsky@chromium.org, kees@kernel.org, tglx@kernel.org,
	linux-fsdevel@vger.kernel.org, linux-aio@kvack.org,
	linux-kernel@vger.kernel.org, linux-ext4@vger.kernel.org,
	linux-riscv@lists.infradead.org
Subject: Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Date: Fri, 4 Sep 2026 14:28:56 +0200	[thread overview]
Message-ID: <apq5iLOF8lAQ_ZVU@pathway.suse.cz> (raw)
In-Reply-To: <20260902074805.398540-1-ruanjinjie@huawei.com>

Adding Risc-V list into Cc.

On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:
> Hi,
> 
> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> smp_store_release()/smp_load_acquire() across various subsystems.
> 
> Background
> ==========
> 
> Many architectures support load acquire and store release instructions
> which can replace explicit memory barriers and save cycles. As noted
> in the ARM architecture reference [1]:
> 
>   "Weaker ordering requirements that are imposed by Load-Acquire and
>    Store-Release instructions allow for micro-architectural
>    optimizations, which could reduce some of the performance impacts
>    that are otherwise imposed by an explicit memory barrier.
> 
>    If the ordering requirement is satisfied using either a Load-Acquire
>    or Store-Release, then it would be preferable to use these
>    instructions instead of a DMB."
> 
> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
> barriers. Replacing the read barrier with smp_load_acquire() reduces
> this to 8 cycles on an Ampere Altra.

I wonder if this is true on all other architectures:

  + It seems that Arm gets the gain because the instruction
    does both load/store + barrier. It helps even when
    the barrier is full.

  + Some other architectures need two instructions. One for the
    load/store and the other for the barrier. But the barrier
    is weaker, it synchronizes just reads or just writes.

For example, I see the following in riscv/include/asm/barrier.h:

<paste riscv/include/asm/barrier.h>
#define smp_mb()	RISCV_FENCE(rw, rw)
#define smp_rmb()	RISCV_FENCE(r, r)
#define smp_wmb()	RISCV_FENCE(w, w)

#define smp_store_release(p, v)						\
do {									\
	RISCV_FENCE(rw, w);						\
	WRITE_ONCE(*p, v);						\
} while (0)

#define smp_load_acquire(p)						\
({									\
	typeof(*p) ___p1 = READ_ONCE(*p);				\
	RISCV_FENCE(r, rw);						\
	___p1;								\
})
</paste riscv/include/asm/barrier.h>

I wonder whether:

  + RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
  + RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)

so it might cause performance regression there...

Best Regards,
Petr

> We also observed significant barrier overhead while profiling Unxibench
> syscall test on arm64: a single getuid() call is ~8ns slower than on
> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
> the use of LDAR, eliminating the measurable overhead.
> 
> This motivated a broader search for existing barrier pairs that can
> be converted to the lighter acquire/release semantics.
> 
> Changes
> =======
> 
> Each patch in this series targets a specific barrier pair where the
> publish/subscribe pattern is already present:
> 
> - Writers populate data, then publish a flag/count/pointer via
>   smp_store_release()
> 
> - Readers load the flag/count/pointer via smp_load_acquire(), then
>   consume the data
> 
> This preserves the existing memory ordering guarantees while allowing
> architectures with native acquire/release instructions (e.g. arm64's
> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
> On architectures without native support, the generated code is
> generally no worse than the explicit barrier pair.
> 
> The conversions are mechanical and no functional change is intended.
> 
> Testing (Kunpeng HIP09 arm64 server)
> ====================================
> 
> 1. UNIXBENCH syscall
> 	Baseline: 715.27
> 	Patched:  718.83
> 	Improvement: +0.50%
> 
> 2. fs/aio (fio + null_blk, 4 jobs):
> 	Baseline: 1441k IOPS, 86.46us
> 	Patched:  1452k IOPS, 85.80us
> 	Improvement: ~0.8%
> 
> Both improvements are consistent across runs and align with the
> expected savings from replacing DMB with LDAR/STLR on arm64.
> 
> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
> 
> Changes in v3:
> - Add Reviewed-by.
> - Split out network patch set as Kuniyuki suggested.
> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
> 
> Changes in v2:
> - Fix pre-existing issue for ext4 and 8021q [3].
> - Fix missing copy_mnt_idmap() udapte [3].
> - Drop nacked isotp patch.
> - Add test data.
> - Add Reviewed-by and update fs patch as Jan suggested.
> 
> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
> 
> Jinjie Ruan (8):
>   user_namespace: Use acquire/release for nr_extents synchronization
>   lib/vsprintf: Use acquire/release for ptr_key publication
>   fs: aio: Use acquire/release for ring->tail publication
>   fs: Use acquire/release for fdtable resize synchronization
>   pidfs: Use test_bit_acquire() for attr flag tests
>   super: Use acquire for SB_BORN check in super_cache_count()
>   ext4: Fix out-of-bounds read in ext4_get_group_info()
>   ext4: Convert group-count barrier protocol to acquire/release
> 
>  fs/aio.c                | 10 ++++------
>  fs/ext4/balloc.c        |  2 +-
>  fs/ext4/ext4.h          | 10 +++-------
>  fs/ext4/mballoc.c       |  6 ++----
>  fs/ext4/resize.c        | 19 +++++++++++--------
>  fs/file.c               | 10 ++++------
>  fs/mnt_idmapping.c      |  5 ++---
>  fs/pidfs.c              |  6 ++----
>  fs/super.c              |  6 ++----
>  kernel/user_namespace.c | 24 +++++++++++++-----------
>  lib/vsprintf.c          | 11 ++++-------
>  11 files changed, 48 insertions(+), 61 deletions(-)
> 
> -- 
> 2.34.1

      parent reply	other threads:[~2026-09-04 12:29 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02  7:47 [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance Jinjie Ruan
2026-09-02  7:47 ` [PATCH v3 1/8] user_namespace: Use acquire/release for nr_extents synchronization Jinjie Ruan
2026-09-02  7:53   ` sashiko-bot
2026-09-02 10:09   ` Bradley Morgan
2026-09-02 13:03   ` Andy Shevchenko
2026-09-02 13:04     ` Andy Shevchenko
2026-09-03  1:42       ` Jinjie Ruan
2026-09-03  2:16     ` Jinjie Ruan
2026-09-02  7:47 ` [PATCH v3 2/8] lib/vsprintf: Use acquire/release for ptr_key publication Jinjie Ruan
2026-09-02  7:52   ` sashiko-bot
2026-09-04 13:13   ` Petr Mladek
2026-09-02  7:48 ` [PATCH v3 3/8] fs: aio: Use acquire/release for ring->tail publication Jinjie Ruan
2026-09-02  7:57   ` sashiko-bot
2026-09-02  7:48 ` [PATCH v3 4/8] fs: Use acquire/release for fdtable resize synchronization Jinjie Ruan
2026-09-02  7:53   ` sashiko-bot
2026-09-02  7:48 ` [PATCH v3 5/8] pidfs: Use test_bit_acquire() for attr flag tests Jinjie Ruan
2026-09-02  7:55   ` sashiko-bot
2026-09-02  7:48 ` [PATCH v3 6/8] super: Use acquire for SB_BORN check in super_cache_count() Jinjie Ruan
2026-09-02  7:58   ` sashiko-bot
2026-09-02  7:48 ` [PATCH v3 7/8] ext4: Fix out-of-bounds read in ext4_get_group_info() Jinjie Ruan
2026-09-02  8:01   ` sashiko-bot
2026-09-02  7:48 ` [PATCH v3 8/8] ext4: Convert group-count barrier protocol to acquire/release Jinjie Ruan
2026-09-02  7:58   ` sashiko-bot
2026-09-02 14:19 ` [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance Theodore Tso
2026-09-02 14:56   ` Jan Kara
2026-09-04  7:20   ` Jinjie Ruan
2026-09-04 12:28 ` Petr Mladek [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=apq5iLOF8lAQ_ZVU@pathway.suse.cz \
    --to=pmladek@suse.com \
    --cc=adilger.kernel@dilger.ca \
    --cc=akpm@linux-foundation.org \
    --cc=andriy.shevchenko@linux.intel.com \
    --cc=bcrl@kvack.org \
    --cc=brauner@kernel.org \
    --cc=jack@suse.cz \
    --cc=kees@kernel.org \
    --cc=libaokun@linux.alibaba.com \
    --cc=linux-aio@kvack.org \
    --cc=linux-ext4@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=linux@rasmusvillemoes.dk \
    --cc=ojaswin@linux.ibm.com \
    --cc=ritesh.list@gmail.com \
    --cc=rostedt@goodmis.org \
    --cc=ruanjinjie@huawei.com \
    --cc=senozhatsky@chromium.org \
    --cc=sforshee@kernel.org \
    --cc=tglx@kernel.org \
    --cc=tytso@mit.edu \
    --cc=viro@zeniv.linux.org.uk \
    --cc=yi.zhang@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox