* [RFC PATCH] ext4: cacheline-align inode cache to avoid false sharing
@ 2026-08-27 7:43 JonasZhou-oc
2026-08-27 7:51 ` sashiko-bot
0 siblings, 1 reply; 3+ messages in thread
From: JonasZhou-oc @ 2026-08-27 7:43 UTC (permalink / raw)
To: linux-ext4
Cc: tytso, adilger.kernel, libaokun, jack, ojaswin, ritesh.list,
yi.zhang, linux-fsdevel, linux-kernel, jianhuizzzzz, JonasZhou
struct ext4_inode_info embeds struct inode, but ext4_inode_cache does
not request cache-line alignment. On the tested x86-64 build,
struct ext4_inode_info is 1072 bytes, so consecutive objects can start
at different offsets within a 64-byte cache line.
The embedded inode starts at offset 232, with i_ctime_nsec, i_blkbits,
i_state, and i_rwsem at offsets 352, 366, 376, and 384, respectively.
When the containing object is not cache-line aligned, inode metadata
and i_rwsem can occupy the same cache line. Metadata updates then
invalidate the line used by CPUs contending on i_rwsem.
Add SLAB_HWCACHE_ALIGN to ext4_inode_cache. This makes the containing
objects start at cache-line boundaries and, with the tested layout,
places i_rwsem at the start of a cache line separate from the preceding
metadata.
Tests were run on Linux 7.2 on a two-socket Intel Xeon Silver 4208
system with 16 CPUs, SMT disabled, and the frequency fixed at 2 GHz.
Ten normal-mode runs of each zjhbench workload gave:
before after change
unixbench.fstime 488330.6 543996.2 +11.40%
unixbench.fsdisk 1396385.0 1501352.5 +7.52%
On this build, the slab allocation size grows from 1072 to 1088 bytes,
an increase of 16 bytes (1.49%) per ext4 inode.
RFC questions:
1. Is using SLAB_HWCACHE_ALIGN for ext4_inode_cache acceptable when
the measured benefit depends on the current struct inode layout?
2. Would it be preferable to enable the alignment only when
vfs_inode.i_rwsem is naturally cache-line aligned within
struct ext4_inode_info, rather than enabling it unconditionally?
3. Should this false sharing instead be addressed in the generic VFS
inode layout, despite the significantly larger memory and
cross-filesystem impact?
4. What additional workloads or configuration coverage would be
required before this could be considered as a non-RFC patch?
Link: https://lore.kernel.org/linux-fsdevel/CAGudoHFgtM8Px4mRNM_fsmi3=vAyCMPC3FBCzk5uE7ma7fdbdQ@mail.gmail.com/
Signed-off-by: JonasZhou <jonaszhou-oc@zhaoxin.com>
---
fs/ext4/super.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/fs/ext4/super.c b/fs/ext4/super.c
index 245f67d10ded..11090fca1613 100644
--- a/fs/ext4/super.c
+++ b/fs/ext4/super.c
@@ -1510,7 +1510,8 @@ static int __init init_inodecache(void)
ext4_inode_cachep = kmem_cache_create("ext4_inode_cache",
sizeof(struct ext4_inode_info),
&args,
- SLAB_RECLAIM_ACCOUNT | SLAB_ACCOUNT);
+ SLAB_RECLAIM_ACCOUNT | SLAB_ACCOUNT |
+ SLAB_HWCACHE_ALIGN);
if (ext4_inode_cachep == NULL)
return -ENOMEM;
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [RFC PATCH] ext4: cacheline-align inode cache to avoid false sharing
2026-08-27 7:43 [RFC PATCH] ext4: cacheline-align inode cache to avoid false sharing JonasZhou-oc
@ 2026-08-27 7:51 ` sashiko-bot
2026-09-07 5:57 ` JonasZhou-oc
0 siblings, 1 reply; 3+ messages in thread
From: sashiko-bot @ 2026-08-27 7:51 UTC (permalink / raw)
To: JonasZhou-oc; +Cc: tytso, linux-ext4
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] Fragile performance optimization relying on configuration-dependent struct layout
--
commit 7acd8c0fa63f737f049b6c32ef756e22fa1aeabb
Author: JonasZhou-oc <jonaszhou-oc@zhaoxin.com>
ext4: cacheline-align inode cache to avoid false sharing
This patch adds SLAB_HWCACHE_ALIGN to the ext4 inode cache creation flags.
The goal is to align the embedded vfs_inode.i_rwsem to a cache-line boundary
to avoid false sharing with other metadata.
Link: https://lore.kernel.org/linux-fsdevel/CAGudoHFgtM8Px4mRNM_fsmi3=vAyCMPC3FBCzk5uE7ma7fdbdQ@mail.gmail.com/
> diff --git a/fs/ext4/super.c b/fs/ext4/super.c
> index bca0dc87d0b7c..49ac97af19945 100644
> --- a/fs/ext4/super.c
> +++ b/fs/ext4/super.c
> @@ -1519,7 +1519,8 @@ static int __init init_inodecache(void)
> ext4_inode_cachep = kmem_cache_create("ext4_inode_cache",
> sizeof(struct ext4_inode_info),
> &args,
> - SLAB_RECLAIM_ACCOUNT | SLAB_ACCOUNT);
> + SLAB_RECLAIM_ACCOUNT | SLAB_ACCOUNT |
> + SLAB_HWCACHE_ALIGN);
[Severity: Medium]
Will this reliably align i_rwsem to a cache-line boundary across different
kernel configurations?
The offset of i_rwsem within struct ext4_inode_info is highly dependent on
kernel configurations. Options like CONFIG_LOCKDEP or CONFIG_DEBUG_SPINLOCK
can alter the sizes of preceding spinlock_t and struct rw_semaphore fields.
While SLAB_HWCACHE_ALIGN ensures the start of struct ext4_inode_info is
cache-line aligned, the varying internal offset of i_rwsem means it may still
arbitrarily span cache-line boundaries or share a line with heavily modified
metadata in other builds.
Does this optimization risk uniformly imposing memory overhead while failing to
provide the performance benefit on varying configurations?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260827074334.616360-1-jonaszhou-oc@zhaoxin.com?part=1
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [RFC PATCH] ext4: cacheline-align inode cache to avoid false sharing
2026-08-27 7:51 ` sashiko-bot
@ 2026-09-07 5:57 ` JonasZhou-oc
0 siblings, 0 replies; 3+ messages in thread
From: JonasZhou-oc @ 2026-09-07 5:57 UTC (permalink / raw)
To: sashiko-bot; +Cc: jonaszhou-oc, linux-ext4, sashiko-reviews, tytso
From: JonasZhou-oc <jonaszhou-oc@zhaoxin.com>
Thanks for pointing this out.
Yes, SLAB_HWCACHE_ALIGN aligns the allocation, but does not guarantee
that the embedded i_rwsem starts at a cache-line boundary in every
configuration. The offset of 384 bytes and the performance results
in the changelog apply only to the tested build. This dependency is
why I raised the layout question in the RFC.
In configurations with a different layout, the patch could increase
the allocation size without separating i_rwsem from the preceding
metadata. The size increase itself is also configuration-dependent;
the reported 16-byte increase applies to the tested build.
One possible refinement is to request SLAB_HWCACHE_ALIGN only when:
IS_ALIGNED(offsetof(struct ext4_inode_info, vfs_inode) +
offsetof(struct inode, i_rwsem),
cache_line_size())
This would restrict the change to layouts where aligning the object
also places i_rwsem at a cache-line boundary. It would separate the
lock from preceding fields, but would not guarantee a performance
improvement or eliminate possible sharing with fields following it.
I have not yet tested this conditional variant. It would still need
configuration coverage and workload validation, including checking
the actual slab memory cost.
Would this layout-dependent opt-in be preferable, or would maintainers
prefer evaluating unconditional alignment with broader performance
and memory measurements?
Thanks,
Jonas
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-07 5:57 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-27 7:43 [RFC PATCH] ext4: cacheline-align inode cache to avoid false sharing JonasZhou-oc
2026-08-27 7:51 ` sashiko-bot
2026-09-07 5:57 ` JonasZhou-oc
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox