From: Christian Brauner <brauner@kernel.org>
To: Jan Kara <jack@suse.cz>
Cc: Pan Deng <pan.deng@intel.com>,
viro@zeniv.linux.org.uk, linux-fsdevel@vger.kernel.org,
tianyou.li@intel.com, tim.c.chen@linux.intel.com,
lipeng.zhu@intel.com
Subject: Re: [PATCH] fs: place f_ref to 3rd cache line in struct file to resolve false sharing
Date: Sat, 1 Mar 2025 11:07:12 +0100 [thread overview]
Message-ID: <20250301-ermahnen-kramen-b6e90ea5b50d@brauner> (raw)
In-Reply-To: <uyqqemnrf46xdht3mr4okv6zw7asfhjabz3fu5fl5yan52ntoh@nflmsbxz6meb>
On Fri, Feb 28, 2025 at 08:51:27PM +0100, Jan Kara wrote:
> On Fri 28-02-25 10:00:59, Pan Deng wrote:
> > When running syscall pread in a high core count system, f_ref contends
> > with the reading of f_mode, f_op, f_mapping, f_inode, f_flags in the
> > same cache line.
>
> Well, but you have to have mulithreaded process using the same struct file
> for the IO, don't you? Otherwise f_ref is not touched...
Yes, it's specifically designed to scale better under high contention.
>
> > This change places f_ref to the 3rd cache line where fields are not
> > updated as frequently as the 1st cache line, and the contention is
> > grealy reduced according to tests. In addition, the size of file
> > object is kept in 3 cache lines.
> >
> > This change has been tested with rocksdb benchmark readwhilewriting case
> > in 1 socket 64 physical core 128 logical core baremetal machine, with
> > build config CONFIG_RANDSTRUCT_NONE=y
> > Command:
> > ./db_bench --benchmarks="readwhilewriting" --threads $cnt --duration 60
> > The throughput(ops/s) is improved up to ~21%.
> > =====
> > thread baseline compare
> > 16 100% +1.3%
> > 32 100% +2.2%
> > 64 100% +7.2%
> > 128 100% +20.9%
> >
> > It was also tested with UnixBench: syscall, fsbuffer, fstime,
> > fsdisk cases that has been used for file struct layout tuning, no
> > regression was observed.
>
> So overall keeping the first cacheline read mostly with important stuff
> makes sense to limit cache traffic. But:
>
> > struct file {
> > - file_ref_t f_ref;
> > spinlock_t f_lock;
> > fmode_t f_mode;
> > const struct file_operations *f_op;
> > @@ -1102,6 +1101,7 @@ struct file {
> > unsigned int f_flags;
> > unsigned int f_iocb_flags;
> > const struct cred *f_cred;
> > + u8 padding[8];
> > /* --- cacheline 1 boundary (64 bytes) --- */
> > struct path f_path;
> > union {
> > @@ -1127,6 +1127,7 @@ struct file {
> > struct file_ra_state f_ra;
> > freeptr_t f_freeptr;
> > };
> > + file_ref_t f_ref;
> > /* --- cacheline 3 boundary (192 bytes) --- */
> > } __randomize_layout
>
> This keeps struct file within 3 cachelines but it actually grows it from
> 184 to 192 bytes (and yes, that changes how many file structs we can fit in
> a slab). So instead of adding 8 bytes of padding, just pick some
> read-mostly element and move it into the hole - f_owner looks like one
> possible candidate.
This is what I did. See vfs-6.15.misc! Thanks!
>
> Also did you test how moving f_ref to the second cache line instead of the
> third one behaves?
>
> Honza
> --
> Jan Kara <jack@suse.com>
> SUSE Labs, CR
next prev parent reply other threads:[~2025-03-01 10:07 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-02-28 2:00 [PATCH] fs: place f_ref to 3rd cache line in struct file to resolve false sharing Pan Deng
2025-02-28 10:15 ` Christian Brauner
2025-02-28 19:51 ` Jan Kara
2025-03-01 10:07 ` Christian Brauner [this message]
2025-03-04 10:36 ` Jan Kara
2025-03-04 10:50 ` Deng, Pan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20250301-ermahnen-kramen-b6e90ea5b50d@brauner \
--to=brauner@kernel.org \
--cc=jack@suse.cz \
--cc=linux-fsdevel@vger.kernel.org \
--cc=lipeng.zhu@intel.com \
--cc=pan.deng@intel.com \
--cc=tianyou.li@intel.com \
--cc=tim.c.chen@linux.intel.com \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox