From: Arnaldo Carvalho de Melo <acme@kernel.org>
To: Namhyung Kim <namhyung@kernel.org>
Cc: Ingo Molnar <mingo@kernel.org>,
Thomas Gleixner <tglx@linutronix.de>,
James Clark <james.clark@linaro.org>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
Clark Williams <williams@redhat.com>,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
Arnaldo Carvalho de Melo <acme@redhat.com>
Subject: Re: [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing
Date: Thu, 1 Oct 2026 11:18:54 +0200 [thread overview]
Message-ID: <ar4lfjssCX9H2CCL@x2> (raw)
In-Reply-To: <ar4LjQ8ryDjWDQ9y@z2>
On Thu, Oct 01, 2026 at 12:28:13AM -0700, Namhyung Kim wrote:
> On Wed, Sep 30, 2026 at 11:37:16PM +0200, Arnaldo Carvalho de Melo wrote:
> > Add a 'perf test -w false_sharing' workload that hammers one shared
> > struct from several CPUs, shaped as a TCP connection: a read-mostly
> > identity (five-tuple) shares a cacheline with per-packet rx counters
> > (the false-sharing line), a second line has packet-path private tx and
> > congestion control counters, and a third the connection config.
> > The packet path runs in the main thread and up to four lookup threads
> > sum the five-tuple and pull the config, reading one volatile shared
> > instance directly so the accesses are PC-relative and resolvable by the
> > data type profiler.
>
> It'd be great if you can share an output of data type profiling with
> cacheline info. Probably like below?
>
> $ perf mem record -- perf test -w false_sharing
>
> $ perf report -s type,typecln -H --group --stdio
Sure, I should've added it there as I usually do :-\
Here it is, will add to the cset message as well:
root@x2:~# perf mem record -- perf test -r5 -w false_sharing
[ perf record: Woken up 19 times to write data ]
[ perf record: Captured and wrote 5.612 MB perf.data (72454 samples) ]
root@x2:~#
root@x2:~# perf report -s type,typecln -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline
# ................... ...............................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
32.55% 0.00% struct net_conn: cache-line 2
0.01% 18.38% struct net_conn: cache-line 1
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.03% Elf64_Addr
0.01% 0.03% Elf64_Addr: cache-line 0
0.01% 0.03% struct sched_entity
0.01% 0.03% struct sched_entity: cache-line 1
0.00% 0.00% struct sched_entity: cache-line 4
0.00% 0.00% struct sched_entity: cache-line 2
0.00% 0.00% struct sched_entity: cache-line 3
0.00% 0.00% struct sched_entity: cache-line 0
0.01% 0.00% struct task_group
0.01% 0.00% struct task_group: cache-line 4
0.00% 0.00% struct task_group: cache-line 5
0.00% 0.00% struct css_rstat_cpu
0.00% 0.00% struct css_rstat_cpu: cache-line 0
root@x2:~#
root@x2:~# perf report -s type,typecln,typeoff -H --group --stdio
# Total Lost Samples: 0
#
# Samples: 72K of events 'cpu/mem-loads,ldlat=30/P, cpu/mem-stores/P'
# Event count (approx.): 1261814616
#
# Overhead Data Type / Data Type Cacheline / Data Type Offset
# ...................... ..................................................
#
88.99% 35.88% struct net_conn
56.43% 17.50% struct net_conn: cache-line 0
9.81% 0.00% struct net_conn +0xa (dport)
9.54% 0.00% struct net_conn +0x8 (sport)
9.53% 0.00% struct net_conn +0xc (state)
9.48% 0.00% struct net_conn +0xd (protocol)
9.27% 0.00% struct net_conn +0x4 (daddr)
8.79% 0.00% struct net_conn +0 (saddr)
0.00% 12.08% struct net_conn +0x10 (bytes_rx)
0.00% 1.71% struct net_conn +0x18 (packets_rx)
0.00% 2.80% struct net_conn +0x20 (rx_queue)
0.00% 0.91% struct net_conn +0x3c (last_ack)
32.55% 0.00% struct net_conn: cache-line 2
5.70% 0.00% struct net_conn +0x83 (rcv_wscale)
5.46% 0.00% struct net_conn +0x84 (keepalive_int)
5.43% 0.00% struct net_conn +0x82 (snd_wscale)
5.39% 0.00% struct net_conn +0x80 (mss)
5.34% 0.00% struct net_conn +0x88 (mark)
5.23% 0.00% struct net_conn +0x8c (priority)
0.01% 18.38% struct net_conn: cache-line 1
0.01% 0.05% struct net_conn +0x5c (retrans)
0.00% 4.83% struct net_conn +0x40 (bytes_tx)
0.00% 10.63% struct net_conn +0x48 (packets_tx)
0.00% 1.54% struct net_conn +0x58 (rtt_us)
0.00% 1.07% struct net_conn +0x50 (cwnd)
0.00% 0.25% struct net_conn +0x54 (ssthresh)
10.08% 63.03% (unknown)
10.08% 63.03% (unknown): cache-line 0
10.08% 63.03% (unknown)
0.82% 0.13% int
0.82% 0.13% int: cache-line 0
0.82% 0.13% int +0 (no field)
0.01% 0.02% struct folio
0.01% 0.02% struct folio: cache-line 0
0.01% 0.00% struct folio +0 (flags.f)
0.00% 0.01% struct folio +0x34 (_refcount.counter)
0.00% 0.00% struct folio +0x18 (mapping)
:
next prev parent reply other threads:[~2026-10-01 9:18 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 21:37 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 1/5] perf config: Move perf_config__set_variable() to util/config.c Arnaldo Carvalho de Melo
2026-09-30 21:48 ` sashiko-bot
2026-10-01 7:18 ` Namhyung Kim
2026-10-01 9:19 ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 2/5] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-30 21:52 ` sashiko-bot
2026-09-30 21:37 ` [PATCH v6 3/5] perf report: Add --no-progress option Arnaldo Carvalho de Melo
2026-09-30 21:58 ` sashiko-bot
2026-10-01 7:01 ` Namhyung Kim
2026-10-01 9:20 ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 4/5] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-09-30 22:01 ` sashiko-bot
2026-10-01 7:24 ` Namhyung Kim
2026-10-01 9:19 ` Arnaldo Carvalho de Melo
2026-09-30 21:37 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-09-30 22:05 ` sashiko-bot
2026-10-01 7:28 ` Namhyung Kim
2026-10-01 9:18 ` Arnaldo Carvalho de Melo [this message]
-- strict thread matches above, loose matches on Subject: below --
2026-09-30 11:24 [PATCH v6 0/5] perf tools: Add progress diagnostics and a false-sharing workload Arnaldo Carvalho de Melo
2026-09-30 11:24 ` [PATCH v6 5/5] perf test: Add false_sharing workload exhibiting cross-CPU false sharing Arnaldo Carvalho de Melo
2026-09-30 11:32 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ar4lfjssCX9H2CCL@x2 \
--to=acme@kernel.org \
--cc=acme@redhat.com \
--cc=adrian.hunter@intel.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=namhyung@kernel.org \
--cc=tglx@linutronix.de \
--cc=williams@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.