All of lore.kernel.org
 help / color / mirror / Atom feed
From: Changbin Du <changbin.du@gmail.com>
To: Namhyung Kim <namhyung@kernel.org>
Cc: Changbin Du <changbin.du@gmail.com>,
	 Peter Zijlstra <peterz@infradead.org>,
	Ingo Molnar <mingo@redhat.com>,
	 Arnaldo Carvalho de Melo <acme@kernel.org>,
	Mark Rutland <mark.rutland@arm.com>,
	 Alexander Shishkin <alexander.shishkin@linux.intel.com>,
	Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
	 Adrian Hunter <adrian.hunter@intel.com>,
	James Clark <james.clark@linaro.org>,
	 linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org
Subject: Re: [PATCH] perf bench: Add atomic CAS benchmark
Date: Wed, 7 Oct 2026 13:27:36 +0800	[thread overview]
Message-ID: <asXXZJzEPoEsMSni@mail.google.com> (raw)
In-Reply-To: <asQlFvclL-vj5xlh@google.com>

Hello,
On Mon, Oct 05, 2026 at 03:30:46PM -0700, Namhyung Kim wrote:
> Hello,
> 
> On Wed, Sep 30, 2026 at 05:16:17PM +0800, Changbin Du wrote:
> > Add a new 'atomic' collection to perf bench for benchmarking
> > compare-and-swap (CAS) atomic operations with multi-threaded
> > contention testing.
> > 
> > The benchmark tests __atomic_compare_exchange_n operations
> > with configurable thread count and iteration count to measure
> > atomic contention effects.
> > 
> > Why this benchmark is needed:
> > - CAS operations are fundamental to lock-free algorithms and data
> >   structures. Understanding their performance characteristics under
> >   contention is critical for designing high-performance concurrent
> >   applications.
> > - The benchmark helps identify atomic operation latency and
> >   scalability issues across different thread counts, revealing
> >   contention patterns that are not visible in single-threaded tests.
> > - Useful for evaluating atomic implementation quality on different
> >   architectures and for regression testing after changes to atomic
> >   primitives or memory ordering.
> > 
> > Measurement methodology:
> > - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> >   after synchronizing on a pthread_barrier, ensuring all threads
> >   begin simultaneously.
> > - Each thread performs a hot loop of atomic compare-and-swap on a
> >   shared u64 counter, incrementing from 0 to iterations.
> > - The shared counter is cache-line aligned (64 bytes) to isolate
> >   contention to the target cache line and avoid false sharing.
> > - The wall-clock time is measured as the max of all per-thread
> >   runtimes (the time for the slowest thread to finish).
> > - The first repeat is excluded from statistics as a warmup phase
> >   to avoid cache-cold effects.
> > Example usage:
> >   $ perf bench atomic cas --threads 2
> >   # Running 'atomic/cas' benchmark:
> > 
> >     Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> >     Avg wall-clock time: 7365.480 msec  (stddev 66.014 msec)
> >     Total ops: 200,000,000
> >     Throughput total: 27,153,697 ops/sec
> >     Per-thread times and throughput (last repeat):
> >       fastest:  7510.031 msec  (13315525 ops/sec)
> >       slowest:  7581.678 msec  (13189692 ops/sec)
> >       avg:      7545.854 msec  (13252310 ops/sec)
> > 
> > Output fields explained:
> > - Threads: number of contending threads
> > - iterations/thread: CAS operations each thread performs
> > - repeats: number of test runs (first is warmup)
> > - Avg wall-clock time: mean time for all threads to complete
> > - stddev: standard deviation across repeats
> > - Total ops: threads x iterations/thread
> > - Throughput total: aggregate ops/sec across all threads
> > - Per-thread times: fastest/slowest/avg thread completion time
> > - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
> 
> Thanks for the contribution!  I think it's very useful.
> Just a few suggestions.
> 
> 1. it'd be nice to add simple atomic_inc benchmark too.
> 2. it'd be nice to have an option to try other ordering requirements
>    than "relaxed".
> 
> Thanks,
> Namhyung

Thanks for the review!
1. Done in v2. The collection now provides atomic inc alongside
cas, measuring contended __atomic_fetch_add() throughput with the
same skeleton (barrier-synchronized start, wall-clock taken from the
slowest thread, warmup repeat). Note that its ops/sec is not
instruction-level comparable with cas — cas counts successful
compare-and-swaps only, not the retries and loads in between — which
the documentation now points out.

2. Adding memory-order variants is not necessary because ordering
has no semantic role in this benchmark. Memory ordering exists to
constrain the visibility of other memory operations relative to an
atomic access; its cost and effect only become meaningful in an
algorithm with additional accesses to order — for example, ordering
the initialization stores of a new node before publishing a pointer
with a release CAS. This benchmark, however, measures a single shared
counter with no other memory operations in the loop, so a stronger
ordering would not order anything of consequence: it would only
measure the marginal cost of the stronger instruction itself. Those
numbers would say nothing about how orderings behave in real
workloads, and would therefore add a configuration knob that invites
misleading comparisons rather than useful measurement.

  reply	other threads:[~2026-10-07  5:28 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  9:16 [PATCH] perf bench: Add atomic CAS benchmark Changbin Du
2026-09-30  9:25 ` sashiko-bot
2026-10-05 22:30 ` Namhyung Kim
2026-10-07  5:27   ` Changbin Du [this message]
2026-10-07 21:35     ` Namhyung Kim
2026-10-08  7:44 ` David Laight
2026-10-09 10:04   ` Changbin Du

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asXXZJzEPoEsMSni@mail.google.com \
    --to=changbin.du@gmail.com \
    --cc=acme@kernel.org \
    --cc=adrian.hunter@intel.com \
    --cc=alexander.shishkin@linux.intel.com \
    --cc=irogers@google.com \
    --cc=james.clark@linaro.org \
    --cc=jolsa@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=mingo@redhat.com \
    --cc=namhyung@kernel.org \
    --cc=peterz@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.