Linux Perf Users
 help / color / mirror / Atom feed
From: Changbin Du <changbin.du@gmail.com>
To: David Laight <david.laight.linux@gmail.com>
Cc: Changbin Du <changbin.du@gmail.com>,
	 Peter Zijlstra <peterz@infradead.org>,
	Ingo Molnar <mingo@redhat.com>,
	 Arnaldo Carvalho de Melo <acme@kernel.org>,
	Namhyung Kim <namhyung@kernel.org>,
	 Mark Rutland <mark.rutland@arm.com>,
	Alexander Shishkin <alexander.shishkin@linux.intel.com>,
	 Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
	 Adrian Hunter <adrian.hunter@intel.com>,
	James Clark <james.clark@linaro.org>,
	 linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org
Subject: Re: [PATCH] perf bench: Add atomic CAS benchmark
Date: Fri, 9 Oct 2026 18:04:39 +0800	[thread overview]
Message-ID: <asi72sgQ2YF3qKFm@mail.google.com> (raw)
In-Reply-To: <20261008084449.5d9971ad@pumpkin>

On Thu, Oct 08, 2026 at 08:44:49AM +0100, David Laight wrote:
> On Wed, 30 Sep 2026 17:16:17 +0800
> Changbin Du <changbin.du@gmail.com> wrote:
> 
> > Add a new 'atomic' collection to perf bench for benchmarking
> > compare-and-swap (CAS) atomic operations with multi-threaded
> > contention testing.
> > 
> > The benchmark tests __atomic_compare_exchange_n operations
> > with configurable thread count and iteration count to measure
> > atomic contention effects.
> > 
> > Why this benchmark is needed:
> > - CAS operations are fundamental to lock-free algorithms and data
> >   structures. Understanding their performance characteristics under
> >   contention is critical for designing high-performance concurrent
> >   applications.
> > - The benchmark helps identify atomic operation latency and
> >   scalability issues across different thread counts, revealing
> >   contention patterns that are not visible in single-threaded tests.
> > - Useful for evaluating atomic implementation quality on different
> >   architectures and for regression testing after changes to atomic
> >   primitives or memory ordering.
> > 
> > Measurement methodology:
> > - Each thread starts a private timer (clock_gettime CLOCK_MONOTONIC)
> >   after synchronizing on a pthread_barrier, ensuring all threads
> >   begin simultaneously.
> > - Each thread performs a hot loop of atomic compare-and-swap on a
> >   shared u64 counter, incrementing from 0 to iterations.
> > - The shared counter is cache-line aligned (64 bytes) to isolate
> >   contention to the target cache line and avoid false sharing.
> > - The wall-clock time is measured as the max of all per-thread
> >   runtimes (the time for the slowest thread to finish).
> > - The first repeat is excluded from statistics as a warmup phase
> >   to avoid cache-cold effects.
> > Example usage:
> >   $ perf bench atomic cas --threads 2
> >   # Running 'atomic/cas' benchmark:
> > 
> >     Threads: 2, iterations/thread: 100000000, repeats: 10 (warmup: 1)
> >     Avg wall-clock time: 7365.480 msec  (stddev 66.014 msec)
> >     Total ops: 200,000,000
> >     Throughput total: 27,153,697 ops/sec
> >     Per-thread times and throughput (last repeat):
> >       fastest:  7510.031 msec  (13315525 ops/sec)
> >       slowest:  7581.678 msec  (13189692 ops/sec)
> >       avg:      7545.854 msec  (13252310 ops/sec)
> > 
> > Output fields explained:
> > - Threads: number of contending threads
> > - iterations/thread: CAS operations each thread performs
> > - repeats: number of test runs (first is warmup)
> > - Avg wall-clock time: mean time for all threads to complete
> > - stddev: standard deviation across repeats
> > - Total ops: threads x iterations/thread
> > - Throughput total: aggregate ops/sec across all threads
> > - Per-thread times: fastest/slowest/avg thread completion time
> > - Per-thread throughput: per-thread ops/sec (shows scheduling imbalance)
> > 
> 
> I think you need to default to one thread per cpu.
> Also try to run the test for a fixed time period rather than a very
> large count.
> You should be able to see that some systems completely fail to make
> progress under very heavy contention.
> (This isn't one thread getting starved, none of them make progress.)
> 
> David
Thanks for the suggestions.

I'll default the thread count to one per online CPU. The benchmark
contends a single shared cache line, so the threads need to run
simultaneously for the result to reflect real cache-coherence
contention; more threads than CPUs get time-sliced, fewer do not
exercise the machine.

For the fixed-time run I'll replace the iteration count with a
runtime in seconds (like the futex benchmarks). Threads synchronize
on a pthread_barrier, spin on the atomic operation while counting
their own completed operations, and stop when the main thread sets a
shared done flag after the runtime; workers poll it every 1024
operations so the check does not perturb the hot loop. Throughput is
the total operation count over the measured wall time, and the
per-thread counts show whether any thread failed to make progress.

-- 
Cheers,
Changbin Du

      reply	other threads:[~2026-10-09 10:04 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  9:16 [PATCH] perf bench: Add atomic CAS benchmark Changbin Du
2026-09-30  9:25 ` sashiko-bot
2026-10-05 22:30 ` Namhyung Kim
2026-10-07  5:27   ` Changbin Du
2026-10-07 21:35     ` Namhyung Kim
2026-10-08  7:44 ` David Laight
2026-10-09 10:04   ` Changbin Du [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asi72sgQ2YF3qKFm@mail.google.com \
    --to=changbin.du@gmail.com \
    --cc=acme@kernel.org \
    --cc=adrian.hunter@intel.com \
    --cc=alexander.shishkin@linux.intel.com \
    --cc=david.laight.linux@gmail.com \
    --cc=irogers@google.com \
    --cc=james.clark@linaro.org \
    --cc=jolsa@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=mingo@redhat.com \
    --cc=namhyung@kernel.org \
    --cc=peterz@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox