From: "Morten Brørup" <mb@smartsharesystems.com>
To: "Bruce Richardson" <bruce.richardson@intel.com>
Cc: <dev@dpdk.org>
Subject: RE: [PATCH] stack: introduce pile
Date: Wed, 12 Aug 2026 18:01:31 +0200 [thread overview]
Message-ID: <98CBD80474FA8B44BF855DF32C47DC35F659DB@smartserver.smartshare.dk> (raw)
In-Reply-To: <anyEWVXfFEHBbsAJ@bricha3-mobl1.ger.corp.intel.com>
> From: Bruce Richardson [mailto:bruce.richardson@intel.com]
> Sent: Wednesday, 12 August 2026 16.34
>
> On Wed, Aug 12, 2026 at 01:47:56PM +0000, Morten Brørup wrote:
> > Added a new high-performance lock-free "pile", using the Stack API.
> > The pile behaves roughly like a stack, but is not strictly LIFO.
> >
> > The pile is optimized for pushing/popping bulks of objects, which
> > it does significantly faster than the lock-free stack.
> >
> > Pushing/popping a number of objects not divisible by the compile time
> > configurable bulk size is handled gracefully, but not as fast as
> > complete bulks.
> >
> > Performance examples, stack_pile_perf_autotest vs. stack_lf_autotest:
> >
> > On a single core, pushing/popping 1 or 8 objects is similar speed.
> > On a single core, pushing/popping 32 objects is 2x faster.
> > On a single core, pushing/popping 512 objects is 10x faster.
> >
> > On four cores, pushing/popping 1, 8 or 32 objects is slightly faster.
> > On four cores, pushing/popping 512 objects is 4x faster.
> >
> How does it differ, or what are the benefits vs the standard stack.
The standard stack is generally faster on single core.
Only pushing/popping 512 objects is slightly faster than the standard stack, but maybe the standard stack could be microoptimized to fix that.
On four cores, it is slightly faster than the standard stack with multiple objects:
stack_perf_autotest
Average cycles per object push/pop (bulk size: 1): 3806.65
Average cycles per object push/pop (bulk size: 8): 554.77
Average cycles per object push/pop (bulk size: 32): 144.70
Average cycles per object push/pop (bulk size: 512): 13.52
stack_pile_perf_autotest
Average cycles per object push/pop (bulk size: 1): 4012.39
Average cycles per object push/pop (bulk size: 8): 459.82
Average cycles per object push/pop (bulk size: 32): 122.40
Average cycles per object push/pop (bulk size: 512): 9.60
> The
> LF
> stack I'd view as a strange edge case, since even with the standard
> MP/MC
> ring mempool driver you still have atomics on the fast-path and you
> will
> still get stalls if a thread gets context switched out in the middle of
> a
> mempool operation.
Yes, the no-stall guarantee is a key benefit of the lock-free stack and the pile.
next prev parent reply other threads:[~2026-08-12 16:01 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-12 13:47 [PATCH] stack: introduce pile Morten Brørup
2026-08-12 14:34 ` Bruce Richardson
2026-08-12 16:01 ` Morten Brørup [this message]
2026-08-12 16:15 ` Bruce Richardson
2026-08-12 16:28 ` Morten Brørup
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=98CBD80474FA8B44BF855DF32C47DC35F659DB@smartserver.smartshare.dk \
--to=mb@smartsharesystems.com \
--cc=bruce.richardson@intel.com \
--cc=dev@dpdk.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox