DPDK-dev Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Morten Brørup" <mb@smartsharesystems.com>
To: "Konstantin Ananyev" <konstantin.ananyev@huawei.com>,
	"Stephen Hemminger" <stephen@networkplumber.org>
Cc: <dev@dpdk.org>, "Bruce Richardson" <bruce.richardson@intel.com>
Subject: RE: [PATCH] eal/x86: optimize memcpy of small 64-byte blocks
Date: Wed, 5 Aug 2026 22:28:05 +0200	[thread overview]
Message-ID: <98CBD80474FA8B44BF855DF32C47DC35F659C3@smartserver.smartshare.dk> (raw)
In-Reply-To: <3a96ca58a1e146758f5314202a04c3de@huawei.com>

> From: Konstantin Ananyev [mailto:konstantin.ananyev@huawei.com]
> Sent: Wednesday, 5 August 2026 19.16
> 
> > > > > > > > +	/* Common way for small copy size of 64-byte blocks.
> > > > > Unlikely, so
> > > > > > > constant size only */
> > > > > > > > +	if (__rte_constant(n) && (n & 63) == 0 && n <=
> > > > > > > RTE_MEMCPY_BLOCK_64_MAX) {
> > > > > > > > +		void *ret = dst;
> > > > > > > > +
> > > > > > >
> > > > > > > Maybe just let compiler decide, it will generate vector
> > > > > instructions in
> > > > > > > most cases.
> > > > > > >
> > > > > > > 	if (__rte_constant(n))
> > > > > > > 		return mempcpy(dst, src, n);
> > > > > >
> > > > > > Maybe in most, but not in all:
> > > > > > https://godbolt.org/z/KvdKqT5rY
> > > > >
> > > > > With '-mavx' or '-mavx512f' it looks like it does for your
> sample
> > > code.
> > > >
> > > > It also does with -msse4.2 when SZ is reduced to 256 bytes.
> > > > Clang switches to inline when SZ is reduced to 128 bytes.
> > > >
> > > > It seems the compiler has a threshold for when to inline and when
> to
> > > call the C
> > > > library's memcpy subroutine.
> > > > The threshold depends on both copy size and vector register size.
> > > > And it is compiler dependent.
> > >
> > > I think there are compiler options to specify desired threshold
> values.
> > > Let say for gcc there is  ' -mmemcpy-strategy=strategy'.
> > > For that example in that particular case
> > > -mmemcpy-strategy=vector_loop:512:align,loop:-1:align
> > > generates sse loads/stores.
> > > Might be we can exploit it somehow?
> >
> > That could give us higher granularity/control over memcpy for
> individual
> > memcpy instances; might be useful for hot code paths where we have
> more
> > knowledge about the copy operation than the compiler can infer.
> > However, pragmas are discouraged in DPDK, and this looks like a very
> similar
> > path.
> 
> Well, right now  rte_memcpy.h is 700+ lines and keeps growing.
> Considering that probably pragmas are not that bad.
> Of course, pragmas have their own issues and it is hard to ensure that
> they will produce same code between different compilers/versions, etc.
> 
> > > I am not really happy that our home-brewed memcpy code-block keeps
> > > growing,
> > > while we keep talking that it would be good to eliminate it
> completely.
> >
> > I agree in principle.
> > However, this rte_memcpy() optimization is for the pile/mempool
> optimizations
> > I'm working on, so there is a specific use case motivating the added
> code.
> 
> I understand that you probably have some specific use-case in mind.
> BTW for this optimization you mentioned above: what is the gain with
> these changes?

IMO, the primary benefit is the much simpler (and smaller) assembly output due to avoiding the address alignment check (and the resulting duplicated code).

I haven't measured the performance gain.
Based on the perf gain in a previous mempool optimization patch [1], it seems avoiding the address alignment check shaves ~2 cycles off the copy operation (for cache-to-cache copy).
I expect that the same gain (from avoiding the address alignment check) applies here.

[1]: https://patchwork.dpdk.org/project/dpdk/patch/20260521185631.116046-1-mb@smartsharesystems.com/


  reply	other threads:[~2026-08-05 20:28 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04 14:33 [PATCH] eal/x86: optimize memcpy of small 64-byte blocks Morten Brørup
2026-08-04 14:33 ` [RFC PATCH] pile stack and mempool driver (resend) Morten Brørup
2026-08-04 14:38   ` Morten Brørup
2026-08-04 15:52 ` [PATCH] eal/x86: optimize memcpy of small 64-byte blocks Stephen Hemminger
2026-08-04 16:25   ` Morten Brørup
2026-08-05  5:45     ` Konstantin Ananyev
2026-08-05  5:56       ` Morten Brørup
2026-08-05  6:50         ` Konstantin Ananyev
2026-08-05  8:36           ` Morten Brørup
2026-08-05 17:16             ` Konstantin Ananyev
2026-08-05 20:28               ` Morten Brørup [this message]
2026-08-06  7:58                 ` Konstantin Ananyev
2026-08-06  9:17                   ` Morten Brørup
2026-08-04 19:11 ` Morten Brørup
2026-08-04 20:42   ` Stephen Hemminger
2026-08-05 17:10 ` Konstantin Ananyev
2026-08-05 20:16   ` Morten Brørup
2026-08-06  7:43     ` Konstantin Ananyev
2026-08-06  8:00       ` Morten Brørup
2026-08-06 10:17 ` [PATCH v2] " Morten Brørup

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=98CBD80474FA8B44BF855DF32C47DC35F659C3@smartserver.smartshare.dk \
    --to=mb@smartsharesystems.com \
    --cc=bruce.richardson@intel.com \
    --cc=dev@dpdk.org \
    --cc=konstantin.ananyev@huawei.com \
    --cc=stephen@networkplumber.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox