From: "Morten Brørup" <mb@smartsharesystems.com>
To: "Konstantin Ananyev" <konstantin.ananyev@huawei.com>,
"Stephen Hemminger" <stephen@networkplumber.org>
Cc: <dev@dpdk.org>, "Bruce Richardson" <bruce.richardson@intel.com>
Subject: RE: [PATCH] eal/x86: optimize memcpy of small 64-byte blocks
Date: Tue, 11 Aug 2026 13:10:35 +0200 [thread overview]
Message-ID: <98CBD80474FA8B44BF855DF32C47DC35F659D3@smartserver.smartshare.dk> (raw)
In-Reply-To: <22720aafe5cf445787940bb319f7e90e@huawei.com>
> From: Konstantin Ananyev [mailto:konstantin.ananyev@huawei.com]
> Sent: Tuesday, 11 August 2026 09.06
>
>
> > > > > > > > > > > > + /* Common way for small copy size of 64-byte
> > > blocks.
> > > > > > > > > Unlikely, so
> > > > > > > > > > > constant size only */
> > > > > > > > > > > > + if (__rte_constant(n) && (n & 63) == 0 && n <=
> > > > > > > > > > > RTE_MEMCPY_BLOCK_64_MAX) {
> > > > > > > > > > > > + void *ret = dst;
> > > > > > > > > > > > +
> > > > > > > > > > >
> > > > > > > > > > > Maybe just let compiler decide, it will generate
> vector
> > > > > > > > > instructions in
> > > > > > > > > > > most cases.
> > > > > > > > > > >
> > > > > > > > > > > if (__rte_constant(n))
> > > > > > > > > > > return mempcpy(dst, src, n);
> > > > > > > > > >
> > > > > > > > > > Maybe in most, but not in all:
> > > > > > > > > > https://godbolt.org/z/KvdKqT5rY
> > > > > > > > >
> > > > > > > > > With '-mavx' or '-mavx512f' it looks like it does for
> your
> > > > > sample
> > > > > > > code.
> > > > > > > >
> > > > > > > > It also does with -msse4.2 when SZ is reduced to 256
> bytes.
> > > > > > > > Clang switches to inline when SZ is reduced to 128 bytes.
> > > > > > > >
> > > > > > > > It seems the compiler has a threshold for when to inline
> and
> > > when
> > > > > to
> > > > > > > call the C
> > > > > > > > library's memcpy subroutine.
> > > > > > > > The threshold depends on both copy size and vector
> register
> > > size.
> > > > > > > > And it is compiler dependent.
> > > > > > >
> > > > > > > I think there are compiler options to specify desired
> threshold
> > > > > values.
> > > > > > > Let say for gcc there is ' -mmemcpy-strategy=strategy'.
> > > > > > > For that example in that particular case
> > > > > > > -mmemcpy-strategy=vector_loop:512:align,loop:-1:align
> > > > > > > generates sse loads/stores.
> > > > > > > Might be we can exploit it somehow?
> > > > > >
> > > > > > That could give us higher granularity/control over memcpy for
> > > > > individual
> > > > > > memcpy instances; might be useful for hot code paths where we
> > > have
> > > > > more
> > > > > > knowledge about the copy operation than the compiler can
> infer.
> > > > > > However, pragmas are discouraged in DPDK, and this looks like
> a
> > > very
> > > > > similar
> > > > > > path.
> > > > >
> > > > > Well, right now rte_memcpy.h is 700+ lines and keeps growing.
> > > > > Considering that probably pragmas are not that bad.
> > > > > Of course, pragmas have their own issues and it is hard to
> ensure
> > > that
> > > > > they will produce same code between different
> compilers/versions,
> > > etc.
> > > > >
> > > > > > > I am not really happy that our home-brewed memcpy code-
> block
> > > keeps
> > > > > > > growing,
> > > > > > > while we keep talking that it would be good to eliminate it
> > > > > completely.
> > > > > >
> > > > > > I agree in principle.
> > > > > > However, this rte_memcpy() optimization is for the
> pile/mempool
> > > > > optimizations
> > > > > > I'm working on, so there is a specific use case motivating
> the
> > > added
> > > > > code.
> > > > >
> > > > > I understand that you probably have some specific use-case in
> mind.
> > > > > BTW for this optimization you mentioned above: what is the gain
> > > with
> > > > > these changes?
> > > >
> > > > IMO, the primary benefit is the much simpler (and smaller)
> assembly
> > > output due
> > > > to avoiding the address alignment check (and the resulting
> duplicated
> > > code).
> > >
> > > I think that should be measurable too: whole binary and/or hot
> path
> > > function size reduction, etc.
> >
> > The size of the generated code for the copy operation is reduced to
> slightly less
> > than half (only one instance of the copy operation instead of two,
> and the
> > address alignment comparison is omitted).
> >
> > >
> > > > I haven't measured the performance gain.
> > > > Based on the perf gain in a previous mempool optimization patch
> [1],
> > > it seems
> > > > avoiding the address alignment check shaves ~2 cycles off the
> copy
> > > operation (for
> > > > cache-to-cache copy).
> > > > I expect that the same gain (from avoiding the address alignment
> > > check) applies
> > > > here.
> > >
> > > Ok, then I suggest we do some measurements first, before going
> forward
> > > with it.
> >
> > I tried memcpy_perf_autotest, but the branch predictor kicks in and
> eliminates
> > the cost of the address alignment comparison. So the results are very
> similar.
>
> Indeed, they are the same.
>
> > BTW, that's a general issue with our perf tests: Repeated testing
> doesn't show
> > the cost of branches, because they are always eliminated by the
> branch
> > predictor.
>
> Might be...
> But we do need some measurable evidence that such change improve
> things,
> otherwise - why to bother?
> If that perf test is not goof enough - let's try to extend it or use
> different one.
> BTW, in your mempool pile RFC, I see you also use rte_memcpy over fixed
> sized buffers.
> Do you see any improvement there (wit/without rte_memcpy optimixation)?
Looking at the generated assembly (objdump -S build/drivers/librte_mempool_stack.so)...
This is the optimized copy loop in pile_dequeue:
1650: 89 ca mov %ecx,%edx
1652: c5 fe 6f 40 40 vmovdqu 0x40(%rax),%ymm0
1657: ff c1 inc %ecx
1659: c1 e2 05 shl $0x5,%edx
165c: 49 8d 14 d4 lea (%r12,%rdx,8),%rdx
1660: c5 fe 7f 02 vmovdqu %ymm0,(%rdx)
1664: c5 fe 6f 48 60 vmovdqu 0x60(%rax),%ymm1
1669: c5 fe 7f 4a 20 vmovdqu %ymm1,0x20(%rdx)
166e: c5 fe 6f 90 80 00 00 vmovdqu 0x80(%rax),%ymm2
1675: 00
1676: c5 fe 7f 52 40 vmovdqu %ymm2,0x40(%rdx)
167b: c5 fe 6f 98 a0 00 00 vmovdqu 0xa0(%rax),%ymm3
1682: 00
1683: c5 fe 7f 5a 60 vmovdqu %ymm3,0x60(%rdx)
1688: c5 fe 6f a0 c0 00 00 vmovdqu 0xc0(%rax),%ymm4
168f: 00
1690: c5 fe 7f a2 80 00 00 vmovdqu %ymm4,0x80(%rdx)
1697: 00
1698: c5 fe 6f a8 e0 00 00 vmovdqu 0xe0(%rax),%ymm5
169f: 00
16a0: c5 fe 7f aa a0 00 00 vmovdqu %ymm5,0xa0(%rdx)
16a7: 00
16a8: c5 fe 6f b0 00 01 00 vmovdqu 0x100(%rax),%ymm6
16af: 00
16b0: c5 fe 7f b2 c0 00 00 vmovdqu %ymm6,0xc0(%rdx)
16b7: 00
16b8: c5 fe 6f b8 20 01 00 vmovdqu 0x120(%rax),%ymm7
16bf: 00
16c0: c5 fe 7f ba e0 00 00 vmovdqu %ymm7,0xe0(%rdx)
16c7: 00
16c8: 48 8b 40 08 mov 0x8(%rax),%rax
16cc: 39 f1 cmp %esi,%ecx
16ce: 75 80 jne 1650 <pile_dequeue+0xe0>
16d0:
This is without the rte_memcpy optimization:
1650: 89 ca mov %ecx,%edx
1652: c5 fe 6f 40 40 vmovdqu 0x40(%rax),%ymm0
1657: 49 89 c1 mov %rax,%r9
165a: ff c1 inc %ecx
165c: c1 e2 05 shl $0x5,%edx
165f: 49 8d 14 d4 lea (%r12,%rdx,8),%rdx
1663: c5 fe 7f 02 vmovdqu %ymm0,(%rdx)
1667: c5 fe 6f 48 60 vmovdqu 0x60(%rax),%ymm1
166c: 49 09 d1 or %rdx,%r9
166f: 41 83 e1 1f and $0x1f,%r9d
1673: c5 fe 7f 4a 20 vmovdqu %ymm1,0x20(%rdx)
1678: c5 fe 6f 90 80 00 00 vmovdqu 0x80(%rax),%ymm2
167f: 00
1680: c5 fe 7f 52 40 vmovdqu %ymm2,0x40(%rdx)
1685: c5 fe 6f 98 a0 00 00 vmovdqu 0xa0(%rax),%ymm3
168c: 00
168d: c5 fe 7f 5a 60 vmovdqu %ymm3,0x60(%rdx)
1692: c5 fe 6f a0 c0 00 00 vmovdqu 0xc0(%rax),%ymm4
1699: 00
169a: c5 fe 7f a2 80 00 00 vmovdqu %ymm4,0x80(%rdx)
16a1: 00
16a2: c5 fe 6f a8 e0 00 00 vmovdqu 0xe0(%rax),%ymm5
16a9: 00
16aa: c5 fe 7f aa a0 00 00 vmovdqu %ymm5,0xa0(%rdx)
16b1: 00
16b2: c5 fe 6f b0 00 01 00 vmovdqu 0x100(%rax),%ymm6
16b9: 00
16ba: c5 fe 7f b2 c0 00 00 vmovdqu %ymm6,0xc0(%rdx)
16c1: 00
16c2: c5 fe 6f b8 20 01 00 vmovdqu 0x120(%rax),%ymm7
16c9: 00
16ca: c5 fe 7f ba e0 00 00 vmovdqu %ymm7,0xe0(%rdx)
16d1: 00
16d2: 48 8b 40 08 mov 0x8(%rax),%rax
16d6: 0f 85 c4 00 00 00 jne 17a0 <pile_dequeue+0x230>
16dc: 39 ce cmp %ecx,%esi
16de: 0f 85 6c ff ff ff jne 1650 <pile_dequeue+0xe0>
16e4:
[...]
17a0: 39 f1 cmp %esi,%ecx
17a2: 0f 85 a8 fe ff ff jne 1650 <pile_dequeue+0xe0>
17a8: e9 37 ff ff ff jmp 16e4 <pile_dequeue+0x174>
Here, the alignment comparison I have optimized away is performed using register %r9.
With the optimization, the generated assembly is less cluttered, and thus easier to review.
(It seems the compiler is clever enough to not duplicate the two instances of the code, but reuse the aligned instance for the unaligned case too. So the reduction in code size is not as great as I previously claimed.)
I agree the performance benefit in CPU cycles is probably insignificant.
But the generated assembly is cleaner.
And if pressured for CPU registers, the optimization also frees up one CPU register for other purposes.
Code size reduction:
The optimized loop is 0x80 bytes of instructions.
The non-optimized is 0x8e bytes, 14 bytes more, in the loop, plus 13 bytes outside the loop.
>
> > ======= ================= ================= =================
> > =================
> > Size Cache to cache Cache to mem Mem to cache
> Mem to mem
> > (bytes) (ticks) (ticks) (ticks)
> (ticks)
> > ------- ----------------- ----------------- ----------------- -------
> ----------
> > ================================= 32B aligned
> > =================================
> > Before
> > C 64 2 - 2( 9.60%) 22 - 21( 4.48%) 33 - 31( 5.15%) 57 -
> 56( 0.28%)
> > C 128 4 - 4( 5.56%) 41 - 40( 3.59%) 53 - 52( 1.49%) 101 -
> 100( 0.71%)
> > C 192 6 - 6( 4.06%) 64 - 60( 5.61%) 68 - 69( -1.68%) 136 -
> 138( -1.49%)
> > C 256 10 - 9( 5.21%) 82 - 81( 1.20%) 93 - 87( 7.47%) 172 -
> 175( -1.46%)
> > After
> > C 64 2 - 2( -0.71%) 22 - 21( 7.62%) 31 - 30( 4.32%) 57 -
> 59( -3.03%)
> > C 128 4 - 4( -0.34%) 40 - 40( 0.60%) 50 - 50( 0.41%) 97 -
> 99( -1.96%)
> > C 192 6 - 6( 5.27%) 59 - 59( -0.27%) 69 - 72( -4.42%) 139 -
> 138( 0.90%)
> > C 256 9 - 9( 9.14%) 77 - 77( 0.25%) 89 - 90( -0.88%) 180 -
> 177( 1.57%)
> > ================================== Unaligned
> > ==================================
> > Before
> > C 64 5 - 5( 0.98%) 32 - 32( 0.72%) 42 - 42( 0.44%) 82 -
> 82( 0.51%)
> > C 128 8 - 9(-11.26%) 51 - 50( 0.53%) 60 - 59( 0.89%) 117 -
> 115( 1.37%)
> > C 192 13 - 12( 1.62%) 66 - 67( -1.12%) 83 - 80( 3.26%) 157 -
> 165( -5.25%)
> > C 256 21 - 22( -5.78%) 94 - 93( 1.65%) 100 - 99( 1.52%) 202 -
> 204( -1.34%)
> > After
> > C 64 5 - 5( 0.05%) 31 - 31( 0.16%) 40 - 39( 1.37%) 77 -
> 77( -0.93%)
> > C 128 9 - 9( -0.69%) 50 - 50( -0.02%) 61 - 60( 2.71%) 118 -
> 122( -3.15%)
> > C 192 12 - 11( 7.68%) 68 - 66( 3.64%) 79 - 81( -2.85%) 164 -
> 164( -0.12%)
> > C 256 21 - 20( 1.80%) 90 - 87( 2.70%) 96 - 97( -0.99%) 202 -
> 201( 0.09%)
> >
> > >
> > > > [1]:
> > >
> https://patchwork.dpdk.org/project/dpdk/patch/20260521185631.116046-1-
> > > > mb@smartsharesystems.com/
next prev parent reply other threads:[~2026-08-11 11:10 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-04 14:33 [PATCH] eal/x86: optimize memcpy of small 64-byte blocks Morten Brørup
2026-08-04 14:33 ` [RFC PATCH] pile stack and mempool driver (resend) Morten Brørup
2026-08-04 14:38 ` Morten Brørup
2026-08-04 15:52 ` [PATCH] eal/x86: optimize memcpy of small 64-byte blocks Stephen Hemminger
2026-08-04 16:25 ` Morten Brørup
2026-08-05 5:45 ` Konstantin Ananyev
2026-08-05 5:56 ` Morten Brørup
2026-08-05 6:50 ` Konstantin Ananyev
2026-08-05 8:36 ` Morten Brørup
2026-08-05 17:16 ` Konstantin Ananyev
2026-08-05 20:28 ` Morten Brørup
2026-08-06 7:58 ` Konstantin Ananyev
2026-08-06 9:17 ` Morten Brørup
2026-08-11 7:05 ` Konstantin Ananyev
2026-08-11 11:10 ` Morten Brørup [this message]
2026-08-04 19:11 ` Morten Brørup
2026-08-04 20:42 ` Stephen Hemminger
2026-08-05 17:10 ` Konstantin Ananyev
2026-08-05 20:16 ` Morten Brørup
2026-08-06 7:43 ` Konstantin Ananyev
2026-08-06 8:00 ` Morten Brørup
2026-08-06 10:17 ` [PATCH v2] " Morten Brørup
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=98CBD80474FA8B44BF855DF32C47DC35F659D3@smartserver.smartshare.dk \
--to=mb@smartsharesystems.com \
--cc=bruce.richardson@intel.com \
--cc=dev@dpdk.org \
--cc=konstantin.ananyev@huawei.com \
--cc=stephen@networkplumber.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox