> From: Konstantin Ananyev [mailto:[email protected]]
> Sent: Tuesday, 18 August 2026 10.15
> 
> > > > > > > > > > > > > > +   /* Common way for small copy size of 64-
> byte
> > > > > blocks.
> > > > > > > > > > > Unlikely, so
> > > > > > > > > > > > > constant size only */
> > > > > > > > > > > > > > +   if (__rte_constant(n) && (n & 63) == 0 &&
> n <=
> > > > > > > > > > > > > RTE_MEMCPY_BLOCK_64_MAX) {
> > > > > > > > > > > > > > +           void *ret = dst;
> > > > > > > > > > > > > > +
> > > > > > > > > > > > >
> > > > > > > > > > > > > Maybe just let compiler decide, it will
> generate
> > > vector
> > > > > > > > > > > instructions in
> > > > > > > > > > > > > most cases.
> > > > > > > > > > > > >
> > > > > > > > > > > > >       if (__rte_constant(n))
> > > > > > > > > > > > >               return mempcpy(dst, src, n);
> > > > > > > > > > > >
> > > > > > > > > > > > Maybe in most, but not in all:
> > > > > > > > > > > > https://godbolt.org/z/KvdKqT5rY
> > > > > > > > > > >
> > > > > > > > > > > With '-mavx' or '-mavx512f' it looks like it does
> for
> > > your
> > > > > > > sample
> > > > > > > > > code.
> > > > > > > > > >
> > > > > > > > > > It also does with -msse4.2 when SZ is reduced to 256
> > > bytes.
> > > > > > > > > > Clang switches to inline when SZ is reduced to 128
> bytes.
> > > > > > > > > >
> > > > > > > > > > It seems the compiler has a threshold for when to
> inline
> > > and
> > > > > when
> > > > > > > to
> > > > > > > > > call the C
> > > > > > > > > > library's memcpy subroutine.
> > > > > > > > > > The threshold depends on both copy size and vector
> > > register
> > > > > size.
> > > > > > > > > > And it is compiler dependent.
> > > > > > > > >
> > > > > > > > > I think there are compiler options to specify desired
> > > threshold
> > > > > > > values.
> > > > > > > > > Let say for gcc there is  ' -mmemcpy-
> strategy=strategy'.
> > > > > > > > > For that example in that particular case
> > > > > > > > > -mmemcpy-strategy=vector_loop:512:align,loop:-1:align
> > > > > > > > > generates sse loads/stores.
> > > > > > > > > Might be we can exploit it somehow?
> > > > > > > >
> > > > > > > > That could give us higher granularity/control over memcpy
> for
> > > > > > > individual
> > > > > > > > memcpy instances; might be useful for hot code paths
> where we
> > > > > have
> > > > > > > more
> > > > > > > > knowledge about the copy operation than the compiler can
> > > infer.
> > > > > > > > However, pragmas are discouraged in DPDK, and this looks
> like
> > > a
> > > > > very
> > > > > > > similar
> > > > > > > > path.
> > > > > > >
> > > > > > > Well, right now  rte_memcpy.h is 700+ lines and keeps
> growing.
> > > > > > > Considering that probably pragmas are not that bad.
> > > > > > > Of course, pragmas have their own issues and it is hard to
> > > ensure
> > > > > that
> > > > > > > they will produce same code between different
> > > compilers/versions,
> > > > > etc.
> > > > > > >
> > > > > > > > > I am not really happy that our home-brewed memcpy code-
> > > block
> > > > > keeps
> > > > > > > > > growing,
> > > > > > > > > while we keep talking that it would be good to
> eliminate it
> > > > > > > completely.
> > > > > > > >
> > > > > > > > I agree in principle.
> > > > > > > > However, this rte_memcpy() optimization is for the
> > > pile/mempool
> > > > > > > optimizations
> > > > > > > > I'm working on, so there is a specific use case
> motivating
> > > the
> > > > > added
> > > > > > > code.
> > > > > > >
> > > > > > > I understand that you probably have some specific use-case
> in
> > > mind.
> > > > > > > BTW for this optimization you mentioned above: what is the
> gain
> > > > > with
> > > > > > > these changes?
> > > > > >
> > > > > > IMO, the primary benefit is the much simpler (and smaller)
> > > assembly
> > > > > output due
> > > > > > to avoiding the address alignment check (and the resulting
> > > duplicated
> > > > > code).
> > > > >
> > > > > I think that should be measurable too: whole binary  and/or hot
> > > path
> > > > > function size reduction, etc.
> > > >
> > > > The size of the generated code for the copy operation is reduced
> to
> > > slightly less
> > > > than half (only one instance of the copy operation instead of
> two,
> > > and the
> > > > address alignment comparison is omitted).
> > > >
> > > > >
> > > > > > I haven't measured the performance gain.
> > > > > > Based on the perf gain in a previous mempool optimization
> patch
> > > [1],
> > > > > it seems
> > > > > > avoiding the address alignment check shaves ~2 cycles off the
> > > copy
> > > > > operation (for
> > > > > > cache-to-cache copy).
> > > > > > I expect that the same gain (from avoiding the address
> alignment
> > > > > check) applies
> > > > > > here.
> > > > >
> > > > > Ok, then I suggest we do some measurements first, before going
> > > forward
> > > > > with it.
> > > >
> > > > I tried memcpy_perf_autotest, but the branch predictor kicks in
> and
> > > eliminates
> > > > the cost of the address alignment comparison. So the results are
> very
> > > similar.
> > >
> > > Indeed, they are the same.
> > >
> > > > BTW, that's a general issue with our perf tests: Repeated testing
> > > doesn't show
> > > > the cost of branches, because they are always eliminated by the
> > > branch
> > > > predictor.
> > >
> > > Might be...
> > > But we do need some measurable evidence that such change improve
> > > things,
> > > otherwise - why to bother?
> > > If that perf test is not goof enough - let's try to extend it or
> use
> > > different one.
> > > BTW, in your mempool pile RFC, I see you also use rte_memcpy over
> fixed
> > > sized buffers.
> > > Do you see any improvement there (wit/without rte_memcpy
> optimixation)?
> >
> > Looking at the generated assembly (objdump -S
> > build/drivers/librte_mempool_stack.so)...
> >
> > This is the optimized copy loop in pile_dequeue:
> >
> >     1650:   89 ca                   mov    %ecx,%edx
> >     1652:   c5 fe 6f 40 40          vmovdqu 0x40(%rax),%ymm0
> >     1657:   ff c1                   inc    %ecx
> >     1659:   c1 e2 05                shl    $0x5,%edx
> >     165c:   49 8d 14 d4             lea    (%r12,%rdx,8),%rdx
> >     1660:   c5 fe 7f 02             vmovdqu %ymm0,(%rdx)
> >     1664:   c5 fe 6f 48 60          vmovdqu 0x60(%rax),%ymm1
> >     1669:   c5 fe 7f 4a 20          vmovdqu %ymm1,0x20(%rdx)
> >     166e:   c5 fe 6f 90 80 00 00    vmovdqu 0x80(%rax),%ymm2
> >     1675:   00
> >     1676:   c5 fe 7f 52 40          vmovdqu %ymm2,0x40(%rdx)
> >     167b:   c5 fe 6f 98 a0 00 00    vmovdqu 0xa0(%rax),%ymm3
> >     1682:   00
> >     1683:   c5 fe 7f 5a 60          vmovdqu %ymm3,0x60(%rdx)
> >     1688:   c5 fe 6f a0 c0 00 00    vmovdqu 0xc0(%rax),%ymm4
> >     168f:   00
> >     1690:   c5 fe 7f a2 80 00 00    vmovdqu %ymm4,0x80(%rdx)
> >     1697:   00
> >     1698:   c5 fe 6f a8 e0 00 00    vmovdqu 0xe0(%rax),%ymm5
> >     169f:   00
> >     16a0:   c5 fe 7f aa a0 00 00    vmovdqu %ymm5,0xa0(%rdx)
> >     16a7:   00
> >     16a8:   c5 fe 6f b0 00 01 00    vmovdqu 0x100(%rax),%ymm6
> >     16af:   00
> >     16b0:   c5 fe 7f b2 c0 00 00    vmovdqu %ymm6,0xc0(%rdx)
> >     16b7:   00
> >     16b8:   c5 fe 6f b8 20 01 00    vmovdqu 0x120(%rax),%ymm7
> >     16bf:   00
> >     16c0:   c5 fe 7f ba e0 00 00    vmovdqu %ymm7,0xe0(%rdx)
> >     16c7:   00
> >     16c8:   48 8b 40 08             mov    0x8(%rax),%rax
> >     16cc:   39 f1                   cmp    %esi,%ecx
> >     16ce:   75 80                   jne    1650 <pile_dequeue+0xe0>
> >     16d0:
> >
> > This is without the rte_memcpy optimization:
> >
> >     1650:   89 ca                   mov    %ecx,%edx
> >     1652:   c5 fe 6f 40 40          vmovdqu 0x40(%rax),%ymm0
> >     1657:   49 89 c1                mov    %rax,%r9
> >     165a:   ff c1                   inc    %ecx
> >     165c:   c1 e2 05                shl    $0x5,%edx
> >     165f:   49 8d 14 d4             lea    (%r12,%rdx,8),%rdx
> >     1663:   c5 fe 7f 02             vmovdqu %ymm0,(%rdx)
> >     1667:   c5 fe 6f 48 60          vmovdqu 0x60(%rax),%ymm1
> >     166c:   49 09 d1                or     %rdx,%r9
> >     166f:   41 83 e1 1f             and    $0x1f,%r9d
> >     1673:   c5 fe 7f 4a 20          vmovdqu %ymm1,0x20(%rdx)
> >     1678:   c5 fe 6f 90 80 00 00    vmovdqu 0x80(%rax),%ymm2
> >     167f:   00
> >     1680:   c5 fe 7f 52 40          vmovdqu %ymm2,0x40(%rdx)
> >     1685:   c5 fe 6f 98 a0 00 00    vmovdqu 0xa0(%rax),%ymm3
> >     168c:   00
> >     168d:   c5 fe 7f 5a 60          vmovdqu %ymm3,0x60(%rdx)
> >     1692:   c5 fe 6f a0 c0 00 00    vmovdqu 0xc0(%rax),%ymm4
> >     1699:   00
> >     169a:   c5 fe 7f a2 80 00 00    vmovdqu %ymm4,0x80(%rdx)
> >     16a1:   00
> >     16a2:   c5 fe 6f a8 e0 00 00    vmovdqu 0xe0(%rax),%ymm5
> >     16a9:   00
> >     16aa:   c5 fe 7f aa a0 00 00    vmovdqu %ymm5,0xa0(%rdx)
> >     16b1:   00
> >     16b2:   c5 fe 6f b0 00 01 00    vmovdqu 0x100(%rax),%ymm6
> >     16b9:   00
> >     16ba:   c5 fe 7f b2 c0 00 00    vmovdqu %ymm6,0xc0(%rdx)
> >     16c1:   00
> >     16c2:   c5 fe 6f b8 20 01 00    vmovdqu 0x120(%rax),%ymm7
> >     16c9:   00
> >     16ca:   c5 fe 7f ba e0 00 00    vmovdqu %ymm7,0xe0(%rdx)
> >     16d1:   00
> >     16d2:   48 8b 40 08             mov    0x8(%rax),%rax
> >     16d6:   0f 85 c4 00 00 00       jne    17a0 <pile_dequeue+0x230>
> >     16dc:   39 ce                   cmp    %ecx,%esi
> >     16de:   0f 85 6c ff ff ff       jne    1650 <pile_dequeue+0xe0>
> >     16e4:
> >     [...]
> >     17a0:   39 f1                   cmp    %esi,%ecx
> >     17a2:   0f 85 a8 fe ff ff       jne    1650 <pile_dequeue+0xe0>
> >     17a8:   e9 37 ff ff ff          jmp    16e4 <pile_dequeue+0x174>
> >
> > Here, the alignment comparison I have optimized away is performed
> using
> > register %r9.
> >
> > With the optimization, the generated assembly is less cluttered, and
> thus easier
> > to review.
> > (It seems the compiler is clever enough to not duplicate the two
> instances of the
> > code, but reuse the aligned instance for the unaligned case too. So
> the reduction
> > in code size is not as great as I previously claimed.)
> >
> > I agree the performance benefit in CPU cycles is probably
> insignificant.
> > But the generated assembly is cleaner.
> > And if pressured for CPU registers, the optimization also frees up
> one CPU
> > register for other purposes.
> >
> > Code size reduction:
> > The optimized loop is 0x80 bytes of instructions.
> > The non-optimized is 0x8e bytes, 14 bytes more, in the loop, plus 13
> bytes
> > outside the loop.
> 
> Honestly, with such insignificant gains, I'd either leave it alone,
> or look more closely at compiler pragmas option we discussed before.

Compiled code is slightly smaller, performance is slightly better.
But source code having slightly more lines of code is more important?

If the source code added was difficult to read, complexity was increased, or 
had side effects or spillover to other modules, I might buy that argument.
But the addition is very simple and completely isolated.

If the objection is about source code readability, I could add more inline 
comments to the added code, elaborating that the added code path is the same 
for all CPU vector sizes, regardless of address alignment.
But I suspect such comments would confuse more than they would help.
And it would add even more lines to the file size.

Adding compiler pragmas should probably be done in source code calling 
rte_memcpy(), like __rte_assume() and the coming __rte_assume_cache_aligned() 
[1].

[1]: 
https://patchwork.dpdk.org/project/dpdk/patch/[email protected]/

If you try experimenting with compiler pragmas, I'm open for reviewing an RFC!
For now, let's take this small step.

> 
> >
> > >
> > > > ======= ================= ================= =================
> > > > =================
> > > >    Size   Cache to cache     Cache to mem      Mem to cache
> > > Mem to mem
> > > > (bytes)          (ticks)          (ticks)           (ticks)
> > > (ticks)
> > > > ------- ----------------- ----------------- ----------------- ---
> ----
> > > ----------
> > > > ================================= 32B aligned
> > > > =================================
> > > > Before
> > > > C    64  2 -  2(  9.60%)  22 - 21(  4.48%)  33 - 31(  5.15%)  57
> -
> > > 56(  0.28%)
> > > > C   128  4 -  4(  5.56%)  41 - 40(  3.59%)  53 - 52(  1.49%) 101
> -
> > > 100(  0.71%)
> > > > C   192  6 -  6(  4.06%)  64 - 60(  5.61%)  68 - 69( -1.68%) 136
> -
> > > 138( -1.49%)
> > > > C   256 10 -  9(  5.21%)  82 - 81(  1.20%)  93 - 87(  7.47%) 172
> -
> > > 175( -1.46%)
> > > > After
> > > > C    64  2 -  2( -0.71%)  22 - 21(  7.62%)  31 - 30(  4.32%)  57
> -
> > > 59( -3.03%)
> > > > C   128  4 -  4( -0.34%)  40 - 40(  0.60%)  50 - 50(  0.41%)  97
> -
> > > 99( -1.96%)
> > > > C   192  6 -  6(  5.27%)  59 - 59( -0.27%)  69 - 72( -4.42%) 139
> -
> > > 138(  0.90%)
> > > > C   256  9 -  9(  9.14%)  77 - 77(  0.25%)  89 - 90( -0.88%) 180
> -
> > > 177(  1.57%)
> > > > ================================== Unaligned
> > > > ==================================
> > > > Before
> > > > C    64  5 -  5(  0.98%)  32 - 32(  0.72%)  42 - 42(  0.44%)  82
> -
> > > 82(  0.51%)
> > > > C   128  8 -  9(-11.26%)  51 - 50(  0.53%)  60 - 59(  0.89%) 117
> -
> > > 115(  1.37%)
> > > > C   192 13 - 12(  1.62%)  66 - 67( -1.12%)  83 - 80(  3.26%) 157
> -
> > > 165( -5.25%)
> > > > C   256 21 - 22( -5.78%)  94 - 93(  1.65%) 100 - 99(  1.52%) 202
> -
> > > 204( -1.34%)
> > > > After
> > > > C    64  5 -  5(  0.05%)  31 - 31(  0.16%)  40 - 39(  1.37%)  77
> -
> > > 77( -0.93%)
> > > > C   128  9 -  9( -0.69%)  50 - 50( -0.02%)  61 - 60(  2.71%) 118
> -
> > > 122( -3.15%)
> > > > C   192 12 - 11(  7.68%)  68 - 66(  3.64%)  79 - 81( -2.85%) 164
> -
> > > 164( -0.12%)
> > > > C   256 21 - 20(  1.80%)  90 - 87(  2.70%)  96 - 97( -0.99%) 202
> -
> > > 201(  0.09%)
> > > >
> > > > >
> > > > > > [1]:
> > > > >
> > >
> https://patchwork.dpdk.org/project/dpdk/patch/20260521185631.116046-1-
> > > > > > [email protected]/

Reply via email to