> > > > > > > > > > > > > +     /* Common way for small copy size of 64-byte
> > > > blocks.
> > > > > > > > > > Unlikely, so
> > > > > > > > > > > > constant size only */
> > > > > > > > > > > > > +     if (__rte_constant(n) && (n & 63) == 0 && n <=
> > > > > > > > > > > > RTE_MEMCPY_BLOCK_64_MAX) {
> > > > > > > > > > > > > +             void *ret = dst;
> > > > > > > > > > > > > +
> > > > > > > > > > > >
> > > > > > > > > > > > Maybe just let compiler decide, it will generate
> > vector
> > > > > > > > > > instructions in
> > > > > > > > > > > > most cases.
> > > > > > > > > > > >
> > > > > > > > > > > >         if (__rte_constant(n))
> > > > > > > > > > > >                 return mempcpy(dst, src, n);
> > > > > > > > > > >
> > > > > > > > > > > Maybe in most, but not in all:
> > > > > > > > > > > https://godbolt.org/z/KvdKqT5rY
> > > > > > > > > >
> > > > > > > > > > With '-mavx' or '-mavx512f' it looks like it does for
> > your
> > > > > > sample
> > > > > > > > code.
> > > > > > > > >
> > > > > > > > > It also does with -msse4.2 when SZ is reduced to 256
> > bytes.
> > > > > > > > > Clang switches to inline when SZ is reduced to 128 bytes.
> > > > > > > > >
> > > > > > > > > It seems the compiler has a threshold for when to inline
> > and
> > > > when
> > > > > > to
> > > > > > > > call the C
> > > > > > > > > library's memcpy subroutine.
> > > > > > > > > The threshold depends on both copy size and vector
> > register
> > > > size.
> > > > > > > > > And it is compiler dependent.
> > > > > > > >
> > > > > > > > I think there are compiler options to specify desired
> > threshold
> > > > > > values.
> > > > > > > > Let say for gcc there is  ' -mmemcpy-strategy=strategy'.
> > > > > > > > For that example in that particular case
> > > > > > > > -mmemcpy-strategy=vector_loop:512:align,loop:-1:align
> > > > > > > > generates sse loads/stores.
> > > > > > > > Might be we can exploit it somehow?
> > > > > > >
> > > > > > > That could give us higher granularity/control over memcpy for
> > > > > > individual
> > > > > > > memcpy instances; might be useful for hot code paths where we
> > > > have
> > > > > > more
> > > > > > > knowledge about the copy operation than the compiler can
> > infer.
> > > > > > > However, pragmas are discouraged in DPDK, and this looks like
> > a
> > > > very
> > > > > > similar
> > > > > > > path.
> > > > > >
> > > > > > Well, right now  rte_memcpy.h is 700+ lines and keeps growing.
> > > > > > Considering that probably pragmas are not that bad.
> > > > > > Of course, pragmas have their own issues and it is hard to
> > ensure
> > > > that
> > > > > > they will produce same code between different
> > compilers/versions,
> > > > etc.
> > > > > >
> > > > > > > > I am not really happy that our home-brewed memcpy code-
> > block
> > > > keeps
> > > > > > > > growing,
> > > > > > > > while we keep talking that it would be good to eliminate it
> > > > > > completely.
> > > > > > >
> > > > > > > I agree in principle.
> > > > > > > However, this rte_memcpy() optimization is for the
> > pile/mempool
> > > > > > optimizations
> > > > > > > I'm working on, so there is a specific use case motivating
> > the
> > > > added
> > > > > > code.
> > > > > >
> > > > > > I understand that you probably have some specific use-case in
> > mind.
> > > > > > BTW for this optimization you mentioned above: what is the gain
> > > > with
> > > > > > these changes?
> > > > >
> > > > > IMO, the primary benefit is the much simpler (and smaller)
> > assembly
> > > > output due
> > > > > to avoiding the address alignment check (and the resulting
> > duplicated
> > > > code).
> > > >
> > > > I think that should be measurable too: whole binary  and/or hot
> > path
> > > > function size reduction, etc.
> > >
> > > The size of the generated code for the copy operation is reduced to
> > slightly less
> > > than half (only one instance of the copy operation instead of two,
> > and the
> > > address alignment comparison is omitted).
> > >
> > > >
> > > > > I haven't measured the performance gain.
> > > > > Based on the perf gain in a previous mempool optimization patch
> > [1],
> > > > it seems
> > > > > avoiding the address alignment check shaves ~2 cycles off the
> > copy
> > > > operation (for
> > > > > cache-to-cache copy).
> > > > > I expect that the same gain (from avoiding the address alignment
> > > > check) applies
> > > > > here.
> > > >
> > > > Ok, then I suggest we do some measurements first, before going
> > forward
> > > > with it.
> > >
> > > I tried memcpy_perf_autotest, but the branch predictor kicks in and
> > eliminates
> > > the cost of the address alignment comparison. So the results are very
> > similar.
> >
> > Indeed, they are the same.
> >
> > > BTW, that's a general issue with our perf tests: Repeated testing
> > doesn't show
> > > the cost of branches, because they are always eliminated by the
> > branch
> > > predictor.
> >
> > Might be...
> > But we do need some measurable evidence that such change improve
> > things,
> > otherwise - why to bother?
> > If that perf test is not goof enough - let's try to extend it or use
> > different one.
> > BTW, in your mempool pile RFC, I see you also use rte_memcpy over fixed
> > sized buffers.
> > Do you see any improvement there (wit/without rte_memcpy optimixation)?
> 
> Looking at the generated assembly (objdump -S
> build/drivers/librte_mempool_stack.so)...
> 
> This is the optimized copy loop in pile_dequeue:
> 
>     1650:     89 ca                   mov    %ecx,%edx
>     1652:     c5 fe 6f 40 40          vmovdqu 0x40(%rax),%ymm0
>     1657:     ff c1                   inc    %ecx
>     1659:     c1 e2 05                shl    $0x5,%edx
>     165c:     49 8d 14 d4             lea    (%r12,%rdx,8),%rdx
>     1660:     c5 fe 7f 02             vmovdqu %ymm0,(%rdx)
>     1664:     c5 fe 6f 48 60          vmovdqu 0x60(%rax),%ymm1
>     1669:     c5 fe 7f 4a 20          vmovdqu %ymm1,0x20(%rdx)
>     166e:     c5 fe 6f 90 80 00 00    vmovdqu 0x80(%rax),%ymm2
>     1675:     00
>     1676:     c5 fe 7f 52 40          vmovdqu %ymm2,0x40(%rdx)
>     167b:     c5 fe 6f 98 a0 00 00    vmovdqu 0xa0(%rax),%ymm3
>     1682:     00
>     1683:     c5 fe 7f 5a 60          vmovdqu %ymm3,0x60(%rdx)
>     1688:     c5 fe 6f a0 c0 00 00    vmovdqu 0xc0(%rax),%ymm4
>     168f:     00
>     1690:     c5 fe 7f a2 80 00 00    vmovdqu %ymm4,0x80(%rdx)
>     1697:     00
>     1698:     c5 fe 6f a8 e0 00 00    vmovdqu 0xe0(%rax),%ymm5
>     169f:     00
>     16a0:     c5 fe 7f aa a0 00 00    vmovdqu %ymm5,0xa0(%rdx)
>     16a7:     00
>     16a8:     c5 fe 6f b0 00 01 00    vmovdqu 0x100(%rax),%ymm6
>     16af:     00
>     16b0:     c5 fe 7f b2 c0 00 00    vmovdqu %ymm6,0xc0(%rdx)
>     16b7:     00
>     16b8:     c5 fe 6f b8 20 01 00    vmovdqu 0x120(%rax),%ymm7
>     16bf:     00
>     16c0:     c5 fe 7f ba e0 00 00    vmovdqu %ymm7,0xe0(%rdx)
>     16c7:     00
>     16c8:     48 8b 40 08             mov    0x8(%rax),%rax
>     16cc:     39 f1                   cmp    %esi,%ecx
>     16ce:     75 80                   jne    1650 <pile_dequeue+0xe0>
>     16d0:
> 
> This is without the rte_memcpy optimization:
> 
>     1650:     89 ca                   mov    %ecx,%edx
>     1652:     c5 fe 6f 40 40          vmovdqu 0x40(%rax),%ymm0
>     1657:     49 89 c1                mov    %rax,%r9
>     165a:     ff c1                   inc    %ecx
>     165c:     c1 e2 05                shl    $0x5,%edx
>     165f:     49 8d 14 d4             lea    (%r12,%rdx,8),%rdx
>     1663:     c5 fe 7f 02             vmovdqu %ymm0,(%rdx)
>     1667:     c5 fe 6f 48 60          vmovdqu 0x60(%rax),%ymm1
>     166c:     49 09 d1                or     %rdx,%r9
>     166f:     41 83 e1 1f             and    $0x1f,%r9d
>     1673:     c5 fe 7f 4a 20          vmovdqu %ymm1,0x20(%rdx)
>     1678:     c5 fe 6f 90 80 00 00    vmovdqu 0x80(%rax),%ymm2
>     167f:     00
>     1680:     c5 fe 7f 52 40          vmovdqu %ymm2,0x40(%rdx)
>     1685:     c5 fe 6f 98 a0 00 00    vmovdqu 0xa0(%rax),%ymm3
>     168c:     00
>     168d:     c5 fe 7f 5a 60          vmovdqu %ymm3,0x60(%rdx)
>     1692:     c5 fe 6f a0 c0 00 00    vmovdqu 0xc0(%rax),%ymm4
>     1699:     00
>     169a:     c5 fe 7f a2 80 00 00    vmovdqu %ymm4,0x80(%rdx)
>     16a1:     00
>     16a2:     c5 fe 6f a8 e0 00 00    vmovdqu 0xe0(%rax),%ymm5
>     16a9:     00
>     16aa:     c5 fe 7f aa a0 00 00    vmovdqu %ymm5,0xa0(%rdx)
>     16b1:     00
>     16b2:     c5 fe 6f b0 00 01 00    vmovdqu 0x100(%rax),%ymm6
>     16b9:     00
>     16ba:     c5 fe 7f b2 c0 00 00    vmovdqu %ymm6,0xc0(%rdx)
>     16c1:     00
>     16c2:     c5 fe 6f b8 20 01 00    vmovdqu 0x120(%rax),%ymm7
>     16c9:     00
>     16ca:     c5 fe 7f ba e0 00 00    vmovdqu %ymm7,0xe0(%rdx)
>     16d1:     00
>     16d2:     48 8b 40 08             mov    0x8(%rax),%rax
>     16d6:     0f 85 c4 00 00 00       jne    17a0 <pile_dequeue+0x230>
>     16dc:     39 ce                   cmp    %ecx,%esi
>     16de:     0f 85 6c ff ff ff       jne    1650 <pile_dequeue+0xe0>
>     16e4:
>     [...]
>     17a0:     39 f1                   cmp    %esi,%ecx
>     17a2:     0f 85 a8 fe ff ff       jne    1650 <pile_dequeue+0xe0>
>     17a8:     e9 37 ff ff ff          jmp    16e4 <pile_dequeue+0x174>
> 
> Here, the alignment comparison I have optimized away is performed using
> register %r9.
> 
> With the optimization, the generated assembly is less cluttered, and thus 
> easier
> to review.
> (It seems the compiler is clever enough to not duplicate the two instances of 
> the
> code, but reuse the aligned instance for the unaligned case too. So the 
> reduction
> in code size is not as great as I previously claimed.)
> 
> I agree the performance benefit in CPU cycles is probably insignificant.
> But the generated assembly is cleaner.
> And if pressured for CPU registers, the optimization also frees up one CPU
> register for other purposes.
> 
> Code size reduction:
> The optimized loop is 0x80 bytes of instructions.
> The non-optimized is 0x8e bytes, 14 bytes more, in the loop, plus 13 bytes
> outside the loop.

Honestly, with such insignificant gains, I'd either leave it alone,
or look more closely at compiler pragmas option we discussed before.

> 
> >
> > > ======= ================= ================= =================
> > > =================
> > >    Size   Cache to cache     Cache to mem      Mem to cache
> > Mem to mem
> > > (bytes)          (ticks)          (ticks)           (ticks)
> > (ticks)
> > > ------- ----------------- ----------------- ----------------- -------
> > ----------
> > > ================================= 32B aligned
> > > =================================
> > > Before
> > > C    64  2 -  2(  9.60%)  22 - 21(  4.48%)  33 - 31(  5.15%)  57 -
> > 56(  0.28%)
> > > C   128  4 -  4(  5.56%)  41 - 40(  3.59%)  53 - 52(  1.49%) 101 -
> > 100(  0.71%)
> > > C   192  6 -  6(  4.06%)  64 - 60(  5.61%)  68 - 69( -1.68%) 136 -
> > 138( -1.49%)
> > > C   256 10 -  9(  5.21%)  82 - 81(  1.20%)  93 - 87(  7.47%) 172 -
> > 175( -1.46%)
> > > After
> > > C    64  2 -  2( -0.71%)  22 - 21(  7.62%)  31 - 30(  4.32%)  57 -
> > 59( -3.03%)
> > > C   128  4 -  4( -0.34%)  40 - 40(  0.60%)  50 - 50(  0.41%)  97 -
> > 99( -1.96%)
> > > C   192  6 -  6(  5.27%)  59 - 59( -0.27%)  69 - 72( -4.42%) 139 -
> > 138(  0.90%)
> > > C   256  9 -  9(  9.14%)  77 - 77(  0.25%)  89 - 90( -0.88%) 180 -
> > 177(  1.57%)
> > > ================================== Unaligned
> > > ==================================
> > > Before
> > > C    64  5 -  5(  0.98%)  32 - 32(  0.72%)  42 - 42(  0.44%)  82 -
> > 82(  0.51%)
> > > C   128  8 -  9(-11.26%)  51 - 50(  0.53%)  60 - 59(  0.89%) 117 -
> > 115(  1.37%)
> > > C   192 13 - 12(  1.62%)  66 - 67( -1.12%)  83 - 80(  3.26%) 157 -
> > 165( -5.25%)
> > > C   256 21 - 22( -5.78%)  94 - 93(  1.65%) 100 - 99(  1.52%) 202 -
> > 204( -1.34%)
> > > After
> > > C    64  5 -  5(  0.05%)  31 - 31(  0.16%)  40 - 39(  1.37%)  77 -
> > 77( -0.93%)
> > > C   128  9 -  9( -0.69%)  50 - 50( -0.02%)  61 - 60(  2.71%) 118 -
> > 122( -3.15%)
> > > C   192 12 - 11(  7.68%)  68 - 66(  3.64%)  79 - 81( -2.85%) 164 -
> > 164( -0.12%)
> > > C   256 21 - 20(  1.80%)  90 - 87(  2.70%)  96 - 97( -0.99%) 202 -
> > 201(  0.09%)
> > >
> > > >
> > > > > [1]:
> > > >
> > https://patchwork.dpdk.org/project/dpdk/patch/20260521185631.116046-1-
> > > > > [email protected]/

Reply via email to