> From: Konstantin Ananyev [mailto:[email protected]] > Sent: Thursday, 6 August 2026 09.58 > > > > > > > > > > > + /* Common way for small copy size of 64-byte > blocks. > > > > > > > Unlikely, so > > > > > > > > > constant size only */ > > > > > > > > > > + if (__rte_constant(n) && (n & 63) == 0 && n <= > > > > > > > > > RTE_MEMCPY_BLOCK_64_MAX) { > > > > > > > > > > + void *ret = dst; > > > > > > > > > > + > > > > > > > > > > > > > > > > > > Maybe just let compiler decide, it will generate vector > > > > > > > instructions in > > > > > > > > > most cases. > > > > > > > > > > > > > > > > > > if (__rte_constant(n)) > > > > > > > > > return mempcpy(dst, src, n); > > > > > > > > > > > > > > > > Maybe in most, but not in all: > > > > > > > > https://godbolt.org/z/KvdKqT5rY > > > > > > > > > > > > > > With '-mavx' or '-mavx512f' it looks like it does for your > > > sample > > > > > code. > > > > > > > > > > > > It also does with -msse4.2 when SZ is reduced to 256 bytes. > > > > > > Clang switches to inline when SZ is reduced to 128 bytes. > > > > > > > > > > > > It seems the compiler has a threshold for when to inline and > when > > > to > > > > > call the C > > > > > > library's memcpy subroutine. > > > > > > The threshold depends on both copy size and vector register > size. > > > > > > And it is compiler dependent. > > > > > > > > > > I think there are compiler options to specify desired threshold > > > values. > > > > > Let say for gcc there is ' -mmemcpy-strategy=strategy'. > > > > > For that example in that particular case > > > > > -mmemcpy-strategy=vector_loop:512:align,loop:-1:align > > > > > generates sse loads/stores. > > > > > Might be we can exploit it somehow? > > > > > > > > That could give us higher granularity/control over memcpy for > > > individual > > > > memcpy instances; might be useful for hot code paths where we > have > > > more > > > > knowledge about the copy operation than the compiler can infer. > > > > However, pragmas are discouraged in DPDK, and this looks like a > very > > > similar > > > > path. > > > > > > Well, right now rte_memcpy.h is 700+ lines and keeps growing. > > > Considering that probably pragmas are not that bad. > > > Of course, pragmas have their own issues and it is hard to ensure > that > > > they will produce same code between different compilers/versions, > etc. > > > > > > > > I am not really happy that our home-brewed memcpy code-block > keeps > > > > > growing, > > > > > while we keep talking that it would be good to eliminate it > > > completely. > > > > > > > > I agree in principle. > > > > However, this rte_memcpy() optimization is for the pile/mempool > > > optimizations > > > > I'm working on, so there is a specific use case motivating the > added > > > code. > > > > > > I understand that you probably have some specific use-case in mind. > > > BTW for this optimization you mentioned above: what is the gain > with > > > these changes? > > > > IMO, the primary benefit is the much simpler (and smaller) assembly > output due > > to avoiding the address alignment check (and the resulting duplicated > code). > > I think that should be measurable too: whole binary and/or hot path > function size reduction, etc.
The size of the generated code for the copy operation is reduced to slightly less than half (only one instance of the copy operation instead of two, and the address alignment comparison is omitted). > > > I haven't measured the performance gain. > > Based on the perf gain in a previous mempool optimization patch [1], > it seems > > avoiding the address alignment check shaves ~2 cycles off the copy > operation (for > > cache-to-cache copy). > > I expect that the same gain (from avoiding the address alignment > check) applies > > here. > > Ok, then I suggest we do some measurements first, before going forward > with it. I tried memcpy_perf_autotest, but the branch predictor kicks in and eliminates the cost of the address alignment comparison. So the results are very similar. BTW, that's a general issue with our perf tests: Repeated testing doesn't show the cost of branches, because they are always eliminated by the branch predictor. ======= ================= ================= ================= ================= Size Cache to cache Cache to mem Mem to cache Mem to mem (bytes) (ticks) (ticks) (ticks) (ticks) ------- ----------------- ----------------- ----------------- ----------------- ================================= 32B aligned ================================= Before C 64 2 - 2( 9.60%) 22 - 21( 4.48%) 33 - 31( 5.15%) 57 - 56( 0.28%) C 128 4 - 4( 5.56%) 41 - 40( 3.59%) 53 - 52( 1.49%) 101 -100( 0.71%) C 192 6 - 6( 4.06%) 64 - 60( 5.61%) 68 - 69( -1.68%) 136 -138( -1.49%) C 256 10 - 9( 5.21%) 82 - 81( 1.20%) 93 - 87( 7.47%) 172 -175( -1.46%) After C 64 2 - 2( -0.71%) 22 - 21( 7.62%) 31 - 30( 4.32%) 57 - 59( -3.03%) C 128 4 - 4( -0.34%) 40 - 40( 0.60%) 50 - 50( 0.41%) 97 - 99( -1.96%) C 192 6 - 6( 5.27%) 59 - 59( -0.27%) 69 - 72( -4.42%) 139 -138( 0.90%) C 256 9 - 9( 9.14%) 77 - 77( 0.25%) 89 - 90( -0.88%) 180 -177( 1.57%) ================================== Unaligned ================================== Before C 64 5 - 5( 0.98%) 32 - 32( 0.72%) 42 - 42( 0.44%) 82 - 82( 0.51%) C 128 8 - 9(-11.26%) 51 - 50( 0.53%) 60 - 59( 0.89%) 117 -115( 1.37%) C 192 13 - 12( 1.62%) 66 - 67( -1.12%) 83 - 80( 3.26%) 157 -165( -5.25%) C 256 21 - 22( -5.78%) 94 - 93( 1.65%) 100 - 99( 1.52%) 202 -204( -1.34%) After C 64 5 - 5( 0.05%) 31 - 31( 0.16%) 40 - 39( 1.37%) 77 - 77( -0.93%) C 128 9 - 9( -0.69%) 50 - 50( -0.02%) 61 - 60( 2.71%) 118 -122( -3.15%) C 192 12 - 11( 7.68%) 68 - 66( 3.64%) 79 - 81( -2.85%) 164 -164( -0.12%) C 256 21 - 20( 1.80%) 90 - 87( 2.70%) 96 - 97( -0.99%) 202 -201( 0.09%) > > > [1]: > https://patchwork.dpdk.org/project/dpdk/patch/20260521185631.116046-1- > > [email protected]/

