> From: Konstantin Ananyev [mailto:[email protected]]
> Sent: Wednesday, 5 August 2026 07.46
> 
> > > From: Stephen Hemminger [mailto:[email protected]]
> > > Sent: Tuesday, 4 August 2026 17.52
> > >
> > > On Tue,  4 Aug 2026 14:33:04 +0000
> > > Morten Brørup <[email protected]> wrote:
> > >
> > > > +       /* Common way for small copy size of 64-byte blocks.
> Unlikely, so
> > > constant size only */
> > > > +       if (__rte_constant(n) && (n & 63) == 0 && n <=
> > > RTE_MEMCPY_BLOCK_64_MAX) {
> > > > +               void *ret = dst;
> > > > +
> > >
> > > Maybe just let compiler decide, it will generate vector
> instructions in
> > > most cases.
> > >
> > >   if (__rte_constant(n))
> > >           return mempcpy(dst, src, n);
> >
> > Maybe in most, but not in all:
> > https://godbolt.org/z/KvdKqT5rY
> 
> With '-mavx' or '-mavx512f' it looks like it does for your sample code.

It also does with -msse4.2 when SZ is reduced to 256 bytes.
Clang switches to inline when SZ is reduced to 128 bytes.

It seems the compiler has a threshold for when to inline and when to call the C 
library's memcpy subroutine.
The threshold depends on both copy size and vector register size.
And it is compiler dependent.

With rte_memcpy() it is always inline.

Reply via email to