The x86 special case exists so that both arms of the weak_barriers test compile to the same thing and the compiler drops the branch. That only holds while the barrier is a plain compiler barrier, which rte_atomic_thread_fence(rte_memory_order_acquire) is not: it grows every function in virtio_rxtx.o and brings the branch back.
Use rte_io_rmb(), which is what the else arm already uses and is a compiler barrier on x86, keeping the generated code unchanged. The ordering is unchanged too: x86 does not reorder loads with loads. Signed-off-by: Stephen Hemminger <[email protected]> --- drivers/net/virtio/virtqueue.h | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/drivers/net/virtio/virtqueue.h b/drivers/net/virtio/virtqueue.h index 1f0e6ae77e..ee21d43068 100644 --- a/drivers/net/virtio/virtqueue.h +++ b/drivers/net/virtio/virtqueue.h @@ -445,15 +445,15 @@ virtqueue_nused(const struct virtqueue *vq) if (vq->hw->weak_barriers) { /** - * x86 prefers to using rte_smp_rmb over rte_atomic_load_explicit as it + * x86 prefers rte_io_rmb over rte_atomic_load_explicit as it * reports a slightly better perf, which comes from the saved - * branch by the compiler. - * The if and else branches are identical with the smp and io - * barriers both defined as compiler barriers on x86. + * branch by the compiler: on x86 this makes the two branches + * identical, since rte_io_rmb() is only a compiler barrier and + * loads are not reordered with other loads. */ #ifdef RTE_ARCH_X86_64 idx = vq->vq_split.ring.used->idx; - rte_smp_rmb(); + rte_io_rmb(); #else idx = rte_atomic_load_explicit(&(vq)->vq_split.ring.used->idx, rte_memory_order_acquire); -- 2.53.0

