The x86 special case exists so that both arms of the weak_barriers
test compile to the same thing and the compiler drops the branch.
That only holds while the barrier is a plain compiler barrier, which
rte_atomic_thread_fence(rte_memory_order_acquire) is not: it grows
every function in virtio_rxtx.o and brings the branch back.

Use rte_io_rmb(), which is what the else arm already uses and is a
compiler barrier on x86, keeping the generated code unchanged. The
ordering is unchanged too: x86 does not reorder loads with loads.

Signed-off-by: Stephen Hemminger <[email protected]>
---
 drivers/net/virtio/virtqueue.h | 10 +++++-----
 1 file changed, 5 insertions(+), 5 deletions(-)

diff --git a/drivers/net/virtio/virtqueue.h b/drivers/net/virtio/virtqueue.h
index 1f0e6ae77e..ee21d43068 100644
--- a/drivers/net/virtio/virtqueue.h
+++ b/drivers/net/virtio/virtqueue.h
@@ -445,15 +445,15 @@ virtqueue_nused(const struct virtqueue *vq)
 
        if (vq->hw->weak_barriers) {
        /**
-        * x86 prefers to using rte_smp_rmb over rte_atomic_load_explicit as it
+        * x86 prefers rte_io_rmb over rte_atomic_load_explicit as it
         * reports a slightly better perf, which comes from the saved
-        * branch by the compiler.
-        * The if and else branches are identical with the smp and io
-        * barriers both defined as compiler barriers on x86.
+        * branch by the compiler: on x86 this makes the two branches
+        * identical, since rte_io_rmb() is only a compiler barrier and
+        * loads are not reordered with other loads.
         */
 #ifdef RTE_ARCH_X86_64
                idx = vq->vq_split.ring.used->idx;
-               rte_smp_rmb();
+               rte_io_rmb();
 #else
                idx = rte_atomic_load_explicit(&(vq)->vq_split.ring.used->idx,
                                rte_memory_order_acquire);
-- 
2.53.0

Reply via email to