BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path. When an skb exceeds the GSO limits of the port it is
sent to, it loses its GSO feature mask and is segmented into individual MSS
sized packets, which that port then sends without TSO.
This series cuts such a TCP GSO skb into GSO skbs which fit the device
limits instead, so the rest of the path keeps using TSO. On a
veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP enabled on the
veth endpoints and left off in the guest, a single iperf3 TCP flow, six
alternating runs per state (-t 15 -O 5, fixed CPU affinity and port tuple):
protocol no BIG TCP mixed, no reseg mixed, resegmented
TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
The middle column is the same tree with the re-segmentation disabled. A
BIG TCP hop which feeds a 64 KiB hop loses 69% of the throughput of a path
which never enables BIG TCP; re-segmentation recovers it.
1/6 moves the GSO size limit lookup into dev.c and makes it follow the
packet's L3 protocol, so where the tag sits does not decide which limit
applies. 2/6 is the preparation which lets the limit tests be skipped for
one caller, and carries no functional change. 3/6 lets the GSO engine
group several MSS into one output skb, and 4/6 clears the whole GSO state
of the last output of such a group.
The new path is taken only when the skb is an unencapsulated TCP GSO skb
without a frag_list, it exceeds gso_max_size or gso_max_segs, and the
device offloads that GSO type. Everything else keeps today's segmentation.
The output obeys the GSO feature and limit contract the device already
advertises, so this needs no new UAPI, no device state and no driver
change, and it applies automatically. max_segs only says how the output is
grouped; an over-limit skb pays one extra ndo_features_check() in exchange
for staying a GSO skb.
Alternatives considered:
- The caller could set skb_shinfo(skb)->gso_size to ~64K and adjust the
gso bits in the shared info afterwards, which would need
skb_unclone(), a repeat of the grouping logic skb_segment() already
has, and a recomputed IPv4 ID for the DF=0 case.
Patch layout:
[1/6] the GSO size limit follows the packet's L3 protocol
[2/6] factor the device limit check out of gso_features_check()
[3/6] let the GSO engine group several MSS into one output skb
[4/6] clear the whole GSO state of the last output segment
[5/6] re-segment oversized TCP GSO skbs in the TX path
[6/6] KUnit coverage for re-segmentation
---
v5:
- patch 3: make max_segs a u8, skb_gso_cb has one byte left
- patch 3: only warn about the input segment count without max_segs
- patch 4: new, clear the whole GSO state of the last output segment
- patch 6: add a plain-tail case, drop the TCP path tests
v4: https://lore.kernel.org/[email protected]/
v3: https://lore.kernel.org/[email protected]/
v2: https://lore.kernel.org/[email protected]/
v1: https://lore.kernel.org/[email protected]/
Wang Zhan (6):
net: core: use the packet's L3 protocol for the GSO size limit
net: core: factor out the GSO device limit check
net: gso: support re-segmentation of TCP GSO skbs
net: gso: clear the GSO state of the last output segment
net: core: re-segment oversized TCP GSO skbs
net: net_test: add tests for TCP re-segmentation
drivers/net/tap.c | 3 +-
include/linux/netdevice.h | 18 ++++----
include/net/gso.h | 7 ++-
include/net/udp.h | 2 +-
net/core/dev.c | 92 +++++++++++++++++++++++++++++++++-----
net/core/gso.c | 6 ++-
net/core/net_test.c | 87 +++++++++++++++++++++++++++++++++++
net/core/skbuff.c | 12 +++--
net/openvswitch/datapath.c | 2 +-
9 files changed, 198 insertions(+), 31 deletions(-)
base-commit: 8df0638138d3e0344fd1fb36cf2d1ca1cf5028f0
--
2.47.3
_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev