From: Neil Ramaswamy <[email protected]> Partial undo can clear the TCPCB_LOST flag on segments already removed from RACK's list, which prevents subsequent RACK loss detection, leading to segments only being retransmitted after the RTO. Restoring them to the RACK list as part of partial undo makes sure we can properly reconsider them for fast retransmission in the future.
To do this, we first sort (by transmission time) the segments whose TCPCB_LOST flag is being cleared and then linearly insert those back into the RACK list, which is sorted by transmission time. My investigation started from seeing repeated TCP stalls in prod and the mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to 96 KiB. It also seems like others have seen similar symptoms before [1]. [1] https://lore.kernel.org/netdev/[email protected]/ Changes in v2: - Added production symptoms - Removed the redundant TCPCB_LOST comparison - Fixed lint/long lines Patch 2 is unchanged. v1: https://lore.kernel.org/netdev/[email protected]/ Neil Ramaswamy (2): tcp: restore RACK list membership when undoing loss selftests: net: packetdrill: test RACK after partial undo net/ipv4/tcp_input.c | 35 ++++++++++++ ...tcp_partial_undo-restores-to-rack-list.pkt | 54 +++++++++++++++++++ 2 files changed, 89 insertions(+) create mode 100644 tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt base-commit: 11536ee3d3e0b1bd35b6f3f8df55a6053eb0c71d -- 2.55.0

