On Wed, Jul 15, 2026 at 5:26 PM Leon Hwang <[email protected]> wrote: > > On 2026/7/15 23:15, Eric Dumazet wrote: > > On Wed, Jul 15, 2026 at 4:54 PM Leon Hwang <[email protected]> wrote: > > [...] > > >> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c > >> index 61045a8886e4..4f1027173e95 100644 > >> --- a/net/ipv4/tcp_input.c > >> +++ b/net/ipv4/tcp_input.c > >> @@ -4853,6 +4853,7 @@ void tcp_done_with_error(struct sock *sk, int err) > >> /* When we get a reset we do this. */ > >> void tcp_reset(struct sock *sk, struct sk_buff *skb) > >> { > >> + const struct net *net = sock_net(sk); > >> int err; > >> > >> trace_tcp_receive_reset(sk); > >> @@ -4869,6 +4870,27 @@ void tcp_reset(struct sock *sk, struct sk_buff *skb) > >> err = ECONNREFUSED; > >> break; > >> case TCP_CLOSE_WAIT: > >> + /* RFC9293 3.10.7.4. Other States > >> + * Second, check the RST bit: > >> + * CLOSE-WAIT STATE > >> + * > >> + * If the RST bit is set, then any outstanding RECEIVEs and > >> + * SEND should receive "reset" responses. All segment > >> queues > >> + * should be flushed. Users should also receive an > >> unsolicited > >> + * general "connection reset" signal. Enter the CLOSED > >> state, > >> + * delete the TCB, and return. > >> + * > >> + * If net.ipv4.tcp_purge_receive_queue is enabled, > >> + * sk_receive_queue will be flushed too. > >> + */ > >> + if > >> (unlikely(READ_ONCE(net->ipv4.sysctl_tcp_purge_receive_queue))) { > >> + struct tcp_sock *tp = tcp_sk(sk); > >> + > >> + skb_queue_purge(&sk->sk_receive_queue); > >> + WRITE_ONCE(tp->copied_seq, tp->rcv_nxt); > >> + WRITE_ONCE(tp->urg_data, 0); > >> + sk_set_peek_off(sk, -1); > >> + } > >> err = EPIPE; > >> break; > >> case TCP_CLOSE: > >> -- > >> 2.55.0 > >> > > > > My thoughts are: > > > > out_of_order_queue has been forgotten. skbs could be there and still > > 'block devmem' > > > > WRITE_ONCE(tp->copied_seq, tp->rcv_nxt) is certainly wrong, because > > read() will return 0, instead of -1 (errno = EPIPE or ECONNRESET) > > So the application will not know a RST was received :/ > > > > I think that BSD and linux implementations have historically retained > > acknowledged, > > buffered receive data upon RST to allow applications to drain data > > already ACKed prior to the reset. > > > > Adding a narrow sysctl specifically for CLOSE_WAIT creates > > inconsistent behavior across TCP states. > > > Got it. I won't pursue this sysctl approach in the future. Thanks for > the review.
My intention was not to kill your proposal, only to start a conversation...

