Netdev List
 help / color / mirror / Atom feed
From: nramaswamy@openai.com
To: netdev@vger.kernel.org
Cc: Neil Ramaswamy <nramaswamy@openai.com>
Subject: [PATCH net 0/2] tcp: preserve RACK tracking across partial undo
Date: Fri, 25 Sep 2026 17:25:21 -0700	[thread overview]
Message-ID: <20260926002520.42955-4-nramaswamy@openai.com> (raw)

From: Neil Ramaswamy <nramaswamy@openai.com>

I've been investigating a bug in TCP where RACK loses track of segments
after partial undo happens. At a high-level, RACK can react to a SACK
by scanning its sorted transmission queue for segments that exceed its
RACK timeout, mark those segments as lost, and remove them from its
sorted queue. However, receiving an ACK with a TSecr less than the
first retansmission timestamp of the recovery episode can trigger
partial undo, which removes the lost flag from segments not yet
cumulatively ACK'd but does not ensure that they are in (or are added to)
the RACK transmission queue.

The symptom that I observe is that a long tail of hole segments can
"escape" RACK and then only get sent out by the retransmission timer,
which can take a long time and even be serial for many lost segments.

A conceptual example below illustrates this situation; I'm intentionally
excluding TSecrs/byte ranges/etc. when not relevant.

   TCP A                            TCP B
   
                             Assume TS.Recent = 200 to start.

    A, TSval = 210 --- <delayed>
    B, TSval = 211 --- <delayed>
    C, TSval = 212 ----------------> arrives
    D, TSval = 213 --- <lost>
    E, TSval = 214 ----------------> arrives
    
                   <---------------- SACK C + E, TSecr = 200
    A', TSval = 216 --- <delayed>
    B', TSval = 217 --------
    D marked lost           |
                            B' ----> arrives, ACK delayed
                    
                    original A ----> A fills leading gap,
                                     TS.Recent = 210
                                     
                   <---------------- ACK through C, SACK E, TSecr = 210
    retrans_out = 0
    210 < 216 permits partial undo
    D lost flag is cleared
    D remains absent from RACK list
    

A bit of commentary on this diagram:

1. The timestamps I'm using are for the purposes of showing the
   partial undo comparison; these aren't an exact schedule with the
   timeouts RACK would use. See the packetdrill for that.
2. A and B are also marked as lost and removed from the RACK list.
   However, PRR only allows A and B to be retransmitted. D
   being marked missing consists of being marked as lost and,
   critically, not being added to the RACK list.
3. I delayed the B' ACK because if it is sent back, this may let the
   sender retransmit D before partial undo happens.
4. At the end of our diagram, D can only be rescued with the
   retransmission timer.

There is also another more catastrophic situation in which partial undo
might happen when a stale TS.Recent is echo replied [1] when handling
out-of-order ACKs, and I've seen this cause the RTO to jump to over 100
seconds. But this combination should not happen if [1] is merged.

I see two options for fixing this, and I provided the first as a patch:

  1. When partial undo runs, we make sure that a segment whose lost flag
     is cleared is added back to the RACK list.
  2. RACK does not remove from the RACK transmission list until a segment
     is acknowledged; I think this is a bad approach because you would
     end up scanning already-marked-as-lost segments every time you do
     loss detection.

Without a fix as such, packet "D" in my included packetdrill takes around
400ms to be retransmitted via RTO. With the attached patch, D is
retransmitted within 50ms of the partial ACK, without an RTO.


[1] https://lore.kernel.org/all/20260921222609.50824-4-jeffjo@openai.com/

Neil Ramaswamy (2):
  tcp: restore RACK list membership when undoing loss
  selftests: net: packetdrill: test RACK after partial undo

 net/ipv4/tcp_input.c                          | 33 ++++++++++++
 ...tcp_partial_undo-restores-to-rack-list.pkt | 54 +++++++++++++++++++
 2 files changed, 87 insertions(+)
 create mode 100644 tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt

base-commit: 11536ee3d3e0b1bd35b6f3f8df55a6053eb0c71d

             reply	other threads:[~2026-09-26  0:25 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-26  0:25 nramaswamy [this message]
2026-09-26  0:25 ` [PATCH net 1/2] tcp: restore RACK list membership when undoing loss nramaswamy
2026-09-30  2:13   ` Jiayuan Chen
2026-10-01  0:30     ` Kuniyuki Iwashima
2026-10-01 23:07   ` Jakub Kicinski
2026-10-01 23:33     ` nramaswamy
2026-09-26  0:25 ` [PATCH net 2/2] selftests: net: packetdrill: test RACK after partial undo nramaswamy
2026-09-28 21:53 ` [PATCH net 0/2] tcp: preserve RACK tracking across " nramaswamy

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260926002520.42955-4-nramaswamy@openai.com \
    --to=nramaswamy@openai.com \
    --cc=netdev@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox