Netdev List
 help / color / mirror / Atom feed
From: Paolo Abeni <pabeni@redhat.com>
To: "Matthieu Baerts (NGI0)" <matttbe@kernel.org>,
	Mat Martineau <martineau@kernel.org>,
	Geliang Tang <geliang@kernel.org>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Simon Horman <horms@kernel.org>
Cc: netdev@vger.kernel.org, mptcp@lists.linux.dev,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH net-next v2 3/5] mptcp: explicitly drop over memory limits
Date: Mon, 3 Aug 2026 15:33:44 +0200	[thread overview]
Message-ID: <e6f39bb0-7c32-4429-a1d8-388783eafb5b@redhat.com> (raw)
In-Reply-To: <20260731-net-next-mptcp-oooq-pruning-v2-3-24838164fa21@kernel.org>

On 7/31/26 4:24 PM, Matthieu Baerts (NGI0) wrote:
> From: Paolo Abeni <pabeni@redhat.com>
> 
> Currently the enforcement of the rcvbuf constraint is implemented
> when moving the skbs into the msk receive or OoO queue, keeping the
> incoming skbs in the subflow queue when over limits.
> 
> Under significant memory pressure the above can cause permanent data
> transfer stalls, as the skb needed to make forward progress can be
> stuck in a subflow queue.
> 
> Over memory limits, drop the incoming skb, relying on MPTCP-level
> retransmissions.
> 
> Note that fallback socket must perform the limit before the skb reaches
> the subflow-level queue, as dropping an in-sequence already acked skb
> would break the stream.
> 
> This is not a complete fix for the stall issue, as the drop strategy
> needs refinements that will come in the next patches.
> 
> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
> Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
> [ Fix typo, comment, and bump LINUX_MIB_TCPRCVQDROP ]
> Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
> ---
> v2:
>  - mib: typo: "constrains" -> "constraints".
>  - mptcp_over_limit: more than 0-win: retrans, dup or old acks.
>  - mptcp_over_limit: bump LINUX_MIB_TCPRCVQDROP.
>  - Note: Sashiko might point to a possible forward-allocated memory
>    leak: this is a temp leak, and releasing additionally allocated fwd
>    memory in the error path will be fix in a patch for -net.
> ---
>  net/mptcp/mib.c      |  2 ++
>  net/mptcp/mib.h      |  2 ++
>  net/mptcp/options.c  | 32 +++++++++++++++++++++++++++++---
>  net/mptcp/protocol.c | 31 +++++++++++++++++++++++--------
>  4 files changed, 56 insertions(+), 11 deletions(-)
> 
> diff --git a/net/mptcp/mib.c b/net/mptcp/mib.c
> index f23fda0c55a7..ef65e2df709f 100644
> --- a/net/mptcp/mib.c
> +++ b/net/mptcp/mib.c
> @@ -85,6 +85,8 @@ static const struct snmp_mib mptcp_snmp_list[] = {
>  	SNMP_MIB_ITEM("SimultConnectFallback", MPTCP_MIB_SIMULTCONNFALLBACK),
>  	SNMP_MIB_ITEM("FallbackFailed", MPTCP_MIB_FALLBACKFAILED),
>  	SNMP_MIB_ITEM("WinProbe", MPTCP_MIB_WINPROBE),
> +	SNMP_MIB_ITEM("BacklogDrop", MPTCP_MIB_BACKLOGDROP),
> +	SNMP_MIB_ITEM("RcvPruned", MPTCP_MIB_RCVPRUNED),
>  };
>  
>  /* mptcp_mib_alloc - allocate percpu mib counters
> diff --git a/net/mptcp/mib.h b/net/mptcp/mib.h
> index 812218b5ed2b..9271205f682e 100644
> --- a/net/mptcp/mib.h
> +++ b/net/mptcp/mib.h
> @@ -88,6 +88,8 @@ enum linux_mptcp_mib_field {
>  	MPTCP_MIB_SIMULTCONNFALLBACK,	/* Simultaneous connect */
>  	MPTCP_MIB_FALLBACKFAILED,	/* Can't fallback due to msk status */
>  	MPTCP_MIB_WINPROBE,		/* MPTCP-level zero window probe */
> +	MPTCP_MIB_BACKLOGDROP,		/* Backlog over memory limit */
> +	MPTCP_MIB_RCVPRUNED,		/* Dropped due to memory constraints */
>  	__MPTCP_MIB_MAX
>  };
>  
> diff --git a/net/mptcp/options.c b/net/mptcp/options.c
> index c664023d37ba..5642277c8b3d 100644
> --- a/net/mptcp/options.c
> +++ b/net/mptcp/options.c
> @@ -1127,8 +1127,34 @@ static bool add_addr_hmac_valid(struct mptcp_sock *msk,
>  	return hmac == mp_opt->ahmac;
>  }
>  
> -/* Return false in case of error (or subflow has been reset),
> - * else return true.
> +static bool mptcp_over_limit(struct sock *sk, struct sock *ssk,
> +			     const struct sk_buff *skb)
> +{
> +	struct mptcp_sock *msk = mptcp_sk(sk);
> +	u64 mem = sk_rmem_alloc_get(sk);
> +
> +	mem += READ_ONCE(msk->backlog_len);
> +	if (likely(mem <= READ_ONCE(sk->sk_rcvbuf)))
> +		return false;

Clashiko noted this is a bit pessimistic/too strict. I *think* it can be
relaxed a bit, but I'm not sure if such option would be actually better.
Unfortunately this is inherently race. I will give a shot.

> +	/* Avoid silently dropping pure acks, fin or already-acked segments. */
> +	if (TCP_SKB_CB(skb)->seq == TCP_SKB_CB(skb)->end_seq ||
> +	    TCP_SKB_CB(skb)->tcp_flags & TCPHDR_FIN ||
> +	    !after(TCP_SKB_CB(skb)->end_seq, tcp_sk(ssk)->rcv_nxt))
> +		return false;
> +
> +	/* Dropped due to memory constraints, schedule an ack. */
> +	inet_csk(ssk)->icsk_ack.pending |= ICSK_ACK_NOMEM | ICSK_ACK_NOW;
> +	inet_csk_schedule_ack(ssk);
> +
> +	/* Plain TCP (fallback) and skb is dropped before the TCP recv queue. */
> +	NET_INC_STATS(sock_net(sk), LINUX_MIB_TCPRCVQDROP);

Clashiko note this account is not correct in case of RST packet with
payload; this could be a follow-up, but given the other feedback I'll
give it a shot.

I took the liberty to ignore minor feedback WRT comment accuracy.

/P


  reply	other threads:[~2026-08-03 13:33 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-31 14:24 [PATCH net-next v2 0/5] mptcp: out-of-order queue pruning Matthieu Baerts (NGI0)
2026-07-31 14:24 ` [PATCH net-next v2 1/5] mptcp: move the retrans loop to a separate helper Matthieu Baerts (NGI0)
2026-07-31 14:24 ` [PATCH net-next v2 2/5] mptcp: let the retrans scheduler do its job Matthieu Baerts (NGI0)
2026-08-03 13:16   ` Paolo Abeni
2026-07-31 14:24 ` [PATCH net-next v2 3/5] mptcp: explicitly drop over memory limits Matthieu Baerts (NGI0)
2026-08-03 13:33   ` Paolo Abeni [this message]
2026-07-31 14:24 ` [PATCH net-next v2 4/5] mptcp: enforce hard limit on backlog flushing Matthieu Baerts (NGI0)
2026-07-31 14:24 ` [PATCH net-next v2 5/5] mptcp: implemented OoO queue pruning Matthieu Baerts (NGI0)
2026-08-03 13:42   ` Paolo Abeni

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=e6f39bb0-7c32-4429-a1d8-388783eafb5b@redhat.com \
    --to=pabeni@redhat.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=geliang@kernel.org \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=martineau@kernel.org \
    --cc=matttbe@kernel.org \
    --cc=mptcp@lists.linux.dev \
    --cc=netdev@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox