Linux Kernel Selftest development
 help / color / mirror / Atom feed
* [PATCH net-next 0/4] mptcp: support MSG_ERRQUEUE
@ 2026-09-18 18:38 Matthieu Baerts (NGI0)
  2026-09-18 18:38 ` [PATCH net-next 4/4] selftests: mptcp: cover IP_RECVERR sockopt propagation Matthieu Baerts (NGI0)
  0 siblings, 1 reply; 3+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-18 18:38 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0),
	David Carlier, Shuah Khan, linux-kselftest

This series lets MPTCP applications use poll(EPOLLERR) and
recvmsg(MSG_ERRQUEUE) on the MPTCP socket to drain TX timestamps
through the standard inet ABI, the same way they would on a plain TCP
socket. ICMP-derived errors stay on the subflow queue: the legacy
RECVERR ABI cannot convey their per-subflow peer identity, and they
are intended for a future MPTCP_RECERR channel.

- Patch 1 splices subflow err-skbs onto the MPTCP's sk_error_queue at
  error-report time. All forwarded events go through sock_queue_err_skb,
  which re-homes skb->sk onto the MPTCP and charges sk_rmem_alloc, so
  the MPTCP's error queue stays bounded by sk_rcvbuf and is dropped under
  rmem pressure, matching tcp's tx-timestamp path and ip_icmp_error() /
  ipv6_icmp_error(). mptcp_recvmsg(MSG_ERRQUEUE) forwards directly to
  inet_recv_error(), and mptcp_poll() advertises EPOLLERR purely on the
  MPTCP's sk_err / sk_error_queue, matching tcp_poll().

- Patch 2 factors the existing inet_flags subflow-propagation hard-coded
  list into a mask, so the next patch can extend it without churn.

- Patch 3 makes IP_RECVERR / IPV6_RECVERR (and the RFC4884 variants)
  propagate to the subflows. The MPTCP stores the bit so MPTCP-aware
  helpers can branch on it.

- Patch 4 is a selftest covering the propagation path.

Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
David Carlier (4):
      mptcp: support MSG_ERRQUEUE on the parent socket
      mptcp: sockopt: factor inet_flags propagation into a mask
      mptcp: propagate RECVERR sockopts to subflows
      selftests: mptcp: cover IP_RECVERR sockopt propagation

 net/mptcp/protocol.c                              |  55 ++++++--
 net/mptcp/sockopt.c                               | 155 ++++++++++++++++++----
 tools/testing/selftests/net/mptcp/mptcp_sockopt.c |  70 ++++++++++
 3 files changed, 246 insertions(+), 34 deletions(-)
---
base-commit: 4bb9710c6a68d35207f123aef55dcd50e7195ec5
change-id: 20260918-net-next-mptcp-msg_errqueue-0e2049a30061

Best regards,
--  
Matthieu Baerts (NGI0) <matttbe@kernel.org>


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH net-next 4/4] selftests: mptcp: cover IP_RECVERR sockopt propagation
  2026-09-18 18:38 [PATCH net-next 0/4] mptcp: support MSG_ERRQUEUE Matthieu Baerts (NGI0)
@ 2026-09-18 18:38 ` Matthieu Baerts (NGI0)
  2026-09-21 18:39   ` netdev-bot+sashiko
  0 siblings, 1 reply; 3+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-18 18:38 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0),
	David Carlier, Shuah Khan, linux-kselftest

From: David Carlier <devnexen@gmail.com>

Exercise setsockopt/getsockopt of IP_RECVERR and IPV6_RECVERR on the
MPTCP parent socket, including the empty-errqueue EAGAIN contract on
MSG_ERRQUEUE|MSG_DONTWAIT.

End-to-end errqueue delivery (ICMP, TX timestamps, zerocopy) depends on
subflow-side producers that are out of scope for this series and will be
covered by follow-up work.

Assisted-by: Codex:gpt-5
Signed-off-by: David Carlier <devnexen@gmail.com>
Acked-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
To: Shuah Khan <shuah@kernel.org>
Cc: linux-kselftest@vger.kernel.org
---
 tools/testing/selftests/net/mptcp/mptcp_sockopt.c | 70 +++++++++++++++++++++++
 1 file changed, 70 insertions(+)

diff --git a/tools/testing/selftests/net/mptcp/mptcp_sockopt.c b/tools/testing/selftests/net/mptcp/mptcp_sockopt.c
index b6e58d936ebe..d68515b7903b 100644
--- a/tools/testing/selftests/net/mptcp/mptcp_sockopt.c
+++ b/tools/testing/selftests/net/mptcp/mptcp_sockopt.c
@@ -183,6 +183,13 @@ static void xgetaddrinfo(const char *node, const char *service,
 	}
 }
 
+static bool expect_all_features(void)
+{
+	char *env = getenv("SELFTESTS_MPTCP_LIB_EXPECT_ALL_FEATURES");
+
+	return env && strcmp(env, "1") == 0;
+}
+
 static int sock_listen_mptcp(const char * const listenaddr,
 			     const char * const port)
 {
@@ -769,6 +776,68 @@ static void test_ip_tos_sockopt(int fd)
 		xerror("expect socklen_t == -1");
 }
 
+static void test_ip_recverr_sockopt(int fd)
+{
+	struct iovec iov = {
+		.iov_base = &(char){ 0 },
+		.iov_len = 1,
+	};
+	struct msghdr msg = {
+		.msg_iov = &iov,
+		.msg_iovlen = 1,
+	};
+	int one = 1, zero = 0, val = -1;
+	socklen_t s = sizeof(val);
+	int level, optname, r;
+
+	switch (pf) {
+	case AF_INET:
+		level = SOL_IP;
+		optname = IP_RECVERR;
+		break;
+	case AF_INET6:
+		level = SOL_IPV6;
+		optname = IPV6_RECVERR;
+		break;
+	default:
+		xerror("Unknown pf %d\n", pf);
+	}
+
+	r = setsockopt(fd, level, optname, &one, sizeof(one));
+	if (r) {
+		/* For older kernels not supporting IP(V6)_RECVERR yet */
+		if (errno == EOPNOTSUPP && !expect_all_features()) {
+			fprintf(stderr, "IP(V6)_RECVERR not supported, SKIP\n");
+			return;
+		}
+
+		die_perror("setsockopt IP(V6)_RECVERR on");
+	}
+
+	r = getsockopt(fd, level, optname, &val, &s);
+	if (r)
+		die_perror("getsockopt IP(V6)_RECVERR on");
+	if (s != sizeof(val) || val != one)
+		xerror("IP(V6)_RECVERR on mismatch val=%d len=%u", val, s);
+
+	r = recvmsg(fd, &msg, MSG_ERRQUEUE | MSG_DONTWAIT);
+	if (r != -1 || errno != EAGAIN)
+		xerror("expected empty errqueue to return EAGAIN, r=%d err=%d",
+		       r, errno);
+
+	r = setsockopt(fd, level, optname, &zero, sizeof(zero));
+	if (r)
+		die_perror("setsockopt IP(V6)_RECVERR off");
+
+	val = -1;
+	s = sizeof(val);
+	r = getsockopt(fd, level, optname, &val, &s);
+	if (r)
+		die_perror("getsockopt IP(V6)_RECVERR off");
+	if (s != sizeof(val) || val != zero)
+		xerror("IP(V6)_RECVERR off mismatch val=%d len=%u", val, s);
+}
+
 static int client(int pipefd)
 {
 	int fd = -1;
@@ -787,6 +856,7 @@ static int client(int pipefd)
 	}
 
 	test_ip_tos_sockopt(fd);
+	test_ip_recverr_sockopt(fd);
 
 	connect_one_server(fd, pipefd);
 

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH net-next 4/4] selftests: mptcp: cover IP_RECVERR sockopt propagation
  2026-09-18 18:38 ` [PATCH net-next 4/4] selftests: mptcp: cover IP_RECVERR sockopt propagation Matthieu Baerts (NGI0)
@ 2026-09-21 18:39   ` netdev-bot+sashiko
  0 siblings, 0 replies; 3+ messages in thread
From: netdev-bot+sashiko @ 2026-09-21 18:39 UTC (permalink / raw)
  To: matttbe
  Cc: martineau, geliang, davem, edumazet, kuba, pabeni, horms, netdev,
	mptcp, linux-kernel, devnexen, shuah, linux-kselftest

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The commit subject claims the patch covers "IP_RECVERR sockopt
  propagation", but the new selftest `test_ip_recverr_sockopt()`…
- [Low] The "empty-errqueue EAGAIN contract" assertion added at
  tools/testing/selftests/net/mptcp/mptcp_sockopt.c:823-826 (`r =…

--

selftests: mptcp: cover IP_RECVERR sockopt propagation

This adds test_ip_recverr_sockopt() to the mptcp_sockopt selftest, which
sets and reads back IP_RECVERR / IPV6_RECVERR on the MPTCP parent socket
and checks that recvmsg() with MSG_ERRQUEUE|MSG_DONTWAIT on an empty
error queue returns EAGAIN.

> diff --git a/tools/testing/selftests/net/mptcp/mptcp_sockopt.c b/tools/testing/selftests/net/mptcp/mptcp_sockopt.c
> index b6e58d936ebe..d68515b7903b 100644
> --- a/tools/testing/selftests/net/mptcp/mptcp_sockopt.c
> +++ b/tools/testing/selftests/net/mptcp/mptcp_sockopt.c

[ ... ]

> @@ -769,6 +776,68 @@ static void test_ip_tos_sockopt(int fd)
>  		xerror("expect socklen_t == -1");
>  }
>  
> +static void test_ip_recverr_sockopt(int fd)
> +{

[Severity: Medium]
The subject says "cover IP_RECVERR sockopt propagation", but does this
test observe any propagation?  Everything here operates on the parent fd
only, and no subflow state is ever queried.

The MPTCP-specific work happens in mptcp_setsockopt_recverr() ->
mptcp_setsockopt_all_sf(), including the skip for IPv6 options on v4
subflows:

net/mptcp/sockopt.c:mptcp_setsockopt_all_sf() {
	...
	if (level == SOL_IPV6 && ssk->sk_family != AF_INET6)
		continue;
	...
}

and in sync_socket_options() for subflows that join later:

net/mptcp/sockopt.c:sync_socket_options() {
	...
	src = READ_ONCE(inet_sk(sk)->inet_flags);
	if (ssk->sk_family != AF_INET6)
		mask &= ~(BIT(INET_FLAGS_RECVERR6) |
			BIT(INET_FLAGS_RECVERR6_RFC4884));
	...
}

The read-back path, however, only reports the parent msk's own flag bit:

net/mptcp/sockopt.c:mptcp_getsockopt_v4() {
	...
	case IP_RECVERR:
		return mptcp_put_int_option(msk, optval, optlen,
				inet_test_bit(RECVERR, sk));
	...
}

with mptcp_getsockopt_v6() doing the same for inet6_test_bit(RECVERR6, sk).

So if the subflow forwarding step and the sync-on-join masking were both
deleted, would this test still pass?  Would it be more accurate to say the
test covers the parent-socket set/get round trip, and to note that subflow
propagation is not asserted?  IP_RECVERR_RFC4884 / IPV6_RECVERR_RFC4884 go
through the same kernel paths and are also not touched here.

> +	struct iovec iov = {
> +		.iov_base = &(char){ 0 },
> +		.iov_len = 1,
> +	};
> +	struct msghdr msg = {
> +		.msg_iov = &iov,
> +		.msg_iovlen = 1,
> +	};
> +	int one = 1, zero = 0, val = -1;
> +	socklen_t s = sizeof(val);
> +	int level, optname, r;
> +
> +	switch (pf) {
> +	case AF_INET:
> +		level = SOL_IP;
> +		optname = IP_RECVERR;
> +		break;
> +	case AF_INET6:
> +		level = SOL_IPV6;
> +		optname = IPV6_RECVERR;
> +		break;
> +	default:
> +		xerror("Unknown pf %d\n", pf);
> +	}
> +
> +	r = setsockopt(fd, level, optname, &one, sizeof(one));
> +	if (r) {
> +		/* For older kernels not supporting IP(V6)_RECVERR yet */
> +		if (errno == EOPNOTSUPP && !expect_all_features()) {
> +			fprintf(stderr, "IP(V6)_RECVERR not supported, SKIP\n");
> +			return;
> +		}
> +
> +		die_perror("setsockopt IP(V6)_RECVERR on");
> +	}
> +
> +	r = getsockopt(fd, level, optname, &val, &s);
> +	if (r)
> +		die_perror("getsockopt IP(V6)_RECVERR on");
> +	if (s != sizeof(val) || val != one)
> +		xerror("IP(V6)_RECVERR on mismatch val=%d len=%u", val, s);
> +
> +	r = recvmsg(fd, &msg, MSG_ERRQUEUE | MSG_DONTWAIT);
> +	if (r != -1 || errno != EAGAIN)
> +		xerror("expected empty errqueue to return EAGAIN, r=%d err=%d",
> +		       r, errno);

[Severity: Low]
Can this check tell the errqueue path apart from the ordinary non-blocking
empty receive queue path?

The branch it aims at is:

net/mptcp/protocol.c:mptcp_recvmsg() {
	...
	if (unlikely(flags & MSG_ERRQUEUE))
		return inet_recv_error(sk, msg, len);
	...
}

If that branch were gone, the normal receive loop in the same function
would still return -EAGAIN for a connected socket with nothing queued:

net/mptcp/protocol.c:mptcp_recvmsg() {
	...
		if (!timeo) {
			copied = -EAGAIN;
			break;
		}
	...
}

Since test_ip_recverr_sockopt() runs before connect_one_server() moves any
payload, the data queue is empty by construction and MSG_DONTWAIT is set,
so both paths return -1/EAGAIN.  Would something like checking that a
queued TX timestamp is actually returned, or that the one byte data buffer
is not consumed, make the assertion discriminating?

> +
> +	r = setsockopt(fd, level, optname, &zero, sizeof(zero));
> +	if (r)
> +		die_perror("setsockopt IP(V6)_RECVERR off");
> +
> +	val = -1;
> +	s = sizeof(val);
> +	r = getsockopt(fd, level, optname, &val, &s);
> +	if (r)
> +		die_perror("getsockopt IP(V6)_RECVERR off");
> +	if (s != sizeof(val) || val != zero)
> +		xerror("IP(V6)_RECVERR off mismatch val=%d len=%u", val, s);
> +}
> +
>  static int client(int pipefd)
>  {
>  	int fd = -1;
> @@ -787,6 +856,7 @@ static int client(int pipefd)
>  	}
>  
>  	test_ip_tos_sockopt(fd);
> +	test_ip_recverr_sockopt(fd);
>  
>  	connect_one_server(fd, pipefd);

Since the option is switched back off here before connect_one_server()
runs, no later part of the test can observe the option state on subflows
either.  Was leaving RECVERR enabled for the rest of the run considered?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260918-net-next-mptcp-msg_errqueue-v1-0-dd77e1738248%40kernel.org

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-21 18:39 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18 18:38 [PATCH net-next 0/4] mptcp: support MSG_ERRQUEUE Matthieu Baerts (NGI0)
2026-09-18 18:38 ` [PATCH net-next 4/4] selftests: mptcp: cover IP_RECVERR sockopt propagation Matthieu Baerts (NGI0)
2026-09-21 18:39   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox