From: Ihor Solodrai <ihor.solodrai@linux.dev>
To: "Alexis Lothoré (eBPF Foundation)" <alexis.lothore@bootlin.com>,
"Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"Andrii Nakryiko" <andrii@kernel.org>,
"Eduard Zingerman" <eddyz87@gmail.com>,
"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
"Martin KaFai Lau" <martin.lau@linux.dev>,
"Song Liu" <song@kernel.org>,
"Yonghong Song" <yonghong.song@linux.dev>,
"Jiri Olsa" <jolsa@kernel.org>,
"Emil Tsalapatis" <emil@etsalapatis.com>,
"Shuah Khan" <shuah@kernel.org>
Cc: ebpf@linuxfoundation.org,
Bastien Curutchet <bastien.curutchet@bootlin.com>,
Thomas Petazzoni <thomas.petazzoni@bootlin.com>,
bpf@vger.kernel.org, linux-kselftest@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH bpf v2] selftests/bpf: keep polling connection that is still in progress
Date: Thu, 6 Aug 2026 11:56:47 -0700 [thread overview]
Message-ID: <71b9f5ae-eed4-4d03-a3b9-2493b525bc34@linux.dev> (raw)
In-Reply-To: <20260803-tc_tunnel_flaky-v2-1-657b287dfa75@bootlin.com>
On 8/3/26 12:36 AM, Alexis Lothoré (eBPF Foundation) wrote:
> Some tests, like tc_tunnel or tc_edt, sporadically fail in CI with the
> following logs:
>
> (network_helpers.c:309: errno: Operation now in progress) \
> Failed to connect to server
> send_and_test_data:FAIL:connect to server unexpected error: -115
>
> This is due to SO_RCVTIMEO and SO_SNDTIMEO being set on the client
> socket (see settimeo() in client_socket()), allowing connect() to return
> an error and to set errno to EINPROGRESS instead of blocking until
> connection result is known. Increasing the timeout value for those tests
> is likely not a good solution (and it has already been done by commit
> 2790db208b44 ("selftests/bpf: Improve tc_tunnel test reliability")):
> they involve subtests that expect the connection to fail, and so
> increasing the timeout value would increase overall test execution
> duration again (not only the connection, but any socket operation).
>
> Another solution, as documented in man 2 connect, is to poll the socket
> for POLLOUT once connect has returned EINPROGRESS, and to get the actual
> connection result through getsockopt: this allows to keep the overall
> timeout values low for the general traffic, while letting a chance to
> the connection to succeed even if CI runners are loaded.
>
> When connect() returns EINPROGRESS, poll the socket for POLLOUT and
> check the connection result via getsockopt(SO_ERROR).
>
> Fixes: 99126abec5e5 ("bpf: selftests: A few improvements to network_helpers.c")
> Signed-off-by: Alexis Lothoré (eBPF Foundation) <alexis.lothore@bootlin.com>
> ---
> Changes in v2:
> - drop unneeded initialization
> - add back error message for immediate connection failure, and slightly
> reword the async connection failure error message
> - Link to v1: https://patch.msgid.link/20260710-tc_tunnel_flaky-v1-1-42aab5399a49@bootlin.com
> ---
> I manage to reproduce the issue locally by running `./test_progs -a
> tc_tunnel` in a qemu machine, while making all my CPUs busy with
> stress-ng on host side; the issue happens pretty quickly. I have not
> been able to reproduce the issue anymore with this fix.
> ---
> tools/testing/selftests/bpf/network_helpers.c | 41 ++++++++++++++++++++++++---
> 1 file changed, 37 insertions(+), 4 deletions(-)
>
> diff --git a/tools/testing/selftests/bpf/network_helpers.c b/tools/testing/selftests/bpf/network_helpers.c
> index db935a9d9fc1..52e7b72f8a77 100644
> --- a/tools/testing/selftests/bpf/network_helpers.c
> +++ b/tools/testing/selftests/bpf/network_helpers.c
> @@ -14,6 +14,7 @@
> #include <sys/types.h>
> #include <sys/un.h>
> #include <sys/eventfd.h>
> +#include <sys/poll.h>
>
> #include <linux/err.h>
> #include <linux/in.h>
> @@ -40,6 +41,8 @@
> #define IPPROTO_MPTCP 262
> #endif
>
> +#define CONNECTION_IN_PROGRESS_TIMEOUT_MS 3000
> +
> #define clean_errno() (errno == 0 ? "None" : strerror(errno))
> #define log_err(MSG, ...) ({ \
> int __save = errno; \
> @@ -294,7 +297,8 @@ int client_socket(int family, int type,
> int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t addrlen,
> const struct network_helper_opts *opts)
> {
> - int fd;
> + socklen_t errlen;
> + int fd, err;
>
> if (!opts)
> opts = &default_opts;
> @@ -305,13 +309,42 @@ int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t add
> return -1;
> }
>
> - if (connect(fd, (const struct sockaddr *)addr, addrlen)) {
> + err = connect(fd, (const struct sockaddr *)addr, addrlen);
> + if (err && errno == EINPROGRESS) {
> + struct pollfd pfd = { .fd = fd, .events = POLLOUT };
> +
> + err = poll(&pfd, 1, CONNECTION_IN_PROGRESS_TIMEOUT_MS);
Hi Alexis, thanks for the patch.
When I ran tc_* selftests in parallel it looked like they hanged.
I think what's happening is that this 3s timeout is additive, making
some tests to wait for too long. For example tc_tunnel has 1s timeout,
but with this change it actually becomes 4s. AI says a successful
tc_tunnel run makes 50+ connections, so it adds up to minutes.
We should probably be using opts->timeout_ms as an absolute budget set
by the caller, and pass it (or remainder?) to the poll().
> +
> + if (err <= 0) {
> + if (err == 0) {
> + log_err("Connection timeout");
> + errno = ETIMEDOUT;
> + } else {
> + log_err("Failed to poll connection status");
> + }
> + goto close;
> + }
Also poll() can return EINTR here. I think we should retry EINTR,
taking into account the absolute deadline.
pw-bot: cr
> +
> + errlen = sizeof(err);
> + if (getsockopt(fd, SOL_SOCKET, SO_ERROR, &err, &errlen) < 0) {
> + log_err("Failed to getsockopt");
> + goto close;
> + }
> +
> + if (err) {
> + log_err("Eventually failed to connect to server");
> + errno = err;
> + goto close;
> + }
> + } else if (err) {
> log_err("Failed to connect to server");
> - save_errno_close(fd);
> - return -1;
> + goto close;
> }
>
> return fd;
> +close:
> + save_errno_close(fd);
> + return -1;
> }
>
> int connect_to_addr_str(int family, int type, const char *addr_str, __u16 port,
>
> ---
> base-commit: 2efc18d4bb9ad28d240b69bd324937f1ec12e93d
> change-id: 20260710-tc_tunnel_flaky-27e9a191bd03
>
> Best regards,
> --
> Alexis Lothoré (eBPF Foundation) <alexis.lothore@bootlin.com>
>
prev parent reply other threads:[~2026-08-06 18:57 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 7:36 [PATCH bpf v2] selftests/bpf: keep polling connection that is still in progress Alexis Lothoré (eBPF Foundation)
2026-08-03 8:35 ` bot+bpf-ci
2026-08-06 18:56 ` Ihor Solodrai [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=71b9f5ae-eed4-4d03-a3b9-2493b525bc34@linux.dev \
--to=ihor.solodrai@linux.dev \
--cc=alexis.lothore@bootlin.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bastien.curutchet@bootlin.com \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=ebpf@linuxfoundation.org \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=thomas.petazzoni@bootlin.com \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox