From: Ihor Solodrai <ihor.solodrai@linux.dev>
To: "Alexis Lothoré (eBPF Foundation)" <alexis.lothore@bootlin.com>,
"Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"Andrii Nakryiko" <andrii@kernel.org>,
"Eduard Zingerman" <eddyz87@gmail.com>,
"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
"Martin KaFai Lau" <martin.lau@linux.dev>,
"Song Liu" <song@kernel.org>,
"Yonghong Song" <yonghong.song@linux.dev>,
"Jiri Olsa" <jolsa@kernel.org>,
"Emil Tsalapatis" <emil@etsalapatis.com>,
"Shuah Khan" <shuah@kernel.org>
Cc: ebpf@linuxfoundation.org,
Bastien Curutchet <bastien.curutchet@bootlin.com>,
Thomas Petazzoni <thomas.petazzoni@bootlin.com>,
bpf@vger.kernel.org, linux-kselftest@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH bpf v2] selftests/bpf: keep polling connection that is still in progress
Date: Thu, 6 Aug 2026 11:56:47 -0700 [thread overview]
Message-ID: <71b9f5ae-eed4-4d03-a3b9-2493b525bc34@linux.dev> (raw)
In-Reply-To: <20260803-tc_tunnel_flaky-v2-1-657b287dfa75@bootlin.com>
On 8/3/26 12:36 AM, Alexis Lothoré (eBPF Foundation) wrote:
> Some tests, like tc_tunnel or tc_edt, sporadically fail in CI with the
> following logs:
>
> (network_helpers.c:309: errno: Operation now in progress) \
> Failed to connect to server
> send_and_test_data:FAIL:connect to server unexpected error: -115
>
> This is due to SO_RCVTIMEO and SO_SNDTIMEO being set on the client
> socket (see settimeo() in client_socket()), allowing connect() to return
> an error and to set errno to EINPROGRESS instead of blocking until
> connection result is known. Increasing the timeout value for those tests
> is likely not a good solution (and it has already been done by commit
> 2790db208b44 ("selftests/bpf: Improve tc_tunnel test reliability")):
> they involve subtests that expect the connection to fail, and so
> increasing the timeout value would increase overall test execution
> duration again (not only the connection, but any socket operation).
>
> Another solution, as documented in man 2 connect, is to poll the socket
> for POLLOUT once connect has returned EINPROGRESS, and to get the actual
> connection result through getsockopt: this allows to keep the overall
> timeout values low for the general traffic, while letting a chance to
> the connection to succeed even if CI runners are loaded.
>
> When connect() returns EINPROGRESS, poll the socket for POLLOUT and
> check the connection result via getsockopt(SO_ERROR).
>
> Fixes: 99126abec5e5 ("bpf: selftests: A few improvements to network_helpers.c")
> Signed-off-by: Alexis Lothoré (eBPF Foundation) <alexis.lothore@bootlin.com>
> ---
> Changes in v2:
> - drop unneeded initialization
> - add back error message for immediate connection failure, and slightly
> reword the async connection failure error message
> - Link to v1: https://patch.msgid.link/20260710-tc_tunnel_flaky-v1-1-42aab5399a49@bootlin.com
> ---
> I manage to reproduce the issue locally by running `./test_progs -a
> tc_tunnel` in a qemu machine, while making all my CPUs busy with
> stress-ng on host side; the issue happens pretty quickly. I have not
> been able to reproduce the issue anymore with this fix.
> ---
> tools/testing/selftests/bpf/network_helpers.c | 41 ++++++++++++++++++++++++---
> 1 file changed, 37 insertions(+), 4 deletions(-)
>
> diff --git a/tools/testing/selftests/bpf/network_helpers.c b/tools/testing/selftests/bpf/network_helpers.c
> index db935a9d9fc1..52e7b72f8a77 100644
> --- a/tools/testing/selftests/bpf/network_helpers.c
> +++ b/tools/testing/selftests/bpf/network_helpers.c
> @@ -14,6 +14,7 @@
> #include <sys/types.h>
> #include <sys/un.h>
> #include <sys/eventfd.h>
> +#include <sys/poll.h>
>
> #include <linux/err.h>
> #include <linux/in.h>
> @@ -40,6 +41,8 @@
> #define IPPROTO_MPTCP 262
> #endif
>
> +#define CONNECTION_IN_PROGRESS_TIMEOUT_MS 3000
> +
> #define clean_errno() (errno == 0 ? "None" : strerror(errno))
> #define log_err(MSG, ...) ({ \
> int __save = errno; \
> @@ -294,7 +297,8 @@ int client_socket(int family, int type,
> int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t addrlen,
> const struct network_helper_opts *opts)
> {
> - int fd;
> + socklen_t errlen;
> + int fd, err;
>
> if (!opts)
> opts = &default_opts;
> @@ -305,13 +309,42 @@ int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t add
> return -1;
> }
>
> - if (connect(fd, (const struct sockaddr *)addr, addrlen)) {
> + err = connect(fd, (const struct sockaddr *)addr, addrlen);
> + if (err && errno == EINPROGRESS) {
> + struct pollfd pfd = { .fd = fd, .events = POLLOUT };
> +
> + err = poll(&pfd, 1, CONNECTION_IN_PROGRESS_TIMEOUT_MS);
Hi Alexis, thanks for the patch.
When I ran tc_* selftests in parallel it looked like they hanged.
I think what's happening is that this 3s timeout is additive, making
some tests to wait for too long. For example tc_tunnel has 1s timeout,
but with this change it actually becomes 4s. AI says a successful
tc_tunnel run makes 50+ connections, so it adds up to minutes.
We should probably be using opts->timeout_ms as an absolute budget set
by the caller, and pass it (or remainder?) to the poll().
> +
> + if (err <= 0) {
> + if (err == 0) {
> + log_err("Connection timeout");
> + errno = ETIMEDOUT;
> + } else {
> + log_err("Failed to poll connection status");
> + }
> + goto close;
> + }
Also poll() can return EINTR here. I think we should retry EINTR,
taking into account the absolute deadline.
pw-bot: cr
> +
> + errlen = sizeof(err);
> + if (getsockopt(fd, SOL_SOCKET, SO_ERROR, &err, &errlen) < 0) {
> + log_err("Failed to getsockopt");
> + goto close;
> + }
> +
> + if (err) {
> + log_err("Eventually failed to connect to server");
> + errno = err;
> + goto close;
> + }
> + } else if (err) {
> log_err("Failed to connect to server");
> - save_errno_close(fd);
> - return -1;
> + goto close;
> }
>
> return fd;
> +close:
> + save_errno_close(fd);
> + return -1;
> }
>
> int connect_to_addr_str(int family, int type, const char *addr_str, __u16 port,
>
> ---
> base-commit: 2efc18d4bb9ad28d240b69bd324937f1ec12e93d
> change-id: 20260710-tc_tunnel_flaky-27e9a191bd03
>
> Best regards,
> --
> Alexis Lothoré (eBPF Foundation) <alexis.lothore@bootlin.com>
>
prev parent reply other threads:[~2026-08-06 18:57 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 7:36 [PATCH bpf v2] selftests/bpf: keep polling connection that is still in progress Alexis Lothoré (eBPF Foundation)
2026-08-03 7:44 ` sashiko-bot
2026-08-03 8:35 ` bot+bpf-ci
2026-08-06 18:56 ` Ihor Solodrai [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=71b9f5ae-eed4-4d03-a3b9-2493b525bc34@linux.dev \
--to=ihor.solodrai@linux.dev \
--cc=alexis.lothore@bootlin.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bastien.curutchet@bootlin.com \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=ebpf@linuxfoundation.org \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=thomas.petazzoni@bootlin.com \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.