The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Ihor Solodrai <ihor.solodrai@linux.dev>
To: "Alexis Lothoré (eBPF Foundation)" <alexis.lothore@bootlin.com>,
	"Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Song Liu" <song@kernel.org>,
	"Yonghong Song" <yonghong.song@linux.dev>,
	"Jiri Olsa" <jolsa@kernel.org>,
	"Emil Tsalapatis" <emil@etsalapatis.com>,
	"Shuah Khan" <shuah@kernel.org>
Cc: ebpf@linuxfoundation.org,
	Bastien Curutchet <bastien.curutchet@bootlin.com>,
	Thomas Petazzoni <thomas.petazzoni@bootlin.com>,
	bpf@vger.kernel.org, linux-kselftest@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH bpf v2] selftests/bpf: keep polling connection that is still in progress
Date: Thu, 6 Aug 2026 11:56:47 -0700	[thread overview]
Message-ID: <71b9f5ae-eed4-4d03-a3b9-2493b525bc34@linux.dev> (raw)
In-Reply-To: <20260803-tc_tunnel_flaky-v2-1-657b287dfa75@bootlin.com>

On 8/3/26 12:36 AM, Alexis Lothoré (eBPF Foundation) wrote:
> Some tests, like tc_tunnel or tc_edt, sporadically fail in CI with the
> following logs:
> 
>   (network_helpers.c:309: errno: Operation now in progress) \
>     Failed to connect to server
>   send_and_test_data:FAIL:connect to server unexpected error: -115
> 
> This is due to SO_RCVTIMEO and SO_SNDTIMEO being set on the client
> socket (see settimeo() in client_socket()), allowing connect() to return
> an error and to set errno to EINPROGRESS instead of blocking until
> connection result is known. Increasing the timeout value for those tests
> is likely not a good solution (and it has already been done by commit
> 2790db208b44 ("selftests/bpf: Improve tc_tunnel test reliability")):
> they involve subtests that expect the connection to fail, and so
> increasing the timeout value would increase overall test execution
> duration again (not only the connection, but any socket operation).
> 
> Another solution, as documented in man 2 connect, is to poll the socket
> for POLLOUT once connect has returned EINPROGRESS, and to get the actual
> connection result through getsockopt: this allows to keep the overall
> timeout values low for the general traffic, while letting a chance to
> the connection to succeed even if CI runners are loaded.
> 
> When connect() returns EINPROGRESS, poll the socket for POLLOUT and
> check the connection result via getsockopt(SO_ERROR).
> 
> Fixes: 99126abec5e5 ("bpf: selftests: A few improvements to network_helpers.c")
> Signed-off-by: Alexis Lothoré (eBPF Foundation) <alexis.lothore@bootlin.com>
> ---
> Changes in v2:
> - drop unneeded initialization
> - add back error message for immediate connection failure, and slightly
>   reword the async connection failure error message
> - Link to v1: https://patch.msgid.link/20260710-tc_tunnel_flaky-v1-1-42aab5399a49@bootlin.com
> ---
> I manage to reproduce the issue locally by running `./test_progs -a
> tc_tunnel` in a qemu machine, while making all my CPUs busy with
> stress-ng on host side; the issue happens pretty quickly. I have not
> been able to reproduce the issue anymore with this fix.
> ---
>  tools/testing/selftests/bpf/network_helpers.c | 41 ++++++++++++++++++++++++---
>  1 file changed, 37 insertions(+), 4 deletions(-)
> 
> diff --git a/tools/testing/selftests/bpf/network_helpers.c b/tools/testing/selftests/bpf/network_helpers.c
> index db935a9d9fc1..52e7b72f8a77 100644
> --- a/tools/testing/selftests/bpf/network_helpers.c
> +++ b/tools/testing/selftests/bpf/network_helpers.c
> @@ -14,6 +14,7 @@
>  #include <sys/types.h>
>  #include <sys/un.h>
>  #include <sys/eventfd.h>
> +#include <sys/poll.h>
>  
>  #include <linux/err.h>
>  #include <linux/in.h>
> @@ -40,6 +41,8 @@
>  #define IPPROTO_MPTCP 262
>  #endif
>  
> +#define CONNECTION_IN_PROGRESS_TIMEOUT_MS	3000
> +
>  #define clean_errno() (errno == 0 ? "None" : strerror(errno))
>  #define log_err(MSG, ...) ({						\
>  			int __save = errno;				\
> @@ -294,7 +297,8 @@ int client_socket(int family, int type,
>  int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t addrlen,
>  		    const struct network_helper_opts *opts)
>  {
> -	int fd;
> +	socklen_t errlen;
> +	int fd, err;
>  
>  	if (!opts)
>  		opts = &default_opts;
> @@ -305,13 +309,42 @@ int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t add
>  		return -1;
>  	}
>  
> -	if (connect(fd, (const struct sockaddr *)addr, addrlen)) {
> +	err = connect(fd, (const struct sockaddr *)addr, addrlen);
> +	if (err && errno == EINPROGRESS) {
> +		struct pollfd pfd = { .fd = fd, .events = POLLOUT };
> +
> +		err = poll(&pfd, 1, CONNECTION_IN_PROGRESS_TIMEOUT_MS);

Hi Alexis, thanks for the patch.

When I ran tc_* selftests in parallel it looked like they hanged.

I think what's happening is that this 3s timeout is additive, making
some tests to wait for too long. For example tc_tunnel has 1s timeout,
but with this change it actually becomes 4s. AI says a successful
tc_tunnel run makes 50+ connections, so it adds up to minutes.

We should probably be using opts->timeout_ms as an absolute budget set
by the caller, and pass it (or remainder?) to the poll().

> +
> +		if (err <= 0) {
> +			if (err == 0) {
> +				log_err("Connection timeout");
> +				errno = ETIMEDOUT;
> +			} else {
> +				log_err("Failed to poll connection status");
> +			}
> +			goto close;
> +		}

Also poll() can return EINTR here. I think we should retry EINTR,
taking into account the absolute deadline.

pw-bot: cr

> +
> +		errlen = sizeof(err);
> +		if (getsockopt(fd, SOL_SOCKET, SO_ERROR, &err, &errlen) < 0) {
> +			log_err("Failed to getsockopt");
> +			goto close;
> +		}
> +
> +		if (err) {
> +			log_err("Eventually failed to connect to server");
> +			errno = err;
> +			goto close;
> +		}
> +	} else if (err) {
>  		log_err("Failed to connect to server");
> -		save_errno_close(fd);
> -		return -1;
> +		goto close;
>  	}
>  
>  	return fd;
> +close:
> +	save_errno_close(fd);
> +	return -1;
>  }
>  
>  int connect_to_addr_str(int family, int type, const char *addr_str, __u16 port,
> 
> ---
> base-commit: 2efc18d4bb9ad28d240b69bd324937f1ec12e93d
> change-id: 20260710-tc_tunnel_flaky-27e9a191bd03
> 
> Best regards,
> --  
> Alexis Lothoré (eBPF Foundation) <alexis.lothore@bootlin.com>
> 


      parent reply	other threads:[~2026-08-06 18:57 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-03  7:36 [PATCH bpf v2] selftests/bpf: keep polling connection that is still in progress Alexis Lothoré (eBPF Foundation)
2026-08-03  8:35 ` bot+bpf-ci
2026-08-06 18:56 ` Ihor Solodrai [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=71b9f5ae-eed4-4d03-a3b9-2493b525bc34@linux.dev \
    --to=ihor.solodrai@linux.dev \
    --cc=alexis.lothore@bootlin.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bastien.curutchet@bootlin.com \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=ebpf@linuxfoundation.org \
    --cc=eddyz87@gmail.com \
    --cc=emil@etsalapatis.com \
    --cc=jolsa@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=shuah@kernel.org \
    --cc=song@kernel.org \
    --cc=thomas.petazzoni@bootlin.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox