From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from dediextern.your-server.de (dediextern.your-server.de [85.10.215.232]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D95D022B8AB; Tue, 10 Jun 2025 14:40:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=85.10.215.232 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1749566459; cv=none; b=OMuYP83xeIy3X9Nz/GBlweHo7Fmxgz+aogCPTT8St8EMmafFvRVRCsK4Uv9MAdwo5allliAggHb3HQWl0K5UiRXy9Zy+u0lkwWdbplfCA0zqP1sdrn0MzPxX2iVyK3/GvWdW+FabImzVUvw67WfFmWvTaI0yM6i6I1KbNse6bnQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1749566459; c=relaxed/simple; bh=oLcHXzZbfxkyWuSpFWR0PuAz9UabAeQdJSISj8eTHuo=; h=Message-ID:Date:MIME-Version:To:Cc:References:From:Subject: In-Reply-To:Content-Type; b=oCookz8XXLob3onVdeRB9dUdvLp0Z1X88CvpL0JZ52gZ+KKYub/nvC98Xrrx3vx6FYpgTSHWelVlgcVeq5/bjxQmPH2M5riGEW6MLqYHbgiQg4wHrlbxmeXv598E3Q9fA2+EyMghY8o3nmxlpccWITcfA3XghNtJmJFgUvbFy+E= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=hetzner-cloud.de; spf=pass smtp.mailfrom=hetzner-cloud.de; arc=none smtp.client-ip=85.10.215.232 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=hetzner-cloud.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=hetzner-cloud.de Received: from sslproxy07.your-server.de ([78.47.199.104]) by dediextern.your-server.de with esmtpsa (TLS1.3) tls TLS_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1uP09d-0008pe-4v; Tue, 10 Jun 2025 16:40:45 +0200 Received: from localhost ([127.0.0.1]) by sslproxy07.your-server.de with esmtpsa (TLS1.3) tls TLS_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1uP09c-000ILZ-26; Tue, 10 Jun 2025 16:40:44 +0200 Message-ID: <3d520952-c343-4ae5-98c0-c9965dc7e320@hetzner-cloud.de> Date: Tue, 10 Jun 2025 16:40:43 +0200 Precedence: bulk X-Mailing-List: netdev@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird To: Eric Dumazet Cc: =?UTF-8?Q?Toke_H=C3=B8iland-J=C3=B8rgensen?= , Jesper Dangaard Brouer , bpf@vger.kernel.org, netdev@vger.kernel.org, Alexei Starovoitov , Daniel Borkmann , John Fastabend , Andrew Lunn , "David S. Miller" , Jakub Kicinski , Paolo Abeni , Jamal Hadi Salim , Cong Wang , Jiri Pirko , linux-kernel@vger.kernel.org References: <9da42688-bfaa-4364-8797-e9271f3bdaef@hetzner-cloud.de> <87zfemtbah.fsf@toke.dk> <5f19b555-b0fa-472a-a5f3-6673c0b69c5c@hetzner-cloud.de> Content-Language: en-US From: Marcus Wichelmann Autocrypt: addr=marcus.wichelmann@hetzner-cloud.de; keydata= xsFNBGJGrHIBEADXeHfBzzMvCfipCSW1oRhksIillcss321wYAvXrQ03a9VN2XJAzwDB/7Sa N2Oqs6JJv4u5uOhaNp1Sx8JlhN6Oippc6MecXuQu5uOmN+DHmSLObKVQNC9I8PqEF2fq87zO DCDViJ7VbYod/X9zUHQrGd35SB0PcDkXE5QaPX3dpz77mXFFWs/TvP6IvM6XVKZce3gitJ98 JO4pQ1gZniqaX4OSmgpHzHmaLCWZ2iU+Kn2M0KD1+/ozr/2bFhRkOwXSMYIdhmOXx96zjqFV vIHa1vBguEt/Ax8+Pi7D83gdMCpyRCQ5AsKVyxVjVml0e/FcocrSb9j8hfrMFplv+Y43DIKu kPVbE6pjHS+rqHf4vnxKBi8yQrfIpQqhgB/fgomBpIJAflu0Phj1nin/QIqKfQatoz5sRJb0 khSnRz8bxVM6Dr/T9i+7Y3suQGNXZQlxmRJmw4CYI/4zPVcjWkZyydq+wKqm39SOo4T512Nw fuHmT6SV9DBD6WWevt2VYKMYSmAXLMcCp7I2EM7aYBEBvn5WbdqkamgZ36tISHBDhJl/k7pz OlXOT+AOh12GCBiuPomnPkyyIGOf6wP/DW+vX6v5416MWiJaUmyH9h8UlhlehkWpEYqw1iCA Wn6TcTXSILx+Nh5smWIel6scvxho84qSZplpCSzZGaidHZRytwARAQABzTZNYXJjdXMgV2lj aGVsbWFubiA8bWFyY3VzLndpY2hlbG1hbm5AaGV0em5lci1jbG91ZC5kZT7CwZgEEwEIAEIW IQQVqNeGYUnoSODnU2dJ0we/n6xHDgUCYkascgIbAwUJEswDAAULCQgHAgMiAgEGFQoJCAsC BBYCAwECHgcCF4AACgkQSdMHv5+sRw4BNxAAlfufPZnHm+WKbvxcPVn6CJyexfuE7E2UkJQl s/JXI+OGRhyqtguFGbQS6j7I06dJs/whj9fOhOBAHxFfMG2UkraqgAOlRUk/YjA98Wm9FvcQ RGZe5DhAekI5Q9I9fBuhxdoAmhhKc/g7E5y/TcS1s2Cs6gnBR5lEKKVcIb0nFzB9bc+oMzfV caStg+PejetxR/lMmcuBYi3s51laUQVCXV52bhnv0ROk0fdSwGwmoi2BDXljGBZl5i5n9wuQ eHMp9hc5FoDF0PHNgr+1y9RsLRJ7sKGabDY6VRGp0MxQP0EDPNWlM5RwuErJThu+i9kU6D0e HAPyJ6i4K7PsjGVE2ZcvOpzEr5e46bhIMKyfWzyMXwRVFuwE7erxvvNrSoM3SzbCUmgwC3P3 Wy30X7NS5xGOCa36p2AtqcY64ZwwoGKlNZX8wM0khaVjPttsynMlwpLcmOulqABwaUpdluUg soqKCqyijBOXCeRSCZ/KAbA1FOvs3NnC9nVqeyCHtkKfuNDzqGY3uiAoD67EM/R9N4QM5w0X HpxgyDk7EC1sCqdnd0N07BBQrnGZACOmz8pAQC2D2coje/nlnZm1xVK1tk18n6fkpYfR5Dnj QvZYxO8MxP6wXamq2H5TRIzfLN1C2ddRsPv4wr9AqmbC9nIvfIQSvPMBx661kznCacANAP/O wU0EYkascgEQAK15Hd7arsIkP7knH885NNcqmeNnhckmu0MoVd11KIO+SSCBXGFfGJ2/a/8M y86SM4iL2774YYMWePscqtGNMPqa8Uk0NU76ojMbWG58gow2dLIyajXj20sQYd9RbNDiQqWp RNmnp0o8K8lof3XgrqjwlSAJbo6JjgdZkun9ZQBQFDkeJtffIv6LFGap9UV7Y3OhU+4ZTWDM XH76ne9u2ipTDu1pm9WeejgJIl6A7Z/7rRVpp6Qlq4Nm39C/ReNvXQIMT2l302wm0xaFQMfK jAhXV/2/8VAAgDzlqxuRGdA8eGfWujAq68hWTP4FzRvk97L4cTu5Tq8WIBMpkjznRahyTzk8 7oev+W5xBhGe03hfvog+pA9rsQIWF5R1meNZgtxR+GBj9bhHV+CUD6Fp+M0ffaevmI5Untyl AqXYdwfuOORcD9wHxw+XX7T/Slxq/Z0CKhfYJ4YlHV2UnjIvEI7EhV2fPhE4WZf0uiFOWw8X XcvPA8u0P1al3EbgeHMBhWLBjh8+Y3/pm0hSOZksKRdNR6PpCksa52ioD+8Z/giTIDuFDCHo p4QMLrv05kA490cNAkwkI/yRjrKL3eGg26FCBh2tQKoUw2H5pJ0TW67/Mn2mXNXjen9hDhAG 7gU40lS90ehhnpJxZC/73j2HjIxSiUkRpkCVKru2pPXx+zDzABEBAAHCwXwEGAEIACYWIQQV qNeGYUnoSODnU2dJ0we/n6xHDgUCYkascgIbDAUJEswDAAAKCRBJ0we/n6xHDsmpD/9/4+pV IsnYMClwfnDXNIU+x6VXTT/8HKiRiotIRFDIeI2skfWAaNgGBWU7iK7FkF/58ys8jKM3EykO D5lvLbGfI/jrTcJVIm9bXX0F1pTiu3SyzOy7EdJur8Cp6CpCrkD+GwkWppNHP51u7da2zah9 CQx6E1NDGM0gSLlCJTciDi6doAkJ14aIX58O7dVeMqmabRAv6Ut45eWqOLvgjzBvdn1SArZm 7AQtxT7KZCz1yYLUgA6TG39bhwkXjtcfT0J4967LuXTgyoKCc969TzmwAT+pX3luMmbXOBl3 mAkwjD782F9sP8D/9h8tQmTAKzi/ON+DXBHjjqGrb8+rCocx2mdWLenDK9sNNsvyLb9oKJoE DdXuCrEQpa3U79RGc7wjXT9h/8VsXmA48LSxhRKn2uOmkf0nCr9W4YmrP+g0RGeCKo3yvFxS +2r2hEb/H7ZTP5PWyJM8We/4ttx32S5ues5+qjlqGhWSzmCcPrwKviErSiBCr4PtcioTBZcW VUssNEOhjUERfkdnHNeuNBWfiABIb1Yn7QC2BUmwOvN2DsqsChyfyuknCbiyQGjAmj8mvfi/ 18FxnhXRoPx3wr7PqGVWgTJD1pscTrbKnoI1jI1/pBCMun+q9v6E7JCgWY181WjxgKSnen0n wySmewx3h/yfMh0aFxHhvLPxrO2IEQ== Subject: Re: [BUG] veth: TX drops with NAPI enabled and crash in combination with qdisc In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Virus-Scanned: Clear (ClamAV 1.0.7/27664/Tue Jun 10 10:41:04 2025) Am 06.06.25 um 11:06 schrieb Eric Dumazet: > On Thu, Jun 5, 2025 at 3:17 PM Marcus Wichelmann > wrote: >> >> Am 06.06.25 um 00:11 schrieb Eric Dumazet: >>> On Thu, Jun 5, 2025 at 9:46 AM Eric Dumazet wrote: >>>> >>>> On Thu, Jun 5, 2025 at 9:15 AM Toke Høiland-Jørgensen wrote: >>>>> >>>>> Marcus Wichelmann writes: >>>>> >>>>>> Hi, >>>>>> >>>>>> while experimenting with XDP_REDIRECT from a veth-pair to another interface, I >>>>>> noticed that the veth-pair looses lots of packets when multiple TCP streams go >>>>>> through it, resulting in stalling TCP connections and noticeable instabilities. >>>>>> >>>>>> This doesn't seem to be an issue with just XDP but rather occurs whenever the >>>>>> NAPI mode of the veth driver is active. >>>>>> I managed to reproduce the same behavior just by bringing the veth-pair into >>>>>> NAPI mode (see commit d3256efd8e8b ("veth: allow enabling NAPI even without >>>>>> XDP")) and running multiple TCP streams through it using a network namespace. >>>>>> >>>>>> Here is how I reproduced it: >>>>>> >>>>>> ip netns add lb >>>>>> ip link add dev to-lb type veth peer name in-lb netns lb >>>>>> >>>>>> # Enable NAPI >>>>>> ethtool -K to-lb gro on >>>>>> ethtool -K to-lb tso off >>>>>> ip netns exec lb ethtool -K in-lb gro on >>>>>> ip netns exec lb ethtool -K in-lb tso off >>>>>> >>>>>> ip link set dev to-lb up >>>>>> ip -netns lb link set dev in-lb up >>>>>> >>>>>> Then run a HTTP server inside the "lb" namespace that serves a large file: >>>>>> >>>>>> fallocate -l 10G testfiles/10GB.bin >>>>>> caddy file-server --root testfiles/ >>>>>> >>>>>> Download this file from within the root namespace multiple times in parallel: >>>>>> >>>>>> curl http://[fe80::...%to-lb]/10GB.bin -o /dev/null >>>>>> >>>>>> In my tests, I ran four parallel curls at the same time and after just a few >>>>>> seconds, three of them stalled while the other one "won" over the full bandwidth >>>>>> and completed the download. >>>>>> >>>>>> This is probably a result of the veth's ptr_ring running full, causing many >>>>>> packet drops on TX, and the TCP congestion control reacting to that. >>>>>> >>>>>> In this context, I also took notice of Jesper's patch which describes a very >>>>>> similar issue and should help to resolve this: >>>>>> commit dc82a33297fc ("veth: apply qdisc backpressure on full ptr_ring to >>>>>> reduce TX drops") >>>>>> >>>>>> But when repeating the above test with latest mainline, which includes this >>>>>> patch, and enabling qdisc via >>>>>> tc qdisc add dev in-lb root sfq perturb 10 >>>>>> the Kernel crashed just after starting the second TCP stream (see output below). >>>>>> >>>>>> So I have two questions: >>>>>> - Is my understanding of the described issue correct and is Jesper's patch >>>>>> sufficient to solve this? >>>>> >>>>> Hmm, yeah, this does sound likely. >>>>> >>>>>> - Is my qdisc configuration to make use of this patch correct and the kernel >>>>>> crash is likely a bug? >>>>>> >>>>>> ------------[ cut here ]------------ >>>>>> UBSAN: array-index-out-of-bounds in net/sched/sch_sfq.c:203:12 >>>>>> index 65535 is out of range for type 'sfq_head [128]' >>>>> >>>>> This (the 'index 65535') kinda screams "integer underflow". So certainly >>>>> looks like a kernel bug, yeah. Don't see any obvious reason why Jesper's >>>>> patch would trigger this; maybe Eric has an idea? >>>>> >>>>> Does this happen with other qdiscs as well, or is it specific to sfq? >>>> >>>> This seems like a bug in sfq, we already had recent fixes in it, and >>>> other fixes in net/sched vs qdisc_tree_reduce_backlog() >>>> >>>> It is possible qdisc_pkt_len() could be wrong in this use case (TSO off ?) >>> >>> This seems to be a very old bug, indeed caused by sch->gso_skb >>> contribution to sch->q.qlen >>> >>> diff --git a/net/sched/sch_sfq.c b/net/sched/sch_sfq.c >>> index b912ad99aa15d95b297fb28d0fd0baa9c21ab5cd..77fa02f2bfcd56a36815199aa2e7987943ea226f >>> 100644 >>> --- a/net/sched/sch_sfq.c >>> +++ b/net/sched/sch_sfq.c >>> @@ -310,7 +310,10 @@ static unsigned int sfq_drop(struct Qdisc *sch, >>> struct sk_buff **to_free) >>> /* It is difficult to believe, but ALL THE SLOTS HAVE >>> LENGTH 1. */ >>> x = q->tail->next; >>> slot = &q->slots[x]; >>> - q->tail->next = slot->next; >>> + if (slot->next == x) >>> + q->tail = NULL; /* no more active slots */ >>> + else >>> + q->tail->next = slot->next; >>> q->ht[slot->hash] = SFQ_EMPTY_SLOT; >>> goto drop; >>> } >>> >> >> Hi, >> >> thank you for looking into it. >> I'll give your patch a try and will also do tests with other qdiscs as well when I'm back >> in office. >> > > I have been using this repro : > > [...] Hi, I can confirm that the sfq qdisc is now stable in this setup, thanks to your fix. I also experimented with other qdiscs and fq_codel works as well. The sfq/fq_codel qdisc works hand-in-hand now with Jesper's patch to resolve the original issue. Multiple TCP connections run very stable, even when NAPI/XDP is active on the veth device, and I can see that the packets are being requeued instead of being dropped in the veth driver. Thank you for your help! Marcus