From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f47.google.com (mail-pj1-f47.google.com [209.85.216.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ADB0D3815E2 for ; Tue, 25 Aug 2026 18:07:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.47 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787681269; cv=none; b=aCt4sWMkx0ZYrM61DW0lip9td+qFpu8jofFUa+73kdhD8d84e01XpHaeqX880lgYNXa/CIEBgigXO4cLgYS5V3R/Wrczfhn6NUOpR/BYHv89NXjUvi6jau2V4Avnfeyn+12H8Lo9GI52oBbfz8YkgH5zU1S+/Dde/3aKXsrOXDA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787681269; c=relaxed/simple; bh=mfvi2AZWpG/1aroz6xmgFc854p2wpbKkK21zgF1vBVM=; h=Mime-Version:Content-Type:Date:Message-Id:Cc:Subject:From:To: References:In-Reply-To; b=l4vwbxl8tflmcfvVWOIXOur8HQEqG1bvDiVC9l1VWHuZr7Hgx5jf8cjXumiCKlI49bR9ZJWtDW3jz9JiXQGr02C47bvb5oCbcMiNr1DNfWWF0Es9OUNGPm8w5BMJUEEHcmYBOL0Vrhzt4xt/nJ5MYqzlRklVtKmLPMvUQLdZvZo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=CAgjqaRr; arc=none smtp.client-ip=209.85.216.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="CAgjqaRr" Received: by mail-pj1-f47.google.com with SMTP id 98e67ed59e1d1-3964dfb5a69so214399a91.1 for ; Tue, 25 Aug 2026 11:07:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787681267; x=1788286067; darn=vger.kernel.org; h=in-reply-to:references:to:from:subject:cc:message-id:date :content-type:content-transfer-encoding:mime-version:from:to:cc :subject:date:message-id:reply-to:content-type; bh=pYrAALZVR+9Z8bDYcIRRZ69pFI6c+MuDXpNxh9XTeIA=; b=CAgjqaRrg2qsOA7lWaoG+B5QNi3vyuWR72eErHOqcoA9zHHbJPTBCNwmsGcHGKExK4 2JK8VRaIkU82VJI3M4Fsk9XxEZiAgJC+JouiQEepYrvM0Cb80Pm8S8L71Bi5xWNJuHi+ cKZ4dGFry4MYTNcwPYpWvaaBnCvHUeYdoafvqwKcrRa2DhYpu2yp7y/CfnB/tpLgDxjx n8Qoovi/q8c7H+V5sqkK3X/gTYesvoeVFiRW9fXIwae1twotrEwzQCsLnQte87lH3veH mQbCyEgpqZV5nsrceSlG5DkHhGg15epIeL+BqJ+rQLMHeVogNdzUrWTBgwCjDSK+Ccng 6NaA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787681267; x=1788286067; h=in-reply-to:references:to:from:subject:cc:message-id:date :content-type:content-transfer-encoding:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=pYrAALZVR+9Z8bDYcIRRZ69pFI6c+MuDXpNxh9XTeIA=; b=Q2i3N87XMiSZaFXuH84Ti2GDsC0iHVDYZprTLwBeKHfRPIJedlVXl/PaHPwe1I7i05 deLs93MXBCSw2mICuF+nZ6mntvyBvVa0MThXtcCfvIOLjhnTyuIXazYy5i0Arj80XeEf if7aPcZL94ZP5Hyxsc289GXPwRg3C6Rj2TsaJ97TkMw3PGpA+UGGXF9WekRPc5A50nnn ra9rKZRkbWt87flqFEY7cmgGyspElaKjSpwlba18HwKJRiqXwIrn+JhiJ/wkoO4Ow8dn Z1y0tqq5ZKH8KirF/Mrno0yzIw1INDWICaigsy95qzGOFkNi0K9gGB+cCRSixaSq78Y5 Uu8g== X-Forwarded-Encrypted: i=1; AHgh+RplmHDWjw6jq1o5SNxBGndyNZJBFHVLe5ErW5vRSaZvyCOnS/ZDFxx9rOtuq78J7bNqKwjM4ifJmyTdNY0=@vger.kernel.org X-Gm-Message-State: AFuF++lMvoiuX2ZE5HHEHtCZjdXEQ+MZHgKwNiMMOlZ9SoBUBtWZpHsx oljHppIcT//pzpp5UZFY2qwDgHcvWrD22CrmX5HFKZM/qywnWC+Egi0T X-Gm-Gg: AR+sD12HeUUF46kuLhsy9nZi6Q7GhlLurbrwoFlk7bjAuOayGP3IUs0PabuolMcwuVU K6uyvunji5SeXbNLQsNXMwW5ZuIFhyYk2NuA8OaaEio6i3mlotPKNk6JLxjXDt4n4bWrIIMCapt yXrpKk+xkgQvS4htdFMfillujAQI2/nOHx9404okRN9qHF//K0qtLs1gNuyQ2Vb1pKRxSUc/9mR N9ae8rmzhjb+0Afdl3Cqei3lZGch9lCkPkXopN0EhbMbhNlBz09c4+RSVZe06MxJpN35jQUQ6Zl ZuRdyJX/BwjBS8dLZfH6yk5t7/mQHXxKL//6SEdWObUChqfCAYHhfo89RPtryh0sF0qlz1ps/t2 85ZSkgoavb6fe0M+Mw/WV2JHwHALfbrC5mdexq3npptGbFbSYBBXA/tJgFp60fcQtEZm3J2s1tR z1Lkt/Epvm0jH/PozNWCO+BvYNa0YH7tPbFkmyYVhogXWDuqBr5knXRe4= X-Received: by 2002:a17:90b:51d2:b0:381:6c5:3f63 with SMTP id 98e67ed59e1d1-3966d21833dmr2061621a91.6.1787681266754; Tue, 25 Aug 2026 11:07:46 -0700 (PDT) Received: from localhost ([13.93.150.60]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-141a9036c85sm805104c88.9.2026.08.25.11.07.45 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 25 Aug 2026 11:07:46 -0700 (PDT) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=UTF-8 Date: Tue, 25 Aug 2026 18:07:44 +0000 Message-Id: Cc: "David Ahern" , "David S . Miller" , "Eric Dumazet" , "Jakub Kicinski" , "Paolo Abeni" , "Simon Horman" , "Arun Ajith S" , "Roopa Prabhu" , "Jaehee Park" , , Subject: Re: [RFC net-next] ipv6: update NUD_FAILED neighbors from NA messages From: "Lawrence Lee" To: "Ido Schimmel" X-Mailer: aerc 0.17.0 References: <20260813233344.445265-1-lfqlee314@gmail.com> <20260825115226.GA1338930@shredder> In-Reply-To: <20260825115226.GA1338930@shredder> On Tue Aug 25, 2026 at 11:52 AM UTC, Ido Schimmel wrote: > On Thu, Aug 13, 2026 at 11:33:44PM +0000, Lawrence Lee wrote: > > I noticed an inconsistency between IPv4 and IPv6 in how the kernel hand= les > > neighbor advertisements/ARP replies for NUD_FAILED neighbor entries > > and would like some guidance/input from maintainers. > >=20 > > For an existing IPv4 neighbor that is in the NUD_FAILED state, receivin= g an > > ARP reply for the neighbor IP will update the entry in the kernel to > > either NUD_STALE or NUD_REACHABLE depending on if the reply is unicast = or > > broadcast. > >=20 > > For an existing IPv6 neighbor that is in the NUD_FAILED state, receivin= g a > > neighbor advertisement (NA) for the neighbor IP does nothing as > > ndisc_recv_na() explicitly ignores NUD_FAILED neighbor entries: > >=20 > > if (READ_ONCE(neigh->nud_state) & NUD_FAILED) > > goto out; > >=20 > > This check was added by commit titled "[IPV6] Don't update FAILED > > entries on receipt of NAs." (Hideaki Yoshifuji, 2005-01-16; pre-git, in > > mainline since v2.6.12-rc2) with the justification "As NAs do not creat= e > > new entries (RFC2461 7.2.5), NA should not change state of FAILED entri= es." > >=20 > > However, RFC9131 introduced a method for NAs to create new neighbor ent= ries > > (implemented as `accept_unsolicited_na` and later renamed to > > `accept_untracked_na`), which means the original justification for igno= ring > > NAs for NUD_FAILED neighbor entries is no longer 100% correct. I think = to > > remain logically consistent, it makes sense to allow NAs to update > > NUD_FAILED entries anytime we allow creating new entries with > > `accept_untracked_na`. > >=20 > > I realize that RFC9131 section 4.2 states the following: > >=20 > > ... routers create a new Neighbor Cache entry upon > > receiving an unsolicited Neighbor Advertisement for an address that > > does not already have a Neighbor Cache entry. These changes do not > > modify the router behavior specified in [RFC4861] for the scenario > > when the corresponding Neighbor Cache entry already exists. > >=20 > > However, I would argue that since NUD_FAILED is purely a kernel constru= ct > > and has no equivalent state defined in RFC4861 section 7.3.2, a neighbo= r in > > state NUD_FAILED does not actually have a valid Neighbor Cache entry as > > defined by RFC4861 and should be treated as if the neighbor entry doesn= 't > > exist; therefore NUD_FAILED neighbors does fall within the scope of > > RFC9131. > >=20 > > The motivation for this question comes from my work on SONiC, a network= OS > > which is built on top of Debian and runs on switching hardware. We have > > encountered an issue where the switch receives traffic for an IPv6 neig= hbor > > before that neighbor is resolvable, which leads to the kernel neighbor > > being set to NUD_FAILED. When the IPv6 neighbor becomes ready to receiv= e > > traffic, it sends an unsolicited NA to the switch which gets ignored > > because the kernel neighbor is NUD_FAILED. Subsequent traffic destined = to > > this neighbor stays entirely within the switch ASIC and isn't visible t= o > > the kernel, so there's no stimulus for the kernel to send neighbor > > solicitations; as a result, the neighbor entry stays unresolved and tra= ffic > > to the neighbor is dropped. > > Why "Subsequent traffic destined to this neighbor stays entirely within > the switch ASIC and isn't visible to the kernel"? If the neighbour is > unresolved and you're relying on the kernel to perform the resolution, > then you should trap these packets and inject them to the kernel's Rx > path. This should provide "stimulus for the kernel to send neighbor > solicitations". Normally, this is what happens. However, I am working on a scenario=20 where we have two switches providing connectivity to a single rack of=20 servers to provide increased redundancy. Each server is connected to=20 both switches using a single cable with 3 ends, and for each server one=20 switch is designated as the 'active' switch and the other as 'standby'=20 at any given time. When we have a FAILED neighbor on one switch, we=20 cannot determine which specific server/switch interface that neighbor=20 should be associated with. For any traffic destined to the FAILED=20 neighbor IP, we route it to the peer switch to maximize the chance that=20 packets can be forwarded successfully (with the hope that the peer is=20 able to resolve the neighbor entry). The routing to the peer takes place=20 entirely within the ASIC, which is why we cannot rely on the kernel to=20 resolve the neighbor in this particular scenario. There is a publicly=20 available HLD on this feature if you are interested:=20 https://github.com/sonic-net/SONiC/blob/master/doc/dualtor/dualtor_active_s= tandby_hld.md#6352-neighbor-miss-due-to-one-side-link-down When we see this particular issue, the traffic is landing on the=20 'active' switch for this particular neighbor IP/server, the switch just=20 doesn't know it yet since the neighbor isn't resolved. The packets get=20 routed to the peer switch, which is 'standby' for the same server. On=20 the peer, these packets will get trapped to the kernel but the peer=20 still cannot resolve the neighbor since it is in 'standby' mode for the=20 associated interface/server (this is a constraint of the physical cable=20 that is used to connect the server to both switches, the cable drops=20 traffic from the standby switch). The routing to the peer was originally intended to still allow=20 forwarding when a physical link/cabling issue prevents one switch from=20 resolving a neighbor entry. It just has the unfortunate effect of=20 preventing the switch from ever resolving the neighbor entry if we get=20 traffic for a neighbor IP before it's ready/resolvable. > If you can't do this for some reason (please explain why), then you can > either: > > 1. Trigger the resolution from user space via NTF_USE. The switch unfortunately has no mechanism to determine when any neighbor=20 becomes ready and resolvable so I'm not sure that this is feasible. =20 SONiC does already have a mechanism to retry resolution of FAILED=20 neighbors but since it only runs periodically, it's not always=20 performant enough for every situation. We are hesitant to increase the=20 retry frequency due to concerns about CPU load. > 2. Configure offloaded neighbours with NTF_EXT_MANAGED so that the > kernel will periodically probe them and keep them reachable when > possible. I believe NTF_EXT_MANAGED neighbors would have similar drawbacks to the=20 existing resolution retry mechanism in SONiC. Retrying too frequently=20 risks excessive CPU load especially with many neighbor entries in the=20 kernel, but reducing the frequency means longer periods of time where=20 traffic cannot be forwarded. > > The main questions I'd like to pose: > >=20 > > 1. When RFC9131 was implemented (`accept_untracked_na`), was an intenti= onal > > choice made to not update the handling of NUD_FAILED neighbors? I > > searched through the discussions for all three commits relevant to t= his > > feature but did not find any mention of NUD_FAILED handling: > > commit f9a2fb73318e ("net/ipv6: Introduce accept_unsolicited_na = knob to implement router-side changes for RFC9131") > > commit 3e0b8f529c10 ("net/ipv6: Expand and rename accept_unsolic= ited_na to accept_untracked_na") > > commit aaa5f515b16b ("net: ipv6: new accept_untracked_na option = to accept na only if in-network") > > AFAIK it wasn't an intentional decision to not update NUD_FAILED > neighbours. > > >=20 > > 2. Should unsolicited NAs be allowed to update NUD_FAILED neighbors whe= n > > `accept_untracked_na` is enabled (this would align with existing > > IPv4/ARP behavior that allows ARP replies to update NUD_FAILED > > neighbors). > > Looks fine, but I suggest first evaluating the alternatives above. They > don't require any kernel changes. Besides, dropping traffic to an > unresolved neighbour instead of trapping to the CPU doesn't make a lot > of sense. Unfortunately some of the platforms that we need to support have fairly=20 weak CPUs. If we tune the alternatives to minimize the amount of time=20 where the neighbor is unresolved and traffic gets dropped, there is a=20 risk of excessive CPU usage which could in turn cause other issues. > Also, note that "align with existing IPv4/ARP behavior" is not accurate: > arp_process() updates an existing NUD_FAILED entry unconditionally and > it can transition from NUD_FAILED to NUD_REACHABLE. So, your patch is > strictly more restrictive than IPv4 rather than aligned with it. Sorry I should have been more specific before. Maybe a better way to=20 phrase it is that this would be analogous to IPv4 in the sense that it=20 allows NUD_FAILED neighbors to change states upon receipt of a NA, but=20 by using NUD_STALE it remains consistent with RFC9131 Section 6.1.1.