Netdev List
 help / color / mirror / Atom feed
* [REGRESSION] r8169: RTL8168h silent RX stall with EEE after "net: phy: realtek: fix EEE advertisement write on the internal PHY MMD path" (6.18.y)
       [not found] <353419280.954808.1790290283530.ref@mail.yahoo.com>
@ 2026-09-24 22:51 ` Dennis Piecha
  2026-09-25  7:13   ` Oleksij Rempel
  0 siblings, 1 reply; 2+ messages in thread
From: Dennis Piecha @ 2026-09-24 22:51 UTC (permalink / raw)
  To: hkallweit1@gmail.com, nic_swsd@realtek.com,
	o.rempel@pengutronix.de
  Cc: andrew@lunn.ch, netdev@vger.kernel.org,
	regressions@lists.linux.dev, stable@vger.kernel.org,
	kuba@kernel.org, nb@tipi-net.de

Hi,

after updating from 6.18.39 to 6.18.52 (Home Assistant OS 18.2 -> 18.3,
close to vanilla stable), the onboard RTL8168h of my NUC silently stops
receiving every ~7-20 minutes. Disabling EEE at runtime makes it stable
again, so I suspect the stable backport of

202fef9bbbf5 ("net: phy: realtek: fix EEE advertisement write on the
internal PHY MMD path")

(0af3afd054e7 in 6.18.y). There are no changes to r8169_main.c between
6.18.39 and 6.18.52, but this PHY fix touches rtlgen_write_mmd(), which
is used by the "Generic FE-GE Realtek PHY" bound to this NIC. My
guess: with the advertisement write now landing, EEE actually gets
negotiated with the link partner, and RTL8168h doesn't handle LPI
correctly on the RX side. The fix itself looks correct to me; it seems
to expose an existing r8169/RTL8168h problem.

Hardware:
Intel NUC11ATKPE, BIOS ATJSLCPX.0039.2023.0221.1502
RTL8168h/8111h, 10ec:8168 (subsys 8086:3027), XID 541
PHY: Generic FE-GE Realtek PHY
link partner: AVM FRITZ!Box 6690 Cable, 1000/Full, flow control off
no VLANs, ASPM already disabled on the device (LnkCtl ASPM bits 0)

EEE state on 6.18.52 (read via ETHTOOL_GEEE):
supported=0x28 advertised=0x28 lp_advertised=0x28
eee_active=1 eee_enabled=1 tx_lpi_enabled=1 tx_lpi_timer=12

Symptom: carrier stays up, nothing at all is logged by r8169 or phylib,
no NETDEV WATCHDOG. rx_packets freezes while TX and the MSI-X interrupt
count keep going; no RX error/missed/dropped counters move:

time rx_packets tx_packets irq bql_inflight tx_timeout
23:08:29 626373 188818 201071 0 0
23:08:31 626458 188908 201236 0 0
23:08:34 626513 189037 201394 0 0 <- stall
23:08:37 626513 189126 201473 0 0
23:08:44 626513 189284 201606 0 0
23:08:51 626513 189517 201788 0 0

It never recovers by itself. "ip link set eno1 down; ip link set eno1 up"
restores traffic within 3 s; flushing the neighbour table does not help.
Disabling TX checksum offload and pinning MTU 1500 (workaround for the
VLAN-related TX timeouts reported by others on the same update) has no
effect here.

EEE test: on a 6.18.52 boot that had stalled at ~7 min and again ~7 min
later, I disabled EEE via ETHTOOL_SEEE (eee_active=0 afterwards). Since
then it has run for 60+ minutes without a stall, nothing else changed.

What I have NOT done yet: verified the EEE state on 6.18.39 for
comparison, and tested 6.18.52 with the commit reverted. I can do the
former easily; building a kernel is harder on this appliance-style OS
but possible if needed. I can also capture a MAC register dump
(ETHTOOL_GREGS) during a stall if that helps.

Original patch thread:
https://lore.kernel.org/all/20260806134716.3511821-1-o.rempel@pengutronix.de/

Downstream report with more detail:
https://github.com/home-assistant/operating-system/issues/5042

#regzbot introduced: 202fef9bbbf5

Thanks,
Dennis Piecha

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [REGRESSION] r8169: RTL8168h silent RX stall with EEE after "net: phy: realtek: fix EEE advertisement write on the internal PHY MMD path" (6.18.y)
  2026-09-24 22:51 ` [REGRESSION] r8169: RTL8168h silent RX stall with EEE after "net: phy: realtek: fix EEE advertisement write on the internal PHY MMD path" (6.18.y) Dennis Piecha
@ 2026-09-25  7:13   ` Oleksij Rempel
  0 siblings, 0 replies; 2+ messages in thread
From: Oleksij Rempel @ 2026-09-25  7:13 UTC (permalink / raw)
  To: Dennis Piecha
  Cc: hkallweit1@gmail.com, nic_swsd@realtek.com, andrew@lunn.ch,
	netdev@vger.kernel.org, regressions@lists.linux.dev,
	stable@vger.kernel.org, kuba@kernel.org, nb@tipi-net.de

H Dennis,

On Thu, Sep 24, 2026 at 10:51:23PM +0000, Dennis Piecha wrote:
> Hi,
> 
> after updating from 6.18.39 to 6.18.52 (Home Assistant OS 18.2 -> 18.3,
> close to vanilla stable), the onboard RTL8168h of my NUC silently stops
> receiving every ~7-20 minutes. Disabling EEE at runtime makes it stable
> again, so I suspect the stable backport of
> 
> 202fef9bbbf5 ("net: phy: realtek: fix EEE advertisement write on the
> internal PHY MMD path")
> 
> (0af3afd054e7 in 6.18.y). There are no changes to r8169_main.c between
> 6.18.39 and 6.18.52, but this PHY fix touches rtlgen_write_mmd(), which
> is used by the "Generic FE-GE Realtek PHY" bound to this NIC. My
> guess: with the advertisement write now landing, EEE actually gets
> negotiated with the link partner, and RTL8168h doesn't handle LPI
> correctly on the RX side. The fix itself looks correct to me; it seems
> to expose an existing r8169/RTL8168h problem.
> 
> Hardware:
> Intel NUC11ATKPE, BIOS ATJSLCPX.0039.2023.0221.1502
> RTL8168h/8111h, 10ec:8168 (subsys 8086:3027), XID 541
> PHY: Generic FE-GE Realtek PHY
> link partner: AVM FRITZ!Box 6690 Cable, 1000/Full, flow control off
> no VLANs, ASPM already disabled on the device (LnkCtl ASPM bits 0)
> 
> EEE state on 6.18.52 (read via ETHTOOL_GEEE):
> supported=0x28 advertised=0x28 lp_advertised=0x28
> eee_active=1 eee_enabled=1 tx_lpi_enabled=1 tx_lpi_timer=12
> 
> Symptom: carrier stays up, nothing at all is logged by r8169 or phylib,
> no NETDEV WATCHDOG. rx_packets freezes while TX and the MSI-X interrupt
> count keep going; no RX error/missed/dropped counters move:
> 
> time rx_packets tx_packets irq bql_inflight tx_timeout
> 23:08:29 626373 188818 201071 0 0
> 23:08:31 626458 188908 201236 0 0
> 23:08:34 626513 189037 201394 0 0 <- stall
> 23:08:37 626513 189126 201473 0 0
> 23:08:44 626513 189284 201606 0 0
> 23:08:51 626513 189517 201788 0 0
> 
> It never recovers by itself. "ip link set eno1 down; ip link set eno1 up"
> restores traffic within 3 s; flushing the neighbour table does not help.
> Disabling TX checksum offload and pinning MTU 1500 (workaround for the
> VLAN-related TX timeouts reported by others on the same update) has no
> effect here.
> 
> EEE test: on a 6.18.52 boot that had stalled at ~7 min and again ~7 min
> later, I disabled EEE via ETHTOOL_SEEE (eee_active=0 afterwards). Since
> then it has run for 60+ minutes without a stall, nothing else changed.
> 
> What I have NOT done yet: verified the EEE state on 6.18.39 for
> comparison, and tested 6.18.52 with the commit reverted. I can do the
> former easily; building a kernel is harder on this appliance-style OS
> but possible if needed. I can also capture a MAC register dump
> (ETHTOOL_GREGS) during a stall if that helps.

Can you please test following patch:
https://lore.kernel.org/all/20260925071040.2137178-1-o.rempel@pengutronix.de/

Best Regards,
Oleksij
-- 
Pengutronix e.K.                           |                             |
Steuerwalder Str. 21                       | http://www.pengutronix.de/  |
31137 Hildesheim, Germany                  | Phone: +49-5121-206917-0    |
Amtsgericht Hildesheim, HRA 2686           | Fax:   +49-5121-206917-5555 |

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-25  7:13 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <353419280.954808.1790290283530.ref@mail.yahoo.com>
2026-09-24 22:51 ` [REGRESSION] r8169: RTL8168h silent RX stall with EEE after "net: phy: realtek: fix EEE advertisement write on the internal PHY MMD path" (6.18.y) Dennis Piecha
2026-09-25  7:13   ` Oleksij Rempel

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox