Netdev List
 help / color / mirror / Atom feed
From: netdev-bot+sashiko@kernel.org
To: jiawenwu@trustnetic.com
Cc: netdev@vger.kernel.org, mengyuanlou@net-swift.com,
	andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
	kuba@kernel.org, pabeni@redhat.com, richardcochran@gmail.com,
	jacob.e.keller@intel.com, kees@kernel.org,
	aleksandr.loktionov@intel.com, vadim.fedorenko@linux.dev,
	sashiko-bot@kernel.org
Subject: Re: [PATCH net v2] net: libwx: fix races in Tx timestamp handling
Date: Thu, 17 Sep 2026 20:02:25 +0000	[thread overview]
Message-ID: <178967534514.22033.6813523642554517754@kernel.org> (raw)
In-Reply-To: <73D0D3F5D96A6928+20260914080020.211580-1-jiawenwu@trustnetic.com>

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:

Critical: 0 · High: 0 (1) · Medium: 1 · Low: 0 (3)

- [Medium] The patch makes wx_ptp_quiesce() (wx_ptp.c:886) perform an
  MMIO register read on a device whose PCIe channel may be frozen or…

Pre-existing issues:
- [High] The NAPI Tx cleanup path can call ptp_schedule_worker() on a PTP
  clock that wx_ptp_stop() is concurrently unregistering and NULLing.
- [Low] wx_ptp_suspend()
  (drivers/net/ethernet/wangxun/libwx/wx_ptp.c:841-852) tears down the…
- [Low] wx_xmit_frame_ring() reads wx->tstamp_config.tx_type inside the
  new ptp_tx_lock critical section (wx_lib.c:1656-1658), but the writer…
- [Low] In wx_xmit_frame_ring()'s out_drop path, wx->tx_hwtstamp_errors++
  is executed after dropping ptp_tx_lock (wx_lib.c:1706-1709), i.e.…

--

net: libwx: fix races in Tx timestamp handling

Adds a dedicated spinlock (wx->ptp_tx_lock) that serializes publication and
consumption of wx->ptp_tx_skb between the Tx path, the PTP auxiliary worker
and the timestamp cleanup paths.  The detach sequence is factored into
__wx_ptp_detach_tx_skb(), and the Tx error path now compares the outstanding
request against its own skb before cancelling it.

Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/6C7EC12D69217315%2B20260818074721.45536-1-jiawenwu%40trustnetic.com

[ ... ]

> diff --git a/drivers/net/ethernet/wangxun/libwx/wx_lib.c b/drivers/net/ethernet/wangxun/libwx/wx_lib.c
> index ed5aad7857bd..940ba2c6150f 100644
> --- a/drivers/net/ethernet/wangxun/libwx/wx_lib.c
> +++ b/drivers/net/ethernet/wangxun/libwx/wx_lib.c

[ ... ]

> @@ -1649,9 +1651,11 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb,
>  
>  	if (unlikely(skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) &&
>  	    wx->ptp_clock) {
> +		unsigned long flags;
> +
> +		spin_lock_irqsave(&wx->ptp_tx_lock, flags);
>  		if (wx->tstamp_config.tx_type == HWTSTAMP_TX_ON &&
> -		    !test_and_set_bit_lock(WX_STATE_PTP_TX_IN_PROGRESS,
> -					   wx->state)) {
> +		    !test_and_set_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state)) {

[Severity: Low]
This is a pre-existing inconsistency and not something introduced here, but
is the read of wx->tstamp_config.tx_type actually covered by the new lock?

The writer side never takes ptp_tx_lock, and it publishes the new config
only after the hardware has already been reprogrammed and cleaned up:

wx_hwtstamp_set()
    wx_ptp_set_timestamp_mode(wx, cfg);   /* clears WX_TSC_1588_CTL_ENABLED,
                                             calls wx_ptp_clear_tx_timestamp() */
    memcpy(&wx->tstamp_config, cfg, sizeof(wx->tstamp_config));

Can a transmit on another CPU still observe HWTSTAMP_TX_ON in that window,
arm WX_STATE_PTP_TX_IN_PROGRESS and set SKBTX_IN_PROGRESS for a frame the
hardware will never timestamp?  The request then pins the skb and its
socket until wx_ptp_tx_hang() cancels it (WX_PTP_TX_TIMEOUT is HZ) with a
bogus tx_hwtstamp_timeouts increment and a "clearing Tx timestamp hang"
warning.  Would publishing tx_type under ptp_tx_lock close that?

>  			skb_shinfo(skb)->tx_flags |= SKBTX_IN_PROGRESS;
>  			tx_flags |= WX_TX_FLAGS_TSTAMP;
>  			wx->ptp_tx_skb = skb_get(skb);

[Severity: Low]
This isn't a bug introduced by this patch either, but can this publication
outlive its only consumer on the suspend path?

wx_ptp_suspend() clears WX_STATE_PTP_RUNNING first, so wx_ptp_do_aux_work()
returns HZ without ever running wx_ptp_tx_hwtstamp_work() or
wx_ptp_tx_hang() again, and then performs the final cleanup while
wx->ptp_clock is still set and Tx has not been stopped yet:

txgbe_close_suspend()
    wx_ptp_suspend(wx);          /* clears RUNNING, wx_ptp_clear_tx_timestamp() */
    txgbe_disable_device(wx);    /* netif_tx_stop_all_queues()/netif_tx_disable() */
    wx_free_resources(wx);

Does an in-flight wx_xmit_frame_ring() that is already past the
queue-stopped check then leak the skb_get() reference (and the socket it
pins), with WX_STATE_PTP_TX_IN_PROGRESS left set and nothing left to
consume or time out the request?

On the PM path this is recovered later by resume -> ndo_open ->
wx_ptp_init() -> wx_ptp_reset() -> wx_ptp_set_timestamp_mode() ->
wx_ptp_clear_tx_timestamp(), so it appears to persist only on the
poweroff/shutdown path.  The ndo_stop path looks unaffected, since
dev_deactivate_many() and netpoll_poll_disable() run before ndo_stop.

>  @@ -1659,6 +1663,7 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb,
>  		} else {
>  			wx->tx_hwtstamp_skipped++;
>  		}
> +		spin_unlock_irqrestore(&wx->ptp_tx_lock, flags);
>  	}
>  
>  	/* record initial flags and protocol */
> @@ -1677,19 +1682,34 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb,
>  		wx->atr(tx_ring, first, ptype);
>  
>  	if (wx_tx_map(tx_ring, first, hdr_len))
> -		goto cleanup_tx_tstamp;
> +		goto out_drop;
>  
>  	return NETDEV_TX_OK;
>  out_drop:
> -	dev_kfree_skb_any(first->skb);
> -	first->skb = NULL;
> -cleanup_tx_tstamp:
> +	/* The hardware will never report a timestamp for a frame it did not
> +	 * transmit, so drop the request. Only do so if it is still ours: the
> +	 * PTP worker may already have completed it and a concurrent transmit
> +	 * may have submitted a new one.
> +	 */
>  	if (unlikely(tx_flags & WX_TX_FLAGS_TSTAMP)) {
> -		dev_kfree_skb_any(wx->ptp_tx_skb);
> -		wx->ptp_tx_skb = NULL;
> -		wx->tx_hwtstamp_errors++;
> -		clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state);
> +		struct sk_buff *ptp_tx_skb = NULL;
> +		unsigned long flags;
> +
> +		spin_lock_irqsave(&wx->ptp_tx_lock, flags);
> +		if (wx->ptp_tx_skb == skb) {
> +			ptp_tx_skb = wx->ptp_tx_skb;
> +			wx->ptp_tx_skb = NULL;
> +			clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state);
> +		}
> +		spin_unlock_irqrestore(&wx->ptp_tx_lock, flags);
> +
> +		if (ptp_tx_skb) {
> +			dev_kfree_skb_any(ptp_tx_skb);
> +			wx->tx_hwtstamp_errors++;
> +		}
>  	}

[Severity: Low]
This is a pre-existing pattern rather than something the patch changes, but
could wx->tx_hwtstamp_errors++ move inside the critical section that is
already held just above?

It is a non-atomic read-modify-write performed after dropping ptp_tx_lock,
and two transmits on different netdev queues share no HARD_TX_LOCK, so can
one increment be lost and make ts_stats->err in wx_get_ts_stats() undercount?
tx_hwtstamp_skipped is incremented under the lock in the arming hunk above.

> +	dev_kfree_skb_any(first->skb);
> +	first->skb = NULL;
>  
>  	return NETDEV_TX_OK;
>  }
> diff --git a/drivers/net/ethernet/wangxun/libwx/wx_ptp.c b/drivers/net/ethernet/wangxun/libwx/wx_ptp.c
> index 4708e7f3958f..f6346d257250 100644
> --- a/drivers/net/ethernet/wangxun/libwx/wx_ptp.c
> +++ b/drivers/net/ethernet/wangxun/libwx/wx_ptp.c
> @@ -129,6 +129,31 @@ static int wx_ptp_settime64(struct ptp_clock_info *ptp,
> +static struct sk_buff *__wx_ptp_detach_tx_skb(struct wx *wx)
> +{
> +	struct sk_buff *skb = wx->ptp_tx_skb;
> +
> +	lockdep_assert_held(&wx->ptp_tx_lock);
> +
> +	rd32ptp(wx, WX_TSC_1588_STMPH);
> +	wx->ptp_tx_skb = NULL;
> +	clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state);
> +
> +	return skb;
> +}

[Severity: Medium]
Is this rd32ptp() safe on the quiesce path?  With this patch,
wx_ptp_quiesce() reaches it through wx_ptp_clear_tx_timestamp(), and
wx_ptp_quiesce() only runs while the PCIe channel is known to be in
trouble:

wx_io_error_detected()          /* including pci_channel_io_frozen and
                                   pci_channel_io_perm_failure */
    set_bit(WX_FLAG_NEED_PCIE_RECOVERY, wx->flags);
    wx_soft_quiesce(wx)
        wx_ptp_quiesce(wx)
            wx_ptp_clear_tx_timestamp(wx)
                __wx_ptp_detach_tx_skb(wx)
                    rd32ptp(wx, WX_TSC_1588_STMPH);   /* readl() */

wx_reset_task() takes the same route after wx_check_pcie_error() has
already decided the link or memory space is gone.

Before this patch wx_ptp_quiesce() performed no register access at all
(it only freed wx->ptp_tx_skb and cleared the state bit), which also
matches commit c023e9769de9 ("net: wangxun: implement soft quiesce for
PCIe error recovery"):

    Function wx_soft_quiesce() provide a lightweight shutdown path during
    PCIe error recovery. It avoids MMIO-dependent operations in PCIe error
    status.

Since the read value is discarded and a dead device returns all-ones, the
unlatch this read exists for does not happen anyway.  Would an MMIO-free
detach for the quiesce path (take ptp_tx_lock, clear wx->ptp_tx_skb and
WX_STATE_PTP_TX_IN_PROGRESS) preserve that property, leaving the
WX_TSC_1588_STMPH unlatch to wx_ptp_suspend(), wx_ptp_tx_hang() and the
worker where the device is known to be alive?

>  /**
>   * wx_ptp_clear_tx_timestamp - utility function to clear Tx timestamp state
>   * @wx: the private board structure
> @@ -139,12 +164,14 @@ static int wx_ptp_settime64(struct ptp_clock_info *ptp,
>   */
>  static void wx_ptp_clear_tx_timestamp(struct wx *wx)
>  {
> -	rd32ptp(wx, WX_TSC_1588_STMPH);
> -	if (wx->ptp_tx_skb) {
> -		dev_kfree_skb_any(wx->ptp_tx_skb);
> -		wx->ptp_tx_skb = NULL;
> -	}
> -	clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state);
> +	struct sk_buff *skb;
> +	unsigned long flags;
> +
> +	spin_lock_irqsave(&wx->ptp_tx_lock, flags);
> +	skb = __wx_ptp_detach_tx_skb(wx);
> +	spin_unlock_irqrestore(&wx->ptp_tx_lock, flags);
> +
> +	dev_kfree_skb_any(skb);
>  }

[ ... ]

> @@ -849,11 +883,7 @@ void wx_ptp_quiesce(struct wx *wx)
>  	if (wx->ptp_clock)
>  		ptp_cancel_worker_sync(wx->ptp_clock);
>  
> -	if (wx->ptp_tx_skb) {
> -		dev_kfree_skb_any(wx->ptp_tx_skb);
> -		wx->ptp_tx_skb = NULL;
> -	}
> -	clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state);
> +	wx_ptp_clear_tx_timestamp(wx);
>  
>  	if (wx->ptp_clock) {
>  		ptp_clock_unregister(wx->ptp_clock);

[Severity: High]
This isn't a bug introduced by this patch, since wx_clean_tx_irq() and the
close ordering are untouched, but while the locking of the in-progress bit
is being reworked, is wx->ptp_clock safe against the NAPI Tx cleanup path?

wx_clean_tx_irq() reads the state bit without ptp_tx_lock and then hands
wx->ptp_clock to the scheduler:

	/* schedule check for Tx timestamp */
	if (unlikely(test_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state)) &&
	    skb_shinfo(tx_buffer->skb)->tx_flags & SKBTX_IN_PROGRESS)
		ptp_schedule_worker(wx->ptp_clock, 0);

and the callee dereferences it right away:

drivers/ptp/ptp_clock.c:ptp_schedule_worker() {
	return kthread_mod_delayed_work(ptp->kworker, &ptp->aux_work, delay);
}

The same clock teardown as in wx_ptp_quiesce() above is done by
wx_ptp_stop(), which runs before NAPI is disabled:

txgbe_close()
    wx_ptp_stop(wx);          /* wx_ptp_suspend(), ptp_clock_unregister(),
                                 wx->ptp_clock = NULL */
    txgbe_down(wx);
        txgbe_disable_device(wx);
            wx_napi_disable_all(wx);

ngbe_close() has the same wx_ptp_stop() before ngbe_down() ordering, and
netif_carrier_off() also happens only inside the disable step, so the
early "if (!netif_carrier_ok())" return in wx_clean_tx_irq() does not
help here.

Can a poll that already passed the test_bit() check then pass a freed or
NULL wx->ptp_clock to ptp_schedule_worker() and oops in softirq context?
Should NAPI be disabled before ptp_clock_unregister(), or should the clock
pointer used by the Tx cleanup path be synchronized with teardown?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/73D0D3F5D96A6928%2B20260914080020.211580-1-jiawenwu%40trustnetic.com

      parent reply	other threads:[~2026-09-17 20:02 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14  8:00 [PATCH net v2] net: libwx: fix races in Tx timestamp handling Jiawen Wu
2026-09-15 23:51 ` Jacob Keller
2026-09-16  2:33   ` Jiawen Wu
2026-09-17 20:02 ` netdev-bot+sashiko [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=178967534514.22033.6813523642554517754@kernel.org \
    --to=netdev-bot+sashiko@kernel.org \
    --cc=aleksandr.loktionov@intel.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=jacob.e.keller@intel.com \
    --cc=jiawenwu@trustnetic.com \
    --cc=kees@kernel.org \
    --cc=kuba@kernel.org \
    --cc=mengyuanlou@net-swift.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=richardcochran@gmail.com \
    --cc=sashiko-bot@kernel.org \
    --cc=vadim.fedorenko@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox