BPF List
 help / color / mirror / Atom feed
From: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
To: <netdev-bot+sashiko@kernel.org>
Cc: <anthony.l.nguyen@intel.com>, <davem@davemloft.net>,
	<kuba@kernel.org>, <pabeni@redhat.com>, <edumazet@google.com>,
	<andrew+netdev@lunn.ch>, <netdev@vger.kernel.org>,
	<zhaochenguang@kylinos.cn>, <magnus.karlsson@intel.com>,
	<jacob.e.keller@intel.com>, <przemyslaw.kitszel@intel.com>,
	<jbrandeb@kernel.org>, <horms@kernel.org>,
	<kerneljasonxing@gmail.com>, <ast@kernel.org>,
	<daniel@iogearbox.net>, <hawk@kernel.org>,
	<john.fastabend@gmail.com>, <sdf@fomichev.me>,
	<bpf@vger.kernel.org>, <sunithax.d.mekala@intel.com>
Subject: Re: [PATCH net 3/8] i40e: make ring pointers unreachable before freeing via rcu
Date: Fri, 25 Sep 2026 17:16:38 +0200	[thread overview]
Message-ID: <araQVrNwgbfhPKqe@boxer> (raw)
In-Reply-To: <179004066858.2160803.3751100830230297898@kernel.org>

On Tue, Sep 22, 2026 at 01:31:08AM +0000, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
> 
> Critical: 0 · High: 1 · Medium: 0 · Low: 0
> 
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
> 
> - [High] Incomplete RCU teardown: the ring-pointer arrays through which
>   RCU readers discover the rings are freed with a plain kfree() and…

this rcu'ing rings but not containers always felt a bit incomplete to me
and while i could say it is a pre-existing issue i think it will be a good
thing to address...unless it's gonna open up this Pandora's box further :<

> 
> --
> 
> i40e: make ring pointers unreachable before freeing via rcu
> 
> The patch reorders i40e_vsi_clear_rings() so the published slots in
> vsi->tx_rings/rx_rings/xdp_rings are cleared before the saved Tx ring
> pointer is handed to kfree_rcu(). The Tx pointer is saved first because it
> is the base of the per-queue-pair allocation block.
> 
> > diff --git a/drivers/net/ethernet/intel/i40e/i40e_main.c b/drivers/net/ethernet/intel/i40e/i40e_main.c
> > index de4c0737f72e..65aa50330aac 100644
> > --- a/drivers/net/ethernet/intel/i40e/i40e_main.c
> > +++ b/drivers/net/ethernet/intel/i40e/i40e_main.c
> > @@ -11693,11 +11693,13 @@ static void i40e_vsi_clear_rings(struct i40e_vsi *vsi)
> >  
> >  	if (vsi->tx_rings && vsi->tx_rings[0]) {
> >  		for (i = 0; i < vsi->alloc_queue_pairs; i++) {
> > -			kfree_rcu(vsi->tx_rings[i], rcu);
> > +			struct i40e_ring *tx_ring = vsi->tx_rings[i];
> > +
> >  			WRITE_ONCE(vsi->tx_rings[i], NULL);
> >  			WRITE_ONCE(vsi->rx_rings[i], NULL);
> >  			if (vsi->xdp_rings)
> >  				WRITE_ONCE(vsi->xdp_rings[i], NULL);
> > +			kfree_rcu(tx_ring, rcu);
> 
> [Severity: High]
> The ring objects are now unpublished before kfree_rcu(), but is the array
> that readers index through itself RCU protected?
> 
> The commit message states that after this change "new readers can no longer
> discover the ring through vsi->tx_rings/rx_rings/xdp_rings before the object
> is queued for deferred freeing".  Does that hold on the i40e_vsi_reinit_setup()
> path, where the arrays are dropped immediately afterwards?
> 
> drivers/net/ethernet/intel/i40e/i40e_main.c:i40e_vsi_reinit_setup() {
> 	i40e_put_lump(pf->qp_pile, vsi->base_queue, vsi->idx);
> 	i40e_vsi_clear_rings(vsi);
> 
> 	i40e_vsi_free_arrays(vsi, false);
> 	i40e_set_num_rings_in_vsi(vsi);
> 	ret = i40e_vsi_alloc_arrays(vsi, false);
> 	...
> }
> 
> and i40e_vsi_free_arrays() uses a plain kfree() plus plain stores, with no
> grace period:
> 
> drivers/net/ethernet/intel/i40e/i40e_main.c:i40e_vsi_free_arrays() {
> 	...
> 	kfree(vsi->tx_rings);
> 	vsi->tx_rings = NULL;
> 	vsi->rx_rings = NULL;
> 	vsi->xdp_rings = NULL;
> }
> 
> Meanwhile i40e_get_netdev_stats_struct() (ndo_get_stats64, reached from
> dev_get_stats() without rtnl) re-evaluates the non-volatile array base on
> every iteration inside the RCU section, and never checks the rx_rings or
> xdp_rings bases at all:
> 
> drivers/net/ethernet/intel/i40e/i40e_main.c:i40e_get_netdev_stats_struct() {
> 	if (!vsi->tx_rings)
> 		return;
> 
> 	rcu_read_lock();
> 	for (i = 0; i < vsi->num_queue_pairs; i++) {
> 		...
> 		ring = READ_ONCE(vsi->tx_rings[i]);
> 		...
> 		if (i40e_enabled_xdp_vsi(vsi)) {
> 			ring = READ_ONCE(vsi->xdp_rings[i]);
> 	...
> }
> 
> i40e_update_vsi_stats() does the same with no base check whatsoever:
> 
> drivers/net/ethernet/intel/i40e/i40e_main.c:i40e_update_vsi_stats() {
> 	rcu_read_lock();
> 	for (q = 0; q < vsi->num_queue_pairs; q++) {
> 		/* locate Tx ring */
> 		p = READ_ONCE(vsi->tx_rings[q]);
> 	...
> }
> 
> So can a reader that is already past the "if (!vsi->tx_rings)" test, or
> inside the loop, reload vsi->tx_rings/rx_rings/xdp_rings after
> i40e_vsi_free_arrays() nulls them and dereference NULL[i]?
> 
> And once the pointer-array slab is recycled, for instance by the
> i40e_vsi_alloc_arrays() call two statements later, or by the SLUB freelist
> pointer written into the freed object, can the same reader read a non-NULL
> wild value out of the freed array and dereference it as struct i40e_ring
> (->syncp, ->stats)?
> 
> The __I40E_VSI_DOWN test at the top of both readers is evaluated before
> rcu_read_lock(), and nothing on this path drains readers that already passed
> it, so it does not appear to close the window.  Should the arrays be released
> through kfree_rcu() (or after a synchronize_rcu()), with the base loaded once
> via rcu_dereference(), for the unpublish ordering established here to actually
> take effect?
> 
> Separately, both readers bound their loops with vsi->num_queue_pairs, which
> i40e_set_num_rings_in_vsi() changes between i40e_vsi_free_arrays() and
> i40e_vsi_alloc_arrays().  Can that let a reader index past the end of the
> freshly allocated array?
> 
> >  		}
> >  	}
> >  }
> 
> -- 
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260918212458.550425-1-anthony.l.nguyen%40intel.com

  reply	other threads:[~2026-09-25 15:17 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-18 21:24 [PATCH net 0/8][pull request] Intel Wired LAN Driver Updates 2026-09-18 (i40e) Tony Nguyen
2026-09-18 21:24 ` [PATCH net 1/8] i40e: unregister netdev before clearing VSI on reinit failure Tony Nguyen
2026-09-22  1:31   ` netdev-bot+sashiko
2026-09-25 13:19     ` Maciej Fijalkowski
2026-09-18 21:24 ` [PATCH net 2/8] i40e: avoid null ptr dereference in i40e_ptp_stop() Tony Nguyen
2026-09-22  1:31   ` netdev-bot+sashiko
2026-09-25 13:30     ` Maciej Fijalkowski
2026-09-18 21:24 ` [PATCH net 3/8] i40e: make ring pointers unreachable before freeing via rcu Tony Nguyen
2026-09-22  1:31   ` netdev-bot+sashiko
2026-09-25 15:16     ` Maciej Fijalkowski [this message]
2026-09-18 21:24 ` [PATCH net 4/8] i40e: avoid deadlock when calling unregister_netdev() Tony Nguyen
2026-09-22  1:31   ` netdev-bot+sashiko
2026-09-25 16:59     ` Maciej Fijalkowski
2026-09-18 21:24 ` [PATCH net 5/8] i40e: fix potential UAF in i40e_vsi_setup()'s error path Tony Nguyen
2026-09-18 21:24 ` [PATCH net 6/8] i40e: do not expose netdev too early Tony Nguyen
2026-09-22  1:31   ` netdev-bot+sashiko
2026-09-18 21:24 ` [PATCH net 7/8] i40e: keep q_vectors array in sync with channel count changes Tony Nguyen
2026-09-22  1:31   ` netdev-bot+sashiko
2026-09-18 21:24 ` [PATCH net 8/8] i40e: xsk: fix multi-buffer XDP_PASS skb construction Tony Nguyen
2026-09-24 11:23 ` [PATCH net 0/8][pull request] Intel Wired LAN Driver Updates 2026-09-18 (i40e) Paolo Abeni
2026-09-24 11:25   ` Paolo Abeni
2026-09-25 12:46     ` Maciej Fijalkowski
2026-09-25 20:00       ` Jakub Kicinski
2026-09-26 12:21         ` Maciej Fijalkowski

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=araQVrNwgbfhPKqe@boxer \
    --to=maciej.fijalkowski@intel.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=anthony.l.nguyen@intel.com \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=hawk@kernel.org \
    --cc=horms@kernel.org \
    --cc=jacob.e.keller@intel.com \
    --cc=jbrandeb@kernel.org \
    --cc=john.fastabend@gmail.com \
    --cc=kerneljasonxing@gmail.com \
    --cc=kuba@kernel.org \
    --cc=magnus.karlsson@intel.com \
    --cc=netdev-bot+sashiko@kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=przemyslaw.kitszel@intel.com \
    --cc=sdf@fomichev.me \
    --cc=sunithax.d.mekala@intel.com \
    --cc=zhaochenguang@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox