Netdev List
 help / color / mirror / Atom feed
From: Nagamani PV <nagamani@linux.ibm.com>
To: netdev-bot+sashiko@kernel.org
Cc: andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
	kuba@kernel.org, pabeni@redhat.com, wintera@linux.ibm.com,
	aswin@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com,
	agordeev@linux.ibm.com, borntraeger@linux.ibm.com,
	svens@linux.ibm.com, kees@kernel.org, linux-s390@vger.kernel.org,
	netdev@vger.kernel.org
Subject: Re: [PATCH net v4 2/2] s390/ctcm: Fix use-after-free in channel_remove()
Date: Mon, 5 Oct 2026 14:18:20 +0530	[thread overview]
Message-ID: <cc794003-f041-4075-a6be-3ba9f2854acb@linux.ibm.com> (raw)
In-Reply-To: <179109824657.434549.2788045083218312906@kernel.org>



On 04/10/26 12:47 PM, netdev-bot+sashiko@kernel.org wrote:
> [Severity: High]
> Do these waits run too late to protect the state that the channel
> callbacks use?
> 
> The only teardown caller of channel_remove() is ctcm_shutdown_device(),
> and it calls it after the net_device and its private state are gone:
> 
> drivers/s390/net/ctcm_main.c:ctcm_shutdown_device() {
>     ...
> 	if (dev) {
> 		unregister_netdev(dev);
> 		ctcm_free_netdevice(dev);
> 	}
>     ...
> 	ccw_device_set_offline(cgdev->cdev[1]);
> 	ccw_device_set_offline(cgdev->cdev[0]);
> 	channel_remove(priv->channel[CTCM_READ]);
> 	channel_remove(priv->channel[CTCM_WRITE]);
>     ...
> }
> 
> ctcm_free_netdevice() does kfree_fsm(grp->fsm), kfree(grp),
> priv->mpcg = NULL, kfree_fsm(priv->fsm) and priv->fsm = NULL. Under MODULE
> it also calls free_netdev(dev). The ccw_device_set_offline() calls in
> between can sleep, so a channel timer or tasklet can fire in that window.
> 
> The channel callbacks reach this state through ch->netdev->ml_priv:
> 
>   ctcm_chx_txretry() (ch->timer):
> 	fsm_event(priv->fsm, DEV_EVENT_TXDOWN, dev);
> 
>   ctcmpc_chx_send_sweep() (sweep_timer):
> 	if (grp->in_sweep == 0)
> 
>   ctcmpc_bh() (ch_tasklet):
> 	while ((fsm_getstate(grp->fsm) != MPCG_STATE_FLOWC) &&
> 
>   ctcm_ccw_check_rc()->ctcm_chx_iofatal():
> 	fsm_event(priv->fsm, ...);
> 
> The circular chain in the commit message also goes through priv->fsm or
> priv->mpcg. For example, mpc_action_go_inop() does this before it reaches
> tasklet_hi_schedule():
> 
> 	grp =  priv->mpcg;
> 	grp->flow_off_called = 0;
> 
> At channel_remove() time priv->fsm and priv->mpcg are already freed or
> NULL. Wouldn't that chain crash before it can re-arm anything?
> 
> If so, timer_shutdown_sync() only waits for a callback that is already
> dereferencing freed or NULL memory. tasklet_kill() also lets a queued
> ch_tasklet run against the freed grp.
> 
> channel_free() moves ch->fsm to CTC_STATE_IDLE, which has no handlers,
> so most late ch->fsm events do nothing. But that state change takes no
> lock. ctcm_irq_handler() runs FSM actions under the ccwdev lock, and for
> MPC ctcm_close() does nothing:
> 
> 	if (!IS_MPC(priv))
> 		fsm_event(priv->fsm, DEV_EVENT_STOP, dev);
> 
> As a result the MPC channels are not halted. An action that races
> channel_free() could leave the channel FSM in TXIDLE or TX with timers
> armed.
> 
> ctcmpc_bh() never checks the ch->fsm state, and ctc_mpc_flow_control()
> can schedule it at any time.
> 
> The ctcm_new_device() out_dev path uses the same order, but its channels
> were never started there.
> 
> Could both channels' timers and tasklets be stopped, and the subchannels
> taken offline, before ctcm_free_netdevice() and kfree_fsm(priv->fsm) run?
> That would mean moving the quiescing out of channel_remove() instead of
> reordering it inside channel_remove().
> 
> [Severity: High]
> This isn't a bug introduced by this patch, but the other fsm_timers in
> this driver seem to follow the same pattern: cancel without a sync, then
> free.
> 
> priv->restart_timer is bound to priv->fsm by fsm_settimer() in
> ctcm_init_netdevice(). grp->timer is bound to grp->fsm by fsm_settimer()
> in ctcmpc_init_mpc_group(). Neither timer is ever cancelled with a sync
> call. The only cancellations are fsm_deltimer(), which is timer_delete().
> 
> mpc_action_go_inop() always re-arms the restart timer:
> 
> 	fsm_deltimer(&priv->restart_timer);
> 	fsm_addtimer(&priv->restart_timer, 500, DEV_EVENT_RESTART, dev);
> 
> dev_action_restart() re-arms it for 1s in the MPC case. So during MPC
> group recovery the restart timer is usually pending.
> 
> For MPC, ctcm_close() does nothing. So ctcm_shutdown_device() then calls
> ctcm_free_netdevice(), which frees grp->fsm, grp (which contains
> grp->timer) and priv->fsm while these timers may be pending or running.
> After that, ctcm_remove_device() does kfree(priv), and priv contains
> restart_timer.
> 
> When one of these timers expires, fsm_expire_timer() calls
> fsm_event(this->fi, ...) on a freed fsm_instance. That reads the freed
> state and jumpmatrix, and calls a function pointer taken from freed
> memory. A timer still pending inside freed grp or priv memory can also
> corrupt the timer wheel.
> 
> There is a similar ordering problem in ctcm_free_netdevice():
> 
> 	if (grp->fsm)
> 		kfree_fsm(grp->fsm);
> 	dev_kfree_skb(grp->xid_skb);
> 	dev_kfree_skb(grp->rcvd_xid_skb);
> 	tasklet_kill(&grp->mpc_tasklet2);
> 	kfree(grp);
> 
> Can a pending mpc_tasklet2 run here after its FSM has been freed?
> 
> Should priv->restart_timer and grp->timer also get timer_shutdown_sync()
> before priv->fsm and grp are freed? And should mpc_tasklet2 be killed
> before kfree_fsm(grp->fsm)?
> 
> [ ... ]
> 

Both are confirmed pre-existing issues in ctcm_shutdown_device() and
ctcm_free_netdevice(), outside the scope of channel_remove(). They are
tracked for a separate follow-up fix.

Nagamani


  reply	other threads:[~2026-10-05  8:48 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  7:02 [PATCH net v4 0/2] s390/ctcm: Fix timer corruption and use-after-free Nagamani PV
2026-09-30  7:02 ` [PATCH net v4 1/2] s390/ctcm: Fix timer corruption in fsm_addtimer() Nagamani PV
2026-10-05 17:16   ` Simon Horman
2026-09-30  7:02 ` [PATCH net v4 2/2] s390/ctcm: Fix use-after-free in channel_remove() Nagamani PV
2026-10-01 11:35   ` Aswin Karuvally
2026-10-04  7:17   ` netdev-bot+sashiko
2026-10-05  8:48     ` Nagamani PV [this message]
2026-10-05 17:16   ` Simon Horman
2026-09-30  7:09 ` [PATCH net v4 0/2] s390/ctcm: Fix timer corruption and use-after-free netdev-bot+sinfo
2026-10-01  8:06   ` Nagamani PV
2026-10-06  0:34     ` Jakub Kicinski
2026-10-06  0:40 ` patchwork-bot+netdevbpf

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cc794003-f041-4075-a6be-3ba9f2854acb@linux.ibm.com \
    --to=nagamani@linux.ibm.com \
    --cc=agordeev@linux.ibm.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=aswin@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=kees@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=netdev-bot+sashiko@kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=svens@linux.ibm.com \
    --cc=wintera@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox