Linux CAN drivers development
 help / color / mirror / Atom feed
* [PATCH v2] tty: ldisc: fix deadlock between ldisc_sem and rtnl_mutex
@ 2026-07-21  8:50 Yun Zhou
  2026-07-22  8:51 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Yun Zhou @ 2026-07-21  8:50 UTC (permalink / raw)
  To: gregkh, jirislaby, socketcan
  Cc: linux-serial, mkl, linux-can, davem, edumazet, kuba, pabeni,
	horms, netdev, linux-kernel, yun.zhou

syzbot reported a circular lock dependency involving tty ldisc_sem and
the networking rtnl_mutex. The full chain is:

  rtnl_mutex --> nft_commit_mutex --> ... --> ep->mtx --> ldisc_sem --> rtnl_mutex

The last edge (ldisc_sem -> rtnl_mutex) is created because tty line
discipline .open() callbacks (slcan, slip) call register_netdev() which
acquires rtnl_mutex, and .open() runs under ldisc_sem write lock in
tty_set_ldisc().

Fix by moving the .open() call outside the ldisc_sem write lock. The
ldisc .open() is initialization of the NEW discipline after the old one
has been closed - there is no need for ldisc_sem protection at this
point since:

 - tty_lock is held throughout, preventing concurrent tty_set_ldisc,
   hangup, or close
 - tty->ldisc is set to NULL during the window. tty_ldisc_ref_wait()
   waits for the transition to complete. tty_ldisc_ref() returns NULL
   which callers already handle.
 - tty buffer data stays queued until the ldisc is installed

The sequence becomes:
  1. Hold ldisc_sem(write): close old ldisc, set tty->ldisc = NULL
  2. Release ldisc_sem(write)
  3. Call new_ldisc->ops->open() without ldisc_sem
  4. Re-acquire ldisc_sem(write): install new ldisc (or restore old)
  5. Release ldisc_sem(write)

Reported-by: syzbot+de610eeef174bd59a8a3@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=de610eeef174bd59a8a3
Signed-off-by: Yun Zhou <yun.zhou@windriver.com>
---
v2:
 - Keep user-visible behavior unchanged: tty_ldisc_ref_wait() now waits
   for the ldisc transition to complete instead of returning NULL (which
   would cause spurious EOF/-EIO to concurrent readers).
 - Fix a race between tty_ldisc_ref_wait() and __tty_hangup() where a
   reader could block forever if it observed ldisc==NULL before
   TTY_HUPPED was set. Add wake_up() after set_bit(TTY_HUPPED).

 drivers/tty/tty_io.c    |  7 +++++++
 drivers/tty/tty_ldisc.c | 34 +++++++++++++++++++++++++++++++---
 2 files changed, 38 insertions(+), 3 deletions(-)

diff --git a/drivers/tty/tty_io.c b/drivers/tty/tty_io.c
index 6b283fd03ff8..e2f82e80f397 100644
--- a/drivers/tty/tty_io.c
+++ b/drivers/tty/tty_io.c
@@ -649,6 +649,13 @@ static void __tty_hangup(struct tty_struct *tty, int exit_session)
 	 */
 	set_bit(TTY_HUPPED, &tty->flags);
 	clear_bit(TTY_HUPPING, &tty->flags);
+
+	/*
+	 * Wake up readers blocked in tty_ldisc_ref_wait() that may have
+	 * seen ldisc == NULL but not yet TTY_HUPPED.
+	 */
+	wake_up(&tty->read_wait);
+
 	tty_unlock(tty);
 
 	if (f)
diff --git a/drivers/tty/tty_ldisc.c b/drivers/tty/tty_ldisc.c
index 27fe8236f662..bd94a1f13c44 100644
--- a/drivers/tty/tty_ldisc.c
+++ b/drivers/tty/tty_ldisc.c
@@ -242,8 +242,20 @@ struct tty_ldisc *tty_ldisc_ref_wait(struct tty_struct *tty)
 
 	ldsem_down_read(&tty->ldisc_sem, MAX_SCHEDULE_TIMEOUT);
 	ld = tty->ldisc;
-	if (!ld)
+	if (!ld) {
 		ldsem_up_read(&tty->ldisc_sem);
+
+		/* ldisc may be NULL during a discipline switch; wait and retry */
+		if (!test_bit(TTY_HUPPED, &tty->flags)) {
+			wait_event(tty->read_wait,
+				   READ_ONCE(tty->ldisc) != NULL ||
+				   test_bit(TTY_HUPPED, &tty->flags));
+			ldsem_down_read(&tty->ldisc_sem, MAX_SCHEDULE_TIMEOUT);
+			ld = tty->ldisc;
+			if (!ld)
+				ldsem_up_read(&tty->ldisc_sem);
+		}
+	}
 	return ld;
 }
 EXPORT_SYMBOL_GPL(tty_ldisc_ref_wait);
@@ -556,15 +568,28 @@ int tty_set_ldisc(struct tty_struct *tty, int disc)
 	/* Shutdown the old discipline. */
 	tty_ldisc_close(tty, old_ldisc);
 
-	/* Now set up the new line discipline. */
-	tty->ldisc = new_ldisc;
+	/* Clear tty->ldisc so concurrent readers back off during transition */
+	tty->ldisc = NULL;
 	tty_set_termios_ldisc(tty, disc);
+	tty_ldisc_unlock(tty);
 
+	/*
+	 * Open the new discipline outside ldisc_sem. The ldisc .open()
+	 * may acquire locks (e.g., rtnl_mutex) that would create circular
+	 * dependencies if taken under ldisc_sem. tty_lock is still held,
+	 * preventing concurrent ldisc changes and hangup.
+	 */
 	retval = tty_ldisc_open(tty, new_ldisc);
+
+	tty_ldisc_lock(tty, MAX_SCHEDULE_TIMEOUT);
+
 	if (retval < 0) {
 		/* Back to the old one or N_TTY if we can't */
 		tty_ldisc_put(new_ldisc);
 		tty_ldisc_restore(tty, old_ldisc);
+	} else {
+		/* Success - install new ldisc */
+		tty->ldisc = new_ldisc;
 	}
 
 	if (tty->ldisc->ops->num != old_ldisc->ops->num && tty->ops->set_ldisc) {
@@ -584,6 +609,9 @@ int tty_set_ldisc(struct tty_struct *tty, int disc)
 out:
 	tty_ldisc_unlock(tty);
 
+	/* Wake up readers waiting for the ldisc transition to complete */
+	wake_up(&tty->read_wait);
+
 	/*
 	 * Restart the work queue in case no characters kick it off. Safe if
 	 * already running
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH v2] tty: ldisc: fix deadlock between ldisc_sem and rtnl_mutex
  2026-07-21  8:50 [PATCH v2] tty: ldisc: fix deadlock between ldisc_sem and rtnl_mutex Yun Zhou
@ 2026-07-22  8:51 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-07-22  8:51 UTC (permalink / raw)
  To: Yun Zhou
  Cc: Oliver Hartkopp, Oleksij Rempel, linux-can, Marc Kleine-Budde,
	Vincent Mailhol

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

New issues:
- [Medium] A Time-of-Check to Time-of-Use (TOCTOU) race condition in `tty_ldisc_ref_wait` can cause spurious EOF or -EIO errors for concurrent readers.

Pre-existing issues:
- [High] The circular lock dependency (ldisc_sem -> rtnl_mutex) is not fully resolved in error fallback, hangup, and teardown paths.
--

--- Patch [1]: [PATCH v2] tty: ldisc: fix deadlock between ldisc_sem and rtnl_mutex ---
commit aa2b45f5a2e0cd2e80f19ce89583432e4e6eb9e4
Author: Yun Zhou <yun.zhou@windriver.com>

tty: ldisc: fix deadlock between ldisc_sem and rtnl_mutex

syzbot reported a circular lock dependency involving tty ldisc_sem and
the networking rtnl_mutex. Fix by moving the .open() call outside the
ldisc_sem write lock.

Closes: https://syzkaller.appspot.com/bug?extid=de610eeef174bd59a8a3

> diff --git a/drivers/tty/tty_ldisc.c b/drivers/tty/tty_ldisc.c
> index 27fe8236f6623..bd94a1f13c44c 100644
> --- a/drivers/tty/tty_ldisc.c
> +++ b/drivers/tty/tty_ldisc.c
> @@ -242,8 +242,20 @@ struct tty_ldisc *tty_ldisc_ref_wait(struct tty_struct *tty)
>  
>  	ldsem_down_read(&tty->ldisc_sem, MAX_SCHEDULE_TIMEOUT);
>  	ld = tty->ldisc;
> -	if (!ld)
> +	if (!ld) {
>  		ldsem_up_read(&tty->ldisc_sem);
> +
> +		/* ldisc may be NULL during a discipline switch; wait and retry */
> +		if (!test_bit(TTY_HUPPED, &tty->flags)) {

[Severity: Medium]
Could this if statement lead to a Time-of-Check to Time-of-Use (TOCTOU)
race condition returning spurious EOF or -EIO to concurrent readers?

If a reader thread waits here for a discipline switch to finish, the writer
finishes, wakes up the reader, and drops tty_lock. Before the reader can
acquire ldisc_sem(read) again, a third thread could initiate another line
discipline switch, acquiring tty_lock and setting tty->ldisc = NULL.

Since this uses an if statement instead of a while loop, wouldn't the reader
see NULL again and incorrectly return NULL without retrying?

> +			wait_event(tty->read_wait,
> +				   READ_ONCE(tty->ldisc) != NULL ||
> +				   test_bit(TTY_HUPPED, &tty->flags));
> +			ldsem_down_read(&tty->ldisc_sem, MAX_SCHEDULE_TIMEOUT);
> +			ld = tty->ldisc;
> +			if (!ld)
> +				ldsem_up_read(&tty->ldisc_sem);
> +		}
> +	}
>  	return ld;
>  }

[ ... ]

> @@ -556,15 +568,28 @@ int tty_set_ldisc(struct tty_struct *tty, int disc)
>  	/* Shutdown the old discipline. */
>  	tty_ldisc_close(tty, old_ldisc);

[Severity: High]
This is a pre-existing issue, but doesn't the circular lock dependency
still exist in the teardown, hangup, and error fallback paths?

Switching away from a network line discipline (like slcan or slip) calls
.close() here while ldisc_sem is held for write. The .close() callback calls
unregister_netdev(), acquiring rtnl_mutex.

Similarly, if opening a new discipline fails, the fallback tty_ldisc_failto()
calls .open() under ldisc_sem, recreating the exact same deadlock. The same
happens in tty_ldisc_reinit().

> -	/* Now set up the new line discipline. */
> -	tty->ldisc = new_ldisc;
> +	/* Clear tty->ldisc so concurrent readers back off during transition */
> +	tty->ldisc = NULL;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260721085035.2485657-1-yun.zhou@windriver.com?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-07-22  8:52 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-21  8:50 [PATCH v2] tty: ldisc: fix deadlock between ldisc_sem and rtnl_mutex Yun Zhou
2026-07-22  8:51 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox