* [PATCH net v2] tls: fix the open record check in the max payload size setsockopt
@ 2026-09-01 7:29 Jiayuan Chen
2026-09-03 9:06 ` Paolo Abeni
0 siblings, 1 reply; 3+ messages in thread
From: Jiayuan Chen @ 2026-09-01 7:29 UTC (permalink / raw)
To: netdev
Cc: Jiayuan Chen, John Fastabend, Jakub Kicinski, Sabrina Dubroca,
David S. Miller, Eric Dumazet, Paolo Abeni, Simon Horman,
Wilfred Mallawa, linux-kernel
do_tls_setsockopt_tx_payload_len() refuses to resize records while one is
open, but the check is broken in two ways.
It reaches for the open record through tls_sw_ctx_tx(), an unchecked cast
of ctx->priv_ctx_tx. Under device offload that pointer is a
tls_offload_context_tx, so the check reads a field of the wrong struct: it
returns EBUSY on whatever happens to be there, and never sees the record
that really is open. Dispatch on tx_conf.
The socket lock alone is also not enough to tell no sender is in flight.
The device tx path drops it in sk_stream_wait_memory() after
tls_push_record() cleared open_record, so setsockopt can slip in
mid-sendmsg and shrink the limit. tls_push_data() latches the limit once
per call, and a record left open by MSG_MORE can then be bigger than the
new limit, making the u32 subtraction in the copy clamp wrap around and
build an oversized record. Take tx_lock like the sendmsg paths do so
setsockopt waits for in-flight senders. Re-reading the limit inside the
tls_push_data() loop could close the wrap too, but a parked sender would
still finish with the old limit, so serializing on tx_lock is simpler.
Fixes: 82cb5be6ad64 ("net/tls: support setting the maximum payload size")
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
---
v1 -> v2:
- also take tx_lock in setsockopt, the socket lock alone can't stop a
device sender parked in sk_stream_wait_memory() (AI review on v1)
v1: https://lore.kernel.org/netdev/20260831025607.62927-1-jiayuan.chen@linux.dev/
base on my netdevsim + tls (in progress)
https://lore.kernel.org/netdev/20260728125658.390500-1-jiayuan.chen@linux.dev/
Previous finding:
b17cf742eaad ("tls: device: fix out-of-bounds write in tls_append_frag()")
---
net/tls/tls_main.c | 27 +++++++++++++++++++++++++--
1 file changed, 25 insertions(+), 2 deletions(-)
diff --git a/net/tls/tls_main.c b/net/tls/tls_main.c
index fbb274287aa5..114e5113bafd 100644
--- a/net/tls/tls_main.c
+++ b/net/tls/tls_main.c
@@ -833,15 +833,29 @@ static int do_tls_setsockopt_no_pad(struct sock *sk, sockptr_t optval,
return rc;
}
+/* priv_ctx_tx holds a different structure on each TX path, so tx_conf has to
+ * say which open record to look at.
+ */
+static bool tls_tx_record_is_open(struct tls_context *ctx)
+{
+ switch (ctx->tx_conf) {
+ case TLS_SW:
+ return !!tls_sw_ctx_tx(ctx)->open_rec;
+ case TLS_HW:
+ return !!tls_offload_ctx_tx(ctx)->open_record;
+ default:
+ return false;
+ }
+}
+
static int do_tls_setsockopt_tx_payload_len(struct sock *sk, sockptr_t optval,
unsigned int optlen)
{
struct tls_context *ctx = tls_get_ctx(sk);
- struct tls_sw_context_tx *sw_ctx = tls_sw_ctx_tx(ctx);
u16 value;
bool tls_13 = ctx->prot_info.version == TLS_1_3_VERSION;
- if (sw_ctx && sw_ctx->open_rec)
+ if (tls_tx_record_is_open(ctx))
return -EBUSY;
if (sockptr_is_null(optval) || optlen != sizeof(value))
@@ -862,6 +876,7 @@ static int do_tls_setsockopt_tx_payload_len(struct sock *sk, sockptr_t optval,
static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
unsigned int optlen)
{
+ struct tls_context *ctx;
int rc = 0;
switch (optname) {
@@ -881,9 +896,17 @@ static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
rc = do_tls_setsockopt_no_pad(sk, optval, optlen);
break;
case TLS_TX_MAX_PAYLOAD_LEN:
+ /* Take tx_lock like the sendmsg paths do, the socket lock is
+ * dropped while a sender waits for memory, with no record open.
+ */
+ ctx = tls_get_ctx(sk);
+ rc = mutex_lock_interruptible(&ctx->tx_lock);
+ if (rc)
+ break;
lock_sock(sk);
rc = do_tls_setsockopt_tx_payload_len(sk, optval, optlen);
release_sock(sk);
+ mutex_unlock(&ctx->tx_lock);
break;
default:
rc = -ENOPROTOOPT;
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [PATCH net v2] tls: fix the open record check in the max payload size setsockopt
2026-09-01 7:29 [PATCH net v2] tls: fix the open record check in the max payload size setsockopt Jiayuan Chen
@ 2026-09-03 9:06 ` Paolo Abeni
2026-09-03 9:40 ` Jiayuan Chen
0 siblings, 1 reply; 3+ messages in thread
From: Paolo Abeni @ 2026-09-03 9:06 UTC (permalink / raw)
To: Jiayuan Chen, netdev
Cc: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
Eric Dumazet, Simon Horman, Wilfred Mallawa, linux-kernel
On 9/1/26 9:29 AM, Jiayuan Chen wrote:
> @@ -862,6 +876,7 @@ static int do_tls_setsockopt_tx_payload_len(struct sock *sk, sockptr_t optval,
> static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
> unsigned int optlen)
> {
> + struct tls_context *ctx;
> int rc = 0;
>
> switch (optname) {
> @@ -881,9 +896,17 @@ static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
> rc = do_tls_setsockopt_no_pad(sk, optval, optlen);
> break;
> case TLS_TX_MAX_PAYLOAD_LEN:
> + /* Take tx_lock like the sendmsg paths do, the socket lock is
> + * dropped while a sender waits for memory, with no record open.
> + */
> + ctx = tls_get_ctx(sk);
> + rc = mutex_lock_interruptible(&ctx->tx_lock);
> + if (rc)
Why using the interruptible variant? the blocking lock just after will
still ignore signals, and this sockopt will now surprisingly fail if a
signal happens at the wrong time.
/P
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH net v2] tls: fix the open record check in the max payload size setsockopt
2026-09-03 9:06 ` Paolo Abeni
@ 2026-09-03 9:40 ` Jiayuan Chen
0 siblings, 0 replies; 3+ messages in thread
From: Jiayuan Chen @ 2026-09-03 9:40 UTC (permalink / raw)
To: Paolo Abeni, netdev
Cc: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
Eric Dumazet, Simon Horman, Wilfred Mallawa, linux-kernel
在 9/3/26 5:06 PM, Paolo Abeni 写道:
> On 9/1/26 9:29 AM, Jiayuan Chen wrote:
>> @@ -862,6 +876,7 @@ static int do_tls_setsockopt_tx_payload_len(struct sock *sk, sockptr_t optval,
>> static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
>> unsigned int optlen)
>> {
>> + struct tls_context *ctx;
>> int rc = 0;
>>
>> switch (optname) {
>> @@ -881,9 +896,17 @@ static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
>> rc = do_tls_setsockopt_no_pad(sk, optval, optlen);
>> break;
>> case TLS_TX_MAX_PAYLOAD_LEN:
>> + /* Take tx_lock like the sendmsg paths do, the socket lock is
>> + * dropped while a sender waits for memory, with no record open.
>> + */
>> + ctx = tls_get_ctx(sk);
>> + rc = mutex_lock_interruptible(&ctx->tx_lock);
>> + if (rc)
> Why using the interruptible variant? the blocking lock just after will
> still ignore signals, and this sockopt will now surprisingly fail if a
> signal happens at the wrong time.
>
> /P
Hi Paolo,
The two waits are very different. The xmit path holds tx_lock across
sk_stream_wait_memory(), which can sleep for an undetermined time (until
the peer reads):
tls_device_sendmsg()
mutex_lock(&tls_ctx->tx_lock);
lock_sock(sk);
tls_push_data()
sk_stream_wait_memory() <- releases sk lock, keeps tx_lock
release_sock(sk);
mutex_unlock(&tls_ctx->tx_lock);
So waiting for tx_lock with plain mutex_lock() can leave the process in
D state for a long time. The lock_sock() after it is fine: the sleeping
sender drops the socket lock, so that wait is only for short critical
sections, never across the long sleep.
tls_sw_sendmsg() also uses mutex_lock_interruptible but
tls_device_sendmsg() still
uses plain mutex_lock() indeed, but that's another topic.
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-03 9:40 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-01 7:29 [PATCH net v2] tls: fix the open record check in the max payload size setsockopt Jiayuan Chen
2026-09-03 9:06 ` Paolo Abeni
2026-09-03 9:40 ` Jiayuan Chen
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).