Netdev List
 help / color / mirror / Atom feed
* [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support
@ 2026-09-17 22:35 Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers Rishikesh Jethwani
                   ` (14 more replies)
  0 siblings, 15 replies; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Hi all,

  This series adds TLS 1.3 hardware offload support including KeyUpdate
  (rekey) and a selftest for validation.

Changes in v17:
  - Addressed review comments

Rishikesh Jethwani (15):
  net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers
  net/mlx5e: add TLS 1.3 hardware offload support
  tls: reject rekey attempts on an existing HW-offloaded connection
  tls: add TLS 1.3 hardware offload support
  tls: split tls_set_sw_offload into init and finalize stages
  tls: prep helpers and refactors for HW offload KeyUpdate
  net: sched: re-validate parked decrypted skbs on requeue
  tcp: fence collapse against rtx-queue tail when write queue is empty
  net: skbuff: add skb->decrypt_failed bit
  net/mlx5e: flag TLS RX records that failed device decryption
  tls: device: add TX KeyUpdate support
  tls: device: add RX KeyUpdate support
  tls: device: add tracepoints for the KeyUpdate path
  selftests: net: add TLS hardware offload test
  tls: document TLS 1.3 hardware offload rekey handling

 Documentation/networking/tls-offload.rst      |  175 ++-
 Documentation/networking/tls.rst              |   17 +
 MAINTAINERS                                   |    2 +
 .../chelsio/inline_crypto/ch_ktls/chcr_ktls.c |    3 +
 .../mellanox/mlx5/core/en_accel/ktls.h        |    8 +-
 .../mellanox/mlx5/core/en_accel/ktls_rx.c     |   13 +-
 .../mellanox/mlx5/core/en_accel/ktls_txrx.c   |   14 +-
 .../net/ethernet/netronome/nfp/crypto/tls.c   |    3 +
 include/linux/skbuff.h                        |    6 +
 include/net/tcp.h                             |    9 +
 include/net/tls.h                             |  149 +-
 include/uapi/linux/snmp.h                     |    6 +
 net/sched/sch_generic.c                       |    9 +
 net/tls/tls.h                                 |   32 +-
 net/tls/tls_device.c                          | 1381 +++++++++++++++--
 net/tls/tls_device_fallback.c                 |  186 ++-
 net/tls/tls_main.c                            |   87 +-
 net/tls/tls_proc.c                            |    6 +
 net/tls/tls_sw.c                              |  191 ++-
 net/tls/trace.h                               |  118 ++
 .../selftests/drivers/net/hw/.gitignore       |    1 +
 .../testing/selftests/drivers/net/hw/Makefile |    2 +
 tools/testing/selftests/drivers/net/hw/config |    2 +
 .../selftests/drivers/net/hw/tls_hw_offload.c | 1132 ++++++++++++++
 .../drivers/net/hw/tls_hw_offload.py          |  446 ++++++
 25 files changed, 3735 insertions(+), 263 deletions(-)
 create mode 100644 tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
 create mode 100755 tools/testing/selftests/drivers/net/hw/tls_hw_offload.py

-- 
2.50.1


^ permalink raw reply	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:55   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 02/15] net/mlx5e: add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (13 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

These drivers only support TLS 1.2. Return early when TLS 1.3
is requested to prevent unsupported hardware offload attempts.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c | 3 +++
 drivers/net/ethernet/netronome/nfp/crypto/tls.c                | 3 +++
 2 files changed, 6 insertions(+)

diff --git a/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c b/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
index f5acd4be1e69..29e108ce6764 100644
--- a/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
+++ b/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
@@ -431,6 +431,9 @@ static int chcr_ktls_dev_add(struct net_device *netdev, struct sock *sk,
 	atomic64_inc(&port_stats->ktls_tx_connection_open);
 	u_ctx = adap->uld[CXGB4_ULD_KTLS].handle;
 
+	if (crypto_info->version != TLS_1_2_VERSION)
+		goto out;
+
 	if (direction == TLS_OFFLOAD_CTX_DIR_RX) {
 		pr_err("not expecting for RX direction\n");
 		goto out;
diff --git a/drivers/net/ethernet/netronome/nfp/crypto/tls.c b/drivers/net/ethernet/netronome/nfp/crypto/tls.c
index 9983d7aa2b9c..13864c6a55dc 100644
--- a/drivers/net/ethernet/netronome/nfp/crypto/tls.c
+++ b/drivers/net/ethernet/netronome/nfp/crypto/tls.c
@@ -287,6 +287,9 @@ nfp_net_tls_add(struct net_device *netdev, struct sock *sk,
 	BUILD_BUG_ON(offsetof(struct nfp_net_tls_offload_ctx, rx_end) >
 		     TLS_DRIVER_STATE_SIZE_RX);
 
+	if (crypto_info->version != TLS_1_2_VERSION)
+		return -EOPNOTSUPP;
+
 	if (!nfp_net_cipher_supported(nn, crypto_info->cipher_type, direction))
 		return -EOPNOTSUPP;
 
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 02/15] net/mlx5e: add TLS 1.3 hardware offload support
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:55   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 03/15] tls: reject rekey attempts on an existing HW-offloaded connection Rishikesh Jethwani
                   ` (12 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Enable TLS 1.3 TX/RX hardware offload on ConnectX-6 Dx and newer
crypto-enabled adapters.
Key changes:
- Add TLS 1.3 capability checking and version validation
- Use MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_3 (0x3) for crypto context
- Handle TLS 1.3 IV format: full 12-byte IV copied to gcm_iv +
  implicit_iv (vs TLS 1.2's 4-byte salt only)

Tested with TLS 1.3 AES-GCM-128 and AES-GCM-256 cipher suites.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
Tested-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
---
 .../ethernet/mellanox/mlx5/core/en_accel/ktls.h    |  8 +++++++-
 .../mellanox/mlx5/core/en_accel/ktls_txrx.c        | 14 +++++++++++---
 2 files changed, 18 insertions(+), 4 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h
index 07a04a142a2e..0469ca6a0762 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h
@@ -30,7 +30,9 @@ static inline bool mlx5e_is_ktls_device(struct mlx5_core_dev *mdev)
 		return false;
 
 	return (MLX5_CAP_TLS(mdev, tls_1_2_aes_gcm_128) ||
-		MLX5_CAP_TLS(mdev, tls_1_2_aes_gcm_256));
+		MLX5_CAP_TLS(mdev, tls_1_2_aes_gcm_256) ||
+		MLX5_CAP_TLS(mdev, tls_1_3_aes_gcm_128) ||
+		MLX5_CAP_TLS(mdev, tls_1_3_aes_gcm_256));
 }
 
 static inline bool mlx5e_ktls_type_check(struct mlx5_core_dev *mdev,
@@ -40,10 +42,14 @@ static inline bool mlx5e_ktls_type_check(struct mlx5_core_dev *mdev,
 	case TLS_CIPHER_AES_GCM_128:
 		if (crypto_info->version == TLS_1_2_VERSION)
 			return MLX5_CAP_TLS(mdev,  tls_1_2_aes_gcm_128);
+		else if (crypto_info->version == TLS_1_3_VERSION)
+			return MLX5_CAP_TLS(mdev,  tls_1_3_aes_gcm_128);
 		break;
 	case TLS_CIPHER_AES_GCM_256:
 		if (crypto_info->version == TLS_1_2_VERSION)
 			return MLX5_CAP_TLS(mdev,  tls_1_2_aes_gcm_256);
+		else if (crypto_info->version == TLS_1_3_VERSION)
+			return MLX5_CAP_TLS(mdev,  tls_1_3_aes_gcm_256);
 		break;
 	}
 
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c
index 570a912dd6fa..f3f1be1d4034 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c
@@ -6,6 +6,7 @@
 
 enum {
 	MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_2 = 0x2,
+	MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_3 = 0x3,
 };
 
 enum {
@@ -15,8 +16,10 @@ enum {
 #define EXTRACT_INFO_FIELDS do { \
 	salt    = info->salt;    \
 	rec_seq = info->rec_seq; \
+	iv      = info->iv;      \
 	salt_sz    = sizeof(info->salt);    \
 	rec_seq_sz = sizeof(info->rec_seq); \
+	iv_sz      = sizeof(info->iv);      \
 } while (0)
 
 static void
@@ -24,9 +27,9 @@ fill_static_params(struct mlx5_wqe_tls_static_params_seg *params,
 		   union mlx5e_crypto_info *crypto_info,
 		   u32 key_id, u32 resync_tcp_sn)
 {
+	u16 salt_sz, rec_seq_sz, iv_sz;
+	char *salt, *rec_seq, *iv;
 	char *initial_rn, *gcm_iv;
-	u16 salt_sz, rec_seq_sz;
-	char *salt, *rec_seq;
 	u8 tls_version;
 	u8 *ctx;
 
@@ -59,7 +62,12 @@ fill_static_params(struct mlx5_wqe_tls_static_params_seg *params,
 	memcpy(gcm_iv,      salt,    salt_sz);
 	memcpy(initial_rn,  rec_seq, rec_seq_sz);
 
-	tls_version = MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_2;
+	if (crypto_info->crypto_info.version == TLS_1_3_VERSION) {
+		memcpy(gcm_iv + salt_sz, iv, iv_sz);
+		tls_version = MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_3;
+	} else {
+		tls_version = MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_2;
+	}
 
 	MLX5_SET(tls_static_params, ctx, tls_version, tls_version);
 	MLX5_SET(tls_static_params, ctx, const_1, 1);
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 03/15] tls: reject rekey attempts on an existing HW-offloaded connection
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 02/15] net/mlx5e: add TLS 1.3 hardware offload support Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 04/15] tls: add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (11 subsequent siblings)
  14 siblings, 0 replies; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

On a TLS 1.3 rekey, do_tls_setsockopt_conf() must not set up a brand-new
device offload: reject the rekey with -EOPNOTSUPP for a HW-offloaded
connection (HW KeyUpdate is not supported yet, and we must not fall back
to software mid-connection), and for a SW-offloaded one re-init the
software crypto state via tls_set_sw_offload() directly.

Do this by moving the reject into the device-offload helpers and shaping
the dispatch in its final form:

  - tls_set_device_offload() and tls_set_device_offload_rx() reject any
    non-initial call (tx_conf / rx_conf != TLS_BASE) with -EOPNOTSUPP.

  - do_tls_setsockopt_conf() always calls the device helper and, on
    failure, either propagates the error (rekey on a HW connection) or
    falls back to tls_set_sw_offload() (initial install, or rekey on a
    SW connection).

Prep for the following patches: "tls: add TLS 1.3 hardware offload
support" drops the TLS_1_2_VERSION guards in those helpers (which today
also happen to reject a TLS 1.3 rekey), and the KeyUpdate patches replace
the reject guards above with real rekey handling in the helpers, leaving
the do_tls_setsockopt_conf() dispatch unchanged.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 net/tls/tls_device.c | 18 ++++++++++++++++++
 net/tls/tls_main.c   | 22 ++++++++++++++++++----
 2 files changed, 36 insertions(+), 4 deletions(-)

diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index f11d0528fc43..f5e1b6b61ce3 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -1077,6 +1077,15 @@ int tls_set_device_offload(struct sock *sk)
 	ctx = tls_get_ctx(sk);
 	prot = &ctx->prot_info;
 
+	/* A rekey (setsockopt on an already-configured socket) is not
+	 * supported on the device offload path yet; reject it here so the
+	 * caller can decide (propagate the error for a HW connection, or
+	 * re-init software crypto for a SW one). KeyUpdate support replaces
+	 * this guard with real rekey handling.
+	 */
+	if (ctx->tx_conf != TLS_BASE)
+		return -EOPNOTSUPP;
+
 	if (ctx->priv_ctx_tx)
 		return -EEXIST;
 
@@ -1202,6 +1211,15 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
 	if (ctx->crypto_recv.info.version != TLS_1_2_VERSION)
 		return -EOPNOTSUPP;
 
+	/* A rekey (setsockopt on an already-configured socket) is not
+	 * supported on the device offload path yet; reject it here so the
+	 * caller can decide (propagate the error for a HW connection, or
+	 * re-init software crypto for a SW one). KeyUpdate support replaces
+	 * this guard with real rekey handling.
+	 */
+	if (ctx->rx_conf != TLS_BASE)
+		return -EOPNOTSUPP;
+
 	netdev = get_netdev_for_sock(sk);
 	if (!netdev) {
 		pr_err_ratelimited("%s: netdev not found\n", __func__);
diff --git a/net/tls/tls_main.c b/net/tls/tls_main.c
index fbb274287aa5..15e83e853f22 100644
--- a/net/tls/tls_main.c
+++ b/net/tls/tls_main.c
@@ -713,8 +713,15 @@ static int do_tls_setsockopt_conf(struct sock *sk, sockptr_t optval,
 		rc = tls_set_device_offload(sk);
 		conf = TLS_HW;
 		if (!rc) {
-			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXDEVICE);
-			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
+			if (!update) {
+				TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXDEVICE);
+				TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
+			}
+		} else if (update && ctx->tx_conf == TLS_HW) {
+			/* HW rekey failed - return the actual error.
+			 * Cannot fall back to SW for an existing HW connection.
+			 */
+			goto err_crypto_info;
 		} else {
 			rc = tls_set_sw_offload(sk, 1,
 						update ? crypto_info : NULL);
@@ -733,8 +740,15 @@ static int do_tls_setsockopt_conf(struct sock *sk, sockptr_t optval,
 		rc = tls_set_device_offload_rx(sk, ctx);
 		conf = TLS_HW;
 		if (!rc) {
-			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXDEVICE);
-			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXDEVICE);
+			if (!update) {
+				TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXDEVICE);
+				TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXDEVICE);
+			}
+		} else if (update && ctx->rx_conf == TLS_HW) {
+			/* HW rekey failed - return the actual error.
+			 * Cannot fall back to SW for an existing HW connection.
+			 */
+			goto err_crypto_info;
 		} else {
 			rc = tls_set_sw_offload(sk, 0,
 						update ? crypto_info : NULL);
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 04/15] tls: add TLS 1.3 hardware offload support
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (2 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 03/15] tls: reject rekey attempts on an existing HW-offloaded connection Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 05/15] tls: split tls_set_sw_offload into init and finalize stages Rishikesh Jethwani
                   ` (10 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Add TLS 1.3 support to the kernel TLS hardware offload infrastructure,
enabling hardware acceleration for TLS 1.3 connections on capable NICs.
The dispatch-side restructure needed to make this safe on the rekey
path lands in a preceding patch ("tls: reject rekey attempts on an
existing HW-offloaded connection"); this patch is limited to the
device / device_fallback side:

  - Drop the TLS_1_2_VERSION guards in tls_set_device_offload() and
    tls_set_device_offload_rx() so 1.3 crypto_info is accepted.

  - tls_device_record_close(): append TLS 1.3's content_type byte
    together with the tag as one tail segment; use the pre-populated
    dummy_page (identity-mapped byte values) as the fallback when
    pfrag allocation fails so the content_type byte lands on the
    dummy path too.

  - tls_device_reencrypt(): use prot->prepend_size instead of an
    open-coded TLS_HEADER_SIZE + iv, so the 1.3 prepend layout is
    handled correctly.

  - tls_device_fallback.c / tls_enc_record(): thread tls_context in
    to reach crypto_send for the 1.3 static IV; select IV source and
    length adjustment based on prot->version; XOR the IV with the
    record sequence via tls_xor_iv_with_seq() for 1.3; use
    prot->aad_size and prot->prepend_size instead of the 1.2-only
    constants. Inline the single caller of tls_init_aead_request().

  - tls_device_init(): pre-populate dummy_page with an identity byte
    map so any record_type used as a page offset yields the correct
    content_type byte on the fallback path (avoids a runtime range
    check).

Tested on Mellanox ConnectX-6 Dx (Crypto Enabled) with TLS 1.3
AES-GCM-128 and AES-GCM-256 cipher suites.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 net/tls/tls_device.c          | 72 +++++++++++++++++++++--------------
 net/tls/tls_device_fallback.c | 58 ++++++++++++++++------------
 2 files changed, 77 insertions(+), 53 deletions(-)

diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index f5e1b6b61ce3..ada66c0bd075 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -317,25 +317,34 @@ static void tls_device_record_close(struct sock *sk,
 				    unsigned char record_type)
 {
 	struct tls_prot_info *prot = &ctx->prot_info;
-	struct page_frag dummy_tag_frag;
-
-	/* append tag
-	 * device will fill in the tag, we just need to append a placeholder
-	 * use socket memory to improve coalescing (re-using a single buffer
-	 * increases frag count)
-	 * if we can't allocate memory now use the dummy page
+	int tail = prot->tag_size + prot->tail_size;
+
+	/* Append tail: tag for TLS 1.2, content_type + tag for TLS 1.3.
+	 * Device fills in the tag, we just need to append a placeholder.
+	 * Use socket memory to improve coalescing (re-using a single buffer
+	 * increases frag count); if allocation fails use dummy_page
+	 * (offset = record_type gives correct content_type byte via
+	 * identity mapping)
 	 */
-	if (unlikely(pfrag->size - pfrag->offset < prot->tag_size) &&
-	    !skb_page_frag_refill(prot->tag_size, pfrag, sk->sk_allocation)) {
-		dummy_tag_frag.page = dummy_page;
-		dummy_tag_frag.offset = 0;
-		pfrag = &dummy_tag_frag;
+	if (unlikely(!pfrag->page || pfrag->size - pfrag->offset < tail) &&
+	    !skb_page_frag_refill(tail, pfrag, sk->sk_allocation)) {
+		struct page_frag dummy_pfrag = {
+			.page = dummy_page,
+			.offset = record_type,
+		};
+		tls_append_frag(record, &dummy_pfrag, tail);
+	} else {
+		if (prot->tail_size) {
+			char *content_type_addr = page_address(pfrag->page) +
+						  pfrag->offset;
+			*content_type_addr = record_type;
+		}
+		tls_append_frag(record, pfrag, tail);
 	}
-	tls_append_frag(record, pfrag, prot->tag_size);
 
 	/* fill prepend */
 	tls_fill_prepend(ctx, skb_frag_address(&record->frags[0]),
-			 record->len - prot->overhead_size,
+			 record->len - prot->overhead_size + prot->tail_size,
 			 record_type);
 }
 
@@ -886,6 +895,7 @@ static int
 tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
 {
 	struct tls_sw_context_rx *sw_ctx = tls_sw_ctx_rx(tls_ctx);
+	struct tls_prot_info *prot = &tls_ctx->prot_info;
 	const struct tls_cipher_desc *cipher_desc;
 	int err, offset, copy, data_len, pos;
 	struct sk_buff *skb, *skb_iter;
@@ -897,7 +907,7 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
 	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
 
 	rxm = strp_msg(tls_strp_msg(sw_ctx));
-	orig_buf = kmalloc(rxm->full_len + TLS_HEADER_SIZE + cipher_desc->iv,
+	orig_buf = kmalloc(rxm->full_len + prot->prepend_size,
 			   sk->sk_allocation);
 	if (!orig_buf)
 		return -ENOMEM;
@@ -912,9 +922,8 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
 	offset = rxm->offset;
 
 	sg_init_table(sg, 1);
-	sg_set_buf(&sg[0], buf,
-		   rxm->full_len + TLS_HEADER_SIZE + cipher_desc->iv);
-	err = skb_copy_bits(skb, offset, buf, TLS_HEADER_SIZE + cipher_desc->iv);
+	sg_set_buf(&sg[0], buf, rxm->full_len + prot->prepend_size);
+	err = skb_copy_bits(skb, offset, buf, prot->prepend_size);
 	if (err)
 		goto free_buf;
 
@@ -1101,11 +1110,6 @@ int tls_set_device_offload(struct sock *sk)
 	}
 
 	crypto_info = &ctx->crypto_send.info;
-	if (crypto_info->version != TLS_1_2_VERSION) {
-		rc = -EOPNOTSUPP;
-		goto release_netdev;
-	}
-
 	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
 	if (!cipher_desc || !cipher_desc->offloadable) {
 		rc = -EINVAL;
@@ -1208,9 +1212,6 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
 	struct net_device *netdev;
 	int rc = 0;
 
-	if (ctx->crypto_recv.info.version != TLS_1_2_VERSION)
-		return -EOPNOTSUPP;
-
 	/* A rekey (setsockopt on an already-configured socket) is not
 	 * supported on the device offload path yet; reject it here so the
 	 * caller can decide (propagate the error for a HW connection, or
@@ -1429,12 +1430,27 @@ static struct notifier_block tls_dev_notifier = {
 
 int __init tls_device_init(void)
 {
-	int err;
+	unsigned char *page_addr;
+	int err, i;
 
-	dummy_page = alloc_page(GFP_KERNEL);
+	dummy_page = alloc_page(GFP_KERNEL | __GFP_ZERO);
 	if (!dummy_page)
 		return -ENOMEM;
 
+	/* Pre-populate the first 256 bytes with an identity map so that,
+	 * when this page is used as the tail-frag fallback (allocation
+	 * failure in tls_device_record_close()), dummy_page[record_type]
+	 * yields the correct TLS 1.3 content_type byte for any record_type
+	 * without runtime validation.
+	 *
+	 * A high record_type pushes the tag placeholder past the identity
+	 * map, so __GFP_ZERO is what keeps tag-placeholder bytes defined
+	 * rather than exposing uninitialized page contents.
+	 */
+	page_addr = page_address(dummy_page);
+	for (i = 0; i < 256; i++)
+		page_addr[i] = (unsigned char)i;
+
 	destruct_wq = alloc_workqueue("ktls_device_destruct", WQ_PERCPU, 0);
 	if (!destruct_wq) {
 		err = -ENOMEM;
diff --git a/net/tls/tls_device_fallback.c b/net/tls/tls_device_fallback.c
index 3b7d0ab2bcf1..1110f7ac6bcb 100644
--- a/net/tls/tls_device_fallback.c
+++ b/net/tls/tls_device_fallback.c
@@ -37,14 +37,15 @@
 
 #include "tls.h"
 
-static int tls_enc_record(struct aead_request *aead_req,
+static int tls_enc_record(struct tls_context *tls_ctx,
+			  struct aead_request *aead_req,
 			  struct crypto_aead *aead, char *aad,
 			  char *iv, __be64 rcd_sn,
 			  struct scatter_walk *in,
-			  struct scatter_walk *out, int *in_len,
-			  struct tls_prot_info *prot)
+			  struct scatter_walk *out, int *in_len)
 {
 	unsigned char buf[TLS_HEADER_SIZE + TLS_MAX_IV_SIZE];
+	struct tls_prot_info *prot = &tls_ctx->prot_info;
 	const struct tls_cipher_desc *cipher_desc;
 	struct scatterlist sg_in[3];
 	struct scatterlist sg_out[3];
@@ -55,7 +56,7 @@ static int tls_enc_record(struct aead_request *aead_req,
 	cipher_desc = get_cipher_desc(prot->cipher_type);
 	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
 
-	buf_size = TLS_HEADER_SIZE + cipher_desc->iv;
+	buf_size = prot->prepend_size;
 	len = min_t(int, *in_len, buf_size);
 
 	memcpy_from_scatterwalk(buf, in, len);
@@ -66,16 +67,27 @@ static int tls_enc_record(struct aead_request *aead_req,
 		return 0;
 
 	len = buf[4] | (buf[3] << 8);
-	len -= cipher_desc->iv;
+	if (prot->version != TLS_1_3_VERSION)
+		len -= cipher_desc->iv;
 
 	tls_make_aad(aad, len - cipher_desc->tag, (char *)&rcd_sn, buf[0], prot);
 
-	memcpy(iv + cipher_desc->salt, buf + TLS_HEADER_SIZE, cipher_desc->iv);
+	if (prot->version == TLS_1_3_VERSION) {
+		void *iv_src = crypto_info_iv(&tls_ctx->crypto_send.info,
+					      cipher_desc);
+
+		memcpy(iv + cipher_desc->salt, iv_src, cipher_desc->iv);
+	} else {
+		memcpy(iv + cipher_desc->salt, buf + TLS_HEADER_SIZE,
+		       cipher_desc->iv);
+	}
+
+	tls_xor_iv_with_seq(prot, iv, (char *)&rcd_sn);
 
 	sg_init_table(sg_in, ARRAY_SIZE(sg_in));
 	sg_init_table(sg_out, ARRAY_SIZE(sg_out));
-	sg_set_buf(sg_in, aad, TLS_AAD_SPACE_SIZE);
-	sg_set_buf(sg_out, aad, TLS_AAD_SPACE_SIZE);
+	sg_set_buf(sg_in, aad, prot->aad_size);
+	sg_set_buf(sg_out, aad, prot->aad_size);
 	scatterwalk_get_sglist(in, sg_in + 1);
 	scatterwalk_get_sglist(out, sg_out + 1);
 
@@ -108,13 +120,6 @@ static int tls_enc_record(struct aead_request *aead_req,
 	return rc;
 }
 
-static void tls_init_aead_request(struct aead_request *aead_req,
-				  struct crypto_aead *aead)
-{
-	aead_request_set_tfm(aead_req, aead);
-	aead_request_set_ad(aead_req, TLS_AAD_SPACE_SIZE);
-}
-
 static struct aead_request *tls_alloc_aead_request(struct crypto_aead *aead,
 						   gfp_t flags)
 {
@@ -124,14 +129,15 @@ static struct aead_request *tls_alloc_aead_request(struct crypto_aead *aead,
 
 	aead_req = kzalloc(req_size, flags);
 	if (aead_req)
-		tls_init_aead_request(aead_req, aead);
+		aead_request_set_tfm(aead_req, aead);
 	return aead_req;
 }
 
-static int tls_enc_records(struct aead_request *aead_req,
+static int tls_enc_records(struct tls_context *tls_ctx,
+			   struct aead_request *aead_req,
 			   struct crypto_aead *aead, struct scatterlist *sg_in,
 			   struct scatterlist *sg_out, char *aad, char *iv,
-			   u64 rcd_sn, int len, struct tls_prot_info *prot)
+			   u64 rcd_sn, int len)
 {
 	struct scatter_walk out, in;
 	int rc;
@@ -140,8 +146,8 @@ static int tls_enc_records(struct aead_request *aead_req,
 	scatterwalk_start(&out, sg_out);
 
 	do {
-		rc = tls_enc_record(aead_req, aead, aad, iv,
-				    cpu_to_be64(rcd_sn), &in, &out, &len, prot);
+		rc = tls_enc_record(tls_ctx, aead_req, aead, aad, iv,
+				    cpu_to_be64(rcd_sn), &in, &out, &len);
 		rcd_sn++;
 
 	} while (rc == 0 && len);
@@ -314,7 +320,10 @@ static struct sk_buff *tls_enc_skb(struct tls_context *tls_ctx,
 	cipher_desc = get_cipher_desc(tls_ctx->crypto_send.info.cipher_type);
 	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
 
-	buf_len = cipher_desc->salt + cipher_desc->iv + TLS_AAD_SPACE_SIZE +
+	aead_request_set_ad(aead_req, tls_ctx->prot_info.aad_size);
+
+	buf_len = cipher_desc->salt + cipher_desc->iv +
+		  tls_ctx->prot_info.aad_size +
 		  sync_size + cipher_desc->tag;
 	buf = kmalloc(buf_len, GFP_ATOMIC);
 	if (!buf)
@@ -324,7 +333,7 @@ static struct sk_buff *tls_enc_skb(struct tls_context *tls_ctx,
 	salt = crypto_info_salt(&tls_ctx->crypto_send.info, cipher_desc);
 	memcpy(iv, salt, cipher_desc->salt);
 	aad = buf + cipher_desc->salt + cipher_desc->iv;
-	dummy_buf = aad + TLS_AAD_SPACE_SIZE;
+	dummy_buf = aad + tls_ctx->prot_info.aad_size;
 
 	nskb = alloc_skb(skb_headroom(skb) + skb->len, GFP_ATOMIC);
 	if (!nskb)
@@ -335,9 +344,8 @@ static struct sk_buff *tls_enc_skb(struct tls_context *tls_ctx,
 	fill_sg_out(sg_out, buf, tls_ctx, nskb, tcp_payload_offset,
 		    payload_len, sync_size, dummy_buf);
 
-	if (tls_enc_records(aead_req, ctx->aead_send, sg_in, sg_out, aad, iv,
-			    rcd_sn, sync_size + payload_len,
-			    &tls_ctx->prot_info) < 0)
+	if (tls_enc_records(tls_ctx, aead_req, ctx->aead_send, sg_in, sg_out,
+			    aad, iv, rcd_sn, sync_size + payload_len) < 0)
 		goto free_nskb;
 
 	complete_skb(nskb, skb, tcp_payload_offset);
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 05/15] tls: split tls_set_sw_offload into init and finalize stages
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (3 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 04/15] tls: add TLS 1.3 hardware offload support Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 06/15] tls: prep helpers and refactors for HW offload KeyUpdate Rishikesh Jethwani
                   ` (9 subsequent siblings)
  14 siblings, 0 replies; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Separate cipher context initialization from key material finalization
to support staged setup for hardware offload fallback paths.

tls_sw_ctx_init() allocates the SW cipher context and installs the key
(crypto_alloc_aead + crypto_aead_setkey); tls_sw_ctx_finalize() then
commits the IV/salt/rec_seq into ctx->{tx,rx}. tls_set_sw_offload() calls
both back-to-back, so its behaviour is unchanged.

In tls_set_device_offload_rx() finalize now runs after tls_dev_add()
rather than as part of the old fused tls_set_sw_offload() call. This is
required by the RX rekey path added later in the series: on rekey the
device is programmed from new_crypto_info (via src_crypto_info), and
tls_sw_ctx_finalize() copies new_crypto_info into ctx->crypto_recv.info
and then memzero_explicit()s the caller's buffer. Running the fused call
before tls_dev_add() would wipe the key material before the NIC is
programmed, so tls_dev_add() would install a zeroed key. Splitting lets
tls_dev_add() consume new_crypto_info while it is still intact, with
finalize committing ctx->rx and scrubbing the buffer afterwards. On the
non-rekey (NULL) path this is a no-op reorder.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 net/tls/tls.h        | 12 +++++++
 net/tls/tls_device.c |  3 +-
 net/tls/tls_sw.c     | 79 ++++++++++++++++++++++++++++++++------------
 3 files changed, 71 insertions(+), 23 deletions(-)

diff --git a/net/tls/tls.h b/net/tls/tls.h
index 60a37bdaaa25..8450492f32ae 100644
--- a/net/tls/tls.h
+++ b/net/tls/tls.h
@@ -147,6 +147,18 @@ void tls_strp_abort_strp(struct tls_strparser *strp, int err);
 int init_prot_info(struct tls_prot_info *prot,
 		   const struct tls_crypto_info *crypto_info,
 		   const struct tls_cipher_desc *cipher_desc);
+/* tls_sw_ctx_init() and tls_sw_ctx_finalize() are two halves of installing
+ * a SW crypto context, split so the device path can attach the NIC between
+ * them. finalize() may only be called after an init() that returned 0, and
+ * both must be called with the same tx and new_crypto_info; on a rekey
+ * (new_crypto_info != NULL) the two must also see the same
+ * new_crypto_info->cipher_type. finalize() commits state and cannot fail,
+ * so violating this leaves the context inconsistent without any error.
+ */
+int tls_sw_ctx_init(struct sock *sk, int tx,
+		    struct tls_crypto_info *new_crypto_info);
+void tls_sw_ctx_finalize(struct sock *sk, int tx,
+			 struct tls_crypto_info *new_crypto_info);
 int tls_set_sw_offload(struct sock *sk, int tx,
 		       struct tls_crypto_info *new_crypto_info);
 void tls_update_rx_zc_capable(struct tls_context *tls_ctx);
diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index ada66c0bd075..a1e22d9f3217 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -1254,7 +1254,7 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
 	context->resync_nh_reset = 1;
 
 	ctx->priv_ctx_rx = context;
-	rc = tls_set_sw_offload(sk, 0, NULL);
+	rc = tls_sw_ctx_init(sk, 0, NULL);
 	if (rc)
 		goto release_ctx;
 
@@ -1268,6 +1268,7 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
 		goto free_sw_resources;
 
 	tls_device_attach(ctx, sk, netdev);
+	tls_sw_ctx_finalize(sk, 0, NULL);
 	up_read(&device_offload_lock);
 
 	dev_put(netdev);
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index d1ad31986cf2..7b593dac2c31 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -2522,20 +2522,19 @@ static void tls_finish_key_update(struct sock *sk, struct tls_context *tls_ctx)
 	ctx->saved_data_ready(sk);
 }
 
-int tls_set_sw_offload(struct sock *sk, int tx,
-		       struct tls_crypto_info *new_crypto_info)
+int tls_sw_ctx_init(struct sock *sk, int tx,
+		    struct tls_crypto_info *new_crypto_info)
 {
 	struct tls_crypto_info *crypto_info, *src_crypto_info;
 	struct tls_sw_context_tx *sw_ctx_tx = NULL;
 	struct tls_sw_context_rx *sw_ctx_rx = NULL;
 	const struct tls_cipher_desc *cipher_desc;
-	char *iv, *rec_seq, *key, *salt;
-	struct cipher_context *cctx;
 	struct tls_prot_info *prot;
 	struct crypto_aead **aead;
 	struct tls_context *ctx;
 	struct crypto_tfm *tfm;
 	int rc = 0;
+	char *key;
 
 	ctx = tls_get_ctx(sk);
 	prot = &ctx->prot_info;
@@ -2556,12 +2555,10 @@ int tls_set_sw_offload(struct sock *sk, int tx,
 	if (tx) {
 		sw_ctx_tx = ctx->priv_ctx_tx;
 		crypto_info = &ctx->crypto_send.info;
-		cctx = &ctx->tx;
 		aead = &sw_ctx_tx->aead_send;
 	} else {
 		sw_ctx_rx = ctx->priv_ctx_rx;
 		crypto_info = &ctx->crypto_recv.info;
-		cctx = &ctx->rx;
 		aead = &sw_ctx_rx->aead_recv;
 	}
 
@@ -2577,10 +2574,7 @@ int tls_set_sw_offload(struct sock *sk, int tx,
 	if (rc)
 		goto free_priv;
 
-	iv = crypto_info_iv(src_crypto_info, cipher_desc);
 	key = crypto_info_key(src_crypto_info, cipher_desc);
-	salt = crypto_info_salt(src_crypto_info, cipher_desc);
-	rec_seq = crypto_info_rec_seq(src_crypto_info, cipher_desc);
 
 	if (!*aead) {
 		*aead = crypto_alloc_aead(cipher_desc->cipher_name, 0, 0);
@@ -2624,19 +2618,6 @@ int tls_set_sw_offload(struct sock *sk, int tx,
 			goto free_aead;
 	}
 
-	memcpy(cctx->iv, salt, cipher_desc->salt);
-	memcpy(cctx->iv + cipher_desc->salt, iv, cipher_desc->iv);
-	memcpy(cctx->rec_seq, rec_seq, cipher_desc->rec_seq);
-
-	if (new_crypto_info) {
-		unsafe_memcpy(crypto_info, new_crypto_info,
-			      cipher_desc->crypto_info,
-			      /* size was checked in do_tls_setsockopt_conf */);
-		memzero_explicit(new_crypto_info, cipher_desc->crypto_info);
-		if (!tx)
-			tls_finish_key_update(sk, ctx);
-	}
-
 	goto out;
 
 free_aead:
@@ -2655,3 +2636,57 @@ int tls_set_sw_offload(struct sock *sk, int tx,
 out:
 	return rc;
 }
+
+void tls_sw_ctx_finalize(struct sock *sk, int tx,
+			 struct tls_crypto_info *new_crypto_info)
+{
+	struct tls_crypto_info *crypto_info, *src_crypto_info;
+	const struct tls_cipher_desc *cipher_desc;
+	struct tls_context *ctx = tls_get_ctx(sk);
+	struct cipher_context *cctx;
+	char *iv, *salt, *rec_seq;
+
+	if (tx) {
+		crypto_info = &ctx->crypto_send.info;
+		cctx = &ctx->tx;
+	} else {
+		crypto_info = &ctx->crypto_recv.info;
+		cctx = &ctx->rx;
+	}
+
+	src_crypto_info = new_crypto_info ?: crypto_info;
+
+	/* Infallible: tls_sw_ctx_init() already validated cipher_type. */
+	cipher_desc = get_cipher_desc(src_crypto_info->cipher_type);
+
+	iv = crypto_info_iv(src_crypto_info, cipher_desc);
+	salt = crypto_info_salt(src_crypto_info, cipher_desc);
+	rec_seq = crypto_info_rec_seq(src_crypto_info, cipher_desc);
+
+	memcpy(cctx->iv, salt, cipher_desc->salt);
+	memcpy(cctx->iv + cipher_desc->salt, iv, cipher_desc->iv);
+	memcpy(cctx->rec_seq, rec_seq, cipher_desc->rec_seq);
+
+	if (new_crypto_info) {
+		unsafe_memcpy(crypto_info, new_crypto_info,
+			      cipher_desc->crypto_info,
+			      /* size was checked in do_tls_setsockopt_conf */);
+		memzero_explicit(new_crypto_info, cipher_desc->crypto_info);
+
+		if (!tx)
+			tls_finish_key_update(sk, ctx);
+	}
+}
+
+int tls_set_sw_offload(struct sock *sk, int tx,
+		       struct tls_crypto_info *new_crypto_info)
+{
+	int rc;
+
+	rc = tls_sw_ctx_init(sk, tx, new_crypto_info);
+	if (rc)
+		return rc;
+
+	tls_sw_ctx_finalize(sk, tx, new_crypto_info);
+	return 0;
+}
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 06/15] tls: prep helpers and refactors for HW offload KeyUpdate
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (4 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 05/15] tls: split tls_set_sw_offload into init and finalize stages Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 07/15] net: sched: re-validate parked decrypted skbs on requeue Rishikesh Jethwani
                   ` (8 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Preparatory refactors for TX and RX HW rekey support; no functional
change.

  - Hoist cipher_context / tls_crypto_context above
    tls_offload_context_tx so they can be embedded in offload
    contexts.

  - Add tls_tx_cipher_ctx() accessor and factor tls_sw_ctx_tx_init()
    so the TX path can redirect to a temporary SW context during
    rekey.

  - Split tls_set_device_offload() into a dispatcher and
    tls_set_device_offload_initial(); a _rekey() sibling follows.

  - Factor tls_device_dev_add_tx() and tls_device_commit_start_marker()
    so the rekey completion path can reuse them.

  - Move crypto_aead_setauthsize() into the !*aead block so a fresh
    AEAD is correctly configured when RX HW rekey allocates one.

  - Export tls_sw_push_pending_record() (drop static) so the device
    TX path can flush a pending SW record while rekey is in flight.

  - Split tls_sw_splice_eof() into tls_sw_splice_eof_locked() plus a
    thin locking wrapper, so the device splice_eof path can reuse the
    inner logic with tx_lock and the socket lock already held.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 include/net/tls.h    |  38 ++++++-----
 net/tls/tls.h        |   3 +
 net/tls/tls_device.c | 157 ++++++++++++++++++++++++++-----------------
 net/tls/tls_sw.c     |  78 ++++++++++++---------
 4 files changed, 168 insertions(+), 108 deletions(-)

diff --git a/include/net/tls.h b/include/net/tls.h
index e57bef58851e..eb258bcd62bc 100644
--- a/include/net/tls.h
+++ b/include/net/tls.h
@@ -155,6 +155,22 @@ struct tls_record_info {
 	skb_frag_t frags[MAX_SKB_FRAGS];
 };
 
+struct cipher_context {
+	char iv[TLS_MAX_IV_SIZE + TLS_MAX_SALT_SIZE];
+	char rec_seq[TLS_MAX_REC_SEQ_SIZE];
+};
+
+union tls_crypto_context {
+	struct tls_crypto_info info;
+	union {
+		struct tls12_crypto_info_aes_gcm_128 aes_gcm_128;
+		struct tls12_crypto_info_aes_gcm_256 aes_gcm_256;
+		struct tls12_crypto_info_chacha20_poly1305 chacha20_poly1305;
+		struct tls12_crypto_info_sm4_gcm sm4_gcm;
+		struct tls12_crypto_info_sm4_ccm sm4_ccm;
+	};
+};
+
 #define TLS_DRIVER_STATE_SIZE_TX	16
 struct tls_offload_context_tx {
 	struct crypto_aead *aead_send;
@@ -195,22 +211,6 @@ enum tls_context_flags {
 	TLS_RX_DEV_CLOSED = 2,
 };
 
-struct cipher_context {
-	char iv[TLS_MAX_IV_SIZE + TLS_MAX_SALT_SIZE];
-	char rec_seq[TLS_MAX_REC_SEQ_SIZE];
-};
-
-union tls_crypto_context {
-	struct tls_crypto_info info;
-	union {
-		struct tls12_crypto_info_aes_gcm_128 aes_gcm_128;
-		struct tls12_crypto_info_aes_gcm_256 aes_gcm_256;
-		struct tls12_crypto_info_chacha20_poly1305 chacha20_poly1305;
-		struct tls12_crypto_info_sm4_gcm sm4_gcm;
-		struct tls12_crypto_info_sm4_ccm sm4_ccm;
-	};
-};
-
 struct tls_prot_info {
 	u16 version;
 	u16 cipher_type;
@@ -392,6 +392,12 @@ static inline struct tls_sw_context_tx *tls_sw_ctx_tx(
 	return (struct tls_sw_context_tx *)tls_ctx->priv_ctx_tx;
 }
 
+static inline struct cipher_context *tls_tx_cipher_ctx(
+		const struct tls_context *tls_ctx)
+{
+	return (struct cipher_context *)&tls_ctx->tx;
+}
+
 static inline struct tls_offload_context_tx *
 tls_offload_ctx_tx(const struct tls_context *tls_ctx)
 {
diff --git a/net/tls/tls.h b/net/tls/tls.h
index 8450492f32ae..920a926e8e68 100644
--- a/net/tls/tls.h
+++ b/net/tls/tls.h
@@ -165,7 +165,10 @@ void tls_update_rx_zc_capable(struct tls_context *tls_ctx);
 void tls_sw_strparser_arm(struct sock *sk, struct tls_context *ctx);
 void tls_sw_strparser_done(struct tls_context *tls_ctx);
 int tls_sw_sendmsg(struct sock *sk, struct msghdr *msg, size_t size);
+void tls_sw_ctx_tx_init(struct sock *sk, struct tls_sw_context_tx *sw_ctx);
+int tls_sw_push_pending_record(struct sock *sk, int flags);
 void tls_sw_splice_eof(struct socket *sock);
+void tls_sw_splice_eof_locked(struct socket *sock);
 void tls_sw_cancel_work_tx(struct tls_context *tls_ctx);
 void tls_sw_release_resources_tx(struct sock *sk);
 void tls_sw_free_ctx_tx(struct tls_context *tls_ctx);
diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index a1e22d9f3217..972c9c7ba7de 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -138,6 +138,41 @@ static struct net_device *get_netdev_for_sock(struct sock *sk)
 	return lowest_dev;
 }
 
+static int tls_device_dev_add_tx(struct sock *sk, struct net_device *netdev,
+				 struct tls_crypto_info *crypto_info,
+				 u32 write_seq)
+{
+	const struct tls_cipher_desc *cipher_desc;
+	char *rec_seq;
+	int rc;
+
+	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
+	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
+
+	rc = netdev->tlsdev_ops->tls_dev_add(netdev, sk, TLS_OFFLOAD_CTX_DIR_TX,
+					     crypto_info, write_seq);
+	rec_seq = crypto_info_rec_seq(crypto_info, cipher_desc);
+	trace_tls_device_offload_set(sk, TLS_OFFLOAD_CTX_DIR_TX,
+				     write_seq, rec_seq, rc);
+	return rc;
+}
+
+static void tls_device_commit_start_marker(struct sock *sk,
+					struct tls_offload_context_tx *offload_ctx,
+					struct tls_record_info *start_marker_record)
+{
+	start_marker_record->end_seq = tcp_sk(sk)->write_seq;
+	start_marker_record->len = 0;
+	start_marker_record->num_frags = 0;
+	list_add_tail_rcu(&start_marker_record->list, &offload_ctx->records_list);
+
+	/* TLS offload is greatly simplified if we don't send
+	 * SKBs where only part of the payload needs to be encrypted.
+	 * So mark the last skb in the write queue as end of record.
+	 */
+	tcp_write_collapse_fence(sk);
+}
+
 static void destroy_record(struct tls_record_info *record)
 {
 	int i;
@@ -1071,66 +1106,31 @@ static struct tls_offload_context_tx *alloc_offload_ctx_tx(struct tls_context *c
 	return offload_ctx;
 }
 
-int tls_set_device_offload(struct sock *sk)
+static int tls_set_device_offload_initial(struct sock *sk,
+					  struct tls_context *ctx,
+					  struct net_device *netdev,
+					  struct tls_crypto_info *crypto_info,
+					  const struct tls_cipher_desc *cipher_desc)
 {
+	struct tls_prot_info *prot = &ctx->prot_info;
 	struct tls_record_info *start_marker_record;
 	struct tls_offload_context_tx *offload_ctx;
-	const struct tls_cipher_desc *cipher_desc;
-	struct tls_crypto_info *crypto_info;
-	struct tls_prot_info *prot;
-	struct net_device *netdev;
-	struct tls_context *ctx;
 	char *iv, *rec_seq;
 	int rc;
 
-	ctx = tls_get_ctx(sk);
-	prot = &ctx->prot_info;
-
-	/* A rekey (setsockopt on an already-configured socket) is not
-	 * supported on the device offload path yet; reject it here so the
-	 * caller can decide (propagate the error for a HW connection, or
-	 * re-init software crypto for a SW one). KeyUpdate support replaces
-	 * this guard with real rekey handling.
-	 */
-	if (ctx->tx_conf != TLS_BASE)
-		return -EOPNOTSUPP;
-
-	if (ctx->priv_ctx_tx)
-		return -EEXIST;
-
-	netdev = get_netdev_for_sock(sk);
-	if (!netdev) {
-		pr_err_ratelimited("%s: netdev not found\n", __func__);
-		return -EINVAL;
-	}
-
-	if (!(netdev->features & NETIF_F_HW_TLS_TX)) {
-		rc = -EOPNOTSUPP;
-		goto release_netdev;
-	}
-
-	crypto_info = &ctx->crypto_send.info;
-	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
-	if (!cipher_desc || !cipher_desc->offloadable) {
-		rc = -EINVAL;
-		goto release_netdev;
-	}
+	iv = crypto_info_iv(crypto_info, cipher_desc);
+	rec_seq = crypto_info_rec_seq(crypto_info, cipher_desc);
 
 	rc = init_prot_info(prot, crypto_info, cipher_desc);
 	if (rc)
-		goto release_netdev;
-
-	iv = crypto_info_iv(crypto_info, cipher_desc);
-	rec_seq = crypto_info_rec_seq(crypto_info, cipher_desc);
+		return rc;
 
 	memcpy(ctx->tx.iv + cipher_desc->salt, iv, cipher_desc->iv);
 	memcpy(ctx->tx.rec_seq, rec_seq, cipher_desc->rec_seq);
 
 	start_marker_record = kmalloc_obj(*start_marker_record);
-	if (!start_marker_record) {
-		rc = -ENOMEM;
-		goto release_netdev;
-	}
+	if (!start_marker_record)
+		return -ENOMEM;
 
 	offload_ctx = alloc_offload_ctx_tx(ctx);
 	if (!offload_ctx) {
@@ -1142,20 +1142,11 @@ int tls_set_device_offload(struct sock *sk)
 	if (rc)
 		goto free_offload_ctx;
 
-	start_marker_record->end_seq = tcp_sk(sk)->write_seq;
-	start_marker_record->len = 0;
-	start_marker_record->num_frags = 0;
-	list_add_tail(&start_marker_record->list, &offload_ctx->records_list);
+	tls_device_commit_start_marker(sk, offload_ctx, start_marker_record);
 
 	clean_acked_data_enable(tcp_sk(sk), &tls_tcp_clean_acked);
 	ctx->push_pending_record = tls_device_push_pending_record;
 
-	/* TLS offload is greatly simplified if we don't send
-	 * SKBs where only part of the payload needs to be encrypted.
-	 * So mark the last skb in the write queue as end of record.
-	 */
-	tcp_write_collapse_fence(sk);
-
 	/* Avoid offloading if the device is down
 	 * We don't want to offload new flows after
 	 * the NETDEV_DOWN event
@@ -1171,11 +1162,8 @@ int tls_set_device_offload(struct sock *sk)
 	}
 
 	ctx->priv_ctx_tx = offload_ctx;
-	rc = netdev->tlsdev_ops->tls_dev_add(netdev, sk, TLS_OFFLOAD_CTX_DIR_TX,
-					     &ctx->crypto_send.info,
-					     tcp_sk(sk)->write_seq);
-	trace_tls_device_offload_set(sk, TLS_OFFLOAD_CTX_DIR_TX,
-				     tcp_sk(sk)->write_seq, rec_seq, rc);
+	rc = tls_device_dev_add_tx(sk, netdev, crypto_info,
+				   tcp_sk(sk)->write_seq);
 	if (rc)
 		goto release_lock;
 
@@ -1187,7 +1175,6 @@ int tls_set_device_offload(struct sock *sk)
 	 * by the netdev's xmit function.
 	 */
 	smp_store_release(&sk->sk_validate_xmit_skb, tls_validate_xmit_skb);
-	dev_put(netdev);
 
 	return 0;
 
@@ -1200,6 +1187,52 @@ int tls_set_device_offload(struct sock *sk)
 	ctx->priv_ctx_tx = NULL;
 free_marker_record:
 	kfree(start_marker_record);
+	return rc;
+}
+
+int tls_set_device_offload(struct sock *sk)
+{
+	const struct tls_cipher_desc *cipher_desc;
+	struct tls_crypto_info *crypto_info;
+	struct net_device *netdev;
+	struct tls_context *ctx;
+	int rc;
+
+	ctx = tls_get_ctx(sk);
+
+	/* A rekey (setsockopt on an already-configured socket) is not
+	 * supported on the device offload path yet; reject it here so the
+	 * caller can decide (propagate the error for a HW connection, or
+	 * re-init software crypto for a SW one). KeyUpdate support replaces
+	 * this guard with real rekey handling.
+	 */
+	if (ctx->tx_conf != TLS_BASE)
+		return -EOPNOTSUPP;
+
+	if (ctx->priv_ctx_tx)
+		return -EEXIST;
+
+	netdev = get_netdev_for_sock(sk);
+	if (!netdev) {
+		pr_err_ratelimited("%s: netdev not found\n", __func__);
+		return -EINVAL;
+	}
+
+	if (!(netdev->features & NETIF_F_HW_TLS_TX)) {
+		rc = -EOPNOTSUPP;
+		goto release_netdev;
+	}
+
+	crypto_info = &ctx->crypto_send.info;
+	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
+	if (!cipher_desc || !cipher_desc->offloadable) {
+		rc = -EINVAL;
+		goto release_netdev;
+	}
+
+	rc = tls_set_device_offload_initial(sk, ctx, netdev, crypto_info,
+					    cipher_desc);
+
 release_netdev:
 	dev_put(netdev);
 	return rc;
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index 7b593dac2c31..5531303dd704 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -555,11 +555,11 @@ static int tls_do_encryption(struct sock *sk,
 		break;
 	}
 
-	memcpy(&rec->iv_data[iv_offset], tls_ctx->tx.iv,
+	memcpy(&rec->iv_data[iv_offset], tls_tx_cipher_ctx(tls_ctx)->iv,
 	       prot->iv_size + prot->salt_size);
 
 	tls_xor_iv_with_seq(prot, rec->iv_data + iv_offset,
-			    tls_ctx->tx.rec_seq);
+			    tls_tx_cipher_ctx(tls_ctx)->rec_seq);
 
 	sge->offset += prot->prepend_size;
 	sge->length -= prot->prepend_size;
@@ -610,7 +610,7 @@ static int tls_do_encryption(struct sock *sk,
 
 	/* Unhook the record from context if encryption is not failure */
 	ctx->open_rec = NULL;
-	tls_advance_record_sn(sk, prot, &tls_ctx->tx);
+	tls_advance_record_sn(sk, prot, tls_tx_cipher_ctx(tls_ctx));
 	return rc;
 }
 
@@ -676,7 +676,7 @@ static int tls_push_record(struct sock *sk, int flags,
 	sg_chain(rec->sg_aead_out, 2, &msg_en->sg.data[i]);
 
 	tls_make_aad(rec->aad_space, msg_pl->sg.size + prot->tail_size,
-		     tls_ctx->tx.rec_seq, record_type, prot);
+		     tls_tx_cipher_ctx(tls_ctx)->rec_seq, record_type, prot);
 
 	tls_fill_prepend(tls_ctx,
 			 page_address(sg_page(&msg_en->sg.data[i])) +
@@ -712,7 +712,7 @@ static int bpf_exec_tx_verdict(struct sk_msg *msg, struct sock *sk,
 	return err;
 }
 
-static int tls_sw_push_pending_record(struct sock *sk, int flags)
+int tls_sw_push_pending_record(struct sock *sk, int flags)
 {
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
 	struct tls_sw_context_tx *ctx = tls_sw_ctx_tx(tls_ctx);
@@ -1027,8 +1027,13 @@ int tls_sw_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
 
 /*
  * Handle unexpected EOF during splice without SPLICE_F_MORE set.
+ *
+ * Inner logic of tls_sw_splice_eof(), factored out so the device
+ * TX path can reuse it with tls_ctx->tx_lock and the socket lock
+ * already held. Callers not already holding both locks must use the
+ * tls_sw_splice_eof() wrapper instead.
  */
-void tls_sw_splice_eof(struct socket *sock)
+void tls_sw_splice_eof_locked(struct socket *sock)
 {
 	struct sock *sk = sock->sk;
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
@@ -1039,21 +1044,15 @@ void tls_sw_splice_eof(struct socket *sock)
 	bool retrying = false;
 	int ret = 0;
 
-	if (!ctx->open_rec)
-		return;
-
-	mutex_lock(&tls_ctx->tx_lock);
-	lock_sock(sk);
-
 retry:
-	/* same checks as in tls_sw_push_pending_record() */
+	/* same open_rec / empty-record checks as tls_sw_push_pending_record() */
 	rec = ctx->open_rec;
 	if (!rec)
-		goto unlock;
+		return;
 
 	msg_pl = &rec->msg_plaintext;
 	if (msg_pl->sg.size == 0)
-		goto unlock;
+		return;
 
 	/* Perform transmission. */
 	ret = bpf_exec_tx_verdict(msg_pl, sk, TLS_RECORD_TYPE_DATA,
@@ -1062,26 +1061,38 @@ void tls_sw_splice_eof(struct socket *sock)
 	case 0:
 	case -EAGAIN:
 		if (retrying)
-			goto unlock;
+			return;
 		retrying = true;
 		goto retry;
 	case -EINPROGRESS:
 		break;
 	default:
-		goto unlock;
+		return;
 	}
 
 	/* Wait for pending encryptions to get completed */
 	if (tls_encrypt_async_wait(ctx))
-		goto unlock;
+		return;
 
 	/* Transmit if any encryptions have completed */
 	if (test_and_clear_bit(BIT_TX_SCHEDULED, &ctx->tx_bitmask)) {
 		cancel_delayed_work(&ctx->tx_work.work);
 		tls_tx_records(sk, 0);
 	}
+}
+
+void tls_sw_splice_eof(struct socket *sock)
+{
+	struct sock *sk = sock->sk;
+	struct tls_context *tls_ctx = tls_get_ctx(sk);
+	struct tls_sw_context_tx *ctx = tls_sw_ctx_tx(tls_ctx);
 
-unlock:
+	if (!ctx->open_rec)
+		return;
+
+	mutex_lock(&tls_ctx->tx_lock);
+	lock_sock(sk);
+	tls_sw_splice_eof_locked(sock);
 	release_sock(sk);
 	mutex_unlock(&tls_ctx->tx_lock);
 }
@@ -2401,6 +2412,15 @@ static void tx_work_handler(struct work_struct *work)
 	}
 }
 
+void tls_sw_ctx_tx_init(struct sock *sk, struct tls_sw_context_tx *sw_ctx)
+{
+	crypto_init_wait(&sw_ctx->async_wait);
+	atomic_set(&sw_ctx->encrypt_pending, 1);
+	INIT_LIST_HEAD(&sw_ctx->tx_list);
+	INIT_DELAYED_WORK(&sw_ctx->tx_work.work, tx_work_handler);
+	sw_ctx->tx_work.sk = sk;
+}
+
 static bool tls_is_tx_ready(struct tls_sw_context_tx *ctx)
 {
 	struct tls_rec *rec;
@@ -2452,11 +2472,7 @@ static struct tls_sw_context_tx *init_ctx_tx(struct tls_context *ctx, struct soc
 		sw_ctx_tx = ctx->priv_ctx_tx;
 	}
 
-	crypto_init_wait(&sw_ctx_tx->async_wait);
-	atomic_set(&sw_ctx_tx->encrypt_pending, 1);
-	INIT_LIST_HEAD(&sw_ctx_tx->tx_list);
-	INIT_DELAYED_WORK(&sw_ctx_tx->tx_work.work, tx_work_handler);
-	sw_ctx_tx->tx_work.sk = sk;
+	tls_sw_ctx_tx_init(sk, sw_ctx_tx);
 
 	return sw_ctx_tx;
 }
@@ -2576,6 +2592,10 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
 
 	key = crypto_info_key(src_crypto_info, cipher_desc);
 
+	/* A rekey normally reuses the existing tfm; the RX HW rekey hands over a
+	 * NULL aead (the old one is retained for the drain), so allocate and
+	 * configure authsize only when a fresh tfm is created here.
+	 */
 	if (!*aead) {
 		*aead = crypto_alloc_aead(cipher_desc->cipher_name, 0, 0);
 		if (IS_ERR(*aead)) {
@@ -2583,6 +2603,10 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
 			*aead = NULL;
 			goto free_priv;
 		}
+
+		rc = crypto_aead_setauthsize(*aead, prot->tag_size);
+		if (rc)
+			goto free_aead;
 	}
 
 	ctx->push_pending_record = tls_sw_push_pending_record;
@@ -2599,12 +2623,6 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
 			goto free_aead;
 	}
 
-	if (!new_crypto_info) {
-		rc = crypto_aead_setauthsize(*aead, prot->tag_size);
-		if (rc)
-			goto free_aead;
-	}
-
 	if (!tx && !new_crypto_info) {
 		tfm = crypto_aead_tfm(sw_ctx_rx->aead_recv);
 
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 07/15] net: sched: re-validate parked decrypted skbs on requeue
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (5 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 06/15] tls: prep helpers and refactors for HW offload KeyUpdate Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 08/15] tcp: fence collapse against rtx-queue tail when write queue is empty Rishikesh Jethwani
                   ` (7 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

A still-cleartext skb of a crypto-offloaded socket was validated against
that socket's offload state (sk->sk_validate_xmit_skb) at the time it was
enqueued. That state can change while the skb is parked on the qdisc,
e.g. a TLS key update or offload teardown, so re-validate it on
requeue, letting the current callback decide how it reaches the wire
instead of emitting now-unencrypted plaintext.

This is a prerequisite for TLS 1.3 device-offload KeyUpdate support,
which swaps the offload state of a live connection.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 net/sched/sch_generic.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/net/sched/sch_generic.c b/net/sched/sch_generic.c
index 6f6a6f0d5eb0..fc8ef0d13f5e 100644
--- a/net/sched/sch_generic.c
+++ b/net/sched/sch_generic.c
@@ -285,6 +285,15 @@ static struct sk_buff *dequeue_skb(struct Qdisc *q, bool *validate,
 		*validate = false;
 		if (xfrm_offload(skb))
 			*validate = true;
+		/* A still-cleartext skb of a crypto-offloaded socket was validated
+		 * against that socket's offload state at the time. That state
+		 * (sk->sk_validate_xmit_skb) can change while the skb is parked here
+		 * e.g. a TLS key update or offload teardown, so re-validate it,
+		 * letting the current callback decide how it reaches the wire instead
+		 * of emitting now-unencrypted plaintext.
+		 */
+		if (skb_is_decrypted(skb))
+			*validate = true;
 		/* check the reason of requeuing without tx lock first */
 		txq = skb_get_tx_queue(txq->dev, skb);
 		if (!netif_xmit_frozen_or_stopped(txq)) {
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 08/15] tcp: fence collapse against rtx-queue tail when write queue is empty
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (6 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 07/15] net: sched: re-validate parked decrypted skbs on requeue Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 09/15] net: skbuff: add skb->decrypt_failed bit Rishikesh Jethwani
                   ` (6 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

tcp_write_collapse_fence() marks the current write-queue tail as
end-of-record so a later tcp_retrans_try_collapse() / tcp_shift_skb_data()
does not merge across the fence. When nothing is queued for transmit the
write queue is empty, and the last skb of the current state is the
retransmit-queue tail (its end_seq == snd_nxt == write_seq); fence that
instead. Otherwise the boundary is left unmarked and a later collapse can
merge it with the first skb of the next state across the fence, since
those paths test only the tail's EOR, not skb->decrypted.

This is a critical fix for both TLS device offload and PSP (PSP Security
Protocol). Both use skb->decrypted to mark encrypted/transformed packets:
pre-key skbs stay decrypted=0, post-key skbs are decrypted=1. A collapse
can merge a post-key skb into a pre-key skb, whose decrypted=0 causes the
driver to send the post-key payload in cleartext.

The write-queue-empty case is the common state at the time encryption
keys are installed (tls_set_device_offload() setsockopt and
psp_sock_set_tx_key() in PSP), when the previous handshake or request
has just finished and nothing is being sent. The existing fence is a
no-op there (write_queue_tail is NULL), leaving the boundary unmarked.
A later retransmit of the final pre-key handshake record can then
collapse against new post-key data via tcp_retrans_try_collapse() or
SACK-driven tcp_shift_skb_data(), leaking the post-key payload in
cleartext.

This fix makes tcp_write_collapse_fence() fence the rtx-queue tail when
the write queue is empty, blocking the merge. TLS 1.3 device-offload
KeyUpdate additionally relies on this behavior to keep old-key and
new-key records in distinct skbs for re-encryption on RX.

Fixes: 1be68a87ab33 ("tcp: add a helper for setting EOR on tail skb")
Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 include/net/tcp.h | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/include/net/tcp.h b/include/net/tcp.h
index 5e5f5f9b89a3..8c6d90e962c4 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -2340,6 +2340,15 @@ static inline void tcp_write_collapse_fence(struct sock *sk)
 {
 	struct sk_buff *skb = tcp_write_queue_tail(sk);
 
+	/* When nothing is queued for transmit, the last skb of the current
+	 * state is the rtx queue tail (its end_seq == snd_nxt == write_seq).
+	 * Fence that instead, otherwise the boundary is left unmarked and a
+	 * later tcp_retrans_try_collapse()/tcp_shift_skb_data() can merge it
+	 * with the first skb of the next state across the fence (they only test
+	 * the tail's EOR, not skb->decrypted).
+	 */
+	if (!skb)
+		skb = tcp_rtx_queue_tail(sk);
 	if (skb)
 		TCP_SKB_CB(skb)->eor = 1;
 }
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 09/15] net: skbuff: add skb->decrypt_failed bit
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (7 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 08/15] tcp: fence collapse against rtx-queue tail when write queue is empty Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 10/15] net/mlx5e: flag TLS RX records that failed device decryption Rishikesh Jethwani
                   ` (5 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Add skb->decrypt_failed, set when hardware could not authenticate an
skb's TLS payload. The payload may have been transformed (XORed) or
left as wire ciphertext, so software must re-authenticate the record
and undo the transform on any XORed fragment before it can be
decrypted. Propagate it in skb_copy_decrypted() alongside
skb->decrypted. Guarded by CONFIG_SKB_DECRYPTED like skb->decrypted.

This is a prerequisite for TLS 1.3 device-offload RX KeyUpdate support,
where the NIC flags records it could not authenticate across a key
change.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 include/linux/skbuff.h | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index 421f6fc45451..5da2c1149d98 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -851,6 +851,10 @@ enum skb_tstamp_type {
  *		unreadable.
  *	@dst_pending_confirm: need to confirm neighbour
  *	@decrypted: Decrypted SKB
+ *	@decrypt_failed: hardware could not authenticate this skb's TLS payload.
+ *		The payload may have been transformed (XORed) or left as wire
+ *		ciphertext, so software must re-authenticate the record and undo the
+ *		transform on any XORed fragment before it can be decrypted
  *	@slow_gro: state present at GRO time, slower prepare step required
  *	@tstamp_type: When set, skb->tstamp has the
  *		delivery_time clock base of skb->tstamp.
@@ -1025,6 +1029,7 @@ struct sk_buff {
 #endif
 #ifdef CONFIG_SKB_DECRYPTED
 	__u8			decrypted:1;
+	__u8			decrypt_failed:1;
 #endif
 	__u8			slow_gro:1;
 #if IS_ENABLED(CONFIG_IP_SCTP)
@@ -1716,6 +1721,7 @@ static inline void skb_copy_decrypted(struct sk_buff *to,
 {
 #ifdef CONFIG_SKB_DECRYPTED
 	to->decrypted = from->decrypted;
+	to->decrypt_failed = from->decrypt_failed;
 #endif
 }
 
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 10/15] net/mlx5e: flag TLS RX records that failed device decryption
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (8 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 09/15] net: skbuff: add skb->decrypt_failed bit Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 11/15] tls: device: add TX KeyUpdate support Rishikesh Jethwani
                   ` (4 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

On CQE_TLS_OFFLOAD_ERROR the device could not authenticate the record's
payload. Depending on where the failure occurred the bytes may have been
transformed (XORed) or left as wire ciphertext, so set the new
skb->decrypt_failed bit (skb->decrypted stays clear) to let the stack
tell the two cases apart, and fall through to the existing tls_err
accounting.

This is consumed by TLS 1.3 device-offload RX KeyUpdate support in a
following patch: the re-encrypt path undoes the transform on any XORed
frag of a mixed record while software re-authenticates, while a
non-mixed record stays wire ciphertext and is decrypted directly.
Without that consumer the flag is simply ignored.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 .../ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c  | 13 ++++++++++++-
 1 file changed, 12 insertions(+), 1 deletion(-)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c
index bca45679e201..8ec40f5fd5b5 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c
@@ -602,7 +602,18 @@ void mlx5e_ktls_handle_rx_skb(struct mlx5e_rq *rq, struct sk_buff *skb,
 		stats->tls_resync_req_pkt++;
 		resync_update_sn(rq, skb);
 		break;
-	default: /* CQE_TLS_OFFLOAD_ERROR: */
+	case CQE_TLS_OFFLOAD_ERROR:
+		/* The device could not authenticate the payload. Depending on
+		 * where the failure occurred the bytes may have been transformed
+		 * (XORed) or left as wire ciphertext. Flag it so that, during a
+		 * TLS 1.3 rekey transition, the re-encrypt path undoes the
+		 * transform on any XORed frag of a mixed record while software
+		 * re-authenticates; a non-mixed record stays wire ciphertext and
+		 * is decrypted directly.
+		 */
+		skb->decrypt_failed = 1;
+		fallthrough;
+	default: /* CQE_TLS_OFFLOAD_NOT_DECRYPTED: */
 		stats->tls_err++;
 		break;
 	}
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 11/15] tls: device: add TX KeyUpdate support
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (9 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 10/15] net/mlx5e: flag TLS RX records that failed device decryption Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 12/15] tls: device: add RX " Rishikesh Jethwani
                   ` (3 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

The NIC key cannot be replaced while HW-offloaded records
are still unacked. tls_device_start_rekey() installs a temporary SW
context with the new key and redirects sendmsg through
tls_sw_sendmsg_locked. If no records are pending,
tls_device_complete_rekey() runs inline during setsockopt; otherwise
tls_tcp_clean_acked sets REKEY_READY once all old-key records are ACKed
and the next sendmsg completes the rekey, flushing SW records and
reinstalling HW offload at the current write_seq. A KeyUpdate
arriving while one is pending re-keys the SW AEAD in place; if the
HW reinstall fails the socket stays in SW mode (REKEY_FAILED).

One side effect touches the non-rekey paths: tx_lock is now taken for
every TLS_TX setsockopt (initial install and SW-only sockets included),
because whether a call is a rekey is only known under lock_sock; this
adds the tx_lock -> lock_sock ordering already used by the data path and
is uncontended during initial setup.

While a rekey is in flight the data path encrypts with the pending key's
SW context, so getsockopt(SOL_TLS, TLS_TX) selects the cipher context
via tls_tx_cipher_ctx(), the same accessor the data path uses, and
reports what sendmsg is actually encrypting with: the pending rekey's
key while one is in flight, otherwise the active key. lock_sock is held,
so rekey.cipher_ctx cannot change under the reader.

Tested on Mellanox ConnectX-6 Dx (Crypto Enabled) with multiple
TLS 1.3 TX KeyUpdate cycles.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 include/net/tls.h             |  83 ++++-
 include/uapi/linux/snmp.h     |   3 +
 net/tls/tls.h                 |   8 +-
 net/tls/tls_device.c          | 631 ++++++++++++++++++++++++++++++++--
 net/tls/tls_device_fallback.c | 128 ++++++-
 net/tls/tls_main.c            |  54 ++-
 net/tls/tls_proc.c            |   3 +
 net/tls/tls_sw.c              |  33 +-
 8 files changed, 895 insertions(+), 48 deletions(-)

diff --git a/include/net/tls.h b/include/net/tls.h
index eb258bcd62bc..b5fc281ff365 100644
--- a/include/net/tls.h
+++ b/include/net/tls.h
@@ -185,6 +185,14 @@ struct tls_offload_context_tx {
 	void (*sk_destruct)(struct sock *sk);
 	struct work_struct destruct_work;
 	struct tls_context *ctx;
+
+	struct {
+		struct tls_sw_context_tx sw;	/* SW context for new key */
+		struct cipher_context tx;	/* IV, rec_seq for new key */
+		union tls_crypto_context crypto_send; /* Crypto for new key */
+		struct tls_record_info *start_marker;
+	} rekey;
+
 	/* The TLS layer reserves room for driver specific state
 	 * Currently the belief is that there is not enough
 	 * driver specific state to justify another layer of indirection
@@ -209,6 +217,28 @@ enum tls_context_flags {
 	 * tls_dev_del call in tls_device_down if it happens simultaneously.
 	 */
 	TLS_RX_DEV_CLOSED = 2,
+	/* TX HW context has been tls_dev_del()'d (mid-rekey before the re-add,
+	 * after a failed re-add, or by tls_device_down()); prevents a second
+	 * tls_dev_del. Cleared when tls_dev_add re-establishes the context.
+	 */
+	TLS_TX_DEV_CLOSED = 3,
+	/* TX rekey is pending, waiting for old-key data to be ACKed.
+	 * While set, new data uses SW path with new key, HW keeps old key
+	 * for retransmissions.
+	 */
+	TLS_TX_REKEY_PENDING = 4,
+	/* All old-key data has been ACKed, ready to install new key in HW. */
+	TLS_TX_REKEY_READY = 5,
+	/* HW rekey failed; TX stays on the SW rekey context until the next
+	 * KeyUpdate re-arms the transition (tls_device_start_rekey()). Also
+	 * stops tls_tcp_clean_acked() from re-setting TLS_TX_REKEY_READY.
+	 */
+	TLS_TX_REKEY_FAILED = 6,
+	/* A rekey has completed on this socket at least once; that arms
+	 * tls_tx_drop_acked_clone() (see its header for the rationale). WARN
+	 * avoidance only.
+	 */
+	TLS_TX_REKEY_FLOOR = 7,
 };
 
 struct tls_prot_info {
@@ -257,6 +287,20 @@ struct tls_context {
 			       */
 	unsigned long flags;
 
+	struct {
+		/* TCP sequence number boundary for pending rekey.
+		 * Packets with seq < this use old key, >= use new key.
+		 */
+		u32 boundary_seq;
+
+		/* SW encryption contexts for the new key, non-NULL only while
+		 * TLS_TX_REKEY_{PENDING,FAILED}; consulted by tls_sw_ctx_tx() and
+		 * tls_tx_cipher_ctx().
+		 */
+		struct tls_sw_context_tx *sw_ctx;
+		struct cipher_context *cipher_ctx;
+	} rekey;
+
 	/* cache cold stuff */
 	struct proto *sk_proto;
 	struct sock *sk;
@@ -356,15 +400,38 @@ tls_validate_xmit_skb(struct sock *sk, struct net_device *dev,
 struct sk_buff *
 tls_validate_xmit_skb_sw(struct sock *sk, struct net_device *dev,
 			 struct sk_buff *skb);
+struct sk_buff *
+tls_validate_xmit_skb_rekey(struct sock *sk, struct net_device *dev,
+			    struct sk_buff *skb);
 
 static inline bool tls_is_skb_tx_device_offloaded(const struct sk_buff *skb)
 {
 #ifdef CONFIG_TLS_DEVICE
 	struct sock *sk = skb->sk;
+	typeof(sk->sk_validate_xmit_skb) validate;
 
-	return sk && sk_fullsock(sk) &&
-	       (smp_load_acquire(&sk->sk_validate_xmit_skb) ==
-	       &tls_validate_xmit_skb);
+	if (!sk || !sk_fullsock(sk))
+		return false;
+
+	/* Pairs with the smp_store_release() that installs or swaps the
+	 * validator (tls_set_device_offload() / tls_device_start_rekey()): the
+	 * pointer read here is published together with the offload state it
+	 * guards, so a non-NULL validator implies that state is visible.
+	 */
+	validate = smp_load_acquire(&sk->sk_validate_xmit_skb);
+	if (likely(validate == &tls_validate_xmit_skb))
+		return true;
+
+	/* A TX rekey (tls_device_start_rekey()) can swap in the rekey validator
+	 * between this skb's validate_xmit_skb(), where the old validator
+	 * passed it through as HW-offload plaintext, and here. A skb->decrypted
+	 * skb under the rekey validator is therefore that straddler: old-key
+	 * plaintext whose HW context is still installed (tls_dev_del() runs in
+	 * tls_device_complete_rekey() only after a synchronize_net() that drains
+	 * this in-flight xmit), so the NIC must still encrypt it. Everything else
+	 * the rekey validator emits is ciphertext (skb->decrypted == 0).
+	 */
+	return validate == &tls_validate_xmit_skb_rekey && skb_is_decrypted(skb);
 #else
 	return false;
 #endif
@@ -389,12 +456,22 @@ static inline struct tls_sw_context_rx *tls_sw_ctx_rx(
 static inline struct tls_sw_context_tx *tls_sw_ctx_tx(
 		const struct tls_context *tls_ctx)
 {
+	struct tls_sw_context_tx *rekey_ctx = READ_ONCE(tls_ctx->rekey.sw_ctx);
+
+	if (unlikely(rekey_ctx))
+		return rekey_ctx;
+
 	return (struct tls_sw_context_tx *)tls_ctx->priv_ctx_tx;
 }
 
 static inline struct cipher_context *tls_tx_cipher_ctx(
 		const struct tls_context *tls_ctx)
 {
+	struct cipher_context *rekey_ctx = READ_ONCE(tls_ctx->rekey.cipher_ctx);
+
+	if (unlikely(rekey_ctx))
+		return rekey_ctx;
+
 	return (struct cipher_context *)&tls_ctx->tx;
 }
 
diff --git a/include/uapi/linux/snmp.h b/include/uapi/linux/snmp.h
index 49f5640092a0..a2e0264641de 100644
--- a/include/uapi/linux/snmp.h
+++ b/include/uapi/linux/snmp.h
@@ -369,6 +369,9 @@ enum
 	LINUX_MIB_TLSTXREKEYOK,			/* TlsTxRekeyOk */
 	LINUX_MIB_TLSTXREKEYERROR,		/* TlsTxRekeyError */
 	LINUX_MIB_TLSRXREKEYRECEIVED,		/* TlsRxRekeyReceived */
+	LINUX_MIB_TLSTXREKEYFALLBACK,		/* TlsTxRekeyFallback */
+	LINUX_MIB_TLSCURRTXREKEY,		/* TlsCurrTxRekey */
+	LINUX_MIB_TLSTXREKEYABORTED,		/* TlsTxRekeyAborted */
 	__LINUX_MIB_TLSMAX
 };
 
diff --git a/net/tls/tls.h b/net/tls/tls.h
index 920a926e8e68..e749f429301a 100644
--- a/net/tls/tls.h
+++ b/net/tls/tls.h
@@ -165,7 +165,10 @@ void tls_update_rx_zc_capable(struct tls_context *tls_ctx);
 void tls_sw_strparser_arm(struct sock *sk, struct tls_context *ctx);
 void tls_sw_strparser_done(struct tls_context *tls_ctx);
 int tls_sw_sendmsg(struct sock *sk, struct msghdr *msg, size_t size);
+int tls_sw_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size);
 void tls_sw_ctx_tx_init(struct sock *sk, struct tls_sw_context_tx *sw_ctx);
+int tls_sw_drain_tx(struct sock *sk, struct tls_context *ctx, int flags);
+int tls_encrypt_async_wait(struct tls_sw_context_tx *ctx);
 int tls_sw_push_pending_record(struct sock *sk, int flags);
 void tls_sw_splice_eof(struct socket *sock);
 void tls_sw_splice_eof_locked(struct socket *sock);
@@ -245,7 +248,8 @@ static inline bool tls_strp_msg_mixed_decrypted(struct tls_sw_context_rx *ctx)
 #ifdef CONFIG_TLS_DEVICE
 int tls_device_init(void);
 void tls_device_cleanup(void);
-int tls_set_device_offload(struct sock *sk);
+int tls_set_device_offload(struct sock *sk,
+			   struct tls_crypto_info *crypto_info);
 void tls_device_free_resources_tx(struct sock *sk);
 int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx);
 void tls_device_offload_cleanup_rx(struct sock *sk);
@@ -256,7 +260,7 @@ static inline int tls_device_init(void) { return 0; }
 static inline void tls_device_cleanup(void) {}
 
 static inline int
-tls_set_device_offload(struct sock *sk)
+tls_set_device_offload(struct sock *sk, struct tls_crypto_info *crypto_info)
 {
 	return -EOPNOTSUPP;
 }
diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index 972c9c7ba7de..f32c1bb6b497 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -57,8 +57,15 @@ static struct page *dummy_page;
 
 static void tls_device_free_ctx(struct tls_context *ctx)
 {
-	if (ctx->tx_conf == TLS_HW)
-		kfree(tls_offload_ctx_tx(ctx));
+	if (ctx->tx_conf == TLS_HW) {
+		struct tls_offload_context_tx *offload_ctx =
+			tls_offload_ctx_tx(ctx);
+
+		kfree(offload_ctx->rekey.start_marker);
+		memzero_explicit(&offload_ctx->rekey,
+				 sizeof(offload_ctx->rekey));
+		kfree(offload_ctx);
+	}
 
 	if (ctx->rx_conf == TLS_HW)
 		kfree(tls_offload_ctx_rx(ctx));
@@ -79,7 +86,9 @@ static void tls_device_tx_del_task(struct work_struct *work)
 	netdev = rcu_dereference_protected(ctx->netdev,
 					   !refcount_read(&ctx->refcount));
 
-	netdev->tlsdev_ops->tls_dev_del(netdev, ctx, TLS_OFFLOAD_CTX_DIR_TX);
+	if (!test_bit(TLS_TX_DEV_CLOSED, &ctx->flags))
+		netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
+						TLS_OFFLOAD_CTX_DIR_TX);
 	dev_put(netdev);
 	ctx->netdev = NULL;
 	tls_device_free_ctx(ctx);
@@ -157,7 +166,10 @@ static int tls_device_dev_add_tx(struct sock *sk, struct net_device *netdev,
 	return rc;
 }
 
-static void tls_device_commit_start_marker(struct sock *sk,
+/* Caller controls locking: initial-offload path is lock-free (pre-publish);
+ * rekey path holds offload_ctx->lock.
+ */
+static void tls_device_add_start_marker(struct sock *sk,
 					struct tls_offload_context_tx *offload_ctx,
 					struct tls_record_info *start_marker_record)
 {
@@ -165,6 +177,13 @@ static void tls_device_commit_start_marker(struct sock *sk,
 	start_marker_record->len = 0;
 	start_marker_record->num_frags = 0;
 	list_add_tail_rcu(&start_marker_record->list, &offload_ctx->records_list);
+}
+
+static void tls_device_commit_start_marker(struct sock *sk,
+					struct tls_offload_context_tx *offload_ctx,
+					struct tls_record_info *start_marker_record)
+{
+	tls_device_add_start_marker(sk, offload_ctx, start_marker_record);
 
 	/* TLS offload is greatly simplified if we don't send
 	 * SKBs where only part of the payload needs to be encrypted.
@@ -194,6 +213,57 @@ static void delete_all_records(struct tls_offload_context_tx *offload_ctx)
 	offload_ctx->retransmit_hint = NULL;
 }
 
+static void tls_device_commit_rekey_marker(struct sock *sk,
+					   struct tls_offload_context_tx *offload_ctx,
+					   struct tls_record_info *start_marker_record)
+{
+	struct tls_record_info *info, *temp;
+	unsigned long flags;
+	__be64 rcd_sn;
+
+	spin_lock_irqsave(&offload_ctx->lock, flags);
+
+	/* The deferred path reaches here with an empty list; the inline
+	 * path may still hold the old start marker (never a real record,
+	 * since tls_has_unacked_records() was false). Only markers are
+	 * ever at the head, so stop at the first non-marker.
+	 */
+	list_for_each_entry_safe(info, temp, &offload_ctx->records_list, list) {
+		if (!tls_record_is_start_marker(info))
+			break;
+		list_del(&info->list);
+		destroy_record(info);
+	}
+	offload_ctx->retransmit_hint = NULL;
+
+	memcpy(&rcd_sn, offload_ctx->rekey.tx.rec_seq, sizeof(rcd_sn));
+	offload_ctx->unacked_record_sn = be64_to_cpu(rcd_sn) - 1;
+
+	tls_device_add_start_marker(sk, offload_ctx, start_marker_record);
+
+	spin_unlock_irqrestore(&offload_ctx->lock, flags);
+
+	tcp_write_collapse_fence(sk);
+}
+
+static bool tls_has_unacked_records(struct tls_offload_context_tx *offload_ctx)
+{
+	struct tls_record_info *info;
+	bool has_unacked = false;
+	unsigned long flags;
+
+	spin_lock_irqsave(&offload_ctx->lock, flags);
+	list_for_each_entry(info, &offload_ctx->records_list, list) {
+		if (!tls_record_is_start_marker(info)) {
+			has_unacked = true;
+			break;
+		}
+	}
+	spin_unlock_irqrestore(&offload_ctx->lock, flags);
+
+	return has_unacked;
+}
+
 static void tls_tcp_clean_acked(struct sock *sk, u32 acked_seq)
 {
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
@@ -222,6 +292,19 @@ static void tls_tcp_clean_acked(struct sock *sk, u32 acked_seq)
 	}
 
 	ctx->unacked_record_sn += deleted_records;
+
+	/* Once all old-key HW records are ACKed, set REKEY_READY to
+	 * let sendmsg know it can finish the rekey and switch back
+	 * to HW offload.
+	 */
+	if (test_bit(TLS_TX_REKEY_PENDING, &tls_ctx->flags) &&
+	    !test_bit(TLS_TX_REKEY_FAILED, &tls_ctx->flags)) {
+		u32 boundary_seq = READ_ONCE(tls_ctx->rekey.boundary_seq);
+
+		if (!before(acked_seq, boundary_seq))
+			set_bit(TLS_TX_REKEY_READY, &tls_ctx->flags);
+	}
+
 	spin_unlock_irqrestore(&ctx->lock, flags);
 }
 
@@ -252,7 +335,15 @@ void tls_device_free_resources_tx(struct sock *sk)
 {
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
 
-	tls_free_partial_record(sk, tls_ctx);
+	if (unlikely(tls_ctx->rekey.sw_ctx))
+		tls_sw_release_resources_tx(sk);
+	else
+		tls_free_partial_record(sk, tls_ctx);
+
+	if (test_bit(TLS_TX_REKEY_PENDING, &tls_ctx->flags)) {
+		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYABORTED);
+		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
+	}
 }
 
 void tls_offload_tx_resync_request(struct sock *sk, u32 got_seq, u32 exp_seq)
@@ -462,6 +553,9 @@ static int tls_device_copy_data(void *addr, size_t bytes, struct iov_iter *i)
 	return 0;
 }
 
+static int tls_device_complete_rekey(struct sock *sk, struct tls_context *ctx,
+				     bool deferred, int push_flags);
+
 static int tls_push_data(struct sock *sk,
 			 struct iov_iter *iter,
 			 size_t size, int flags,
@@ -607,18 +701,46 @@ static int tls_push_data(struct sock *sk,
 	return rc;
 }
 
+/* True while TX is routed through the temporary SW rekey context: a rekey is in
+ * progress (PENDING) or has failed and the socket stays pinned to SW (FAILED).
+ */
+static bool tls_device_tx_uses_sw(const struct tls_context *ctx)
+{
+	return test_bit(TLS_TX_REKEY_PENDING, &ctx->flags) ||
+	       test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
+}
+
 int tls_device_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
 {
 	unsigned char record_type = TLS_RECORD_TYPE_DATA;
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
 	int rc;
 
+	/* Reject unsupported flags up front. tls_push_data() enforces the same
+	 * set, but during a rekey the send is routed to tls_sw_sendmsg_locked(),
+	 * which is the _locked variant and does not re-check; without this,
+	 * MSG_ZEROCOPY / MSG_OOB etc. would reach tcp_sendmsg_locked() on the
+	 * kernel-owned record pages while PENDING/FAILED.
+	 */
+	if (msg->msg_flags & ~(MSG_MORE | MSG_DONTWAIT | MSG_NOSIGNAL |
+			       MSG_SPLICE_PAGES | MSG_EOR))
+		return -EOPNOTSUPP;
+
 	if (!tls_ctx->zerocopy_sendfile)
 		msg->msg_flags &= ~MSG_SPLICE_PAGES;
 
 	mutex_lock(&tls_ctx->tx_lock);
 	lock_sock(sk);
 
+	/* Old-key records all ACKed; switch back to HW. */
+	if (test_bit(TLS_TX_REKEY_READY, &tls_ctx->flags))
+		tls_device_complete_rekey(sk, tls_ctx, true, msg->msg_flags);
+
+	if (tls_device_tx_uses_sw(tls_ctx)) {
+		rc = tls_sw_sendmsg_locked(sk, msg, size);
+		goto out;
+	}
+
 	if (unlikely(msg->msg_controllen)) {
 		rc = tls_process_cmsg(sk, msg, &record_type);
 		if (rc)
@@ -647,8 +769,10 @@ void tls_device_splice_eof(struct socket *sock)
 	mutex_lock(&tls_ctx->tx_lock);
 	lock_sock(sk);
 
-	if (tls_is_partially_sent_record(tls_ctx) ||
-	    tls_is_pending_open_record(tls_ctx)) {
+	if (tls_device_tx_uses_sw(tls_ctx)) {
+		tls_sw_splice_eof_locked(sock);
+	} else if (tls_is_partially_sent_record(tls_ctx) ||
+		   tls_is_pending_open_record(tls_ctx)) {
 		iov_iter_bvec(&iter, ITER_SOURCE, NULL, 0, 0);
 		tls_push_data(sk, &iter, 0, 0, TLS_RECORD_TYPE_DATA);
 	}
@@ -719,14 +843,30 @@ EXPORT_SYMBOL(tls_get_record);
 
 static int tls_device_push_pending_record(struct sock *sk, int flags)
 {
+	struct tls_context *tls_ctx = tls_get_ctx(sk);
 	struct iov_iter iter;
 
+	if (tls_device_tx_uses_sw(tls_ctx))
+		return tls_sw_push_pending_record(sk, flags);
+
 	iov_iter_kvec(&iter, ITER_SOURCE, NULL, 0, 0);
 	return tls_push_data(sk, &iter, 0, flags, TLS_RECORD_TYPE_DATA);
 }
 
 void tls_device_write_space(struct sock *sk, struct tls_context *ctx)
 {
+	if (tls_device_tx_uses_sw(ctx)) {
+		struct tls_offload_context_tx *offload_ctx;
+		unsigned long flags;
+
+		offload_ctx = tls_offload_ctx_tx(ctx);
+		spin_lock_irqsave(&offload_ctx->lock, flags);
+		if (tls_device_tx_uses_sw(ctx))
+			tls_sw_write_space(sk, ctx);
+		spin_unlock_irqrestore(&offload_ctx->lock, flags);
+		return;
+	}
+
 	if (tls_is_partially_sent_record(ctx)) {
 		gfp_t sk_allocation = sk->sk_allocation;
 
@@ -1106,6 +1246,425 @@ static struct tls_offload_context_tx *alloc_offload_ctx_tx(struct tls_context *c
 	return offload_ctx;
 }
 
+/* Build a fresh AEAD tfm for the rekey with the given key, so it can be
+ * swapped in only on success. Re-keying a live tfm in place is not atomic:
+ * a failed crypto_aead_setkey() leaves it with CRYPTO_TFM_NEED_KEY set,
+ * destroying the previous key. Returns an ERR_PTR() on failure.
+ */
+static struct crypto_aead *tls_device_build_rekey_aead(
+				const struct tls_cipher_desc *cipher_desc,
+				char *key, u32 alg_flags)
+{
+	struct crypto_aead *aead;
+	int rc;
+
+	aead = crypto_alloc_aead(cipher_desc->cipher_name, 0, alg_flags);
+	if (IS_ERR(aead))
+		return aead;
+
+	rc = crypto_aead_setkey(aead, key, cipher_desc->key);
+	if (!rc)
+		rc = crypto_aead_setauthsize(aead, cipher_desc->tag);
+	if (rc) {
+		crypto_free_aead(aead);
+		return ERR_PTR(rc);
+	}
+
+	return aead;
+}
+
+static void tls_device_copy_rekey_iv_seq(
+				struct tls_offload_context_tx *offload_ctx,
+				const struct tls_cipher_desc *cipher_desc,
+				char *salt, char *iv, char *rec_seq)
+{
+	memcpy(offload_ctx->rekey.tx.iv, salt, cipher_desc->salt);
+	memcpy(offload_ctx->rekey.tx.iv + cipher_desc->salt, iv,
+	       cipher_desc->iv);
+	memcpy(offload_ctx->rekey.tx.rec_seq, rec_seq, cipher_desc->rec_seq);
+}
+
+static int tls_device_init_rekey_sw(struct sock *sk,
+				    struct tls_context *ctx,
+				    struct tls_offload_context_tx *offload_ctx,
+				    struct tls_crypto_info *new_crypto_info)
+{
+	struct tls_sw_context_tx *sw_ctx = &offload_ctx->rekey.sw;
+	const struct tls_cipher_desc *cipher_desc;
+	char *key;
+	int rc;
+
+	cipher_desc = get_cipher_desc(new_crypto_info->cipher_type);
+	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
+
+	memset(sw_ctx, 0, sizeof(*sw_ctx));
+	tls_sw_ctx_tx_init(sk, sw_ctx);
+
+	key = crypto_info_key(new_crypto_info, cipher_desc);
+	sw_ctx->aead_send = tls_device_build_rekey_aead(cipher_desc, key, 0);
+	if (IS_ERR(sw_ctx->aead_send)) {
+		rc = PTR_ERR(sw_ctx->aead_send);
+		sw_ctx->aead_send = NULL;
+		return rc;
+	}
+
+	return 0;
+}
+
+static int tls_device_start_rekey(struct sock *sk,
+				  struct tls_context *ctx,
+				  struct tls_offload_context_tx *offload_ctx,
+				  struct tls_crypto_info *new_crypto_info)
+{
+	bool rekey_pending = test_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
+	bool rekey_failed = test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
+	const struct tls_cipher_desc *cipher_desc;
+	struct crypto_aead *new_aead, *old_aead;
+	char *key, *iv, *rec_seq, *salt;
+	int push_flags = MSG_NOSIGNAL;
+	unsigned long flags;
+	int rc;
+
+	cipher_desc = get_cipher_desc(new_crypto_info->cipher_type);
+	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
+
+	key = crypto_info_key(new_crypto_info, cipher_desc);
+	iv = crypto_info_iv(new_crypto_info, cipher_desc);
+	rec_seq = crypto_info_rec_seq(new_crypto_info, cipher_desc);
+	salt = crypto_info_salt(new_crypto_info, cipher_desc);
+
+	/* The record flushes below hand the open/partially sent HW record to
+	 * TCP and may have to wait for send buffer space. Honour the socket's
+	 * non-blocking mode so an O_NONBLOCK application is not put to sleep
+	 * inside setsockopt(): it gets -EAGAIN and retries once the socket is
+	 * writable. Kernel sockets (no backing file, e.g. nvme-tcp) keep the
+	 * blocking semantics, matching how they call sendmsg().
+	 */
+	if (sk->sk_socket && sk->sk_socket->file &&
+	    (sk->sk_socket->file->f_flags & O_NONBLOCK))
+		push_flags |= MSG_DONTWAIT;
+
+	if (rekey_pending || rekey_failed) {
+		/* Flush any SW open_record before swapping the key. -EINPROGRESS
+		 * means an async AEAD accepted the record for encryption; it is a
+		 * success, waited for by tls_encrypt_async_wait() just below (as
+		 * tls_process_cmsg()/tls_sw_drain_tx() also treat it).
+		 */
+		if (tls_is_pending_open_record(ctx)) {
+			rc = ctx->push_pending_record(sk, push_flags);
+			if (rc < 0 && rc != -EINPROGRESS)
+				return rc;
+		}
+
+		/* Wait for in-flight async encryptions submitted to this tfm
+		 * with the previous key before changing it.
+		 */
+		rc = tls_encrypt_async_wait(&offload_ctx->rekey.sw);
+		if (rc)
+			return rc;
+
+		/* Build the new key into a fresh tfm and swap it in only on
+		 * success; A failed rekey here must leave the SW fallback
+		 * path able to encrypt.
+		 */
+		new_aead = tls_device_build_rekey_aead(cipher_desc, key, 0);
+		if (IS_ERR(new_aead))
+			return PTR_ERR(new_aead);
+
+		old_aead = offload_ctx->rekey.sw.aead_send;
+		offload_ctx->rekey.sw.aead_send = new_aead;
+		crypto_free_aead(old_aead);
+
+		tls_device_copy_rekey_iv_seq(offload_ctx, cipher_desc,
+					     salt, iv, rec_seq);
+
+		if (rekey_failed) {
+			/* Re-arm FAILED -> PENDING under device_offload_lock. The
+			 * PENDING set and FAILED clear are two stores to ctx->flags,
+			 * and tls_device_down() tests !PENDING && !FAILED as two
+			 * separate loads; without the lock those loads could straddle
+			 * the flip and see neither bit, letting tls_device_down()
+			 * install tls_validate_xmit_skb_sw with PENDING set (dropping
+			 * all new-key ciphertext). The lock keeps PENDING || FAILED
+			 * observable throughout. Non-blocking, so no NETDEV_DOWN stall.
+			 */
+			down_read(&device_offload_lock);
+			spin_lock_irqsave(&offload_ctx->lock, flags);
+			WRITE_ONCE(ctx->rekey.boundary_seq, tcp_sk(sk)->snd_una);
+			set_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
+			spin_unlock_irqrestore(&offload_ctx->lock, flags);
+			/* Release pairs with test_bit_acquire() in the validator:
+			 * a TX seeing FAILED clear must see the fresh boundary_seq.
+			 */
+			clear_bit_unlock(TLS_TX_REKEY_FAILED, &ctx->flags);
+			up_read(&device_offload_lock);
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXSW);
+			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
+		}
+	} else {
+		/* Drain partially sent record and flush open HW record
+		 * before switching to SW.
+		 */
+		if (tls_is_partially_sent_record(ctx)) {
+			rc = tls_push_partial_record(sk, ctx,
+						     MSG_SENDPAGE_DECRYPTED |
+						     push_flags);
+			if (rc < 0)
+				return rc;
+		}
+		if (tls_is_pending_open_record(ctx)) {
+			rc = ctx->push_pending_record(sk, push_flags);
+			if (rc < 0)
+				return rc;
+		}
+
+		rc = tls_device_init_rekey_sw(sk, ctx, offload_ctx,
+					      new_crypto_info);
+		if (rc)
+			return rc;
+
+		tls_device_copy_rekey_iv_seq(offload_ctx, cipher_desc,
+					     salt, iv, rec_seq);
+
+		/* Publish the rekey under device_offload_lock so that setting
+		 * TLS_TX_REKEY_PENDING and installing the rekey validator is
+		 * atomic against tls_device_down(), which under down_write() tests
+		 * !PENDING and installs tls_validate_xmit_skb_sw. Otherwise the two
+		 * validator stores could interleave to leave PENDING set with the
+		 * SW validator, and every new-key ciphertext (never on the offload
+		 * records_list) would then be dropped by tls_sw_fallback(). The
+		 * blocking flush and crypto_alloc above deliberately run WITHOUT
+		 * this lock, so a stalled peer cannot hold up NETDEV_DOWN (which
+		 * takes down_write() under RTNL) or any other down_read() user.
+		 */
+		down_read(&device_offload_lock);
+
+		/* Prevent a partial record straddling the SW/HW boundary. */
+		tcp_write_collapse_fence(sk);
+
+		WRITE_ONCE(ctx->rekey.sw_ctx, &offload_ctx->rekey.sw);
+		WRITE_ONCE(ctx->rekey.cipher_ctx, &offload_ctx->rekey.tx);
+
+		spin_lock_irqsave(&offload_ctx->lock, flags);
+		WRITE_ONCE(ctx->rekey.boundary_seq, tcp_sk(sk)->write_seq);
+		set_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
+		spin_unlock_irqrestore(&offload_ctx->lock, flags);
+
+		/* Switch to rekey validator; new sends won't use HW offload */
+		smp_store_release(&sk->sk_validate_xmit_skb,
+				  tls_validate_xmit_skb_rekey);
+
+		up_read(&device_offload_lock);
+	}
+
+	unsafe_memcpy(&offload_ctx->rekey.crypto_send.info, new_crypto_info,
+		      cipher_desc->crypto_info,
+		      /* checked in do_tls_setsockopt_conf */);
+	memzero_explicit(new_crypto_info, cipher_desc->crypto_info);
+
+	return 0;
+}
+
+static int tls_device_complete_rekey(struct sock *sk, struct tls_context *ctx,
+				     bool deferred, int push_flags)
+{
+	struct tls_offload_context_tx *offload_ctx = tls_offload_ctx_tx(ctx);
+	struct crypto_aead *new_aead, *old_aead, *old_sw_aead;
+	const struct tls_cipher_desc *cipher_desc;
+	struct net_device *netdev;
+	unsigned long flags;
+	char *key;
+	int rc;
+
+	cipher_desc = get_cipher_desc(offload_ctx->rekey.crypto_send.info.cipher_type);
+	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
+
+	DEBUG_NET_WARN_ON_ONCE(!offload_ctx->rekey.start_marker);
+
+	rc = tls_sw_drain_tx(sk, ctx, push_flags);
+	/* -EAGAIN (sndbuf full) and a signal (-EINTR/-ERESTARTSYS from
+	 * sk_stream_wait_memory()) are transient: leave the rekey PENDING and
+	 * retry on the next sendmsg rather than permanently dropping HW offload.
+	 * tls_tx_records() likewise passes these through without aborting.
+	 */
+	if (rc == -EAGAIN || rc == -EINTR || rc == -ERESTARTSYS)
+		return rc;
+	if (rc)
+		goto rekey_fallback;	/* hard failure: fall back to SW */
+
+	down_read(&device_offload_lock);
+
+	netdev = rcu_dereference_protected(ctx->netdev,
+					   lockdep_is_held(&device_offload_lock));
+	if (!netdev) {
+		rc = -ENODEV;
+		goto release_lock;
+	}
+
+	/* Drain in-flight xmit users before tls_dev_del() and before freeing the
+	 * old fallback aead_send: (1) under the rekey validator a decrypted
+	 * straddler may still be inside the driver on the HW context (same swap ->
+	 * synchronize_net -> dev_del order as tls_device_down(), which also keeps a
+	 * decrypted skb from reaching a torn-down context); (2) pre-boundary
+	 * retransmits routed to tls_sw_fallback() read aead_send locklessly. No new
+	 * fallback can start here: every pre-boundary record is ACKed and freed, so
+	 * fill_sg_in() bails.
+	 */
+	synchronize_net();
+
+	if (!test_bit(TLS_TX_DEV_CLOSED, &ctx->flags)) {
+		netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
+						TLS_OFFLOAD_CTX_DIR_TX);
+		set_bit(TLS_TX_DEV_CLOSED, &ctx->flags);
+	}
+
+	/* Build the new SW-fallback key into a fresh tfm and swap it in only
+	 * on success. Doing this while the HW context is torn down
+	 * (TLS_TX_DEV_CLOSED set) means a failure falls into rekey_fallback
+	 * with HW off, so the SW fallback is coherent, same as a dev_add
+	 * failure.
+	 */
+	key = crypto_info_key(&offload_ctx->rekey.crypto_send.info, cipher_desc);
+	new_aead = tls_device_build_rekey_aead(cipher_desc, key, CRYPTO_ALG_ASYNC);
+	if (IS_ERR(new_aead)) {
+		rc = PTR_ERR(new_aead);
+		goto release_lock;
+	}
+
+	/* crypto_send.info.rec_seq is frozen at setsockopt time; the SW context
+	 * advanced rekey.tx.rec_seq for every record it sent, so hand the NIC the
+	 * live record number (mirrors the RX deferred add).
+	 */
+	memcpy(crypto_info_rec_seq(&offload_ctx->rekey.crypto_send.info, cipher_desc),
+	       offload_ctx->rekey.tx.rec_seq, cipher_desc->rec_seq);
+
+	rc = tls_device_dev_add_tx(sk, netdev, &offload_ctx->rekey.crypto_send.info,
+				   tcp_sk(sk)->write_seq);
+	if (rc) {
+		crypto_free_aead(new_aead);
+		goto release_lock;
+	}
+
+	/* Point of no return: HW is live with the new key. Swap in the new
+	 * fallback tfm and drop the old one; the remaining steps cannot fail.
+	 */
+	old_aead = offload_ctx->aead_send;
+	offload_ctx->aead_send = new_aead;
+	crypto_free_aead(old_aead);
+	clear_bit(TLS_TX_DEV_CLOSED, &ctx->flags);
+
+	memcpy(ctx->tx.iv, offload_ctx->rekey.tx.iv,
+	       cipher_desc->salt + cipher_desc->iv);
+	memcpy(ctx->tx.rec_seq, offload_ctx->rekey.tx.rec_seq,
+	       cipher_desc->rec_seq);
+	unsafe_memcpy(&ctx->crypto_send.info,
+		      &offload_ctx->rekey.crypto_send.info,
+		      cipher_desc->crypto_info,
+		      /* checked during rekey setup */);
+
+	/* Start marker: the NIC passes through everything before
+	 * write_seq untouched (it is already SW-encrypted ciphertext),
+	 * same as during initial offload setup. Also drops the stale
+	 * marker and rebases unacked_record_sn so the record-sequence
+	 * bookkeeping stays consistent on the inline path.
+	 */
+	tls_device_commit_rekey_marker(sk, offload_ctx,
+				       offload_ctx->rekey.start_marker);
+
+	old_sw_aead = tls_sw_ctx_tx(ctx)->aead_send;
+
+	spin_lock_irqsave(&offload_ctx->lock, flags);
+	clear_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
+	clear_bit(TLS_TX_REKEY_READY, &ctx->flags);
+	clear_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
+
+	/* Arm the drop floor before restoring the HW validator: from now on
+	 * tls_validate_xmit_skb() drops payload retransmits of fully-ACKed data, so
+	 * a stale clone whose record was purged here does not reach the NIC and trip
+	 * its WARN on the new start marker. The cleartext leak on that path is closed
+	 * separately by the skb_is_decrypted() gate in tls_sw_fallback(); this is
+	 * only WARN avoidance. Set once; stays set for the socket's life.
+	 */
+	set_bit(TLS_TX_REKEY_FLOOR, &ctx->flags);
+
+	/* Switch back to HW offload validator */
+	smp_store_release(&sk->sk_validate_xmit_skb, tls_validate_xmit_skb);
+
+	WRITE_ONCE(ctx->rekey.sw_ctx, NULL);
+	WRITE_ONCE(ctx->rekey.cipher_ctx, NULL);
+	spin_unlock_irqrestore(&offload_ctx->lock, flags);
+
+	memzero_explicit(&offload_ctx->rekey, sizeof(offload_ctx->rekey));
+	crypto_free_aead(old_sw_aead);
+
+	up_read(&device_offload_lock);
+
+	if (deferred)
+		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
+	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYOK);
+	return 0;
+
+release_lock:
+	up_read(&device_offload_lock);
+
+rekey_fallback:
+	kfree(offload_ctx->rekey.start_marker);
+	offload_ctx->rekey.start_marker = NULL;
+	spin_lock_irqsave(&offload_ctx->lock, flags);
+	set_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
+	clear_bit(TLS_TX_REKEY_READY, &ctx->flags);
+	clear_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
+	spin_unlock_irqrestore(&offload_ctx->lock, flags);
+	if (deferred)
+		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
+	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYFALLBACK);
+	TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
+	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXSW);
+
+	return 0;
+}
+
+static int tls_set_device_offload_rekey(struct sock *sk,
+					struct tls_context *ctx,
+					struct tls_crypto_info *new_crypto_info)
+{
+	struct tls_offload_context_tx *offload_ctx = tls_offload_ctx_tx(ctx);
+	bool rekey_pending = test_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
+	bool rekey_failed = test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
+	bool defer = true;
+	int rc;
+
+	/* Defer the switch back to HW until any in-flight old-key records are
+	 * ACKed. A partially_sent_record needs no separate check: its record is
+	 * on records_list before it is sent (tls_push_record()) and stays there
+	 * until ACKed, so tls_has_unacked_records() already covers it.
+	 */
+	if (!rekey_pending && !rekey_failed)
+		defer = tls_has_unacked_records(offload_ctx) ||
+			tls_is_pending_open_record(ctx);
+
+	if (!offload_ctx->rekey.start_marker) {
+		offload_ctx->rekey.start_marker =
+			kmalloc_obj(*offload_ctx->rekey.start_marker);
+		if (!offload_ctx->rekey.start_marker)
+			return -ENOMEM;
+	}
+
+	rc = tls_device_start_rekey(sk, ctx, offload_ctx, new_crypto_info);
+	if (rc)
+		return rc;
+
+	if (defer) {
+		if (!rekey_pending)
+			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
+		else
+			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYOK);
+		return 0;
+	}
+
+	return tls_device_complete_rekey(sk, ctx, false, 0);
+}
+
 static int tls_set_device_offload_initial(struct sock *sk,
 					  struct tls_context *ctx,
 					  struct net_device *netdev,
@@ -1190,25 +1749,39 @@ static int tls_set_device_offload_initial(struct sock *sk,
 	return rc;
 }
 
-int tls_set_device_offload(struct sock *sk)
+int tls_set_device_offload(struct sock *sk,
+			   struct tls_crypto_info *new_crypto_info)
 {
+	struct tls_crypto_info *crypto_info, *src_crypto_info;
 	const struct tls_cipher_desc *cipher_desc;
-	struct tls_crypto_info *crypto_info;
 	struct net_device *netdev;
 	struct tls_context *ctx;
 	int rc;
 
 	ctx = tls_get_ctx(sk);
 
-	/* A rekey (setsockopt on an already-configured socket) is not
-	 * supported on the device offload path yet; reject it here so the
-	 * caller can decide (propagate the error for a HW connection, or
-	 * re-init software crypto for a SW one). KeyUpdate support replaces
-	 * this guard with real rekey handling.
+	/* A rekey of a SW-offloaded socket belongs to tls_set_sw_offload(). */
+	if (new_crypto_info && ctx->tx_conf != TLS_HW)
+		return -EINVAL;
+
+	crypto_info = &ctx->crypto_send.info;
+	src_crypto_info = new_crypto_info ?: crypto_info;
+	cipher_desc = get_cipher_desc(src_crypto_info->cipher_type);
+	if (!cipher_desc || !cipher_desc->offloadable)
+		return -EINVAL;
+
+	/* A rekey targets the device already holding the HW TX context
+	 * (ctx->netdev), which can differ from the socket's current route after
+	 * a route change or bond/team failover; tls_set_device_offload_rekey()
+	 * and tls_device_complete_rekey() resolve it from ctx->netdev under
+	 * device_offload_lock. Only the initial install needs the route device.
 	 */
-	if (ctx->tx_conf != TLS_BASE)
-		return -EOPNOTSUPP;
+	if (new_crypto_info)
+		return tls_set_device_offload_rekey(sk, ctx, src_crypto_info);
 
+	/* Initial install: a HW TX context must not already exist, otherwise
+	 * alloc_offload_ctx_tx() below would silently overwrite it.
+	 */
 	if (ctx->priv_ctx_tx)
 		return -EEXIST;
 
@@ -1223,14 +1796,7 @@ int tls_set_device_offload(struct sock *sk)
 		goto release_netdev;
 	}
 
-	crypto_info = &ctx->crypto_send.info;
-	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
-	if (!cipher_desc || !cipher_desc->offloadable) {
-		rc = -EINVAL;
-		goto release_netdev;
-	}
-
-	rc = tls_set_device_offload_initial(sk, ctx, netdev, crypto_info,
+	rc = tls_set_device_offload_initial(sk, ctx, netdev, src_crypto_info,
 					    cipher_desc);
 
 release_netdev:
@@ -1370,10 +1936,16 @@ static int tls_device_down(struct net_device *netdev)
 	spin_unlock_irqrestore(&tls_device_lock, flags);
 
 	list_for_each_entry_safe(ctx, tmp, &list, list)	{
-		/* Stop offloaded TX and switch to the fallback.
-		 * tls_is_skb_tx_device_offloaded will return false.
+		/* Stop offloaded TX and switch to the fallback. For a socket not
+		 * mid-rekey, tls_is_skb_tx_device_offloaded() then returns false; a
+		 * PENDING/FAILED socket keeps the rekey validator (under which only a
+		 * decrypted straddler still offloads), and the synchronize_net()
+		 * below drains any such in-flight skb before tls_dev_del().
 		 */
-		WRITE_ONCE(ctx->sk->sk_validate_xmit_skb, tls_validate_xmit_skb_sw);
+		if (!test_bit(TLS_TX_REKEY_PENDING, &ctx->flags) &&
+		    !test_bit(TLS_TX_REKEY_FAILED, &ctx->flags))
+			WRITE_ONCE(ctx->sk->sk_validate_xmit_skb,
+				   tls_validate_xmit_skb_sw);
 
 		/* Stop the RX and TX resync.
 		 * tls_dev_resync must not be called after tls_dev_del.
@@ -1390,9 +1962,12 @@ static int tls_device_down(struct net_device *netdev)
 		synchronize_net();
 
 		/* Release the offload context on the driver side. */
-		if (ctx->tx_conf == TLS_HW)
+		if (ctx->tx_conf == TLS_HW &&
+		    !test_bit(TLS_TX_DEV_CLOSED, &ctx->flags)) {
 			netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
 							TLS_OFFLOAD_CTX_DIR_TX);
+			set_bit(TLS_TX_DEV_CLOSED, &ctx->flags);
+		}
 		if (ctx->rx_conf == TLS_HW &&
 		    !test_bit(TLS_RX_DEV_CLOSED, &ctx->flags))
 			netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
diff --git a/net/tls/tls_device_fallback.c b/net/tls/tls_device_fallback.c
index 1110f7ac6bcb..f2a0ae827bb2 100644
--- a/net/tls/tls_device_fallback.c
+++ b/net/tls/tls_device_fallback.c
@@ -190,6 +190,14 @@ static void complete_skb(struct sk_buff *nskb, struct sk_buff *skb, int headln)
 
 	skb_copy_header(nskb, skb);
 
+	/* nskb now carries ciphertext, but skb_copy_header() inherited
+	 * skb->decrypted from the plaintext original. Clear it so the bit keeps
+	 * meaning "still-plaintext, needs an encryptor": otherwise a requeued
+	 * nskb would be needlessly re-validated (and re-encrypted) and would trip
+	 * the NIC's decrypted-vs-start-marker WARN.
+	 */
+	nskb->decrypted = 0;
+
 	skb_put(nskb, skb->len);
 	memcpy(nskb->data, skb->data, headln);
 
@@ -396,8 +404,17 @@ static struct sk_buff *tls_sw_fallback(struct sock *sk, struct sk_buff *skb)
 	sg_init_table(sg_out, ARRAY_SIZE(sg_out));
 
 	if (fill_sg_in(sg_in, skb, ctx, &rcd_sn, &sync_size, &resync_sgs)) {
-		/* bypass packets before kernel TLS socket option was set */
-		if (sync_size < 0 && payload_len <= -sync_size)
+		/* Below the record range (start marker / already-freed record).
+		 * Pass through only cleartext that was never offload-encrypted
+		 * (skb->decrypted == 0): genuine pre-TLS bytes sent before the
+		 * socket option was set, or SW-encrypted rekey ciphertext. A
+		 * decrypted=1 skb here is offload-record plaintext whose record was
+		 * purged (e.g. a rekey installed a new start marker above its seq);
+		 * it must never reach the wire in the clear, so continue on and
+		 * drop it (nskb stays NULL).
+		 */
+		if (sync_size < 0 && payload_len <= -sync_size &&
+		    !skb_is_decrypted(skb))
 			nskb = skb_get(skb);
 		goto put_sg;
 	}
@@ -416,11 +433,57 @@ static struct sk_buff *tls_sw_fallback(struct sock *sk, struct sk_buff *skb)
 	return nskb;
 }
 
+/* Post-rekey drop floor. Once a rekey has completed (TLS_TX_REKEY_FLOOR set), a
+ * stale retransmit clone of already-ACKed data may still be dequeued from a
+ * qdisc; if its offload record was purged at completion it now maps to a rekey
+ * start marker. The cleartext leak on that path is closed unconditionally by
+ * the skb_is_decrypted() gate in tls_sw_fallback(); this floor additionally
+ * drops the clone before it reaches the NIC, avoiding the driver's WARN
+ * (mlx5e_ktls_handle_tx_skb() SKIP_NO_DATA) on an otherwise-legitimate race.
+ * Only needed by tls_validate_xmit_skb() (the restored HW-offload validator):
+ * only there can a purged-record clone reach the NIC and hit the new start
+ * marker. Under the rekey/SW validators the only skb the NIC offloads is a
+ * decrypted straddler whose record is still present (no SKIP_NO_DATA), and a
+ * stale clone is dropped by the skb_is_decrypted() gate in tls_sw_fallback().
+ * Such a clone is exactly a payload skb whose end_seq <= snd_una: the peer has
+ * already ACKed that data, so dropping it is always safe. Live/unacked data
+ * (including a legitimate retransmit, or a straddler ending past snd_una) is
+ * never touched; pure ACKs and zero-window probes carry no payload and pass.
+ */
+static bool tls_tx_drop_acked_clone(struct sock *sk, struct sk_buff *skb)
+{
+	int payload_len = skb->len - skb_tcp_all_headers(skb);
+	u32 end_seq;
+
+	if (likely(!test_bit(TLS_TX_REKEY_FLOOR, &tls_get_ctx(sk)->flags)))
+		return false;
+
+	if (payload_len <= 0)
+		return false;
+
+	/* Drop only when the whole payload is already ACKed (end_seq <= snd_una):
+	 * such a skb is purely a stale retransmit clone the peer already has. A
+	 * clone straddling snd_una still carries unacked bytes, so leave it to the
+	 * normal paths (a live record is re-encrypted; a marker/freed-record hit is
+	 * dropped there too). Both the leak (skb_is_decrypted() gate) and the mlx5
+	 * WARN only concern the fully-ACKed case handled here.
+	 */
+	end_seq = ntohl(tcp_hdr(skb)->seq) + payload_len;
+	return !after(end_seq, READ_ONCE(tcp_sk(sk)->snd_una));
+}
+
 struct sk_buff *tls_validate_xmit_skb(struct sock *sk,
 				      struct net_device *dev,
 				      struct sk_buff *skb)
 {
-	if (dev == rcu_dereference_bh(tls_get_ctx(sk)->netdev) ||
+	struct tls_context *tls_ctx = tls_get_ctx(sk);
+
+	if (unlikely(tls_tx_drop_acked_clone(sk, skb))) {
+		kfree_skb(skb);
+		return NULL;
+	}
+
+	if (dev == rcu_dereference_bh(tls_ctx->netdev) ||
 	    netif_is_bond_master(dev))
 		return skb;
 
@@ -435,6 +498,65 @@ struct sk_buff *tls_validate_xmit_skb_sw(struct sock *sk,
 	return tls_sw_fallback(sk, skb);
 }
 
+struct sk_buff *tls_validate_xmit_skb_rekey(struct sock *sk,
+					    struct net_device *dev,
+					    struct sk_buff *skb)
+{
+	struct tls_context *tls_ctx = tls_get_ctx(sk);
+	u32 tcp_seq = ntohl(tcp_hdr(skb)->seq);
+	u32 pivot_seq;
+
+	/* acquire pairs with clear_bit_unlock() on re-arm; makes the refreshed
+	 * boundary_seq visible in the else branch below.
+	 */
+	if (test_bit_acquire(TLS_TX_REKEY_FAILED, &tls_ctx->flags)) {
+		int payload_len = skb->len - skb_tcp_all_headers(skb);
+		u32 snd_una = READ_ONCE(tcp_sk(sk)->snd_una);
+
+		/* FAILED: HW context gone and all old-key plaintext ACKed
+		 * (snd_una >= boundary_seq). seq < boundary_seq is old-key data
+		 * whose records are freed, so tls_sw_fallback() drops it. seq >=
+		 * boundary_seq is SW ciphertext with no record. A retransmit is
+		 * built at seq == snd_una (tcp_trim_head()), so an ACK landing
+		 * before we run can move snd_una past seq while the tail is
+		 * unacked; pivoting on snd_una alone would drop that live data
+		 * and force an RTO. Pass through any non-decrypted skb ending
+		 * past snd_una (mirrors tls_tx_drop_acked_clone()); fully-ACKed
+		 * clones fall to the pivot and are dropped.
+		 */
+		if (payload_len > 0 && !skb_is_decrypted(skb) &&
+		    after(tcp_seq + payload_len, snd_una))
+			return skb;
+
+		pivot_seq = snd_una;
+	} else {
+		/* PENDING: new-key data is SW-encrypted at seq >= boundary_seq;
+		 * old-key data below it is still unacked.
+		 *
+		 * On the first arm, boundary_seq is published by the
+		 * smp_store_release() of sk_validate_xmit_skb in
+		 * tls_device_start_rekey(); the xmit path loads that pointer with a
+		 * plain read (net/core/dev.c), so pair it here with an smp_rmb()
+		 * before reading boundary_seq. A stale boundary_seq (0) would pass an
+		 * unacked old-key plaintext skb through; tls_is_skb_tx_device_offloaded()
+		 * would still HW-encrypt it with the installed old key, so not a leak,
+		 * but the barrier keeps the pivot accurate.
+		 */
+		smp_rmb();
+		pivot_seq = READ_ONCE(tls_ctx->rekey.boundary_seq);
+	}
+
+	/* At or after the pivot: already correctly encrypted, pass through */
+	if (!before(tcp_seq, pivot_seq))
+		return skb;
+
+	/* Below the pivot: retransmit of old data, SW fallback with old key */
+	return tls_sw_fallback(sk, skb);
+}
+
+/* Address taken by tls_is_skb_tx_device_offloaded() in the offload drivers. */
+EXPORT_SYMBOL_GPL(tls_validate_xmit_skb_rekey);
+
 struct sk_buff *tls_encrypt_skb(struct sk_buff *skb)
 {
 	return tls_sw_fallback(skb->sk, skb);
diff --git a/net/tls/tls_main.c b/net/tls/tls_main.c
index 15e83e853f22..3dd3a4ce8209 100644
--- a/net/tls/tls_main.c
+++ b/net/tls/tls_main.c
@@ -347,8 +347,14 @@ static void tls_sk_proto_cleanup(struct sock *sk,
 		tls_sw_release_resources_tx(sk);
 		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXSW);
 	} else if (ctx->tx_conf == TLS_HW) {
+		bool rekey_failed = test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
+
 		tls_device_free_resources_tx(sk);
-		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
+
+		if (rekey_failed)
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXSW);
+		else
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
 	}
 
 	if (ctx->rx_conf == TLS_SW) {
@@ -369,6 +375,8 @@ static void tls_sk_proto_close(struct sock *sk, long timeout)
 
 	if (ctx->tx_conf == TLS_SW)
 		tls_sw_cancel_work_tx(ctx);
+	else if (ctx->tx_conf == TLS_HW && ctx->rekey.sw_ctx)
+		tls_sw_cancel_work_tx(ctx);
 
 	lock_sock(sk);
 	free_ctx = ctx->tx_conf != TLS_HW && ctx->rx_conf != TLS_HW;
@@ -445,8 +453,17 @@ static int do_tls_getsockopt_conf(struct sock *sk, sockopt_t *opt, int tx)
 
 	/* get user crypto info */
 	if (tx) {
-		crypto_info = &ctx->crypto_send.info;
-		cctx = &ctx->tx;
+		/* Select the cipher context via the same accessor the data path
+		 * uses, so getsockopt reports the IV/rec_seq that sendmsg encrypts
+		 * with (the pending rekey's while one is in flight, else the
+		 * active key). crypto_info has no accessor; select it the same way.
+		 * lock_sock is held, so rekey.cipher_ctx cannot change under us.
+		 */
+		cctx = tls_tx_cipher_ctx(ctx);
+		if (ctx->rekey.cipher_ctx)
+			crypto_info = &tls_offload_ctx_tx(ctx)->rekey.crypto_send.info;
+		else
+			crypto_info = &ctx->crypto_send.info;
 	} else {
 		crypto_info = &ctx->crypto_recv.info;
 		cctx = &ctx->rx;
@@ -710,7 +727,7 @@ static int do_tls_setsockopt_conf(struct sock *sk, sockptr_t optval,
 	}
 
 	if (tx) {
-		rc = tls_set_device_offload(sk);
+		rc = tls_set_device_offload(sk, update ? crypto_info : NULL);
 		conf = TLS_HW;
 		if (!rc) {
 			if (!update) {
@@ -787,7 +804,11 @@ static int do_tls_setsockopt_conf(struct sock *sk, sockptr_t optval,
 	return 0;
 
 err_crypto_info:
-	if (update) {
+	/* -EAGAIN is a transient sndbuf-full condition on a non-blocking rekey,
+	 * not a failed KeyUpdate: the old key stays installed and userspace
+	 * retries once the socket is writable, so don't count it as an error.
+	 */
+	if (update && rc != -EAGAIN) {
 		TLS_INC_STATS(sock_net(sk), tx ? LINUX_MIB_TLSTXREKEYERROR
 					       : LINUX_MIB_TLSRXREKEYERROR);
 	}
@@ -880,12 +901,29 @@ static int do_tls_setsockopt(struct sock *sk, int optname, sockptr_t optval,
 
 	switch (optname) {
 	case TLS_TX:
-	case TLS_RX:
+	case TLS_RX: {
+		/* tls_device_sendmsg() holds tx_lock across the lock_sock drop
+		 * in sk_stream_wait_memory() with a half-built open_record
+		 * exposed. A concurrent HW-offload rekey (tls_device_start_rekey())
+		 * would flush that record and swap the key under the sender,
+		 * corrupting record framing. Serialize TX setsockopt against
+		 * the data path with tx_lock, unconditionally for TLS_TX,
+		 * since during initial setup there is no sender contending it.
+		 */
+		bool tx = optname == TLS_TX;
+
+		if (tx) {
+			rc = mutex_lock_interruptible(&tls_get_ctx(sk)->tx_lock);
+			if (rc)
+				break;
+		}
 		lock_sock(sk);
-		rc = do_tls_setsockopt_conf(sk, optval, optlen,
-					    optname == TLS_TX);
+		rc = do_tls_setsockopt_conf(sk, optval, optlen, tx);
 		release_sock(sk);
+		if (tx)
+			mutex_unlock(&tls_get_ctx(sk)->tx_lock);
 		break;
+	}
 	case TLS_TX_ZEROCOPY_RO:
 		lock_sock(sk);
 		rc = do_tls_setsockopt_tx_zc(sk, optval, optlen);
diff --git a/net/tls/tls_proc.c b/net/tls/tls_proc.c
index 4012c4372d4c..4bb1e3727e28 100644
--- a/net/tls/tls_proc.c
+++ b/net/tls/tls_proc.c
@@ -27,6 +27,9 @@ static const struct snmp_mib tls_mib_list[] = {
 	SNMP_MIB_ITEM("TlsTxRekeyOk", LINUX_MIB_TLSTXREKEYOK),
 	SNMP_MIB_ITEM("TlsTxRekeyError", LINUX_MIB_TLSTXREKEYERROR),
 	SNMP_MIB_ITEM("TlsRxRekeyReceived", LINUX_MIB_TLSRXREKEYRECEIVED),
+	SNMP_MIB_ITEM("TlsTxRekeyFallback", LINUX_MIB_TLSTXREKEYFALLBACK),
+	SNMP_MIB_ITEM("TlsCurrTxRekey", LINUX_MIB_TLSCURRTXREKEY),
+	SNMP_MIB_ITEM("TlsTxRekeyAborted", LINUX_MIB_TLSTXREKEYABORTED),
 };
 
 static int tls_statistics_seq_show(struct seq_file *seq, void *v)
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index 5531303dd704..fd162d8f1d64 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -522,7 +522,7 @@ static void tls_encrypt_done(void *data, int err)
 		complete(&ctx->async_wait.completion);
 }
 
-static int tls_encrypt_async_wait(struct tls_sw_context_tx *ctx)
+int tls_encrypt_async_wait(struct tls_sw_context_tx *ctx)
 {
 	if (!atomic_dec_and_test(&ctx->encrypt_pending))
 		crypto_wait_req(-EINPROGRESS, &ctx->async_wait);
@@ -763,8 +763,7 @@ static int tls_sw_sendmsg_splice(struct sock *sk, struct msghdr *msg,
 	return 0;
 }
 
-static int tls_sw_sendmsg_locked(struct sock *sk, struct msghdr *msg,
-				 size_t size)
+int tls_sw_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size)
 {
 	long timeo = sock_sndtimeo(sk, msg->msg_flags & MSG_DONTWAIT);
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
@@ -2421,6 +2420,31 @@ void tls_sw_ctx_tx_init(struct sock *sk, struct tls_sw_context_tx *sw_ctx)
 	sw_ctx->tx_work.sk = sk;
 }
 
+int tls_sw_drain_tx(struct sock *sk, struct tls_context *ctx, int flags)
+{
+	struct tls_sw_context_tx *sw_ctx = tls_sw_ctx_tx(ctx);
+	int rc;
+
+	flags = (flags & MSG_DONTWAIT) | MSG_NOSIGNAL;
+
+	if (sw_ctx->open_rec)
+		tls_sw_push_pending_record(sk, flags);
+	rc = tls_encrypt_async_wait(sw_ctx);
+	if (rc)
+		return rc;
+	rc = tls_tx_records(sk, flags);
+	if (rc < 0 || tls_is_partially_sent_record(ctx) ||
+	    tls_is_pending_open_record(ctx) ||
+	    !list_empty(&sw_ctx->tx_list))
+		return rc < 0 ? rc : -EAGAIN;
+
+	tls_free_open_rec(sk);
+
+	cancel_delayed_work_sync(&sw_ctx->tx_work.work);
+	clear_bit(BIT_TX_SCHEDULED, &sw_ctx->tx_bitmask);
+	return 0;
+}
+
 static bool tls_is_tx_ready(struct tls_sw_context_tx *ctx)
 {
 	struct tls_rec *rec;
@@ -2609,7 +2633,8 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
 			goto free_aead;
 	}
 
-	ctx->push_pending_record = tls_sw_push_pending_record;
+	if (tx)
+		ctx->push_pending_record = tls_sw_push_pending_record;
 
 	/* setkey is the last operation that could fail during a
 	 * rekey. if it succeeds, we can start modifying the
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 12/15] tls: device: add RX KeyUpdate support
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (10 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 11/15] tls: device: add TX KeyUpdate support Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 13/15] tls: device: add tracepoints for the KeyUpdate path Rishikesh Jethwani
                   ` (2 subsequent siblings)
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

On RX, the NIC may have already decrypted in-flight records with
the old key before the peer's KeyUpdate is parsed, so the old
AEAD, IV and rec_seq are retained on tls_offload_context_rx.
tls_device_rx_del_key() is called from tls_check_pending_rekey()
when a KeyUpdate record is decoded; otherwise post-KeyUpdate records
(carrying new-key wire encryption) would be decrypted with the retired key.
tls_device_decrypted() classifies records by old_nic_boundary:

  - after the boundary: new-key record; drop the old key.
  - before, fully encrypted (including a non-mixed record flagged
    decrypt_failed, which was not transformed): advance old_rec_seq,
    let SW AEAD decrypt.
  - before, mixed - partially decrypted, or with XORed frags flagged
    decrypt_failed: reencrypt with the old key so SW AEAD can decrypt.

rec_start_seq is the TCP sequence of the record's first byte, used both
for the trace_tls_device_decrypted() tracepoint and the old_nic_boundary
classification above. Because copied_seq is advanced at different points
in the two strparser modes, the record start is computed differently: in
copy_mode the record has already been dequeued (tcp_read_done() in
tls_strp_msg_cow() advanced copied_seq past it), so full_len is
subtracted; in non-copy mode copied_seq still points at the record start
and is used directly. This also corrects the tracepoint's first argument,
which previously subtracted full_len unconditionally and was off by one
record on the non-copy path.

On CQE_TLS_OFFLOAD_ERROR the NIC could not authenticate a record; the
driver sets a new skb->decrypt_failed bit while skb->decrypted stays
clear. Whether the payload was transformed depends on the record: in a
mixed record the XORed frags carry skb->decrypt_failed and must be undone,
so tls_device_decrypted() routes mixed records through the
reencrypt-with-old-key path, which undoes the transform per frag and lets
the SW AEAD re-authenticate and decrypt. A non-mixed record with
skb->decrypt_failed set was not transformed; it is still wire ciphertext,
classified as encrypted and decrypted directly after advancing old_rec_seq.

The new key's tls_dev_add is deferred until the old key is fully
consumed: tls_set_device_offload_rx() sets dev_add_pending while
old_aead_recv is retained, and tls_device_deferred_dev_add_rx()
installs the new key when the first post-boundary record is seen. It
anchors the NIC on that record's start (rec_start_seq, already
copy_mode-adjusted as above) paired with the new key's starting record
number, read from crypto_recv.info after tls_sw_ctx_finalize() stored it
there. Handing the NIC that (TCP seq, rec_seq) pair, rather than the
raw copied_seq, which in copy_mode has advanced past the record end,
keeps the sequence mapping the NIC tracks consistent in both strparser
modes.

Tested on Mellanox ConnectX-6 Dx (Crypto Enabled) with multiple
TLS 1.3 RX KeyUpdate cycles.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 include/net/tls.h         |  28 +-
 include/uapi/linux/snmp.h |   3 +
 net/tls/tls.h             |   9 +-
 net/tls/tls_device.c      | 572 ++++++++++++++++++++++++++++++++++----
 net/tls/tls_main.c        |  11 +-
 net/tls/tls_proc.c        |   3 +
 net/tls/tls_sw.c          |   1 +
 7 files changed, 562 insertions(+), 65 deletions(-)

diff --git a/include/net/tls.h b/include/net/tls.h
index b5fc281ff365..6844a685d6e0 100644
--- a/include/net/tls.h
+++ b/include/net/tls.h
@@ -211,10 +211,14 @@ enum tls_context_flags {
 	 * to be atomic.
 	 */
 	TLS_TX_SYNC_SCHED = 1,
-	/* tls_dev_del was called for the RX side, device state was released,
-	 * but tls_ctx->netdev might still be kept, because TX-side driver
-	 * resources might not be released yet. Used to prevent the second
-	 * tls_dev_del call in tls_device_down if it happens simultaneously.
+	/* tls_dev_del was called for the RX side, releasing the NIC's RX
+	 * offload context, while tls_ctx->netdev is still kept (TX-side driver
+	 * resources may not be released yet, or a rekey is about to re-add the
+	 * context). Set in that case, and during a rekey before re-add, and
+	 * cleared when tls_dev_add re-establishes the context. Readers use it to
+	 * avoid a second tls_dev_del and to suppress resync while the NIC has no
+	 * key. tls_device_down() sets it too, so the rekey paths can test the bit
+	 * alone.
 	 */
 	TLS_RX_DEV_CLOSED = 2,
 	/* TX HW context has been tls_dev_del()'d (mid-rekey before the re-add,
@@ -239,6 +243,14 @@ enum tls_context_flags {
 	 * avoidance only.
 	 */
 	TLS_TX_REKEY_FLOOR = 7,
+	/* The RX side fell back to SW decryption during a rekey (tls_dev_add()
+	 * failed, or the netdev is gone) and the socket has been moved from the
+	 * TlsCurrRxDevice to the TlsCurrRxSw gauge while rx_conf stays TLS_HW.
+	 * Accounting only: the functional state is TLS_RX_DEV_{DEGRADED,CLOSED}.
+	 * Cleared, moving the socket back, when a later rekey re-adds the NIC
+	 * context. Mirrors TLS_TX_REKEY_FAILED for the close-time decrement.
+	 */
+	TLS_RX_REKEY_FAILED = 8,
 };
 
 struct tls_prot_info {
@@ -359,6 +371,14 @@ struct tls_offload_context_rx {
 	u8 resync_nh_reset:1;
 	/* CORE_NEXT_HINT-only member, but use the hole here */
 	u8 resync_nh_do_now:1;
+	/* tls_dev_add deferred until old key is freed */
+	u8 dev_add_pending:1;
+	struct {
+		struct crypto_aead *old_aead_recv; /* old key AEAD cipher */
+		char old_iv[TLS_MAX_IV_SIZE + TLS_MAX_SALT_SIZE]; /* old key IV */
+		char old_rec_seq[TLS_MAX_REC_SEQ_SIZE]; /* old key TLS record seq */
+		u32 old_nic_boundary; /* TCP seq below which the NIC may have used the old key */
+	} rekey;
 	union {
 		/* TLS_OFFLOAD_SYNC_TYPE_DRIVER_REQ */
 		struct {
diff --git a/include/uapi/linux/snmp.h b/include/uapi/linux/snmp.h
index a2e0264641de..423aec9ae4ca 100644
--- a/include/uapi/linux/snmp.h
+++ b/include/uapi/linux/snmp.h
@@ -370,8 +370,11 @@ enum
 	LINUX_MIB_TLSTXREKEYERROR,		/* TlsTxRekeyError */
 	LINUX_MIB_TLSRXREKEYRECEIVED,		/* TlsRxRekeyReceived */
 	LINUX_MIB_TLSTXREKEYFALLBACK,		/* TlsTxRekeyFallback */
+	LINUX_MIB_TLSRXREKEYFALLBACK,		/* TlsRxRekeyFallback */
 	LINUX_MIB_TLSCURRTXREKEY,		/* TlsCurrTxRekey */
+	LINUX_MIB_TLSCURRRXREKEY,		/* TlsCurrRxRekey */
 	LINUX_MIB_TLSTXREKEYABORTED,		/* TlsTxRekeyAborted */
+	LINUX_MIB_TLSRXREKEYABORTED,		/* TlsRxRekeyAborted */
 	__LINUX_MIB_TLSMAX
 };
 
diff --git a/net/tls/tls.h b/net/tls/tls.h
index e749f429301a..5d8f4d458df8 100644
--- a/net/tls/tls.h
+++ b/net/tls/tls.h
@@ -251,8 +251,10 @@ void tls_device_cleanup(void);
 int tls_set_device_offload(struct sock *sk,
 			   struct tls_crypto_info *crypto_info);
 void tls_device_free_resources_tx(struct sock *sk);
-int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx);
+int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
+			      struct tls_crypto_info *crypto_info);
 void tls_device_offload_cleanup_rx(struct sock *sk);
+void tls_device_rx_del_key(struct sock *sk, struct tls_context *ctx);
 void tls_device_rx_resync_new_rec(struct sock *sk, u32 rcd_len, u32 seq);
 int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx);
 #else
@@ -268,13 +270,16 @@ tls_set_device_offload(struct sock *sk, struct tls_crypto_info *crypto_info)
 static inline void tls_device_free_resources_tx(struct sock *sk) {}
 
 static inline int
-tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
+tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
+			  struct tls_crypto_info *crypto_info)
 {
 	return -EOPNOTSUPP;
 }
 
 static inline void tls_device_offload_cleanup_rx(struct sock *sk) {}
 static inline void
+tls_device_rx_del_key(struct sock *sk, struct tls_context *ctx) {}
+static inline void
 tls_device_rx_resync_new_rec(struct sock *sk, u32 rcd_len, u32 seq) {}
 
 static inline int
diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index f32c1bb6b497..ac09f356cff9 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -67,8 +67,18 @@ static void tls_device_free_ctx(struct tls_context *ctx)
 		kfree(offload_ctx);
 	}
 
-	if (ctx->rx_conf == TLS_HW)
-		kfree(tls_offload_ctx_rx(ctx));
+	if (ctx->rx_conf == TLS_HW) {
+		struct tls_offload_context_rx *offload_ctx =
+			tls_offload_ctx_rx(ctx);
+
+		/* Normally freed and NULLed in tls_device_offload_cleanup_rx();
+		 * free defensively here so a future path can't leak the tfm.
+		 */
+		crypto_free_aead(offload_ctx->rekey.old_aead_recv);
+		memzero_explicit(&offload_ctx->rekey,
+				 sizeof(offload_ctx->rekey));
+		kfree(offload_ctx);
+	}
 
 	tls_ctx_free(NULL, ctx);
 }
@@ -192,6 +202,129 @@ static void tls_device_commit_start_marker(struct sock *sk,
 	tcp_write_collapse_fence(sk);
 }
 
+/* Account a rekey that could not (re)install the RX key on the NIC. The event
+ * counter is bumped every time; the gauges move only on the first fallback
+ * since the socket was last offloaded, so the recurring post-NETDEV_DOWN
+ * rekeys and repeated failed adds do not drift them. The matching move back is
+ * in tls_device_dev_add_rx(); the close-time decrement keys off the bit.
+ */
+static void tls_device_rx_rekey_fallback(struct sock *sk,
+					 struct tls_context *tls_ctx)
+{
+	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYFALLBACK);
+	if (!test_and_set_bit(TLS_RX_REKEY_FAILED, &tls_ctx->flags)) {
+		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXDEVICE);
+		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXSW);
+	}
+}
+
+static int tls_device_dev_add_rx(struct sock *sk, struct tls_context *tls_ctx,
+				 struct net_device *netdev,
+				 struct tls_crypto_info *crypto_info,
+				 u32 cur_seq, bool is_rekey)
+{
+	const struct tls_cipher_desc *cipher_desc;
+	char *rec_seq;
+	int rc;
+
+	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
+	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
+
+	rc = netdev->tlsdev_ops->tls_dev_add(netdev, sk,
+					     TLS_OFFLOAD_CTX_DIR_RX,
+					     crypto_info, cur_seq);
+	rec_seq = crypto_info_rec_seq(crypto_info, cipher_desc);
+	trace_tls_device_offload_set(sk, TLS_OFFLOAD_CTX_DIR_RX,
+				     cur_seq, rec_seq, rc);
+	if (!rc) {
+		clear_bit(TLS_RX_DEV_DEGRADED, &tls_ctx->flags);
+		clear_bit(TLS_RX_DEV_CLOSED, &tls_ctx->flags);
+		/* Back on the NIC after an earlier SW fallback: undo its move. */
+		if (test_and_clear_bit(TLS_RX_REKEY_FAILED, &tls_ctx->flags)) {
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXSW);
+			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXDEVICE);
+		}
+		if (is_rekey)
+			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYOK);
+	} else if (is_rekey) {
+		set_bit(TLS_RX_DEV_DEGRADED, &tls_ctx->flags);
+		set_bit(TLS_RX_DEV_CLOSED, &tls_ctx->flags);
+		tls_device_rx_rekey_fallback(sk, tls_ctx);
+	}
+	return rc;
+}
+
+static void tls_device_deferred_dev_add_rx(struct sock *sk,
+					   struct tls_context *tls_ctx,
+					   struct tls_offload_context_rx *ctx,
+					   u32 rec_start_seq)
+{
+	const struct tls_cipher_desc *cipher_desc;
+	union tls_crypto_context crypto_ctx;
+	struct net_device *netdev;
+
+	ctx->dev_add_pending = 0;
+
+	/* crypto_recv.info.rec_seq is frozen at the value setsockopt() passed
+	 * in: the new key's first record number. The records that drained
+	 * between setsockopt() and this boundary crossing were SW-decrypted
+	 * under the new key and advanced tls_ctx->rx.rec_seq, so the record
+	 * starting at rec_start_seq, the one being decrypted right now,
+	 * before tls_rx_one_record() calls tls_advance_record_sn(), is
+	 * numbered by rx.rec_seq, not by the blob. Hand the NIC the live
+	 * (TCP seq, record number) pair, as getsockopt(TLS_RX) already does.
+	 */
+	cipher_desc = get_cipher_desc(tls_ctx->crypto_recv.info.cipher_type);
+	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
+	crypto_ctx = tls_ctx->crypto_recv;
+	memcpy(crypto_info_rec_seq(&crypto_ctx.info, cipher_desc),
+	       tls_ctx->rx.rec_seq, cipher_desc->rec_seq);
+
+	down_read(&device_offload_lock);
+	netdev = rcu_dereference_protected(tls_ctx->netdev,
+					   lockdep_is_held(&device_offload_lock));
+	if (netdev)
+		tls_device_dev_add_rx(sk, tls_ctx, netdev,
+				      &crypto_ctx.info,
+				      rec_start_seq, true);
+	else
+		tls_device_rx_rekey_fallback(sk, tls_ctx);
+	up_read(&device_offload_lock);
+	memzero_explicit(&crypto_ctx, sizeof(crypto_ctx));
+	TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXREKEY);
+}
+
+/* Retire the NIC's RX key when a KeyUpdate record is decoded (from
+ * tls_check_pending_rekey(), lock_sock held). The NIC must lose the old key
+ * now, before it transforms further post-KeyUpdate records that are new-key on
+ * the wire. TLS_RX_DEV_CLOSED is re-tested under device_offload_lock because
+ * tls_device_down() can run in between; synchronize_net() drains the RX path
+ * before the driver frees its context.
+ */
+void tls_device_rx_del_key(struct sock *sk, struct tls_context *ctx)
+{
+	struct net_device *netdev;
+
+	if (ctx->rx_conf != TLS_HW)
+		return;
+	if (test_bit(TLS_RX_DEV_CLOSED, &ctx->flags))
+		return;
+
+	down_read(&device_offload_lock);
+	netdev = rcu_dereference_protected(ctx->netdev,
+					   lockdep_is_held(&device_offload_lock));
+	if (!netdev || test_bit(TLS_RX_DEV_CLOSED, &ctx->flags)) {
+		up_read(&device_offload_lock);
+		return;
+	}
+
+	set_bit(TLS_RX_DEV_CLOSED, &ctx->flags);
+	synchronize_net();
+	netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
+					TLS_OFFLOAD_CTX_DIR_RX);
+	up_read(&device_offload_lock);
+}
+
 static void destroy_record(struct tls_record_info *record)
 {
 	int i;
@@ -969,6 +1102,8 @@ void tls_device_rx_resync_new_rec(struct sock *sk, u32 rcd_len, u32 seq)
 		return;
 	if (unlikely(test_bit(TLS_RX_DEV_DEGRADED, &tls_ctx->flags)))
 		return;
+	if (unlikely(test_bit(TLS_RX_DEV_CLOSED, &tls_ctx->flags)))
+		return;
 
 	prot = &tls_ctx->prot_info;
 	rx_ctx = tls_offload_ctx_rx(tls_ctx);
@@ -1114,7 +1249,7 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
 	if (skb_pagelen(skb) > offset) {
 		copy = min_t(int, skb_pagelen(skb) - offset, data_len);
 
-		if (skb->decrypted) {
+		if (skb->decrypted || skb->decrypt_failed) {
 			err = skb_store_bits(skb, offset, buf, copy);
 			if (err)
 				goto free_buf;
@@ -1141,7 +1276,7 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
 		copy = min_t(int, skb_iter->len - frag_pos,
 			     data_len + rxm->offset - offset);
 
-		if (skb_iter->decrypted) {
+		if (skb_iter->decrypted || skb_iter->decrypt_failed) {
 			err = skb_store_bits(skb_iter, frag_pos, buf, copy);
 			if (err)
 				goto free_buf;
@@ -1158,6 +1293,77 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
 	return err;
 }
 
+/*
+ * Reconstruct a boundary record whose frags the NIC XORed with the old key,
+ * then hand it to the SW AEAD under the current (new) key.
+ *
+ * These are deliberately two different keys: the sender has already done its
+ * TX KeyUpdate, so the record on the wire is AEAD-encrypted with the new key,
+ * but the RX NIC still holds the old key and CTR-XORed some frags with the old
+ * keystream. tls_device_reencrypt() must undo that XOR with the *old* key to
+ * restore the pristine new-key ciphertext, so swap the old key in only for the
+ * reconstruction and restore the current key before returning; the SW AEAD
+ * decrypt that follows then runs under the new key, matching the wire record.
+ */
+static int tls_device_reencrypt_old_key(struct sock *sk,
+					struct tls_offload_context_rx *ctx,
+					struct tls_sw_context_rx *sw_ctx,
+					struct tls_context *tls_ctx)
+{
+	struct crypto_aead *saved_aead = sw_ctx->aead_recv;
+	char saved_iv[TLS_MAX_IV_SIZE + TLS_MAX_SALT_SIZE];
+	char saved_rec_seq[TLS_MAX_REC_SEQ_SIZE];
+	int ret;
+
+	memcpy(saved_iv, tls_ctx->rx.iv, sizeof(saved_iv));
+	memcpy(saved_rec_seq, tls_ctx->rx.rec_seq, sizeof(saved_rec_seq));
+
+	sw_ctx->aead_recv = ctx->rekey.old_aead_recv;
+	memcpy(tls_ctx->rx.iv, ctx->rekey.old_iv, sizeof(ctx->rekey.old_iv));
+	memcpy(tls_ctx->rx.rec_seq, ctx->rekey.old_rec_seq,
+	       sizeof(ctx->rekey.old_rec_seq));
+
+	ret = tls_device_reencrypt(sk, tls_ctx);
+
+	memcpy(ctx->rekey.old_rec_seq, tls_ctx->rx.rec_seq,
+	       sizeof(ctx->rekey.old_rec_seq));
+
+	sw_ctx->aead_recv = saved_aead;
+	memcpy(tls_ctx->rx.iv, saved_iv, sizeof(saved_iv));
+	memcpy(tls_ctx->rx.rec_seq, saved_rec_seq, sizeof(saved_rec_seq));
+
+	if (ret)
+		return ret;
+
+	tls_bigint_increment(ctx->rekey.old_rec_seq,
+			     tls_ctx->prot_info.rec_seq_size);
+	ctx->resync_nh_reset = 1;
+
+	return 0;
+}
+
+/*
+ * TCP sequence of the first byte of the record the strparser currently holds
+ * or is still collecting. In non-copy mode tcp_sk(sk)->copied_seq is left at
+ * the record start until tls_strp_msg_consume(). In copy mode
+ * tls_strp_read_copy() zeroes stm.offset and anchor->len and then
+ * tls_strp_read_copyin() -> tcp_read_sock() advances copied_seq by every byte
+ * it appends to the anchor, a complete parsed-ahead record, a partial one
+ * under rmem pressure, or only header bytes, so subtract anchor->len to get
+ * back to the record start. Both the recv path and the setsockopt rekey path
+ * must classify records against the same start, so share this helper.
+ */
+static u32 tls_device_rx_rec_start(struct sock *sk,
+				   struct tls_sw_context_rx *sw_ctx)
+{
+	u32 copied_seq = tcp_sk(sk)->copied_seq;
+
+	if (sw_ctx->strp.copy_mode)
+		return copied_seq - sw_ctx->strp.anchor->len;
+
+	return copied_seq;
+}
+
 int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
 {
 	struct tls_offload_context_rx *ctx = tls_offload_ctx_rx(tls_ctx);
@@ -1165,6 +1371,7 @@ int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
 	struct sk_buff *skb = tls_strp_msg(sw_ctx);
 	struct strp_msg *rxm = strp_msg(skb);
 	int is_decrypted, is_encrypted;
+	u32 rec_start_seq;
 
 	if (!tls_strp_msg_mixed_decrypted(sw_ctx)) {
 		is_decrypted = skb->decrypted;
@@ -1174,10 +1381,72 @@ int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
 		is_encrypted = 0;
 	}
 
-	trace_tls_device_decrypted(sk, tcp_sk(sk)->copied_seq - rxm->full_len,
+	rec_start_seq = tls_device_rx_rec_start(sk, sw_ctx);
+
+	trace_tls_device_decrypted(sk, rec_start_seq,
 				   tls_ctx->rx.rec_seq, rxm->full_len,
 				   is_encrypted, is_decrypted);
 
+	if (unlikely(ctx->rekey.old_aead_recv)) {
+		bool nic_touched = !is_encrypted || skb->decrypt_failed;
+		bool before_nic_boundary;
+
+		/* old_nic_boundary is the TCP stack's view at setsockopt time
+		 * (rcv_nxt plus the out-of-order tail), not the NIC's last
+		 * transformed byte. A segment the NIC transformed with the old
+		 * key before tls_dev_del returned can still be in the RQ/CQ, in
+		 * a GRO list or in the socket backlog when that snapshot is
+		 * taken and reach TCP later, above it. While old_aead_recv is
+		 * held the NIC has no RX context for this socket at all: the
+		 * old one was deleted before old_aead_recv was set and the new
+		 * one is only installed once it is freed below. So a NIC mark
+		 * seen here can only be the old key's transform, wherever the
+		 * record sits relative to the snapshot. Slide the boundary out
+		 * over such a record instead of retiring the old key on it; the
+		 * old key is retired only on a record the NIC never saw.
+		 */
+		if (nic_touched &&
+		    !before(rec_start_seq, ctx->rekey.old_nic_boundary))
+			ctx->rekey.old_nic_boundary = rec_start_seq + rxm->full_len;
+
+		before_nic_boundary =
+			before(rec_start_seq, ctx->rekey.old_nic_boundary);
+
+		if (before_nic_boundary) {
+			/* Non-mixed (skb->decrypted clear) is untouched wire
+			 * ciphertext even if skb->decrypt_failed is set, so advance
+			 * old_rec_seq and let the SW AEAD decrypt it directly.
+			 * old_rec_seq tracks the stream's record number, which the
+			 * NIC also advances for records it did not transform, so
+			 * keeping it in step lets a later NIC-touched record be undone
+			 * with the right nonce. A mixed record carries NIC-XORed frags
+			 * (skb->decrypt_failed or skb->decrypted) and takes the
+			 * old-key reencrypt path below, which undoes the transform per
+			 * frag before the SW AEAD decrypts.
+			 */
+			if (is_encrypted) {
+				tls_bigint_increment(ctx->rekey.old_rec_seq,
+						     tls_ctx->prot_info.rec_seq_size);
+				return 0;
+			}
+
+			return tls_device_reencrypt_old_key(sk, ctx,
+							    sw_ctx, tls_ctx);
+		}
+
+		crypto_free_aead(ctx->rekey.old_aead_recv);
+		ctx->rekey.old_aead_recv = NULL;
+
+		/* Anchor the NIC on the start of this first post-boundary
+		 * record. rec_start_seq already accounts for copy_mode, where
+		 * copied_seq has advanced past the record end; using it keeps
+		 * the (TCP seq, record number) pair consistent in both modes.
+		 */
+		if (ctx->dev_add_pending)
+			tls_device_deferred_dev_add_rx(sk, tls_ctx, ctx,
+						       rec_start_seq);
+	}
+
 	if (unlikely(test_bit(TLS_RX_DEV_DEGRADED, &tls_ctx->flags))) {
 		if (likely(is_encrypted || is_decrypted))
 			return is_decrypted;
@@ -1804,73 +2073,224 @@ int tls_set_device_offload(struct sock *sk,
 	return rc;
 }
 
-int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
+int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
+			      struct tls_crypto_info *new_crypto_info)
 {
-	struct tls12_crypto_info_aes_gcm_128 *info;
+	struct tls_crypto_info *crypto_info, *src_crypto_info;
+	const struct tls_cipher_desc *cipher_desc;
+	u32 drain_start = tcp_sk(sk)->copied_seq;
 	struct tls_offload_context_rx *context;
 	struct net_device *netdev;
+	bool was_dev_add_pending;
+	bool moved_aead_recv = false;
+	bool retired_pending = false;
+	bool put_netdev = false;
 	int rc = 0;
 
-	/* A rekey (setsockopt on an already-configured socket) is not
-	 * supported on the device offload path yet; reject it here so the
-	 * caller can decide (propagate the error for a HW connection, or
-	 * re-init software crypto for a SW one). KeyUpdate support replaces
-	 * this guard with real rekey handling.
-	 */
-	if (ctx->rx_conf != TLS_BASE)
-		return -EOPNOTSUPP;
+	/* A rekey of a SW-offloaded socket belongs to tls_set_sw_offload(). */
+	if (new_crypto_info && ctx->rx_conf != TLS_HW)
+		return -EINVAL;
 
-	netdev = get_netdev_for_sock(sk);
-	if (!netdev) {
-		pr_err_ratelimited("%s: netdev not found\n", __func__);
+	crypto_info = &ctx->crypto_recv.info;
+	src_crypto_info = new_crypto_info ?: crypto_info;
+	cipher_desc = get_cipher_desc(src_crypto_info->cipher_type);
+	if (!cipher_desc || !cipher_desc->offloadable)
 		return -EINVAL;
-	}
 
-	if (!(netdev->features & NETIF_F_HW_TLS_RX)) {
-		rc = -EOPNOTSUPP;
-		goto release_netdev;
-	}
+	if (new_crypto_info) {
+		/* Rekey targets the device holding the HW RX context, which
+		 * can differ from the socket's route after a route change or
+		 * bond/team failover. Resolve it from ctx->netdev under
+		 * device_offload_lock, like the other del/add-key paths, not
+		 * via get_netdev_for_sock(). The context owns the reference,
+		 * so don't take an extra one here.
+		 *
+		 * A NULL netdev means tls_device_down() already ran: the HW RX
+		 * context is deleted, TLS_RX_DEV_{DEGRADED,CLOSED} are set and
+		 * every record is decrypted in SW, but rx_conf stays TLS_HW.
+		 * The rekey is still required, the peer's KeyUpdate was parsed
+		 * and recvmsg() returns -EKEYEXPIRED until the new key lands,
+		 * so run the same state machine (queued records may still carry
+		 * the deleted NIC context's old-key XOR) and account the new key
+		 * as a SW fallback in place of the tls_dev_del()/tls_dev_add()
+		 * steps, mirroring the TX side (tls_device_complete_rekey()).
+		 * Do not fail the setsockopt.
+		 */
+		down_read(&device_offload_lock);
+		netdev = rcu_dereference_protected(ctx->netdev,
+						   lockdep_is_held(&device_offload_lock));
+	} else {
+		netdev = get_netdev_for_sock(sk);
+		if (!netdev) {
+			pr_err_ratelimited("%s: netdev not found\n", __func__);
+			return -EINVAL;
+		}
+		put_netdev = true;
 
-	/* Avoid offloading if the device is down
-	 * We don't want to offload new flows after
-	 * the NETDEV_DOWN event
-	 *
-	 * device_offload_lock is taken in tls_devices's NETDEV_DOWN
-	 * handler thus protecting from the device going down before
-	 * ctx was added to tls_device_list.
-	 */
-	down_read(&device_offload_lock);
-	if (!(netdev->flags & IFF_UP)) {
-		rc = -EINVAL;
-		goto release_lock;
+		if (!(netdev->features & NETIF_F_HW_TLS_RX)) {
+			rc = -EOPNOTSUPP;
+			goto release_netdev;
+		}
+
+		/* Avoid offloading if the device is down
+		 * We don't want to offload new flows after
+		 * the NETDEV_DOWN event
+		 *
+		 * device_offload_lock is taken in tls_devices's NETDEV_DOWN
+		 * handler thus protecting from the device going down before
+		 * ctx was added to tls_device_list.
+		 */
+		down_read(&device_offload_lock);
+		if (!(netdev->flags & IFF_UP)) {
+			rc = -EINVAL;
+			goto release_lock;
+		}
 	}
 
-	context = kzalloc_obj(*context);
-	if (!context) {
-		rc = -ENOMEM;
-		goto release_lock;
+	if (!new_crypto_info) {
+		context = kzalloc_obj(*context);
+		if (!context) {
+			rc = -ENOMEM;
+			goto release_lock;
+		}
+		ctx->priv_ctx_rx = context;
+	} else {
+		context = tls_offload_ctx_rx(ctx);
 	}
+	was_dev_add_pending = context->dev_add_pending;
 	context->resync_nh_reset = 1;
 
-	ctx->priv_ctx_rx = context;
-	rc = tls_sw_ctx_init(sk, 0, NULL);
+	if (new_crypto_info) {
+		struct tls_sw_context_rx *sw_ctx = tls_sw_ctx_rx(ctx);
+
+		/* Classify against the record start, not the raw copied_seq: in
+		 * strparser copy mode tcp_read_sock() has already advanced
+		 * copied_seq past a parsed-ahead (possibly partial) record the
+		 * user has not received, which may still carry the old NIC key's
+		 * XOR. tls_device_decrypted() compensates the same way; keeping
+		 * both in sync is what lets a drained-vs-still-draining decision
+		 * here match the reencrypt-key decision there.
+		 */
+		drain_start = tls_device_rx_rec_start(sk, sw_ctx);
+
+		/* netdev is NULL only after tls_device_down(), which already
+		 * deleted the HW RX context and set TLS_RX_DEV_CLOSED; the
+		 * netdev check just makes that dependency explicit.
+		 */
+		if (netdev && !test_bit(TLS_RX_DEV_CLOSED, &ctx->flags)) {
+			set_bit(TLS_RX_DEV_CLOSED, &ctx->flags);
+			synchronize_net();
+			netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
+							TLS_OFFLOAD_CTX_DIR_RX);
+		}
+
+		if (context->rekey.old_aead_recv &&
+		    before(drain_start, context->rekey.old_nic_boundary)) {
+			/* Previous rekey still draining. Keep rekey.old_aead_recv,
+			 * it is the only key that can undo the NIC-XOR on queued
+			 * records. sw_ctx->aead_recv may be re-setkey'd by
+			 * tls_sw_ctx_init(); that intermediate key was never on
+			 * the NIC and its wire era is drained, so it is needed
+			 * for neither undo nor AEAD. Defer dev_add; the new key
+			 * is installed once drain_start crosses rekey.old_nic_boundary.
+			 */
+			context->dev_add_pending = 1;
+		} else {
+			struct tcp_sock *tp = tcp_sk(sk);
+			u32 nic_end;
+
+			if (context->rekey.old_aead_recv) {
+				crypto_free_aead(context->rekey.old_aead_recv);
+				context->rekey.old_aead_recv = NULL;
+			}
+
+			/* Flush the backlog so TCP's view is current, then take the
+			 * highest byte TCP holds, including the out-of-order tail:
+			 * a NIC-transformed segment behind a host-side drop sits
+			 * above rcv_nxt until the retransmit fills the hole and
+			 * must still be classified against the old key. This is
+			 * still only the stack's view, a transformed segment the
+			 * NIC has not delivered yet is caught in-band by
+			 * tls_device_decrypted(), which slides the boundary.
+			 */
+			__sk_flush_backlog(sk);
+			nic_end = tp->rcv_nxt;
+			if (!RB_EMPTY_ROOT(&tp->out_of_order_queue) &&
+			    after(TCP_SKB_CB(tp->ooo_last_skb)->end_seq, nic_end))
+				nic_end = TCP_SKB_CB(tp->ooo_last_skb)->end_seq;
+
+			if (before(drain_start, nic_end)) {
+				context->rekey.old_aead_recv = sw_ctx->aead_recv;
+				/* NULL so tls_sw_ctx_init() allocates a fresh tfm
+				 * for the new key instead of re-keying the one we
+				 * must keep for the drain.
+				 */
+				sw_ctx->aead_recv = NULL;
+				moved_aead_recv = true;
+				memcpy(context->rekey.old_iv, ctx->rx.iv,
+				       sizeof(context->rekey.old_iv));
+				memcpy(context->rekey.old_rec_seq, ctx->rx.rec_seq,
+				       sizeof(context->rekey.old_rec_seq));
+				context->rekey.old_nic_boundary = nic_end;
+				context->dev_add_pending = 1;
+			} else if (was_dev_add_pending) {
+				/* A prior rekey's deferred dev_add can no longer
+				 * run: its trigger (old_aead_recv) was just freed
+				 * above and no new drain replaces it. Its era
+				 * drained successfully (drain_start is already past
+				 * old_nic_boundary), so retire it and let the new
+				 * key install immediately below. retired_pending
+				 * defers its OK/gauge accounting to the post-init
+				 * block, past the error goto, so a failed
+				 * tls_sw_ctx_init() needs no counter undo.
+				 */
+				context->dev_add_pending = 0;
+				retired_pending = true;
+			}
+		}
+	}
+
+	rc = tls_sw_ctx_init(sk, 0, new_crypto_info);
 	if (rc)
 		goto release_ctx;
 
-	rc = netdev->tlsdev_ops->tls_dev_add(netdev, sk, TLS_OFFLOAD_CTX_DIR_RX,
-					     &ctx->crypto_recv.info,
-					     tcp_sk(sk)->copied_seq);
-	info = (void *)&ctx->crypto_recv.info;
-	trace_tls_device_offload_set(sk, TLS_OFFLOAD_CTX_DIR_RX,
-				     tcp_sk(sk)->copied_seq, info->rec_seq, rc);
-	if (rc)
-		goto free_sw_resources;
+	if (!context->dev_add_pending) {
+		if (retired_pending) {
+			/* Account the superseded rekey that drained OK, mirroring
+			 * the deferred-add path: one RXREKEYOK and release its
+			 * in-flight gauge. The new key's own OK/FALLBACK is counted
+			 * by tls_device_dev_add_rx() just below.
+			 */
+			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYOK);
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXREKEY);
+		}
+		if (netdev) {
+			rc = tls_device_dev_add_rx(sk, ctx, netdev,
+						   src_crypto_info, drain_start,
+						   !!new_crypto_info);
+		} else {
+			/* No device after tls_device_down(); the SW path keeps
+			 * decrypting.
+			 */
+			tls_device_rx_rekey_fallback(sk, ctx);
+		}
+		if (!new_crypto_info) {
+			if (rc)
+				goto free_sw_resources;
+			tls_device_attach(ctx, sk, netdev);
+		}
+	} else if (!was_dev_add_pending) {
+		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXREKEY);
+	} else {
+		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYOK);
+	}
+
+	tls_sw_ctx_finalize(sk, 0, new_crypto_info);
 
-	tls_device_attach(ctx, sk, netdev);
-	tls_sw_ctx_finalize(sk, 0, NULL);
 	up_read(&device_offload_lock);
 
-	dev_put(netdev);
+	if (put_netdev)
+		dev_put(netdev);
 
 	return 0;
 
@@ -1879,17 +2299,39 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
 	tls_sw_free_resources_rx(sk);
 	down_read(&device_offload_lock);
 release_ctx:
-	ctx->priv_ctx_rx = NULL;
+	if (!new_crypto_info) {
+		ctx->priv_ctx_rx = NULL;
+	} else {
+		/* A failed RX rekey is terminal, so there is no HW state to roll
+		 * back to. KeyUpdate is directional and the peer's TX has already
+		 * switched keys, so once the new RX key fails to install the old
+		 * SW key restored below cannot decrypt any further record; the
+		 * socket is dead and the app must close it. The half-torn HW
+		 * context (tls_dev_del already ran) and any dangling
+		 * dev_add_pending / old_aead_recv are reclaimed by
+		 * tls_device_offload_cleanup_rx() on close.
+		 */
+		context->dev_add_pending = was_dev_add_pending;
+		if (moved_aead_recv) {
+			struct tls_sw_context_rx *sw_ctx = tls_sw_ctx_rx(ctx);
+
+			crypto_free_aead(sw_ctx->aead_recv);
+			sw_ctx->aead_recv = context->rekey.old_aead_recv;
+			context->rekey.old_aead_recv = NULL;
+		}
+	}
 release_lock:
 	up_read(&device_offload_lock);
 release_netdev:
-	dev_put(netdev);
+	if (put_netdev)
+		dev_put(netdev);
 	return rc;
 }
 
 void tls_device_offload_cleanup_rx(struct sock *sk)
 {
 	struct tls_context *tls_ctx = tls_get_ctx(sk);
+	struct tls_offload_context_rx *rx_ctx;
 	struct net_device *netdev;
 
 	down_read(&device_offload_lock);
@@ -1898,8 +2340,9 @@ void tls_device_offload_cleanup_rx(struct sock *sk)
 	if (!netdev)
 		goto out;
 
-	netdev->tlsdev_ops->tls_dev_del(netdev, tls_ctx,
-					TLS_OFFLOAD_CTX_DIR_RX);
+	if (!test_bit(TLS_RX_DEV_CLOSED, &tls_ctx->flags))
+		netdev->tlsdev_ops->tls_dev_del(netdev, tls_ctx,
+						TLS_OFFLOAD_CTX_DIR_RX);
 
 	if (tls_ctx->tx_conf != TLS_HW) {
 		dev_put(netdev);
@@ -1909,6 +2352,19 @@ void tls_device_offload_cleanup_rx(struct sock *sk)
 	}
 out:
 	up_read(&device_offload_lock);
+
+	rx_ctx = tls_offload_ctx_rx(tls_ctx);
+	if (rx_ctx && rx_ctx->rekey.old_aead_recv) {
+		crypto_free_aead(rx_ctx->rekey.old_aead_recv);
+		rx_ctx->rekey.old_aead_recv = NULL;
+	}
+
+	if (rx_ctx && rx_ctx->dev_add_pending) {
+		rx_ctx->dev_add_pending = 0;
+		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYABORTED);
+		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXREKEY);
+	}
+
 	tls_sw_release_resources_rx(sk);
 }
 
@@ -1969,9 +2425,11 @@ static int tls_device_down(struct net_device *netdev)
 			set_bit(TLS_TX_DEV_CLOSED, &ctx->flags);
 		}
 		if (ctx->rx_conf == TLS_HW &&
-		    !test_bit(TLS_RX_DEV_CLOSED, &ctx->flags))
+		    !test_bit(TLS_RX_DEV_CLOSED, &ctx->flags)) {
 			netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
 							TLS_OFFLOAD_CTX_DIR_RX);
+			set_bit(TLS_RX_DEV_CLOSED, &ctx->flags);
+		}
 
 		dev_put(netdev);
 
diff --git a/net/tls/tls_main.c b/net/tls/tls_main.c
index 3dd3a4ce8209..0a9e7d15fa95 100644
--- a/net/tls/tls_main.c
+++ b/net/tls/tls_main.c
@@ -361,8 +361,14 @@ static void tls_sk_proto_cleanup(struct sock *sk,
 		tls_sw_release_resources_rx(sk);
 		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXSW);
 	} else if (ctx->rx_conf == TLS_HW) {
+		bool rekey_failed = test_bit(TLS_RX_REKEY_FAILED, &ctx->flags);
+
 		tls_device_offload_cleanup_rx(sk);
-		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXDEVICE);
+
+		if (rekey_failed)
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXSW);
+		else
+			TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRRXDEVICE);
 	}
 }
 
@@ -754,7 +760,8 @@ static int do_tls_setsockopt_conf(struct sock *sk, sockptr_t optval,
 			conf = TLS_SW;
 		}
 	} else {
-		rc = tls_set_device_offload_rx(sk, ctx);
+		rc = tls_set_device_offload_rx(sk, ctx,
+					       update ? crypto_info : NULL);
 		conf = TLS_HW;
 		if (!rc) {
 			if (!update) {
diff --git a/net/tls/tls_proc.c b/net/tls/tls_proc.c
index 4bb1e3727e28..6255f7b07eb7 100644
--- a/net/tls/tls_proc.c
+++ b/net/tls/tls_proc.c
@@ -28,8 +28,11 @@ static const struct snmp_mib tls_mib_list[] = {
 	SNMP_MIB_ITEM("TlsTxRekeyError", LINUX_MIB_TLSTXREKEYERROR),
 	SNMP_MIB_ITEM("TlsRxRekeyReceived", LINUX_MIB_TLSRXREKEYRECEIVED),
 	SNMP_MIB_ITEM("TlsTxRekeyFallback", LINUX_MIB_TLSTXREKEYFALLBACK),
+	SNMP_MIB_ITEM("TlsRxRekeyFallback", LINUX_MIB_TLSRXREKEYFALLBACK),
 	SNMP_MIB_ITEM("TlsCurrTxRekey", LINUX_MIB_TLSCURRTXREKEY),
+	SNMP_MIB_ITEM("TlsCurrRxRekey", LINUX_MIB_TLSCURRRXREKEY),
 	SNMP_MIB_ITEM("TlsTxRekeyAborted", LINUX_MIB_TLSTXREKEYABORTED),
+	SNMP_MIB_ITEM("TlsRxRekeyAborted", LINUX_MIB_TLSRXREKEYABORTED),
 };
 
 static int tls_statistics_seq_show(struct seq_file *seq, void *v)
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index fd162d8f1d64..d546091dd524 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -1561,6 +1561,7 @@ static int tls_check_pending_rekey(struct sock *sk, struct tls_context *ctx,
 	if (hs_type == TLS_HANDSHAKE_KEYUPDATE) {
 		struct tls_sw_context_rx *rx_ctx = ctx->priv_ctx_rx;
 
+		tls_device_rx_del_key(sk, ctx);
 		WRITE_ONCE(rx_ctx->key_update_pending, true);
 		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYRECEIVED);
 	}
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 13/15] tls: device: add tracepoints for the KeyUpdate path
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (11 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 12/15] tls: device: add RX " Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 14/15] selftests: net: add TLS hardware offload test Rishikesh Jethwani
  2026-09-17 22:35 ` [PATCH net-next v17 15/15] tls: document TLS 1.3 hardware offload rekey handling Rishikesh Jethwani
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Add five trace events covering the rekey state machine in
tls_device.c:

  tls_device_rekey_start: rekey accepted, inflight=1 marks a
                          drain window where old-key data is still
                          queued and dev_add is deferred, whether
                          this rekey opened the window or landed
                          behind a still-draining prior one.
                          inflight=0 means the key was hot-swapped
                          with no drain; nic_boundary is then the
                          candidate receive frontier, not a stored
                          boundary.
  tls_device_rekey_reencrypt: old-key undo pass for a boundary
                              record
  tls_device_rekey_done: a drain window's boundary was crossed and
                         old_aead_recv freed, either by a record
                         arriving past old_nic_boundary or by a new
                         rekey superseding a prior one whose era had
                         already drained. deferred dev_add is issued
                         if pending.
  tls_device_complete_rekey_retry: TX rekey completion hit the
                                   transient -EAGAIN retry in sendmsg,
                                   the next sendmsg retries.
  tls_device_complete_rekey_fail: TX rekey completion gave up and
                                  fell back to SW encryption. Emitted
                                  from the fallback path itself, since
                                  the sendmsg call site only sees the
                                  transient retry.

These are independent event markers, not a paired begin/end span. A
rekey_start is not 1:1 with a rekey_done: a rekey that lands behind a
still-draining prior one reports inflight=1 without opening a window of
its own, so several starts can map to a single window and a single done.
The invariant is per-window - each drain window that is opened closes
with exactly one rekey_done. A window can also be abandoned without a
done when a rekey is aborted (tls_sw_ctx_init() failing after the window
opened, or the socket being torn down mid-drain, the latter counted by
TLSRXREKEYABORTED); those paths are not traced.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 net/tls/tls_device.c |  36 ++++++++++++-
 net/tls/trace.h      | 118 +++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 152 insertions(+), 2 deletions(-)

diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index ac09f356cff9..5f45c097bad3 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -866,8 +866,16 @@ int tls_device_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
 	lock_sock(sk);
 
 	/* Old-key records all ACKed; switch back to HW. */
-	if (test_bit(TLS_TX_REKEY_READY, &tls_ctx->flags))
-		tls_device_complete_rekey(sk, tls_ctx, true, msg->msg_flags);
+	if (test_bit(TLS_TX_REKEY_READY, &tls_ctx->flags)) {
+		rc = tls_device_complete_rekey(sk, tls_ctx, true, msg->msg_flags);
+		/* Non-zero here is the transient -EAGAIN retry,
+		 * the next sendmsg retries. Hard failures return 0 after
+		 * falling back to SW and emit tls_device_complete_rekey_fail
+		 * from the fallback path.
+		 */
+		if (rc)
+			trace_tls_device_complete_rekey_retry(sk);
+	}
 
 	if (tls_device_tx_uses_sw(tls_ctx)) {
 		rc = tls_sw_sendmsg_locked(sk, msg, size);
@@ -1430,10 +1438,15 @@ int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
 				return 0;
 			}
 
+			trace_tls_device_rekey_reencrypt(sk, rec_start_seq,
+							 ctx->rekey.old_nic_boundary);
+
 			return tls_device_reencrypt_old_key(sk, ctx,
 							    sw_ctx, tls_ctx);
 		}
 
+		trace_tls_device_rekey_done(sk, rec_start_seq,
+					    ctx->rekey.old_nic_boundary);
 		crypto_free_aead(ctx->rekey.old_aead_recv);
 		ctx->rekey.old_aead_recv = NULL;
 
@@ -1890,6 +1903,13 @@ static int tls_device_complete_rekey(struct sock *sk, struct tls_context *ctx,
 	TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
 	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXSW);
 
+	/* Hard failure: HW rekey gave up and the connection is now pinned to
+	 * SW encryption. The call site only sees the transient -EAGAIN retry
+	 * (rc is not propagated here), so emit the trace from the fallback
+	 * path itself; rc still holds the originating error.
+	 */
+	trace_tls_device_complete_rekey_fail(sk, rc);
+
 	return 0;
 }
 
@@ -2195,11 +2215,21 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
 			 * is installed once drain_start crosses rekey.old_nic_boundary.
 			 */
 			context->dev_add_pending = 1;
+			trace_tls_device_rekey_start(sk, drain_start,
+						     context->rekey.old_nic_boundary,
+						     true);
 		} else {
 			struct tcp_sock *tp = tcp_sk(sk);
 			u32 nic_end;
 
 			if (context->rekey.old_aead_recv) {
+				/* Prior rekey's era already drained (drain_start is
+				 * past old_nic_boundary), so retiring its key here
+				 * is a boundary crossing, same as the free in
+				 * tls_device_decrypted(); mark it done.
+				 */
+				trace_tls_device_rekey_done(sk, drain_start,
+							    context->rekey.old_nic_boundary);
 				crypto_free_aead(context->rekey.old_aead_recv);
 				context->rekey.old_aead_recv = NULL;
 			}
@@ -2247,6 +2277,8 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
 				context->dev_add_pending = 0;
 				retired_pending = true;
 			}
+			trace_tls_device_rekey_start(sk, drain_start, nic_end,
+						     before(drain_start, nic_end));
 		}
 	}
 
diff --git a/net/tls/trace.h b/net/tls/trace.h
index 2d8ce4ff3265..5b9c1f86d82d 100644
--- a/net/tls/trace.h
+++ b/net/tls/trace.h
@@ -192,6 +192,124 @@ TRACE_EVENT(tls_device_tx_resync_send,
 	)
 );
 
+TRACE_EVENT(tls_device_rekey_start,
+
+	TP_PROTO(struct sock *sk, u32 copied_seq, u32 nic_boundary,
+		 bool inflight),
+
+	TP_ARGS(sk, copied_seq, nic_boundary, inflight),
+
+	TP_STRUCT__entry(
+		__field(	struct sock *,	sk		)
+		__field(	u32,		copied_seq	)
+		__field(	u32,		nic_boundary	)
+		__field(	bool,		inflight	)
+	),
+
+	TP_fast_assign(
+		__entry->sk = sk;
+		__entry->copied_seq = copied_seq;
+		__entry->nic_boundary = nic_boundary;
+		__entry->inflight = inflight;
+	),
+
+	TP_printk(
+		"sk=%p copied_seq=%u nic_boundary=%u inflight=%d",
+		__entry->sk, __entry->copied_seq, __entry->nic_boundary,
+		__entry->inflight
+	)
+);
+
+TRACE_EVENT(tls_device_rekey_reencrypt,
+
+	TP_PROTO(struct sock *sk, u32 tcp_seq, u32 nic_boundary),
+
+	TP_ARGS(sk, tcp_seq, nic_boundary),
+
+	TP_STRUCT__entry(
+		__field(	struct sock *,	sk		)
+		__field(	u32,		tcp_seq		)
+		__field(	u32,		nic_boundary	)
+	),
+
+	TP_fast_assign(
+		__entry->sk = sk;
+		__entry->tcp_seq = tcp_seq;
+		__entry->nic_boundary = nic_boundary;
+	),
+
+	TP_printk(
+		"sk=%p tcp_seq=%u nic_boundary=%u",
+		__entry->sk, __entry->tcp_seq, __entry->nic_boundary
+	)
+);
+
+TRACE_EVENT(tls_device_rekey_done,
+
+	TP_PROTO(struct sock *sk, u32 tcp_seq, u32 nic_boundary),
+
+	TP_ARGS(sk, tcp_seq, nic_boundary),
+
+	TP_STRUCT__entry(
+		__field(	struct sock *,	sk		)
+		__field(	u32,		tcp_seq		)
+		__field(	u32,		nic_boundary	)
+	),
+
+	TP_fast_assign(
+		__entry->sk = sk;
+		__entry->tcp_seq = tcp_seq;
+		__entry->nic_boundary = nic_boundary;
+	),
+
+	TP_printk(
+		"sk=%p tcp_seq=%u nic_boundary=%u",
+		__entry->sk, __entry->tcp_seq, __entry->nic_boundary
+	)
+);
+
+TRACE_EVENT(tls_device_complete_rekey_fail,
+
+	TP_PROTO(struct sock *sk, int rc),
+
+	TP_ARGS(sk, rc),
+
+	TP_STRUCT__entry(
+		__field(	struct sock *,	sk	)
+		__field(	int,		rc	)
+	),
+
+	TP_fast_assign(
+		__entry->sk = sk;
+		__entry->rc = rc;
+	),
+
+	TP_printk(
+		"sk=%p rc=%d",
+		__entry->sk, __entry->rc
+	)
+);
+
+TRACE_EVENT(tls_device_complete_rekey_retry,
+
+	TP_PROTO(struct sock *sk),
+
+	TP_ARGS(sk),
+
+	TP_STRUCT__entry(
+		__field(	struct sock *,	sk	)
+	),
+
+	TP_fast_assign(
+		__entry->sk = sk;
+	),
+
+	TP_printk(
+		"sk=%p",
+		__entry->sk
+	)
+);
+
 #endif /* _TLS_TRACE_H_ */
 
 #undef TRACE_INCLUDE_PATH
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 14/15] selftests: net: add TLS hardware offload test
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (12 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 13/15] tls: device: add tracepoints for the KeyUpdate path Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  2026-09-17 22:35 ` [PATCH net-next v17 15/15] tls: document TLS 1.3 hardware offload rekey handling Rishikesh Jethwani
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Two-node kTLS HW offload test using NetDrvEpEnv. A C helper binary
acts as TLS client or server; a Python harness drives it and verifies
TLS stat counters (RekeyOk, RekeyReceived, RekeyFallback,
CurrRekey, RekeyAborted, RekeyError, DecryptError).

Covers TLS 1.2/1.3 with AES-GCM-128/256, rekey with various buffer
sizes, and burst variants that stress TX rekey (temporary SW phase,
HW reinstall) and RX rekey (boundary tracking, old-key reencryption,
deferred dev_add).

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 MAINTAINERS                                   |    2 +
 .../selftests/drivers/net/hw/.gitignore       |    1 +
 .../testing/selftests/drivers/net/hw/Makefile |    2 +
 tools/testing/selftests/drivers/net/hw/config |    2 +
 .../selftests/drivers/net/hw/tls_hw_offload.c | 1132 +++++++++++++++++
 .../drivers/net/hw/tls_hw_offload.py          |  446 +++++++
 6 files changed, 1585 insertions(+)
 create mode 100644 tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
 create mode 100755 tools/testing/selftests/drivers/net/hw/tls_hw_offload.py

diff --git a/MAINTAINERS b/MAINTAINERS
index 0e04d92d1b09..4f1645bf2ee5 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -19255,6 +19255,8 @@ F:	Documentation/networking/tls*
 F:	include/net/tls.h
 F:	include/uapi/linux/tls.h
 F:	net/tls/
+F:	tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
+F:	tools/testing/selftests/drivers/net/hw/tls_hw_offload.py
 F:	tools/testing/selftests/net/tls.c
 
 NETWORKING [SOCKETS]
diff --git a/tools/testing/selftests/drivers/net/hw/.gitignore b/tools/testing/selftests/drivers/net/hw/.gitignore
index 46540468a775..911a9bfeaf41 100644
--- a/tools/testing/selftests/drivers/net/hw/.gitignore
+++ b/tools/testing/selftests/drivers/net/hw/.gitignore
@@ -1,4 +1,5 @@
 # SPDX-License-Identifier: GPL-2.0-only
 iou-zcrx
 ncdevmem
+tls_hw_offload
 toeplitz
diff --git a/tools/testing/selftests/drivers/net/hw/Makefile b/tools/testing/selftests/drivers/net/hw/Makefile
index 8aebdc6feb17..b3831c2d09ea 100644
--- a/tools/testing/selftests/drivers/net/hw/Makefile
+++ b/tools/testing/selftests/drivers/net/hw/Makefile
@@ -46,6 +46,7 @@ TEST_PROGS = \
 	rss_drv.py \
 	rss_flow_label.py \
 	rss_input_xfrm.py \
+	tls_hw_offload.py \
 	toeplitz.py \
 	tso.py \
 	userns_devmem.py \
@@ -80,6 +81,7 @@ YNL_GEN_FILES := \
 # end of YNL_GEN_FILES
 TEST_GEN_FILES += $(YNL_GEN_FILES)
 TEST_GEN_FILES += $(patsubst %.c,%.o,$(wildcard *.bpf.c))
+TEST_GEN_FILES += tls_hw_offload
 
 include ../../../lib.mk
 
diff --git a/tools/testing/selftests/drivers/net/hw/config b/tools/testing/selftests/drivers/net/hw/config
index d89a9ba17655..169e608516bd 100644
--- a/tools/testing/selftests/drivers/net/hw/config
+++ b/tools/testing/selftests/drivers/net/hw/config
@@ -22,6 +22,8 @@ CONFIG_NET_IPIP=y
 CONFIG_NETKIT=y
 CONFIG_NET_SCH_INGRESS=y
 CONFIG_SYNC_FILE=y
+CONFIG_TLS=y
+CONFIG_TLS_DEVICE=y
 CONFIG_UDMABUF=y
 CONFIG_USER_NS=y
 CONFIG_VXLAN=y
diff --git a/tools/testing/selftests/drivers/net/hw/tls_hw_offload.c b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
new file mode 100644
index 000000000000..303c6752ace2
--- /dev/null
+++ b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
@@ -0,0 +1,1132 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * TLS Hardware Offload Two-Node Test
+ *
+ * Tests kTLS hardware offload between two physical nodes using
+ * hardcoded keys. Supports TLS 1.2/1.3, AES-GCM-128/256, and rekey.
+ */
+
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <errno.h>
+#include <limits.h>
+#include <time.h>
+#include <sys/time.h>
+#include <signal.h>
+#include <sys/types.h>
+#include <sys/socket.h>
+#include <netinet/in.h>
+#include <netinet/tcp.h>
+#include <netdb.h>
+#include <linux/tls.h>
+
+#define TLS_RECORD_TYPE_HANDSHAKE		22
+#define TLS_HANDSHAKE_KEY_UPDATE		0x18
+
+/* Large enough for a TLS 1.3 KeyUpdate handshake record's plaintext. */
+#define MIN_BUF_SIZE   16
+
+/* Initial key material */
+static struct tls12_crypto_info_aes_gcm_128 tls_info_key0_128 = {
+	.info = {
+		.version = TLS_1_3_VERSION,
+		.cipher_type = TLS_CIPHER_AES_GCM_128,
+	},
+	.iv = { 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08 },
+	.key = { 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08,
+		 0x09, 0x0a, 0x0b, 0x0c, 0x0d, 0x0e, 0x0f, 0x10 },
+	.salt = { 0x01, 0x02, 0x03, 0x04 },
+	.rec_seq = { 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00 },
+};
+
+static struct tls12_crypto_info_aes_gcm_256 tls_info_key0_256 = {
+	.info = {
+		.version = TLS_1_3_VERSION,
+		.cipher_type = TLS_CIPHER_AES_GCM_256,
+	},
+	.iv = { 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08 },
+	.key = { 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08,
+		 0x09, 0x0a, 0x0b, 0x0c, 0x0d, 0x0e, 0x0f, 0x10,
+		 0x11, 0x12, 0x13, 0x14, 0x15, 0x16, 0x17, 0x18,
+		 0x19, 0x1a, 0x1b, 0x1c, 0x1d, 0x1e, 0x1f, 0x20 },
+	.salt = { 0x01, 0x02, 0x03, 0x04 },
+	.rec_seq = { 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00 },
+};
+
+static int num_rekeys;
+static int num_iterations = 100;
+static int cipher_type = TLS_CIPHER_AES_GCM_128;
+static int tls_version = TLS_1_3_VERSION;
+static int server_port = 4433;
+static char *server_ip;
+/* Address family to force: AF_UNSPEC (any), AF_INET (-4), AF_INET6 (-6). */
+static int force_family = AF_UNSPEC;
+
+static int send_size = 16384;
+static int random_size_max;
+/* Burst mode: sender keeps pushing records without reading from the peer;
+ * receiver drains without echoing back. Only the client initiates rekey.
+ */
+static int burst_mode;
+static int zc_rx;
+
+/* XOR each byte with the generation so both endpoints derive the
+ * same per-generation key without a real KDF. Generation 0 leaves
+ * the base key unchanged.
+ */
+static void derive_key_fields(unsigned char *key, int key_size,
+			      unsigned char *iv, int iv_size,
+			      unsigned char *salt, int salt_size,
+			      unsigned char *rec_seq, int rec_seq_size,
+			      int generation)
+{
+	int i;
+
+	for (i = 0; i < key_size; i++)
+		key[i] ^= generation;
+	for (i = 0; i < iv_size; i++)
+		iv[i] ^= generation;
+	for (i = 0; i < salt_size; i++)
+		salt[i] ^= generation;
+	memset(rec_seq, 0, rec_seq_size);
+}
+
+static void derive_key_128(struct tls12_crypto_info_aes_gcm_128 *key,
+			   int generation)
+{
+	memcpy(key, &tls_info_key0_128, sizeof(*key));
+	key->info.version = tls_version;
+	derive_key_fields(key->key, TLS_CIPHER_AES_GCM_128_KEY_SIZE,
+			  key->iv, TLS_CIPHER_AES_GCM_128_IV_SIZE,
+			  key->salt, TLS_CIPHER_AES_GCM_128_SALT_SIZE,
+			  key->rec_seq, TLS_CIPHER_AES_GCM_128_REC_SEQ_SIZE,
+			  generation);
+}
+
+static void derive_key_256(struct tls12_crypto_info_aes_gcm_256 *key,
+			   int generation)
+{
+	memcpy(key, &tls_info_key0_256, sizeof(*key));
+	key->info.version = tls_version;
+	derive_key_fields(key->key, TLS_CIPHER_AES_GCM_256_KEY_SIZE,
+			  key->iv, TLS_CIPHER_AES_GCM_256_IV_SIZE,
+			  key->salt, TLS_CIPHER_AES_GCM_256_SALT_SIZE,
+			  key->rec_seq, TLS_CIPHER_AES_GCM_256_REC_SEQ_SIZE,
+			  generation);
+}
+
+static const char *cipher_name(int cipher)
+{
+	switch (cipher) {
+	case TLS_CIPHER_AES_GCM_128: return "AES-GCM-128";
+	case TLS_CIPHER_AES_GCM_256: return "AES-GCM-256";
+	default: return "unknown";
+	}
+}
+
+static const char *version_name(int version)
+{
+	switch (version) {
+	case TLS_1_2_VERSION: return "TLS 1.2";
+	case TLS_1_3_VERSION: return "TLS 1.3";
+	default: return "unknown";
+	}
+}
+
+static int setup_tls_ulp(int fd)
+{
+	int ret;
+
+	ret = setsockopt(fd, IPPROTO_TCP, TCP_ULP, "tls", sizeof("tls"));
+	if (ret < 0) {
+		printf("SETUP ERROR: TCP_ULP failed: %s\n", strerror(errno));
+		return -1;
+	}
+	return 0;
+}
+
+/* Echo (non-burst) mode drives both directions from a single thread: the
+ * client pushes a whole payload with one blocking send() and only reads the
+ * echo afterwards, while the server blocks in send() mid-echo. If a payload
+ * exceeds the peer's receive window the two sides deadlock - client stuck in
+ * send(), server stuck echoing, neither draining the other. Size the socket
+ * buffers so a full payload always fits in the peer's window (the forward
+ * send() then completes without needing the peer to read concurrently); the
+ * send/recv timeouts armed by set_io_timeouts() turn any residual stall into a
+ * loud EAGAIN instead of a hang.
+ */
+static void configure_echo_socket(int fd, int payload)
+{
+	int want = payload;
+
+	if (want < MIN_BUF_SIZE)
+		want = MIN_BUF_SIZE;
+
+	/* SO_*BUFFORCE bypasses the rmem_max/wmem_max sysctl caps (needs
+	 * CAP_NET_ADMIN); fall back to the best-effort, cap-limited option
+	 * when unprivileged - the timeouts below still turn any resulting
+	 * stall into a loud failure rather than a hang.
+	 */
+	if (setsockopt(fd, SOL_SOCKET, SO_RCVBUFFORCE, &want, sizeof(want)) < 0)
+		setsockopt(fd, SOL_SOCKET, SO_RCVBUF, &want, sizeof(want));
+	if (setsockopt(fd, SOL_SOCKET, SO_SNDBUFFORCE, &want, sizeof(want)) < 0)
+		setsockopt(fd, SOL_SOCKET, SO_SNDBUF, &want, sizeof(want));
+}
+
+/* Arm send/recv timeouts so any unexpected stall fails loudly with EAGAIN
+ * instead of hanging until the harness SIGKILLs us. Wanted in both echo and
+ * burst modes - burst mode has no other stall guard.
+ */
+static void set_io_timeouts(int fd)
+{
+	struct timeval tv = { .tv_sec = 8, .tv_usec = 0 };
+
+	setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv));
+	setsockopt(fd, SOL_SOCKET, SO_SNDTIMEO, &tv, sizeof(tv));
+}
+
+/* Send the whole buffer, looping over short counts. A blocking SOCK_STREAM
+ * send() may return fewer bytes than requested (e.g. when SO_SNDTIMEO fires
+ * after partial progress) without setting errno, so a short count is not an
+ * error - only a negative return is. Looping also sends each iteration as one
+ * uninterrupted run of bytes, which the peer's userspace reassembly in burst
+ * mode counts on to keep iterations aligned.
+ */
+static int send_all(int fd, const char *buf, ssize_t len)
+{
+	ssize_t sent = 0;
+	ssize_t ret;
+
+	while (sent < len) {
+		ret = send(fd, buf + sent, len - sent, 0);
+		if (ret < 0) {
+			printf("FAIL: send failed: %s\n", strerror(errno));
+			return -1;
+		}
+		sent += ret;
+	}
+	return 0;
+}
+
+static int set_zc_rx(int fd)
+{
+	int val = 1;
+
+	if (setsockopt(fd, SOL_TLS, TLS_RX_EXPECT_NO_PAD, &val,
+		       sizeof(val)) < 0) {
+		printf("SETUP ERROR: TLS_RX_EXPECT_NO_PAD failed: %s\n",
+		       strerror(errno));
+		return -1;
+	}
+	return 0;
+}
+
+/* Send a TLS 1.3 KeyUpdate handshake record. The kernel only
+ * inspects the HandshakeType byte to detect KeyUpdate, so don't
+ * bother with the 3-byte length or request_update fields.
+ */
+static int send_tls_key_update(int fd)
+{
+	char cmsg_buf[CMSG_SPACE(sizeof(unsigned char))];
+	unsigned char key_update_msg = TLS_HANDSHAKE_KEY_UPDATE;
+	struct msghdr msg = {0};
+	struct cmsghdr *cmsg;
+	struct iovec iov;
+
+	iov.iov_base = &key_update_msg;
+	iov.iov_len = sizeof(key_update_msg);
+
+	msg.msg_iov = &iov;
+	msg.msg_iovlen = 1;
+	msg.msg_control = cmsg_buf;
+	msg.msg_controllen = sizeof(cmsg_buf);
+
+	cmsg = CMSG_FIRSTHDR(&msg);
+	cmsg->cmsg_level = SOL_TLS;
+	cmsg->cmsg_type = TLS_SET_RECORD_TYPE;
+	cmsg->cmsg_len = CMSG_LEN(sizeof(unsigned char));
+	*CMSG_DATA(cmsg) = TLS_RECORD_TYPE_HANDSHAKE;
+	msg.msg_controllen = cmsg->cmsg_len;
+
+	if (sendmsg(fd, &msg, 0) < 0) {
+		printf("sendmsg KeyUpdate failed: %s\n", strerror(errno));
+		return -1;
+	}
+
+	printf("Sent TLS KeyUpdate handshake message\n");
+	return 0;
+}
+
+static int recv_tls_message(int fd, char *buf, size_t buflen, int *record_type,
+			    int flags)
+{
+	char cmsg_buf[CMSG_SPACE(sizeof(unsigned char))];
+	struct msghdr msg = {0};
+	struct cmsghdr *cmsg;
+	struct iovec iov;
+	int ret;
+
+	iov.iov_base = buf;
+	iov.iov_len = buflen;
+
+	msg.msg_iov = &iov;
+	msg.msg_iovlen = 1;
+	msg.msg_control = cmsg_buf;
+	msg.msg_controllen = sizeof(cmsg_buf);
+
+	ret = recvmsg(fd, &msg, flags);
+	if (ret <= 0)
+		return ret;
+
+	cmsg = CMSG_FIRSTHDR(&msg);
+	if (cmsg && cmsg->cmsg_level == SOL_TLS &&
+	    cmsg->cmsg_type == TLS_GET_RECORD_TYPE)
+		*record_type = *((unsigned char *)CMSG_DATA(cmsg));
+
+	return ret;
+}
+
+/* Confirm a handshake record starting with HandshakeType KeyUpdate. */
+static int check_keyupdate(const char *buf, int len, int record_type)
+{
+	if (record_type != TLS_RECORD_TYPE_HANDSHAKE) {
+		printf("Expected handshake record (0x%02x), got 0x%02x\n",
+		       TLS_RECORD_TYPE_HANDSHAKE, record_type);
+		return -1;
+	}
+	if (len < 1 || (unsigned char)buf[0] != TLS_HANDSHAKE_KEY_UPDATE) {
+		printf("Expected KeyUpdate (0x%02x), got 0x%02x\n",
+		       TLS_HANDSHAKE_KEY_UPDATE,
+		       len ? (unsigned char)buf[0] : 0);
+		return -1;
+	}
+	printf("Received TLS KeyUpdate\n");
+	return 0;
+}
+
+static int recv_tls_keyupdate(int fd)
+{
+	char buf[MIN_BUF_SIZE];
+	int record_type = 0;
+	int ret;
+
+	ret = recv_tls_message(fd, buf, sizeof(buf), &record_type, 0);
+	if (ret < 0) {
+		printf("recv_tls_message failed: %s\n", strerror(errno));
+		return -1;
+	}
+
+	return check_keyupdate(buf, ret, record_type);
+}
+
+static int check_ekeyexpired(int fd)
+{
+	char buf[MIN_BUF_SIZE];
+	int ret;
+
+	ret = recv(fd, buf, sizeof(buf), MSG_DONTWAIT);
+	if (ret == -1 && errno == EKEYEXPIRED) {
+		printf("recv() returned EKEYEXPIRED as expected\n");
+		return 0;
+	}
+	if (ret > 0) {
+		printf("FAIL: recv() returned %d bytes, expected EKEYEXPIRED\n",
+		       ret);
+		return -1;
+	}
+	if (ret == 0) {
+		printf("FAIL: connection closed during rekey\n");
+		return -1;
+	}
+	printf("FAIL: recv() returned unexpected error: %s\n",
+	       strerror(errno));
+	return -1;
+}
+
+static int do_tls_rekey(int fd, int direction, int generation, int cipher)
+{
+	const char *dir = direction == TLS_TX ? "TX" : "RX";
+	int ret;
+
+	printf("%s TLS_%s %s gen %d...\n",
+	       generation ? "Rekeying" : "Installing",
+	       dir, cipher_name(cipher), generation);
+
+	if (cipher == TLS_CIPHER_AES_GCM_256) {
+		struct tls12_crypto_info_aes_gcm_256 key;
+
+		derive_key_256(&key, generation);
+		ret = setsockopt(fd, SOL_TLS, direction, &key, sizeof(key));
+	} else {
+		struct tls12_crypto_info_aes_gcm_128 key;
+
+		derive_key_128(&key, generation);
+		ret = setsockopt(fd, SOL_TLS, direction, &key, sizeof(key));
+	}
+
+	if (ret < 0) {
+		printf("%sTLS_%s %s gen %d failed: %s\n",
+		       generation ? "" : "SETUP ERROR: ", dir,
+		       cipher_name(cipher), generation, strerror(errno));
+		return -1;
+	}
+	printf("TLS_%s %s gen %d installed\n",
+	       dir, cipher_name(cipher), generation);
+	return 0;
+}
+
+/* Open a TCP connection to server_ip:server_port, switch to the TLS
+ * ULP, and install initial generation-0 TX/RX keys. Works over IPv4 or
+ * IPv6: getaddrinfo() resolves server_ip (honouring any -4/-6 forced
+ * family and %zone scope IDs in link-local addresses). Returns the fd on
+ * success, -1 on error (with the fd already closed).
+ */
+static int client_connect_tls(void)
+{
+	struct addrinfo hints = {0}, *res, *rp;
+	char port_str[16];
+	int csk = -1;
+	int ret;
+
+	hints.ai_family = force_family;
+	hints.ai_socktype = SOCK_STREAM;
+	hints.ai_protocol = IPPROTO_TCP;
+	snprintf(port_str, sizeof(port_str), "%d", server_port);
+
+	ret = getaddrinfo(server_ip, port_str, &hints, &res);
+	if (ret) {
+		printf("SETUP ERROR: getaddrinfo(%s): %s\n", server_ip,
+		       gai_strerror(ret));
+		return -1;
+	}
+
+	printf("Connecting to %s:%d...\n", server_ip, server_port);
+	for (rp = res; rp; rp = rp->ai_next) {
+		csk = socket(rp->ai_family, rp->ai_socktype, rp->ai_protocol);
+		if (csk < 0)
+			continue;
+		if (connect(csk, rp->ai_addr, rp->ai_addrlen) == 0)
+			break;
+		close(csk);
+		csk = -1;
+	}
+	freeaddrinfo(res);
+
+	if (csk < 0) {
+		printf("SETUP ERROR: connect to %s:%d failed: %s\n",
+		       server_ip, server_port, strerror(errno));
+		return -1;
+	}
+	printf("Connected!\n");
+
+	if (setup_tls_ulp(csk) < 0)
+		goto err;
+
+	if (do_tls_rekey(csk, TLS_TX, 0, cipher_type) < 0 ||
+	    do_tls_rekey(csk, TLS_RX, 0, cipher_type) < 0)
+		goto err;
+
+	set_io_timeouts(csk);
+	if (!burst_mode)
+		configure_echo_socket(csk, random_size_max > 0 ?
+					    random_size_max : send_size);
+
+	return csk;
+err:
+	close(csk);
+	return -1;
+}
+
+/* Drain `len` echoed bytes from the server and verify they match the
+ * payload we just sent.
+ */
+static int client_recv_echo(int fd, const char *sent, char *echo_buf,
+			    ssize_t len)
+{
+	ssize_t total = 0;
+	ssize_t n;
+
+	while (total < len) {
+		n = recv(fd, echo_buf + total, len - total, 0);
+		if (n < 0) {
+			printf("FAIL: Echo recv failed: %s\n", strerror(errno));
+			return -1;
+		}
+		if (n == 0) {
+			printf("FAIL: Connection closed during echo\n");
+			return -1;
+		}
+		total += n;
+	}
+
+	if (memcmp(sent, echo_buf, len) != 0) {
+		printf("FAIL: Echo data mismatch!\n");
+		return -1;
+	}
+	printf("Received echo %zd bytes (ok)\n", total);
+	return 0;
+}
+
+/* Client side of a rekey: send KeyUpdate and rotate TX. In echo mode
+ * also wait for the peer's KeyUpdate and rotate RX.
+ */
+static int client_rekey(int fd, int generation)
+{
+	if (send_tls_key_update(fd) < 0) {
+		printf("FAIL: send KeyUpdate\n");
+		return -1;
+	}
+
+	if (do_tls_rekey(fd, TLS_TX, generation, cipher_type) < 0)
+		return -1;
+
+	if (burst_mode)
+		return 0;
+
+	if (recv_tls_keyupdate(fd) < 0) {
+		printf("FAIL: recv KeyUpdate from server\n");
+		return -1;
+	}
+
+	if (check_ekeyexpired(fd) < 0)
+		return -1;
+
+	return do_tls_rekey(fd, TLS_RX, generation, cipher_type);
+}
+
+static int do_client(void)
+{
+	char *buf = NULL, *echo_buf = NULL;
+	int max_size, rekey_interval;
+	int csk = -1, i;
+	int test_result = -1;
+	int current_gen = 0;
+	int next_rekey_at;
+	ssize_t n;
+
+	max_size = random_size_max > 0 ? random_size_max : send_size;
+	if (max_size < MIN_BUF_SIZE)
+		max_size = MIN_BUF_SIZE;
+	buf = malloc(max_size);
+	if (!burst_mode)
+		echo_buf = malloc(max_size);
+	if (!buf || (!burst_mode && !echo_buf)) {
+		printf("SETUP ERROR: failed to allocate buffers\n");
+		goto out;
+	}
+
+	csk = client_connect_tls();
+	if (csk < 0)
+		goto out;
+
+	if (num_rekeys)
+		printf("TLS %s setup complete. Will perform %d rekey(s).\n",
+		       cipher_name(cipher_type), num_rekeys);
+	else
+		printf("TLS setup complete.\n");
+
+	if (random_size_max > 0)
+		printf("Sending %d messages of random size (1..%d bytes)...\n",
+		       num_iterations, random_size_max);
+	else
+		printf("Sending %d messages of %d bytes...\n",
+		       num_iterations, send_size);
+
+	rekey_interval = num_iterations / (num_rekeys + 1);
+	next_rekey_at = rekey_interval;
+
+	for (i = 1; i <= num_iterations; i++) {
+		int this_size;
+
+		if (random_size_max > 0)
+			this_size = (rand() % random_size_max) + 1;
+		else
+			this_size = send_size;
+
+		/* In burst mode, use a per-iteration fill pattern so the
+		 * receiver can detect any plaintext corruption without a
+		 * round-trip echo.
+		 */
+		if (burst_mode) {
+			memset(buf, i & 0xFF, this_size);
+		} else {
+			int j;
+
+			for (j = 0; j < this_size; j++)
+				buf[j] = rand() & 0xFF;
+		}
+
+		if (send_all(csk, buf, this_size) < 0)
+			goto out;
+		n = this_size;
+
+		if (!burst_mode) {
+			printf("Sent %zd bytes (iteration %d)\n", n, i);
+			if (client_recv_echo(csk, buf, echo_buf, n) < 0)
+				goto out;
+		}
+
+		/* Rekey at intervals. In echo mode this is a full bidirectional
+		 * exchange; in burst mode the client only rotates its TX key
+		 * and sends KeyUpdate - the peer is expected to follow.
+		 */
+		if (num_rekeys && current_gen < num_rekeys &&
+		    i == next_rekey_at) {
+			current_gen++;
+			printf("\n=== Client Rekey gen %d ===\n", current_gen);
+
+			if (client_rekey(csk, current_gen) < 0)
+				goto out;
+
+			next_rekey_at += rekey_interval;
+			printf("=== Client Rekey gen %d Complete ===\n\n",
+			       current_gen);
+		}
+	}
+
+	test_result = 0;
+out:
+	if (num_rekeys)
+		printf("Rekeys completed: %d/%d\n", current_gen, num_rekeys);
+	if (csk >= 0)
+		close(csk);
+	free(buf);
+	free(echo_buf);
+	return test_result;
+}
+
+/* Bind/listen on server_port, accept one client, switch to the TLS ULP
+ * and install initial generation-0 keys (plus zc_rx if requested).
+ * Returns the connected fd on success and writes the listener fd to
+ * *lsk_out so the caller can close it. Returns -1 on error, with all
+ * intermediate fds already closed and *lsk_out left at -1.
+ */
+static int server_accept_tls(int *lsk_out)
+{
+	struct addrinfo hints = {0}, *res, *rp;
+	int lsk = -1, csk, one = 1;
+	char port_str[16];
+	int ret;
+
+	*lsk_out = -1;
+
+	/* AI_PASSIVE gives a wildcard bind address for the chosen family
+	 * (0.0.0.0 / ::). The family is forced by -4/-6; when unspecified,
+	 * bind the first entry that works.
+	 */
+	hints.ai_family = force_family;
+	hints.ai_socktype = SOCK_STREAM;
+	hints.ai_protocol = IPPROTO_TCP;
+	hints.ai_flags = AI_PASSIVE;
+	snprintf(port_str, sizeof(port_str), "%d", server_port);
+
+	ret = getaddrinfo(NULL, port_str, &hints, &res);
+	if (ret) {
+		printf("SETUP ERROR: getaddrinfo(port %d): %s\n", server_port,
+		       gai_strerror(ret));
+		return -1;
+	}
+
+	for (rp = res; rp; rp = rp->ai_next) {
+		lsk = socket(rp->ai_family, rp->ai_socktype, rp->ai_protocol);
+		if (lsk < 0)
+			continue;
+		setsockopt(lsk, SOL_SOCKET, SO_REUSEADDR, &one, sizeof(one));
+		if (bind(lsk, rp->ai_addr, rp->ai_addrlen) == 0)
+			break;
+		close(lsk);
+		lsk = -1;
+	}
+	freeaddrinfo(res);
+
+	if (lsk < 0) {
+		printf("SETUP ERROR: failed to bind port %d: %s\n",
+		       server_port, strerror(errno));
+		return -1;
+	}
+
+	if (listen(lsk, 1) < 0) {
+		printf("SETUP ERROR: listen failed: %s\n", strerror(errno));
+		close(lsk);
+		return -1;
+	}
+
+	printf("Server listening on port %d\n", server_port);
+	printf("Waiting for client connection...\n");
+
+	/* Bound accept() so a client that never connects (a deploy or connect
+	 * failure on the peer) does not block the server forever and leak the
+	 * process past the harness timeout. accept() honours SO_RCVTIMEO on the
+	 * listening socket; the client connects right after wait_port_listen(),
+	 * so 30s is generous.
+	 */
+	{
+		struct timeval tv = { .tv_sec = 30, .tv_usec = 0 };
+
+		setsockopt(lsk, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv));
+	}
+
+	csk = accept(lsk, (struct sockaddr *)NULL, (socklen_t *)NULL);
+	if (csk < 0) {
+		if (errno == EAGAIN || errno == EWOULDBLOCK)
+			printf("SETUP ERROR: accept timed out; client never connected\n");
+		else
+			printf("SETUP ERROR: accept failed: %s\n", strerror(errno));
+		close(lsk);
+		return -1;
+	}
+	printf("Client connected!\n");
+
+	if (setup_tls_ulp(csk) < 0)
+		goto err;
+
+	if (do_tls_rekey(csk, TLS_TX, 0, cipher_type) < 0 ||
+	    do_tls_rekey(csk, TLS_RX, 0, cipher_type) < 0)
+		goto err;
+
+	if (zc_rx && set_zc_rx(csk) < 0)
+		goto err;
+
+	set_io_timeouts(csk);
+	if (!burst_mode)
+		configure_echo_socket(csk, random_size_max > 0 ?
+					    random_size_max : send_size);
+
+	*lsk_out = lsk;
+	return csk;
+err:
+	close(csk);
+	close(lsk);
+	return -1;
+}
+
+/* Server side of a rekey: confirm recv() reports EKEYEXPIRED, then rotate RX.
+ * In echo mode also send a KeyUpdate back and rotate TX.
+ */
+static int server_rekey(int fd, int generation)
+{
+	if (check_ekeyexpired(fd) < 0)
+		return -1;
+
+	if (do_tls_rekey(fd, TLS_RX, generation, cipher_type) < 0)
+		return -1;
+
+	if (burst_mode)
+		return 0;
+
+	if (send_tls_key_update(fd) < 0) {
+		printf("FAIL: send KeyUpdate\n");
+		return -1;
+	}
+
+	return do_tls_rekey(fd, TLS_TX, generation, cipher_type);
+}
+
+/* Burst mode: verify one reassembled iteration of send_size plaintext bytes,
+ * each filled with (send_iter & 0xff). Catches decrypt-succeeded-but-
+ * plaintext-corrupt bugs that AEAD counters alone would miss.
+ */
+static int server_verify_burst(const char *buf, int send_iter)
+{
+	unsigned char expect = send_iter & 0xFF;
+	int j;
+
+	for (j = 0; j < send_size; j++) {
+		if ((unsigned char)buf[j] != expect) {
+			printf("FAIL: data mismatch iter %d off %d: exp 0x%02x got 0x%02x\n",
+			       send_iter, j, expect, (unsigned char)buf[j]);
+			return -1;
+		}
+	}
+	return 0;
+}
+
+static int do_server(void)
+{
+	int lsk = -1, csk = -1;
+	ssize_t n, total = 0;
+	int test_result = -1;
+	int current_gen = 0;
+	int recv_count = 0;
+	int send_iter = 1;
+	char *buf = NULL;
+	int record_type = 0;
+	int filled = 0;
+	int buf_size;
+
+	buf_size = send_size;
+	if (buf_size < MIN_BUF_SIZE)
+		buf_size = MIN_BUF_SIZE;
+	buf = malloc(buf_size);
+	if (!buf) {
+		printf("SETUP ERROR: failed to allocate buffer\n");
+		goto out;
+	}
+
+	csk = server_accept_tls(&lsk);
+	if (csk < 0)
+		goto out;
+
+	printf("TLS %s setup complete. Receiving...\n",
+	       cipher_name(cipher_type));
+
+	/* Burst mode: reassemble one iteration (send_size bytes) in userspace
+	 * from however much each recv returns, rather than demanding a full
+	 * send_size batch in a single MSG_WAITALL call. A blocking MSG_WAITALL
+	 * of send_size deadlocks when the client's last record of an iteration
+	 * is still partly in flight as its socket buffer fills: the server
+	 * waits for bytes the client cannot send until the server reads, and
+	 * the server will not read until it has the whole batch. Draining
+	 * whatever is available keeps the receive window open and breaks that
+	 * cycle. kTLS never splits a record and returns data and control
+	 * (KeyUpdate) records separately, and each iteration is a whole number
+	 * of records, so capping each recv at the iteration boundary keeps the
+	 * reassembly aligned and delivers a KeyUpdate on its own.
+	 */
+
+	/* Main receive loop */
+	while (1) {
+		char *dst = burst_mode ? buf + filled : buf;
+		size_t want = burst_mode ? (size_t)(send_size - filled)
+					 : (size_t)buf_size;
+
+		n = recv_tls_message(csk, dst, want, &record_type, 0);
+		if (n == 0) {
+			/* A clean close on an iteration boundary is success;
+			 * one with a partial iteration still buffered means the
+			 * peer dropped the tail - the truncated-data case this
+			 * test exists to catch, so fail loudly.
+			 */
+			if (burst_mode && filled) {
+				printf("FAIL: closed mid-iteration (%d/%d bytes buffered)\n",
+				       filled, send_size);
+				goto out;
+			}
+			printf("Connection closed by client\n");
+			break;
+		}
+		if (n < 0) {
+			printf("FAIL: recv failed: %s\n", strerror(errno));
+			goto out;
+		}
+
+		/* Handle KeyUpdate. In echo mode the server mirrors the
+		 * rekey back to the peer; in burst mode it only rotates its
+		 * RX key and keeps draining. A KeyUpdate always lands on a
+		 * send_size boundary, so no partial iteration must be buffered
+		 * when one arrives.
+		 */
+		if (record_type == TLS_RECORD_TYPE_HANDSHAKE) {
+			/* Check for a partial iteration before validating the
+			 * KeyUpdate, so a mid-iteration arrival fails with this
+			 * message rather than a misleading KeyUpdate-OK line.
+			 */
+			if (burst_mode && filled) {
+				printf("FAIL: KeyUpdate mid-iteration (%d/%d bytes buffered)\n",
+				       filled, send_size);
+				goto out;
+			}
+			if (check_keyupdate(dst, n, record_type) < 0)
+				goto out;
+			current_gen++;
+			printf("\n=== Server Rekey gen %d ===\n", current_gen);
+
+			if (server_rekey(csk, current_gen) < 0)
+				goto out;
+
+			printf("=== Server Rekey gen %d Complete ===\n\n",
+			       current_gen);
+			continue;
+		}
+
+		total += n;
+
+		if (burst_mode) {
+			filled += n;
+			if (filled < send_size)
+				continue;
+			if (server_verify_burst(buf, send_iter) < 0)
+				goto out;
+			recv_count++;
+			send_iter++;
+			filled = 0;
+			continue;
+		}
+
+		recv_count++;
+		printf("Received %zd bytes (total: %zd, count: %d)\n",
+		       n, total, recv_count);
+
+		if (send_all(csk, buf, n) < 0)
+			goto out;
+		printf("Echoed %zd bytes back to client\n", n);
+	}
+
+	test_result = 0;
+out:
+	printf("Connection closed. Total received: %zd bytes\n", total);
+	if (num_rekeys)
+		printf("Rekeys completed: %d\n", current_gen);
+
+	if (csk >= 0)
+		close(csk);
+	if (lsk >= 0)
+		close(lsk);
+	free(buf);
+	return test_result;
+}
+
+static int parse_int_arg(const char *arg, int min, int max,
+			 const char *name, int *out)
+{
+	char *endp;
+	long val;
+
+	errno = 0;
+	val = strtol(arg, &endp, 10);
+	if (errno || endp == arg || *endp != '\0' || val < min || val > max) {
+		if (max == INT_MAX)
+			printf("ERROR: Invalid %s '%s'. Must be >= %d.\n",
+			       name, arg, min);
+		else
+			printf("ERROR: Invalid %s '%s'. Must be %d..%d.\n",
+			       name, arg, min, max);
+		return -1;
+	}
+	*out = (int)val;
+	return 0;
+}
+
+static int parse_cipher_option(const char *arg)
+{
+	if (strcmp(arg, "128") == 0) {
+		cipher_type = TLS_CIPHER_AES_GCM_128;
+		return 0;
+	} else if (strcmp(arg, "256") == 0) {
+		cipher_type = TLS_CIPHER_AES_GCM_256;
+		return 0;
+	}
+	printf("ERROR: Invalid cipher '%s'. Must be 128 or 256.\n", arg);
+	return -1;
+}
+
+static int parse_version_option(const char *arg)
+{
+	if (strcmp(arg, "1.2") == 0) {
+		tls_version = TLS_1_2_VERSION;
+		return 0;
+	} else if (strcmp(arg, "1.3") == 0) {
+		tls_version = TLS_1_3_VERSION;
+		return 0;
+	}
+	printf("ERROR: Invalid TLS version '%s'. Must be 1.2 or 1.3.\n", arg);
+	return -1;
+}
+
+static void print_usage(const char *prog)
+{
+	printf("TLS Hardware Offload Two-Node Test\n\n");
+	printf("Usage:\n");
+	printf("  %s server [OPTIONS]\n", prog);
+	printf("  %s client -s <ip> [OPTIONS]\n", prog);
+	printf("\nOptions:\n");
+	printf("  -s <ip>       Server IP address, v4 or v6 (client, required)\n");
+	printf("  -p <port>     Server port (default: 4433)\n");
+	printf("  -4            Force IPv4 (default: auto/either)\n");
+	printf("  -6            Force IPv6 (default: auto/either)\n");
+	printf("  -b <size>     Send buffer size in bytes (default: 16384)\n");
+	printf("  -r <max>      Use random send buffer sizes (1..<max>)\n");
+	printf("  -v <version>  TLS version: 1.2 or 1.3 (default: 1.3)\n");
+	printf("  -c <cipher>   Cipher: 128 or 256 (default: 128)\n");
+	printf("  -n <N>        Number of send/echo iterations (default: 100)\n");
+	printf("  -k <N>        Perform N rekeys (client only, TLS 1.3; N < iterations)\n");
+	printf("  -B            Burst mode: client sends continuously without echo;\n");
+	printf("                server drains and handles KeyUpdate without responding.\n");
+	printf("  -Z            Set TLS_RX_EXPECT_NO_PAD on the server: TLS 1.3\n");
+	printf("                opt-in to the zero-copy RX fast path. Not needed\n");
+	printf("                for TLS 1.2 (always eligible). Server only.\n");
+	printf("  -h            Show this help message\n");
+	printf("\nExample:\n");
+	printf("  Node A: %s server\n", prog);
+	printf("  Node B: %s client -s 192.168.20.2\n", prog);
+	printf("\nRekey Example (3 rekeys, TLS 1.3 only):\n");
+	printf("  Node A: %s server\n", prog);
+	printf("  Node B: %s client -s 192.168.20.2 -k 3\n", prog);
+	printf("\nBurst Mode Example (client stresses TX rekey under load):\n");
+	printf("  Node A: %s server -B\n", prog);
+	printf("  Node B: %s client -s 192.168.20.2 -B -k 3\n", prog);
+	printf("\nIPv6 Example:\n");
+	printf("  Node A: %s server -6\n", prog);
+	printf("  Node B: %s client -6 -s fd00::2\n", prog);
+}
+
+int main(int argc, char *argv[])
+{
+	int send_size_set = 0;
+	int is_server;
+	int opt;
+
+	/* When the peer aborts a TLS connection (e.g. tls_err_abort() on a
+	 * failed decrypt), a send() here would raise SIGPIPE and kill us by
+	 * signal, so the harness sees only a bare non-zero exit with no
+	 * "FAIL:" line. Ignore it and let send()/sendmsg() return EPIPE, which
+	 * send_all()/send_tls_key_update() report.
+	 */
+	signal(SIGPIPE, SIG_IGN);
+
+	if (argc < 2 ||
+	    (strcmp(argv[1], "server") && strcmp(argv[1], "client"))) {
+		print_usage(argv[0]);
+		return 1;
+	}
+	is_server = !strcmp(argv[1], "server");
+
+	optind = 2; /* skip subcommand */
+	while ((opt = getopt(argc, argv, "s:p:b:r:c:v:k:n:BZ46h")) != -1) {
+		switch (opt) {
+		case 's':
+			server_ip = optarg;
+			break;
+		case '4':
+			if (force_family == AF_INET6) {
+				printf("ERROR: -4 and -6 are mutually exclusive\n");
+				return 1;
+			}
+			force_family = AF_INET;
+			break;
+		case '6':
+			if (force_family == AF_INET) {
+				printf("ERROR: -4 and -6 are mutually exclusive\n");
+				return 1;
+			}
+			force_family = AF_INET6;
+			break;
+		case 'B':
+			burst_mode = 1;
+			break;
+		case 'Z':
+			zc_rx = 1;
+			break;
+		case 'p':
+			if (parse_int_arg(optarg, 1, 65535, "port",
+					  &server_port) < 0)
+				return 1;
+			break;
+		case 'b':
+			if (parse_int_arg(optarg, 1, INT_MAX, "buffer size",
+					  &send_size) < 0)
+				return 1;
+			send_size_set = 1;
+			break;
+		case 'r':
+			if (parse_int_arg(optarg, 1, INT_MAX, "random size",
+					  &random_size_max) < 0)
+				return 1;
+			break;
+		case 'c':
+			if (parse_cipher_option(optarg) < 0)
+				return 1;
+			break;
+		case 'v':
+			if (parse_version_option(optarg) < 0)
+				return 1;
+			break;
+		case 'k':
+			if (parse_int_arg(optarg, 1, 255, "rekey count",
+					  &num_rekeys) < 0)
+				return 1;
+			break;
+		case 'n':
+			if (parse_int_arg(optarg, 1, INT_MAX, "iteration count",
+					  &num_iterations) < 0)
+				return 1;
+			break;
+		case 'h':
+			print_usage(argv[0]);
+			return 0;
+		default:
+			print_usage(argv[0]);
+			return 1;
+		}
+	}
+
+	if (send_size_set && random_size_max > 0) {
+		printf("ERROR: -b and -r are mutually exclusive\n");
+		return 1;
+	}
+
+	if (zc_rx && tls_version != TLS_1_3_VERSION) {
+		printf("ERROR: -Z (TLS_RX_EXPECT_NO_PAD) requires TLS 1.3\n");
+		return 1;
+	}
+
+	if (burst_mode && random_size_max > 0) {
+		printf("ERROR: -B and -r are mutually exclusive\n");
+		return 1;
+	}
+
+	if (burst_mode && send_size < MIN_BUF_SIZE) {
+		printf("ERROR: -b must be >= %d in burst mode (-B)\n",
+		       MIN_BUF_SIZE);
+		return 1;
+	}
+
+	if (is_server) {
+		if (server_ip) {
+			printf("warning: -s is ignored in server mode\n");
+			server_ip = NULL;
+		}
+		if (random_size_max > 0) {
+			printf("warning: -r is ignored in server mode\n");
+			random_size_max = 0;
+		}
+		if (num_rekeys) {
+			printf("warning: -k is ignored in server mode\n");
+			num_rekeys = 0;
+		}
+	} else {
+		if (!server_ip) {
+			printf("ERROR: Client requires -s <ip> option\n");
+			return 1;
+		}
+		if (tls_version == TLS_1_2_VERSION && num_rekeys) {
+			printf("ERROR: TLS 1.2 does not support rekey\n");
+			return 1;
+		}
+		if (num_rekeys >= num_iterations) {
+			printf("ERROR: num_rekeys (%d) must be < num_iterations (%d)\n",
+			       num_rekeys, num_iterations);
+			return 1;
+		}
+		if (zc_rx) {
+			printf("ERROR: -Z applies to the server (receiver) only\n");
+			return 1;
+		}
+	}
+
+	printf("TLS Version: %s\n", version_name(tls_version));
+	printf("Cipher: %s\n", cipher_name(cipher_type));
+	printf("Address family: %s\n",
+	       force_family == AF_INET ? "IPv4" :
+	       force_family == AF_INET6 ? "IPv6" : "auto");
+	if (random_size_max > 0)
+		printf("Buffer size: random (1..%d)\n", random_size_max);
+	else
+		printf("Buffer size: %d\n", send_size);
+
+	if (num_rekeys)
+		printf("Rekey testing ENABLED: %d rekey(s)\n", num_rekeys);
+	if (burst_mode)
+		printf("Burst mode ENABLED\n");
+	if (zc_rx)
+		printf("TLS_RX_EXPECT_NO_PAD ENABLED\n");
+
+	srand(time(NULL));
+
+	if (is_server)
+		return do_server() ? 1 : 0;
+
+	return do_client() ? 1 : 0;
+}
diff --git a/tools/testing/selftests/drivers/net/hw/tls_hw_offload.py b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.py
new file mode 100755
index 000000000000..99ae5b3b8996
--- /dev/null
+++ b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.py
@@ -0,0 +1,446 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+
+"""Test kTLS hardware offload using a C helper binary."""
+
+from collections import defaultdict
+
+from lib.py import ksft_run, ksft_exit, ksft_pr, KsftSkipEx
+from lib.py import ksft_ge, ksft_eq
+from lib.py import ksft_variants, KsftNamedVariant
+from lib.py import NetDrvEpEnv
+from lib.py import cmd, bkg, wait_port_listen, rand_port
+from lib.py import CmdExitFailure
+
+# Burst variants push hundreds of MB and perform many rekeys, so they
+# need far longer than the default cmd() timeout.
+BURST_TIMEOUT_S = 180
+REKEY_TIMEOUT_S = 90
+
+# Reading /proc/net/tls_stat is trivial locally, but on the remote it runs
+# over ssh, where connection setup can occasionally spike past the short
+# default cmd() timeout. Give these tiny reads plenty of headroom so a slow
+# ssh round-trip doesn't fail an otherwise-good variant.
+STATS_TIMEOUT_S = 30
+
+# Per-packet HW crypto counters exposed via `ethtool -S` on the DUT NIC,
+# keyed by the `ethtool -i` driver name. TlsTxDevice/TlsRxDevice in
+# /proc/net/tls_stat only prove tls_dev_add() accepted the offload; these
+# increment once per packet the NIC actually encrypted/decrypted (mlx5 counts
+# gso_segs, not records), so they prove the HW crypto path was exercised.
+# Names are driver-specific, so the
+# check only runs on drivers listed here and is skipped (not failed) on
+# others, keeping the test portable across NICs.
+HW_CRYPTO_COUNTERS = {
+    'mlx5_core': {'Tx': 'tx_tls_encrypted_packets',
+                  'Rx': 'rx_tls_decrypted_packets'},
+}
+
+
+def check_tls_support(cfg):
+    """Skip the suite unless both hosts have kTLS and the DUT HW offload."""
+    # The tls module is autoloaded lazily on the first TCP_ULP="tls"
+    # setsockopt, so /proc/net/tls_stat (created from the module's pernet
+    # init) may not exist yet on a freshly booted host. Load the module
+    # explicitly before probing for it.
+    try:
+        cmd("modprobe tls")
+        cmd("modprobe tls", host=cfg.remote)
+        cmd("test -f /proc/net/tls_stat")
+        cmd("test -f /proc/net/tls_stat", host=cfg.remote)
+    except CmdExitFailure as e:
+        raise KsftSkipEx(f"kTLS not supported: {e}") from e
+
+    try:
+        features = cmd(f"ethtool -k {cfg.ifname}").stdout
+        if 'tls-hw-tx-offload: on' not in features:
+            raise KsftSkipEx("Device does not support TLS HW TX offload")
+        if 'tls-hw-rx-offload: on' not in features:
+            raise KsftSkipEx("Device does not support TLS HW RX offload")
+    except CmdExitFailure as e:
+        raise KsftSkipEx(f"Cannot determine TLS HW offload support: {e}") from e
+
+
+def read_tls_stats(host=None):
+    """Snapshot the per-netns TLS MIB from /proc/net/tls_stat as a dict."""
+    # /proc/net/tls_stat exposes the per-netns TLS MIB (TLS_INC_STATS on
+    # sock_net(sk)). The test runs on a real NIC in the host namespace, so
+    # these counters are shared with anything else doing kTLS there. The
+    # strict before/after delta checks (exact rekey-outcome sums, zero error
+    # counters) assume no other kTLS activity in this namespace during a
+    # variant's window; concurrent kTLS users would perturb the deltas and
+    # cause spurious failures. Don't run other kTLS workloads alongside this
+    # test.
+    stats = defaultdict(int)
+    output = cmd("cat /proc/net/tls_stat", host=host, timeout=STATS_TIMEOUT_S)
+    for line in output.stdout.strip().split('\n'):
+        parts = line.split()
+        if len(parts) == 2:
+            stats[parts[0]] = int(parts[1])
+    return stats
+
+
+def nic_driver(cfg):
+    """DUT NIC driver name from `ethtool -i`, or None if undetermined."""
+    try:
+        output = cmd(f"ethtool -i {cfg.ifname}").stdout
+    except CmdExitFailure:
+        return None
+    for line in output.splitlines():
+        if line.startswith('driver:'):
+            return line.split(':', 1)[1].strip()
+    return None
+
+
+def read_nic_stats(cfg):
+    """Snapshot the DUT NIC's `ethtool -S` counters as a dict."""
+    # Driver per-record TLS counters from `ethtool -S` on the DUT NIC. Same
+    # before/after-delta caveat as read_tls_stats(): these are device-wide,
+    # so concurrent kTLS traffic on this NIC would perturb the deltas.
+    stats = defaultdict(int)
+    output = cmd(f"ethtool -S {cfg.ifname}").stdout
+    for line in output.strip().split('\n'):
+        key, sep, val = line.partition(':')
+        if sep and val.strip().isdigit():
+            stats[key.strip()] = int(val.strip())
+    return stats
+
+
+def stat_diff(before, after, key):
+    """Return the delta of counter `key` between two stat snapshots."""
+    return after[key] - before[key]
+
+
+def check_hw_crypto(cfg, before, after, with_tx, with_rx):
+    """DUT-side ethtool -S check: the NIC actually crypto'd records in HW.
+
+    Complements the TlsTxDevice/TlsRxDevice MIBs, which only confirm the
+    offload was installed, not that any record was processed in hardware.
+    Driver-specific; skipped (without failing) on drivers not in
+    HW_CRYPTO_COUNTERS so the test stays portable.
+    """
+    counters = HW_CRYPTO_COUNTERS.get(cfg.nic_driver)
+    if not counters:
+        ksft_pr(f"NOTE: DUT driver '{cfg.nic_driver}' has no known per-record "
+                f"HW crypto counters, skipping ethtool -S check")
+        return
+
+    for direction, active in (('Tx', with_tx), ('Rx', with_rx)):
+        if not active:
+            continue
+        key = counters[direction]
+        if key not in after:
+            ksft_pr(f"NOTE: DUT {direction}: counter '{key}' not exposed by "
+                    f"{cfg.nic_driver}, skipping")
+            continue
+        got = stat_diff(before, after, key)
+        ksft_ge(got, 1,
+                comment=f"DUT {direction}: NIC reported no HW crypto "
+                        f"({key}={got})")
+
+
+def check_path(before, after, direction, role, require_hw):
+    """On the DUT, require HW offload; on the remote, HW or SW is fine."""
+    dev = stat_diff(before, after, f'Tls{direction}Device')
+    sw = stat_diff(before, after, f'Tls{direction}Sw')
+    if require_hw:
+        ksft_ge(dev, 1,
+                comment=f"{role} {direction}: HW offload not engaged "
+                        f"(Device={dev}, Sw={sw})")
+    else:
+        ksft_ge(dev + sw, 1,
+                comment=f"{role} {direction}: no TLS activity "
+                        f"(Device={dev}, Sw={sw})")
+
+
+def verify_tls_counters(stats_before, stats_after, expected_rekeys,
+                        tls_role, is_dut, burst=False, allow_fallback=False):
+    """Verify TLS counters on one side of the connection.
+
+    tls_role: 'client' or 'server' (TLS role this side played).
+    is_dut: True for the local DUT; requires HW offload counters.
+    burst: burst mode - only the TLS client rotates its TX key; the TLS
+           server only follows with an RX rotation on KeyUpdate receipt.
+    allow_fallback: tolerate rekeys completing in SW (TlsRx/TxRekeyFallback).
+           Default False: a rekey on an up, offload-capable device must stay
+           in HW, so any fallback is a regression. Set True only where SW
+           fallback is expected (e.g. a mid-connection link-flap variant, or
+           the peer, whose offload state is not under test).
+    """
+    role = 'DUT' if is_dut else 'Peer'
+
+    def diff(key):
+        return stat_diff(stats_before, stats_after, key)
+
+    # In burst mode the TLS client only TXs and the TLS server only RXs.
+    # In echo mode both sides drive both directions.
+    with_tx = not burst or tls_role == 'client'
+    with_rx = not burst or tls_role != 'client'
+
+    if with_tx:
+        check_path(stats_before, stats_after, 'Tx', role, require_hw=is_dut)
+    if with_rx:
+        check_path(stats_before, stats_after, 'Rx', role, require_hw=is_dut)
+
+    if expected_rekeys > 0:
+        if with_tx:
+            # Each KeyUpdate yields exactly one terminal outcome, so
+            #   TlsTxRekeyOk + TlsTxRekeyAborted + TlsTxRekeyFallback == N.
+            # At most one rekey can be PENDING at socket close (single
+            # TLS_TX_REKEY_PENDING bit), so at most one lands in
+            # TlsTxRekeyAborted. TlsTxRekeyFallback is a legitimate, graceful
+            # degradation: the device did not (re)install the HW context for
+            # that rekey (device gone, dev_add rejected, or a transient
+            # crypto/alloc error) so it completed in SW while the kernel
+            # returned success. It is recoverable - the next KeyUpdate
+            # re-attempts HW offload (tls_device_start_rekey() clears
+            # TLS_TX_REKEY_FAILED). It is folded into the outcome sum below; on
+            # the DUT it must be 0 (allow_fallback=False), on the peer it is
+            # only NOTEd. A genuine rekey bug still surfaces as TlsTxRekeyError.
+            ksft_ge(1, diff('TlsTxRekeyAborted'),
+                    comment=f"{role} Tx: TlsTxRekeyAborted expected <= 1")
+            ksft_eq(diff('TlsTxRekeyOk') + diff('TlsTxRekeyAborted') +
+                    diff('TlsTxRekeyFallback'), expected_rekeys,
+                    comment=f"{role} Tx: rekey outcomes must sum to "
+                            f"{expected_rekeys}")
+            fallback = diff('TlsTxRekeyFallback')
+            if allow_fallback:
+                if fallback:
+                    ksft_pr(f"NOTE: {role} Tx: {fallback} rekey(s) completed "
+                            f"in SW (TlsTxRekeyFallback); HW not re-installed")
+            else:
+                ksft_eq(fallback, 0,
+                        comment=f"{role} Tx: TlsTxRekeyFallback expected 0 "
+                                f"(rekey must stay in HW offload)")
+            ksft_eq(diff('TlsTxRekeyError'), 0,
+                    comment=f"{role} Tx: TlsTxRekeyError expected 0")
+            ksft_eq(diff('TlsCurrTxRekey'), 0,
+                    comment=f"{role} Tx: TlsCurrTxRekey expected 0")
+        if with_rx:
+            # As on TX, each received KeyUpdate yields one terminal outcome:
+            #   TlsRxRekeyOk + TlsRxRekeyAborted + TlsRxRekeyFallback == N.
+            # At most one rekey can be deferred (single dev_add_pending) at
+            # socket close, landing in TlsRxRekeyAborted. TlsRxRekeyFallback
+            # is a recoverable, graceful degradation (dev_add failed or the
+            # device was gone, so RX temporarily dropped to SW; the next
+            # KeyUpdate re-adds the HW context and clears TLS_RX_DEV_DEGRADED).
+            # It is folded into the outcome sum below; on the DUT it must be 0
+            # (allow_fallback=False), on the peer it is only NOTEd. A genuine
+            # rekey bug still surfaces as TlsRxRekeyError.
+            ksft_ge(1, diff('TlsRxRekeyAborted'),
+                    comment=f"{role} Rx: TlsRxRekeyAborted expected <= 1")
+            ksft_eq(diff('TlsRxRekeyOk') + diff('TlsRxRekeyAborted') +
+                    diff('TlsRxRekeyFallback'), expected_rekeys,
+                    comment=f"{role} Rx: rekey outcomes must sum to "
+                            f"{expected_rekeys}")
+            ksft_eq(diff('TlsRxRekeyReceived'), expected_rekeys,
+                    comment=f"{role} Rx: TlsRxRekeyReceived expected "
+                            f"{expected_rekeys}")
+            fallback = diff('TlsRxRekeyFallback')
+            if allow_fallback:
+                if fallback:
+                    ksft_pr(f"NOTE: {role} Rx: {fallback} rekey(s) completed "
+                            f"in SW (TlsRxRekeyFallback); HW not re-installed")
+            else:
+                ksft_eq(fallback, 0,
+                        comment=f"{role} Rx: TlsRxRekeyFallback expected 0 "
+                                f"(rekey must stay in HW offload)")
+            ksft_eq(diff('TlsRxRekeyError'), 0,
+                    comment=f"{role} Rx: TlsRxRekeyError expected 0")
+            ksft_eq(diff('TlsCurrRxRekey'), 0,
+                    comment=f"{role} Rx: TlsCurrRxRekey expected 0")
+
+    ksft_eq(diff('TlsDecryptError'), 0,
+            comment=f"{role}: TlsDecryptError expected 0")
+
+
+def run_tls_test(cfg, cipher="128", tls_version="1.3", rekey=0,
+                 buffer_size=None, random_max=None, burst=False, zc=False,
+                 dut_role="client", num_iterations=None, ipver="4"):
+    """Run the TLS offload test.
+
+    dut_role: 'client' (default) - DUT runs the TLS client, remote the server.
+              'server' - swap: DUT listens, remote connects. Used for burst_rx
+              so the DUT's RX path is the one under rekey pressure.
+
+    ipver: '4' or '6' - IP version to run over. The C helper is forced to the
+           matching family with -4/-6 and connects to the peer's v4/v6 address.
+           Variants requesting '6' skip cleanly when the environment lacks IPv6
+           connectivity (require_ipver()).
+
+    The DUT (local) is the kernel under test; the remote is just a traffic
+    source/sink and may run any kernel without HW offload. Both sides run
+    kTLS because TLS is pairwise, but verify_tls_counters() requires HW
+    offload only on the DUT (is_dut=True); the peer may use SW kTLS.
+
+    Rekey/burst variants additionally require the peer to support TLS 1.3
+    KeyUpdate (as the RX or TX side of the rotation). SW KeyUpdate and its
+    MIB counters landed together in v6.14; an older peer cannot follow the
+    rotation, so those variants are skipped rather than failed when the peer
+    lacks the rekey counters (see the probe below).
+    """
+    cfg.require_ipver(ipver)
+
+    port = rand_port()
+    send_size = random_max or buffer_size
+
+    if dut_role == "client":
+        server_bin, server_host = cfg.bin_remote, cfg.remote
+        client_bin, client_host = cfg.bin_local, None
+        client_target = cfg.remote_addr_v[ipver]
+    else:
+        server_bin, server_host = cfg.bin_local, None
+        client_bin, client_host = cfg.bin_remote, cfg.remote
+        client_target = cfg.addr_v[ipver]
+
+    server_parts = [f"{server_bin} server -p {port} -c {cipher}",
+                    f"-v {tls_version}", f"-{ipver}"]
+    if burst:
+        server_parts.append("-B")
+    if zc:
+        server_parts.append("-Z")
+    if send_size:
+        server_parts.append(f"-b {send_size}")
+    server_cmd = " ".join(server_parts)
+
+    client_parts = [f"{client_bin} client -s {client_target}",
+                    f"-p {port} -c {cipher} -v {tls_version} -{ipver}"]
+    if rekey:
+        client_parts.append(f"-k {rekey}")
+    if burst:
+        client_parts.append("-B")
+    if num_iterations:
+        client_parts.append(f"-n {num_iterations}")
+    if random_max:
+        client_parts.append(f"-r {random_max}")
+    elif buffer_size:
+        client_parts.append(f"-b {buffer_size}")
+    client_cmd = " ".join(client_parts)
+
+    if burst:
+        cmd_timeout = BURST_TIMEOUT_S
+    elif rekey:
+        cmd_timeout = REKEY_TIMEOUT_S
+    else:
+        cmd_timeout = 20
+
+    stats_before_local = read_tls_stats()
+    stats_before_remote = read_tls_stats(host=cfg.remote)
+    nic_before = read_nic_stats(cfg)
+
+    # /proc/net/tls_stat lists every MIB the running kernel knows (0 or not),
+    # so a missing name means the peer predates that counter. The base rekey
+    # counters (TlsRxRekeyReceived, Tls{Rx,Tx}RekeyOk, Tls{Rx,Tx}RekeyError)
+    # shipped with SW KeyUpdate in v6.14; a peer without them cannot follow a
+    # KeyUpdate, so the rekey/burst variants can't run against it. Skip cleanly
+    # here rather than letting the peer-side rekey-sum / RxRekeyReceived checks
+    # report a confusing "expected N, got 0" later. TlsRxRekeyReceived is a
+    # reliable probe: the peer must bump it to have processed the rotation at all.
+    #
+    # Only a base v6.14 counter is probed. The newer HW-path MIBs (Aborted,
+    # Fallback, CurrRekey) are structurally 0 on a SW-only peer and defaultdict
+    # returns 0 for absent names, so the peer-side checks hold either way.
+    if rekey and 'TlsRxRekeyReceived' not in stats_before_remote:
+        raise KsftSkipEx("Peer kernel lacks TLS 1.3 KeyUpdate support "
+                         "(no rekey MIB counters); required for rekey tests")
+
+    with bkg(server_cmd, host=server_host, exit_wait=True):
+        wait_port_listen(port, host=server_host)
+        # Start the client in the background so we keep a handle to it. A
+        # foreground cmd() raises TimeoutExpired from inside its constructor
+        # if the client hangs, and since the child is not killed on timeout
+        # it would be left running with no handle to reap it. A leaked
+        # client keeps bumping the per-netns TLS counters (TlsTxRekeyAborted,
+        # TlsDecryptError, ...) and would corrupt the before/after
+        # measurement window of a later variant. The finally clause reaps it
+        # within this variant's window instead.
+        client = cmd(client_cmd, host=client_host, background=True)
+        try:
+            client.process(terminate=False, fail=True, timeout=cmd_timeout)
+        finally:
+            if client.proc.poll() is None:
+                client.process(terminate=True, fail=False, timeout=5)
+
+    stats_after_local = read_tls_stats()
+    stats_after_remote = read_tls_stats(host=cfg.remote)
+    nic_after = read_nic_stats(cfg)
+
+    peer_tls_role = 'server' if dut_role == 'client' else 'client'
+
+    # Which directions the DUT drives (mirrors verify_tls_counters()): in
+    # burst mode the TLS client only TXs and the server only RXs; echo mode
+    # drives both.
+    dut_with_tx = not burst or dut_role == 'client'
+    dut_with_rx = not burst or dut_role != 'client'
+
+    verify_tls_counters(stats_before_local, stats_after_local,
+                        rekey, dut_role, is_dut=True, burst=burst)
+    check_hw_crypto(cfg, nic_before, nic_after, dut_with_tx, dut_with_rx)
+    verify_tls_counters(stats_before_remote, stats_after_remote,
+                        rekey, peer_tls_role, is_dut=False, burst=burst,
+                        allow_fallback=True)
+
+
+# The cipher/version matrix runs over IPv4; the socket setup is the only
+# IP-version-specific code path, so a single representative variant over
+# IPv6 is enough to cover it (it skips cleanly without v6 connectivity).
+# The rekey and burst suites below likewise stay on IPv4 to bound runtime.
+@ksft_variants([
+    KsftNamedVariant("tls13_aes128", "128", "1.3", "4"),
+    KsftNamedVariant("tls13_aes256", "256", "1.3", "4"),
+    KsftNamedVariant("tls12_aes128", "128", "1.2", "4"),
+    KsftNamedVariant("tls12_aes256", "256", "1.2", "4"),
+    KsftNamedVariant("tls13_aes128_ip6", "128", "1.3", "6"),
+])
+def test_tls_offload(cfg, cipher, tls_version, ipver):
+    """Cipher/version matrix over the HW offload data path, no rekey."""
+    run_tls_test(cfg, cipher=cipher, tls_version=tls_version, ipver=ipver)
+
+
+@ksft_variants([
+    KsftNamedVariant("single", 1),
+    KsftNamedVariant("multiple", 99),
+    KsftNamedVariant("small_buf", 30, 512),
+    KsftNamedVariant("large_buf", 10, 2097152),
+    KsftNamedVariant("random_buf", 20, None, 8192),
+])
+def test_tls_offload_rekey(cfg, rekey, buffer_size=None, random_max=None):
+    """Echo-mode TLS 1.3 KeyUpdate rekeys across a range of buffer sizes."""
+    run_tls_test(cfg, cipher="128", tls_version="1.3", rekey=rekey,
+                 buffer_size=buffer_size, random_max=random_max)
+
+
+# Columns:                                          dut_role  zc     interval rekeys buffer_size
+@ksft_variants([
+    KsftNamedVariant("burst_tx_rekey_every_1",        "client", False, 1,       50,    65536),
+    KsftNamedVariant("burst_tx_rekey_every_1000",     "client", False, 1000,    3,     65536),
+    KsftNamedVariant("burst_rx_rekey_every_10",       "server", False, 10,      20,    65536),
+    KsftNamedVariant("burst_rx_rekey_every_10000",    "server", False, 10000,   1,     32768),
+    KsftNamedVariant("burst_rx_zc_rekey_every_100",   "server", True,  100,     10,    65536),
+    KsftNamedVariant("burst_rx_zc_rekey_every_20000", "server", True,  20000,   1,     16384),
+])
+def test_tls_offload_burst(cfg, dut_role, zc, interval, rekeys, buffer_size):
+    """High-volume one-directional traffic with frequent rekeys."""
+    run_tls_test(cfg, cipher="128", tls_version="1.3", rekey=rekeys,
+                 buffer_size=buffer_size, burst=True, zc=zc, dut_role=dut_role,
+                 num_iterations=interval * (rekeys + 1))
+
+
+def main() -> None:
+    """Set up the DUT/peer environment and run the offload test suites."""
+    with NetDrvEpEnv(__file__, nsim_test=False) as cfg:
+        cfg.bin_local = cfg.test_dir / "tls_hw_offload"
+        if not cfg.bin_local.exists():
+            raise KsftSkipEx(f"tls_hw_offload binary not found at {cfg.bin_local}")
+        cfg.bin_remote = cfg.remote.deploy(cfg.bin_local)
+        cfg.require_ipver("4")
+        check_tls_support(cfg)
+        cfg.nic_driver = nic_driver(cfg)
+
+        ksft_run([test_tls_offload, test_tls_offload_rekey,
+                  test_tls_offload_burst], args=(cfg, ))
+    ksft_exit()
+
+
+if __name__ == "__main__":
+    main()
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH net-next v17 15/15] tls: document TLS 1.3 hardware offload rekey handling
  2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
                   ` (13 preceding siblings ...)
  2026-09-17 22:35 ` [PATCH net-next v17 14/15] selftests: net: add TLS hardware offload test Rishikesh Jethwani
@ 2026-09-17 22:35 ` Rishikesh Jethwani
  2026-09-22  1:56   ` netdev-bot+sashiko
  14 siblings, 1 reply; 29+ messages in thread
From: Rishikesh Jethwani @ 2026-09-17 22:35 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd, davem,
	pabeni, edumazet, leon, andrew.gospodarek, Rishikesh Jethwani

Document TLS 1.3 support in the hardware offload path.

In Documentation/networking/tls-offload.rst, replace the stale note that
offload does not support TLS 1.3 with a description of the KeyUpdate
handling on both TX and RX, including how old-key records in flight are
bridged in software: each record is classified relative to the boundary at
which the NIC stopped using the old key, and partially transformed (mixed)
records are re-encrypted with the old key. In a mixed record the fragments
the NIC transformed but could not authenticate carry skb->decrypt_failed; a
non-mixed record with that flag was not transformed and is decrypted
directly.

In Documentation/networking/tls.rst, add the per-namespace rekey counters:
TlsCurrTxRekey/TlsCurrRxRekey for sessions currently in a deferred rekey,
TlsTxRekeyFallback/TlsRxRekeyFallback for rekeys that could not be
offloaded and fell back to software cryptography, and
TlsTxRekeyAborted/TlsRxRekeyAborted for deferred rekeys still pending when
the socket was destroyed.

Signed-off-by: Rishikesh Jethwani <rjethwani@purestorage.com>
---
 Documentation/networking/tls-offload.rst | 175 +++++++++++++++++++++--
 Documentation/networking/tls.rst         |  17 +++
 2 files changed, 183 insertions(+), 9 deletions(-)

diff --git a/Documentation/networking/tls-offload.rst b/Documentation/networking/tls-offload.rst
index e5802bcd4d22..cdf84f4b817a 100644
--- a/Documentation/networking/tls-offload.rst
+++ b/Documentation/networking/tls-offload.rst
@@ -99,9 +99,8 @@ at the end of kernel structures (see :c:member:`driver_state` members
 in ``include/net/tls.h``) to avoid additional allocations and pointer
 dereferences.
 
-When the offloaded connection is destroyed the core calls
-the :c:member:`tls_dev_del` callback so the driver can release per-direction
-state:
+The core calls the :c:member:`tls_dev_del` callback so the driver can release
+per-direction state:
 
 .. code-block:: c
 
@@ -109,7 +108,14 @@ state:
 			    struct tls_context *ctx,
 			    enum tls_offload_ctx_dir direction);
 
-``tls_dev_del`` is mandatory whenever ``tls_dev_add`` is provided.
+``tls_dev_del`` is called either when the offloaded connection is destroyed or,
+for a TLS 1.3 connection, when the old key is retired during a rekey (see the
+`Rekey`_ section). It operates on a single ``direction``, so the driver must
+release only the state for that direction and must not free state shared
+between directions or the socket as a whole. After a rekey ``tls_dev_del``,
+``tls_dev_add`` may be called again for the same socket and direction to
+install the new key. ``tls_dev_del`` is mandatory whenever ``tls_dev_add`` is
+provided.
 
 The third TLS device callback is :c:member:`tls_dev_resync`, called by the core
 to synchronize the TCP stream with the record boundaries:
@@ -205,7 +211,10 @@ Upon reception of a TLS offloaded packet, the driver sets
 the :c:member:`decrypted` mark in :c:type:`struct sk_buff <sk_buff>`
 corresponding to the segment. Networking stack makes sure decrypted
 and non-decrypted segments do not get coalesced (e.g. by GRO or socket layer)
-and takes care of partial decryption.
+and takes care of partial decryption. A segment the device processed but
+could not authenticate may instead carry the :c:member:`decrypt_failed`
+mark; see the `Error handling`_ section for what the mark implies about
+the payload.
 
 Resync handling
 ===============
@@ -404,8 +413,121 @@ records, then after 4 records, after 8, after 16... up until every
 Rekey
 =====
 
-Offload does not currently support TLS 1.3, therefore key rotation
-is not a concern for offloaded connections at this point.
+TLS 1.3 allows traffic keys to be updated mid-connection using the
+KeyUpdate message. Offloaded TLS 1.3 connections must therefore switch
+keys without tearing down the offload. The device cannot simply be given
+the new key because records encrypted (TX) or transformed (RX) with the
+old key may still be in flight. The stack retains the necessary old-key
+state and bridges the transition in software.
+
+TX
+--
+
+On TX, the new key is installed in a temporary software context, and
+sendmsg is routed through the software path. If no hardware-offloaded
+records remain unacknowledged, the switch completes inline during
+setsockopt. Otherwise the rekey is left pending and is completed later,
+on the sender's next ``sendmsg()`` after all old-key records have been
+ACKed (see `Completing a deferred rekey`_). Completion calls
+:c:func:`tls_dev_del` for the old key and reinstalls hardware offload
+with the new key at the current TCP write sequence. If reinstallation
+fails, the connection keeps encrypting in software with the new key; the
+next KeyUpdate re-arms the transition and retries the hardware
+installation.
+
+Unlike the software path, a ``TLS_TX`` setsockopt on an offloaded
+connection first flushes the open and partially sent hardware records to
+TCP before installing the new key. It therefore behaves like a blocking
+``send()`` of that record: it may wait for send buffer space (bounded by
+``SO_SNDTIMEO``), and on a non-blocking socket it fails with ``-EAGAIN``
+and must be retried once the socket is writable. The new key is not
+installed until the call succeeds; the connection keeps using the old key
+in the meantime.
+
+Completing a deferred rekey
+~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+A deferred rekey is completed by the sender, not by the ACK path. When
+the last old-key record is acknowledged the stack only marks the rekey
+as ready; the device is not touched. The switch itself,
+:c:func:`tls_dev_del` of the old key followed by :c:func:`tls_dev_add`
+of the new one, runs at the start of the next ``sendmsg()`` on the
+socket, and that ``sendmsg()`` is the first to be encrypted by hardware
+again. No other event completes it: ``splice_eof()``, write-space
+wakeups, retransmissions and pure ACKs all leave the connection on the
+software path.
+
+This is intentional. Completion has to flush the software context's
+open record to TCP and may sleep for send buffer space, which rules out
+the ACK and write-space paths. Beyond that, the stack only switches when
+it has new data to hand to the device: the software path is fully
+correct with the new key, so deferring the switch costs host CPU but
+nothing else, and it keeps the device from being programmed for a
+connection that may never send again.
+
+Two consequences follow. A connection that stops sending after a
+KeyUpdate stays in the deferred state until it is closed: it is
+encrypted in software with the new key, it is counted in
+``TlsCurrTxRekey``, and at close it is reported as
+``TlsTxRekeyAborted``. That counter therefore includes senders that
+simply had nothing more to send, not only sockets torn down
+mid-transition, and is not by itself an error indication. And the return
+to hardware is delayed by at least one ACK round trip after the last
+old-key record, plus however long the application waits before its next
+``sendmsg()``. A sender that wants the hardware path back promptly can
+issue a small ``sendmsg()`` once its old data has been acknowledged.
+
+Completion can fail transiently or permanently. If the software flush
+cannot get send buffer space (``-EAGAIN``, or a signal on a blocking
+socket) the rekey stays pending, the ``sendmsg()`` proceeds in software,
+and the next ``sendmsg()`` retries; the ``tls_device_complete_rekey_retry``
+tracepoint fires. A hard failure (:c:func:`tls_dev_add` rejected, or the
+netdev gone) is terminal for this KeyUpdate: the connection is pinned to
+software encryption with the new key, counted in ``TlsTxRekeyFallback``
+and moved from ``TlsCurrTxDevice`` to ``TlsCurrTxSw``; the
+``tls_device_complete_rekey_fail`` tracepoint fires. The next ``TLS_TX``
+setsockopt re-arms the transition and retries.
+
+The decision to defer is taken at the start of the ``TLS_TX``
+setsockopt, before the open hardware record is flushed to TCP. That
+flush may block for send buffer space, and old-key records acknowledged
+while it sleeps do not change the decision: the rekey is still deferred
+and completes on a following ``sendmsg()`` rather than inline. This is
+conservative, not a correctness issue. The boundary is fixed at the
+write sequence after the flush, so the acknowledgment of the flushed
+record itself arms completion; the cost is one more ACK round trip and
+one more ``sendmsg()``. Applications should not expect an inline switch
+whenever the socket has unacknowledged data at the time of the
+setsockopt.
+
+RX
+--
+
+On RX, the NIC may already have transformed in-flight records with the
+old key before the peer's KeyUpdate is parsed. When the KeyUpdate is
+decoded, the stack removes the old key from the NIC but retains the old
+AEAD, IV, and record sequence in the software offload context.
+
+Each record is classified by the TCP sequence of its first byte relative
+to the boundary at which the NIC stopped using the old key. Records
+starting after that boundary carry new-key wire encryption, so the old
+software AEAD state can be released. Records before the boundary that
+remain fully encrypted are passed to the software path. Records that
+were partially transformed by the NIC are re-encrypted with the old key
+to restore the new-key ciphertext, allowing the software AEAD to decrypt
+them with the new key.
+
+If old-key records are still queued, installation of the new key through
+:c:func:`tls_dev_add` is deferred until those records have been consumed;
+otherwise it occurs immediately. When the NIC cannot authenticate a record
+processed during the transition, the affected fragments are delivered with
+``skb->decrypt_failed`` set, following the contract described in the
+`Error handling`_ section. In a mixed record such a fragment was
+transformed (XORed) with the old key, and the re-encrypt path uses this to
+undo the transform on those fragments with the old key while leaving
+untouched fragments intact. A non-mixed record carrying
+``skb->decrypt_failed`` was not transformed; it is still wire ciphertext
+and is decrypted directly by the software AEAD under the new key.
 
 Error handling
 ==============
@@ -442,8 +564,43 @@ to the host's stack as it was on the wire (recovering original packet in the
 driver if device provides precise error is sufficient).
 
 The Linux networking stack does not provide a way of reporting per-packet
-decryption and authentication errors, packets with errors must simply not
-have the :c:member:`decrypted` mark set.
+decryption and authentication errors. A packet with errors must not have
+the :c:member:`decrypted` mark set. In addition, the driver may set the
+:c:member:`decrypt_failed` mark on a segment the device matched to an
+offloaded connection and processed but could not authenticate. The two
+marks are mutually exclusive.
+
+The stack interprets :c:member:`decrypt_failed` per record, relative to the
+:c:member:`decrypted` mark of the other segments making up the same record.
+Coalescing (GRO, socket layer) and record classification are keyed on
+:c:member:`decrypted` alone, so :c:member:`decrypt_failed` segments may be
+merged with unmarked ones. A driver setting the mark must therefore honour
+the following contract:
+
+ * In a record none of whose segments carry :c:member:`decrypted`, every
+   segment, including one with :c:member:`decrypt_failed` set, must hold
+   the payload exactly as it was on the wire. This is the general rule
+   above: if the device did not successfully decrypt any part of a record
+   it must hand the whole record over untouched. The stack passes such a
+   record to software decryption directly and does not consult
+   :c:member:`decrypt_failed`.
+
+ * In a record where some segments carry :c:member:`decrypted` (a mixed
+   record), a segment with :c:member:`decrypt_failed` set must hold payload
+   the device has already transformed (XORed with the cipher keystream) but
+   failed to authenticate, and a segment with neither mark must hold the
+   payload as it was on the wire. The stack re-encrypts the
+   :c:member:`decrypted` and :c:member:`decrypt_failed` segments to restore
+   the ciphertext, leaves the unmarked segments intact, and authenticates
+   the whole record in software.
+
+A transformed segment delivered without :c:member:`decrypt_failed`, or an
+untransformed segment of a mixed record delivered with it, is restored
+incorrectly and the record fails software authentication. A device which
+cannot tell the driver whether a failed segment was transformed must
+recover the original packet before handing it to the stack, as described
+above, and leave both marks clear. During a TLS 1.3 rekey the mark also
+tells the stack which key the device applied; see the `Rekey`_ section.
 
 A packet should also not be handled by the TLS offload if it contains
 incorrect checksums.
diff --git a/Documentation/networking/tls.rst b/Documentation/networking/tls.rst
index 980c442d7161..cf05543260d8 100644
--- a/Documentation/networking/tls.rst
+++ b/Documentation/networking/tls.rst
@@ -314,6 +314,11 @@ TLS implementation exposes the following per-namespace statistics
   number of TX and RX sessions currently installed where NIC handles
   cryptography
 
+- ``TlsCurrTxRekey``, ``TlsCurrRxRekey`` -
+  number of TX and RX sessions currently undergoing a deferred rekey,
+  i.e. a rekey which could not be applied immediately and is waiting for
+  in-flight records to drain before the new key is installed in hardware
+
 - ``TlsTxSw``, ``TlsRxSw`` -
   number of TX and RX sessions opened with host cryptography
 
@@ -344,3 +349,15 @@ TLS implementation exposes the following per-namespace statistics
 - ``TlsRxRekeyReceived`` -
   number of received KeyUpdate handshake messages, requiring userspace
   to provide a new RX key
+
+- ``TlsTxRekeyFallback``, ``TlsRxRekeyFallback`` -
+  number of rekeys on existing sessions for TX and RX which could not be
+  offloaded to the NIC and fell back to software cryptography
+
+- ``TlsTxRekeyAborted``, ``TlsRxRekeyAborted`` -
+  number of deferred rekeys for TX and RX which were still pending when
+  the socket was destroyed, and so never completed. For TX hardware
+  offload this includes senders that sent nothing further after the
+  KeyUpdate, since the switch back to hardware only happens on
+  ``sendmsg()`` (see the Rekey section of
+  Documentation/networking/tls-offload.rst)
-- 
2.50.1


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers
  2026-09-17 22:35 ` [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers Rishikesh Jethwani
@ 2026-09-22  1:55   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:55 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Claim: the two new `crypto_info->version != TLS_1_2_VERSION`
  guards in chcr_ktls_dev_add() and nfp_net_tls_add() are dead code,…
- [Low] Claim: the patch implements the same rejection with two different
  semantics — nfp_net_tls_add() returns -EOPNOTSUPP while…

--

net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers

The patch adds an early version check to the ->tls_dev_add callbacks of
the chcr_ktls and nfp drivers so that a crypto_info whose version is not
TLS_1_2_VERSION is refused before any hardware state is programmed.

A couple of questions on the two new checks below.

> diff --git a/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c b/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
> index f5acd4be1e69d..29e108ce67645 100644
> --- a/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
> +++ b/drivers/net/ethernet/chelsio/inline_crypto/ch_ktls/chcr_ktls.c
> @@ -431,6 +431,9 @@ static int chcr_ktls_dev_add(struct net_device *netdev, struct sock *sk,
>  	atomic64_inc(&port_stats->ktls_tx_connection_open);
>  	u_ctx = adap->uld[CXGB4_ULD_KTLS].handle;
>  
> +	if (crypto_info->version != TLS_1_2_VERSION)
> +		goto out;
> +

[Severity: Low]
Is this new check in chcr_ktls_dev_add() reachable as of this commit?  At
this point net/tls still filters the version before calling
->tls_dev_add:

net/tls/tls_device.c:tls_set_device_offload() {
	...
	if (crypto_info->version != TLS_1_2_VERSION) {
		rc = -EOPNOTSUPP;
		goto release_netdev;
	}
	...
}

and the same for the receive side in tls_set_device_offload_rx():

	if (ctx->crypto_recv.info.version != TLS_1_2_VERSION)
		return -EOPNOTSUPP;

The guards do become live once those two core checks are dropped later in
the series, and the ordering here looks bisect-safe, so this is only a
changelog question: could the commit message say that this is a
prerequisite for the later core change that starts accepting TLS 1.3
crypto_info, rather than describing it as preventing attempts that the
core currently already blocks?

For nfp the guard is also not covered by nfp_net_cipher_supported(), which
only looks at cipher_type, so a TLS 1.3 AES-GCM-128 crypto_info would pass
that helper once the core stops filtering versions.

[Severity: Low]
This isn't a bug, but the two halves of the patch report the same
condition differently.  nfp_net_tls_add() returns -EOPNOTSUPP while
chcr_ktls_dev_add() jumps to the shared failure label:

out:
	atomic64_inc(&port_stats->ktls_tx_connection_fail);
	return -1;

so chcr reports -EPERM for an unsupported version and also counts the
request in both ktls_tx_connection_open and ktls_tx_connection_fail.

The errno does not reach userspace on the only reachable path, since
do_tls_setsockopt_conf() treats any nonzero return from
tls_set_device_offload() on an initial install as a software fallback:

net/tls/tls_main.c:do_tls_setsockopt_conf() {
	...
	rc = tls_set_device_offload(sk);
	conf = TLS_HW;
	if (!rc) {
	...
	} else {
		rc = tls_set_sw_offload(sk, 1, update ? crypto_info : NULL);
	...
}

and the counter behaviour matches the existing reject branches in the same
function (RX direction, chcr_get_ktls_tx_info(), u_ctx->detach), which all
use the same out: label.  Would it still be preferable to return
-EOPNOTSUPP here to match nfp_net_tls_add() and the core?

>  	if (direction == TLS_OFFLOAD_CTX_DIR_RX) {
>  		pr_err("not expecting for RX direction\n");
>  		goto out;
> diff --git a/drivers/net/ethernet/netronome/nfp/crypto/tls.c b/drivers/net/ethernet/netronome/nfp/crypto/tls.c
> index 9983d7aa2b9cd..13864c6a55dce 100644
> --- a/drivers/net/ethernet/netronome/nfp/crypto/tls.c
> +++ b/drivers/net/ethernet/netronome/nfp/crypto/tls.c
> @@ -287,6 +287,9 @@ nfp_net_tls_add(struct net_device *netdev, struct sock *sk,
>  	BUILD_BUG_ON(offsetof(struct nfp_net_tls_offload_ctx, rx_end) >
>  		     TLS_DRIVER_STATE_SIZE_RX);
>  
> +	if (crypto_info->version != TLS_1_2_VERSION)
> +		return -EOPNOTSUPP;
> +
>  	if (!nfp_net_cipher_supported(nn, crypto_info->cipher_type, direction))
>  		return -EOPNOTSUPP;
>

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 02/15] net/mlx5e: add TLS 1.3 hardware offload support
  2026-09-17 22:35 ` [PATCH net-next v17 02/15] net/mlx5e: add TLS 1.3 hardware offload support Rishikesh Jethwani
@ 2026-09-22  1:55   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:55 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Concern (INVALIDATED by series context): 'net/mlx5e: add TLS 1.3
  hardware offload support' is claimed to be dead code because net/tls…
- [Low] Concern (no defect proven): fill_static_params() writes the TLS
  1.3 8-byte IV tail into the device's implicit_iv field by pointer…

--

net/mlx5e: add TLS 1.3 hardware offload support

This adds TLS 1.3 capability checks to mlx5e_is_ktls_device() and
mlx5e_ktls_type_check(), and teaches fill_static_params() to program the
TLS 1.3 static params context including the 8-byte IV tail that follows
the 4-byte salt.

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h
> index 07a04a142a2ea..0469ca6a0762e 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls.h
> @@ -30,7 +30,9 @@ static inline bool mlx5e_is_ktls_device(struct mlx5_core_dev *mdev)
>  		return false;
>  
>  	return (MLX5_CAP_TLS(mdev, tls_1_2_aes_gcm_128) ||
> -		MLX5_CAP_TLS(mdev, tls_1_2_aes_gcm_256));
> +		MLX5_CAP_TLS(mdev, tls_1_2_aes_gcm_256) ||
> +		MLX5_CAP_TLS(mdev, tls_1_3_aes_gcm_128) ||
> +		MLX5_CAP_TLS(mdev, tls_1_3_aes_gcm_256));
>  }
>  
>  static inline bool mlx5e_ktls_type_check(struct mlx5_core_dev *mdev,
> @@ -40,10 +42,14 @@ static inline bool mlx5e_ktls_type_check(struct mlx5_core_dev *mdev,
>  	case TLS_CIPHER_AES_GCM_128:
>  		if (crypto_info->version == TLS_1_2_VERSION)
>  			return MLX5_CAP_TLS(mdev,  tls_1_2_aes_gcm_128);
> +		else if (crypto_info->version == TLS_1_3_VERSION)
> +			return MLX5_CAP_TLS(mdev,  tls_1_3_aes_gcm_128);
>  		break;
>  	case TLS_CIPHER_AES_GCM_256:
>  		if (crypto_info->version == TLS_1_2_VERSION)
>  			return MLX5_CAP_TLS(mdev,  tls_1_2_aes_gcm_256);
> +		else if (crypto_info->version == TLS_1_3_VERSION)
> +			return MLX5_CAP_TLS(mdev,  tls_1_3_aes_gcm_256);
>  		break;
>  	}

[Severity: Low]
This isn't a bug, but could the commit message mention the core dependency?

At this commit the new TLS_1_3_VERSION arms in mlx5e_ktls_type_check() and
the TLS 1.3 branch in fill_static_params() cannot be reached, because
net/tls still rejects 1.3 for device offload:

net/tls/tls_device.c:tls_set_device_offload() {
	crypto_info = &ctx->crypto_send.info;
	if (crypto_info->version != TLS_1_2_VERSION) {
		rc = -EOPNOTSUPP;
		goto release_netdev;
	}
}

net/tls/tls_device.c:tls_set_device_offload_rx() {
	if (ctx->crypto_recv.info.version != TLS_1_2_VERSION)
		return -EOPNOTSUPP;
}

Those two guards are dropped by "tls: add TLS 1.3 hardware offload support"
later in the same series, so the driver code does become live by the end of
the series and there is no functional problem here.

Since the message says "Tested with TLS 1.3 AES-GCM-128 and AES-GCM-256
cipher suites", would it help to state that the test depends on the later
core patch, or to reorder so the core change lands first?

Related, mlx5e_is_ktls_device() starts advertising kTLS on hardware that
only reports tls_1_3_aes_gcm_128/256 caps while the core still offers only
1.2 sessions, which mlx5e_ktls_type_check() then refuses. Is that
intermediate state intentional for the two commits it lasts?

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c
> index 570a912dd6faf..f3f1be1d40343 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_txrx.c

[ ... ]

> @@ -15,8 +16,10 @@ enum {
>  #define EXTRACT_INFO_FIELDS do { \
>  	salt    = info->salt;    \
>  	rec_seq = info->rec_seq; \
> +	iv      = info->iv;      \
>  	salt_sz    = sizeof(info->salt);    \
>  	rec_seq_sz = sizeof(info->rec_seq); \
> +	iv_sz      = sizeof(info->iv);      \
>  } while (0)

[ ... ]

> @@ -59,7 +62,12 @@ fill_static_params(struct mlx5_wqe_tls_static_params_seg *params,
>  	memcpy(gcm_iv,      salt,    salt_sz);
>  	memcpy(initial_rn,  rec_seq, rec_seq_sz);
>  
> -	tls_version = MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_2;
> +	if (crypto_info->crypto_info.version == TLS_1_3_VERSION) {
> +		memcpy(gcm_iv + salt_sz, iv, iv_sz);
                        ^^^^^^^^^^^^^^^^

[Severity: Low]
This isn't a bug today, but would it be clearer to address implicit_iv by
name rather than walking off the end of gcm_iv?

The layout is:

include/linux/mlx5/mlx5_ifc.h:
struct mlx5_ifc_tls_static_params_bits {
	...
	u8         gcm_iv[0x20];

	u8         implicit_iv[0x40];

	u8         reserved_at_100[0x8];
	u8         dek_index[0x18];
	...
};

gcm_iv is 4 bytes, so gcm_iv + salt_sz with salt_sz == 4 and iv_sz == 8
lands exactly on the 8 bytes of implicit_iv for both AES-GCM-128 and
AES-GCM-256, which is why nothing is currently overwritten.

The write does depend on salt_sz staying 4 and on gcm_iv/implicit_iv staying
adjacent though. Would using

	implicit_iv = MLX5_ADDR_OF(tls_static_params, ctx, implicit_iv);

plus a BUILD_BUG_ON() tying sizeof(info->salt) and sizeof(info->iv) to the
field widths be preferable, so a future cipher suite with a different salt
size or a PRM layout change fails to build instead of writing into
dek_index?

> +		tls_version = MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_3;
> +	} else {
> +		tls_version = MLX5E_STATIC_PARAMS_CONTEXT_TLS_1_2;
> +	}

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 04/15] tls: add TLS 1.3 hardware offload support
  2026-09-17 22:35 ` [PATCH net-next v17 04/15] tls: add TLS 1.3 hardware offload support Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] tls: add TLS 1.3 hardware offload support drops the
  TLS_1_2_VERSION guards in tls_set_device_offload() and…
- [Low] Documentation/networking/tls-offload.rst (Rekey section, lines
  404-408 at the parent commit) states that offload does not support TLS…

--

tls: add TLS 1.3 hardware offload support

This patch enables TLS 1.3 on the kTLS device offload path by dropping the
TLS_1_2_VERSION checks in tls_set_device_offload() and
tls_set_device_offload_rx(), appending the 1.3 content_type byte with the
tag in tls_device_record_close(), and teaching the device fallback
encryption path about the 1.3 IV, AAD and prepend layouts.

> diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
> index f5e1b6b61ce38..ada66c0bd0756 100644
> --- a/net/tls/tls_device.c
> +++ b/net/tls/tls_device.c

[ ... ]

> @@ -1101,11 +1110,6 @@ int tls_set_device_offload(struct sock *sk)
>  	}
>  
>  	crypto_info = &ctx->crypto_send.info;
> -	if (crypto_info->version != TLS_1_2_VERSION) {
> -		rc = -EOPNOTSUPP;
> -		goto release_netdev;
> -	}
> -
>  	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
>  	if (!cipher_desc || !cipher_desc->offloadable) {
>  		rc = -EINVAL;
> @@ -1208,9 +1212,6 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx)
>  	struct net_device *netdev;
>  	int rc = 0;
>  
> -	if (ctx->crypto_recv.info.version != TLS_1_2_VERSION)
> -		return -EOPNOTSUPP;
> -
>  	/* A rekey (setsockopt on an already-configured socket) is not
>  	 * supported on the device offload path yet; reject it here so the
>  	 * caller can decide (propagate the error for a HW connection, or

[Severity: Medium]

At this point in the series, TLS 1.3 becomes offloadable in both directions
while the rekey rejection right below in tls_set_device_offload_rx() is
still in place.  Can a HW-offloaded TLS 1.3 RX socket get permanently
stuck if the peer sends a KeyUpdate?

The peer KeyUpdate makes tls_check_pending_rekey() in net/tls/tls_sw.c set
sw_ctx->key_update_pending, and tls_rx_rec_wait() then bails out
unconditionally:

	/* a rekey is pending, let userspace deal with it */
	if (unlikely(ctx->key_update_pending))
		return -EKEYEXPIRED;

tls_sk_poll() also masks EPOLLIN in that state.  The only place that clears
the flag is tls_finish_key_update(), reached from the SW setup path, but
the setsockopt(TLS_RX) that would install the new key is rejected in
tls_set_device_offload_rx() for rx_conf != TLS_BASE and the error is
propagated by do_tls_setsockopt_conf() without a SW fallback.  So recvmsg()
keeps returning -EKEYEXPIRED with no way for userspace to recover.

TX side, a TLS 1.3 KeyUpdate that used to work via software now turns into
a hard setsockopt failure on an offload-capable NIC.

The later patches in this series ("tls: device: add TX KeyUpdate support"
and "tls: device: add RX KeyUpdate support") implement the rekey and make
tls_set_device_offload_rx() end in tls_sw_ctx_finalize(), which clears
key_update_pending, so the end state of the series looks fine.  Would it
make sense to reorder so this patch does not leave a bisectable window
where a remote peer can wedge an offloaded 1.3 socket?

[Severity: Low]

Should the Rekey section of Documentation/networking/tls-offload.rst be
updated in this patch?  It currently says:

	Rekey
	=====

	Offload does not currently support TLS 1.3, therefore key rotation
	is not a concern for offloaded connections at this point.

which no longer holds once the version guards above are gone and 1.3
crypto_info is programmed into the NIC for both directions.  The "TLS 1.3
Key Updates" section of Documentation/networking/tls.rst also describes
installing a new key with setsockopt(TLS_TX/TLS_RX) without mentioning the
hardware offload limitation that this patch temporarily leaves behind.

The later patch "tls: document TLS 1.3 hardware offload rekey handling"
does rewrite that section with the TX and RX rekey description, so this is
only a mismatch in the intermediate state of the series.

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 06/15] tls: prep helpers and refactors for HW offload KeyUpdate
  2026-09-17 22:35 ` [PATCH net-next v17 06/15] tls: prep helpers and refactors for HW offload KeyUpdate Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] Concern (validated as NOT actionable): moving
  crypto_aead_setauthsize() into the `if (!*aead)` block of…

--

tls: prep helpers and refactors for HW offload KeyUpdate

This patch does preparatory refactoring for TX and RX HW rekey support:
hoisting cipher_context / tls_crypto_context, adding tls_tx_cipher_ctx(),
splitting tls_set_device_offload() and tls_sw_splice_eof(), and moving
crypto_aead_setauthsize() into the !*aead block of tls_sw_ctx_init().

One question below about the relocated crypto_aead_setauthsize() call.

> diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
> index 7b593dac2c31..5531303dd704 100644
> --- a/net/tls/tls_sw.c
> +++ b/net/tls/tls_sw.c

[ ... ]

> @@ -2576,6 +2592,10 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
>  
>  	key = crypto_info_key(src_crypto_info, cipher_desc);
>  
> +	/* A rekey normally reuses the existing tfm; the RX HW rekey hands over a
> +	 * NULL aead (the old one is retained for the drain), so allocate and
> +	 * configure authsize only when a fresh tfm is created here.
> +	 */
>  	if (!*aead) {
>  		*aead = crypto_alloc_aead(cipher_desc->cipher_name, 0, 0);
>  		if (IS_ERR(*aead)) {
> @@ -2583,6 +2603,10 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
>  			*aead = NULL;
>  			goto free_priv;
>  		}
> +
> +		rc = crypto_aead_setauthsize(*aead, prot->tag_size);
> +		if (rc)
> +			goto free_aead;
>  	}

[Severity: Medium]
Could the contract for this new error path be spelled out in the comment?

This adds a fallible step inside the !*aead block, and its unwind label
does:

net/tls/tls_sw.c:tls_sw_ctx_init() {
	...
free_aead:
	crypto_free_aead(*aead);
	*aead = NULL;
free_priv:
	if (!new_crypto_info) {
	...
}

So when a caller combines new_crypto_info != NULL with *aead == NULL, a
setauthsize failure leaves the live sw context with aead_recv == NULL and
free_priv skipped.

The same shape applies to the setkey branch just below, where a freshly
allocated tfm is left installed but unkeyed:

	rc = crypto_aead_setkey(*aead, key, cipher_desc->key);
	if (rc) {
		if (new_crypto_info)
			goto out;

At this commit neither case looks reachable: the only rekey caller reaches
tls_sw_ctx_init() with an already-allocated *aead, and both device helpers
still return -EOPNOTSUPP when tx_conf / rx_conf is not TLS_BASE, so
!*aead is never taken together with new_crypto_info != NULL.

The later RX rekey work in this series is the first user that hands over a
NULL aead, and it does restore context->rekey.old_aead_recv on every
tls_sw_ctx_init() failure, so the ownership rule seems intentional. Would
it be worth stating in the comment that a caller passing *aead == NULL with
new_crypto_info != NULL owns restoring the previous tfm on error, so a
future caller does not inherit a NULL or unkeyed aead_recv?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 07/15] net: sched: re-validate parked decrypted skbs on requeue
  2026-09-17 22:35 ` [PATCH net-next v17 07/15] net: sched: re-validate parked decrypted skbs on requeue Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 1 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] net/sched/sch_generic.c:dequeue_skb() — the new `if
  (skb_is_decrypted(skb)) *validate = true;` is not idempotent for kTLS…
- [Medium] net/sched/sch_generic.c:dequeue_skb() — the patch is presented
  purely as "a prerequisite for TLS 1.3 device-offload KeyUpdate…

--

net: sched: re-validate parked decrypted skbs on requeue

The patch forces *validate = true in dequeue_skb() for skbs sitting on
q->gso_skb when skb_is_decrypted(skb) is set, so the socket's current
sk_validate_xmit_skb callback is consulted again after a requeue. It is
described as a prerequisite for TLS 1.3 device-offload KeyUpdate support.

> diff --git a/net/sched/sch_generic.c b/net/sched/sch_generic.c
> index 6f6a6f0d5eb0..fc8ef0d13f5e 100644
> --- a/net/sched/sch_generic.c
> +++ b/net/sched/sch_generic.c
> @@ -285,6 +285,15 @@ static struct sk_buff *dequeue_skb(struct Qdisc *q, bool *validate,
>  		*validate = false;
>  		if (xfrm_offload(skb))
>  			*validate = true;
> +		/* A still-cleartext skb of a crypto-offloaded socket was validated
> +		 * against that socket's offload state at the time. That state
> +		 * (sk->sk_validate_xmit_skb) can change while the skb is parked here
> +		 * e.g. a TLS key update or offload teardown, so re-validate it,
> +		 * letting the current callback decide how it reaches the wire instead
> +		 * of emitting now-unencrypted plaintext.
> +		 */
> +		if (skb_is_decrypted(skb))
> +			*validate = true;

[Severity: High]
Is skb_is_decrypted() a reliable test for "still cleartext" here?  For kTLS
TX, the software fallback's output skb appears to inherit the bit from the
plaintext original:

net/tls/tls_device_fallback.c:complete_skb() {
	skb_copy_header(nskb, skb);
	...
}

and skb_copy_header() -> __copy_skb_header() copies the whole headers group:

net/core/skbuff.c:__copy_skb_header() {
	memcpy(&new->headers, &old->headers, sizeof(new->headers));
}

decrypted:1 is a member of struct_group(headers, ...) in struct sk_buff, and
nothing in complete_skb() clears it at this commit, so the ciphertext skb
also has skb->decrypted == 1.

Can that ciphertext skb reach this new gate?  sch_direct_xmit() validates
before transmit and requeues onto q->gso_skb on back-pressure:

net/sched/sch_generic.c:sch_direct_xmit() {
	if (validate)
		skb = validate_xmit_skb_list(skb, dev, &again);
	...
		dev_requeue_skb(skb, q);
}

which is why the existing code sets *validate = false with the comment "skb
in gso_skb were already validated".

If so, what stops a second tls_sw_fallback() pass over the already-encrypted
payload?  Unlike the neighbouring xfrm_offload() case, which is guarded by
XFRM_XMIT, tls_validate_xmit_skb_sw() has no "already encrypted" check:

net/tls/tls_device_fallback.c {
struct sk_buff *tls_validate_xmit_skb_sw(...)
{
	return tls_sw_fallback(sk, skb);
}
}

and fill_sg_in() re-resolves the record from the unchanged TCP sequence:

net/tls/tls_device_fallback.c:fill_sg_in() {
	u32 tcp_seq = ntohl(tcp_hdr(skb)->seq);
	record = tls_get_record(ctx, tcp_seq, rcd_sn);
	if (!record) {
		spin_unlock_irqrestore(&ctx->lock, flags);
		return -EINVAL;
	}
}

Does that mean the same record sequence number, and therefore the same AEAD
nonce and keystream, is used a second time with the previous ciphertext as
input?  For the CTR-based AEADs used by the offload this would put
C1 XOR KS, i.e. the plaintext, on the wire, which is the outcome the
changelog says it is preventing.  And when the record has already been
freed, fill_sg_in() returns -EINVAL, tls_sw_fallback() does kfree_skb() and
returns NULL, so a mid-stream segment that has already been removed from the
qdisc accounting is dropped silently.

For reference, a later commit in the same series, 3f66fe17313fe ("tls:
device: add TX KeyUpdate support"), adds nskb->decrypted = 0 in
complete_skb() with a comment describing this exact failure mode.  Could
that clearing be moved into this patch, or ahead of it, so the tree is not
left in this state for several commits?

[Severity: Medium]
Should this carry a Fixes: tag?  The changelog presents the change purely as
"a prerequisite for TLS 1.3 device-offload KeyUpdate support", but it also
changes behaviour for offload users already in the tree, since before this
patch sk->sk_validate_xmit_skb was never consulted again for anything parked
on q->gso_skb.

Two in-tree writers can change that state while the skb is parked:

net/tls/tls_device.c:tls_device_down() {
		WRITE_ONCE(ctx->sk->sk_validate_xmit_skb, tls_validate_xmit_skb_sw);
}

which also clears ctx->netdev, and:

net/psp/psp_sock.c:psp_validate_xmit() {
	pas = psp_skb_get_assoc_rcu(skb);
	good = !pas || rcu_access_pointer(dev->psp_dev) == pas->psd;
}

where the dev->psp_dev relation can change under psp device teardown.  If
either path is reachable today, this is a confidentiality fix for current
users and backporters get no signal from the changelog.  The transition in
tls_device_down() came in c55dcdd435aa6, which would be the Fixes: candidate.

The other direction is also unmentioned: forcing *validate = true on requeue
introduces new drop paths for skbs already accounted out of the qdisc.
psp_validate_xmit() can now consume a requeued skb with
SKB_DROP_REASON_PSP_OUTPUT, and an orphaned socket hits:

net/core/dev.c:sk_validate_xmit_skb() {
	} else if (unlikely(skb_is_decrypted(skb))) {
		pr_warn_ratelimited("unencrypted skb with no associated socket - dropping\n");
		kfree_skb(skb);
		skb = NULL;
	}
}

Could the changelog mention that requeued decrypted skbs can now be dropped
rather than transmitted?  Alternatively, if none of the above paths turns out
to be reachable in-tree, the "emitting now-unencrypted plaintext" wording in
both the changelog and the new in-code comment overstates the present risk
and could be reworded.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 08/15] tcp: fence collapse against rtx-queue tail when write queue is empty
  2026-09-17 22:35 ` [PATCH net-next v17 08/15] tcp: fence collapse against rtx-queue tail when write queue is empty Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Changelog claim about TLS 1.3 device-offload KeyUpdate: the
  concern states that the final paragraph ("TLS 1.3 device-offload…
- [Low] Stale call-site comment: net/tls/tls_device.c:169-173
  (tls_device_commit_start_marker) states "So mark the last skb in the…

--

tcp: fence collapse against rtx-queue tail when write queue is empty

This makes tcp_write_collapse_fence() fall back to the retransmit queue
tail when the write queue is empty, so the end-of-record boundary is still
marked at the point encryption keys are installed.  The stated goal is to
stop tcp_retrans_try_collapse() and tcp_shift_skb_data() from merging a
post-key skb into a pre-key skb.

No Link: tags are present in the commit header.

Two small notes, both about wording rather than behaviour.

> tcp: fence collapse against rtx-queue tail when write queue is empty

[Severity: Low]
The changelog ends with:

    "TLS 1.3 device-offload KeyUpdate additionally relies on this behavior
     to keep old-key and new-key records in distinct skbs for
     re-encryption on RX."

At this commit there is no device-offload KeyUpdate path that can reach
the helper, since tls_set_device_offload() and tls_set_device_offload_rx()
in net/tls/tls_device.c still reject a rekey on an already-configured
socket:

net/tls/tls_device.c:tls_set_device_offload() {
	/* A rekey (setsockopt on an already-configured socket) is not
	 * supported on the device offload path yet; ...
	 */
	if (ctx->tx_conf != TLS_BASE)
		return -EOPNOTSUPP;

The KeyUpdate consumer that does call tcp_write_collapse_fence() arrives
in the later patches of the series.  Would it be worth phrasing that
paragraph as a forward reference to the upcoming KeyUpdate patches, so
the dependency ordering is clear to someone reading this commit alone?

> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index 5e5f5f9b89a3..8c6d90e962c4 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -2340,6 +2340,15 @@ static inline void tcp_write_collapse_fence(struct sock *sk)
>  {
>  	struct sk_buff *skb = tcp_write_queue_tail(sk);
>  
> +	/* When nothing is queued for transmit, the last skb of the current
> +	 * state is the rtx queue tail (its end_seq == snd_nxt == write_seq).
> +	 * Fence that instead, otherwise the boundary is left unmarked and a
> +	 * later tcp_retrans_try_collapse()/tcp_shift_skb_data() can merge it
> +	 * with the first skb of the next state across the fence (they only test
> +	 * the tail's EOR, not skb->decrypted).
> +	 */
> +	if (!skb)
> +		skb = tcp_rtx_queue_tail(sk);

[Severity: Low]
With this fallback the helper can now set eor on an already-transmitted,
unacked skb from the retransmit queue, not just on a write-queue skb.

The comment at the tls_device.c call site still describes the old
behaviour:

net/tls/tls_device.c:tls_device_commit_start_marker() {
	/* TLS offload is greatly simplified if we don't send
	 * SKBs where only part of the payload needs to be encrypted.
	 * So mark the last skb in the write queue as end of record.
	 */
	tcp_write_collapse_fence(sk);
}

That text is unchanged through the end of the series.  Since the commit
message describes the empty-write-queue case as the common state at key
install time, the fallback is the path most readers will hit.

Should that comment be updated to mention the retransmit-queue tail as
well?

>  	if (skb)
>  		TCP_SKB_CB(skb)->eor = 1;
>  }

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 09/15] net: skbuff: add skb->decrypt_failed bit
  2026-09-17 22:35 ` [PATCH net-next v17 09/15] net: skbuff: add skb->decrypt_failed bit Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The patch makes `decrypt_failed` a per-SKB (per-segment)
  crypto-state bit that is propagated by `skb_copy_decrypted()`…

--

net: skbuff: add skb->decrypt_failed bit

This adds a new skb->decrypt_failed bit under CONFIG_SKB_DECRYPTED, set
when the NIC could not authenticate an skb's TLS payload, and propagates
it in skb_copy_decrypted() next to skb->decrypted.  It is described as a
prerequisite for TLS 1.3 device-offload RX KeyUpdate support.

> diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
> index 421f6fc45451..5da2c1149d98 100644
> --- a/include/linux/skbuff.h
> +++ b/include/linux/skbuff.h
> @@ -851,6 +851,10 @@ enum skb_tstamp_type {
>   *		unreadable.
>   *	@dst_pending_confirm: need to confirm neighbour
>   *	@decrypted: Decrypted SKB
> + *	@decrypt_failed: hardware could not authenticate this skb's TLS payload.
> + *		The payload may have been transformed (XORed) or left as wire
> + *		ciphertext, so software must re-authenticate the record and undo the
> + *		transform on any XORed fragment before it can be decrypted
>   *	@slow_gro: state present at GRO time, slower prepare step required
>   *	@tstamp_type: When set, skb->tstamp has the
>   *		delivery_time clock base of skb->tstamp.
> @@ -1025,6 +1029,7 @@ struct sk_buff {
>  #endif
>  #ifdef CONFIG_SKB_DECRYPTED
>  	__u8			decrypted:1;
> +	__u8			decrypt_failed:1;
>  #endif
>  	__u8			slow_gro:1;
>  #if IS_ENABLED(CONFIG_IP_SCTP)
> @@ -1716,6 +1721,7 @@ static inline void skb_copy_decrypted(struct sk_buff *to,
>  {
>  #ifdef CONFIG_SKB_DECRYPTED
>  	to->decrypted = from->decrypted;
> +	to->decrypt_failed = from->decrypt_failed;
>  #endif
>  }

[Severity: Medium]
Should skb_cmp_decrypted() be extended with the new bit as well?  After
this patch the copy side carries two bits while the compare side still
only looks at ->decrypted:

include/linux/skbuff.h:
static inline int skb_cmp_decrypted(const struct sk_buff *skb1,
				    const struct sk_buff *skb2)
{
#ifdef CONFIG_SKB_DECRYPTED
	return skb2->decrypted - skb1->decrypted;
#else
	return 0;
#endif
}

skb_cmp_decrypted() is the only merge-compatibility gate the receive side
has:

net/ipv4/tcp_offload.c:tcp_gro_receive()
	flush |= skb_cmp_decrypted(p, skb);

include/net/tcp.h:tcp_skb_can_collapse_rx()  (tcp_try_coalesce(),
tcp_collapse())
	return likely(mptcp_skb_can_collapse(to, from) &&
		      !skb_cmp_decrypted(to, from));

net/tls/tls_strp.c:tls_strp_copyin()
	strp->mixed_decrypted |= !!skb_cmp_decrypted(skb, in_skb);

net/tls/tls_strp.c:tls_strp_check_queue_ok()
	if (skb_cmp_decrypted(first, skb))
		return false;

net/core/skbuff.c:skb_shift()
	DEBUG_NET_WARN_ON_ONCE(skb_cmp_decrypted(tgt, skb));

A decrypt_failed segment and an untouched segment both have
decrypted == 0, so all of the above treat them as compatible.  Can two
such skbs then be coalesced into a single skb that carries only the
head's decrypt_failed value, and can the same mismatch leave
strp->mixed_decrypted clear so tls_strp_copyin() flattens the bytes?

The consumer added later in the series works at fragment granularity:

net/tls/tls_device.c:tls_device_reencrypt()
	if (skb_iter->decrypted || skb_iter->decrypt_failed) {
		err = skb_store_bits(skb_iter, frag_pos, buf, copy);

If a merged skb reports decrypt_failed for a byte range that was never
XORed, or loses the mark for a range that was, does this restore the
wrong bytes and make the record fail software authentication with
-EBADMSG?

tls_device_decrypted() reads the mark from the message head only:

	bool nic_touched = !is_encrypted || skb->decrypt_failed;

If the head is the unmarked segment of such a merge, can nic_touched come
out false for a record the NIC did transform, retiring
rekey.old_aead_recv early so a later XORed record is reconstructed with
the wrong key?

The documentation patch in this series states the asymmetry is intended:

	Coalescing (GRO, socket layer) and record classification are keyed on
	:c:member:`decrypted` alone, so :c:member:`decrypt_failed` segments may be
	merged with unmarked ones.

but the same contract then requires that, within a mixed record, a
decrypt_failed segment holds XOR-transformed payload while a segment with
neither mark holds the payload as it was on the wire.  How is that
per-segment property preserved once the merge described above happens?

For what it is worth, this does not look like an authentication bypass:
mlx5e sets decrypt_failed without also setting decrypted, and an skb with
decrypted == 0 is still never merged into a decrypted == 1 skb, so the
record is always authenticated in software.  The failure mode appears to
be a rejected record and a torn-down kTLS stream instead.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 10/15] net/mlx5e: flag TLS RX records that failed device decryption
  2026-09-17 22:35 ` [PATCH net-next v17 10/15] net/mlx5e: flag TLS RX records that failed device decryption Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 0 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] mlx5e_ktls_handle_rx_skb()
  (drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c:614) sets…

--

net/mlx5e: flag TLS RX records that failed device decryption

The patch splits CQE_TLS_OFFLOAD_ERROR out of the default case in
mlx5e_ktls_handle_rx_skb() and sets the new skb->decrypt_failed bit while
leaving skb->decrypted clear, then falls through to the existing tls_err
accounting. The bit is intended to be consumed by the TLS 1.3 device-offload
RX KeyUpdate support added later in the series.

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c
> index bca45679e201..8ec40f5fd5b5 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c
> @@ -602,7 +602,18 @@ void mlx5e_ktls_handle_rx_skb(struct mlx5e_rq *rq, struct sk_buff *skb,
>  		stats->tls_resync_req_pkt++;
>  		resync_update_sn(rq, skb);
>  		break;
> -	default: /* CQE_TLS_OFFLOAD_ERROR: */
> +	case CQE_TLS_OFFLOAD_ERROR:
> +		/* The device could not authenticate the payload. Depending on
> +		 * where the failure occurred the bytes may have been transformed
> +		 * (XORed) or left as wire ciphertext. Flag it so that, during a
> +		 * TLS 1.3 rekey transition, the re-encrypt path undoes the
> +		 * transform on any XORed frag of a mixed record while software
> +		 * re-authenticates; a non-mixed record stays wire ciphertext and
> +		 * is decrypted directly.
> +		 */
> +		skb->decrypt_failed = 1;

[Severity: High]

Can this per-skb distinction survive coalescing?  The bit is set here with
skb->decrypted left at 0, but the only comparator the stack uses to decide
whether two skbs may be merged still looks at decrypted alone:

include/linux/skbuff.h:skb_cmp_decrypted() {
	...
	return skb2->decrypted - skb1->decrypted;
	...
}

So a segment flagged here and a segment the device never touched compare
equal.  That comparator gates the merge in tcp_gro_receive():

	flush |= skb_cmp_decrypted(p, skb);

and in tcp_skb_can_collapse_rx(), used by tcp_try_coalesce(),
tcp_collapse() and tcp_add_backlog():

	return likely(mptcp_skb_can_collapse(to, from) &&
		      !skb_cmp_decrypted(to, from));

After such a merge only one flag value survives for bytes contributed by
several skbs, since tcp_collapse() copies it from the first skb only:

		skb_copy_decrypted(nskb, skb);

Note that the prerequisite patch extended skb_copy_decrypted() to propagate
decrypt_failed but left skb_cmp_decrypted() unchanged.  Is that asymmetry
intentional?

The same comparator feeds the mixed-record notion this commit message relies
on:

net/tls/tls_strp.c:tls_strp_copyin() {
	...
	strp->mixed_decrypted |= !!skb_cmp_decrypted(skb, in_skb);
	...
}

If two segments of one record differ only in decrypt_failed, the record is
never classified as mixed.  Does the consumer added later in the series then
act on the wrong value?  It restores a whole frag based on the single bit:

net/tls/tls_device.c:tls_device_reencrypt() {
	...
		if (skb_iter->decrypted || skb_iter->decrypt_failed) {
			err = skb_store_bits(skb_iter, frag_pos, buf, copy);
	...
}

For a merged skb covering both XORed bytes and untouched wire ciphertext,
does this either undo the transform on bytes the device never touched, or
leave the transform in place on bytes it did XOR, so that software AEAD
authentication fails with -EBADMSG?

A second consequence is in the rekey boundary tracking:

net/tls/tls_device.c:tls_device_decrypted() {
	...
		bool nic_touched = !is_encrypted || skb->decrypt_failed;
	...
}

If a record the device never processed inherits decrypt_failed from a
GRO- or collapse-merged neighbour, does ctx->rekey.old_nic_boundary slide
past that record, keeping the retired old AEAD alive and postponing the
deferred tls_dev_add()?

The series documentation appears to describe exactly this mismatch while at
the same time permitting the merge:

Documentation/networking/tls-offload.rst
	Coalescing (GRO, socket layer) and record classification are keyed on
	decrypted alone, so decrypt_failed segments may be merged with unmarked
	ones. ... A transformed segment delivered without decrypt_failed, or an
	untransformed segment of a mixed record delivered with it, is restored
	incorrectly and the record fails software authentication.

Since GRO and TCP coalescing run after mlx5e_ktls_handle_rx_skb() has set
the bit, can the driver honour that contract at all?  Would extending
skb_cmp_decrypted() to also compare decrypt_failed (and thus forbidding
coalescing across the new bit) be needed before this flag can be relied on?

The segmentation and arrival pattern is driven by peer traffic and a
peer-initiated KeyUpdate, so the merge condition is reachable from a remote
endpoint.

For reference, skb_cmp_decrypted() is unchanged at the end of the series, so
the gap does not appear to be closed by a later patch.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 11/15] tls: device: add TX KeyUpdate support
  2026-09-17 22:35 ` [PATCH net-next v17 11/15] tls: device: add TX KeyUpdate support Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 10 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 6 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] tls_validate_xmit_skb_rekey() ignores its `dev` argument, unlike
  its peer tls_validate_xmit_skb(), which passes the skb through when…
- [Medium] tls_tx_drop_acked_clone() (and the FAILED-state pass-through
  test in tls_validate_xmit_skb_rekey()) compute the skb's end of…
- [Medium] tls_device_sendmsg() evaluates TLS_TX_REKEY_READY once, before
  entering tls_sw_sendmsg_locked(), and the SW send loop never re-checks…
- [Medium] The deferred-rekey completion trigger can be lost, leaving the
  socket pinned to the temporary SW context.
- [Medium] The nskb->decrypted = 0 clearing added to complete_skb()
  repairs a data-corruption window that the earlier series commit…
- [Medium] The commit message describes only the rekey state machine and
  the tx_lock/getsockopt side effects, but the patch also changes the…
- [Medium] tls_sw_ctx_init() is changed so that ctx->push_pending_record
  is only assigned when tx != 0.
- [Low] The commit message states that a second KeyUpdate arriving while
  one is pending "re-keys the SW AEAD in place", while the code…
- [Low] Assessed as RESOLVED LATER IN THE SERIES (retained for policy
  combination).
- [Low] Assessed as RESOLVED LATER IN THE SERIES (retained for policy
  combination).

--

tls: device: add TX KeyUpdate support

This adds TX KeyUpdate support for the device-offload path: a temporary SW
context encrypts with the new key while old-key HW records drain, and HW
offload is reinstalled at the current write_seq once everything is ACKed.
A few shared fallback paths (complete_skb(), tls_sw_fallback(),
tls_validate_xmit_skb()) are touched along the way.

> diff --git a/net/tls/tls.h b/net/tls/tls.h
> index 920a926e8e68..e749f429301a 100644

[ ... ]

> diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
> index 972c9c7ba7de..f32c1bb6b497 100644
> --- a/net/tls/tls_device.c
> +++ b/net/tls/tls_device.c

[ ... ]

> @@ -607,18 +701,46 @@ static int tls_push_data(struct sock *sk,
>  	return rc;
>  }
>  
> +/* True while TX is routed through the temporary SW rekey context: a rekey is in
> + * progress (PENDING) or has failed and the socket stays pinned to SW (FAILED).
> + */
> +static bool tls_device_tx_uses_sw(const struct tls_context *ctx)
> +{
> +	return test_bit(TLS_TX_REKEY_PENDING, &ctx->flags) ||
> +	       test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
> +}
> +
>  int tls_device_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
>  {
>  	unsigned char record_type = TLS_RECORD_TYPE_DATA;
>  	struct tls_context *tls_ctx = tls_get_ctx(sk);
>  	int rc;
>  
> +	/* Reject unsupported flags up front. tls_push_data() enforces the same
> +	 * set, but during a rekey the send is routed to tls_sw_sendmsg_locked(),
> +	 * which is the _locked variant and does not re-check; without this,
> +	 * MSG_ZEROCOPY / MSG_OOB etc. would reach tcp_sendmsg_locked() on the
> +	 * kernel-owned record pages while PENDING/FAILED.
> +	 */
> +	if (msg->msg_flags & ~(MSG_MORE | MSG_DONTWAIT | MSG_NOSIGNAL |
> +			       MSG_SPLICE_PAGES | MSG_EOR))
> +		return -EOPNOTSUPP;
> +
>  	if (!tls_ctx->zerocopy_sendfile)
>  		msg->msg_flags &= ~MSG_SPLICE_PAGES;
>  
>  	mutex_lock(&tls_ctx->tx_lock);
>  	lock_sock(sk);
>  
> +	/* Old-key records all ACKed; switch back to HW. */
> +	if (test_bit(TLS_TX_REKEY_READY, &tls_ctx->flags))
> +		tls_device_complete_rekey(sk, tls_ctx, true, msg->msg_flags);
> +
> +	if (tls_device_tx_uses_sw(tls_ctx)) {
> +		rc = tls_sw_sendmsg_locked(sk, msg, size);
> +		goto out;
> +	}

[Severity: Medium]
Can a single very large send outrun the frozen boundary_seq here?

TLS_TX_REKEY_READY is evaluated once, before entering
tls_sw_sendmsg_locked(), and the SW loop never re-checks it, so
rekey.boundary_seq stays fixed for the whole call while the TCP sequence
advances.  The rekey provably cannot complete during the call either, since
tx_lock is held across the send and both completion sites require it.

Once the wire sequence advances more than 2^31 past boundary_seq, the signed
comparison in tls_validate_xmit_skb_rekey():

	if (!before(tcp_seq, pivot_seq))
		return skb;

	return tls_sw_fallback(sk, skb);

flips, and freshly produced SW ciphertext is treated as pre-boundary old-key
data.  In tls_sw_fallback() the old record can still be matched via
retransmit_hint, fill_sg_in() then returns non-zero with sync_size == 0, and:

	if (sync_size < 0 && payload_len <= -sync_size &&
	    !skb_is_decrypted(skb))
		nskb = skb_get(skb);
	goto put_sg;

leaves nskb NULL, so the skb is freed rather than transmitted.  Retransmits
are misclassified the same way, so the connection would stall.  A single
writev approaching MAX_RW_COUNT (~2 GiB) plus ~22 bytes of framing per 16KB
record is enough to exceed 2^31 of wire sequence space.  Would re-checking
READY inside the SW send loop, or refreshing the pivot, avoid this?

[ ... ]

> @@ -1106,6 +1246,425 @@ static struct tls_offload_context_tx *alloc_offload_ctx_tx(struct tls_context *c
>  	return offload_ctx;
>  }
>  

[ ... ]

> +static int tls_device_start_rekey(struct sock *sk,
> +				  struct tls_context *ctx,
> +				  struct tls_offload_context_tx *offload_ctx,
> +				  struct tls_crypto_info *new_crypto_info)
> +{

[ ... ]

> +		old_aead = offload_ctx->rekey.sw.aead_send;
> +		offload_ctx->rekey.sw.aead_send = new_aead;
> +		crypto_free_aead(old_aead);

[Severity: Low]
This isn't a bug, but the commit message says:

  "A KeyUpdate arriving while one is pending re-keys the SW AEAD in place"

The code does the opposite: tls_device_build_rekey_aead() allocates a fresh
transform, and the comment on that helper explains why in-place re-keying is
avoided:

  /* ... Re-keying a live tfm in place is not atomic:
   * a failed crypto_aead_setkey() leaves it with CRYPTO_TFM_NEED_KEY set,
   * destroying the previous key.

Then the swap above installs the new transform only on success and frees the
old one.  Could the changelog wording be adjusted, since "in place" is the
term the code's own comment uses for the approach it rejects?

> +
> +		tls_device_copy_rekey_iv_seq(offload_ctx, cipher_desc,
> +					     salt, iv, rec_seq);
> +
> +		if (rekey_failed) {

[ ... ]

> +			down_read(&device_offload_lock);
> +			spin_lock_irqsave(&offload_ctx->lock, flags);
> +			WRITE_ONCE(ctx->rekey.boundary_seq, tcp_sk(sk)->snd_una);
> +			set_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
> +			spin_unlock_irqrestore(&offload_ctx->lock, flags);

[Severity: Medium]
Can the deferred completion trigger be lost here, pinning the socket to the
temporary SW context?

Two cases look reachable.

First, the FAILED to PENDING re-arm above publishes
boundary_seq = tcp_sk(sk)->snd_una, a value that is already satisfied at the
moment it is written.  The only place that sets READY is
tls_tcp_clean_acked():

	if (test_bit(TLS_TX_REKEY_PENDING, &tls_ctx->flags) &&
	    !test_bit(TLS_TX_REKEY_FAILED, &tls_ctx->flags)) {
		u32 boundary_seq = READ_ONCE(tls_ctx->rekey.boundary_seq);

		if (!before(acked_seq, boundary_seq))
			set_bit(TLS_TX_REKEY_READY, &tls_ctx->flags);
	}

and that only runs when snd_una advances, so READY requires a *later* ACK
rather than the already-reached boundary.

Second, on the initial arm, tls_set_device_offload_rekey() decides defer from
tls_has_unacked_records()/tls_is_pending_open_record() before
tls_device_start_rekey() publishes boundary_seq and PENDING, and the arming
path sleeps in between (record flush, then crypto_alloc_aead() inside
tls_device_init_rekey_sw()).  If the ACK covering the boundary is processed in
that window, tls_tcp_clean_acked() skips the new block because PENDING is not
set yet, and nothing re-evaluates snd_una against boundary_seq afterwards.

The result is a socket that keeps encrypting in SW with the new key,
TlsCurrTxRekey staying elevated, and eventually TlsTxRekeyAborted.  Would a
post-arm check of snd_una versus boundary_seq close both cases?

The last patch in the series ("tls: document TLS 1.3 hardware offload rekey
handling") justifies this race by saying the acknowledgment of the flushed
record itself arms completion, but in the sub-case above there is no flushed
record still in flight when the arming ACK already arrived.

> +			/* Release pairs with test_bit_acquire() in the validator:
> +			 * a TX seeing FAILED clear must see the fresh boundary_seq.
> +			 */
> +			clear_bit_unlock(TLS_TX_REKEY_FAILED, &ctx->flags);
> +			up_read(&device_offload_lock);

[ ... ]

> +static int tls_set_device_offload_rekey(struct sock *sk,
> +					struct tls_context *ctx,
> +					struct tls_crypto_info *new_crypto_info)
> +{
> +	struct tls_offload_context_tx *offload_ctx = tls_offload_ctx_tx(ctx);
> +	bool rekey_pending = test_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
> +	bool rekey_failed = test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
> +	bool defer = true;
> +	int rc;
> +
> +	/* Defer the switch back to HW until any in-flight old-key records are
> +	 * ACKed. A partially_sent_record needs no separate check: its record is
> +	 * on records_list before it is sent (tls_push_record()) and stays there
> +	 * until ACKed, so tls_has_unacked_records() already covers it.
> +	 */
> +	if (!rekey_pending && !rekey_failed)
> +		defer = tls_has_unacked_records(offload_ctx) ||
> +			tls_is_pending_open_record(ctx);

[ ... ]

> diff --git a/net/tls/tls_device_fallback.c b/net/tls/tls_device_fallback.c
> index 1110f7ac6bcb..f2a0ae827bb2 100644
> --- a/net/tls/tls_device_fallback.c
> +++ b/net/tls/tls_device_fallback.c
> @@ -190,6 +190,14 @@ static void complete_skb(struct sk_buff *nskb, struct sk_buff *skb, int headln)
>  
>  	skb_copy_header(nskb, skb);
>  
> +	/* nskb now carries ciphertext, but skb_copy_header() inherited
> +	 * skb->decrypted from the plaintext original. Clear it so the bit keeps
> +	 * meaning "still-plaintext, needs an encryptor": otherwise a requeued
> +	 * nskb would be needlessly re-validated (and re-encrypted) and would trip
> +	 * the NIC's decrypted-vs-start-marker WARN.
> +	 */
> +	nskb->decrypted = 0;

[Severity: Medium]
Should this clearing live in an earlier commit of the series?

The earlier series commit d209263206bf ("net: sched: re-validate parked
decrypted skbs on requeue") makes dequeue_skb() set validate for any requeued
skb with skb_is_decrypted():

	if (skb_is_decrypted(skb))
		*validate = true;

Until this commit, complete_skb() copied the plaintext original's headers into
the ciphertext nskb (skb_copy_header() -> __copy_skb_header() copies the
headers struct_group, which contains the decrypted bit) and set nskb->sk.  A
requeued fallback output would therefore be re-validated,
tls_validate_xmit_skb() would again see dev != ctx->netdev (the condition that
caused the first fallback) and tls_sw_fallback() would encrypt the ciphertext a
second time.

Because the second pass reuses the same record, rcd_sn, IV and sync_size, the
CTR keystream is applied twice, so the wire payload becomes the original
plaintext with a bogus tag.

Neither changelog mentions the dependency, so commits d209263206bf..3f66fe17313f
carry that window.  Could this hunk move into (or before) the sched commit, or
could both changelogs state the ordering requirement?

>  	skb_put(nskb, skb->len);
>  	memcpy(nskb->data, skb->data, headln);
>  
> @@ -396,8 +404,17 @@ static struct sk_buff *tls_sw_fallback(struct sock *sk, struct sk_buff *skb)
>  	sg_init_table(sg_out, ARRAY_SIZE(sg_out));
>  
>  	if (fill_sg_in(sg_in, skb, ctx, &rcd_sn, &sync_size, &resync_sgs)) {
> -		/* bypass packets before kernel TLS socket option was set */
> -		if (sync_size < 0 && payload_len <= -sync_size)
> +		/* Below the record range (start marker / already-freed record).
> +		 * Pass through only cleartext that was never offload-encrypted
> +		 * (skb->decrypted == 0): genuine pre-TLS bytes sent before the
> +		 * socket option was set, or SW-encrypted rekey ciphertext. A
> +		 * decrypted=1 skb here is offload-record plaintext whose record was
> +		 * purged (e.g. a rekey installed a new start marker above its seq);
> +		 * it must never reach the wire in the clear, so continue on and
> +		 * drop it (nskb stays NULL).
> +		 */
> +		if (sync_size < 0 && payload_len <= -sync_size &&
> +		    !skb_is_decrypted(skb))
>  			nskb = skb_get(skb);
>  		goto put_sg;
>  	}

[Severity: Medium]
Could the commit message cover the shared fallback-path changes too?

The changelog describes the rekey state machine, the tx_lock ordering and the
getsockopt accessor, but this patch also changes behaviour for every
device-offloaded socket that ever falls back (route change, bond member,
post-NETDEV_DOWN, funeth/nfp tls_encrypt_skb()):

 - complete_skb() now clears nskb->decrypted on all SW-fallback encryptions.
 - the hunk above turns a pass-through into a drop, and the comment describes
   it as preventing offload-record plaintext from reaching the wire in the
   clear.
 - tls_validate_xmit_skb() gains tls_tx_drop_acked_clone(), a per-skb
   header/sequence computation plus a silent kfree_skb() rule that stays armed
   for the socket's life once TLS_TX_REKEY_FLOOR is set.

The skb_is_decrypted() gate in particular looks like a standalone fix that a
stable backporter would want to find on its own.  Could it be split out, or at
least described in the changelog?

>  
> @@ -416,11 +433,57 @@ static struct sk_buff *tls_sw_fallback(struct sock *sk, struct sk_buff *skb)
>  	return nskb;
>  }
>  
> +/* Post-rekey drop floor. Once a rekey has completed (TLS_TX_REKEY_FLOOR set), a

[ ... ]

> +static bool tls_tx_drop_acked_clone(struct sock *sk, struct sk_buff *skb)
> +{
> +	int payload_len = skb->len - skb_tcp_all_headers(skb);
> +	u32 end_seq;
> +
> +	if (likely(!test_bit(TLS_TX_REKEY_FLOOR, &tls_get_ctx(sk)->flags)))
> +		return false;
> +
> +	if (payload_len <= 0)
> +		return false;
> +

[ ... ]

> +	end_seq = ntohl(tcp_hdr(skb)->seq) + payload_len;
> +	return !after(end_seq, READ_ONCE(tcp_sk(sk)->snd_una));
> +}

[Severity: Medium]
Does end_seq need to account for the FIN's sequence-space byte?

A FIN consumes one sequence number, so for a data+FIN skb whose payload is
fully ACKed but whose FIN is not (snd_una == seq + payload_len), the test
above is true and tls_validate_xmit_skb() frees the whole skb:

	if (unlikely(tls_tx_drop_acked_clone(sk, skb))) {
		kfree_skb(skb);
		return NULL;
	}

tcp_send_fin() can attach the FIN to an already-transmitted skb:

net/ipv4/tcp_output.c:tcp_send_fin() {
	...
	tskb = tail;
	if (!tskb && tcp_under_memory_pressure(sk))
		tskb = skb_rb_last(&sk->tcp_rtx_queue);
	...
}

so a clone parked in a qdisc can carry data plus a live FIN while the peer has
ACKed only the data.  Shutdown would then wait for the RTO.  The FAILED-state
pass-through in tls_validate_xmit_skb_rekey() makes the same payload-only
comparison:

		if (payload_len > 0 && !skb_is_decrypted(skb) &&
		    after(tcp_seq + payload_len, snd_una))
			return skb;

Would adding the FIN/SYN bit to end_seq (or skipping skbs that carry flags)
be appropriate here?

> +
>  struct sk_buff *tls_validate_xmit_skb(struct sock *sk,
>  				      struct net_device *dev,
>  				      struct sk_buff *skb)
>  {
> -	if (dev == rcu_dereference_bh(tls_get_ctx(sk)->netdev) ||
> +	struct tls_context *tls_ctx = tls_get_ctx(sk);
> +
> +	if (unlikely(tls_tx_drop_acked_clone(sk, skb))) {
> +		kfree_skb(skb);
> +		return NULL;
> +	}
> +
> +	if (dev == rcu_dereference_bh(tls_ctx->netdev) ||
>  	    netif_is_bond_master(dev))
>  		return skb;
>  
> @@ -435,6 +498,65 @@ struct sk_buff *tls_validate_xmit_skb_sw(struct sock *sk,
>  	return tls_sw_fallback(sk, skb);
>  }
>  
> +struct sk_buff *tls_validate_xmit_skb_rekey(struct sock *sk,
> +					    struct net_device *dev,
> +					    struct sk_buff *skb)
> +{
> +	struct tls_context *tls_ctx = tls_get_ctx(sk);
> +	u32 tcp_seq = ntohl(tcp_hdr(skb)->seq);
> +	u32 pivot_seq;

[Severity: High]
Should this validator consult dev the way tls_validate_xmit_skb() does?

The peer function passes the skb through when the device is the offload device
or a bond master:

	if (dev == rcu_dereference_bh(tls_ctx->netdev) ||
	    netif_is_bond_master(dev))
		return skb;

That netif_is_bond_master() branch exists because sk_validate_xmit_skb() runs
twice for a bonded socket: once at the master level and again on the slave.
Since ctx->netdev is the lowest device (get_netdev_for_sock() ->
netdev_sk_get_lowest_dev()), dev at the master level is bond0 != ctx->netdev.

Under the rekey validator a pre-pivot (old-key) retransmit then takes this
path twice:

tls_validate_xmit_skb_rekey(dev=bond0)
  before(tcp_seq, pivot_seq)
    tls_sw_fallback()            /* produces ciphertext nskb */
      complete_skb()             /* nskb->decrypted = 0, nskb->sk = sk */
bond_dev_queue_xmit()
  __dev_queue_xmit(slave)
    validate_xmit_skb()
      sk_validate_xmit_skb()
        tls_validate_xmit_skb_rekey(dev=slave)
          before(tcp_seq, pivot_seq)   /* still true */
            tls_sw_fallback()          /* encrypts the ciphertext again */

The record is still on records_list (it is unacked, which is why the rekey is
PENDING), so the sync_size < 0 bypass does not catch it, and the same record,
rcd_sn, IV and sync_size mean the same CTR keystream is XORed back over the
ciphertext.  Can this put the original plaintext on the wire with an invalid
tag?

tls_validate_xmit_skb_sw() has the same omission, but that one is only
reachable after tls_device_down(), whereas the rekey validator is on the
normal KeyUpdate path.

> +
> +	/* acquire pairs with clear_bit_unlock() on re-arm; makes the refreshed
> +	 * boundary_seq visible in the else branch below.
> +	 */
> +	if (test_bit_acquire(TLS_TX_REKEY_FAILED, &tls_ctx->flags)) {
> +		int payload_len = skb->len - skb_tcp_all_headers(skb);
> +		u32 snd_una = READ_ONCE(tcp_sk(sk)->snd_una);
> +

[ ... ]

> +		if (payload_len > 0 && !skb_is_decrypted(skb) &&
> +		    after(tcp_seq + payload_len, snd_una))
> +			return skb;
> +
> +		pivot_seq = snd_una;
> +	} else {

[ ... ]

> +		smp_rmb();
> +		pivot_seq = READ_ONCE(tls_ctx->rekey.boundary_seq);
> +	}
> +
> +	/* At or after the pivot: already correctly encrypted, pass through */
> +	if (!before(tcp_seq, pivot_seq))
> +		return skb;
> +
> +	/* Below the pivot: retransmit of old data, SW fallback with old key */
> +	return tls_sw_fallback(sk, skb);
> +}

[ ... ]

> diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
> index 5531303dd704..fd162d8f1d64 100644
> --- a/net/tls/tls_sw.c
> +++ b/net/tls/tls_sw.c

[ ... ]

> @@ -2609,7 +2633,8 @@ int tls_sw_ctx_init(struct sock *sk, int tx,
>  			goto free_aead;
>  	}
>  
> -	ctx->push_pending_record = tls_sw_push_pending_record;
> +	if (tx)
> +		ctx->push_pending_record = tls_sw_push_pending_record;

[Severity: Medium]
Is this guard a separate fix that deserves its own patch and a Fixes: tag?

Before this change, an RX SW setsockopt on a socket already running HW TX
offload replaced the device dispatcher:

tls_set_device_offload_rx()
  tls_sw_ctx_init(sk, 0, NULL)
    ctx->push_pending_record = tls_sw_push_pending_record;

After that, with pending_open_record_frags set by tls_push_data(),
tls_handle_open_record():

	if (tls_is_pending_open_record(ctx))
		return ctx->push_pending_record(sk, flags);

calls tls_sw_push_pending_record(), which reads ctx->open_rec at the
tls_sw_context_tx offset while priv_ctx_tx is actually a
struct tls_offload_context_tx (landing inside sg_tx_data[]), and can then
dispatch bpf_exec_tx_verdict() on that.

The guard is also a prerequisite for the rekey flush, which calls
ctx->push_pending_record() and needs it to still be the device dispatcher.
Could this be split out with a Fixes: tag, or at least called out in the
changelog?

Two documentation concerns were also raised and are resolved later in the
series; noting them only for completeness.

The three new counters (TlsTxRekeyFallback, TlsCurrTxRekey,
TlsTxRekeyAborted) are not described in Documentation/networking/tls.rst at
this commit.

[Severity: Low]
The later commit "tls: document TLS 1.3 hardware offload rekey handling" adds
the TlsCurrTxRekey, TlsTxRekeyFallback and TlsTxRekeyAborted entries, so this
is addressed by the end of the series.

Documentation/networking/tls-offload.rst still says key rotation is not a
concern for offloaded connections and describes tls_dev_del() as destroy-time
cleanup only, while tls_device_complete_rekey() does a tls_dev_del() followed
by tls_dev_add() on a live socket.

[Severity: Low]
The same later documentation commit rewrites the tls_dev_del() lifecycle text
and the Rekey section, so this is addressed by the end of the series.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 12/15] tls: device: add RX KeyUpdate support
  2026-09-17 22:35 ` [PATCH net-next v17 12/15] tls: device: add RX " Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 9 potential
issue(s) to consider.

Critical: 0 · High: 4 · Medium: 3 · Low: 2

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] A record whose payload the NIC transformed (XORed) with the
  retired key but could not authenticate can be classified as untouched…
- [High] The inverse of the same ambiguity: tls_device_reencrypt() now
  rewrites every fragment whose skb->decrypt_failed is set (`if…
- [High] Re-adding an mlx5 RX offload context on a live socket (now
  possible because this patch calls tls_dev_del() followed by…
- [High] mlx5e_ktls_add_rx() publishes the new priv_rx into the TLS
  context (mlx5e_set_ktls_rx_priv_ctx()) before creating the TIR and…
- [Medium] Commit message contradicts the code for the two sequence
  computations the whole RX rekey classification rests on.
- [Medium] During the old-key drain, tls_device_decrypted() advances
  ctx->rekey.old_rec_seq before software decryption of the record has…
- [Medium] In tls_set_device_offload_rx()'s rekey path, an empty TCP
  receive snapshot is treated as proof that no old-key-transformed data…
- [Low] The patch adds three user-visible SNMP counters
  (TlsRxRekeyFallback, TlsCurrRxRekey, TlsRxRekeyAborted) exposed via…
- [Low] The patch silently fixes a pre-existing, userspace-observable
  tracepoint bug (trace_tls_device_decrypted() passed…

--

tls: device: add RX KeyUpdate support

This adds RX KeyUpdate handling to the TLS device offload path: the old
AEAD, IV and record sequence are kept on tls_offload_context_rx, the NIC
key is deleted when a KeyUpdate record is decoded, and records are then
classified against old_nic_boundary until the old key can be retired.
Installation of the new key on the NIC is deferred until the old key's
records have drained, with new SNMP counters and a TLS_RX_REKEY_FAILED
flag for the SW fallback accounting.

A few questions below, first about the commit message and then about the
code.

> rec_start_seq is the TCP sequence of the record's first byte, used both
> for the trace_tls_device_decrypted() tracepoint and the old_nic_boundary
> classification above. Because copied_seq is advanced at different points
> in the two strparser modes, the record start is computed differently: in
> copy_mode the record has already been dequeued (tcp_read_done() in
> tls_strp_msg_cow() advanced copied_seq past it), so full_len is
> subtracted; in non-copy mode copied_seq still points at the record start
> and is used directly.

[ ... ]

> installs the new key when the first post-boundary record is seen. It
> anchors the NIC on that record's start (rec_start_seq, already
> copy_mode-adjusted as above) paired with the new key's starting record
> number, read from crypto_recv.info after tls_sw_ctx_finalize() stored it
> there.

[Severity: Medium]
Do these two paragraphs describe something different from what the code
does?

tls_device_rx_rec_start() subtracts the anchor length, not full_len:

	if (sw_ctx->strp.copy_mode)
		return copied_seq - sw_ctx->strp.anchor->len;

and the helper's own comment attributes the copy_mode advance to
tls_strp_read_copy() -> tls_strp_read_copyin() -> tcp_read_sock(), not to
tcp_read_done() in tls_strp_msg_cow().

The two quantities coincide only at the tls_device_decrypted() call site,
where msg_ready implies anchor->len == full_len.  At the other call site,
drain_start in tls_set_device_offload_rx(), tls_strp_read_copy() has
zeroed anchor->len and tcp_read_sock() may have appended a partial record
or only header bytes, so full_len would be the wrong quantity there.

For the second paragraph, tls_device_deferred_dev_add_rx() deliberately
overwrites the frozen crypto_recv.info record number with the live one:

	crypto_ctx = tls_ctx->crypto_recv;
	memcpy(crypto_info_rec_seq(&crypto_ctx.info, cipher_desc),
	       tls_ctx->rx.rec_seq, cipher_desc->rec_seq);

and the comment right above it explains why the value from
crypto_recv.info must not be used.  Could the message be updated to match,
so a later simplification toward the described behaviour does not
desynchronise the (TCP seq, record number) pair handed to the NIC?

> diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
> index f32c1bb6b497..ac09f356cff9 100644
> --- a/net/tls/tls_device.c
> +++ b/net/tls/tls_device.c

[ ... ]

> @@ -192,6 +202,129 @@ static void tls_device_commit_start_marker(struct sock *sk,
>  	tcp_write_collapse_fence(sk);
>  }
>  

[ ... ]

> +static int tls_device_dev_add_rx(struct sock *sk, struct tls_context *tls_ctx,
> +				 struct net_device *netdev,
> +				 struct tls_crypto_info *crypto_info,
> +				 u32 cur_seq, bool is_rekey)
> +{
> +	const struct tls_cipher_desc *cipher_desc;
> +	char *rec_seq;
> +	int rc;
> +
> +	cipher_desc = get_cipher_desc(crypto_info->cipher_type);
> +	DEBUG_NET_WARN_ON_ONCE(!cipher_desc || !cipher_desc->offloadable);
> +
> +	rc = netdev->tlsdev_ops->tls_dev_add(netdev, sk,
> +					     TLS_OFFLOAD_CTX_DIR_RX,
> +					     crypto_info, cur_seq);

[Severity: High]
Calling tls_dev_add() for RX on a live socket is new here (previously add
happened once and del only at close).  Can a GET_PSV completion belonging
to the retired mlx5 context then operate on the new one?

mlx5e_ktls_del_rx() states that it does not wait for those completions:

	/* priv_rx should normally be freed here, but if there is an outstanding
	 * GET_PSV, deallocation will be delayed until the CQE for GET_PSV is
	 * processed. */

and mlx5e_ktls_handle_get_psv_completion() resolves the resync state
through the socket, not through its own retired priv_rx:

	rx_ctx = tls_offload_ctx_rx(tls_get_ctx(priv_rx->sk));
	async_resync = rx_ctx->resync_async;

After a del+add, that pointer is the new generation's &resync->core, so
the stale completion's DELETING branch would call
tls_offload_rx_resync_async_request_cancel() on the new generation's
pending request, leaving the NIC unsynced.

Can the same completion also write through freed memory?
mlx5e_ktls_add_rx() publishes

	tls_offload_ctx_rx(tls_ctx)->resync_async = &resync->core;

before the fallible post_rx_param_wqes(), and the error path does
kfree(priv_rx) without clearing that pointer.

[Severity: High]
Related question on the same call: is the replacement priv_rx visible to
NAPI before it is fully initialised?  mlx5e_ktls_add_rx() calls
mlx5e_set_ktls_rx_priv_ctx() before mlx5e_rx_res_tls_tir_create() and
resync_init(), so at that moment resync->priv is still NULL (kzalloc'd),
resync->work is uninitialised and the refcount is 0.

A RESYNC CQE generated before tls_dev_del() but polled after the re-add
finds the socket by 5-tuple in resync_update_sn() and reaches
resync_queue_get_psv(), which only checks pointer presence and
MLX5E_PRIV_RX_FLAG_DELETING:

	mlx5e_ktls_priv_rx_get(priv_rx);
	if (unlikely(!queue_work(resync->priv->tls->rx_wq, &resync->work)))

Would that dereference a NULL resync->priv in softirq context, on top of a
refcount_inc() from 0 and a queue_work() on an uninitialised work_struct?
The synchronize_net() in mlx5e_ktls_del_rx() only waits for a running
poll, not for CQEs already sitting in the CQ, and the new
TLS_RX_DEV_CLOSED test only gates the core parser path.

[ ... ]

> +void tls_device_rx_del_key(struct sock *sk, struct tls_context *ctx)
> +{

[ ... ]

> +	set_bit(TLS_RX_DEV_CLOSED, &ctx->flags);
> +	synchronize_net();
> +	netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
> +					TLS_OFFLOAD_CTX_DIR_RX);
> +	up_read(&device_offload_lock);
> +}
> +

[ ... ]

> @@ -1114,7 +1249,7 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
>  	if (skb_pagelen(skb) > offset) {
>  		copy = min_t(int, skb_pagelen(skb) - offset, data_len);
>  
> -		if (skb->decrypted) {
> +		if (skb->decrypted || skb->decrypt_failed) {
>  			err = skb_store_bits(skb, offset, buf, copy);
>  			if (err)
>  				goto free_buf;
> @@ -1141,7 +1276,7 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
>  		copy = min_t(int, skb_iter->len - frag_pos,
>  			     data_len + rxm->offset - offset);
>  
> -		if (skb_iter->decrypted) {
> +		if (skb_iter->decrypted || skb_iter->decrypt_failed) {
>  			err = skb_store_bits(skb_iter, frag_pos, buf, copy);
>  			if (err)
>  				goto free_buf;

[Severity: High]
Can decrypt_failed be used as proof that these bytes were XORed?  The
kernel-doc for the new bit says it cannot:

 * The payload may have been transformed (XORed) or left as wire
 *		ciphertext, so software must re-authenticate the record and undo the
 *		transform on any XORed fragment

If a fragment marked decrypt_failed was left as wire ciphertext, does
skb_store_bits() here overwrite pristine new-key ciphertext, so the SW
AEAD that follows rejects an otherwise valid record and
tls_rx_one_record() calls tls_err_abort(sk, -EBADMSG)?

There is also no way to tell the two kinds of fragment apart at this
point, since skb_cmp_decrypted() ignores the bit:

	return skb2->decrypted - skb1->decrypted;

so tcp_gro_receive() (flush |= skb_cmp_decrypted(p, skb)) and
tls_strp_copyin() can merge an XORed decrypt_failed segment with an
untouched one into a single skb or anchor fragment that carries one mark.

Note this hunk also affects sockets that never rekey, via the
TLS_RX_DEV_DEGRADED path in tls_device_decrypted().

> @@ -1158,6 +1293,77 @@ tls_device_reencrypt(struct sock *sk, struct tls_context *tls_ctx)
>  	return err;
>  }
>  

[ ... ]

> +static u32 tls_device_rx_rec_start(struct sock *sk,
> +				   struct tls_sw_context_rx *sw_ctx)
> +{
> +	u32 copied_seq = tcp_sk(sk)->copied_seq;
> +
> +	if (sw_ctx->strp.copy_mode)
> +		return copied_seq - sw_ctx->strp.anchor->len;
> +
> +	return copied_seq;
> +}
> +
>  int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
>  {
> @@ -1165,6 +1371,7 @@ int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
>  	struct sk_buff *skb = tls_strp_msg(sw_ctx);
>  	struct strp_msg *rxm = strp_msg(skb);
>  	int is_decrypted, is_encrypted;
> +	u32 rec_start_seq;
>  
> @@ -1174,10 +1381,72 @@ int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
>  		is_encrypted = 0;
>  	}
>  
> -	trace_tls_device_decrypted(sk, tcp_sk(sk)->copied_seq - rxm->full_len,
> +	rec_start_seq = tls_device_rx_rec_start(sk, sw_ctx);
> +
> +	trace_tls_device_decrypted(sk, rec_start_seq,
>  				   tls_ctx->rx.rec_seq, rxm->full_len,
>  				   is_encrypted, is_decrypted);
>  

[Severity: Low]
This change also fixes a user-visible tracepoint bug that predates the
patch: the old expression subtracted rxm->full_len unconditionally, and on
the non-copy path tcp_sk(sk)->copied_seq is still at the record start, so
the tracepoint reported the previous record's start.

Would it be worth splitting this into its own patch with a Fixes: tag, so
it can be backported independently of the feature?

> +	if (unlikely(ctx->rekey.old_aead_recv)) {
> +		bool nic_touched = !is_encrypted || skb->decrypt_failed;
> +		bool before_nic_boundary;
> +

[ ... ]

> +		if (nic_touched &&
> +		    !before(rec_start_seq, ctx->rekey.old_nic_boundary))
> +			ctx->rekey.old_nic_boundary = rec_start_seq + rxm->full_len;
> +
> +		before_nic_boundary =
> +			before(rec_start_seq, ctx->rekey.old_nic_boundary);
> +
> +		if (before_nic_boundary) {

[ ... ]

> +			if (is_encrypted) {
> +				tls_bigint_increment(ctx->rekey.old_rec_seq,
> +						     tls_ctx->prot_info.rec_seq_size);
> +				return 0;
> +			}
> +
> +			return tls_device_reencrypt_old_key(sk, ctx,
> +							    sw_ctx, tls_ctx);
> +		}

[Severity: High]
Is a record whose payload the NIC XORed with the retired key always
flagged mixed, so that it takes the reencrypt path rather than this
is_encrypted shortcut?

mixed_decrypted is derived solely from skb_cmp_decrypted() in
tls_strp_copyin():

	strp->mixed_decrypted |= !!skb_cmp_decrypted(skb, in_skb);

and skb_cmp_decrypted() looks only at skb->decrypted:

	return skb2->decrypted - skb1->decrypted;

So a record whose every segment carries only decrypt_failed is non-mixed,
is_encrypted is 1, and the branch above returns 0 without ever calling
tls_device_reencrypt_old_key().

The commit message says such a record "was not transformed", but
mlx5e_ktls_handle_rx_skb() sets the flag unconditionally and says the
opposite:

	/* The device could not authenticate the payload. Depending on
	 * where the failure occurred the bytes may have been transformed
	 * (XORed) or left as wire ciphertext. */
	skb->decrypt_failed = 1;

If the payload was transformed, tls_decrypt_sw() authenticates XORed bytes
and tls_rx_one_record() does:

	if (err < 0) {
		tls_err_abort(sk, -EBADMSG);
		return err;
	}

which kills a healthy TLS 1.3 connection during the rekey the patch is
meant to support.

Independently of what the device does, does GRO destroy the per-segment
distinction this classification depends on?  tcp_gro_receive() has:

	flush |= skb_cmp_decrypted(p, skb);

which is 0 for a (unmarked, decrypt_failed) pair, so the two are merged
into one skb carrying only the head's flags.

[Severity: Medium]
Can old_rec_seq end up one record ahead if software decryption of this
record fails transiently?  The increment above happens before
tls_decrypt_sw() runs, and on a failure tls_rx_one_record() returns
without tls_advance_record_sn() and without consuming the strparser
record.

On the next recvmsg(), tls_rx_rec_wait() returns 1 immediately because
tls_strp_msg_ready() is already true - its sk_err checks live inside the

	while (!tls_strp_msg_ready(ctx))

loop - so tls_device_decrypted() runs again for the same record and
advances ctx->rekey.old_rec_seq a second time.  Would a later mixed record
then be reconstructed by tls_device_reencrypt_old_key() with the wrong
nonce?

[ ... ]

> @@ -1804,73 +2073,224 @@ int tls_set_device_offload(struct sock *sk,

[ ... ]

> +		} else {
> +			struct tcp_sock *tp = tcp_sk(sk);
> +			u32 nic_end;
> +
> +			if (context->rekey.old_aead_recv) {
> +				crypto_free_aead(context->rekey.old_aead_recv);
> +				context->rekey.old_aead_recv = NULL;
> +			}
> +
> +			/* Flush the backlog so TCP's view is current, then take the
> +			 * highest byte TCP holds, including the out-of-order tail:
> +			 * a NIC-transformed segment behind a host-side drop sits
> +			 * above rcv_nxt until the retransmit fills the hole and
> +			 * must still be classified against the old key. This is
> +			 * still only the stack's view, a transformed segment the
> +			 * NIC has not delivered yet is caught in-band by
> +			 * tls_device_decrypted(), which slides the boundary.
> +			 */
> +			__sk_flush_backlog(sk);
> +			nic_end = tp->rcv_nxt;
> +			if (!RB_EMPTY_ROOT(&tp->out_of_order_queue) &&
> +			    after(TCP_SKB_CB(tp->ooo_last_skb)->end_seq, nic_end))
> +				nic_end = TCP_SKB_CB(tp->ooo_last_skb)->end_seq;
> +
> +			if (before(drain_start, nic_end)) {
> +				context->rekey.old_aead_recv = sw_ctx->aead_recv;

[Severity: Medium]
What happens when this snapshot finds nothing queued?  Then
before(drain_start, nic_end) is false, no old AEAD/IV/rec_seq is retained
(any previous one has just been freed above) and the new key goes to the
NIC immediately.

The comment in tls_device_decrypted() says a late arrival is possible:

	 * A segment the NIC transformed with the old
	 * key before tls_dev_del returned can still be in the RQ/CQ, in
	 * a GRO list or in the socket backlog when that snapshot is
	 * taken and reach TCP later, above it.

but the boundary-sliding code that compensates for that runs only inside
if (unlikely(ctx->rekey.old_aead_recv)).  With no old key retained, would
a late mixed record fall through to the tail of tls_device_decrypted() and
be reconstructed by tls_device_reencrypt() with the current key, so
authentication fails and the connection is aborted with -EBADMSG?

The synchronize_net() in mlx5e_ktls_del_rx() removes the steering rule and
TIR but does not drain already-transformed RX packets into TCP, so is
there anything that rules out this asymmetry between the retain and
no-retain branches?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 13/15] tls: device: add tracepoints for the KeyUpdate path
  2026-09-17 22:35 ` [PATCH net-next v17 13/15] tls: device: add tracepoints for the KeyUpdate path Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] New tracepoint `tls_device_rekey_start`
  (net/tls/trace.h:195-221) declares and prints its sequence field as…
- [Low] The new comment added at the tls_device_sendmsg() call site
  (net/tls/tls_device.c:869-877) states "Non-zero here is the transient…

--

tls: device: add tracepoints for the KeyUpdate path

This adds five trace events to net/tls/trace.h covering the device
KeyUpdate/rekey state machine: rekey_start, rekey_reencrypt and
rekey_done on the RX side, plus complete_rekey_retry and
complete_rekey_fail on the TX side. The RX events are emitted from
tls_set_device_offload_rx() and tls_device_decrypted(), the TX ones from
tls_device_sendmsg() and the fallback path of
tls_device_complete_rekey().

No Link: tags were present in the commit header.

> diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
> index ac09f356cff9..5f45c097bad3 100644
> --- a/net/tls/tls_device.c
> +++ b/net/tls/tls_device.c
> @@ -866,8 +866,16 @@ int tls_device_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
>  	lock_sock(sk);
>  
>  	/* Old-key records all ACKed; switch back to HW. */
> -	if (test_bit(TLS_TX_REKEY_READY, &tls_ctx->flags))
> -		tls_device_complete_rekey(sk, tls_ctx, true, msg->msg_flags);
> +	if (test_bit(TLS_TX_REKEY_READY, &tls_ctx->flags)) {
> +		rc = tls_device_complete_rekey(sk, tls_ctx, true, msg->msg_flags);
> +		/* Non-zero here is the transient -EAGAIN retry,
> +		 * the next sendmsg retries. Hard failures return 0 after
> +		 * falling back to SW and emit tls_device_complete_rekey_fail
> +		 * from the fallback path.
> +		 */
> +		if (rc)
> +			trace_tls_device_complete_rekey_retry(sk);
> +	}

[Severity: Low]
Is "Non-zero here is the transient -EAGAIN retry" the full set of
transient returns here?

tls_device_complete_rekey() passes three errors back to this caller:

net/tls/tls_device.c:tls_device_complete_rekey() {
	...
	/* -EAGAIN (sndbuf full) and a signal (-EINTR/-ERESTARTSYS from
	 * sk_stream_wait_memory()) are transient: leave the rekey PENDING and
	 * retry on the next sendmsg rather than permanently dropping HW offload.
	 */
	if (rc == -EAGAIN || rc == -EINTR || rc == -ERESTARTSYS)
		return rc;
	...
}

So a signal delivered to a sender blocked in sk_stream_wait_memory()
during a KeyUpdate also emits trace_tls_device_complete_rekey_retry(),
not only sndbuf pressure. Should the comment here (and the matching
"TX rekey completion hit the transient -EAGAIN retry in sendmsg" text in
the changelog) name -EINTR/-ERESTARTSYS too?

Related: the retry event carries only sk, while its sibling
tls_device_complete_rekey_fail(sk, rc) carries rc. Without rc on the
retry event, can a trace distinguish repeated sndbuf exhaustion from
repeated signal delivery when a TX rekey stays PENDING? Would adding rc
to tls_device_complete_rekey_retry make that visible?

>  
>  	if (tls_device_tx_uses_sw(tls_ctx)) {
>  		rc = tls_sw_sendmsg_locked(sk, msg, size);
> @@ -1430,10 +1438,15 @@ int tls_device_decrypted(struct sock *sk, struct tls_context *tls_ctx)
>  				return 0;
>  			}
>  
> +			trace_tls_device_rekey_reencrypt(sk, rec_start_seq,
> +							 ctx->rekey.old_nic_boundary);
> +
>  			return tls_device_reencrypt_old_key(sk, ctx,
>  							    sw_ctx, tls_ctx);
>  		}
>  
> +		trace_tls_device_rekey_done(sk, rec_start_seq,
> +					    ctx->rekey.old_nic_boundary);
>  		crypto_free_aead(ctx->rekey.old_aead_recv);
>  		ctx->rekey.old_aead_recv = NULL;
>  

[ ... ]

> @@ -1890,6 +1903,13 @@ static int tls_device_complete_rekey(struct sock *sk, struct tls_context *ctx,
>  	TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXDEVICE);
>  	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXSW);
>  
> +	/* Hard failure: HW rekey gave up and the connection is now pinned to
> +	 * SW encryption. The call site only sees the transient -EAGAIN retry
> +	 * (rc is not propagated here), so emit the trace from the fallback
> +	 * path itself; rc still holds the originating error.
> +	 */
> +	trace_tls_device_complete_rekey_fail(sk, rc);
> +
>  	return 0;
>  }
>  
> @@ -2195,11 +2215,21 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
>  			 * is installed once drain_start crosses rekey.old_nic_boundary.
>  			 */
>  			context->dev_add_pending = 1;
> +			trace_tls_device_rekey_start(sk, drain_start,
> +						     context->rekey.old_nic_boundary,
> +						     true);
>  		} else {
>  			struct tcp_sock *tp = tcp_sk(sk);
>  			u32 nic_end;
>  
>  			if (context->rekey.old_aead_recv) {
> +				/* Prior rekey's era already drained (drain_start is
> +				 * past old_nic_boundary), so retiring its key here
> +				 * is a boundary crossing, same as the free in
> +				 * tls_device_decrypted(); mark it done.
> +				 */
> +				trace_tls_device_rekey_done(sk, drain_start,
> +							    context->rekey.old_nic_boundary);
>  				crypto_free_aead(context->rekey.old_aead_recv);
>  				context->rekey.old_aead_recv = NULL;
>  			}

[ ... ]

> @@ -2247,6 +2277,8 @@ int tls_set_device_offload_rx(struct sock *sk, struct tls_context *ctx,
>  				context->dev_add_pending = 0;
>  				retired_pending = true;
>  			}
> +			trace_tls_device_rekey_start(sk, drain_start, nic_end,
> +						     before(drain_start, nic_end));
>  		}
>  	}
>  
> diff --git a/net/tls/trace.h b/net/tls/trace.h
> index 2d8ce4ff3265..5b9c1f86d82d 100644
> --- a/net/tls/trace.h
> +++ b/net/tls/trace.h
> @@ -192,6 +192,124 @@ TRACE_EVENT(tls_device_tx_resync_send,
>  	)
>  );
>  
> +TRACE_EVENT(tls_device_rekey_start,
> +
> +	TP_PROTO(struct sock *sk, u32 copied_seq, u32 nic_boundary,
> +		 bool inflight),
> +
> +	TP_ARGS(sk, copied_seq, nic_boundary, inflight),
> +
> +	TP_STRUCT__entry(
> +		__field(	struct sock *,	sk		)
> +		__field(	u32,		copied_seq	)
> +		__field(	u32,		nic_boundary	)
> +		__field(	bool,		inflight	)
> +	),
> +
> +	TP_fast_assign(
> +		__entry->sk = sk;
> +		__entry->copied_seq = copied_seq;
> +		__entry->nic_boundary = nic_boundary;
> +		__entry->inflight = inflight;
> +	),
> +
> +	TP_printk(
> +		"sk=%p copied_seq=%u nic_boundary=%u inflight=%d",
> +		__entry->sk, __entry->copied_seq, __entry->nic_boundary,
> +		__entry->inflight
> +	)
> +);

[Severity: Medium]
Is copied_seq the right name for this field? Both producers in
tls_set_device_offload_rx() pass drain_start, which comes from
tls_device_rx_rec_start():

net/tls/tls_device.c:tls_device_rx_rec_start() {
	u32 copied_seq = tcp_sk(sk)->copied_seq;

	if (sw_ctx->strp.copy_mode)
		return copied_seq - sw_ctx->strp.anchor->len;

	return copied_seq;
}

In strparser copy mode that is the record start, not tp->copied_seq,
which is also what the comment above drain_start says:

	/* Classify against the record start, not the raw copied_seq: ...

The same quantity is exported as tcp_seq by the pre-existing
tls_device_decrypted event and by the two RX events added here
(tls_device_rekey_reencrypt/tls_device_rekey_done, both fed
rec_start_seq). Since one header would then label one value two
different ways, would naming this field tcp_seq keep the interface
self-consistent for anyone correlating rekey_start with
tls_device_decrypted?

Separately, would the rx_/tx_ direction qualifier used by the existing
events in this header (tls_device_rx_resync_send,
tls_device_tx_resync_req, tls_device_tx_resync_send) be worth applying
here? rekey_start, rekey_done and rekey_reencrypt are only emitted from
the RX path (tls_set_device_offload_rx() and tls_device_decrypted()),
while complete_rekey_retry/_fail are only emitted from the TX path, so
enabling tls_device_rekey_start while chasing a TX KeyUpdate yields no
events.

> +
> +TRACE_EVENT(tls_device_rekey_reencrypt,
> +
> +	TP_PROTO(struct sock *sk, u32 tcp_seq, u32 nic_boundary),
> +
> +	TP_ARGS(sk, tcp_seq, nic_boundary),

[ ... ]

> +TRACE_EVENT(tls_device_complete_rekey_fail,
> +
> +	TP_PROTO(struct sock *sk, int rc),
> +
> +	TP_ARGS(sk, rc),

[ ... ]

> +TRACE_EVENT(tls_device_complete_rekey_retry,
> +
> +	TP_PROTO(struct sock *sk),
> +
> +	TP_ARGS(sk),
> +
> +	TP_STRUCT__entry(
> +		__field(	struct sock *,	sk	)
> +	),

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 14/15] selftests: net: add TLS hardware offload test
  2026-09-17 22:35 ` [PATCH net-next v17 14/15] selftests: net: add TLS hardware offload test Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 9 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 6 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] check_tls_support() in
  tools/testing/selftests/drivers/net/hw/tls_hw_offload.py runs…
- [Medium] check_tls_support() only inspects the direction-wide `tls-hw-
  tx-offload` / `tls-hw-rx-offload` netdev feature bits, which say…
- [Medium] The comment added in run_tls_test() asserts that the `finally`
  clause reaps a hung helper 'within this variant's window' so it cannot…
- [Medium] verify_tls_counters() accepts an aborted rekey as a legitimate
  terminal outcome: it asserts only `TlsTxRekeyAborted <= 1` and…
- [Medium] run_tls_test() starts the server with bkg(..., exit_wait=True)
  - which captures stdout/stderr into pipes and does not call…
- [Medium] In burst mode the server never validates how many complete
  iterations it received.
- [Low] HW_CRYPTO_COUNTERS maps only 'mlx5_core' to
  tx_tls_encrypted_packets / rx_tls_decrypted_packets, so…
- [Low] parse_int_arg(optarg, 1, INT_MAX, "iteration count",
  &num_iterations) (tls_hw_offload.c:1042) explicitly accepts -n…
- [Low] server_accept_tls() arms a 30 s SO_RCVTIMEO for accept()
  (tls_hw_offload.c:666) while the harness joins the background server…

--

selftests: net: add TLS hardware offload test

This adds a two-node kTLS hardware offload selftest: a C helper binary that
acts as TLS client or server, plus a Python harness built on NetDrvEpEnv
that snapshots and verifies the TLS MIB counters around each run. Variants
cover TLS 1.2/1.3 with AES-GCM-128/256, echo-mode rekeys over several
buffer sizes, and burst variants that stress TX and RX KeyUpdate handling.

A few questions about the test logic and the harness below.

> diff --git a/tools/testing/selftests/drivers/net/hw/config b/tools/testing/selftests/drivers/net/hw/config
> index d89a9ba17655..169e608516bd 100644
> --- a/tools/testing/selftests/drivers/net/hw/config
> +++ b/tools/testing/selftests/drivers/net/hw/config
> @@ -22,6 +22,8 @@ CONFIG_NET_IPIP=y
>  CONFIG_NETKIT=y
>  CONFIG_NET_SCH_INGRESS=y
>  CONFIG_SYNC_FILE=y
> +CONFIG_TLS=y
> +CONFIG_TLS_DEVICE=y
>  CONFIG_UDMABUF=y
>  CONFIG_USER_NS=y
>  CONFIG_VXLAN=y

[ ... ]

> diff --git a/tools/testing/selftests/drivers/net/hw/tls_hw_offload.c b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
> new file mode 100644
> index 000000000000..303c6752ace2
> --- /dev/null
> +++ b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.c
> @@ -0,0 +1,1132 @@

[ ... ]

> +	rekey_interval = num_iterations / (num_rekeys + 1);
> +	next_rekey_at = rekey_interval;
> +
> +	for (i = 1; i <= num_iterations; i++) {

[Severity: Low]
Is -n INT_MAX actually a usable value here?  main() accepts it:

	case 'n':
		if (parse_int_arg(optarg, 1, INT_MAX, "iteration count",
				  &num_iterations) < 0)

but do_client() drives a signed int with an inclusive bound, so
terminating the loop requires evaluating INT_MAX + 1, which is signed
overflow; with wrapping arithmetic i goes negative and the loop keeps
going.  do_server() has the same shape with send_iter++ in the burst
path.  The harness never passes anything near that, so this is only
reachable from the command line the usage text advertises - would
capping the accepted range (or making the counters unsigned/long) be
preferable?

> +		int this_size;

[ ... ]

> +	/* Bound accept() so a client that never connects (a deploy or connect
> +	 * failure on the peer) does not block the server forever and leak the
> +	 * process past the harness timeout. accept() honours SO_RCVTIMEO on the
> +	 * listening socket; the client connects right after wait_port_listen(),
> +	 * so 30s is generous.
> +	 */
> +	{
> +		struct timeval tv = { .tv_sec = 30, .tv_usec = 0 };
> +
> +		setsockopt(lsk, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv));
> +	}

[Severity: Low]
Does this 30 s bound in server_accept_tls() actually keep the server
inside the variant that started it?  run_tls_test() joins the server
through bkg(..., exit_wait=True), which reaches cmd.process() with no
timeout argument, i.e. the 20 s default in
tools/testing/selftests/net/lib/py/utils.py:

    def process(self, terminate=True, fail=None, expect_fail=False, timeout=20)

So on a client-side setup failure the helper's own accept() bound
outlives the harness's join window, and the listening socket plus the
process can still be around while the next variant picks a rand_port()
and takes its before/after counter snapshots.  Would 30 s here want to
be shorter than the harness join timeout, or the join given an explicit
timeout?

[ ... ]

> +	/* Main receive loop */
> +	while (1) {
> +		char *dst = burst_mode ? buf + filled : buf;
> +		size_t want = burst_mode ? (size_t)(send_size - filled)
> +					 : (size_t)buf_size;
> +
> +		n = recv_tls_message(csk, dst, want, &record_type, 0);
> +		if (n == 0) {
> +			/* A clean close on an iteration boundary is success;
> +			 * one with a partial iteration still buffered means the
> +			 * peer dropped the tail - the truncated-data case this
> +			 * test exists to catch, so fail loudly.
> +			 */
> +			if (burst_mode && filled) {
> +				printf("FAIL: closed mid-iteration (%d/%d bytes buffered)\n",
> +				       filled, send_size);
> +				goto out;
> +			}
> +			printf("Connection closed by client\n");
> +			break;
> +		}

[Severity: Medium]
Does this catch the truncation case the comment describes?  The EOF
check only rejects a partial iteration (filled != 0).  If a whole number
of send_size iterations is lost at the end of the stream, filled is 0,
do_server() breaks out and sets test_result = 0.

The count is tracked but never checked: recv_count++ happens per
completed iteration and is never compared against an expected total,
and run_tls_test() passes -n only to the client:

    if num_iterations:
        client_parts.append(f"-n {num_iterations}")

So losing a suffix of complete iterations - for instance after the last
KeyUpdate, where the rekey deltas already match and check_hw_crypto()
only needs one packet - passes on both helpers and in the harness.
Would passing the expected iteration count to the server and comparing
recv_count against it close that?

> +		if (n < 0) {
> +			printf("FAIL: recv failed: %s\n", strerror(errno));
> +			goto out;
> +		}

[ ... ]

> +		recv_count++;
> +		printf("Received %zd bytes (total: %zd, count: %d)\n",
> +		       n, total, recv_count);
> +
> +		if (send_all(csk, buf, n) < 0)
> +			goto out;
> +		printf("Echoed %zd bytes back to client\n", n);
> +	}

[Severity: Medium]
Can this per-recv logging stall the echo path?  The server is started
with bkg(server_cmd, host=server_host, exit_wait=True), which captures
stdout/stderr into pipes and only calls communicate() when the context
exits, while the harness waits for the client inside that context:

    with bkg(server_cmd, host=server_host, exit_wait=True):
        wait_port_listen(port, host=server_host)
        ...
        client.process(terminate=False, fail=True, timeout=cmd_timeout)

Two lines per recv from do_server() fill a 64 KiB pipe after roughly
1500 recvs.  Once printf() blocks, the server stops calling recv() and
send_all(), the client sits in client_recv_echo() waiting for an echo,
and nobody drains the server's pipe until the client finishes.  The 8 s
SO_RCVTIMEO turns that into "FAIL: Echo recv failed" rather than a hang,
but the failure has nothing to do with the kernel.  Would draining the
server's output concurrently, redirecting it to a file, or dropping the
per-iteration logging avoid this for the high-volume echo variants
(large_buf sends 100 x 2 MiB)?

[ ... ]

> diff --git a/tools/testing/selftests/drivers/net/hw/tls_hw_offload.py b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.py
> new file mode 100755
> index 000000000000..99ae5b3b8996
> --- /dev/null
> +++ b/tools/testing/selftests/drivers/net/hw/tls_hw_offload.py
> @@ -0,0 +1,446 @@

[ ... ]

> +# Per-packet HW crypto counters exposed via `ethtool -S` on the DUT NIC,
> +# keyed by the `ethtool -i` driver name. TlsTxDevice/TlsRxDevice in
> +# /proc/net/tls_stat only prove tls_dev_add() accepted the offload; these
> +# increment once per packet the NIC actually encrypted/decrypted (mlx5 counts
> +# gso_segs, not records), so they prove the HW crypto path was exercised.
> +# Names are driver-specific, so the
> +# check only runs on drivers listed here and is skipped (not failed) on
> +# others, keeping the test portable across NICs.
> +HW_CRYPTO_COUNTERS = {
> +    'mlx5_core': {'Tx': 'tx_tls_encrypted_packets',
> +                  'Rx': 'rx_tls_decrypted_packets'},
> +}

[Severity: Low]
Are these names really driver-specific?  Documentation/networking/tls-offload.rst
lists tx_tls_encrypted_packets and rx_tls_decrypted_packets under the
"minimum set of TLS-related statistics [that] should be reported by the
driver", and nfp (nfp_net_ethtool.c), cxgb4 and funeth export exactly
those strings.

With only mlx5_core in the table, check_hw_crypto() prints a NOTE and
returns on every other driver, leaving no assertion that the NIC did any
crypto.  check_path() does not substitute for it, since
do_tls_setsockopt_conf() bumps TlsTxDevice/TlsRxDevice at context
install time, before any record is processed.  Could the lookup fall
back to those documented names for any driver that exposes them?

> +def check_tls_support(cfg):
> +    """Skip the suite unless both hosts have kTLS and the DUT HW offload."""
> +    # The tls module is autoloaded lazily on the first TCP_ULP="tls"
> +    # setsockopt, so /proc/net/tls_stat (created from the module's pernet
> +    # init) may not exist yet on a freshly booted host. Load the module
> +    # explicitly before probing for it.
> +    try:
> +        cmd("modprobe tls")
> +        cmd("modprobe tls", host=cfg.remote)
> +        cmd("test -f /proc/net/tls_stat")
> +        cmd("test -f /proc/net/tls_stat", host=cfg.remote)
> +    except CmdExitFailure as e:
> +        raise KsftSkipEx(f"kTLS not supported: {e}") from e

[Severity: Medium]
Should the modprobe calls be inside this try block?  cmd() defaults to
fail=True and raises CmdExitFailure on a non-zero exit, so any modprobe
failure unrelated to kTLS - no modules.dep for a locally built kernel,
no modprobe on a minimal peer rootfs, kmod refusing a builtin - is
reported as "kTLS not supported" and skips the whole suite even though
/proc/net/tls_stat exists and offload works.

The config fragment this patch adds asks for CONFIG_TLS=y /
CONFIG_TLS_DEVICE=y, i.e. the intended configuration is built-in, where
the modprobe is unnecessary and the following test -f probe is already
authoritative.  Would making the modprobe best-effort (fail=False, or
its own try/except) avoid a permanent skip in that configuration?

> +
> +    try:
> +        features = cmd(f"ethtool -k {cfg.ifname}").stdout
> +        if 'tls-hw-tx-offload: on' not in features:
> +            raise KsftSkipEx("Device does not support TLS HW TX offload")
> +        if 'tls-hw-rx-offload: on' not in features:
> +            raise KsftSkipEx("Device does not support TLS HW RX offload")
> +    except CmdExitFailure as e:
> +        raise KsftSkipEx(f"Cannot determine TLS HW offload support: {e}") from e

[Severity: Medium]
Are the per-direction feature bits enough to gate the whole matrix?
They say nothing about which version/cipher tuples the device can
install, and drivers advertise them while rejecting tuples:

drivers/net/ethernet/netronome/nfp/crypto/tls.c:nfp_net_tls_add() {
	...
	if (crypto_info->version != TLS_1_2_VERSION)
		return -EOPNOTSUPP;
	...
}

mlx5 similarly checks separate firmware caps.  On an initial install
failure do_tls_setsockopt_conf() falls back to software:

net/tls/tls_main.c:do_tls_setsockopt_conf() {
	...
		} else {
			rc = tls_set_sw_offload(sk, 1, update ? crypto_info : NULL);
			...
			conf = TLS_SW;
	...
}

so the run proceeds in SW and check_path(..., require_hw=True) then
fails with "HW offload not engaged".  On an nfp or chcr NIC that means
the TLS 1.3 and AES-GCM-256 variants report failures rather than skips -
could the harness probe each tuple (for example a throwaway socket
install, or a TlsTxDevice check used to skip) instead of asserting?

[ ... ]

> +    if expected_rekeys > 0:
> +        if with_tx:

[ ... ]

> +            ksft_ge(1, diff('TlsTxRekeyAborted'),
> +                    comment=f"{role} Tx: TlsTxRekeyAborted expected <= 1")
> +            ksft_eq(diff('TlsTxRekeyOk') + diff('TlsTxRekeyAborted') +
> +                    diff('TlsTxRekeyFallback'), expected_rekeys,
> +                    comment=f"{role} Tx: rekey outcomes must sum to "
> +                            f"{expected_rekeys}")

[Severity: Medium]
Should an aborted rekey count as a satisfied expectation?  The Aborted
counters are emitted when a still-pending rekey is destroyed at teardown,
i.e. when it never completed:

net/tls/tls_device.c:tls_device_free_resources_tx() {
	...
	if (test_bit(TLS_TX_REKEY_PENDING, &tls_ctx->flags)) {
		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYABORTED);
	...
}

net/tls/tls_device.c:tls_device_offload_cleanup_rx() {
	...
	if (rx_ctx && rx_ctx->dev_add_pending) {
		rx_ctx->dev_add_pending = 0;
		TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSRXREKEYABORTED);
	...
}

With Aborted folded into the sum, the single-rekey variants ("single",
burst_rx_rekey_every_10000, burst_rx_zc_rekey_every_20000) pass with
RekeyOk=0, RekeyAborted=1, Fallback=0, Error=0, CurrRekey=0 - no
successful HW rekey at all.  The RX block has the same shape.

Multi-rekey variants have a related hole, since a superseded pending
rekey bumps Ok without a HW reinstall:

net/tls/tls_device.c:tls_set_device_offload_rekey() {
	...
	if (defer) {
		if (!rekey_pending)
			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
		else
			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYOK);
	...
}

so N-1 Ok plus 1 Aborted also passes.  check_hw_crypto() spans the whole
connection, so records sent under the initial key satisfy it too.  Would
requiring RekeyOk == expected_rekeys on the DUT (and Aborted == 0) be
the stricter oracle intended here?

[ ... ]

> +    with bkg(server_cmd, host=server_host, exit_wait=True):
> +        wait_port_listen(port, host=server_host)
> +        # Start the client in the background so we keep a handle to it. A
> +        # foreground cmd() raises TimeoutExpired from inside its constructor
> +        # if the client hangs, and since the child is not killed on timeout
> +        # it would be left running with no handle to reap it. A leaked
> +        # client keeps bumping the per-netns TLS counters (TlsTxRekeyAborted,
> +        # TlsDecryptError, ...) and would corrupt the before/after
> +        # measurement window of a later variant. The finally clause reaps it
> +        # within this variant's window instead.
> +        client = cmd(client_cmd, host=client_host, background=True)
> +        try:
> +            client.process(terminate=False, fail=True, timeout=cmd_timeout)
> +        finally:
> +            if client.proc.poll() is None:
> +                client.process(terminate=True, fail=False, timeout=5)

[Severity: Medium]
Does the finally clause reap a client that runs on the peer?  With the
ssh backend, cmd.proc is the local ssh client:

tools/testing/selftests/drivers/net/lib/py/remote_ssh.py:Remote.cmd() {
    def cmd(self, comm):
        return subprocess.Popen(["ssh", "-q", self.name, comm],
                                stdout=subprocess.PIPE, stderr=subprocess.PIPE)
}

There is no pty and no remote PID tracking, and cmd._process_terminate()
only signals that handle:

    def _process_terminate(self, terminate, timeout):
        if terminate:
            self.proc.terminate()
        stdout, stderr = self.proc.communicate(timeout=timeout)

so SIGTERM reaches ssh, not tls_hw_offload on the peer, and there is no
SIGKILL fallback when communicate(timeout=5) expires.  The helper also
no longer dies when its stdout pipe closes, since main() does:

	signal(SIGPIPE, SIG_IGN);

printf() just returns EPIPE and the send/recv loop continues.

This covers the paths the new cleanup code is aimed at: the burst_rx_*
variants set dut_role="server", so the client runs remotely for up to
BURST_TIMEOUT_S, and the dut_role="client" variants run the server
remotely under bkg().  A surviving remote helper keeps driving kTLS
while the next variant takes its /proc/net/tls_stat and ethtool -S
snapshots, which breaks the exact equalities (rekey outcome sums,
TlsCurrTxRekey == 0, TlsDecryptError == 0).  Would the remote helper
need an explicit kill (pkill over ssh, ssh -tt, or a recorded remote
pid) for the comment's guarantee to hold?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH net-next v17 15/15] tls: document TLS 1.3 hardware offload rekey handling
  2026-09-17 22:35 ` [PATCH net-next v17 15/15] tls: document TLS 1.3 hardware offload rekey handling Rishikesh Jethwani
@ 2026-09-22  1:56   ` netdev-bot+sashiko
  0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22  1:56 UTC (permalink / raw)
  To: rjethwani
  Cc: netdev, saeedm, tariqt, mbloch, borisp, john.fastabend, kuba, sd,
	davem, pabeni, edumazet, leon, andrew.gospodarek

Thank you for your contribution! Sashiko AI review found 7 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 3 · Low: 4

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The new decrypt_failed contract states, for a record none of
  whose segments carry `decrypted`, that "The stack passes such a record…
- [Medium] The newly documented decrypt_failed contract requires the
  driver to know whether the device transformed (XORed) the failed…
- [Medium] The new documentation permits coalescing of decrypt_failed
  segments with unmarked ones ("Coalescing (GRO, socket layer) and…
- [Low] The documented meanings of the rekey counters in
  Documentation/networking/tls.rst do not match their…
- [Low] The rewritten, exhaustively worded description of when
  tls_dev_del is called ("either when the offloaded connection is…
- [Low] In the "Error handling" section, the retained blanket statement
  "The Linux networking stack does not provide a way of reporting per-…
- [Low] tls-offload.rst states the TX switch "completes inline during
  setsockopt" if "no hardware-offloaded records remain unacknowledged".

--

tls: document TLS 1.3 hardware offload rekey handling

This replaces the stale "offload does not support TLS 1.3" note in
Documentation/networking/tls-offload.rst with a description of KeyUpdate
handling on TX and RX, including the skb->decrypt_failed driver contract,
and documents the new rekey counters in Documentation/networking/tls.rst.

No Link: tags are present in the commit header.

A few questions about the new text where it does not seem to line up with
net/tls/tls_device.c and the only in-tree producer of the new mark.

> diff --git a/Documentation/networking/tls-offload.rst b/Documentation/networking/tls-offload.rst
> index e5802bcd4d22..cdf84f4b817a 100644
> --- a/Documentation/networking/tls-offload.rst
> +++ b/Documentation/networking/tls-offload.rst

[ ... ]

> @@ -109,7 +108,14 @@ state:
>  			    struct tls_context *ctx,
>  			    enum tls_offload_ctx_dir direction);
>  
> -``tls_dev_del`` is mandatory whenever ``tls_dev_add`` is provided.
> +``tls_dev_del`` is called either when the offloaded connection is destroyed or,
> +for a TLS 1.3 connection, when the old key is retired during a rekey (see the
> +`Rekey`_ section). It operates on a single ``direction``, so the driver must

[Severity: Low]

This reads as an exhaustive "either ... or" list of the call sites. Is the
device teardown case missing? tls_device_down() calls the callback for each
live direction of every offloaded socket on the disappearing netdev:

	if (ctx->tx_conf == TLS_HW &&
	    !test_bit(TLS_TX_DEV_CLOSED, &ctx->flags)) {
		netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
						TLS_OFFLOAD_CTX_DIR_TX);
	...
	if (ctx->rx_conf == TLS_HW &&
	    !test_bit(TLS_RX_DEV_CLOSED, &ctx->flags)) {
		netdev->tlsdev_ops->tls_dev_del(netdev, ctx,
						TLS_OFFLOAD_CTX_DIR_RX);

Those connections are neither being destroyed nor rekeying, so should the
NETDEV_DOWN path be mentioned alongside the other two?

> +release only the state for that direction and must not free state shared
> +between directions or the socket as a whole. After a rekey ``tls_dev_del``,
> +``tls_dev_add`` may be called again for the same socket and direction to
> +install the new key. ``tls_dev_del`` is mandatory whenever ``tls_dev_add`` is
> +provided.

[ ... ]

> @@ -404,8 +413,121 @@ records, then after 4 records, after 8, after 16... up until every
>  Rekey
>  =====
>  
> -Offload does not currently support TLS 1.3, therefore key rotation
> -is not a concern for offloaded connections at this point.
> +TLS 1.3 allows traffic keys to be updated mid-connection using the
> +KeyUpdate message. Offloaded TLS 1.3 connections must therefore switch
> +keys without tearing down the offload. The device cannot simply be given
> +the new key because records encrypted (TX) or transformed (RX) with the
> +old key may still be in flight. The stack retains the necessary old-key
> +state and bridges the transition in software.
> +
> +TX
> +--
> +
> +On TX, the new key is installed in a temporary software context, and
> +sendmsg is routed through the software path. If no hardware-offloaded
> +records remain unacknowledged, the switch completes inline during
> +setsockopt. Otherwise the rekey is left pending and is completed later,

[Severity: Low]

Is unacknowledged old-key data really the only reason for deferring?
tls_set_device_offload_rekey() starts from defer = true and only narrows it
in one case:

	bool rekey_pending = test_bit(TLS_TX_REKEY_PENDING, &ctx->flags);
	bool rekey_failed = test_bit(TLS_TX_REKEY_FAILED, &ctx->flags);
	bool defer = true;
	...
	if (!rekey_pending && !rekey_failed)
		defer = tls_has_unacked_records(offload_ctx) ||
			tls_is_pending_open_record(ctx);

So a socket with a buffered open hardware record (MSG_MORE / cork) but
nothing unacknowledged also takes the deferred path, and a socket that is
already pending or has previously failed always does. Should those two
cases be listed here?

Relatedly, the new TlsCurrTxRekey text in tls.rst says such sessions are
"waiting for in-flight records to drain". In the re-armed case after a
hardware fallback, tls_device_start_rekey() does:

			WRITE_ONCE(ctx->rekey.boundary_seq, tcp_sk(sk)->snd_una);
			set_bit(TLS_TX_REKEY_PENDING, &ctx->flags);

which leaves nothing draining, yet the session is counted in the gauge.

[ ... ]

>  
>  Error handling
>  ==============
> @@ -442,8 +564,43 @@ to the host's stack as it was on the wire (recovering original packet in the
>  driver if device provides precise error is sufficient).
>  
>  The Linux networking stack does not provide a way of reporting per-packet
> -decryption and authentication errors, packets with errors must simply not
> -have the :c:member:`decrypted` mark set.
> +decryption and authentication errors. A packet with errors must not have
> +the :c:member:`decrypted` mark set. In addition, the driver may set the
> +:c:member:`decrypt_failed` mark on a segment the device matched to an
> +offloaded connection and processed but could not authenticate. The two
> +marks are mutually exclusive.

[Severity: Low]

These two sentences seem to contradict each other. The first says there is
no way to report per-packet decryption and authentication errors, and the
next one introduces decrypt_failed, which include/linux/skbuff.h describes
as "hardware could not authenticate this skb's TLS payload", i.e. exactly a
per-segment authentication failure indication.

Could the text spell out the distinction it intends, for instance that
decrypt_failed is an internal hint about what the device did to the payload
rather than an error reported to the application?

> +
> +The stack interprets :c:member:`decrypt_failed` per record, relative to the
> +:c:member:`decrypted` mark of the other segments making up the same record.
> +Coalescing (GRO, socket layer) and record classification are keyed on
> +:c:member:`decrypted` alone, so :c:member:`decrypt_failed` segments may be
> +merged with unmarked ones. A driver setting the mark must therefore honour
> +the following contract:

[Severity: Medium]

If a decrypt_failed segment can be merged with an unmarked one, how does the
per-segment contract survive the merge?

skb_cmp_decrypted() compares only the decrypted bit:

include/linux/skbuff.h:
	return skb2->decrypted - skb1->decrypted;

and both coalescing paths gate on it alone:

net/ipv4/tcp_offload.c:tcp_gro_receive()
	flush |= skb_cmp_decrypted(p, skb);

include/net/tcp.h:tcp_skb_can_collapse_rx()
	return likely(mptcp_skb_can_collapse(to, from) &&
		      !skb_cmp_decrypted(to, from));

used by tcp_try_coalesce() -> skb_try_coalesce(), which copies the payload
without carrying the source skb's decrypt_failed bit.

After such a merge one skb can hold both device-XORed bytes and untouched
wire ciphertext, while tls_device_reencrypt() decides per head skb and per
fraglist entry:

		if (skb->decrypted || skb->decrypt_failed) {
			err = skb_store_bits(skb, offset, buf, copy);
	...
		if (skb_iter->decrypted || skb_iter->decrypt_failed) {
			err = skb_store_bits(skb_iter, frag_pos, buf, copy);

Would the merged payload then be either entirely re-encrypted or entirely
skipped, so the record fails software authentication even though the driver
followed the documented contract?

> +
> + * In a record none of whose segments carry :c:member:`decrypted`, every
> +   segment, including one with :c:member:`decrypt_failed` set, must hold
> +   the payload exactly as it was on the wire. This is the general rule
> +   above: if the device did not successfully decrypt any part of a record
> +   it must hand the whole record over untouched. The stack passes such a
> +   record to software decryption directly and does not consult
> +   :c:member:`decrypt_failed`.

[Severity: Medium]

Is "does not consult :c:member:`decrypt_failed`" accurate for this case?
tls_device_decrypted() reads the mark whenever the old RX AEAD is still
held, including when the record is fully encrypted:

net/tls/tls_device.c:tls_device_decrypted()
	if (unlikely(ctx->rekey.old_aead_recv)) {
		bool nic_touched = !is_encrypted || skb->decrypt_failed;
	...
		if (nic_touched &&
		    !before(rec_start_seq, ctx->rekey.old_nic_boundary))
			ctx->rekey.old_nic_boundary = rec_start_seq + rxm->full_len;

Each such record slides old_nic_boundary past itself, so the
crypto_free_aead(ctx->rekey.old_aead_recv) arm and the deferred
tls_dev_add() of the new RX key are not reached while records with the mark
keep arriving.

Does that mean a driver which follows this bullet and sets the mark on
untransformed records can keep an RX socket in software decryption
indefinitely? The record itself still decrypts, but offload is never
reinstalled.

> +
> + * In a record where some segments carry :c:member:`decrypted` (a mixed
> +   record), a segment with :c:member:`decrypt_failed` set must hold payload
> +   the device has already transformed (XORed with the cipher keystream) but
> +   failed to authenticate, and a segment with neither mark must hold the
> +   payload as it was on the wire. The stack re-encrypts the
> +   :c:member:`decrypted` and :c:member:`decrypt_failed` segments to restore
> +   the ciphertext, leaves the unmarked segments intact, and authenticates
> +   the whole record in software.
> +
> +A transformed segment delivered without :c:member:`decrypt_failed`, or an
> +untransformed segment of a mixed record delivered with it, is restored
> +incorrectly and the record fails software authentication. A device which
> +cannot tell the driver whether a failed segment was transformed must
> +recover the original packet before handing it to the stack, as described
> +above, and leave both marks clear. During a TLS 1.3 rekey the mark also
> +tells the stack which key the device applied; see the `Rekey`_ section.

[Severity: Medium]

Can the only in-tree producer of the mark satisfy this contract? mlx5 sets
it for every TLS offload error CQE, and its own comment says the
transformation state is not known:

drivers/net/ethernet/mellanox/mlx5/core/en_accel/ktls_rx.c:
	case CQE_TLS_OFFLOAD_ERROR:
		/* The device could not authenticate the payload. Depending on
		 * where the failure occurred the bytes may have been transformed
		 * (XORed) or left as wire ciphertext. ...
		 */
		skb->decrypt_failed = 1;

That looks like the case this paragraph says must leave both marks clear.
For a non-mixed record the stack then treats the payload as untouched wire
ciphertext:

net/tls/tls_device.c:tls_device_decrypted()
			if (is_encrypted) {
				tls_bigint_increment(ctx->rekey.old_rec_seq,
						     tls_ctx->prot_info.rec_seq_size);
				return 0;
			}

and for a mixed record tls_device_reencrypt() rewrites the marked fragment.
Either way one of the two possibilities the mlx5 comment admits is restored
incorrectly.

Should the documented contract be relaxed to match what ConnectX error CQEs
actually express, or should the driver stop setting the mark?

>  
>  A packet should also not be handled by the TLS offload if it contains
>  incorrect checksums.
> diff --git a/Documentation/networking/tls.rst b/Documentation/networking/tls.rst
> index 980c442d7161..cf05543260d8 100644
> --- a/Documentation/networking/tls.rst
> +++ b/Documentation/networking/tls.rst
> @@ -314,6 +314,11 @@ TLS implementation exposes the following per-namespace statistics
>    number of TX and RX sessions currently installed where NIC handles
>    cryptography
>  
> +- ``TlsCurrTxRekey``, ``TlsCurrRxRekey`` -
> +  number of TX and RX sessions currently undergoing a deferred rekey,
> +  i.e. a rekey which could not be applied immediately and is waiting for
> +  in-flight records to drain before the new key is installed in hardware

[Severity: Low]

Two questions about how these descriptions match the increment and
decrement sites.

Does the TX gauge really only cover sessions waiting for in-flight records
to drain? When the boundary is acknowledged, tls_tcp_clean_acked() only
arms completion:

		if (!before(acked_seq, boundary_seq))
			set_bit(TLS_TX_REKEY_READY, &tls_ctx->flags);

TLS_TX_REKEY_PENDING and the gauge stay set until a later sendmsg() runs
tls_device_complete_rekey(), which is where the decrement lives:

	if (deferred)
		TLS_DEC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
	TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYOK);

The re-armed-after-fallback case has nothing draining either, as noted
above. The tls-offload.rst text added by this patch says as much ("A
connection that stops sending after a KeyUpdate stays in the deferred state
until it is closed ... it is counted in ``TlsCurrTxRekey``"), which reads
differently from the wording here.

Second, should the unchanged TlsTxRekeyOk entry ("number of successful
rekeys on existing sessions for TX and RX") be updated as well?
tls_set_device_offload_rekey() bumps it when a new KeyUpdate supersedes a
still-pending deferred rekey:

	if (defer) {
		if (!rekey_pending)
			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSCURRTXREKEY);
		else
			TLS_INC_STATS(sock_net(sk), LINUX_MIB_TLSTXREKEYOK);
		return 0;
	}

In that branch no tls_dev_del()/tls_dev_add() ran for the superseded key.

> +
>  - ``TlsTxSw``, ``TlsRxSw`` -
>    number of TX and RX sessions opened with host cryptography
>  

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917224355.2288021-1-rjethwani%40purestorage.com

^ permalink raw reply	[flat|nested] 29+ messages in thread

end of thread, other threads:[~2026-09-22  1:56 UTC | newest]

Thread overview: 29+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 22:35 [PATCH net-next v17 00/15] tls: Add TLS 1.3 hardware offload support Rishikesh Jethwani
2026-09-17 22:35 ` [PATCH net-next v17 01/15] net: tls: reject TLS 1.3 offload in chcr_ktls and nfp drivers Rishikesh Jethwani
2026-09-22  1:55   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 02/15] net/mlx5e: add TLS 1.3 hardware offload support Rishikesh Jethwani
2026-09-22  1:55   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 03/15] tls: reject rekey attempts on an existing HW-offloaded connection Rishikesh Jethwani
2026-09-17 22:35 ` [PATCH net-next v17 04/15] tls: add TLS 1.3 hardware offload support Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 05/15] tls: split tls_set_sw_offload into init and finalize stages Rishikesh Jethwani
2026-09-17 22:35 ` [PATCH net-next v17 06/15] tls: prep helpers and refactors for HW offload KeyUpdate Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 07/15] net: sched: re-validate parked decrypted skbs on requeue Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 08/15] tcp: fence collapse against rtx-queue tail when write queue is empty Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 09/15] net: skbuff: add skb->decrypt_failed bit Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 10/15] net/mlx5e: flag TLS RX records that failed device decryption Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 11/15] tls: device: add TX KeyUpdate support Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 12/15] tls: device: add RX " Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 13/15] tls: device: add tracepoints for the KeyUpdate path Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 14/15] selftests: net: add TLS hardware offload test Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko
2026-09-17 22:35 ` [PATCH net-next v17 15/15] tls: document TLS 1.3 hardware offload rekey handling Rishikesh Jethwani
2026-09-22  1:56   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox