From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm2-f12.google.com (mail-wm2-f12.google.com [74.125.225.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 692093E1230 for ; Wed, 30 Sep 2026 19:21:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.140 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790796089; cv=none; b=KlYR7I7/7Go/eJOC9dsAb9MURJk7RX6ltTh5LmAhN02+w5CAqcXOtDO0C7o0tYxgqqewFbZNMYHpeBPIIbNQphj54+HoBX0TG1abv31nlQgohgqEdrvp/2rZQ9Rlkid54Icx4kwbivAy/n7kjD2P2UAG7YSwC/30XnJKFSqIzZk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790796089; c=relaxed/simple; bh=MyorId+LNXDQ3tBXVrKOG34/XwgfgE8q2c/oj8aOYQ8=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=L+JaZYrz8Ez0AdYllyqlQg/fagkv/16JVY9eFberlq1DBsxGdg6ZpFgoEdRZxGSqJ2B/bFFXP4j5M4dV/te4ZgN1xbBNlI3LGewHvs+FrTAjgzQPfk2NObVQrPg08Fcp9abdfd4y94YxqqeKkxrLUz48ZcDRIZGnD0mUFnZvBAk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=d1YyOr6j; arc=none smtp.client-ip=74.125.225.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="d1YyOr6j" Received: by mail-wm2-f12.google.com with SMTP id 5b1f17b1804b1-49cd5462b69so32432505e9.1 for ; Wed, 30 Sep 2026 12:21:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790796085; x=1791400885; darn=lists.linux.dev; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=nf3oV330ckZwsw/LTGlRTnwET/Uw246suYw2VNp67Og=; b=d1YyOr6jyq5jT1bqG3icpXINOYaLTZWk1Wlty6CNXFb4s1Cfb48+5+rDkGEB4v4Bo3 h1Lnxu5avy221YFv+ll8q/wieDfv/PC2zZfO2ocWxNmnKwcAl/erNFHr1rb2gT9rDJWY LPFU9ZXS2AnB0rICp1mndn/Pohwoz+vqnryW56OAtfMNy/byCIK3Fk/vlWt3vw26mr+j v0JrSIK58y+9yY7Rs+BKWUqdwUbXYmxyeFEZ24T8YV4fUUakrM34TKtU1Txdp9olPv7y 3mbsKrmUQ+la1mUv5+tgyMKGi4RBdjCL3GUPGRadoXrkPTOM/u5RSdvFKa2Cf2H0SFep HlAQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790796085; x=1791400885; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=nf3oV330ckZwsw/LTGlRTnwET/Uw246suYw2VNp67Og=; b=pxw2g+orrKRIcDoRyA821b6gbJQjuFQXrKK3sXszZEg3jJwkpibRt08ANmg6pTCCzF BiLXXF6I/5iRsTEMeY4XzP6LCkKMpfepAXELzWk1Hii/oiy5JzDvIGt8aBO1PsmEBXGf DQfZjPPqofAuGO/SLgsrW48w7SxGj2Qvvu5IJOeC68KBm6Qt+wM9xF+SIwGVDulfFSgj LTnOTeLKHaqlYscuo9A5OJSELvkMz8DIHdvsoiWqy8OnPPqMFO+OdQr4iBSd+HgXB5X3 zMsIlV6t3Dtkcm3HFvbiKSiKSVStFkudGvdA5Hr2ZipWHwTiNX7YqfsyLh4GnfEWl/0J jwNA== X-Forwarded-Encrypted: i=1; AKwUvByYZ/mqSJnEFfOxZjxthhWUh+Fs52R4NuKWZGI5doG/5D8u+SOSy4PKBtdXtYSonYT8eVsr6BwfPD8LbzWFpQ==@lists.linux.dev X-Gm-Message-State: AFuF++lRJH1ggen3ww7P67Do/m5FVNmCES9cVK+XLpTYjc47kj7ZXYL6 3Mrr8aM2zOYXjRUap904tFg3esEmIyMyhDBqgw7DqRzs+kT6pXiqmAE3 X-Gm-Gg: AYBFou23DxlVw2XKI9ELi4qrJqaCb2+BGUIuWETGQUOZQxRhqUgGav3Ok1oK8+lgWSx CFzOKh+NBdDWRYKdJwsFsBgT6c2zFpiZOrQqdQrL6xOF0aU6GG3fa1PJRqUxFVMUbPnR4iaoxVl ZxaFgTBQx1TCADqmKn8g/U8Ot4+W8px5L3CsvLKgYk0m9RABAJI/CDk7VbBn0nvvOGREQgISosC SopkNvSIt1czNNA14iFFZCBE7elwClRUln8dqaWOxgv38JU1b/DNk7OPM9qdLFqdseQ/xN29/QZ Sx80CpAnTBi1nxe8DTlSMJhpmP0qgga+fzznXlca/mwjFavhkX0UJ5x2Jl+g9rACrUaN/JabPWI SgvKLfrVho4ZtyOiYhf2Kat+QlDkj8UM6hdPOOLFCMZONYF4X1Z2jDPeoWagEQcy9HYq/3doMHg aM9m5JCdhdjQyDNT0IsN1G9kzuPnz81TaCbZuGoiGcAffWsPPrO5tQ73zoWCCHsXR4F+UJyBstP fwDqJMjRLgILKa3RavXEtBxh6g651Z+Db11Ceeb/5bizkkl8JwzvD0gfeWHSPNcHJVKxd8eB+D7 aeatUzL1GWDRs9dcThkGfEvm1NFeXVRBANuojbGx9lRf315NE3tqRn4e9YasY9Hrg0tu7EV4lv3 J X-Received: by 2002:a05:600c:37c4:b0:49f:ce78:356b with SMTP id 5b1f17b1804b1-4a01b125598mr35582525e9.28.1790796084483; Wed, 30 Sep 2026 12:21:24 -0700 (PDT) Received: from localhost.localdomain (dynamic-2a02-3100-b2b3-0101-b51c-a607-3dcc-3979.310.pool.telefonica.de. [2a02:3100:b2b3:101:b51c:a607:3dcc:3979]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4a01f9248ccsm3513165e9.3.2026.09.30.12.21.23 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Wed, 30 Sep 2026 12:21:23 -0700 (PDT) From: Karl Mehltretter To: netdev@vger.kernel.org Cc: Karl Mehltretter , "David S. Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt , Stephen Hemminger , Arend van Spriel , linux-kernel@vger.kernel.org, linux-rt-devel@lists.linux.dev, stable@vger.kernel.org Subject: [PATCH net v2 1/2] netpoll: use a raw lock for the deferred transmit queue Date: Wed, 30 Sep 2026 21:21:02 +0200 Message-Id: <20260930192103.62973-2-kmehltretter@gmail.com> X-Mailer: git-send-email 2.39.5 (Apple Git-154) In-Reply-To: <20260928064239.32456-1-kmehltretter@gmail.com> References: <20260928064239.32456-1-kmehltretter@gmail.com> Precedence: bulk X-Mailing-List: linux-rt-devel@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit netpoll_send_skb() calls __netpoll_send_skb() with hard interrupts disabled. When direct transmission cannot complete, the latter queues the skb with skb_queue_tail(). The sk_buff_head lock may sleep on PREEMPT_RT: BUG: sleeping function called from invalid context in_atomic(): 0, irqs_disabled(): 1, non_block: 0 rt_spin_lock skb_queue_tail netpoll_send_skb The delayed transmit worker has the same problem when it requeues a busy skb with skb_queue_head() after disabling interrupts. Add a dedicated raw spinlock and use the unlocked skb queue helpers under it. Keep raw critical sections limited to queue operations. During cleanup, splice the queue to a private list before freeing its skbs. Fixes: b6cd27ed3388 ("netpoll per device txq") Cc: stable@vger.kernel.org # 6.12+ Assisted-by: LLM Signed-off-by: Karl Mehltretter --- include/linux/netpoll.h | 1 + net/core/netpoll.c | 75 +++++++++++++++++++++++++++++++++++++---- 2 files changed, 70 insertions(+), 6 deletions(-) diff --git a/include/linux/netpoll.h b/include/linux/netpoll.h index 1c6b1eec5efd6..e20e0592e9349 100644 --- a/include/linux/netpoll.h +++ b/include/linux/netpoll.h @@ -47,6 +47,7 @@ struct netpoll_info { struct semaphore dev_lock; struct sk_buff_head txq; + raw_spinlock_t txq_lock; struct delayed_work tx_work; diff --git a/net/core/netpoll.c b/net/core/netpoll.c index fe1e0cda5d6bf..e0cfcb05468e2 100644 --- a/net/core/netpoll.c +++ b/net/core/netpoll.c @@ -79,6 +79,68 @@ static netdev_tx_t netpoll_start_xmit(struct sk_buff *skb, return status; } +/* + * Transmit paths can access txq with hard IRQs disabled. Use a raw lock + * because the skb queue lock may sleep on PREEMPT_RT. + */ +static bool netpoll_txq_empty(struct netpoll_info *npinfo) +{ + unsigned long flags; + bool empty; + + raw_spin_lock_irqsave(&npinfo->txq_lock, flags); + empty = skb_queue_empty(&npinfo->txq); + raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags); + + return empty; +} + +static struct sk_buff *netpoll_txq_dequeue(struct netpoll_info *npinfo) +{ + unsigned long flags; + struct sk_buff *skb; + + raw_spin_lock_irqsave(&npinfo->txq_lock, flags); + skb = __skb_dequeue(&npinfo->txq); + raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags); + + return skb; +} + +static void netpoll_txq_queue_head(struct netpoll_info *npinfo, + struct sk_buff *skb) +{ + unsigned long flags; + + raw_spin_lock_irqsave(&npinfo->txq_lock, flags); + __skb_queue_head(&npinfo->txq, skb); + raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags); +} + +static void netpoll_txq_queue_tail(struct netpoll_info *npinfo, + struct sk_buff *skb) +{ + unsigned long flags; + + raw_spin_lock_irqsave(&npinfo->txq_lock, flags); + __skb_queue_tail(&npinfo->txq, skb); + raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags); +} + +static void netpoll_txq_purge(struct netpoll_info *npinfo) +{ + struct sk_buff_head purge; + unsigned long flags; + + __skb_queue_head_init(&purge); + + raw_spin_lock_irqsave(&npinfo->txq_lock, flags); + skb_queue_splice_init(&npinfo->txq, &purge); + raw_spin_unlock_irqrestore(&npinfo->txq_lock, flags); + + __skb_queue_purge(&purge); +} + static void queue_process(struct work_struct *work) { struct netpoll_info *npinfo = @@ -86,7 +148,7 @@ static void queue_process(struct work_struct *work) struct sk_buff *skb; unsigned long flags; - while ((skb = skb_dequeue(&npinfo->txq))) { + while ((skb = netpoll_txq_dequeue(npinfo))) { struct net_device *dev = skb->dev; struct netdev_queue *txq; unsigned int q_index; @@ -107,7 +169,7 @@ static void queue_process(struct work_struct *work) HARD_TX_LOCK(dev, txq, smp_processor_id()); if (netif_xmit_frozen_or_stopped(txq) || !dev_xmit_complete(netpoll_start_xmit(skb, dev, txq))) { - skb_queue_head(&npinfo->txq, skb); + netpoll_txq_queue_head(npinfo, skb); HARD_TX_UNLOCK(dev, txq); local_irq_restore(flags); @@ -282,7 +344,7 @@ static netdev_tx_t __netpoll_send_skb(struct netpoll *np, struct sk_buff *skb) } /* don't get messages out of order, and no recursion */ - if (skb_queue_len(&npinfo->txq) == 0 && !netpoll_owner_active(dev)) { + if (netpoll_txq_empty(npinfo) && !netpoll_owner_active(dev)) { struct netdev_queue *txq; txq = netdev_core_pick_tx(dev, skb, NULL); @@ -314,7 +376,7 @@ static netdev_tx_t __netpoll_send_skb(struct netpoll *np, struct sk_buff *skb) } if (!dev_xmit_complete(status)) { - skb_queue_tail(&npinfo->txq, skb); + netpoll_txq_queue_tail(npinfo, skb); schedule_delayed_work(&npinfo->tx_work,0); } ret = NETDEV_TX_OK; @@ -362,7 +424,8 @@ int __netpoll_setup(struct netpoll *np, struct net_device *ndev) } sema_init(&npinfo->dev_lock, 1); - skb_queue_head_init(&npinfo->txq); + __skb_queue_head_init(&npinfo->txq); + raw_spin_lock_init(&npinfo->txq_lock); INIT_DELAYED_WORK(&npinfo->tx_work, queue_process); refcount_set(&npinfo->refcnt, 1); @@ -397,7 +460,7 @@ static void rcu_cleanup_netpoll_info(struct rcu_head *rcu_head) struct netpoll_info *npinfo = container_of(rcu_head, struct netpoll_info, rcu); - skb_queue_purge(&npinfo->txq); + netpoll_txq_purge(npinfo); kfree(npinfo); } -- 2.53.0