From: Simon Horman <horms@kernel.org>
To: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
Cc: netdev@vger.kernel.org, Haiyang Zhang <haiyangz@microsoft.com>,
Wei Liu <wei.liu@kernel.org>, Dexuan Cui <decui@microsoft.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
"David S. Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@google.com>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
Konstantin Taranov <kotaranov@microsoft.com>,
Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Jesper Dangaard Brouer <hawk@kernel.org>,
John Fastabend <john.fastabend@gmail.com>,
Erni Sri Satya Vennela <ernis@linux.microsoft.com>,
Aditya Garg <gargaditya@linux.microsoft.com>,
Dipayaan Roy <dipayanroy@linux.microsoft.com>,
Breno Leitao <leitao@debian.org>,
Jacob Keller <jacob.e.keller@intel.com>,
Saurabh Sengar <ssengar@linux.microsoft.com>,
linux-hyperv@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-rdma@vger.kernel.org, bpf@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH net v2] net: mana: reserve RX buffer headroom to fix forwarding performance
Date: Wed, 7 Oct 2026 16:09:11 +0100 [thread overview]
Message-ID: <20261007150911.GT83879@horms.kernel.org> (raw)
In-Reply-To: <20261003013647.2051416-1-hamzamahfooz@linux.microsoft.com>
On Fri, Oct 02, 2026 at 09:36:47PM -0400, Hamza Mahfooz wrote:
> Commit 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers
> instead of full pages to improve memory efficiency.") started handing
> out RX buffers with zero headroom so that two buffers fit into one page
> at the default MTU.
>
> The MANA TX path, however, stores the per scatter-gather entry DMA
> mappings in `struct mana_skb_head` at skb->head, and mana_start_xmit()
> therefore calls skb_cow_head(skb, MANA_HEADROOM). The port advertises
> this requirement as ndev->needed_headroom = MANA_HEADROOM.
>
> As a result every packet that is received and then forwarded out of a
> MANA port fails the skb_cow() in ip_forward() and gets reallocated and
> copied by pskb_expand_head(). This is invisible to a plain RX or TX
> workload, but it puts a full skb reallocation plus memcpy on the hot
> path of every single forwarded packet, which is exactly what a
> router/NVA workload does.
>
> Restore the headroom. Note that reserving MANA_HEADROOM (232) is not
> enough: ip_forward() asks for LL_RESERVED_SPACE(dev), which rounds
> hard_header_len + needed_headroom up to HH_DATA_MOD and is 256 bytes on
> ethernet. Use LL_RESERVED_SPACE() directly so the value keeps tracking
> both constants. Also, since LL_RESERVED_SPACE() tracks MANA_HEADROOM,
> it grows with MAX_SKB_FRAGS and for MAX_SKB_FRAGS >= 19 it is greater
> than 256, so we have to account for that by using the headroom the RX
> queue actually uses (instead of assuming XDP_PACKET_HEADROOM) and
> turning MANA_XDP_MTU_MAX into MANA_XDP_MTU_MAX(ndev) (note that at the
> default CONFIG_MAX_SKB_FRAGS=17 they are equivalent).
>
> At the default MTU on a 4K page this means a buffer no longer fits twice
> into a page (SKB_DATA_ALIGN(1500 + MANA_RXBUF_PAD + 256) = 2112), so the
> frag-vs-single decision is now made by computing the real buffer size
> instead of comparing the MTU against PAGE_SIZE / 2. The page_pool
> fragment path is still used wherever at least two buffers genuinely fit,
> e.g. on 16K and 64K page sizes.
>
> Measured on an Azure VM with a MANA NIC acting as a forwarding NVA (UDP,
> 1400 byte payload, 4 streams, 8 Gbps offered, only the forwarding
> node's kernel differs), 8 runs each, median:
>
> forwarded pps throughput
> before 272,830 3.06 Gbps
> after 390,560 4.37 Gbps (+43%)
>
> perf on the forwarding node, same workload:
>
> memset_orig __pi_memcpy pskb_expand_head
> before 10.07% 3.96% present
> after 0.94% 0.64% gone
>
> Cc: stable@vger.kernel.org
> Fixes: 730ff06d3f5c ("net: mana: Use page pool fragments for RX buffers instead of full pages to improve memory efficiency.")
> Signed-off-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
> ---
> v2:
> - Fix the XDP headroom mismatch with CONFIG_MAX_SKB_FRAGS >= 19,
> by passing rxq->headroom to xdp_prepare_buff().
> mana_build_skb() then picks up the correct offset via
> xdp->data - xdp->data_hard_start. Also, turn MANA_XDP_MTU_MAX
> into MANA_XDP_MTU_MAX(ndev) to account for the headroom,
> since it is no longer a constant. (Narcisa, Sashiko)
> - Use the new mana_single_rxbuf_per_page_forced() helper in
> mana_set_priv_flags(). (Sashiko)
> - Trim the comment above mana_get_rxbuf_headroom() and drop the stale
> "XDP headroom" wording from the comment above mana_get_rxbuf_cfg().
> (Narcisa, Sashiko)
Reviewed-by: Simon Horman <horms@kernel.org>
next prev parent reply other threads:[~2026-10-07 15:09 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-03 1:36 [PATCH net v2] net: mana: reserve RX buffer headroom to fix forwarding performance Hamza Mahfooz
2026-10-03 1:47 ` sashiko-bot
2026-10-07 15:09 ` Simon Horman [this message]
2026-10-08 17:50 ` patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261007150911.GT83879@horms.kernel.org \
--to=horms@kernel.org \
--cc=andrew+netdev@lunn.ch \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=decui@microsoft.com \
--cc=dipayanroy@linux.microsoft.com \
--cc=edumazet@google.com \
--cc=ernis@linux.microsoft.com \
--cc=gargaditya@linux.microsoft.com \
--cc=haiyangz@microsoft.com \
--cc=hamzamahfooz@linux.microsoft.com \
--cc=hawk@kernel.org \
--cc=jacob.e.keller@intel.com \
--cc=john.fastabend@gmail.com \
--cc=kotaranov@microsoft.com \
--cc=kuba@kernel.org \
--cc=leitao@debian.org \
--cc=linux-hyperv@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=ssengar@linux.microsoft.com \
--cc=stable@vger.kernel.org \
--cc=wei.liu@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox