From: Jerome Tollet <jtollet@cisco.com>
To: spdk@lists.linux.dev
Subject: [SPDK] RFC: dynamic NVMf/TCP receive vectors for socket RX zero-copy
Date: Fri, 04 Sep 2026 21:20:43 +0200 [thread overview]
Message-ID: <e83f73bf79fa3966df4ca1cf96590b4f@cisco.com> (raw)
Hi,
I am experimenting with the SPDK socket RX zero-copy review series
27033, 27037, 27048 and 27049 using an architecture where SPDK is embedded
inside VPP as an in-process plugin. The SPDK socket backend uses the VPP
HostStack instead of POSIX sockets.
The initial architecture and BlueField-3 results are described here:
https://medium.com/fd-io-vpp/spdk-inside-vpp-accelerating-nvme-tcp-on-bluefield-3-bc0419048b1e
The current work extends that integration with RX zero-copy.
The VPP HostStack backend can expose a received TCP payload as many separate
segments. With the existing inline limits, fragmented NVMe/TCP payloads can
exceed the 33 request iovecs or 16 PDU iovecs and fall back to copying.
A simple initial workaround expanded the inline arrays to a full 256-entry
receive vector. However, this substantially increases the permanent size of
every request and PDU object, while the current uint8_t iovec count still
limits the usable vector to 255 segments.
I posted three dependent changes:
- 29255 fixes two correctness issues in the copied fallback: it preserves
the number of bytes already consumed and allocates the materialized buffer
with SPDK_MALLOC_DMA.
https://review.spdk.io/c/spdk/spdk/+/29255
- 29257 lets an NVMf transport provide an extended request iovec while
keeping the existing 33-entry inline array as the default.
https://review.spdk.io/c/spdk/spdk/+/29257
- 29258 makes NVMf/TCP grow its request, segment-reference and PDU vectors
geometrically and only when their inline capacity is exceeded. Expanded
vectors are retained for object reuse, while the copied fallback is
preserved for allocation failures or the uint8_t iovec-count limit.
https://review.spdk.io/c/spdk/spdk/+/29258
On BlueField-3, the integration delivered almost 100% of RX payload bytes
directly to SPDK. In a controlled 128 KiB write test, throughput increased
from 3.104 to 4.804 GB/s compared with the copied path.
A fixed-256 versus dynamic-vector comparison showed essentially identical
throughput at 128 KiB and a 1.05% improvement at 1 MiB, while the dynamic
implementation keeps the normal object footprint close to the original
sizes.
I would particularly appreciate feedback on whether the optional extended
request vector belongs in the generic NVMf layer, and whether retaining
expanded allocations for object reuse is the preferred lifecycle.
Best regards,
Jerome
reply other threads:[~2026-09-04 19:21 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e83f73bf79fa3966df4ca1cf96590b4f@cisco.com \
--to=jtollet@cisco.com \
--cc=spdk@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox