* [PATCH v3 2/9] net/gve: delay adding mbuf head to software ring
From: Joshua Washington @ 2026-07-07 16:40 UTC (permalink / raw)
To: Jeroen de Borst, Joshua Washington, Xiaoyun Li, Junfeng Guo
Cc: dev, stable, Jasper Tran O'Leary
In-Reply-To: <20260707164020.2936476-1-joshwash@google.com>
The GQ TX datapath was set up to write the mbuf head into the sw_ring
before writing the descriptors. This poses a problem because it's
possible for the packet to be dropped due to lacking the FIFO space to
do a proper TX. In such a case, the packet won't be sent, and will lead
to leaked mbufs in the subsequent segments.
There is also no real reason that the head mbuf must be set in the
sw_ring separately from the others; the mbuf chain is not actually
walked as part of GQ TX.
Fixes: a46583cf43c8 ("net/gve: support Rx/Tx")
Cc: stable@dpdk.org
Signed-off-by: Joshua Washington <joshwash@google.com>
Reviewed-by: Jasper Tran O'Leary <jtranoleary@google.com>
---
drivers/net/gve/gve_tx.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/net/gve/gve_tx.c b/drivers/net/gve/gve_tx.c
index 5c73c21b8d..59c82b04ed 100644
--- a/drivers/net/gve/gve_tx.c
+++ b/drivers/net/gve/gve_tx.c
@@ -301,7 +301,6 @@ gve_tx_burst_qpl(void *tx_queue, struct rte_mbuf **tx_pkts, uint16_t nb_pkts)
(uint32_t)(tx_offload.l2_len + tx_offload.l3_len + tx_offload.l4_len) :
tx_pkt->pkt_len;
- sw_ring[sw_id] = tx_pkt;
if (!is_fifo_avail(txq, hlen)) {
gve_tx_clean(txq);
if (!is_fifo_avail(txq, hlen))
@@ -344,13 +343,14 @@ gve_tx_burst_qpl(void *tx_queue, struct rte_mbuf **tx_pkts, uint16_t nb_pkts)
}
/* record mbuf in sw_ring for free */
- for (i = 1; i < first->nb_segs; i++) {
+ for (i = 0; i < first->nb_segs; i++) {
+ if (!tx_pkt)
+ break;
+ sw_ring[sw_id] = tx_pkt;
sw_id = (sw_id + 1) & mask;
tx_pkt = tx_pkt->next;
- sw_ring[sw_id] = tx_pkt;
}
- sw_id = (sw_id + 1) & mask;
tx_id = (tx_id + 1) & mask;
txq->nb_free -= nb_used;
--
2.55.0.rc2.803.g1fd1e6609c-goog
^ permalink raw reply related
* [PATCH v3 1/9] net/gve: clear out shared memory region for stats report
From: Joshua Washington @ 2026-07-07 16:40 UTC (permalink / raw)
To: Jeroen de Borst, Joshua Washington, Ferruh Yigit, Rushil Gupta
Cc: dev, stable, Mark Blasko
In-Reply-To: <20260707164020.2936476-1-joshwash@google.com>
The stats report memzone is allocated from hugepage memory which could
possibly have had sensitive data from a previous DPDK invocation.
Clear out the buffer before sharing the memory region with the virtual
device to protect guest memory.
Fixes: 458b53dec01e ("net/gve: enable imissed stats for GQ format")
Cc: stable@dpdk.org
Signed-off-by: Joshua Washington <joshwash@google.com>
Reviewed-by: Mark Blasko <blasko@google.com>
---
drivers/net/gve/gve_ethdev.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/drivers/net/gve/gve_ethdev.c b/drivers/net/gve/gve_ethdev.c
index 0b02dcb3ad..f73784a109 100644
--- a/drivers/net/gve/gve_ethdev.c
+++ b/drivers/net/gve/gve_ethdev.c
@@ -306,6 +306,8 @@ gve_alloc_stats_report(struct gve_priv *priv,
if (!priv->stats_report_mem)
return -ENOMEM;
+ memset(priv->stats_report_mem->addr, 0, priv->stats_report_mem->len);
+
/* offset by skipping stats written by gve. */
priv->stats_start_idx = (GVE_TX_STATS_REPORT_NUM * nb_tx_queues) +
(GVE_RX_STATS_REPORT_NUM * nb_rx_queues);
--
2.55.0.rc2.803.g1fd1e6609c-goog
^ permalink raw reply related
* [PATCH v3 0/9] Stability fixes for GVE
From: Joshua Washington @ 2026-07-07 16:40 UTC (permalink / raw)
Cc: dev, Joshua Washington
In-Reply-To: <20260703131308.2507403-1-joshwash@google.com>
This patch series consists of mostly unrelated fixes in the GVE driver.
Joshua Washington (9):
net/gve: clear out shared memory region for stats report
net/gve: delay adding mbuf head to software ring
net/gve: copy data to QPL buffer when mbuf read does not
net/gve: validate buf ID before processing Rx packet
net/gve: set mbuf to null in software ring after use
net/gve: free ctx mbuf if packet dropped after first segment
net/gve: increase range of DMA memzone ids to 64 bits
net/gve: don't reset ring size bounds to default on reset
net/gve: restrict max ring size in GQ QPL to 2K
drivers/net/gve/base/gve_adminq.c | 12 ++++++--
drivers/net/gve/base/gve_osdep.h | 4 +--
drivers/net/gve/gve_ethdev.c | 8 ++++--
drivers/net/gve/gve_ethdev.h | 1 +
drivers/net/gve/gve_rx.c | 3 ++
drivers/net/gve/gve_rx_dqo.c | 6 ++++
drivers/net/gve/gve_tx.c | 46 ++++++++++++++++++-------------
7 files changed, 53 insertions(+), 27 deletions(-)
---
v3:
* Fix 32-bit complication issue
v2:
* Remove unused definition
--
2.55.0.rc2.803.g1fd1e6609c-goog
^ permalink raw reply
* Re: [PATCH v2 0/9] Stability fixes for GVE
From: Joshua Washington @ 2026-07-07 16:36 UTC (permalink / raw)
To: David Marchand; +Cc: Stephen Hemminger, dev
In-Reply-To: <CAJFAV8wWC9VbkDT0tqk+EkZGfxJm1weiV5u=HwQhOk3sC2L4eg@mail.gmail.com>
On Tue, Jul 7, 2026 at 12:10 PM David Marchand
<david.marchand@redhat.com> wrote:
>
> On Tue, 7 Jul 2026 at 18:02, Joshua Washington <joshwash@google.com> wrote:
> >
> > >
> > > Better but fails on 32 bit build during post-merge testing
> > > DPDK 26.07.0-rc2
> >
> > Will fix in v3. Is there a way I can test 32-bit compile using
> > test-meson-builds.sh? It doesn't seem to be covered there by default.
> >
>
> You'll need support in your toolchain, but this script does support
> 32-bit compilation.
>
> On debian for example, you'll need (at least):
> dpkg --add-architecture i386
> apt install -y gcc-multilib g++-multilib libnuma-dev:i386
Ah, thanks. I installed the required dependencies to build 32-bit
manually, and the script picked it up automatically as expected.
>
> If you have a personal github account, it is also possible to try your
> changes after enabling github actions in your dpdk fork.
>
>
> --
> David Marchand
>
^ permalink raw reply
* Re: [PATCH v4] dts: update test suite names to be clear and consistent
From: Luca Vizzarro @ 2026-07-07 16:36 UTC (permalink / raw)
To: Andrew Bailey; +Cc: patrickrobb1997, dev, lylavoie, knimoji, ahassick
In-Reply-To: <CABJ3N2VPwsAoDyY4x=k=JQoU74UkdbQQRAy2LQgw_W9r-q8awQ@mail.gmail.com>
On 07/07/2026 15:09, Andrew Bailey wrote:
> Hi Luca,
>
> I thought that I did this previously but tried again just to make
> sure, referring to "git mv". The commit is clean when I change the name
> of the files but when I squash my commit that actually changes the
> contents of the files it makes the overview incredibly messy as before.
> I tried doing the moves and changes in different commits and
> squashing them, as well as just doing them in all one commit with the
> same result. I did not submit two separate patches because then the
> first commit alone would break doc builds. Is there something else you
> would like me to try in order to make this cleaner?
Mmmmh, that's rather odd. You can try to reset two commits. So if your
worktree has the 2 commits (before squashing) do:
git reset --soft HEAD~2
This should keep all the changes of the two commits uncommitted. And
then you can fix the renaming here and then do a final commit.
Otherwise... if nothing else works, will accept v4.
Luca
^ permalink raw reply
* Re: [PATCH v2 0/9] Stability fixes for GVE
From: David Marchand @ 2026-07-07 16:10 UTC (permalink / raw)
To: Joshua Washington; +Cc: Stephen Hemminger, dev
In-Reply-To: <CALuQH+UE2kJdfbny0qLSxB+izFSPTQJXVCwDNbUotrcQL4LYhw@mail.gmail.com>
On Tue, 7 Jul 2026 at 18:02, Joshua Washington <joshwash@google.com> wrote:
>
> >
> > Better but fails on 32 bit build during post-merge testing
> > DPDK 26.07.0-rc2
>
> Will fix in v3. Is there a way I can test 32-bit compile using
> test-meson-builds.sh? It doesn't seem to be covered there by default.
>
You'll need support in your toolchain, but this script does support
32-bit compilation.
On debian for example, you'll need (at least):
dpkg --add-architecture i386
apt install -y gcc-multilib g++-multilib libnuma-dev:i386
If you have a personal github account, it is also possible to try your
changes after enabling github actions in your dpdk fork.
--
David Marchand
^ permalink raw reply
* Re: [PATCH v2 0/9] Stability fixes for GVE
From: Joshua Washington @ 2026-07-07 16:02 UTC (permalink / raw)
To: Stephen Hemminger; +Cc: dev
In-Reply-To: <20260703072721.18c5987a@phoenix.local>
>
> Better but fails on 32 bit build during post-merge testing
> DPDK 26.07.0-rc2
Will fix in v3. Is there a way I can test 32-bit compile using
test-meson-builds.sh? It doesn't seem to be covered there by default.
^ permalink raw reply
* [PATCH] vhost: fix null dereference in async packed dequeue
From: Anton Vanda @ 2026-07-07 13:50 UTC (permalink / raw)
To: Thomas Monjalon, Maxime Coquelin, Chenbo Xia, Yuan Wang,
Cheng Jiang
Cc: dev, stable, cheng1.jiang, Anton Vanda
In the batch path of the asynchronous packed ring dequeue, the address
of the virtio net header is obtained from vhost_iova_to_vva(), which
returns 0 when a guest-provided descriptor address cannot be fully
translated. The batch check only validates that the descriptor address
is non-zero and that the length is consistent. A malicious or buggy
guest could therefore trigger a NULL pointer dereference and crash the
vhost application (denial of service).
Check the translation result and leave the batch fast path with an error
on failure, so the single-packet path handles the invalid descriptor, as
is already done for the non-batch async dequeue path.
Perform the header translation before the DMA iovec setup so that the
early return cannot leave the async iterator state partially updated.
Fixes: c2fa52bf1e5d ("vhost: add batch dequeue in async vhost packed ring")
Cc: stable@dpdk.org
Signed-off-by: Anton Vanda <avanda@ptsecurity.com>
---
.mailmap | 1 +
lib/vhost/virtio_net.c | 26 +++++++++++++++++---------
2 files changed, 18 insertions(+), 9 deletions(-)
diff --git a/.mailmap b/.mailmap
index 05a55c0bd6..9215581a29 100644
--- a/.mailmap
+++ b/.mailmap
@@ -141,6 +141,7 @@ Antara Ganesh Kolar <antara.ganesh.kolar@intel.com>
Anthony Fee <anthonyx.fee@intel.com>
Anthony Harivel <aharivel@redhat.com>
Antonio Fischetti <antonio.fischetti@intel.com>
+Anton Vanda <avanda@ptsecurity.com>
Anup Prabhu <aprabhu@marvell.com>
Anupam Kapoor <anupam.kapoor@gmail.com>
Anurag Mandal <anurag.mandal@intel.com>
diff --git a/lib/vhost/virtio_net.c b/lib/vhost/virtio_net.c
index 0658b81de5..6528c06ea4 100644
--- a/lib/vhost/virtio_net.c
+++ b/lib/vhost/virtio_net.c
@@ -4039,6 +4039,23 @@ virtio_dev_tx_async_packed_batch(struct virtio_net *dev,
vhost_for_each_try_unroll(i, 0, PACKED_BATCH_SIZE)
rte_prefetch0((void *)(uintptr_t)desc_addrs[i]);
+ if (virtio_net_with_host_offload(dev)) {
+ vhost_for_each_try_unroll(i, 0, PACKED_BATCH_SIZE) {
+ desc_vva = vhost_iova_to_vva(dev, vq, desc_addrs[i],
+ &lens[i], VHOST_ACCESS_RO);
+ /*
+ * A malformed or unmapped guest descriptor makes the
+ * IOVA translation fail (returns 0). Bail out of the
+ * batch fast path so the single-packet path handles the
+ * error, instead of dereferencing a NULL header.
+ */
+ if (unlikely(!desc_vva))
+ return -1;
+ hdr = (struct virtio_net_hdr *)(uintptr_t)desc_vva;
+ pkts_info[slot_idx + i].nethdr = *hdr;
+ }
+ }
+
vhost_for_each_try_unroll(i, 0, PACKED_BATCH_SIZE) {
host_iova[i] = (void *)(uintptr_t)gpa_to_first_hpa(dev,
desc_addrs[i] + buf_offset, pkts[i]->pkt_len, &mapped_len[i]);
@@ -4053,15 +4070,6 @@ virtio_dev_tx_async_packed_batch(struct virtio_net *dev,
async->iter_idx++;
}
- if (virtio_net_with_host_offload(dev)) {
- vhost_for_each_try_unroll(i, 0, PACKED_BATCH_SIZE) {
- desc_vva = vhost_iova_to_vva(dev, vq, desc_addrs[i],
- &lens[i], VHOST_ACCESS_RO);
- hdr = (struct virtio_net_hdr *)(uintptr_t)desc_vva;
- pkts_info[slot_idx + i].nethdr = *hdr;
- }
- }
-
vq_inc_last_avail_packed(vq, PACKED_BATCH_SIZE);
vhost_async_shadow_dequeue_packed_batch(vq, ids);
--
2.51.0
^ permalink raw reply related
* Re: [PATCH] build: fix libpcap detection with unusable header
From: Bruce Richardson @ 2026-07-07 14:16 UTC (permalink / raw)
To: Maayan Kashani; +Cc: dev, rasland, stable, David Marchand, Stephen Hemminger
In-Reply-To: <20260707133028.236546-1-mkashani@nvidia.com>
On Tue, Jul 07, 2026 at 04:30:28PM +0300, Maayan Kashani wrote:
> The libpcap detection only checked that the library was present, that
> <pcap.h> could be included, and that a trivial program links. None of
> these fail when a distribution ships a libpcap-devel package whose
> <pcap.h> is empty (present but declaring nothing) while libpcap.so is
> installed: has_header() passes on an empty file and the link test does
> not reference any pcap symbol. RTE_HAS_LIBPCAP was then set and the
> pcap-dependent code (bpf_convert.c, port source/sink, dumpcap) failed
> to build with errors such as:
>
Why would a distribution ship a devel package with an empty header file for
pcap.h? Is that expected behaviour in some scenario that I'm unaware of,
because it seems strange to me?
/Bruce
> bpf_convert.c: error: invalid use of undefined type
> 'const struct bpf_insn'
>
> Require an actual pcap declaration (pcap_create) to be visible in the
> header before enabling libpcap support, so a broken or empty header
> disables the optional pcap features instead of breaking the build.
>
> Fixes: d6024c0a6757 ("build: cleanup libpcap dependent components")
> Cc: stable@dpdk.org
>
> Signed-off-by: Maayan Kashani <mkashani@nvidia.com>
> ---
> config/meson.build | 5 ++++-
> 1 file changed, 4 insertions(+), 1 deletion(-)
>
> diff --git a/config/meson.build b/config/meson.build
> index d7f5e55c18f..4bfb8535781 100644
> --- a/config/meson.build
> +++ b/config/meson.build
> @@ -290,7 +290,10 @@ if not pcap_dep.found()
> # pcap got a pkg-config file only in 1.9.0
> pcap_dep = cc.find_library(pcap_lib, required: false)
> endif
> -if (pcap_dep.found() and cc.has_header('pcap.h', dependencies: pcap_dep)
> +# has_header() passes even when pcap.h is empty; require a real declaration.
> +if (pcap_dep.found()
> + and cc.has_header_symbol('pcap.h', 'pcap_create', dependencies: pcap_dep,
> + args: '-D_GNU_SOURCE')
> and cc.links(min_c_code, dependencies: pcap_dep))
> dpdk_conf.set('RTE_HAS_LIBPCAP', 1)
> dpdk_extra_ldflags += '-l@0@'.format(pcap_lib)
> --
> 2.21.0
>
^ permalink raw reply
* Re: [PATCH v4] dts: update test suite names to be clear and consistent
From: Andrew Bailey @ 2026-07-07 14:09 UTC (permalink / raw)
To: Luca Vizzarro; +Cc: patrickrobb1997, dev, lylavoie, knimoji, ahassick
In-Reply-To: <CABJ3N2UPaeLyD7Mv1vtoQSvRJLK0MUfC0LHvvuz-kvjx1YKwcQ@mail.gmail.com>
[-- Attachment #1: Type: text/plain, Size: 613 bytes --]
Hi Luca,
I thought that I did this previously but tried again just to make sure,
referring to "git mv". The commit is clean when I change the name of the
files but when I squash my commit that actually changes the contents of the
files it makes the overview incredibly messy as before. I tried doing the
moves and changes in different commits and squashing them, as well as just
doing them in all one commit with the same result. I did not submit two
separate patches because then the first commit alone would break doc
builds. Is there something else you would like me to try in order to make
this cleaner?
[-- Attachment #2: Type: text/html, Size: 682 bytes --]
^ permalink raw reply
* [PATCH] net/nfp: fix UB in BAR size shift operations
From: Alexey Simakov @ 2026-07-07 14:01 UTC (permalink / raw)
To: Chaoyong He, Alejandro Lucero; +Cc: dev, stable, Alexey Simakov
The literal '1' is a signed 32-bit int. Shifting it by bar->bitsize
is undefined behavior when bitsize >= 31, and sign-extends when
bitsize == 31 (producing a wrong upper-bound check). BAR aperture
sizes from hardware can exceed this range.
Fix by using RTE_BIT64() which produces a 64-bit unsigned value,
matching the type of the operands (uint64_t base, uint64_t offset).
Fixes: c7e9729da6b5 ("net/nfp: support CPP")
Fixes: 1fbe51cd9c3a ("net/nfp: extend usage of BAR from 8 to 24")
Cc: stable@dpdk.org
Signed-off-by: Alexey Simakov <bigalex934@gmail.com>
---
drivers/net/nfp/nfpcore/nfp6000_pcie.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/net/nfp/nfpcore/nfp6000_pcie.c b/drivers/net/nfp/nfpcore/nfp6000_pcie.c
index 83b7116097..e03c4776d0 100644
--- a/drivers/net/nfp/nfpcore/nfp6000_pcie.c
+++ b/drivers/net/nfp/nfpcore/nfp6000_pcie.c
@@ -55,11 +55,11 @@
*/
#define NFP_PCI_MIN_MAP_SIZE 0x080000 /* 512K */
-#define NFP_PCIE_P2C_FIXED_SIZE(bar) (1 << (bar)->bitsize)
-#define NFP_PCIE_P2C_BULK_SIZE(bar) (1 << (bar)->bitsize)
+#define NFP_PCIE_P2C_FIXED_SIZE(bar) RTE_BIT64((bar)->bitsize)
+#define NFP_PCIE_P2C_BULK_SIZE(bar) RTE_BIT64((bar)->bitsize)
#define NFP_PCIE_P2C_GENERAL_TARGET_OFFSET(bar, x) ((x) << ((bar)->bitsize - 2))
#define NFP_PCIE_P2C_GENERAL_TOKEN_OFFSET(bar, x) ((x) << ((bar)->bitsize - 4))
-#define NFP_PCIE_P2C_GENERAL_SIZE(bar) (1 << ((bar)->bitsize - 4))
+#define NFP_PCIE_P2C_GENERAL_SIZE(bar) RTE_BIT64(((bar)->bitsize - 4))
#define NFP_PCIE_P2C_EXPBAR_OFFSET(bar_index) ((bar_index) * 4)
@@ -443,7 +443,7 @@ matching_bar_exist(struct nfp_bar *bar,
(bar_token < 0 || bar_token == token) &&
bar_action == action &&
bar->base <= offset &&
- (bar->base + (1 << bar->bitsize)) >= (offset + size))
+ (bar->base + RTE_BIT64(bar->bitsize)) >= (offset + size))
return true;
/* No match */
--
2.34.1
^ permalink raw reply related
* [PATCH] build: fix libpcap detection with unusable header
From: Maayan Kashani @ 2026-07-07 13:30 UTC (permalink / raw)
To: dev
Cc: mkashani, rasland, stable, Bruce Richardson, David Marchand,
Stephen Hemminger
The libpcap detection only checked that the library was present, that
<pcap.h> could be included, and that a trivial program links. None of
these fail when a distribution ships a libpcap-devel package whose
<pcap.h> is empty (present but declaring nothing) while libpcap.so is
installed: has_header() passes on an empty file and the link test does
not reference any pcap symbol. RTE_HAS_LIBPCAP was then set and the
pcap-dependent code (bpf_convert.c, port source/sink, dumpcap) failed
to build with errors such as:
bpf_convert.c: error: invalid use of undefined type
'const struct bpf_insn'
Require an actual pcap declaration (pcap_create) to be visible in the
header before enabling libpcap support, so a broken or empty header
disables the optional pcap features instead of breaking the build.
Fixes: d6024c0a6757 ("build: cleanup libpcap dependent components")
Cc: stable@dpdk.org
Signed-off-by: Maayan Kashani <mkashani@nvidia.com>
---
config/meson.build | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/config/meson.build b/config/meson.build
index d7f5e55c18f..4bfb8535781 100644
--- a/config/meson.build
+++ b/config/meson.build
@@ -290,7 +290,10 @@ if not pcap_dep.found()
# pcap got a pkg-config file only in 1.9.0
pcap_dep = cc.find_library(pcap_lib, required: false)
endif
-if (pcap_dep.found() and cc.has_header('pcap.h', dependencies: pcap_dep)
+# has_header() passes even when pcap.h is empty; require a real declaration.
+if (pcap_dep.found()
+ and cc.has_header_symbol('pcap.h', 'pcap_create', dependencies: pcap_dep,
+ args: '-D_GNU_SOURCE')
and cc.links(min_c_code, dependencies: pcap_dep))
dpdk_conf.set('RTE_HAS_LIBPCAP', 1)
dpdk_extra_ldflags += '-l@0@'.format(pcap_lib)
--
2.21.0
^ permalink raw reply related
* Re: [PATCH v4] dts: update test suite names to be clear and consistent
From: Andrew Bailey @ 2026-07-07 13:07 UTC (permalink / raw)
To: Luca Vizzarro; +Cc: patrickrobb1997, dev, lylavoie, knimoji, ahassick
In-Reply-To: <e135faa6-020a-42f1-a09f-e21e90dde010@arm.com>
[-- Attachment #1: Type: text/plain, Size: 49 bytes --]
Not a problem, I will send a new version shortly
[-- Attachment #2: Type: text/html, Size: 91 bytes --]
^ permalink raw reply
* [RFC v4 11/11] doc: add release notes for VDUSE live migration support
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
Document the new VDUSE API version 1 features added to support
live migration:
- Address Space ID (ASID) support
- Virtqueue groups
- Queue ready feature (VDUSE_F_QUEUE_READY)
- Suspend feature (VDUSE_F_SUSPEND)
- Link status support (VIRTIO_NET_F_STATUS)
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
doc/guides/rel_notes/release_26_07.rst | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 6352ef27ab05..22c4da3efed6 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -131,6 +131,23 @@ New Features
to support adding and removing memory regions without resetting
the whole guest memory map.
+* **Added VDUSE API version 1 support in vhost library.**
+
+ Updated VDUSE (vDPA Device in Userspace) support with API version 1 features
+ to enable live migration:
+
+ * Added Address Space ID (ASID) support for multiple independent address spaces
+ per device, enabling better isolation and support for virtqueue groups.
+ * Added virtqueue group support to assign different virtqueues to separate
+ address spaces (e.g., control queue vs data queues).
+ * Added ``VDUSE_F_QUEUE_READY`` feature for explicit dataplane queue
+ readiness signaling, ensuring proper device initialization ordering
+ during live migration.
+ * Added ``VDUSE_F_SUSPEND`` feature to reliably suspend virtqueue processing
+ and fetch virtqueue state during live migration.
+ * Added ``VIRTIO_NET_F_STATUS`` support for link status reporting and
+ gratuitous ARP signaling in VDUSE network devices.
+
* **Added LinkData sxe2 ethernet driver.**
Added network driver for the LinkData network adapters.
--
2.55.0
^ permalink raw reply related
* [RFC v4 10/11] vhost: Support vduse suspend feature
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
Add support for the VDUSE_F_SUSPEND feature.
The suspend feature allows the driver to stop the device from processing
the virtqueues. This ensures that the virtqueue state can be fetched
reliably in a live migration.
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 35 ++++++++++++++++++++++++++++++++++-
lib/vhost/vhost.h | 1 +
2 files changed, 35 insertions(+), 1 deletion(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index cb0bbed96012..7156ec4facd7 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -46,7 +46,8 @@ static const char * const vduse_reqs_str[] = {
(id < RTE_DIM(vduse_reqs_str) ? \
vduse_reqs_str[id] : "Unknown")
-static const uint64_t supported_vduse_features = RTE_BIT64(VDUSE_F_QUEUE_READY);
+static const uint64_t supported_vduse_features =
+ RTE_BIT64(VDUSE_F_QUEUE_READY) | RTE_BIT64(VDUSE_F_SUSPEND);
static uint64_t vduse_vq_to_group(struct virtio_net *dev, struct vhost_virtqueue *vq)
{
@@ -412,6 +413,7 @@ vduse_device_stop(struct virtio_net *dev)
vhost_destroy_device_notify(dev);
dev->flags &= ~VIRTIO_DEV_READY;
+ dev->vduse_suspended = false;
for (i = 0; i < dev->nr_vring; i++)
vduse_vring_cleanup(dev, i);
@@ -518,6 +520,12 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
}
+ if (dev->vduse_suspended) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "SET_VQ_READY received on suspended device");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
i = req.vq_ready.num;
if (i >= dev->nr_vring) {
@@ -548,6 +556,31 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
vq->enabled = req.vq_ready.ready;
resp.result = VDUSE_REQ_RESULT_OK;
break;
+ case VDUSE_SUSPEND:
+ if (!(dev->vduse_features & RTE_BIT64(VDUSE_F_SUSPEND))) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Unnegotiated suspend message");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+ if (!(dev->status & VIRTIO_DEVICE_STATUS_DRIVER_OK)) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Unexpected suspend message with no DRIVER_OK");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+ for (i = 0; dev->notify_ops != NULL &&
+ dev->notify_ops->vring_state_changed != NULL &&
+ i < dev->nr_vring; i++) {
+ if (dev->virtqueue[i] == dev->cvq)
+ continue;
+
+ dev->notify_ops->vring_state_changed(dev->vid, i, false);
+ }
+ dev->vduse_suspended = true;
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
+
default:
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index c2b645c2d4b0..f0e7953e3bdc 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -534,6 +534,7 @@ struct __rte_cache_aligned virtio_net {
int vduse_dev_fd;
uint64_t vduse_api_ver;
uint64_t vduse_features;
+ bool vduse_suspended;
struct vhost_virtqueue *cvq;
--
2.55.0
^ permalink raw reply related
* [RFC v4 09/11] vhost: Support VDUSE QUEUE_READY feature
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
Add support for the VDUSE_F_QUEUE_READY feature.
In VDUSE, the dataplane is enabled only after control virtqueue so the
device is fully configured in the destination of a live migration before
the dataplane starts. This message signals the VDUSE device when the
dataplane queues should be enabled.
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
v4:
* Use new VDUSE_SET_FEATURES ioctl.
* Clarified some error messages and check for more conditions, like
invalid queue index.
v3:
* Replace incorrect '%lx' DEBUG print format specifier with PRIx64
v2:
* Following latest comments on kernel series about VDUSE features, not
checking API version but only check if VDUSE_GET_FEATURES success.
---
lib/vhost/vduse.c | 85 +++++++++++++++++++++++++++++++++++++++++++++--
lib/vhost/vhost.h | 1 +
2 files changed, 84 insertions(+), 2 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index a94cbb40a1d8..cb0bbed96012 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -39,12 +39,15 @@ static const char * const vduse_reqs_str[] = {
"VDUSE_SET_STATUS",
"VDUSE_UPDATE_IOTLB",
"VDUSE_SET_VQ_GROUP_ASID",
+ "VDUSE_SET_VQ_READY",
};
#define vduse_req_id_to_str(id) \
(id < RTE_DIM(vduse_reqs_str) ? \
vduse_reqs_str[id] : "Unknown")
+static const uint64_t supported_vduse_features = RTE_BIT64(VDUSE_F_QUEUE_READY);
+
static uint64_t vduse_vq_to_group(struct virtio_net *dev, struct vhost_virtqueue *vq)
{
if (dev->vduse_api_ver < 1)
@@ -500,6 +503,51 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
}
resp.result = VDUSE_REQ_RESULT_OK;
break;
+ case VDUSE_SET_VQ_READY:
+ if (!(dev->status & VIRTIO_DEVICE_STATUS_DRIVER_OK)) {
+ /*
+ * dev->notify_ops is NULL if !S_DRIVER_OK,
+ * vduse_device_start will check the queue readiness.
+ */
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
+ }
+ if (!(dev->vduse_features & RTE_BIT64(VDUSE_F_QUEUE_READY))) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Unexpected ready message with no ready feature acked");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ i = req.vq_ready.num;
+ if (i >= dev->nr_vring) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Invalid virtqueue index: %u", i);
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ vq = dev->virtqueue[i];
+ if (dev->notify_ops == NULL || dev->notify_ops->vring_state_changed == NULL) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "No ops->vring_state_changed");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ ret = dev->notify_ops->vring_state_changed(dev->vid, i,
+ req.vq_ready.ready);
+ VHOST_CONFIG_LOG(dev->ifname, INFO,
+ "\t\t VQ %d gets ready %d ok %d", i,
+ req.vq_ready.ready, ret);
+ if (ret != 0) {
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ vq->enabled = req.vq_ready.ready;
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
default:
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
@@ -517,7 +565,8 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
if ((old_status ^ dev->status) & VIRTIO_DEVICE_STATUS_DRIVER_OK) {
if (dev->status & VIRTIO_DEVICE_STATUS_DRIVER_OK) {
/* Poll virtqueues ready states before starting device */
- ret = vduse_wait_for_virtqueues_ready(dev);
+ ret = dev->vduse_features & RTE_BIT64(VDUSE_F_QUEUE_READY) ? 0
+ : vduse_wait_for_virtqueues_ready(dev);
if (ret < 0) {
VHOST_CONFIG_LOG(dev->ifname, ERR,
"Failed to wait for virtqueues ready, aborting device start");
@@ -723,6 +772,26 @@ vduse_reconnect_start_device(struct virtio_net *dev)
return ret;
}
+/* If some error occurs just continue as if the kernel exposed no features */
+static uint64_t
+vduse_device_get_vduse_features(int control_fd, const char *log_name)
+{
+ uint64_t vduse_kernel_features;
+ int ret;
+
+ ret = ioctl(control_fd, VDUSE_GET_FEATURES, &vduse_kernel_features);
+ if (ret < 0) {
+ VHOST_CONFIG_LOG(log_name, INFO,
+ "Failed to get kernel VDUSE features, assuming not supported: %d(%s)",
+ errno, strerror(errno));
+ return 0;
+ }
+
+ VHOST_CONFIG_LOG(log_name, DEBUG, "Setting vhost kernel features: %"PRIx64,
+ vduse_kernel_features & supported_vduse_features);
+ return vduse_kernel_features & supported_vduse_features;
+}
+
int
vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool linearbuf)
{
@@ -731,7 +800,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
struct virtio_net *dev;
struct virtio_net_config vnet_config = {{ 0 }};
uint64_t ver;
- uint64_t features;
+ uint64_t features, vduse_features = 0;
const char *name = path + strlen("/dev/vduse/");
bool reconnect = false;
@@ -815,9 +884,20 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
dev_config->ngroups = 2;
dev_config->nas = 2;
}
+
+ vduse_features = vduse_device_get_vduse_features(control_fd, name);
dev_config->config_size = sizeof(struct virtio_net_config);
memcpy(dev_config->config, &vnet_config, sizeof(vnet_config));
+ if (vduse_features) {
+ ret = ioctl(control_fd, VDUSE_SET_FEATURES, &vduse_features);
+ if (ret < 0) {
+ VHOST_CONFIG_LOG(name, ERR, "Failed to set VDUSE features: %s",
+ strerror(errno));
+ goto out_ctrl_close;
+ }
+ }
+
ret = ioctl(control_fd, VDUSE_CREATE_DEV, dev_config);
free(dev_config);
dev_config = NULL;
@@ -865,6 +945,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
dev->vduse_ctrl_fd = control_fd;
dev->vduse_dev_fd = dev_fd;
dev->vduse_api_ver = ver;
+ dev->vduse_features = vduse_features;
ret = vduse_reconnect_log_map(dev, !reconnect);
if (ret < 0)
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index 5cbd64d539b9..c2b645c2d4b0 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -533,6 +533,7 @@ struct __rte_cache_aligned virtio_net {
int vduse_ctrl_fd;
int vduse_dev_fd;
uint64_t vduse_api_ver;
+ uint64_t vduse_features;
struct vhost_virtqueue *cvq;
--
2.55.0
^ permalink raw reply related
* [RFC v4 08/11] uapi: Align vduse.h for enable and suspend VDUSE messages
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
This is a prerequisite for next patches to use them.
These features are required to properly support live migration and proper
device initialization ordering.
These are not in maintainer's branch at the moment:
https://lore.kernel.org/lkml/20260310190759.1097506-1-eperezma@redhat.com
https://lore.kernel.org/lkml/20260310191019.1099757-1-eperezma@redhat.com
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
kernel/linux/uapi/linux/vduse.h | 28 ++++++++++++++++++++++++++++
1 file changed, 28 insertions(+)
diff --git a/kernel/linux/uapi/linux/vduse.h b/kernel/linux/uapi/linux/vduse.h
index 361eea511c21..b7f8c04a0a44 100644
--- a/kernel/linux/uapi/linux/vduse.h
+++ b/kernel/linux/uapi/linux/vduse.h
@@ -14,6 +14,12 @@
#define VDUSE_API_VERSION_1 1
+/* The VDUSE instance expects a request for vq ready */
+#define VDUSE_F_QUEUE_READY 0
+
+/* The VDUSE instance expects a request for suspend */
+#define VDUSE_F_SUSPEND 1
+
/*
* Get the version of VDUSE API that kernel supported (VDUSE_API_VERSION).
* This is used for future extension.
@@ -63,6 +69,12 @@ struct vduse_dev_config {
*/
#define VDUSE_DESTROY_DEV _IOW(VDUSE_BASE, 0x03, char[VDUSE_NAME_MAX])
+/* Get the VDUSE supported features */
+#define VDUSE_GET_FEATURES _IOR(VDUSE_BASE, 0x04, __u64)
+
+/* Set the VDUSE features */
+#define VDUSE_SET_FEATURES _IOW(VDUSE_BASE, 0x05, __u64)
+
/* The ioctls for VDUSE device (/dev/vduse/$NAME) */
/**
@@ -325,6 +337,8 @@ enum vduse_req_type {
VDUSE_SET_STATUS,
VDUSE_UPDATE_IOTLB,
VDUSE_SET_VQ_GROUP_ASID,
+ VDUSE_SET_VQ_READY,
+ VDUSE_SUSPEND,
};
/**
@@ -372,6 +386,15 @@ struct vduse_iova_range_v2 {
__u32 padding;
};
+/**
+ * struct vduse_vq_ready - Virtqueue ready request message
+ * @num: Virtqueue number
+ */
+struct vduse_vq_ready {
+ __u32 num;
+ __u32 ready;
+};
+
/**
* struct vduse_dev_request - control request
* @type: request type
@@ -382,6 +405,7 @@ struct vduse_iova_range_v2 {
* @iova: IOVA range for updating
* @iova_v2: IOVA range for updating if API_VERSION >= 1
* @vq_group_asid: ASID of a virtqueue group
+ * @vq_ready: Virtqueue ready request
* @padding: padding
*
* Structure used by read(2) on /dev/vduse/$NAME.
@@ -399,6 +423,10 @@ struct vduse_dev_request {
*/
struct vduse_iova_range_v2 iova_v2;
struct vduse_vq_group_asid vq_group_asid;
+
+ /* Only if VDUSE_F_QUEUE_READY is negotiated */
+ struct vduse_vq_ready vq_ready;
+
__u32 padding[32];
};
};
--
2.55.0
^ permalink raw reply related
* [RFC v4 07/11] vhost: add net status feature to VDUSE
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Enable the VIRTIO_NET_F_STATUS feature for VDUSE devices.
This allows the device to report link status (e.g.,
VIRTIO_NET_S_LINK_UP). It also allows the device to signal the driver
that it needs to send gratuitous ARP with VIRTIO_NET_S_ANNOUNCE.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 1 +
lib/vhost/vduse.h | 3 ++-
2 files changed, 3 insertions(+), 1 deletion(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index cac1956fbc03..a94cbb40a1d8 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -801,6 +801,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
goto out_ctrl_close;
}
+ vnet_config.status = VIRTIO_NET_S_LINK_UP;
vnet_config.max_virtqueue_pairs = max_queue_pairs;
memset(dev_config, 0, sizeof(struct vduse_dev_config));
diff --git a/lib/vhost/vduse.h b/lib/vhost/vduse.h
index b2515bb9df76..d697f85be5cc 100644
--- a/lib/vhost/vduse.h
+++ b/lib/vhost/vduse.h
@@ -7,7 +7,8 @@
#include "vhost.h"
-#define VDUSE_NET_SUPPORTED_FEATURES VIRTIO_NET_SUPPORTED_FEATURES
+#define VDUSE_NET_SUPPORTED_FEATURES (VIRTIO_NET_SUPPORTED_FEATURES | \
+ (1ULL << VIRTIO_NET_F_STATUS))
int vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool linearbuf);
int vduse_device_destroy(const char *path);
--
2.55.0
^ permalink raw reply related
* [RFC v4 06/11] vhost: claim VDUSE support for API version 1
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 27609f28c1a3..cac1956fbc03 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -25,7 +25,7 @@
#include "vhost.h"
#include "virtio_net_ctrl.h"
-#define VHOST_VDUSE_API_VERSION 0ULL
+#define VHOST_VDUSE_API_VERSION 1ULL
#define VDUSE_CTRL_PATH "/dev/vduse/control"
struct vduse {
--
2.55.0
^ permalink raw reply related
* [RFC v4 05/11] vhost: add ASID support to VDUSE IOTLB operations
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Make use of the newly introduced address space ID when
calling Vhost IOTLB API.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 28 ++++++++++++++++++++++------
1 file changed, 22 insertions(+), 6 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 5cfe569f7d2a..27609f28c1a3 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -71,7 +71,7 @@ vduse_iotlb_remove_notify(uint64_t addr, uint64_t offset, uint64_t size)
static int
vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm __rte_unused)
{
- struct vduse_iotlb_entry entry;
+ struct vduse_iotlb_entry_v2 entry = {};
uint64_t size, page_size;
struct stat stat;
void *mmap_addr;
@@ -79,8 +79,9 @@ vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm _
entry.start = iova;
entry.last = iova + 1;
+ entry.asid = asid;
- ret = ioctl(dev->vduse_dev_fd, VDUSE_IOTLB_GET_FD, &entry);
+ ret = ioctl(dev->vduse_dev_fd, VDUSE_IOTLB_GET_FD2, &entry);
if (ret < 0) {
VHOST_CONFIG_LOG(dev->ifname, ERR, "Failed to get IOTLB entry for 0x%" PRIx64,
iova);
@@ -90,6 +91,7 @@ vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm _
fd = ret;
VHOST_CONFIG_LOG(dev->ifname, DEBUG, "New IOTLB entry:");
+ VHOST_CONFIG_LOG(dev->ifname, DEBUG, "\tASID: %d", entry.asid);
VHOST_CONFIG_LOG(dev->ifname, DEBUG, "\tIOVA: %" PRIx64 " - %" PRIx64,
(uint64_t)entry.start, (uint64_t)entry.last);
VHOST_CONFIG_LOG(dev->ifname, DEBUG, "\toffset: %" PRIx64, (uint64_t)entry.offset);
@@ -458,10 +460,24 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
resp.result = VDUSE_REQ_RESULT_OK;
break;
case VDUSE_UPDATE_IOTLB:
- VHOST_CONFIG_LOG(dev->ifname, INFO, "\tIOVA range: %" PRIx64 " - %" PRIx64,
- (uint64_t)req.iova.start, (uint64_t)req.iova.last);
- vhost_user_iotlb_cache_remove(dev, 0, req.iova.start,
- req.iova.last - req.iova.start + 1);
+ {
+ uint64_t start, last;
+ uint32_t asid;
+
+ if (dev->vduse_api_ver < 1) {
+ start = req.iova.start;
+ last = req.iova.last;
+ asid = 0;
+ } else {
+ start = req.iova_v2.start;
+ last = req.iova_v2.last;
+ asid = req.iova_v2.asid;
+ }
+
+ VHOST_CONFIG_LOG(dev->ifname, INFO, "\t(ASID %d) IOVA range: %" PRIx64 " - %" PRIx64,
+ asid, start, last);
+ vhost_user_iotlb_cache_remove(dev, asid, start, last - start + 1);
+ }
resp.result = VDUSE_REQ_RESULT_OK;
break;
case VDUSE_SET_VQ_GROUP_ASID:
--
2.55.0
^ permalink raw reply related
* [RFC v4 04/11] vhost: add virtqueues groups support to VDUSE
From: Eugenio Pérez @ 2026-07-07 12:45 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
VDUSE API version 1 introduces the notion of virtqueue
groups, which once supported, enables the support of
multiple addresses spaces.
For VDUSE networking devices, we need two groups, one for
the datapath queues, and one for the control queue.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 44 ++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 42 insertions(+), 2 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 1891c4a9bbc3..5cfe569f7d2a 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -38,12 +38,24 @@ static const char * const vduse_reqs_str[] = {
"VDUSE_GET_VQ_STATE",
"VDUSE_SET_STATUS",
"VDUSE_UPDATE_IOTLB",
+ "VDUSE_SET_VQ_GROUP_ASID",
};
#define vduse_req_id_to_str(id) \
(id < RTE_DIM(vduse_reqs_str) ? \
vduse_reqs_str[id] : "Unknown")
+static uint64_t vduse_vq_to_group(struct virtio_net *dev, struct vhost_virtqueue *vq)
+{
+ if (dev->vduse_api_ver < 1)
+ return 0;
+
+ if (vq == dev->cvq)
+ return 1;
+
+ return 0;
+}
+
static int
vduse_inject_irq(struct virtio_net *dev, struct vhost_virtqueue *vq)
{
@@ -271,6 +283,7 @@ vduse_vring_cleanup(struct virtio_net *dev, unsigned int index)
vq->size = 0;
vq->last_used_idx = 0;
vq->last_avail_idx = 0;
+ vq->asid = 0;
}
/*
@@ -410,6 +423,7 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
struct vduse_dev_response resp;
struct vhost_virtqueue *vq;
uint8_t old_status = dev->status;
+ uint32_t i;
int ret;
memset(&resp, 0, sizeof(resp));
@@ -450,6 +464,26 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
req.iova.last - req.iova.start + 1);
resp.result = VDUSE_REQ_RESULT_OK;
break;
+ case VDUSE_SET_VQ_GROUP_ASID:
+ if (dev->vduse_api_ver < 1) {
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ VHOST_CONFIG_LOG(dev->ifname, INFO, "\tAssigning ASID %d to group %d",
+ req.vq_group_asid.asid, req.vq_group_asid.group);
+
+ for (i = 0; i < dev->nr_vring; i++) {
+ vq = dev->virtqueue[i];
+
+ if (vduse_vq_to_group(dev, vq) == req.vq_group_asid.group) {
+ vq->asid = req.vq_group_asid.asid;
+ VHOST_CONFIG_LOG(dev->ifname, INFO, "\t\tVQ %d gets ASID %d",
+ i, req.vq_group_asid.asid);
+ }
+ }
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
default:
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
@@ -760,6 +794,10 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
dev_config->features = features;
dev_config->vq_num = total_queues;
dev_config->vq_align = rte_mem_page_size();
+ if (ver >= 1) {
+ dev_config->ngroups = 2;
+ dev_config->nas = 2;
+ }
dev_config->config_size = sizeof(struct virtio_net_config);
memcpy(dev_config->config, &vnet_config, sizeof(vnet_config));
@@ -848,11 +886,15 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
vq = dev->virtqueue[i];
vq->reconnect_log = &dev->reconnect_log->vring[i];
+ if (i == max_queue_pairs * 2)
+ dev->cvq = vq;
+
if (reconnect)
continue;
vq_cfg.index = i;
vq_cfg.max_size = 1024;
+ vq_cfg.group = vduse_vq_to_group(dev, vq);
ret = ioctl(dev->vduse_dev_fd, VDUSE_VQ_SETUP, &vq_cfg);
if (ret) {
@@ -861,8 +903,6 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
}
}
- dev->cvq = dev->virtqueue[max_queue_pairs * 2];
-
ret = fdset_add(vduse.fdset, dev->vduse_dev_fd, vduse_events_handler, NULL, dev);
if (ret) {
VHOST_CONFIG_LOG(name, ERR, "Failed to add fd %d to vduse fdset",
--
2.55.0
^ permalink raw reply related
* [RFC v4 03/11] vhost: add VDUSE API version negotiation
From: Eugenio Pérez @ 2026-07-07 12:44 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
As preliminary step to support new VDUSE API version
introducing ASID support, this patch adds API version
negotiation to keep compatibility with older kernels.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 14 ++++++++++++--
lib/vhost/vhost.h | 1 +
2 files changed, 13 insertions(+), 2 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 7217895d1b9b..1891c4a9bbc3 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -25,7 +25,7 @@
#include "vhost.h"
#include "virtio_net_ctrl.h"
-#define VHOST_VDUSE_API_VERSION 0
+#define VHOST_VDUSE_API_VERSION 0ULL
#define VDUSE_CTRL_PATH "/dev/vduse/control"
struct vduse {
@@ -680,7 +680,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
uint32_t i, max_queue_pairs, total_queues;
struct virtio_net *dev;
struct virtio_net_config vnet_config = {{ 0 }};
- uint64_t ver = VHOST_VDUSE_API_VERSION;
+ uint64_t ver;
uint64_t features;
const char *name = path + strlen("/dev/vduse/");
bool reconnect = false;
@@ -700,6 +700,15 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
return -1;
}
+ if (ioctl(control_fd, VDUSE_GET_API_VERSION, &ver)) {
+ VHOST_CONFIG_LOG(name, ERR, "Failed to get API version: %s", strerror(errno));
+ ret = -1;
+ goto out_ctrl_close;
+ }
+
+ ver = RTE_MIN(ver, VHOST_VDUSE_API_VERSION);
+ VHOST_CONFIG_LOG(name, INFO, "Using VDUSE API version %" PRIu64 "", ver);
+
if (ioctl(control_fd, VDUSE_SET_API_VERSION, &ver)) {
VHOST_CONFIG_LOG(name, ERR, "Failed to set API version: %" PRIu64 ": %s",
ver, strerror(errno));
@@ -800,6 +809,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
strncpy(dev->ifname, path, IF_NAME_SZ - 1);
dev->vduse_ctrl_fd = control_fd;
dev->vduse_dev_fd = dev_fd;
+ dev->vduse_api_ver = ver;
ret = vduse_reconnect_log_map(dev, !reconnect);
if (ret < 0)
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index 0b3fca0ab769..5cbd64d539b9 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -532,6 +532,7 @@ struct __rte_cache_aligned virtio_net {
int postcopy_listening;
int vduse_ctrl_fd;
int vduse_dev_fd;
+ uint64_t vduse_api_ver;
struct vhost_virtqueue *cvq;
--
2.55.0
^ permalink raw reply related
* [RFC v4 02/11] vhost: introduce ASID support
From: Eugenio Pérez @ 2026-07-07 12:44 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Set all ASID = 0 as it is the default when not set explicitly.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/iotlb.c | 233 +++++++++++++++++++++++++----------------
lib/vhost/iotlb.h | 14 +--
lib/vhost/vduse.c | 9 +-
lib/vhost/vhost.c | 16 +--
lib/vhost/vhost.h | 13 +--
lib/vhost/vhost_user.c | 13 +--
6 files changed, 177 insertions(+), 121 deletions(-)
diff --git a/lib/vhost/iotlb.c b/lib/vhost/iotlb.c
index f2c275a7d77e..e5e69d11cb65 100644
--- a/lib/vhost/iotlb.c
+++ b/lib/vhost/iotlb.c
@@ -11,6 +11,16 @@
#include "iotlb.h"
#include "vhost.h"
+struct iotlb {
+ rte_rwlock_t pending_lock;
+ struct vhost_iotlb_entry *pool;
+ TAILQ_HEAD(, vhost_iotlb_entry) list;
+ TAILQ_HEAD(, vhost_iotlb_entry) pending_list;
+ int cache_nr;
+ rte_spinlock_t free_lock;
+ SLIST_HEAD(, vhost_iotlb_entry) free_list;
+};
+
struct vhost_iotlb_entry {
TAILQ_ENTRY(vhost_iotlb_entry) next;
SLIST_ENTRY(vhost_iotlb_entry) next_free;
@@ -85,78 +95,78 @@ vhost_user_iotlb_clear_dump(struct virtio_net *dev, struct vhost_iotlb_entry *no
}
static struct vhost_iotlb_entry *
-vhost_user_iotlb_pool_get(struct virtio_net *dev)
+vhost_user_iotlb_pool_get(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node;
- rte_spinlock_lock(&dev->iotlb_free_lock);
- node = SLIST_FIRST(&dev->iotlb_free_list);
+ rte_spinlock_lock(&dev->iotlb[asid]->free_lock);
+ node = SLIST_FIRST(&dev->iotlb[asid]->free_list);
if (node != NULL)
- SLIST_REMOVE_HEAD(&dev->iotlb_free_list, next_free);
- rte_spinlock_unlock(&dev->iotlb_free_lock);
+ SLIST_REMOVE_HEAD(&dev->iotlb[asid]->free_list, next_free);
+ rte_spinlock_unlock(&dev->iotlb[asid]->free_lock);
return node;
}
static void
-vhost_user_iotlb_pool_put(struct virtio_net *dev, struct vhost_iotlb_entry *node)
+vhost_user_iotlb_pool_put(struct virtio_net *dev, int asid, struct vhost_iotlb_entry *node)
{
- rte_spinlock_lock(&dev->iotlb_free_lock);
- SLIST_INSERT_HEAD(&dev->iotlb_free_list, node, next_free);
- rte_spinlock_unlock(&dev->iotlb_free_lock);
+ rte_spinlock_lock(&dev->iotlb[asid]->free_lock);
+ SLIST_INSERT_HEAD(&dev->iotlb[asid]->free_list, node, next_free);
+ rte_spinlock_unlock(&dev->iotlb[asid]->free_lock);
}
static void
-vhost_user_iotlb_cache_random_evict(struct virtio_net *dev);
+vhost_user_iotlb_cache_random_evict(struct virtio_net *dev, int asid);
static void
-vhost_user_iotlb_pending_remove_all(struct virtio_net *dev)
+vhost_user_iotlb_pending_remove_all(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node, *temp_node;
- rte_rwlock_write_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_lock(&dev->iotlb[asid]->pending_lock);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_pending_list, next, temp_node) {
- TAILQ_REMOVE(&dev->iotlb_pending_list, node, next);
- vhost_user_iotlb_pool_put(dev, node);
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->pending_list, next, temp_node) {
+ TAILQ_REMOVE(&dev->iotlb[asid]->pending_list, node, next);
+ vhost_user_iotlb_pool_put(dev, asid, node);
}
- rte_rwlock_write_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_unlock(&dev->iotlb[asid]->pending_lock);
}
bool
-vhost_user_iotlb_pending_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_user_iotlb_pending_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm)
{
struct vhost_iotlb_entry *node;
bool found = false;
- rte_rwlock_read_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_read_lock(&dev->iotlb[asid]->pending_lock);
- TAILQ_FOREACH(node, &dev->iotlb_pending_list, next) {
+ TAILQ_FOREACH(node, &dev->iotlb[asid]->pending_list, next) {
if ((node->iova == iova) && (node->perm == perm)) {
found = true;
break;
}
}
- rte_rwlock_read_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_read_unlock(&dev->iotlb[asid]->pending_lock);
return found;
}
void
-vhost_user_iotlb_pending_insert(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_user_iotlb_pending_insert(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm)
{
struct vhost_iotlb_entry *node;
- node = vhost_user_iotlb_pool_get(dev);
+ node = vhost_user_iotlb_pool_get(dev, asid);
if (node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, DEBUG,
"IOTLB pool empty, clear entries for pending insertion");
- if (!TAILQ_EMPTY(&dev->iotlb_pending_list))
- vhost_user_iotlb_pending_remove_all(dev);
+ if (!TAILQ_EMPTY(&dev->iotlb[asid]->pending_list))
+ vhost_user_iotlb_pending_remove_all(dev, asid);
else
- vhost_user_iotlb_cache_random_evict(dev);
- node = vhost_user_iotlb_pool_get(dev);
+ vhost_user_iotlb_cache_random_evict(dev, asid);
+ node = vhost_user_iotlb_pool_get(dev, asid);
if (node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, ERR,
"IOTLB pool still empty, pending insertion failure");
@@ -167,21 +177,22 @@ vhost_user_iotlb_pending_insert(struct virtio_net *dev, uint64_t iova, uint8_t p
node->iova = iova;
node->perm = perm;
- rte_rwlock_write_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_lock(&dev->iotlb[asid]->pending_lock);
- TAILQ_INSERT_TAIL(&dev->iotlb_pending_list, node, next);
+ TAILQ_INSERT_TAIL(&dev->iotlb[asid]->pending_list, node, next);
- rte_rwlock_write_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_unlock(&dev->iotlb[asid]->pending_lock);
}
void
-vhost_user_iotlb_pending_remove(struct virtio_net *dev, uint64_t iova, uint64_t size, uint8_t perm)
+vhost_user_iotlb_pending_remove(struct virtio_net *dev, int asid,
+ uint64_t iova, uint64_t size, uint8_t perm)
{
struct vhost_iotlb_entry *node, *temp_node;
- rte_rwlock_write_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_lock(&dev->iotlb[asid]->pending_lock);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_pending_list, next,
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->pending_list, next,
temp_node) {
if (node->iova < iova)
continue;
@@ -189,53 +200,53 @@ vhost_user_iotlb_pending_remove(struct virtio_net *dev, uint64_t iova, uint64_t
continue;
if ((node->perm & perm) != node->perm)
continue;
- TAILQ_REMOVE(&dev->iotlb_pending_list, node, next);
- vhost_user_iotlb_pool_put(dev, node);
+ TAILQ_REMOVE(&dev->iotlb[asid]->pending_list, node, next);
+ vhost_user_iotlb_pool_put(dev, asid, node);
}
- rte_rwlock_write_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_unlock(&dev->iotlb[asid]->pending_lock);
}
static void
-vhost_user_iotlb_cache_remove_all(struct virtio_net *dev)
+vhost_user_iotlb_cache_remove_all(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node, *temp_node;
vhost_user_iotlb_wr_lock_all(dev);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_list, next, temp_node) {
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->list, next, temp_node) {
vhost_user_iotlb_clear_dump(dev, node, NULL, NULL);
- TAILQ_REMOVE(&dev->iotlb_list, node, next);
+ TAILQ_REMOVE(&dev->iotlb[asid]->list, node, next);
vhost_user_iotlb_remove_notify(dev, node);
- vhost_user_iotlb_pool_put(dev, node);
+ vhost_user_iotlb_pool_put(dev, asid, node);
}
- dev->iotlb_cache_nr = 0;
+ dev->iotlb[asid]->cache_nr = 0;
vhost_user_iotlb_wr_unlock_all(dev);
}
static void
-vhost_user_iotlb_cache_random_evict(struct virtio_net *dev)
+vhost_user_iotlb_cache_random_evict(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node, *temp_node, *prev_node = NULL;
int entry_idx;
vhost_user_iotlb_wr_lock_all(dev);
- entry_idx = rte_rand() % dev->iotlb_cache_nr;
+ entry_idx = rte_rand() % dev->iotlb[asid]->cache_nr;
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_list, next, temp_node) {
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->list, next, temp_node) {
if (!entry_idx) {
struct vhost_iotlb_entry *next_node = RTE_TAILQ_NEXT(node, next);
vhost_user_iotlb_clear_dump(dev, node, prev_node, next_node);
- TAILQ_REMOVE(&dev->iotlb_list, node, next);
+ TAILQ_REMOVE(&dev->iotlb[asid]->list, node, next);
vhost_user_iotlb_remove_notify(dev, node);
- vhost_user_iotlb_pool_put(dev, node);
- dev->iotlb_cache_nr--;
+ vhost_user_iotlb_pool_put(dev, asid, node);
+ dev->iotlb[asid]->cache_nr--;
break;
}
prev_node = node;
@@ -246,20 +257,20 @@ vhost_user_iotlb_cache_random_evict(struct virtio_net *dev)
}
void
-vhost_user_iotlb_cache_insert(struct virtio_net *dev, uint64_t iova, uint64_t uaddr,
+vhost_user_iotlb_cache_insert(struct virtio_net *dev, int asid, uint64_t iova, uint64_t uaddr,
uint64_t uoffset, uint64_t size, uint64_t page_size, uint8_t perm)
{
struct vhost_iotlb_entry *node, *new_node;
- new_node = vhost_user_iotlb_pool_get(dev);
+ new_node = vhost_user_iotlb_pool_get(dev, asid);
if (new_node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, DEBUG,
"IOTLB pool empty, clear entries for cache insertion");
- if (!TAILQ_EMPTY(&dev->iotlb_list))
- vhost_user_iotlb_cache_random_evict(dev);
+ if (!TAILQ_EMPTY(&dev->iotlb[asid]->list))
+ vhost_user_iotlb_cache_random_evict(dev, asid);
else
- vhost_user_iotlb_pending_remove_all(dev);
- new_node = vhost_user_iotlb_pool_get(dev);
+ vhost_user_iotlb_pending_remove_all(dev, asid);
+ new_node = vhost_user_iotlb_pool_get(dev, asid);
if (new_node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, ERR,
"IOTLB pool still empty, cache insertion failed");
@@ -276,36 +287,36 @@ vhost_user_iotlb_cache_insert(struct virtio_net *dev, uint64_t iova, uint64_t ua
vhost_user_iotlb_wr_lock_all(dev);
- TAILQ_FOREACH(node, &dev->iotlb_list, next) {
+ TAILQ_FOREACH(node, &dev->iotlb[asid]->list, next) {
/*
* Entries must be invalidated before being updated.
* So if iova already in list, assume identical.
*/
if (node->iova == new_node->iova) {
- vhost_user_iotlb_pool_put(dev, new_node);
+ vhost_user_iotlb_pool_put(dev, asid, new_node);
goto unlock;
} else if (node->iova > new_node->iova) {
vhost_user_iotlb_set_dump(dev, new_node);
TAILQ_INSERT_BEFORE(node, new_node, next);
- dev->iotlb_cache_nr++;
+ dev->iotlb[asid]->cache_nr++;
goto unlock;
}
}
vhost_user_iotlb_set_dump(dev, new_node);
- TAILQ_INSERT_TAIL(&dev->iotlb_list, new_node, next);
- dev->iotlb_cache_nr++;
+ TAILQ_INSERT_TAIL(&dev->iotlb[asid]->list, new_node, next);
+ dev->iotlb[asid]->cache_nr++;
unlock:
- vhost_user_iotlb_pending_remove(dev, iova, size, perm);
+ vhost_user_iotlb_pending_remove(dev, asid, iova, size, perm);
vhost_user_iotlb_wr_unlock_all(dev);
}
void
-vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t size)
+vhost_user_iotlb_cache_remove(struct virtio_net *dev, int asid, uint64_t iova, uint64_t size)
{
struct vhost_iotlb_entry *node, *temp_node, *prev_node = NULL;
@@ -314,7 +325,7 @@ vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t si
vhost_user_iotlb_wr_lock_all(dev);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_list, next, temp_node) {
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->list, next, temp_node) {
/* Sorted list */
if (unlikely(iova + size < node->iova))
break;
@@ -324,10 +335,10 @@ vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t si
vhost_user_iotlb_clear_dump(dev, node, prev_node, next_node);
- TAILQ_REMOVE(&dev->iotlb_list, node, next);
+ TAILQ_REMOVE(&dev->iotlb[asid]->list, node, next);
vhost_user_iotlb_remove_notify(dev, node);
- vhost_user_iotlb_pool_put(dev, node);
- dev->iotlb_cache_nr--;
+ vhost_user_iotlb_pool_put(dev, asid, node);
+ dev->iotlb[asid]->cache_nr--;
} else {
prev_node = node;
}
@@ -337,7 +348,8 @@ vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t si
}
uint64_t
-vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova, uint64_t *size, uint8_t perm)
+vhost_user_iotlb_cache_find(struct virtio_net *dev, int asid,
+ uint64_t iova, uint64_t *size, uint8_t perm)
{
struct vhost_iotlb_entry *node;
uint64_t offset, vva = 0, mapped = 0;
@@ -345,7 +357,7 @@ vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova, uint64_t *siz
if (unlikely(!*size))
goto out;
- TAILQ_FOREACH(node, &dev->iotlb_list, next) {
+ TAILQ_FOREACH(node, &dev->iotlb[asid]->list, next) {
/* List sorted by iova */
if (unlikely(iova < node->iova))
break;
@@ -378,25 +390,28 @@ vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova, uint64_t *siz
}
void
-vhost_user_iotlb_flush_all(struct virtio_net *dev)
+vhost_user_iotlb_flush_all(struct virtio_net *dev, int asid)
{
- vhost_user_iotlb_cache_remove_all(dev);
- vhost_user_iotlb_pending_remove_all(dev);
+ vhost_user_iotlb_cache_remove_all(dev, asid);
+ vhost_user_iotlb_pending_remove_all(dev, asid);
}
-int
-vhost_user_iotlb_init(struct virtio_net *dev)
+static int
+vhost_user_iotlb_init_one(struct virtio_net *dev, int asid)
{
unsigned int i;
int socket = 0;
- if (dev->iotlb_pool) {
- /*
- * The cache has already been initialized,
- * just drop all cached and pending entries.
- */
- vhost_user_iotlb_flush_all(dev);
- rte_free(dev->iotlb_pool);
+ if (dev->iotlb[asid] != NULL) {
+ if (dev->iotlb[asid]->pool != NULL) {
+ /*
+ * The cache has already been initialized,
+ * just drop all cached and pending entries.
+ */
+ vhost_user_iotlb_flush_all(dev, asid);
+ rte_free(dev->iotlb[asid]->pool);
+ }
+ rte_free(dev->iotlb[asid]);
}
#ifdef RTE_LIBRTE_VHOST_NUMA
@@ -404,31 +419,73 @@ vhost_user_iotlb_init(struct virtio_net *dev)
socket = 0;
#endif
- rte_spinlock_init(&dev->iotlb_free_lock);
- rte_rwlock_init(&dev->iotlb_pending_lock);
+ dev->iotlb[asid] = rte_malloc_socket("iotlb", sizeof(struct iotlb), 0, socket);
+ if (!dev->iotlb[asid]) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR, "Failed to allocate IOTLB");
+ return -1;
+ }
+
+ rte_spinlock_init(&dev->iotlb[asid]->free_lock);
+ rte_rwlock_init(&dev->iotlb[asid]->pending_lock);
- SLIST_INIT(&dev->iotlb_free_list);
- TAILQ_INIT(&dev->iotlb_list);
- TAILQ_INIT(&dev->iotlb_pending_list);
+ SLIST_INIT(&dev->iotlb[asid]->free_list);
+ TAILQ_INIT(&dev->iotlb[asid]->list);
+ TAILQ_INIT(&dev->iotlb[asid]->pending_list);
if (dev->flags & VIRTIO_DEV_SUPPORT_IOMMU) {
- dev->iotlb_pool = rte_calloc_socket("iotlb", IOTLB_CACHE_SIZE,
+ dev->iotlb[asid]->pool = rte_calloc_socket("iotlb_pool", IOTLB_CACHE_SIZE,
sizeof(struct vhost_iotlb_entry), 0, socket);
- if (!dev->iotlb_pool) {
+ if (!dev->iotlb[asid]->pool) {
VHOST_CONFIG_LOG(dev->ifname, ERR, "Failed to create IOTLB cache pool");
- return -1;
+ goto free_iotlb;
}
for (i = 0; i < IOTLB_CACHE_SIZE; i++)
- vhost_user_iotlb_pool_put(dev, &dev->iotlb_pool[i]);
+ vhost_user_iotlb_pool_put(dev, asid, &dev->iotlb[asid]->pool[i]);
}
- dev->iotlb_cache_nr = 0;
+ dev->iotlb[asid]->cache_nr = 0;
+
+ return 0;
+
+free_iotlb:
+ rte_free(dev->iotlb[asid]);
+ dev->iotlb[asid] = NULL;
+ return -1;
+}
+
+int
+vhost_user_iotlb_init(struct virtio_net *dev)
+{
+ int i;
+
+ for (i = 0; i < IOTLB_MAX_ASID; i++)
+ if (vhost_user_iotlb_init_one(dev, i) < 0)
+ goto fail;
return 0;
+fail:
+ while (i--) {
+ rte_free(dev->iotlb[i]->pool);
+ dev->iotlb[i]->pool = NULL;
+ rte_free(dev->iotlb[i]);
+ dev->iotlb[i] = NULL;
+ }
+
+ return -1;
}
void
vhost_user_iotlb_destroy(struct virtio_net *dev)
{
- rte_free(dev->iotlb_pool);
+ int i;
+
+ for (i = 0; i < IOTLB_MAX_ASID; i++) {
+ if (dev->iotlb[i]) {
+ rte_free(dev->iotlb[i]->pool);
+ dev->iotlb[i]->pool = NULL;
+
+ rte_free(dev->iotlb[i]);
+ dev->iotlb[i] = NULL;
+ }
+ }
}
diff --git a/lib/vhost/iotlb.h b/lib/vhost/iotlb.h
index 72232b0dcf08..52963d6c4de0 100644
--- a/lib/vhost/iotlb.h
+++ b/lib/vhost/iotlb.h
@@ -57,16 +57,16 @@ vhost_user_iotlb_wr_unlock_all(struct virtio_net *dev)
rte_rwlock_write_unlock(&dev->virtqueue[i]->iotlb_lock);
}
-void vhost_user_iotlb_cache_insert(struct virtio_net *dev, uint64_t iova, uint64_t uaddr,
+void vhost_user_iotlb_cache_insert(struct virtio_net *dev, int asid, uint64_t iova, uint64_t uaddr,
uint64_t uoffset, uint64_t size, uint64_t page_size, uint8_t perm);
-void vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t size);
-uint64_t vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova,
+void vhost_user_iotlb_cache_remove(struct virtio_net *dev, int asid, uint64_t iova, uint64_t size);
+uint64_t vhost_user_iotlb_cache_find(struct virtio_net *dev, int asid, uint64_t iova,
uint64_t *size, uint8_t perm);
-bool vhost_user_iotlb_pending_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm);
-void vhost_user_iotlb_pending_insert(struct virtio_net *dev, uint64_t iova, uint8_t perm);
-void vhost_user_iotlb_pending_remove(struct virtio_net *dev, uint64_t iova,
+bool vhost_user_iotlb_pending_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm);
+void vhost_user_iotlb_pending_insert(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm);
+void vhost_user_iotlb_pending_remove(struct virtio_net *dev, int asid, uint64_t iova,
uint64_t size, uint8_t perm);
-void vhost_user_iotlb_flush_all(struct virtio_net *dev);
+void vhost_user_iotlb_flush_all(struct virtio_net *dev, int asid);
int vhost_user_iotlb_init(struct virtio_net *dev);
void vhost_user_iotlb_destroy(struct virtio_net *dev);
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 0b5d158feeb9..7217895d1b9b 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -57,7 +57,7 @@ vduse_iotlb_remove_notify(uint64_t addr, uint64_t offset, uint64_t size)
}
static int
-vduse_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm __rte_unused)
+vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm __rte_unused)
{
struct vduse_iotlb_entry entry;
uint64_t size, page_size;
@@ -102,7 +102,7 @@ vduse_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm __rte_unuse
}
page_size = (uint64_t)stat.st_blksize;
- vhost_user_iotlb_cache_insert(dev, entry.start, (uint64_t)(uintptr_t)mmap_addr,
+ vhost_user_iotlb_cache_insert(dev, asid, entry.start, (uint64_t)(uintptr_t)mmap_addr,
entry.offset, size, page_size, entry.perm);
ret = 0;
@@ -398,7 +398,8 @@ vduse_device_stop(struct virtio_net *dev)
for (i = 0; i < dev->nr_vring; i++)
vduse_vring_cleanup(dev, i);
- vhost_user_iotlb_flush_all(dev);
+ for (i = 0; i < IOTLB_MAX_ASID; i++)
+ vhost_user_iotlb_flush_all(dev, i);
}
static void
@@ -445,7 +446,7 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
case VDUSE_UPDATE_IOTLB:
VHOST_CONFIG_LOG(dev->ifname, INFO, "\tIOVA range: %" PRIx64 " - %" PRIx64,
(uint64_t)req.iova.start, (uint64_t)req.iova.last);
- vhost_user_iotlb_cache_remove(dev, req.iova.start,
+ vhost_user_iotlb_cache_remove(dev, 0, req.iova.start,
req.iova.last - req.iova.start + 1);
resp.result = VDUSE_REQ_RESULT_OK;
break;
diff --git a/lib/vhost/vhost.c b/lib/vhost/vhost.c
index 7e68b2c3be92..cb3af28671cc 100644
--- a/lib/vhost/vhost.c
+++ b/lib/vhost/vhost.c
@@ -62,9 +62,9 @@ static const struct vhost_vq_stats_name_off vhost_vq_stat_strings[] = {
#define VHOST_NB_VQ_STATS RTE_DIM(vhost_vq_stat_strings)
static int
-vhost_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm)
{
- return dev->backend_ops->iotlb_miss(dev, iova, perm);
+ return dev->backend_ops->iotlb_miss(dev, asid, iova, perm);
}
uint64_t
@@ -78,7 +78,7 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
tmp_size = *size;
- vva = vhost_user_iotlb_cache_find(dev, iova, &tmp_size, perm);
+ vva = vhost_user_iotlb_cache_find(dev, vq->asid, iova, &tmp_size, perm);
if (tmp_size == *size) {
if (dev->flags & VIRTIO_DEV_STATS_ENABLED)
vq->stats.iotlb_hits++;
@@ -90,7 +90,7 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
iova += tmp_size;
- if (!vhost_user_iotlb_pending_miss(dev, iova, perm)) {
+ if (!vhost_user_iotlb_pending_miss(dev, vq->asid, iova, perm)) {
/*
* iotlb_lock is read-locked for a full burst,
* but it only protects the iotlb cache.
@@ -100,12 +100,12 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
*/
vhost_user_iotlb_rd_unlock(vq);
- vhost_user_iotlb_pending_insert(dev, iova, perm);
- if (vhost_iotlb_miss(dev, iova, perm)) {
+ vhost_user_iotlb_pending_insert(dev, vq->asid, iova, perm);
+ if (vhost_iotlb_miss(dev, vq->asid, iova, perm)) {
VHOST_DATA_LOG(dev->ifname, ERR,
"IOTLB miss req failed for IOVA 0x%" PRIx64,
iova);
- vhost_user_iotlb_pending_remove(dev, iova, 1, perm);
+ vhost_user_iotlb_pending_remove(dev, vq->asid, iova, 1, perm);
}
vhost_user_iotlb_rd_lock(vq);
@@ -113,7 +113,7 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
tmp_size = *size;
/* Retry in case of VDUSE, as it is synchronous */
- vva = vhost_user_iotlb_cache_find(dev, iova, &tmp_size, perm);
+ vva = vhost_user_iotlb_cache_find(dev, vq->asid, iova, &tmp_size, perm);
if (tmp_size == *size)
return vva;
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index ee61f7415ee3..0b3fca0ab769 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -85,7 +85,7 @@ struct vhost_virtqueue;
typedef void (*vhost_iotlb_remove_notify)(uint64_t addr, uint64_t off, uint64_t size);
-typedef int (*vhost_iotlb_miss_cb)(struct virtio_net *dev, uint64_t iova, uint8_t perm);
+typedef int (*vhost_iotlb_miss_cb)(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm);
typedef int (*vhost_vring_inject_irq_cb)(struct virtio_net *dev, struct vhost_virtqueue *vq);
/**
@@ -326,6 +326,7 @@ struct __rte_cache_aligned vhost_virtqueue {
uint16_t batch_copy_nb_elems;
struct batch_copy_elem *batch_copy_elems;
int numa_node;
+ int asid;
bool used_wrap_counter;
bool avail_wrap_counter;
@@ -483,6 +484,8 @@ struct inflight_mem_info {
uint64_t size;
};
+#define IOTLB_MAX_ASID 2
+
/**
* Device structure contains all configuration information relating
* to the device.
@@ -504,13 +507,7 @@ struct __rte_cache_aligned virtio_net {
int linearbuf;
struct vhost_virtqueue *virtqueue[VHOST_MAX_VRING];
- rte_rwlock_t iotlb_pending_lock;
- struct vhost_iotlb_entry *iotlb_pool;
- TAILQ_HEAD(, vhost_iotlb_entry) iotlb_list;
- TAILQ_HEAD(, vhost_iotlb_entry) iotlb_pending_list;
- int iotlb_cache_nr;
- rte_spinlock_t iotlb_free_lock;
- SLIST_HEAD(, vhost_iotlb_entry) iotlb_free_list;
+ struct iotlb *iotlb[IOTLB_MAX_ASID];
struct inflight_mem_info *inflight_info;
#define IF_NAME_SZ (PATH_MAX > IFNAMSIZ ? PATH_MAX : IFNAMSIZ)
diff --git a/lib/vhost/vhost_user.c b/lib/vhost/vhost_user.c
index 020c993b2991..e623651fc01e 100644
--- a/lib/vhost/vhost_user.c
+++ b/lib/vhost/vhost_user.c
@@ -1530,7 +1530,8 @@ vhost_user_set_mem_table(struct virtio_net **pdev,
/* Flush IOTLB cache as previous HVAs are now invalid */
if (dev->features & (1ULL << VIRTIO_F_IOMMU_PLATFORM))
- vhost_user_iotlb_flush_all(dev);
+ for (i = 0; i < IOTLB_MAX_ASID; i++)
+ vhost_user_iotlb_flush_all(dev, i);
free_all_mem_regions(dev);
rte_free(dev->mem);
@@ -1830,7 +1831,7 @@ vhost_user_rem_mem_reg(struct virtio_net **pdev,
if (dev->async_copy && rte_vfio_is_enabled("vfio"))
async_dma_map_region(dev, current_region, false);
if (dev->features & (1ULL << VIRTIO_F_IOMMU_PLATFORM))
- vhost_user_iotlb_cache_remove(dev,
+ vhost_user_iotlb_cache_remove(dev, 0,
current_region->guest_phys_addr,
current_region->size);
remove_guest_pages(dev, current_region);
@@ -2581,7 +2582,7 @@ vhost_user_get_vring_base(struct virtio_net **pdev,
ctx->msg.size = sizeof(ctx->msg.payload.state);
ctx->fd_num = 0;
- vhost_user_iotlb_flush_all(dev);
+ vhost_user_iotlb_flush_all(dev, vq->asid);
rte_rwlock_write_lock(&vq->access_lock);
vring_invalidate(dev, vq);
@@ -3030,7 +3031,7 @@ vhost_user_iotlb_msg(struct virtio_net **pdev,
pg_sz = hua_to_alignment(dev->mem, (void *)(uintptr_t)vva);
- vhost_user_iotlb_cache_insert(dev, imsg->iova, vva, 0, len, pg_sz, imsg->perm);
+ vhost_user_iotlb_cache_insert(dev, 0, imsg->iova, vva, 0, len, pg_sz, imsg->perm);
for (i = 0; i < dev->nr_vring; i++) {
struct vhost_virtqueue *vq = dev->virtqueue[i];
@@ -3047,7 +3048,7 @@ vhost_user_iotlb_msg(struct virtio_net **pdev,
}
break;
case VHOST_IOTLB_INVALIDATE:
- vhost_user_iotlb_cache_remove(dev, imsg->iova, imsg->size);
+ vhost_user_iotlb_cache_remove(dev, 0, imsg->iova, imsg->size);
for (i = 0; i < dev->nr_vring; i++) {
struct vhost_virtqueue *vq = dev->virtqueue[i];
@@ -3640,7 +3641,7 @@ vhost_user_msg_handler(int vid, int fd)
}
static int
-vhost_user_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_user_iotlb_miss(struct virtio_net *dev, int asid __rte_unused, uint64_t iova, uint8_t perm)
{
int ret;
struct vhu_msg_context ctx = {
--
2.55.0
^ permalink raw reply related
* [RFC v4 01/11] uapi: align VDUSE header for ASID
From: Eugenio Pérez @ 2026-07-07 12:44 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
In-Reply-To: <20260707124507.251729-1-eperezma@redhat.com>
Add all the ioctls and argument struct definitions so we can call them
in next patches.
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
kernel/linux/uapi/linux/vduse.h | 93 +++++++++++++++++++++++++++++----
1 file changed, 84 insertions(+), 9 deletions(-)
diff --git a/kernel/linux/uapi/linux/vduse.h b/kernel/linux/uapi/linux/vduse.h
index f46269af349a..361eea511c21 100644
--- a/kernel/linux/uapi/linux/vduse.h
+++ b/kernel/linux/uapi/linux/vduse.h
@@ -1,6 +1,6 @@
/* SPDX-License-Identifier: ((GPL-2.0 WITH Linux-syscall-note) OR BSD-3-Clause) */
-#ifndef _VDUSE_H_
-#define _VDUSE_H_
+#ifndef _UAPI_VDUSE_H_
+#define _UAPI_VDUSE_H_
#include <linux/types.h>
@@ -10,6 +10,10 @@
#define VDUSE_API_VERSION 0
+/* VQ groups and ASID support */
+
+#define VDUSE_API_VERSION_1 1
+
/*
* Get the version of VDUSE API that kernel supported (VDUSE_API_VERSION).
* This is used for future extension.
@@ -27,6 +31,8 @@
* @features: virtio features
* @vq_num: the number of virtqueues
* @vq_align: the allocation alignment of virtqueue's metadata
+ * @ngroups: number of vq groups that VDUSE device declares
+ * @nas: number of address spaces that VDUSE device declares
* @reserved: for future use, needs to be initialized to zero
* @config_size: the size of the configuration space
* @config: the buffer of the configuration space
@@ -41,7 +47,9 @@ struct vduse_dev_config {
__u64 features;
__u32 vq_num;
__u32 vq_align;
- __u32 reserved[13];
+ __u32 ngroups; /* if VDUSE_API_VERSION >= 1 */
+ __u32 nas; /* if VDUSE_API_VERSION >= 1 */
+ __u32 reserved[11];
__u32 config_size;
__u8 config[];
};
@@ -118,14 +126,18 @@ struct vduse_config_data {
* struct vduse_vq_config - basic configuration of a virtqueue
* @index: virtqueue index
* @max_size: the max size of virtqueue
- * @reserved: for future use, needs to be initialized to zero
+ * @reserved1: for future use, needs to be initialized to zero
+ * @group: virtqueue group
+ * @reserved2: for future use, needs to be initialized to zero
*
* Structure used by VDUSE_VQ_SETUP ioctl to setup a virtqueue.
*/
struct vduse_vq_config {
__u32 index;
__u16 max_size;
- __u16 reserved[13];
+ __u16 reserved1;
+ __u32 group;
+ __u16 reserved2[10];
};
/*
@@ -156,6 +168,16 @@ struct vduse_vq_state_packed {
__u16 last_used_idx;
};
+/**
+ * struct vduse_vq_group_asid - virtqueue group ASID
+ * @group: Index of the virtqueue group
+ * @asid: Address space ID of the group
+ */
+struct vduse_vq_group_asid {
+ __u32 group;
+ __u32 asid;
+};
+
/**
* struct vduse_vq_info - information of a virtqueue
* @index: virtqueue index
@@ -215,6 +237,7 @@ struct vduse_vq_eventfd {
* @uaddr: start address of userspace memory, it must be aligned to page size
* @iova: start of the IOVA region
* @size: size of the IOVA region
+ * @asid: Address space ID of the IOVA region
* @reserved: for future use, needs to be initialized to zero
*
* Structure used by VDUSE_IOTLB_REG_UMEM and VDUSE_IOTLB_DEREG_UMEM
@@ -224,7 +247,8 @@ struct vduse_iova_umem {
__u64 uaddr;
__u64 iova;
__u64 size;
- __u64 reserved[3];
+ __u32 asid;
+ __u32 reserved[5];
};
/* Register userspace memory for IOVA regions */
@@ -237,7 +261,8 @@ struct vduse_iova_umem {
* struct vduse_iova_info - information of one IOVA region
* @start: start of the IOVA region
* @last: last of the IOVA region
- * @capability: capability of the IOVA regsion
+ * @capability: capability of the IOVA region
+ * @asid: Address space ID of the IOVA region, only if device API version >= 1
* @reserved: for future use, needs to be initialized to zero
*
* Structure used by VDUSE_IOTLB_GET_INFO ioctl to get information of
@@ -248,7 +273,8 @@ struct vduse_iova_info {
__u64 last;
#define VDUSE_IOVA_CAP_UMEM (1 << 0)
__u64 capability;
- __u64 reserved[3];
+ __u32 asid; /* Only if device API version >= 1 */
+ __u32 reserved[5];
};
/*
@@ -257,6 +283,32 @@ struct vduse_iova_info {
*/
#define VDUSE_IOTLB_GET_INFO _IOWR(VDUSE_BASE, 0x1a, struct vduse_iova_info)
+/**
+ * struct vduse_iotlb_entry_v2 - entry of IOTLB to describe one IOVA region
+ *
+ * @v1: the original vduse_iotlb_entry
+ * @asid: address space ID of the IOVA region
+ * @reserved: for future use, needs to be initialized to zero
+ *
+ * Structure used by VDUSE_IOTLB_GET_FD2 ioctl to find an overlapped IOVA region.
+ */
+struct vduse_iotlb_entry_v2 {
+ __u64 offset;
+ __u64 start;
+ __u64 last;
+ __u8 perm;
+ __u8 padding[7];
+ __u32 asid;
+ __u32 reserved[11];
+};
+
+/*
+ * Same as VDUSE_IOTLB_GET_FD but with vduse_iotlb_entry_v2 argument that
+ * support extra fields.
+ */
+#define VDUSE_IOTLB_GET_FD2 _IOWR(VDUSE_BASE, 0x1b, struct vduse_iotlb_entry_v2)
+
+
/* The control messages definition for read(2)/write(2) on /dev/vduse/$NAME */
/**
@@ -265,11 +317,14 @@ struct vduse_iova_info {
* @VDUSE_SET_STATUS: set the device status
* @VDUSE_UPDATE_IOTLB: Notify userspace to update the memory mapping for
* specified IOVA range via VDUSE_IOTLB_GET_FD ioctl
+ * @VDUSE_SET_VQ_GROUP_ASID: Notify userspace to update the address space of a
+ * virtqueue group.
*/
enum vduse_req_type {
VDUSE_GET_VQ_STATE,
VDUSE_SET_STATUS,
VDUSE_UPDATE_IOTLB,
+ VDUSE_SET_VQ_GROUP_ASID,
};
/**
@@ -304,6 +359,19 @@ struct vduse_iova_range {
__u64 last;
};
+/**
+ * struct vduse_iova_range_v2 - IOVA range [start, last] if API_VERSION >= 1
+ * @start: start of the IOVA range
+ * @last: last of the IOVA range
+ * @asid: address space ID of the IOVA range
+ */
+struct vduse_iova_range_v2 {
+ __u64 start;
+ __u64 last;
+ __u32 asid;
+ __u32 padding;
+};
+
/**
* struct vduse_dev_request - control request
* @type: request type
@@ -312,6 +380,8 @@ struct vduse_iova_range {
* @vq_state: virtqueue state, only index field is available
* @s: device status
* @iova: IOVA range for updating
+ * @iova_v2: IOVA range for updating if API_VERSION >= 1
+ * @vq_group_asid: ASID of a virtqueue group
* @padding: padding
*
* Structure used by read(2) on /dev/vduse/$NAME.
@@ -324,6 +394,11 @@ struct vduse_dev_request {
struct vduse_vq_state vq_state;
struct vduse_dev_status s;
struct vduse_iova_range iova;
+ /* Following members but padding exist only if vduse api
+ * version >= 1
+ */
+ struct vduse_iova_range_v2 iova_v2;
+ struct vduse_vq_group_asid vq_group_asid;
__u32 padding[32];
};
};
@@ -350,4 +425,4 @@ struct vduse_dev_response {
};
};
-#endif /* _VDUSE_H_ */
+#endif /* _UAPI_VDUSE_H_ */
--
2.55.0
^ permalink raw reply related
* [RFC v4 00/11] Add vduse live migration features
From: Eugenio Pérez @ 2026-07-07 12:44 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: Yongji Xie, david.marchand, dev, mst, jasowangio, chenbox
This series introduces features to the VDUSE (vDPA Device in Userspace) driver
to support Live Migration.
Currently, DPDK does not support VDUSE devices live migration because the
driver lacks a mechanism to suspend the device and quiesce the rings to
initiate the switchover. This series implements the suspend operation to
address this limitation.
Furthermore, enabling Live Migration for devices with control virtqueue needs
two additional features. Both of them are included in this series.
* Address Spaces (ASID) support: This allows QEMU to isolate and intercept the
device's CVQ. By doing so, QEMU is able to migrate the device status
transparently, without requiring the device to support state save and
restore.
* QUEUE_ENABLE: This allows QEMU to control when the dataplane virtqueues are
enabled. This ensures the dataplane is started after the device
configuration has been fully restores via the CVQ.
Last but not least, it enables the VIRTIO_NET_F_STATUS feature. This allows the
device to signal the driver that it needs to send gratuitous ARP with
VIRTIO_NET_S_ANNOUNCE, reducing the Live Migration downtime.
v4:
* Sync headers with Linux's latest existing and proposed UAPI. Both
files constants and new ioctl VDUSE_SET_FEATURES.
* Check for more error conditions and clarified some error messages in
ready message processing.
* Add relevant release notes.
* Fix cosmetic whitespaces & checkpath errors.
* Fix error path of vhost_user_iotlb_init and vhost_user_iotlb_init_one.
* Fix commits author.
v3:
* Replace incorrect '%lx' DEBUG print format specifier with PRIx64
v2:
* Following latest comments on kernel series about VDUSE features, not checking
API version but only check if VDUSE_GET_FEATURES success.
* Move the start and last declarations in the braces as gcc 8 does not like
them interleaved with statements. Actually, I think the move was a mistake in
the first version.
https://mails.dpdk.org/archives/test-report/2026-February/958175.html
Eugenio Pérez (4):
uapi: align VDUSE header for ASID
vhost: Support VDUSE QUEUE_READY feature
vhost: Support vduse suspend feature
doc: add release notes for VDUSE live migration support
Maxime Coquelin (7):
vhost: introduce ASID support
vhost: add VDUSE API version negotiation
vhost: add virtqueues groups support to VDUSE
vhost: add ASID support to VDUSE IOTLB operations
vhost: claim VDUSE support for API version 1
vhost: add net status feature to VDUSE
uapi: Align vduse.h for enable and suspend VDUSE messages
doc/guides/rel_notes/release_26_07.rst | 17 ++
kernel/linux/uapi/linux/vduse.h | 121 ++++++++++++-
lib/vhost/iotlb.c | 233 +++++++++++++++----------
lib/vhost/iotlb.h | 14 +-
lib/vhost/vduse.c | 212 ++++++++++++++++++++--
lib/vhost/vduse.h | 3 +-
lib/vhost/vhost.c | 16 +-
lib/vhost/vhost.h | 16 +-
lib/vhost/vhost_user.c | 13 +-
9 files changed, 503 insertions(+), 142 deletions(-)
--
2.55.0
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox