* [PATCH v2 00/11] Add vduse live migration features
@ 2026-09-28 13:29 Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers Eugenio Pérez
` (10 more replies)
0 siblings, 11 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
This series introduces features to the VDUSE (vDPA Device in Userspace) driver
to support Live Migration.
Currently, DPDK does not support VDUSE devices live migration because the
driver lacks a mechanism to suspend the device and quiesce the rings to
initiate the switchover. This series implements the suspend operation to
address this limitation.
Furthermore, enabling Live Migration for devices with control virtqueue
requires two additional features. Both of them are included in this series.
* Address Spaces ID (ASID) support: This allows QEMU to isolate and intercept
the device's CVQ. By doing so, QEMU is able to migrate the device status
transparently, without requiring the device to support state save and
restore.
* QUEUE_READY: This allows QEMU to control when the dataplane virtqueues are
enabled. This ensures the dataplane is started after the device
configuration has been fully restored via the CVQ.
It also enables the VIRTIO_NET_F_STATUS feature. This allows the device to
signal the driver that it needs to send gratuitous ARP with
VIRTIO_NET_S_ANNOUNCE, reducing the Live Migration downtime.
Note that kernel headers are update to v7.3-rc3, not a stable version. This
would allow to ack the patches while the kernel uapi is published in 7.3.
v2:
* Add -D__counted_by(X)= if the compiler does not support it. It is
pulled by the kernel headers.
PATCH v1:
* Move to VDUSE_SET_FEATURES ioctl instead of config space field.
* Add a few bounds control.
* Update kernel headers to v7.3-rc3.
RFC v4:
* Sync headers with Linux's latest existing and proposed UAPI. Both
files constants and new ioctl VDUSE_SET_FEATURES.
* Check for more error conditions and clarified some error messages in
ready message processing.
* Add relevant release notes.
* Fix cosmetic whitespaces & checkpath errors.
* Fix error path of vhost_user_iotlb_init and vhost_user_iotlb_init_one.
* Fix commits author.
RFC v3:
* Replace incorrect '%lx' DEBUG print format specifier with PRIx64
RFC v2:
* Following latest comments on kernel series about VDUSE features, not checking
API version but only check if VDUSE_GET_FEATURES success.
* Move the start and last declarations in the braces as gcc 8 does not like
them interleaved with statements. Actually, I think the move was a mistake in
the first version.
https://mails.dpdk.org/archives/test-report/2026-February/958175.html
Eugenio Pérez (5):
build: define __counted_by as empty on unsupported compilers
uapi: import VDUSE and VFIO header from v7.3-rc3 kernel
vhost: Support VDUSE QUEUE_READY feature
vhost: Support vduse suspend feature
doc: add release notes for VDUSE live migration support
Maxime Coquelin (6):
vhost: introduce ASID support
vhost: add VDUSE API version negotiation
vhost: add virtqueues groups support to VDUSE
vhost: add ASID support to VDUSE IOTLB operations
vhost: claim VDUSE support for API version 1
vhost: add net status feature to VDUSE
config/meson.build | 6 +
doc/guides/rel_notes/release_26_07.rst | 8 +
kernel/linux/uapi/linux/vduse.h | 115 +++++++++++-
kernel/linux/uapi/linux/vfio.h | 91 +++++++++-
kernel/linux/uapi/version | 2 +-
lib/vhost/iotlb.c | 233 +++++++++++++++---------
lib/vhost/iotlb.h | 14 +-
lib/vhost/vduse.c | 239 +++++++++++++++++++++++--
lib/vhost/vduse.h | 3 +-
lib/vhost/vhost.c | 16 +-
lib/vhost/vhost.h | 16 +-
lib/vhost/vhost_user.c | 13 +-
12 files changed, 614 insertions(+), 142 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 23:42 ` Stephen Hemminger
2026-09-28 13:29 ` [PATCH v2 02/11] uapi: import VDUSE and VFIO header from v7.3-rc3 kernel Eugenio Pérez
` (9 subsequent siblings)
10 siblings, 1 reply; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
__counted_by is a GCC 14 / Clang 18 extension used in imported kernel
uapi headers. Older compilers reject it with an error. Define it as an
empty function-like macro when the compiler does not support it.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---
config/meson.build | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/config/meson.build b/config/meson.build
index 344f68822be3..9777713ea434 100644
--- a/config/meson.build
+++ b/config/meson.build
@@ -474,6 +474,12 @@ dpdk_arch_headers += files('rte_config.h')
# specify -D_GNU_SOURCE unconditionally
add_project_arguments('-D_GNU_SOURCE', language: 'c')
+# define __counted_by as empty on compilers that don't support it
+if not cc.compiles('struct s { int n; int a[] __counted_by(n); };',
+ name : '__counted_by support')
+ add_project_arguments('-D__counted_by(a)=', language : 'c')
+endif
+
# specify -D__BSD_VISIBLE for FreeBSD
if is_freebsd
add_project_arguments('-D__BSD_VISIBLE', language: 'c')
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 02/11] uapi: import VDUSE and VFIO header from v7.3-rc3 kernel
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 03/11] vhost: introduce ASID support Eugenio Pérez
` (8 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
This header will be used by the Vhost library.
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
kernel/linux/uapi/linux/vduse.h | 115 ++++++++++++++++++++++++++++++--
kernel/linux/uapi/linux/vfio.h | 91 ++++++++++++++++++++++++-
kernel/linux/uapi/version | 2 +-
3 files changed, 199 insertions(+), 9 deletions(-)
diff --git a/kernel/linux/uapi/linux/vduse.h b/kernel/linux/uapi/linux/vduse.h
index f46269af349a..bab47129db63 100644
--- a/kernel/linux/uapi/linux/vduse.h
+++ b/kernel/linux/uapi/linux/vduse.h
@@ -10,6 +10,16 @@
#define VDUSE_API_VERSION 0
+/* VQ groups and ASID support */
+
+#define VDUSE_API_VERSION_1 1
+
+/* The VDUSE instance expects a request for vq ready */
+#define VDUSE_F_QUEUE_READY 0
+
+/* The VDUSE instance expects a request for suspend */
+#define VDUSE_F_SUSPEND 1
+
/*
* Get the version of VDUSE API that kernel supported (VDUSE_API_VERSION).
* This is used for future extension.
@@ -27,6 +37,8 @@
* @features: virtio features
* @vq_num: the number of virtqueues
* @vq_align: the allocation alignment of virtqueue's metadata
+ * @ngroups: number of vq groups that VDUSE device declares
+ * @nas: number of address spaces that VDUSE device declares
* @reserved: for future use, needs to be initialized to zero
* @config_size: the size of the configuration space
* @config: the buffer of the configuration space
@@ -41,7 +53,9 @@ struct vduse_dev_config {
__u64 features;
__u32 vq_num;
__u32 vq_align;
- __u32 reserved[13];
+ __u32 ngroups; /* if VDUSE_API_VERSION >= 1 */
+ __u32 nas; /* if VDUSE_API_VERSION >= 1 */
+ __u32 reserved[11];
__u32 config_size;
__u8 config[];
};
@@ -55,6 +69,12 @@ struct vduse_dev_config {
*/
#define VDUSE_DESTROY_DEV _IOW(VDUSE_BASE, 0x03, char[VDUSE_NAME_MAX])
+/* Get the VDUSE supported features */
+#define VDUSE_GET_FEATURES _IOR(VDUSE_BASE, 0x04, __u64)
+
+/* Set the VDUSE features */
+#define VDUSE_SET_FEATURES _IOW(VDUSE_BASE, 0x05, __u64)
+
/* The ioctls for VDUSE device (/dev/vduse/$NAME) */
/**
@@ -118,14 +138,18 @@ struct vduse_config_data {
* struct vduse_vq_config - basic configuration of a virtqueue
* @index: virtqueue index
* @max_size: the max size of virtqueue
- * @reserved: for future use, needs to be initialized to zero
+ * @reserved1: for future use, needs to be initialized to zero
+ * @group: virtqueue group
+ * @reserved2: for future use, needs to be initialized to zero
*
* Structure used by VDUSE_VQ_SETUP ioctl to setup a virtqueue.
*/
struct vduse_vq_config {
__u32 index;
__u16 max_size;
- __u16 reserved[13];
+ __u16 reserved1;
+ __u32 group;
+ __u16 reserved2[10];
};
/*
@@ -156,6 +180,16 @@ struct vduse_vq_state_packed {
__u16 last_used_idx;
};
+/**
+ * struct vduse_vq_group_asid - virtqueue group ASID
+ * @group: Index of the virtqueue group
+ * @asid: Address space ID of the group
+ */
+struct vduse_vq_group_asid {
+ __u32 group;
+ __u32 asid;
+};
+
/**
* struct vduse_vq_info - information of a virtqueue
* @index: virtqueue index
@@ -215,6 +249,7 @@ struct vduse_vq_eventfd {
* @uaddr: start address of userspace memory, it must be aligned to page size
* @iova: start of the IOVA region
* @size: size of the IOVA region
+ * @asid: Address space ID of the IOVA region
* @reserved: for future use, needs to be initialized to zero
*
* Structure used by VDUSE_IOTLB_REG_UMEM and VDUSE_IOTLB_DEREG_UMEM
@@ -224,7 +259,8 @@ struct vduse_iova_umem {
__u64 uaddr;
__u64 iova;
__u64 size;
- __u64 reserved[3];
+ __u32 asid;
+ __u32 reserved[5];
};
/* Register userspace memory for IOVA regions */
@@ -237,7 +273,8 @@ struct vduse_iova_umem {
* struct vduse_iova_info - information of one IOVA region
* @start: start of the IOVA region
* @last: last of the IOVA region
- * @capability: capability of the IOVA regsion
+ * @capability: capability of the IOVA region
+ * @asid: Address space ID of the IOVA region, only if device API version >= 1
* @reserved: for future use, needs to be initialized to zero
*
* Structure used by VDUSE_IOTLB_GET_INFO ioctl to get information of
@@ -248,7 +285,8 @@ struct vduse_iova_info {
__u64 last;
#define VDUSE_IOVA_CAP_UMEM (1 << 0)
__u64 capability;
- __u64 reserved[3];
+ __u32 asid; /* Only if device API version >= 1 */
+ __u32 reserved[5];
};
/*
@@ -257,6 +295,32 @@ struct vduse_iova_info {
*/
#define VDUSE_IOTLB_GET_INFO _IOWR(VDUSE_BASE, 0x1a, struct vduse_iova_info)
+/**
+ * struct vduse_iotlb_entry_v2 - entry of IOTLB to describe one IOVA region
+ *
+ * @v1: the original vduse_iotlb_entry
+ * @asid: address space ID of the IOVA region
+ * @reserved: for future use, needs to be initialized to zero
+ *
+ * Structure used by VDUSE_IOTLB_GET_FD2 ioctl to find an overlapped IOVA region.
+ */
+struct vduse_iotlb_entry_v2 {
+ __u64 offset;
+ __u64 start;
+ __u64 last;
+ __u8 perm;
+ __u8 padding[7];
+ __u32 asid;
+ __u32 reserved[11];
+};
+
+/*
+ * Same as VDUSE_IOTLB_GET_FD but with vduse_iotlb_entry_v2 argument that
+ * support extra fields.
+ */
+#define VDUSE_IOTLB_GET_FD2 _IOWR(VDUSE_BASE, 0x1b, struct vduse_iotlb_entry_v2)
+
+
/* The control messages definition for read(2)/write(2) on /dev/vduse/$NAME */
/**
@@ -265,11 +329,16 @@ struct vduse_iova_info {
* @VDUSE_SET_STATUS: set the device status
* @VDUSE_UPDATE_IOTLB: Notify userspace to update the memory mapping for
* specified IOVA range via VDUSE_IOTLB_GET_FD ioctl
+ * @VDUSE_SET_VQ_GROUP_ASID: Notify userspace to update the address space of a
+ * virtqueue group.
*/
enum vduse_req_type {
VDUSE_GET_VQ_STATE,
VDUSE_SET_STATUS,
VDUSE_UPDATE_IOTLB,
+ VDUSE_SET_VQ_GROUP_ASID,
+ VDUSE_SET_VQ_READY,
+ VDUSE_SUSPEND,
};
/**
@@ -304,6 +373,28 @@ struct vduse_iova_range {
__u64 last;
};
+/**
+ * struct vduse_iova_range_v2 - IOVA range [start, last] if API_VERSION >= 1
+ * @start: start of the IOVA range
+ * @last: last of the IOVA range
+ * @asid: address space ID of the IOVA range
+ */
+struct vduse_iova_range_v2 {
+ __u64 start;
+ __u64 last;
+ __u32 asid;
+ __u32 padding;
+};
+
+/**
+ * struct vduse_vq_ready - Virtqueue ready request message
+ * @num: Virtqueue number
+ */
+struct vduse_vq_ready {
+ __u32 num;
+ __u32 ready;
+};
+
/**
* struct vduse_dev_request - control request
* @type: request type
@@ -312,6 +403,9 @@ struct vduse_iova_range {
* @vq_state: virtqueue state, only index field is available
* @s: device status
* @iova: IOVA range for updating
+ * @iova_v2: IOVA range for updating if API_VERSION >= 1
+ * @vq_group_asid: ASID of a virtqueue group
+ * @vq_ready: Virtqueue ready request
* @padding: padding
*
* Structure used by read(2) on /dev/vduse/$NAME.
@@ -324,6 +418,15 @@ struct vduse_dev_request {
struct vduse_vq_state vq_state;
struct vduse_dev_status s;
struct vduse_iova_range iova;
+ /* Following members but padding exist only if vduse api
+ * version >= 1
+ */
+ struct vduse_iova_range_v2 iova_v2;
+ struct vduse_vq_group_asid vq_group_asid;
+
+ /* Only if VDUSE_F_QUEUE_READY is negotiated */
+ struct vduse_vq_ready vq_ready;
+
__u32 padding[32];
};
};
diff --git a/kernel/linux/uapi/linux/vfio.h b/kernel/linux/uapi/linux/vfio.h
index 79bf8c0cc5e4..c85dcbfe302b 100644
--- a/kernel/linux/uapi/linux/vfio.h
+++ b/kernel/linux/uapi/linux/vfio.h
@@ -14,6 +14,7 @@
#include <linux/types.h>
#include <linux/ioctl.h>
+#include <linux/stddef.h>
#define VFIO_API_VERSION 0
@@ -140,7 +141,7 @@ struct vfio_info_cap_header {
*
* Retrieve information about the group. Fills in provided
* struct vfio_group_info. Caller sets argsz.
- * Return: 0 on succes, -errno on failure.
+ * Return: 0 on success, -errno on failure.
* Availability: Always
*/
struct vfio_group_status {
@@ -905,10 +906,12 @@ struct vfio_device_feature {
* VFIO_DEVICE_BIND_IOMMUFD - _IOR(VFIO_TYPE, VFIO_BASE + 18,
* struct vfio_device_bind_iommufd)
* @argsz: User filled size of this data.
- * @flags: Must be 0.
+ * @flags: Must be 0 or a bit flags of VFIO_DEVICE_BIND_*
* @iommufd: iommufd to bind.
* @out_devid: The device id generated by this bind. devid is a handle for
* this device/iommufd bond and can be used in IOMMUFD commands.
+ * @token_uuid_ptr: Valid if VFIO_DEVICE_BIND_FLAG_TOKEN. Points to a 16 byte
+ * UUID in the same format as VFIO_DEVICE_FEATURE_PCI_VF_TOKEN.
*
* Bind a vfio_device to the specified iommufd.
*
@@ -917,13 +920,21 @@ struct vfio_device_feature {
*
* Unbind is automatically conducted when device fd is closed.
*
+ * A token is sometimes required to open the device, unless this is known to be
+ * needed VFIO_DEVICE_BIND_FLAG_TOKEN should not be set and token_uuid_ptr is
+ * ignored. The only case today is a PF/VF relationship where the VF bind must
+ * be provided the same token as VFIO_DEVICE_FEATURE_PCI_VF_TOKEN provided to
+ * the PF.
+ *
* Return: 0 on success, -errno on failure.
*/
struct vfio_device_bind_iommufd {
__u32 argsz;
__u32 flags;
+#define VFIO_DEVICE_BIND_FLAG_TOKEN (1 << 0)
__s32 iommufd;
__u32 out_devid;
+ __aligned_u64 token_uuid_ptr;
};
#define VFIO_DEVICE_BIND_IOMMUFD _IO(VFIO_TYPE, VFIO_BASE + 18)
@@ -953,6 +964,10 @@ struct vfio_device_bind_iommufd {
* hwpt corresponding to the given pt_id.
*
* Return: 0 on success, -errno on failure.
+ *
+ * When a device is resetting, -EBUSY will be returned to reject any concurrent
+ * attachment to the resetting device itself or any sibling device in the IOMMU
+ * group having the resetting device.
*/
struct vfio_device_attach_iommufd_pt {
__u32 argsz;
@@ -1251,6 +1266,19 @@ enum vfio_device_mig_state {
* The initial_bytes field indicates the amount of initial precopy
* data available from the device. This field should have a non-zero initial
* value and decrease as migration data is read from the device.
+ * The presence of the VFIO_PRECOPY_INFO_REINIT output flag indicates
+ * that new initial data is present on the stream.
+ * The new initial data may result, for example, from device reconfiguration
+ * during migration that requires additional initialization data.
+ * In that case initial_bytes may report a non-zero value irrespective of
+ * any previously reported values, which progresses towards zero as precopy
+ * data is read from the data stream. dirty_bytes is also reset
+ * to zero and represents the state change of the device relative to the new
+ * initial_bytes.
+ * VFIO_PRECOPY_INFO_REINIT can be reported only after userspace opts in to
+ * VFIO_DEVICE_FEATURE_MIG_PRECOPY_INFOv2. Without this opt-in, the flags field
+ * of struct vfio_precopy_info is reserved for bug-compatibility reasons.
+ *
* It is recommended to leave PRE_COPY for STOP_COPY only after this field
* reaches zero. Leaving PRE_COPY earlier might make things slower.
*
@@ -1286,6 +1314,7 @@ enum vfio_device_mig_state {
struct vfio_precopy_info {
__u32 argsz;
__u32 flags;
+#define VFIO_PRECOPY_INFO_REINIT (1 << 0) /* output - new initial data is present */
__aligned_u64 initial_bytes;
__aligned_u64 dirty_bytes;
};
@@ -1468,6 +1497,64 @@ struct vfio_device_feature_bus_master {
};
#define VFIO_DEVICE_FEATURE_BUS_MASTER 10
+/**
+ * Upon VFIO_DEVICE_FEATURE_GET create a dma_buf fd for the
+ * regions selected.
+ *
+ * open_flags are the typical flags passed to open(2), eg O_RDWR, O_CLOEXEC,
+ * etc. offset/length specify a slice of the region to create the dmabuf from.
+ * nr_ranges is the total number of (P2P DMA) ranges that comprise the dmabuf.
+ *
+ * flags should be 0.
+ *
+ * Return: The fd number on success, -1 and errno is set on failure.
+ */
+#define VFIO_DEVICE_FEATURE_DMA_BUF 11
+
+struct vfio_region_dma_range {
+ __u64 offset;
+ __u64 length;
+};
+
+struct vfio_device_feature_dma_buf {
+ __u32 region_index;
+ __u32 open_flags;
+ __u32 flags;
+ __u32 nr_ranges;
+ struct vfio_region_dma_range dma_ranges[] __counted_by(nr_ranges);
+};
+
+/*
+ * Enables the migration precopy_info_v2 behaviour.
+ *
+ * VFIO_DEVICE_FEATURE_MIG_PRECOPY_INFOv2.
+ *
+ * On SET, enables the v2 pre_copy_info behaviour, where the
+ * vfio_precopy_info.flags is a valid output field.
+ */
+#define VFIO_DEVICE_FEATURE_MIG_PRECOPY_INFOv2 12
+
+/**
+ * VFIO_DEVICE_FEATURE_ZPCI_ERROR feature provides PCI error information to
+ * userspace for vfio-pci devices on s390. On s390, PCI error recovery
+ * involves platform firmware and notification to operating systems is done
+ * by architecture specific mechanism. Exposing this information to
+ * userspace allows it to take appropriate actions to handle an
+ * error on the device.
+ *
+ * Userspace provides an opaque buffer of fixed length, and the kernel
+ * fills it with the zpci_ccdf_err data structure. The length of
+ * zpci_ccdf_err is provided to userspace via the
+ * VFIO_DEVICE_INFO_CAP_ZPCI_BASE capability.
+ *
+ * The ioctl returns -ENOMSG if there are no pending PCI errors.
+ */
+struct vfio_device_feature_zpci_err {
+ __aligned_u64 data;
+};
+
+#define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13
+
/* -------- API for Type1 VFIO IOMMU -------- */
/**
diff --git a/kernel/linux/uapi/version b/kernel/linux/uapi/version
index 966a9983019b..4e528e220b99 100644
--- a/kernel/linux/uapi/version
+++ b/kernel/linux/uapi/version
@@ -1 +1 @@
-v6.16
+v7.3-rc3
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 03/11] vhost: introduce ASID support
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 02/11] uapi: import VDUSE and VFIO header from v7.3-rc3 kernel Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 04/11] vhost: add VDUSE API version negotiation Eugenio Pérez
` (7 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Set all ASID = 0 as it is the default when not set explicitly.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/iotlb.c | 233 +++++++++++++++++++++++++----------------
lib/vhost/iotlb.h | 14 +--
lib/vhost/vduse.c | 9 +-
lib/vhost/vhost.c | 16 +--
lib/vhost/vhost.h | 13 +--
lib/vhost/vhost_user.c | 13 +--
6 files changed, 177 insertions(+), 121 deletions(-)
diff --git a/lib/vhost/iotlb.c b/lib/vhost/iotlb.c
index f2c275a7d77e..e5e69d11cb65 100644
--- a/lib/vhost/iotlb.c
+++ b/lib/vhost/iotlb.c
@@ -11,6 +11,16 @@
#include "iotlb.h"
#include "vhost.h"
+struct iotlb {
+ rte_rwlock_t pending_lock;
+ struct vhost_iotlb_entry *pool;
+ TAILQ_HEAD(, vhost_iotlb_entry) list;
+ TAILQ_HEAD(, vhost_iotlb_entry) pending_list;
+ int cache_nr;
+ rte_spinlock_t free_lock;
+ SLIST_HEAD(, vhost_iotlb_entry) free_list;
+};
+
struct vhost_iotlb_entry {
TAILQ_ENTRY(vhost_iotlb_entry) next;
SLIST_ENTRY(vhost_iotlb_entry) next_free;
@@ -85,78 +95,78 @@ vhost_user_iotlb_clear_dump(struct virtio_net *dev, struct vhost_iotlb_entry *no
}
static struct vhost_iotlb_entry *
-vhost_user_iotlb_pool_get(struct virtio_net *dev)
+vhost_user_iotlb_pool_get(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node;
- rte_spinlock_lock(&dev->iotlb_free_lock);
- node = SLIST_FIRST(&dev->iotlb_free_list);
+ rte_spinlock_lock(&dev->iotlb[asid]->free_lock);
+ node = SLIST_FIRST(&dev->iotlb[asid]->free_list);
if (node != NULL)
- SLIST_REMOVE_HEAD(&dev->iotlb_free_list, next_free);
- rte_spinlock_unlock(&dev->iotlb_free_lock);
+ SLIST_REMOVE_HEAD(&dev->iotlb[asid]->free_list, next_free);
+ rte_spinlock_unlock(&dev->iotlb[asid]->free_lock);
return node;
}
static void
-vhost_user_iotlb_pool_put(struct virtio_net *dev, struct vhost_iotlb_entry *node)
+vhost_user_iotlb_pool_put(struct virtio_net *dev, int asid, struct vhost_iotlb_entry *node)
{
- rte_spinlock_lock(&dev->iotlb_free_lock);
- SLIST_INSERT_HEAD(&dev->iotlb_free_list, node, next_free);
- rte_spinlock_unlock(&dev->iotlb_free_lock);
+ rte_spinlock_lock(&dev->iotlb[asid]->free_lock);
+ SLIST_INSERT_HEAD(&dev->iotlb[asid]->free_list, node, next_free);
+ rte_spinlock_unlock(&dev->iotlb[asid]->free_lock);
}
static void
-vhost_user_iotlb_cache_random_evict(struct virtio_net *dev);
+vhost_user_iotlb_cache_random_evict(struct virtio_net *dev, int asid);
static void
-vhost_user_iotlb_pending_remove_all(struct virtio_net *dev)
+vhost_user_iotlb_pending_remove_all(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node, *temp_node;
- rte_rwlock_write_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_lock(&dev->iotlb[asid]->pending_lock);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_pending_list, next, temp_node) {
- TAILQ_REMOVE(&dev->iotlb_pending_list, node, next);
- vhost_user_iotlb_pool_put(dev, node);
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->pending_list, next, temp_node) {
+ TAILQ_REMOVE(&dev->iotlb[asid]->pending_list, node, next);
+ vhost_user_iotlb_pool_put(dev, asid, node);
}
- rte_rwlock_write_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_unlock(&dev->iotlb[asid]->pending_lock);
}
bool
-vhost_user_iotlb_pending_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_user_iotlb_pending_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm)
{
struct vhost_iotlb_entry *node;
bool found = false;
- rte_rwlock_read_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_read_lock(&dev->iotlb[asid]->pending_lock);
- TAILQ_FOREACH(node, &dev->iotlb_pending_list, next) {
+ TAILQ_FOREACH(node, &dev->iotlb[asid]->pending_list, next) {
if ((node->iova == iova) && (node->perm == perm)) {
found = true;
break;
}
}
- rte_rwlock_read_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_read_unlock(&dev->iotlb[asid]->pending_lock);
return found;
}
void
-vhost_user_iotlb_pending_insert(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_user_iotlb_pending_insert(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm)
{
struct vhost_iotlb_entry *node;
- node = vhost_user_iotlb_pool_get(dev);
+ node = vhost_user_iotlb_pool_get(dev, asid);
if (node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, DEBUG,
"IOTLB pool empty, clear entries for pending insertion");
- if (!TAILQ_EMPTY(&dev->iotlb_pending_list))
- vhost_user_iotlb_pending_remove_all(dev);
+ if (!TAILQ_EMPTY(&dev->iotlb[asid]->pending_list))
+ vhost_user_iotlb_pending_remove_all(dev, asid);
else
- vhost_user_iotlb_cache_random_evict(dev);
- node = vhost_user_iotlb_pool_get(dev);
+ vhost_user_iotlb_cache_random_evict(dev, asid);
+ node = vhost_user_iotlb_pool_get(dev, asid);
if (node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, ERR,
"IOTLB pool still empty, pending insertion failure");
@@ -167,21 +177,22 @@ vhost_user_iotlb_pending_insert(struct virtio_net *dev, uint64_t iova, uint8_t p
node->iova = iova;
node->perm = perm;
- rte_rwlock_write_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_lock(&dev->iotlb[asid]->pending_lock);
- TAILQ_INSERT_TAIL(&dev->iotlb_pending_list, node, next);
+ TAILQ_INSERT_TAIL(&dev->iotlb[asid]->pending_list, node, next);
- rte_rwlock_write_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_unlock(&dev->iotlb[asid]->pending_lock);
}
void
-vhost_user_iotlb_pending_remove(struct virtio_net *dev, uint64_t iova, uint64_t size, uint8_t perm)
+vhost_user_iotlb_pending_remove(struct virtio_net *dev, int asid,
+ uint64_t iova, uint64_t size, uint8_t perm)
{
struct vhost_iotlb_entry *node, *temp_node;
- rte_rwlock_write_lock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_lock(&dev->iotlb[asid]->pending_lock);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_pending_list, next,
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->pending_list, next,
temp_node) {
if (node->iova < iova)
continue;
@@ -189,53 +200,53 @@ vhost_user_iotlb_pending_remove(struct virtio_net *dev, uint64_t iova, uint64_t
continue;
if ((node->perm & perm) != node->perm)
continue;
- TAILQ_REMOVE(&dev->iotlb_pending_list, node, next);
- vhost_user_iotlb_pool_put(dev, node);
+ TAILQ_REMOVE(&dev->iotlb[asid]->pending_list, node, next);
+ vhost_user_iotlb_pool_put(dev, asid, node);
}
- rte_rwlock_write_unlock(&dev->iotlb_pending_lock);
+ rte_rwlock_write_unlock(&dev->iotlb[asid]->pending_lock);
}
static void
-vhost_user_iotlb_cache_remove_all(struct virtio_net *dev)
+vhost_user_iotlb_cache_remove_all(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node, *temp_node;
vhost_user_iotlb_wr_lock_all(dev);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_list, next, temp_node) {
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->list, next, temp_node) {
vhost_user_iotlb_clear_dump(dev, node, NULL, NULL);
- TAILQ_REMOVE(&dev->iotlb_list, node, next);
+ TAILQ_REMOVE(&dev->iotlb[asid]->list, node, next);
vhost_user_iotlb_remove_notify(dev, node);
- vhost_user_iotlb_pool_put(dev, node);
+ vhost_user_iotlb_pool_put(dev, asid, node);
}
- dev->iotlb_cache_nr = 0;
+ dev->iotlb[asid]->cache_nr = 0;
vhost_user_iotlb_wr_unlock_all(dev);
}
static void
-vhost_user_iotlb_cache_random_evict(struct virtio_net *dev)
+vhost_user_iotlb_cache_random_evict(struct virtio_net *dev, int asid)
{
struct vhost_iotlb_entry *node, *temp_node, *prev_node = NULL;
int entry_idx;
vhost_user_iotlb_wr_lock_all(dev);
- entry_idx = rte_rand() % dev->iotlb_cache_nr;
+ entry_idx = rte_rand() % dev->iotlb[asid]->cache_nr;
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_list, next, temp_node) {
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->list, next, temp_node) {
if (!entry_idx) {
struct vhost_iotlb_entry *next_node = RTE_TAILQ_NEXT(node, next);
vhost_user_iotlb_clear_dump(dev, node, prev_node, next_node);
- TAILQ_REMOVE(&dev->iotlb_list, node, next);
+ TAILQ_REMOVE(&dev->iotlb[asid]->list, node, next);
vhost_user_iotlb_remove_notify(dev, node);
- vhost_user_iotlb_pool_put(dev, node);
- dev->iotlb_cache_nr--;
+ vhost_user_iotlb_pool_put(dev, asid, node);
+ dev->iotlb[asid]->cache_nr--;
break;
}
prev_node = node;
@@ -246,20 +257,20 @@ vhost_user_iotlb_cache_random_evict(struct virtio_net *dev)
}
void
-vhost_user_iotlb_cache_insert(struct virtio_net *dev, uint64_t iova, uint64_t uaddr,
+vhost_user_iotlb_cache_insert(struct virtio_net *dev, int asid, uint64_t iova, uint64_t uaddr,
uint64_t uoffset, uint64_t size, uint64_t page_size, uint8_t perm)
{
struct vhost_iotlb_entry *node, *new_node;
- new_node = vhost_user_iotlb_pool_get(dev);
+ new_node = vhost_user_iotlb_pool_get(dev, asid);
if (new_node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, DEBUG,
"IOTLB pool empty, clear entries for cache insertion");
- if (!TAILQ_EMPTY(&dev->iotlb_list))
- vhost_user_iotlb_cache_random_evict(dev);
+ if (!TAILQ_EMPTY(&dev->iotlb[asid]->list))
+ vhost_user_iotlb_cache_random_evict(dev, asid);
else
- vhost_user_iotlb_pending_remove_all(dev);
- new_node = vhost_user_iotlb_pool_get(dev);
+ vhost_user_iotlb_pending_remove_all(dev, asid);
+ new_node = vhost_user_iotlb_pool_get(dev, asid);
if (new_node == NULL) {
VHOST_CONFIG_LOG(dev->ifname, ERR,
"IOTLB pool still empty, cache insertion failed");
@@ -276,36 +287,36 @@ vhost_user_iotlb_cache_insert(struct virtio_net *dev, uint64_t iova, uint64_t ua
vhost_user_iotlb_wr_lock_all(dev);
- TAILQ_FOREACH(node, &dev->iotlb_list, next) {
+ TAILQ_FOREACH(node, &dev->iotlb[asid]->list, next) {
/*
* Entries must be invalidated before being updated.
* So if iova already in list, assume identical.
*/
if (node->iova == new_node->iova) {
- vhost_user_iotlb_pool_put(dev, new_node);
+ vhost_user_iotlb_pool_put(dev, asid, new_node);
goto unlock;
} else if (node->iova > new_node->iova) {
vhost_user_iotlb_set_dump(dev, new_node);
TAILQ_INSERT_BEFORE(node, new_node, next);
- dev->iotlb_cache_nr++;
+ dev->iotlb[asid]->cache_nr++;
goto unlock;
}
}
vhost_user_iotlb_set_dump(dev, new_node);
- TAILQ_INSERT_TAIL(&dev->iotlb_list, new_node, next);
- dev->iotlb_cache_nr++;
+ TAILQ_INSERT_TAIL(&dev->iotlb[asid]->list, new_node, next);
+ dev->iotlb[asid]->cache_nr++;
unlock:
- vhost_user_iotlb_pending_remove(dev, iova, size, perm);
+ vhost_user_iotlb_pending_remove(dev, asid, iova, size, perm);
vhost_user_iotlb_wr_unlock_all(dev);
}
void
-vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t size)
+vhost_user_iotlb_cache_remove(struct virtio_net *dev, int asid, uint64_t iova, uint64_t size)
{
struct vhost_iotlb_entry *node, *temp_node, *prev_node = NULL;
@@ -314,7 +325,7 @@ vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t si
vhost_user_iotlb_wr_lock_all(dev);
- RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb_list, next, temp_node) {
+ RTE_TAILQ_FOREACH_SAFE(node, &dev->iotlb[asid]->list, next, temp_node) {
/* Sorted list */
if (unlikely(iova + size < node->iova))
break;
@@ -324,10 +335,10 @@ vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t si
vhost_user_iotlb_clear_dump(dev, node, prev_node, next_node);
- TAILQ_REMOVE(&dev->iotlb_list, node, next);
+ TAILQ_REMOVE(&dev->iotlb[asid]->list, node, next);
vhost_user_iotlb_remove_notify(dev, node);
- vhost_user_iotlb_pool_put(dev, node);
- dev->iotlb_cache_nr--;
+ vhost_user_iotlb_pool_put(dev, asid, node);
+ dev->iotlb[asid]->cache_nr--;
} else {
prev_node = node;
}
@@ -337,7 +348,8 @@ vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t si
}
uint64_t
-vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova, uint64_t *size, uint8_t perm)
+vhost_user_iotlb_cache_find(struct virtio_net *dev, int asid,
+ uint64_t iova, uint64_t *size, uint8_t perm)
{
struct vhost_iotlb_entry *node;
uint64_t offset, vva = 0, mapped = 0;
@@ -345,7 +357,7 @@ vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova, uint64_t *siz
if (unlikely(!*size))
goto out;
- TAILQ_FOREACH(node, &dev->iotlb_list, next) {
+ TAILQ_FOREACH(node, &dev->iotlb[asid]->list, next) {
/* List sorted by iova */
if (unlikely(iova < node->iova))
break;
@@ -378,25 +390,28 @@ vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova, uint64_t *siz
}
void
-vhost_user_iotlb_flush_all(struct virtio_net *dev)
+vhost_user_iotlb_flush_all(struct virtio_net *dev, int asid)
{
- vhost_user_iotlb_cache_remove_all(dev);
- vhost_user_iotlb_pending_remove_all(dev);
+ vhost_user_iotlb_cache_remove_all(dev, asid);
+ vhost_user_iotlb_pending_remove_all(dev, asid);
}
-int
-vhost_user_iotlb_init(struct virtio_net *dev)
+static int
+vhost_user_iotlb_init_one(struct virtio_net *dev, int asid)
{
unsigned int i;
int socket = 0;
- if (dev->iotlb_pool) {
- /*
- * The cache has already been initialized,
- * just drop all cached and pending entries.
- */
- vhost_user_iotlb_flush_all(dev);
- rte_free(dev->iotlb_pool);
+ if (dev->iotlb[asid] != NULL) {
+ if (dev->iotlb[asid]->pool != NULL) {
+ /*
+ * The cache has already been initialized,
+ * just drop all cached and pending entries.
+ */
+ vhost_user_iotlb_flush_all(dev, asid);
+ rte_free(dev->iotlb[asid]->pool);
+ }
+ rte_free(dev->iotlb[asid]);
}
#ifdef RTE_LIBRTE_VHOST_NUMA
@@ -404,31 +419,73 @@ vhost_user_iotlb_init(struct virtio_net *dev)
socket = 0;
#endif
- rte_spinlock_init(&dev->iotlb_free_lock);
- rte_rwlock_init(&dev->iotlb_pending_lock);
+ dev->iotlb[asid] = rte_malloc_socket("iotlb", sizeof(struct iotlb), 0, socket);
+ if (!dev->iotlb[asid]) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR, "Failed to allocate IOTLB");
+ return -1;
+ }
+
+ rte_spinlock_init(&dev->iotlb[asid]->free_lock);
+ rte_rwlock_init(&dev->iotlb[asid]->pending_lock);
- SLIST_INIT(&dev->iotlb_free_list);
- TAILQ_INIT(&dev->iotlb_list);
- TAILQ_INIT(&dev->iotlb_pending_list);
+ SLIST_INIT(&dev->iotlb[asid]->free_list);
+ TAILQ_INIT(&dev->iotlb[asid]->list);
+ TAILQ_INIT(&dev->iotlb[asid]->pending_list);
if (dev->flags & VIRTIO_DEV_SUPPORT_IOMMU) {
- dev->iotlb_pool = rte_calloc_socket("iotlb", IOTLB_CACHE_SIZE,
+ dev->iotlb[asid]->pool = rte_calloc_socket("iotlb_pool", IOTLB_CACHE_SIZE,
sizeof(struct vhost_iotlb_entry), 0, socket);
- if (!dev->iotlb_pool) {
+ if (!dev->iotlb[asid]->pool) {
VHOST_CONFIG_LOG(dev->ifname, ERR, "Failed to create IOTLB cache pool");
- return -1;
+ goto free_iotlb;
}
for (i = 0; i < IOTLB_CACHE_SIZE; i++)
- vhost_user_iotlb_pool_put(dev, &dev->iotlb_pool[i]);
+ vhost_user_iotlb_pool_put(dev, asid, &dev->iotlb[asid]->pool[i]);
}
- dev->iotlb_cache_nr = 0;
+ dev->iotlb[asid]->cache_nr = 0;
+
+ return 0;
+
+free_iotlb:
+ rte_free(dev->iotlb[asid]);
+ dev->iotlb[asid] = NULL;
+ return -1;
+}
+
+int
+vhost_user_iotlb_init(struct virtio_net *dev)
+{
+ int i;
+
+ for (i = 0; i < IOTLB_MAX_ASID; i++)
+ if (vhost_user_iotlb_init_one(dev, i) < 0)
+ goto fail;
return 0;
+fail:
+ while (i--) {
+ rte_free(dev->iotlb[i]->pool);
+ dev->iotlb[i]->pool = NULL;
+ rte_free(dev->iotlb[i]);
+ dev->iotlb[i] = NULL;
+ }
+
+ return -1;
}
void
vhost_user_iotlb_destroy(struct virtio_net *dev)
{
- rte_free(dev->iotlb_pool);
+ int i;
+
+ for (i = 0; i < IOTLB_MAX_ASID; i++) {
+ if (dev->iotlb[i]) {
+ rte_free(dev->iotlb[i]->pool);
+ dev->iotlb[i]->pool = NULL;
+
+ rte_free(dev->iotlb[i]);
+ dev->iotlb[i] = NULL;
+ }
+ }
}
diff --git a/lib/vhost/iotlb.h b/lib/vhost/iotlb.h
index 72232b0dcf08..52963d6c4de0 100644
--- a/lib/vhost/iotlb.h
+++ b/lib/vhost/iotlb.h
@@ -57,16 +57,16 @@ vhost_user_iotlb_wr_unlock_all(struct virtio_net *dev)
rte_rwlock_write_unlock(&dev->virtqueue[i]->iotlb_lock);
}
-void vhost_user_iotlb_cache_insert(struct virtio_net *dev, uint64_t iova, uint64_t uaddr,
+void vhost_user_iotlb_cache_insert(struct virtio_net *dev, int asid, uint64_t iova, uint64_t uaddr,
uint64_t uoffset, uint64_t size, uint64_t page_size, uint8_t perm);
-void vhost_user_iotlb_cache_remove(struct virtio_net *dev, uint64_t iova, uint64_t size);
-uint64_t vhost_user_iotlb_cache_find(struct virtio_net *dev, uint64_t iova,
+void vhost_user_iotlb_cache_remove(struct virtio_net *dev, int asid, uint64_t iova, uint64_t size);
+uint64_t vhost_user_iotlb_cache_find(struct virtio_net *dev, int asid, uint64_t iova,
uint64_t *size, uint8_t perm);
-bool vhost_user_iotlb_pending_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm);
-void vhost_user_iotlb_pending_insert(struct virtio_net *dev, uint64_t iova, uint8_t perm);
-void vhost_user_iotlb_pending_remove(struct virtio_net *dev, uint64_t iova,
+bool vhost_user_iotlb_pending_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm);
+void vhost_user_iotlb_pending_insert(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm);
+void vhost_user_iotlb_pending_remove(struct virtio_net *dev, int asid, uint64_t iova,
uint64_t size, uint8_t perm);
-void vhost_user_iotlb_flush_all(struct virtio_net *dev);
+void vhost_user_iotlb_flush_all(struct virtio_net *dev, int asid);
int vhost_user_iotlb_init(struct virtio_net *dev);
void vhost_user_iotlb_destroy(struct virtio_net *dev);
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index f8a4a8edcbf2..3c72bafcf511 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -57,7 +57,7 @@ vduse_iotlb_remove_notify(uint64_t addr, uint64_t offset, uint64_t size)
}
static int
-vduse_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm __rte_unused)
+vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm __rte_unused)
{
struct vduse_iotlb_entry entry;
uint64_t size, page_size;
@@ -102,7 +102,7 @@ vduse_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm __rte_unuse
}
page_size = (uint64_t)stat.st_blksize;
- vhost_user_iotlb_cache_insert(dev, entry.start, (uint64_t)(uintptr_t)mmap_addr,
+ vhost_user_iotlb_cache_insert(dev, asid, entry.start, (uint64_t)(uintptr_t)mmap_addr,
entry.offset, size, page_size, entry.perm);
ret = 0;
@@ -398,7 +398,8 @@ vduse_device_stop(struct virtio_net *dev)
for (i = 0; i < dev->nr_vring; i++)
vduse_vring_cleanup(dev, i);
- vhost_user_iotlb_flush_all(dev);
+ for (i = 0; i < IOTLB_MAX_ASID; i++)
+ vhost_user_iotlb_flush_all(dev, i);
}
static void
@@ -445,7 +446,7 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
case VDUSE_UPDATE_IOTLB:
VHOST_CONFIG_LOG(dev->ifname, INFO, "\tIOVA range: %" PRIx64 " - %" PRIx64,
(uint64_t)req.iova.start, (uint64_t)req.iova.last);
- vhost_user_iotlb_cache_remove(dev, req.iova.start,
+ vhost_user_iotlb_cache_remove(dev, 0, req.iova.start,
req.iova.last - req.iova.start + 1);
resp.result = VDUSE_REQ_RESULT_OK;
break;
diff --git a/lib/vhost/vhost.c b/lib/vhost/vhost.c
index fde8acb00cde..8bfdbee894b5 100644
--- a/lib/vhost/vhost.c
+++ b/lib/vhost/vhost.c
@@ -62,9 +62,9 @@ static const struct vhost_vq_stats_name_off vhost_vq_stat_strings[] = {
#define VHOST_NB_VQ_STATS RTE_DIM(vhost_vq_stat_strings)
static int
-vhost_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm)
{
- return dev->backend_ops->iotlb_miss(dev, iova, perm);
+ return dev->backend_ops->iotlb_miss(dev, asid, iova, perm);
}
uint64_t
@@ -78,7 +78,7 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
tmp_size = *size;
- vva = vhost_user_iotlb_cache_find(dev, iova, &tmp_size, perm);
+ vva = vhost_user_iotlb_cache_find(dev, vq->asid, iova, &tmp_size, perm);
if (tmp_size == *size) {
if (dev->flags & VIRTIO_DEV_STATS_ENABLED)
vq->stats.iotlb_hits++;
@@ -90,7 +90,7 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
iova += tmp_size;
- if (!vhost_user_iotlb_pending_miss(dev, iova, perm)) {
+ if (!vhost_user_iotlb_pending_miss(dev, vq->asid, iova, perm)) {
/*
* iotlb_lock is read-locked for a full burst,
* but it only protects the iotlb cache.
@@ -100,12 +100,12 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
*/
vhost_user_iotlb_rd_unlock(vq);
- vhost_user_iotlb_pending_insert(dev, iova, perm);
- if (vhost_iotlb_miss(dev, iova, perm)) {
+ vhost_user_iotlb_pending_insert(dev, vq->asid, iova, perm);
+ if (vhost_iotlb_miss(dev, vq->asid, iova, perm)) {
VHOST_DATA_LOG(dev->ifname, ERR,
"IOTLB miss req failed for IOVA 0x%" PRIx64,
iova);
- vhost_user_iotlb_pending_remove(dev, iova, 1, perm);
+ vhost_user_iotlb_pending_remove(dev, vq->asid, iova, 1, perm);
}
vhost_user_iotlb_rd_lock(vq);
@@ -113,7 +113,7 @@ __vhost_iova_to_vva(struct virtio_net *dev, struct vhost_virtqueue *vq,
tmp_size = *size;
/* Retry in case of VDUSE, as it is synchronous */
- vva = vhost_user_iotlb_cache_find(dev, iova, &tmp_size, perm);
+ vva = vhost_user_iotlb_cache_find(dev, vq->asid, iova, &tmp_size, perm);
if (tmp_size == *size)
return vva;
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index bb4708aed5ed..e0441adb6193 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -85,7 +85,7 @@ struct vhost_virtqueue;
typedef void (*vhost_iotlb_remove_notify)(uint64_t addr, uint64_t off, uint64_t size);
-typedef int (*vhost_iotlb_miss_cb)(struct virtio_net *dev, uint64_t iova, uint8_t perm);
+typedef int (*vhost_iotlb_miss_cb)(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm);
typedef int (*vhost_vring_inject_irq_cb)(struct virtio_net *dev, struct vhost_virtqueue *vq);
/**
@@ -326,6 +326,7 @@ struct __rte_cache_aligned vhost_virtqueue {
uint16_t batch_copy_nb_elems;
struct batch_copy_elem *batch_copy_elems;
int numa_node;
+ int asid;
bool used_wrap_counter;
bool avail_wrap_counter;
@@ -483,6 +484,8 @@ struct inflight_mem_info {
uint64_t size;
};
+#define IOTLB_MAX_ASID 2
+
/**
* Device structure contains all configuration information relating
* to the device.
@@ -504,13 +507,7 @@ struct __rte_cache_aligned virtio_net {
int linearbuf;
struct vhost_virtqueue *virtqueue[VHOST_MAX_VRING];
- rte_rwlock_t iotlb_pending_lock;
- struct vhost_iotlb_entry *iotlb_pool;
- TAILQ_HEAD(, vhost_iotlb_entry) iotlb_list;
- TAILQ_HEAD(, vhost_iotlb_entry) iotlb_pending_list;
- int iotlb_cache_nr;
- rte_spinlock_t iotlb_free_lock;
- SLIST_HEAD(, vhost_iotlb_entry) iotlb_free_list;
+ struct iotlb *iotlb[IOTLB_MAX_ASID];
struct inflight_mem_info *inflight_info;
#define IF_NAME_SZ (PATH_MAX > IFNAMSIZ ? PATH_MAX : IFNAMSIZ)
diff --git a/lib/vhost/vhost_user.c b/lib/vhost/vhost_user.c
index 020c993b2991..e623651fc01e 100644
--- a/lib/vhost/vhost_user.c
+++ b/lib/vhost/vhost_user.c
@@ -1530,7 +1530,8 @@ vhost_user_set_mem_table(struct virtio_net **pdev,
/* Flush IOTLB cache as previous HVAs are now invalid */
if (dev->features & (1ULL << VIRTIO_F_IOMMU_PLATFORM))
- vhost_user_iotlb_flush_all(dev);
+ for (i = 0; i < IOTLB_MAX_ASID; i++)
+ vhost_user_iotlb_flush_all(dev, i);
free_all_mem_regions(dev);
rte_free(dev->mem);
@@ -1830,7 +1831,7 @@ vhost_user_rem_mem_reg(struct virtio_net **pdev,
if (dev->async_copy && rte_vfio_is_enabled("vfio"))
async_dma_map_region(dev, current_region, false);
if (dev->features & (1ULL << VIRTIO_F_IOMMU_PLATFORM))
- vhost_user_iotlb_cache_remove(dev,
+ vhost_user_iotlb_cache_remove(dev, 0,
current_region->guest_phys_addr,
current_region->size);
remove_guest_pages(dev, current_region);
@@ -2581,7 +2582,7 @@ vhost_user_get_vring_base(struct virtio_net **pdev,
ctx->msg.size = sizeof(ctx->msg.payload.state);
ctx->fd_num = 0;
- vhost_user_iotlb_flush_all(dev);
+ vhost_user_iotlb_flush_all(dev, vq->asid);
rte_rwlock_write_lock(&vq->access_lock);
vring_invalidate(dev, vq);
@@ -3030,7 +3031,7 @@ vhost_user_iotlb_msg(struct virtio_net **pdev,
pg_sz = hua_to_alignment(dev->mem, (void *)(uintptr_t)vva);
- vhost_user_iotlb_cache_insert(dev, imsg->iova, vva, 0, len, pg_sz, imsg->perm);
+ vhost_user_iotlb_cache_insert(dev, 0, imsg->iova, vva, 0, len, pg_sz, imsg->perm);
for (i = 0; i < dev->nr_vring; i++) {
struct vhost_virtqueue *vq = dev->virtqueue[i];
@@ -3047,7 +3048,7 @@ vhost_user_iotlb_msg(struct virtio_net **pdev,
}
break;
case VHOST_IOTLB_INVALIDATE:
- vhost_user_iotlb_cache_remove(dev, imsg->iova, imsg->size);
+ vhost_user_iotlb_cache_remove(dev, 0, imsg->iova, imsg->size);
for (i = 0; i < dev->nr_vring; i++) {
struct vhost_virtqueue *vq = dev->virtqueue[i];
@@ -3640,7 +3641,7 @@ vhost_user_msg_handler(int vid, int fd)
}
static int
-vhost_user_iotlb_miss(struct virtio_net *dev, uint64_t iova, uint8_t perm)
+vhost_user_iotlb_miss(struct virtio_net *dev, int asid __rte_unused, uint64_t iova, uint8_t perm)
{
int ret;
struct vhu_msg_context ctx = {
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 04/11] vhost: add VDUSE API version negotiation
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (2 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 03/11] vhost: introduce ASID support Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 05/11] vhost: add virtqueues groups support to VDUSE Eugenio Pérez
` (6 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
From: Maxime Coquelin <maxime.coquelin@redhat.com>
As preliminary step to support new VDUSE API version
introducing ASID support, this patch adds API version
negotiation to keep compatibility with older kernels.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 14 ++++++++++++--
lib/vhost/vhost.h | 1 +
2 files changed, 13 insertions(+), 2 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 3c72bafcf511..ceef207ad2df 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -25,7 +25,7 @@
#include "vhost.h"
#include "virtio_net_ctrl.h"
-#define VHOST_VDUSE_API_VERSION 0
+#define VHOST_VDUSE_API_VERSION 0ULL
#define VDUSE_CTRL_PATH "/dev/vduse/control"
struct vduse {
@@ -680,7 +680,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
uint32_t i, max_queue_pairs, total_queues;
struct virtio_net *dev;
struct virtio_net_config vnet_config = {{ 0 }};
- uint64_t ver = VHOST_VDUSE_API_VERSION;
+ uint64_t ver;
uint64_t features;
const char *name = path + strlen("/dev/vduse/");
bool reconnect = false;
@@ -700,6 +700,15 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
return -1;
}
+ if (ioctl(control_fd, VDUSE_GET_API_VERSION, &ver)) {
+ VHOST_CONFIG_LOG(name, ERR, "Failed to get API version: %s", strerror(errno));
+ ret = -1;
+ goto out_ctrl_close;
+ }
+
+ ver = RTE_MIN(ver, VHOST_VDUSE_API_VERSION);
+ VHOST_CONFIG_LOG(name, INFO, "Using VDUSE API version %" PRIu64 "", ver);
+
if (ioctl(control_fd, VDUSE_SET_API_VERSION, &ver)) {
VHOST_CONFIG_LOG(name, ERR, "Failed to set API version: %" PRIu64 ": %s",
ver, strerror(errno));
@@ -800,6 +809,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
strlcpy(dev->ifname, path, sizeof(dev->ifname));
dev->vduse_ctrl_fd = control_fd;
dev->vduse_dev_fd = dev_fd;
+ dev->vduse_api_ver = ver;
ret = vduse_reconnect_log_map(dev, !reconnect);
if (ret < 0)
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index e0441adb6193..36d65bc2c162 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -532,6 +532,7 @@ struct __rte_cache_aligned virtio_net {
int postcopy_listening;
int vduse_ctrl_fd;
int vduse_dev_fd;
+ uint64_t vduse_api_ver;
struct vhost_virtqueue *cvq;
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 05/11] vhost: add virtqueues groups support to VDUSE
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (3 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 04/11] vhost: add VDUSE API version negotiation Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 06/11] vhost: add ASID support to VDUSE IOTLB operations Eugenio Pérez
` (5 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
From: Maxime Coquelin <maxime.coquelin@redhat.com>
VDUSE API version 1 introduces the notion of virtqueue
groups, which once supported, enables the support of
multiple addresses spaces.
For VDUSE networking devices, we need two groups, one for
the datapath queues, and one for the control queue.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 52 +++++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 50 insertions(+), 2 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index ceef207ad2df..8c031168bf77 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -38,12 +38,24 @@ static const char * const vduse_reqs_str[] = {
"VDUSE_GET_VQ_STATE",
"VDUSE_SET_STATUS",
"VDUSE_UPDATE_IOTLB",
+ "VDUSE_SET_VQ_GROUP_ASID",
};
#define vduse_req_id_to_str(id) \
(id < RTE_DIM(vduse_reqs_str) ? \
vduse_reqs_str[id] : "Unknown")
+static uint64_t vduse_vq_to_group(struct virtio_net *dev, struct vhost_virtqueue *vq)
+{
+ if (dev->vduse_api_ver < 1)
+ return 0;
+
+ if (vq == dev->cvq)
+ return 1;
+
+ return 0;
+}
+
static int
vduse_inject_irq(struct virtio_net *dev, struct vhost_virtqueue *vq)
{
@@ -271,6 +283,7 @@ vduse_vring_cleanup(struct virtio_net *dev, unsigned int index)
vq->size = 0;
vq->last_used_idx = 0;
vq->last_avail_idx = 0;
+ vq->asid = 0;
}
/*
@@ -410,6 +423,7 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
struct vduse_dev_response resp;
struct vhost_virtqueue *vq;
uint8_t old_status = dev->status;
+ uint32_t i;
int ret;
memset(&resp, 0, sizeof(resp));
@@ -450,6 +464,34 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
req.iova.last - req.iova.start + 1);
resp.result = VDUSE_REQ_RESULT_OK;
break;
+ case VDUSE_SET_VQ_GROUP_ASID:
+ if (dev->vduse_api_ver < 1) {
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ if (req.vq_group_asid.asid >= RTE_DIM(dev->iotlb)) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Invalid ASID %u for group %u",
+ req.vq_group_asid.asid, req.vq_group_asid.group);
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ VHOST_CONFIG_LOG(dev->ifname, INFO, "\tAssigning ASID %d to group %d",
+ req.vq_group_asid.asid, req.vq_group_asid.group);
+
+ for (i = 0; i < dev->nr_vring; i++) {
+ vq = dev->virtqueue[i];
+
+ if (vduse_vq_to_group(dev, vq) == req.vq_group_asid.group) {
+ vq->asid = req.vq_group_asid.asid;
+ VHOST_CONFIG_LOG(dev->ifname, INFO, "\t\tVQ %d gets ASID %d",
+ i, req.vq_group_asid.asid);
+ }
+ }
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
default:
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
@@ -760,6 +802,10 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
dev_config->features = features;
dev_config->vq_num = total_queues;
dev_config->vq_align = rte_mem_page_size();
+ if (ver >= 1) {
+ dev_config->ngroups = 2;
+ dev_config->nas = 2;
+ }
dev_config->config_size = sizeof(struct virtio_net_config);
memcpy(dev_config->config, &vnet_config, sizeof(vnet_config));
@@ -848,11 +894,15 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
vq = dev->virtqueue[i];
vq->reconnect_log = &dev->reconnect_log->vring[i];
+ if (i == max_queue_pairs * 2)
+ dev->cvq = vq;
+
if (reconnect)
continue;
vq_cfg.index = i;
vq_cfg.max_size = 1024;
+ vq_cfg.group = vduse_vq_to_group(dev, vq);
ret = ioctl(dev->vduse_dev_fd, VDUSE_VQ_SETUP, &vq_cfg);
if (ret) {
@@ -861,8 +911,6 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
}
}
- dev->cvq = dev->virtqueue[max_queue_pairs * 2];
-
ret = fdset_add(vduse.fdset, dev->vduse_dev_fd, vduse_events_handler, NULL, dev);
if (ret) {
VHOST_CONFIG_LOG(name, ERR, "Failed to add fd %d to vduse fdset",
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 06/11] vhost: add ASID support to VDUSE IOTLB operations
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (4 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 05/11] vhost: add virtqueues groups support to VDUSE Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 07/11] vhost: claim VDUSE support for API version 1 Eugenio Pérez
` (4 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Make use of the newly introduced address space ID when
calling Vhost IOTLB API.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 39 +++++++++++++++++++++++++++++++++------
1 file changed, 33 insertions(+), 6 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 8c031168bf77..2d3d8a536773 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -71,7 +71,7 @@ vduse_iotlb_remove_notify(uint64_t addr, uint64_t offset, uint64_t size)
static int
vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm __rte_unused)
{
- struct vduse_iotlb_entry entry;
+ struct vduse_iotlb_entry_v2 entry = {};
uint64_t size, page_size;
struct stat stat;
void *mmap_addr;
@@ -79,8 +79,13 @@ vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm _
entry.start = iova;
entry.last = iova + 1;
+ entry.asid = asid;
- ret = ioctl(dev->vduse_dev_fd, VDUSE_IOTLB_GET_FD, &entry);
+ if (dev->vduse_api_ver < 1) {
+ ret = ioctl(dev->vduse_dev_fd, VDUSE_IOTLB_GET_FD, &entry);
+ } else {
+ ret = ioctl(dev->vduse_dev_fd, VDUSE_IOTLB_GET_FD2, &entry);
+ }
if (ret < 0) {
VHOST_CONFIG_LOG(dev->ifname, ERR, "Failed to get IOTLB entry for 0x%" PRIx64,
iova);
@@ -90,6 +95,7 @@ vduse_iotlb_miss(struct virtio_net *dev, int asid, uint64_t iova, uint8_t perm _
fd = ret;
VHOST_CONFIG_LOG(dev->ifname, DEBUG, "New IOTLB entry:");
+ VHOST_CONFIG_LOG(dev->ifname, DEBUG, "\tASID: %d", entry.asid);
VHOST_CONFIG_LOG(dev->ifname, DEBUG, "\tIOVA: %" PRIx64 " - %" PRIx64,
(uint64_t)entry.start, (uint64_t)entry.last);
VHOST_CONFIG_LOG(dev->ifname, DEBUG, "\toffset: %" PRIx64, (uint64_t)entry.offset);
@@ -458,10 +464,31 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
resp.result = VDUSE_REQ_RESULT_OK;
break;
case VDUSE_UPDATE_IOTLB:
- VHOST_CONFIG_LOG(dev->ifname, INFO, "\tIOVA range: %" PRIx64 " - %" PRIx64,
- (uint64_t)req.iova.start, (uint64_t)req.iova.last);
- vhost_user_iotlb_cache_remove(dev, 0, req.iova.start,
- req.iova.last - req.iova.start + 1);
+ {
+ uint64_t start, last;
+ uint32_t asid;
+
+ if (dev->vduse_api_ver < 1) {
+ start = req.iova.start;
+ last = req.iova.last;
+ asid = 0;
+ } else {
+ start = req.iova_v2.start;
+ last = req.iova_v2.last;
+ asid = req.iova_v2.asid;
+ }
+
+ if (asid >= RTE_DIM(dev->iotlb)) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Invalid ASID %u in IOTLB update", asid);
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ VHOST_CONFIG_LOG(dev->ifname, INFO, "\t(ASID %d) IOVA range: %" PRIx64 " - %" PRIx64,
+ asid, start, last);
+ vhost_user_iotlb_cache_remove(dev, asid, start, last - start + 1);
+ }
resp.result = VDUSE_REQ_RESULT_OK;
break;
case VDUSE_SET_VQ_GROUP_ASID:
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 07/11] vhost: claim VDUSE support for API version 1
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (5 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 06/11] vhost: add ASID support to VDUSE IOTLB operations Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 08/11] vhost: add net status feature to VDUSE Eugenio Pérez
` (3 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 2d3d8a536773..4950a5da01dd 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -25,7 +25,7 @@
#include "vhost.h"
#include "virtio_net_ctrl.h"
-#define VHOST_VDUSE_API_VERSION 0ULL
+#define VHOST_VDUSE_API_VERSION 1ULL
#define VDUSE_CTRL_PATH "/dev/vduse/control"
struct vduse {
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 08/11] vhost: add net status feature to VDUSE
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (6 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 07/11] vhost: claim VDUSE support for API version 1 Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 09/11] vhost: Support VDUSE QUEUE_READY feature Eugenio Pérez
` (2 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
From: Maxime Coquelin <maxime.coquelin@redhat.com>
Enable the VIRTIO_NET_F_STATUS feature for VDUSE devices.
This allows the device to report link status (e.g.,
VIRTIO_NET_S_LINK_UP). It also allows the device to signal the driver
that it needs to send gratuitous ARP with VIRTIO_NET_S_ANNOUNCE.
Signed-off-by: Maxime Coquelin <maxime.coquelin@redhat.com>
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 1 +
lib/vhost/vduse.h | 3 ++-
2 files changed, 3 insertions(+), 1 deletion(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index 4950a5da01dd..b51776cf3ff4 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -820,6 +820,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
goto out_ctrl_close;
}
+ vnet_config.status = VIRTIO_NET_S_LINK_UP;
vnet_config.max_virtqueue_pairs = max_queue_pairs;
memset(dev_config, 0, sizeof(struct vduse_dev_config));
diff --git a/lib/vhost/vduse.h b/lib/vhost/vduse.h
index b2515bb9df76..d697f85be5cc 100644
--- a/lib/vhost/vduse.h
+++ b/lib/vhost/vduse.h
@@ -7,7 +7,8 @@
#include "vhost.h"
-#define VDUSE_NET_SUPPORTED_FEATURES VIRTIO_NET_SUPPORTED_FEATURES
+#define VDUSE_NET_SUPPORTED_FEATURES (VIRTIO_NET_SUPPORTED_FEATURES | \
+ (1ULL << VIRTIO_NET_F_STATUS))
int vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool linearbuf);
int vduse_device_destroy(const char *path);
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 09/11] vhost: Support VDUSE QUEUE_READY feature
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (7 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 08/11] vhost: add net status feature to VDUSE Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 10/11] vhost: Support vduse suspend feature Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 11/11] doc: add release notes for VDUSE live migration support Eugenio Pérez
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
Add support for the VDUSE_F_QUEUE_READY feature.
In VDUSE, the dataplane is enabled only after control virtqueue so the
device is fully configured in the destination of a live migration before
the dataplane starts. This message signals the VDUSE device when the
dataplane queues should be enabled.
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 86 +++++++++++++++++++++++++++++++++++++++++++++--
lib/vhost/vhost.h | 1 +
2 files changed, 85 insertions(+), 2 deletions(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index b51776cf3ff4..bcbd9fcd6d52 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -39,12 +39,15 @@ static const char * const vduse_reqs_str[] = {
"VDUSE_SET_STATUS",
"VDUSE_UPDATE_IOTLB",
"VDUSE_SET_VQ_GROUP_ASID",
+ "VDUSE_SET_VQ_READY",
};
#define vduse_req_id_to_str(id) \
(id < RTE_DIM(vduse_reqs_str) ? \
vduse_reqs_str[id] : "Unknown")
+static const uint64_t supported_vduse_features = RTE_BIT64(VDUSE_F_QUEUE_READY);
+
static uint64_t vduse_vq_to_group(struct virtio_net *dev, struct vhost_virtqueue *vq)
{
if (dev->vduse_api_ver < 1)
@@ -519,6 +522,51 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
}
resp.result = VDUSE_REQ_RESULT_OK;
break;
+ case VDUSE_SET_VQ_READY:
+ if (!(dev->status & VIRTIO_DEVICE_STATUS_DRIVER_OK)) {
+ /*
+ * dev->notify_ops is NULL if !S_DRIVER_OK,
+ * vduse_device_start will check the queue readiness.
+ */
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
+ }
+ if (!(dev->vduse_features & RTE_BIT64(VDUSE_F_QUEUE_READY))) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Unexpected ready message with no ready feature acked");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ i = req.vq_ready.num;
+ if (i >= dev->nr_vring) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Invalid virtqueue index: %u", i);
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ vq = dev->virtqueue[i];
+ if (dev->notify_ops == NULL || dev->notify_ops->vring_state_changed == NULL) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "No ops->vring_state_changed");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ ret = dev->notify_ops->vring_state_changed(dev->vid, i,
+ req.vq_ready.ready);
+ VHOST_CONFIG_LOG(dev->ifname, INFO,
+ "\t\t VQ %d gets ready %d ok %d", i,
+ req.vq_ready.ready, ret);
+ if (ret != 0) {
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+
+ vq->enabled = req.vq_ready.ready;
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
default:
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
@@ -536,7 +584,8 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
if ((old_status ^ dev->status) & VIRTIO_DEVICE_STATUS_DRIVER_OK) {
if (dev->status & VIRTIO_DEVICE_STATUS_DRIVER_OK) {
/* Poll virtqueues ready states before starting device */
- ret = vduse_wait_for_virtqueues_ready(dev);
+ ret = dev->vduse_features & RTE_BIT64(VDUSE_F_QUEUE_READY) ? 0
+ : vduse_wait_for_virtqueues_ready(dev);
if (ret < 0) {
VHOST_CONFIG_LOG(dev->ifname, ERR,
"Failed to wait for virtqueues ready, aborting device start");
@@ -742,6 +791,26 @@ vduse_reconnect_start_device(struct virtio_net *dev)
return ret;
}
+/* If some error occurs just continue as if the kernel exposed no features */
+static uint64_t
+vduse_device_get_vduse_features(int control_fd, const char *log_name)
+{
+ uint64_t vduse_kernel_features;
+ int ret;
+
+ ret = ioctl(control_fd, VDUSE_GET_FEATURES, &vduse_kernel_features);
+ if (ret < 0) {
+ VHOST_CONFIG_LOG(log_name, INFO,
+ "Failed to get kernel VDUSE features, assuming not supported: %d(%s)",
+ errno, strerror(errno));
+ return 0;
+ }
+
+ VHOST_CONFIG_LOG(log_name, DEBUG, "Setting vhost kernel features: %"PRIx64,
+ vduse_kernel_features & supported_vduse_features);
+ return vduse_kernel_features & supported_vduse_features;
+}
+
int
vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool linearbuf)
{
@@ -750,7 +819,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
struct virtio_net *dev;
struct virtio_net_config vnet_config = {{ 0 }};
uint64_t ver;
- uint64_t features;
+ uint64_t features, vduse_features = 0;
const char *name = path + strlen("/dev/vduse/");
bool reconnect = false;
@@ -834,6 +903,18 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
dev_config->ngroups = 2;
dev_config->nas = 2;
}
+
+ vduse_features = vduse_device_get_vduse_features(control_fd, name);
+ if (vduse_features) {
+ ret = ioctl(control_fd, VDUSE_SET_FEATURES, &vduse_features);
+ if (ret < 0) {
+ VHOST_CONFIG_LOG(name, ERR, "Failed to set VDUSE features: %s",
+ strerror(errno));
+ free(dev_config);
+ goto out_ctrl_close;
+ }
+ }
+
dev_config->config_size = sizeof(struct virtio_net_config);
memcpy(dev_config->config, &vnet_config, sizeof(vnet_config));
@@ -884,6 +965,7 @@ vduse_device_create(const char *path, bool compliant_ol_flags, bool extbuf, bool
dev->vduse_ctrl_fd = control_fd;
dev->vduse_dev_fd = dev_fd;
dev->vduse_api_ver = ver;
+ dev->vduse_features = vduse_features;
ret = vduse_reconnect_log_map(dev, !reconnect);
if (ret < 0)
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index 36d65bc2c162..369ccf399146 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -533,6 +533,7 @@ struct __rte_cache_aligned virtio_net {
int vduse_ctrl_fd;
int vduse_dev_fd;
uint64_t vduse_api_ver;
+ uint64_t vduse_features;
struct vhost_virtqueue *cvq;
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 10/11] vhost: Support vduse suspend feature
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (8 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 09/11] vhost: Support VDUSE QUEUE_READY feature Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 11/11] doc: add release notes for VDUSE live migration support Eugenio Pérez
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
Add support for the VDUSE_F_SUSPEND feature.
The suspend feature allows the driver to stop the device from processing
the virtqueues. This ensures that the virtqueue state can be fetched
reliably in a live migration.
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
lib/vhost/vduse.c | 42 +++++++++++++++++++++++++++++++++++++++++-
lib/vhost/vhost.h | 1 +
2 files changed, 42 insertions(+), 1 deletion(-)
diff --git a/lib/vhost/vduse.c b/lib/vhost/vduse.c
index bcbd9fcd6d52..cdff4df84a84 100644
--- a/lib/vhost/vduse.c
+++ b/lib/vhost/vduse.c
@@ -46,7 +46,8 @@ static const char * const vduse_reqs_str[] = {
(id < RTE_DIM(vduse_reqs_str) ? \
vduse_reqs_str[id] : "Unknown")
-static const uint64_t supported_vduse_features = RTE_BIT64(VDUSE_F_QUEUE_READY);
+static const uint64_t supported_vduse_features =
+ RTE_BIT64(VDUSE_F_QUEUE_READY) | RTE_BIT64(VDUSE_F_SUSPEND);
static uint64_t vduse_vq_to_group(struct virtio_net *dev, struct vhost_virtqueue *vq)
{
@@ -416,6 +417,7 @@ vduse_device_stop(struct virtio_net *dev)
vhost_destroy_device_notify(dev);
dev->flags &= ~VIRTIO_DEV_READY;
+ dev->vduse_suspended = false;
for (i = 0; i < dev->nr_vring; i++)
vduse_vring_cleanup(dev, i);
@@ -537,6 +539,12 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
}
+ if (dev->vduse_suspended) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "SET_VQ_READY received on suspended device");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
i = req.vq_ready.num;
if (i >= dev->nr_vring) {
@@ -567,11 +575,43 @@ vduse_events_handler(int fd, void *arg, int *close __rte_unused)
vq->enabled = req.vq_ready.ready;
resp.result = VDUSE_REQ_RESULT_OK;
break;
+ case VDUSE_SUSPEND:
+ if (!(dev->vduse_features & RTE_BIT64(VDUSE_F_SUSPEND))) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Unnegotiated suspend message");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+ if (!(dev->status & VIRTIO_DEVICE_STATUS_DRIVER_OK)) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Unexpected suspend message with no DRIVER_OK");
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ break;
+ }
+ for (i = 0; dev->notify_ops != NULL &&
+ dev->notify_ops->vring_state_changed != NULL &&
+ i < dev->nr_vring; i++) {
+ if (dev->virtqueue[i] == dev->cvq)
+ continue;
+
+ ret = dev->notify_ops->vring_state_changed(dev->vid, i, false);
+ if (ret) {
+ VHOST_CONFIG_LOG(dev->ifname, ERR,
+ "Failed to disable VQ %u for suspend", i);
+ resp.result = VDUSE_REQ_RESULT_FAILED;
+ goto out;
+ }
+ }
+ dev->vduse_suspended = true;
+ resp.result = VDUSE_REQ_RESULT_OK;
+ break;
+
default:
resp.result = VDUSE_REQ_RESULT_FAILED;
break;
}
+out:
resp.request_id = req.request_id;
ret = write(dev->vduse_dev_fd, &resp, sizeof(resp));
diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h
index 369ccf399146..d77e2003d4bd 100644
--- a/lib/vhost/vhost.h
+++ b/lib/vhost/vhost.h
@@ -534,6 +534,7 @@ struct __rte_cache_aligned virtio_net {
int vduse_dev_fd;
uint64_t vduse_api_ver;
uint64_t vduse_features;
+ bool vduse_suspended;
struct vhost_virtqueue *cvq;
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 11/11] doc: add release notes for VDUSE live migration support
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
` (9 preceding siblings ...)
2026-09-28 13:29 ` [PATCH v2 10/11] vhost: Support vduse suspend feature Eugenio Pérez
@ 2026-09-28 13:29 ` Eugenio Pérez
10 siblings, 0 replies; 14+ messages in thread
From: Eugenio Pérez @ 2026-09-28 13:29 UTC (permalink / raw)
To: Maxime Coquelin; +Cc: dev, mst, Yongji Xie, jasowangio, chenbox, david.marchand
Document the new VDUSE API version 1 features added to support
live migration:
- Address Space ID (ASID) support
- Virtqueue groups
- Queue ready feature (VDUSE_F_QUEUE_READY)
- Suspend feature (VDUSE_F_SUSPEND)
- Link status support (VIRTIO_NET_F_STATUS)
Signed-off-by: Eugenio Pérez <eperezma@redhat.com>
---
doc/guides/rel_notes/release_26_07.rst | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 3d18ba2dd04e..4943b84f8c9b 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -104,6 +104,14 @@ New Features
Added network driver for the LinkData network adapters.
+* **Added VDUSE API version 1 support in vhost library.**
+
+ Updated VDUSE (vDPA Device in Userspace) support with API version 1 features
+ to enable live migration:
+
+ * Added Address Space ID (ASID) support for multiple independent address spaces
+ per device, enabling better isolation and support for virtqueue groups.
+
* **Updated Microsoft mana driver.**
Added device reset support to the MANA PMD,
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers
2026-09-28 13:29 ` [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers Eugenio Pérez
@ 2026-09-28 23:42 ` Stephen Hemminger
2026-09-29 7:26 ` Maxime Coquelin
0 siblings, 1 reply; 14+ messages in thread
From: Stephen Hemminger @ 2026-09-28 23:42 UTC (permalink / raw)
To: Eugenio Pérez
Cc: Maxime Coquelin, dev, mst, Yongji Xie, jasowangio, chenbox,
david.marchand
On Mon, 28 Sep 2026 15:29:03 +0200
Eugenio Pérez <eperezma@redhat.com> wrote:
> __counted_by is a GCC 14 / Clang 18 extension used in imported kernel
> uapi headers. Older compilers reject it with an error. Define it as an
> empty function-like macro when the compiler does not support it.
>
> Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
> ---
NAK
Missing Signed-Of-By and the right way to handle counted_by is to
introduce and a new macro in rte_common.h (aka __rte_counted_by).
There are lots of places in DPDK that could use it.
If the problem is in vhost, then get it from stddef.h
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers
2026-09-28 23:42 ` Stephen Hemminger
@ 2026-09-29 7:26 ` Maxime Coquelin
0 siblings, 0 replies; 14+ messages in thread
From: Maxime Coquelin @ 2026-09-29 7:26 UTC (permalink / raw)
To: Stephen Hemminger
Cc: Eugenio Pérez, dev, mst, Yongji Xie, jasowangio, chenbox,
david.marchand
[-- Attachment #1: Type: text/plain, Size: 874 bytes --]
On Tue, Sep 29, 2026 at 1:42 AM Stephen Hemminger <
stephen@networkplumber.org> wrote:
> On Mon, 28 Sep 2026 15:29:03 +0200
> Eugenio Pérez <eperezma@redhat.com> wrote:
>
> > __counted_by is a GCC 14 / Clang 18 extension used in imported kernel
> > uapi headers. Older compilers reject it with an error. Define it as an
> > empty function-like macro when the compiler does not support it.
> >
> > Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
> > ---
>
> NAK
>
> Missing Signed-Of-By and the right way to handle counted_by is to
> introduce and a new macro in rte_common.h (aka __rte_counted_by).
> There are lots of places in DPDK that could use it.
>
> If the problem is in vhost, then get it from stddef.h
>
>
To me the solution is to import stddef.h, I have not tried but it is quite
small and self-contained.
Regards,
Maxime
[-- Attachment #2: Type: text/html, Size: 1444 bytes --]
^ permalink raw reply [flat|nested] 14+ messages in thread
end of thread, other threads:[~2026-09-29 7:26 UTC | newest]
Thread overview: 14+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 13:29 [PATCH v2 00/11] Add vduse live migration features Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 01/11] build: define __counted_by as empty on unsupported compilers Eugenio Pérez
2026-09-28 23:42 ` Stephen Hemminger
2026-09-29 7:26 ` Maxime Coquelin
2026-09-28 13:29 ` [PATCH v2 02/11] uapi: import VDUSE and VFIO header from v7.3-rc3 kernel Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 03/11] vhost: introduce ASID support Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 04/11] vhost: add VDUSE API version negotiation Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 05/11] vhost: add virtqueues groups support to VDUSE Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 06/11] vhost: add ASID support to VDUSE IOTLB operations Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 07/11] vhost: claim VDUSE support for API version 1 Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 08/11] vhost: add net status feature to VDUSE Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 09/11] vhost: Support VDUSE QUEUE_READY feature Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 10/11] vhost: Support vduse suspend feature Eugenio Pérez
2026-09-28 13:29 ` [PATCH v2 11/11] doc: add release notes for VDUSE live migration support Eugenio Pérez
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox