Kernel KVM virtualization development
 help / color / mirror / Atom feed
* [PATCH kvmtool 0/7] Fix --vcpu-affinity
@ 2026-09-17 15:49 Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 1/7] arm64: Pass the number of elements as the first argument to calloc() Alexandru Elisei
                   ` (7 more replies)
  0 siblings, 8 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

This is my attempt at fixing --vcpu-affinity, which I introduced. Patches #1-#4
are straightforward fixes, and they don't modify the current (broken) behaviour
in any way. Patch #5 fixes is most of the fix: with it, the VCPU threads run on
the physical CPUs passed to --vcpu-affinity, and most of the kvmtool threads,
except the virtio threads, use the main thread affinity.

The virtio threads are special because they are spawned from the VCPU threads,
and they inherit the VCPU thread affinity. Note that there might be other cases
where this happens.

Patches #6-#7 attempt to fix the remaining issue: #6 switch the entire codebase
to kvm_create_thread() instead of pthread_create(). In #7, kvm_create_thread()
will use the main thread affinity when creating new threads. These two patches
are RFC because they cause a lot of churn and I would like some feedback if it's
worth it.

Why is it useful to run the VCPUs on different physical CPUs than the other
threads? I can think of two situations:

1. Testing heterogenous hardware configurations, where the VM creation and setup
happens on a set of physical CPUs, and the VCPUs are run on a different set. I
used these patches for the KVM SPE series.

2. I/O testing. Maybe?

Tested on an Orion board and a x86 machine. On the x86 machine, there's was this
one thread called 'kvm-nx-lpage-re' which had affinity of the VCPUs, but I
couldn't figure out where it was created. I grep'ed for each part of the name in
the source node, no luck. I used strace, and couldn't see a clone call that
returned the corresponding pid. Any clues would be appreciated.

Alexandru Elisei (7):
  arm64: Pass the number of elements as the first argument to calloc()
  arm64: Consistently treat vcpu_affinity_cpuset as dynamically
    allocated
  arm64: Free the temporary cpumask in vcpu_affinity_parser()
  arm64/pmu: Consider kvmtool's affinity when searching for a PMU
  Correctly apply --vcpu-affinity to the VCPU threads
  Introduce kvm_create_thread()
  Don't apply --vcpu-affinity to threads spawned from VCPUs

 arm64/include/kvm/kvm-arch.h        |  2 --
 arm64/include/kvm/kvm-config-arch.h |  5 ----
 arm64/kvm-cpu.c                     |  9 ------
 arm64/kvm.c                         | 31 --------------------
 arm64/pmu.c                         | 31 +++++++++++---------
 builtin-run.c                       | 41 +++++++++++++++++++++++++--
 disk/aio.c                          |  3 +-
 disk/blk.c                          | 12 ++++++--
 disk/core.c                         | 29 ++++++++++---------
 disk/qcow.c                         | 44 ++++++++++++++++++++---------
 disk/raw.c                          | 21 +++++++++++---
 epoll.c                             |  2 +-
 include/kvm/disk-image.h            | 28 ++++++++++++------
 include/kvm/kvm-config.h            |  1 +
 include/kvm/kvm.h                   |  2 ++
 include/kvm/qcow.h                  |  3 +-
 include/kvm/uip.h                   |  2 ++
 include/kvm/util.h                  |  6 ++++
 kvm.c                               | 18 ++++++++++++
 net/uip/tcp.c                       |  7 ++---
 net/uip/udp.c                       |  5 ++--
 term.c                              |  4 +--
 ui/gtk3.c                           |  2 +-
 ui/sdl.c                            |  2 +-
 ui/vnc.c                            |  2 +-
 util/threadpool.c                   |  9 +++---
 util/util.c                         | 42 +++++++++++++++++++++++++++
 virtio/blk.c                        |  2 +-
 virtio/net.c                        | 26 +++++++++++------
 29 files changed, 258 insertions(+), 133 deletions(-)


base-commit: f67bc0bdae9433a9cfd05e65ea2c1bb6102566d9
-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH kvmtool 1/7] arm64: Pass the number of elements as the first argument to calloc()
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 2/7] arm64: Consistently treat vcpu_affinity_cpuset as dynamically allocated Alexandru Elisei
                   ` (6 subsequent siblings)
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

man 3 calloc is pretty clear that the first argument to calloc() is the
number of elements to be allocated, and the second argument is the size
of one element. Fix the instances where the order was inversed.

Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 arm64/kvm.c | 2 +-
 arm64/pmu.c | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/arm64/kvm.c b/arm64/kvm.c
index 36b3284e4a92..cdd673542aea 100644
--- a/arm64/kvm.c
+++ b/arm64/kvm.c
@@ -454,7 +454,7 @@ int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
 
 	kvm->cfg.arch.vcpu_affinity = cpulist;
 
-	cpumask = calloc(1, cpumask_size());
+	cpumask = calloc(cpumask_size(), 1);
 	if (!cpumask)
 		die_perror("calloc");
 
diff --git a/arm64/pmu.c b/arm64/pmu.c
index ef7faca1b71d..49d58ddd529b 100644
--- a/arm64/pmu.c
+++ b/arm64/pmu.c
@@ -192,7 +192,7 @@ static int find_pmu(struct kvm *kvm)
 	cpumask_t *cpumask;
 	int i, this_cpu;
 
-	cpumask = calloc(1, cpumask_size());
+	cpumask = calloc(cpumask_size(), 1);
 	if (!cpumask)
 		die_perror("calloc");
 
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH kvmtool 2/7] arm64: Consistently treat vcpu_affinity_cpuset as dynamically allocated
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 1/7] arm64: Pass the number of elements as the first argument to calloc() Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 3/7] arm64: Free the temporary cpumask in vcpu_affinity_parser() Alexandru Elisei
                   ` (5 subsequent siblings)
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

man 3 CPU_ALLOC says that for dynamically allocated cpusets, one should use
the *_S macros to manipulate the cpuset, and the example program also
demonstrates this. Use CPU_SET_S and CPU_ISSET_S to comply.

Fixes: 4639b72f61a3 ("arm64: Add --vcpu-affinity command line argument")
Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 arm64/kvm.c | 6 ++++--
 arm64/pmu.c | 6 ++++--
 2 files changed, 8 insertions(+), 4 deletions(-)

diff --git a/arm64/kvm.c b/arm64/kvm.c
index cdd673542aea..5f3ed9f6d201 100644
--- a/arm64/kvm.c
+++ b/arm64/kvm.c
@@ -450,6 +450,7 @@ int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
 	struct kvm *kvm = opt->ptr;
 	const char *cpulist = arg;
 	cpumask_t *cpumask;
+	size_t setsize;
 	int cpu, ret;
 
 	kvm->cfg.arch.vcpu_affinity = cpulist;
@@ -467,10 +468,11 @@ int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
 	kvm->arch.vcpu_affinity_cpuset = CPU_ALLOC(NR_CPUS);
 	if (!kvm->arch.vcpu_affinity_cpuset)
 		die_perror("CPU_ALLOC");
-	CPU_ZERO_S(CPU_ALLOC_SIZE(NR_CPUS), kvm->arch.vcpu_affinity_cpuset);
 
+	setsize = CPU_ALLOC_SIZE(NR_CPUS);
+	CPU_ZERO_S(setsize, kvm->arch.vcpu_affinity_cpuset);
 	for_each_cpu(cpu, cpumask)
-		CPU_SET(cpu, kvm->arch.vcpu_affinity_cpuset);
+		CPU_SET_S(cpu, setsize, kvm->arch.vcpu_affinity_cpuset);
 
 	return 0;
 }
diff --git a/arm64/pmu.c b/arm64/pmu.c
index 49d58ddd529b..e5a07abcd317 100644
--- a/arm64/pmu.c
+++ b/arm64/pmu.c
@@ -191,6 +191,7 @@ static int find_pmu(struct kvm *kvm)
 {
 	cpumask_t *cpumask;
 	int i, this_cpu;
+	size_t setsize;
 
 	cpumask = calloc(cpumask_size(), 1);
 	if (!cpumask)
@@ -202,8 +203,9 @@ static int find_pmu(struct kvm *kvm)
 			return -errno;
 		cpumask_set_cpu(this_cpu, cpumask);
 	} else {
-		for (i = 0; i < CPU_SETSIZE; i ++) {
-			if (CPU_ISSET(i, kvm->arch.vcpu_affinity_cpuset))
+		setsize = CPU_ALLOC_SIZE(NR_CPUS);
+		for (i = 0; i < NR_CPUS; i ++) {
+			if (CPU_ISSET_S(i, setsize, kvm->arch.vcpu_affinity_cpuset))
 				cpumask_set_cpu(i, cpumask);
 		}
 	}
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH kvmtool 3/7] arm64: Free the temporary cpumask in vcpu_affinity_parser()
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 1/7] arm64: Pass the number of elements as the first argument to calloc() Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 2/7] arm64: Consistently treat vcpu_affinity_cpuset as dynamically allocated Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 4/7] arm64/pmu: Consider kvmtool's affinity when searching for a PMU Alexandru Elisei
                   ` (4 subsequent siblings)
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

Free the local variable 'cpumask' when the function is successful.

Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 arm64/kvm.c | 13 +++++++------
 1 file changed, 7 insertions(+), 6 deletions(-)

diff --git a/arm64/kvm.c b/arm64/kvm.c
index 5f3ed9f6d201..88b3622766bf 100644
--- a/arm64/kvm.c
+++ b/arm64/kvm.c
@@ -451,7 +451,8 @@ int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
 	const char *cpulist = arg;
 	cpumask_t *cpumask;
 	size_t setsize;
-	int cpu, ret;
+	int ret = 0;
+	int cpu;
 
 	kvm->cfg.arch.vcpu_affinity = cpulist;
 
@@ -460,10 +461,8 @@ int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
 		die_perror("calloc");
 
 	ret = cpulist_parse(cpulist, cpumask);
-	if (ret) {
-		free(cpumask);
-		return ret;
-	}
+	if (ret)
+		goto out_free;
 
 	kvm->arch.vcpu_affinity_cpuset = CPU_ALLOC(NR_CPUS);
 	if (!kvm->arch.vcpu_affinity_cpuset)
@@ -474,7 +473,9 @@ int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
 	for_each_cpu(cpu, cpumask)
 		CPU_SET_S(cpu, setsize, kvm->arch.vcpu_affinity_cpuset);
 
-	return 0;
+out_free:
+	free(cpumask);
+	return ret;
 }
 
 void kvm__arch_validate_cfg(struct kvm *kvm)
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH kvmtool 4/7] arm64/pmu: Consider kvmtool's affinity when searching for a PMU
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
                   ` (2 preceding siblings ...)
  2026-09-17 15:49 ` [PATCH kvmtool 3/7] arm64: Free the temporary cpumask in vcpu_affinity_parser() Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-17 15:49 ` [PATCH kvmtool 5/7] Correctly apply --vcpu-affinity to the VCPU threads Alexandru Elisei
                   ` (3 subsequent siblings)
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

When --vcpu-affinity is set, find_pmu() -> find_pmu_cpumask() attempts to
find a PMU instance that is shared by all CPUs in the specified affinity.

If --vcpu-affinity is not set, find_pmu() will only attempt to find a PMU
instance associated with the *current* physical CPU. This misses the fact
that it might be possible that the other threads that kvmtool creates
(which include the VCPU threads), might be executed on physical CPUs which
have a different PMU instance than the *current* physical CPU (or none at
all). This can lead to hard to reproduce and diagnose errors.

Improve things by teaching find_pmu() to use kvmtool process CPU affinity
when --vcpu-affinity is not set.

Also fix a memory leak by freeing the local variable 'cpumask'.

Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 arm64/pmu.c | 38 +++++++++++++++++++++++++++-----------
 1 file changed, 27 insertions(+), 11 deletions(-)

diff --git a/arm64/pmu.c b/arm64/pmu.c
index e5a07abcd317..d3761689c907 100644
--- a/arm64/pmu.c
+++ b/arm64/pmu.c
@@ -189,28 +189,44 @@ out_free:
  */
 static int find_pmu(struct kvm *kvm)
 {
+	cpu_set_t *affinity;
 	cpumask_t *cpumask;
-	int i, this_cpu;
 	size_t setsize;
+	int i, ret = 0;
 
 	cpumask = calloc(cpumask_size(), 1);
 	if (!cpumask)
 		die_perror("calloc");
 
-	if (!kvm->arch.vcpu_affinity_cpuset) {
-		this_cpu = sched_getcpu();
-		if (this_cpu < 0)
-			return -errno;
-		cpumask_set_cpu(this_cpu, cpumask);
+	setsize = CPU_ALLOC_SIZE(NR_CPUS);
+
+	if (kvm->arch.vcpu_affinity_cpuset) {
+		affinity = kvm->arch.vcpu_affinity_cpuset;
 	} else {
-		setsize = CPU_ALLOC_SIZE(NR_CPUS);
-		for (i = 0; i < NR_CPUS; i ++) {
-			if (CPU_ISSET_S(i, setsize, kvm->arch.vcpu_affinity_cpuset))
-				cpumask_set_cpu(i, cpumask);
+		affinity = CPU_ALLOC(NR_CPUS);
+		if (!affinity)
+			die_perror("CPU_ALLOC");
+		CPU_ZERO_S(setsize, affinity);
+
+		ret = sched_getaffinity(0, setsize, affinity);
+		if (ret < 0) {
+			ret = -errno;
+			goto out_free;
 		}
 	}
 
-	return find_pmu_cpumask(kvm, cpumask);
+	for (i = 0; i < NR_CPUS; i ++) {
+		if (CPU_ISSET_S(i, setsize, affinity))
+			cpumask_set_cpu(i, cpumask);
+	}
+
+	ret = find_pmu_cpumask(kvm, cpumask);
+
+out_free:
+	free(cpumask);
+	if (!kvm->arch.vcpu_affinity_cpuset)
+		CPU_FREE(affinity);
+	return ret;
 }
 
 void pmu__generate_fdt_nodes(void *fdt, struct kvm *kvm)
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH kvmtool 5/7] Correctly apply --vcpu-affinity to the VCPU threads
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
                   ` (3 preceding siblings ...)
  2026-09-17 15:49 ` [PATCH kvmtool 4/7] arm64/pmu: Consider kvmtool's affinity when searching for a PMU Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-17 15:49 ` [RFC PATCH kvmtool 6/7] Introduce kvm_create_thread() Alexandru Elisei
                   ` (2 subsequent siblings)
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

Commit 639b72f61a3 ("arm64: Add --vcpu-affinity command line argument")
added the --vcpu-affinity command line argument to run the VCPU threads on
the specified physical CPUs. The goal was to allow the rest of the threads
created by kvmtool to be scheduled freely by the operating system.

The affinity is set from kvm_cpu__init() -> kvm_cpu__reset_vcpu(), which is
called from the main thread, before the VCPU threads are created.  This
means that the main thread is pinned to the CPU list specified with
--vcpu-affinity, as well as all the other threads created after this point
(which includes the VCPUs), which defeats the purpose of the
--vcpu-affinity.

Make this right by setting the allowed CPUs only for the VCPU threads, when
they are created. Make --vcpu-affinity an arch-independent option, since
it's not something tied to a particular architecture, and this makes the
fix cleaner.

Also do proper cleanup by freeing the affinity cpuset when struct kvm is
freed.

Fixes: 639b72f61a3 ("arm64: Add --vcpu-affinity command line argument")
Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 arm64/include/kvm/kvm-arch.h        |  2 --
 arm64/include/kvm/kvm-config-arch.h |  5 ----
 arm64/kvm-cpu.c                     |  9 -------
 arm64/kvm.c                         | 34 ------------------------
 arm64/pmu.c                         |  6 ++---
 builtin-run.c                       | 41 +++++++++++++++++++++++++++--
 include/kvm/kvm-config.h            |  1 +
 include/kvm/kvm.h                   |  1 +
 include/kvm/util.h                  |  4 +++
 kvm.c                               |  3 +++
 util/util.c                         | 29 ++++++++++++++++++++
 11 files changed, 80 insertions(+), 55 deletions(-)

diff --git a/arm64/include/kvm/kvm-arch.h b/arm64/include/kvm/kvm-arch.h
index e7dd52692935..f0223132e60c 100644
--- a/arm64/include/kvm/kvm-arch.h
+++ b/arm64/include/kvm/kvm-arch.h
@@ -112,8 +112,6 @@ struct kvm_arch {
 	u64	initrd_guest_start;
 	u64	initrd_size;
 	u64	dtb_guest_start;
-
-	cpu_set_t *vcpu_affinity_cpuset;
 };
 
 struct kvm_cpu *kvm__arch_mpidr_to_vcpu(struct kvm *kvm, u64 target_mpidr);
diff --git a/arm64/include/kvm/kvm-config-arch.h b/arm64/include/kvm/kvm-config-arch.h
index d8a8ef7fd490..adda6403401e 100644
--- a/arm64/include/kvm/kvm-config-arch.h
+++ b/arm64/include/kvm/kvm-config-arch.h
@@ -5,7 +5,6 @@
 
 struct kvm_config_arch {
 	const char	*dump_dtb_filename;
-	const char	*vcpu_affinity;
 	unsigned int	force_cntfrq;
 	bool		aarch32_guest;
 	bool		has_pmuv3;
@@ -24,7 +23,6 @@ struct kvm_config_arch {
 };
 
 int irqchip_parser(const struct option *opt, const char *arg, int unset);
-int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset);
 int sve_vl_parser(const struct option *opt, const char *arg, int unset);
 
 #define OPT_ARCH_RUN(pfx, cfg)							\
@@ -37,9 +35,6 @@ int sve_vl_parser(const struct option *opt, const char *arg, int unset);
 			" main thread, unless --vcpu-affinity is set"),		\
 	OPT_BOOLEAN('\0', "disable-mte", &(cfg)->mte_disabled,			\
 			"Disable Memory Tagging Extension"),			\
-	OPT_CALLBACK('\0', "vcpu-affinity", kvm, "cpulist",  			\
-			"Specify the CPU affinity that will apply to "		\
-			"all VCPUs", vcpu_affinity_parser, kvm),		\
 	OPT_U64('\0', "kaslr-seed", &(cfg)->kaslr_seed,				\
 			"Specify random seed for Kernel Address Space "		\
 			"Layout Randomization (KASLR)"),			\
diff --git a/arm64/kvm-cpu.c b/arm64/kvm-cpu.c
index 3aa76843fda2..9063fbcdcb64 100644
--- a/arm64/kvm-cpu.c
+++ b/arm64/kvm-cpu.c
@@ -376,15 +376,6 @@ int sve_vl_parser(const struct option *opt, const char *arg, int unset)
 void kvm_cpu__reset_vcpu(struct kvm_cpu *vcpu)
 {
 	struct kvm *kvm = vcpu->kvm;
-	cpu_set_t *affinity;
-	int ret;
-
-	affinity = kvm->arch.vcpu_affinity_cpuset;
-	if (affinity) {
-		ret = sched_setaffinity(0, sizeof(cpu_set_t), affinity);
-		if (ret == -1)
-			die_perror("sched_setaffinity");
-	}
 
 	if (kvm->cfg.arch.aarch32_guest)
 		return reset_vcpu_aarch32(vcpu);
diff --git a/arm64/kvm.c b/arm64/kvm.c
index 88b3622766bf..23c630f33aa5 100644
--- a/arm64/kvm.c
+++ b/arm64/kvm.c
@@ -11,7 +11,6 @@
 #include "asm/smccc.h"
 
 #include <linux/byteorder.h>
-#include <linux/cpumask.h>
 #include <linux/kernel.h>
 #include <linux/kvm.h>
 #include <linux/sizes.h>
@@ -445,39 +444,6 @@ int kvm__arch_setup_firmware(struct kvm *kvm)
 	return 0;
 }
 
-int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
-{
-	struct kvm *kvm = opt->ptr;
-	const char *cpulist = arg;
-	cpumask_t *cpumask;
-	size_t setsize;
-	int ret = 0;
-	int cpu;
-
-	kvm->cfg.arch.vcpu_affinity = cpulist;
-
-	cpumask = calloc(cpumask_size(), 1);
-	if (!cpumask)
-		die_perror("calloc");
-
-	ret = cpulist_parse(cpulist, cpumask);
-	if (ret)
-		goto out_free;
-
-	kvm->arch.vcpu_affinity_cpuset = CPU_ALLOC(NR_CPUS);
-	if (!kvm->arch.vcpu_affinity_cpuset)
-		die_perror("CPU_ALLOC");
-
-	setsize = CPU_ALLOC_SIZE(NR_CPUS);
-	CPU_ZERO_S(setsize, kvm->arch.vcpu_affinity_cpuset);
-	for_each_cpu(cpu, cpumask)
-		CPU_SET_S(cpu, setsize, kvm->arch.vcpu_affinity_cpuset);
-
-out_free:
-	free(cpumask);
-	return ret;
-}
-
 void kvm__arch_validate_cfg(struct kvm *kvm)
 {
 
diff --git a/arm64/pmu.c b/arm64/pmu.c
index d3761689c907..bd3f225842be 100644
--- a/arm64/pmu.c
+++ b/arm64/pmu.c
@@ -200,8 +200,8 @@ static int find_pmu(struct kvm *kvm)
 
 	setsize = CPU_ALLOC_SIZE(NR_CPUS);
 
-	if (kvm->arch.vcpu_affinity_cpuset) {
-		affinity = kvm->arch.vcpu_affinity_cpuset;
+	if (kvm->vcpu_affinity) {
+		affinity = kvm->vcpu_affinity;
 	} else {
 		affinity = CPU_ALLOC(NR_CPUS);
 		if (!affinity)
@@ -224,7 +224,7 @@ static int find_pmu(struct kvm *kvm)
 
 out_free:
 	free(cpumask);
-	if (!kvm->arch.vcpu_affinity_cpuset)
+	if (!kvm->vcpu_affinity)
 		CPU_FREE(affinity);
 	return ret;
 }
diff --git a/builtin-run.c b/builtin-run.c
index 81f255f911b3..127245f9a6b8 100644
--- a/builtin-run.c
+++ b/builtin-run.c
@@ -34,6 +34,7 @@
 #include "kvm/kvm-ipc.h"
 #include "kvm/builtin-debug.h"
 
+#include <linux/cpumask.h>
 #include <linux/types.h>
 #include <linux/err.h>
 #include <linux/sizes.h>
@@ -166,6 +167,39 @@ static int loglevel_parser(const struct option *opt, const char *arg, int unset)
 	return 0;
 }
 
+static int vcpu_affinity_parser(const struct option *opt, const char *arg, int unset)
+{
+	struct kvm *kvm = opt->ptr;
+	const char *cpulist = arg;
+	cpumask_t *cpumask;
+	size_t setsize;
+	int ret = 0;
+	int cpu;
+
+	kvm->cfg.vcpu_affinity = cpulist;
+
+	cpumask = calloc(cpumask_size(), 1);
+	if (!cpumask)
+		die_perror("calloc");
+
+	ret = cpulist_parse(cpulist, cpumask);
+	if (ret)
+		goto out_free;
+
+	kvm->vcpu_affinity = CPU_ALLOC(NR_CPUS);
+	if (!kvm->vcpu_affinity)
+		die_perror("CPU_ALLOC");
+
+	setsize = CPU_ALLOC_SIZE(NR_CPUS);
+	CPU_ZERO_S(setsize, kvm->vcpu_affinity);
+	for_each_cpu(cpu, cpumask)
+		CPU_SET_S(cpu, setsize, kvm->vcpu_affinity);
+
+out_free:
+	free(cpumask);
+	return 0;
+}
+
 #ifndef OPT_ARCH_RUN
 #define OPT_ARCH_RUN(...)
 #endif
@@ -237,6 +271,9 @@ static int loglevel_parser(const struct option *opt, const char *arg, int unset)
 		     virtio_transport_parser, NULL),			\
 	OPT_CALLBACK('\0', "loglevel", NULL, "[error|warning|info|debug]",\
 			"Set the verbosity level", loglevel_parser, NULL),\
+	OPT_CALLBACK('\0', "vcpu-affinity", kvm, "cpulist",  		\
+			"Specify the CPU affinity that will apply to "	\
+			"all VCPUs", vcpu_affinity_parser, kvm),	\
 									\
 	OPT_GROUP("Kernel options:"),					\
 	OPT_STRING('k', "kernel", &(cfg)->kernel_filename, "kernel",	\
@@ -834,8 +871,8 @@ static int kvm_cmd_run_work(struct kvm *kvm)
 	int i;
 
 	for (i = 0; i < kvm->nrcpus; i++) {
-		if (pthread_create(&kvm->cpus[i]->thread, NULL, kvm_cpu_thread, kvm->cpus[i]) != 0)
-			die("unable to create KVM VCPU thread");
+		if (kvm_create_vcpu_thread(kvm, &kvm->cpus[i]->thread, kvm_cpu_thread, kvm->cpus[i]))
+			die_perror("unable to create KVM VCPU thread");
 	}
 
 	/* Only VCPU #0 is going to exit by itself when shutting down */
diff --git a/include/kvm/kvm-config.h b/include/kvm/kvm-config.h
index 592b035785c9..3f636cb2ea0f 100644
--- a/include/kvm/kvm-config.h
+++ b/include/kvm/kvm-config.h
@@ -52,6 +52,7 @@ struct kvm_config {
 	const char *hugetlbfs_path;
 	const char *custom_rootfs_name;
 	const char *real_cmdline;
+	const char *vcpu_affinity;
 	struct virtio_net_params *net_params;
 	bool single_step;
 	bool vnc;
diff --git a/include/kvm/kvm.h b/include/kvm/kvm.h
index a9376b6dd67e..01c4f5fa952d 100644
--- a/include/kvm/kvm.h
+++ b/include/kvm/kvm.h
@@ -102,6 +102,7 @@ struct kvm {
 	int                     nr_disks;
 
 	int			vm_state;
+	cpu_set_t		*vcpu_affinity;
 
 #ifdef KVM_BRLOCK_DEBUG
 	pthread_rwlock_t	brlock_sem;
diff --git a/include/kvm/util.h b/include/kvm/util.h
index 9e23431ecfe2..0f5a4bba5714 100644
--- a/include/kvm/util.h
+++ b/include/kvm/util.h
@@ -19,6 +19,7 @@
 #include <signal.h>
 #include <errno.h>
 #include <limits.h>
+#include <pthread.h>
 #include <sys/param.h>
 #include <sys/types.h>
 #include <linux/types.h>
@@ -150,4 +151,7 @@ void *mmap_hugetlbfs(struct kvm *kvm, const char *htlbfs_path, u64 size);
 void *mmap_anon_or_hugetlbfs(struct kvm *kvm, const char *hugetlbfs_path, u64 size);
 void *mmap_guest_memfd(struct kvm *kvm, u64 size);
 
+int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
+			   void *(*start_routine)(void *), void *arg);
+
 #endif /* KVM__UTIL_H */
diff --git a/kvm.c b/kvm.c
index 96583f916442..e29b73d27f33 100644
--- a/kvm.c
+++ b/kvm.c
@@ -182,6 +182,9 @@ int kvm__exit(struct kvm *kvm)
 		free(bank);
 	}
 
+	if (kvm->vcpu_affinity)
+		CPU_FREE(kvm->vcpu_affinity);
+
 	free(kvm);
 	return 0;
 }
diff --git a/util/util.c b/util/util.c
index 51a45916fd94..ff888bf9e061 100644
--- a/util/util.c
+++ b/util/util.c
@@ -206,3 +206,32 @@ void *mmap_guest_memfd(struct kvm *kvm, u64 size)
 
 	return addr;
 }
+
+int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
+			   void *(*start_routine)(void *), void *arg)
+{
+	pthread_attr_t attr;
+	cpu_set_t *affinity;
+	int ret;
+
+	pthread_attr_init(&attr);
+
+	affinity = kvm->vcpu_affinity;
+	if (affinity) {
+		ret = pthread_attr_setaffinity_np(&attr, sizeof(cpu_set_t), affinity);
+		if (ret) {
+			/* Set errno so the caller can call die_perror(). */
+			errno = ret;
+			/* kvmtool treats a negative return value as an error. */
+			return -ret;
+		}
+	}
+
+	ret = pthread_create(thread, &attr, start_routine, arg);
+	if (ret)
+		errno = ret;
+
+	pthread_attr_destroy(&attr);
+
+	return -ret;
+}
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [RFC PATCH kvmtool 6/7] Introduce kvm_create_thread()
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
                   ` (4 preceding siblings ...)
  2026-09-17 15:49 ` [PATCH kvmtool 5/7] Correctly apply --vcpu-affinity to the VCPU threads Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-17 15:49 ` [RFC PATCH kvmtool 7/7] Don't apply --vcpu-affinity to threads spawned from VCPUs Alexandru Elisei
  2026-09-18 12:22 ` [PATCH kvmtool 0/7] Fix --vcpu-affinity Marc Zyngier
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

Switch over all pthread_create() invocations to kvm_create_thread(). Also
add error checking where it was missing.

Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 disk/aio.c               |  3 +--
 disk/blk.c               | 12 +++++++++--
 disk/core.c              | 29 +++++++++++++-------------
 disk/qcow.c              | 44 ++++++++++++++++++++++++++++------------
 disk/raw.c               | 21 +++++++++++++++----
 epoll.c                  |  2 +-
 include/kvm/disk-image.h | 28 +++++++++++++++++--------
 include/kvm/qcow.h       |  3 ++-
 include/kvm/uip.h        |  2 ++
 include/kvm/util.h       |  2 ++
 net/uip/tcp.c            |  7 +++----
 net/uip/udp.c            |  5 +++--
 term.c                   |  4 ++--
 ui/gtk3.c                |  2 +-
 ui/sdl.c                 |  2 +-
 ui/vnc.c                 |  2 +-
 util/threadpool.c        |  9 ++++----
 util/util.c              | 12 +++++++++++
 virtio/blk.c             |  2 +-
 virtio/net.c             | 26 ++++++++++++++++--------
 20 files changed, 146 insertions(+), 71 deletions(-)

diff --git a/disk/aio.c b/disk/aio.c
index a7418c8c261d..365f075272dd 100644
--- a/disk/aio.c
+++ b/disk/aio.c
@@ -127,9 +127,8 @@ int disk_aio_setup(struct disk_image *disk)
 		return -errno;
 
 	io_setup(AIO_MAX, &disk->ctx);
-	r = pthread_create(&disk->thread, NULL, disk_aio_thread, disk);
+	r = kvm_create_thread(disk->kvm, &disk->thread, disk_aio_thread, disk);
 	if (r) {
-		r = -errno;
 		close(disk->evt);
 		return r;
 	}
diff --git a/disk/blk.c b/disk/blk.c
index b4c9fba3bcec..82e1937a1906 100644
--- a/disk/blk.c
+++ b/disk/blk.c
@@ -35,8 +35,14 @@ static bool is_mounted(struct stat *st)
 	return false;
 }
 
-struct disk_image *blkdev__probe(const char *filename, int flags, struct stat *st)
+struct disk_image *blkdev__probe(const char *filename, int flags, struct stat *st,
+				 struct kvm *kvm)
 {
+	struct new_disk_image ndi = {
+		.ops	= &blk_dev_ops,
+		.kvm	= kvm,
+		.flags	= DISK_IMAGE_REGULAR,
+	};
 	int fd, r;
 	u64 size;
 
@@ -56,17 +62,19 @@ struct disk_image *blkdev__probe(const char *filename, int flags, struct stat *s
 	fd = open(filename, flags);
 	if (fd < 0)
 		return ERR_PTR(fd);
+	ndi.fd = fd;
 
 	if (ioctl(fd, BLKGETSIZE64, &size) < 0) {
 		r = -errno;
 		close(fd);
 		return ERR_PTR(r);
 	}
+	ndi.size = size;
 
 	/*
 	 * FIXME: This will not work on 32-bit host because we can not
 	 * mmap large disk. There is not enough virtual address space
 	 * in 32-bit host. However, this works on 64-bit host.
 	 */
-	return disk_image__new(fd, size, &blk_dev_ops, DISK_IMAGE_REGULAR);
+	return disk_image__new(&ndi);
 }
diff --git a/disk/core.c b/disk/core.c
index b232eece9e73..5a1692fc7951 100644
--- a/disk/core.c
+++ b/disk/core.c
@@ -59,9 +59,7 @@ int disk_img_name_parser(const struct option *opt, const char *arg, int unset)
 	return 0;
 }
 
-struct disk_image *disk_image__new(int fd, u64 size,
-				   struct disk_image_operations *ops,
-				   int use_mmap)
+struct disk_image *disk_image__new(struct new_disk_image *ndi)
 {
 	struct disk_image *disk;
 	int r;
@@ -71,16 +69,18 @@ struct disk_image *disk_image__new(int fd, u64 size,
 		return ERR_PTR(-ENOMEM);
 
 	*disk = (struct disk_image) {
-		.fd	= fd,
-		.size	= size,
-		.ops	= ops,
+		.fd	= ndi->fd,
+		.size	= ndi->size,
+		.ops	= ndi->ops,
+		.kvm	= ndi->kvm,
 	};
 
-	if (use_mmap == DISK_IMAGE_MMAP) {
+	if (ndi->flags == DISK_IMAGE_MMAP) {
 		/*
 		 * The write to disk image will be discarded
 		 */
-		disk->priv = mmap(NULL, size, PROT_RW, MAP_PRIVATE | MAP_NORESERVE, fd, 0);
+		disk->priv = mmap(NULL, ndi->size, PROT_RW,
+				  MAP_PRIVATE | MAP_NORESERVE, ndi->fd, 0);
 		if (disk->priv == MAP_FAILED) {
 			r = -errno;
 			goto err_free_disk;
@@ -95,13 +95,14 @@ struct disk_image *disk_image__new(int fd, u64 size,
 
 err_unmap_disk:
 	if (disk->priv)
-		munmap(disk->priv, size);
+		munmap(disk->priv, ndi->size);
 err_free_disk:
 	free(disk);
 	return ERR_PTR(r);
 }
 
-static struct disk_image *disk_image__open(const char *filename, bool readonly, bool direct)
+static struct disk_image *disk_image__open(const char *filename, bool readonly, bool direct,
+					   struct kvm *kvm)
 {
 	struct disk_image *disk;
 	struct stat st;
@@ -118,7 +119,7 @@ static struct disk_image *disk_image__open(const char *filename, bool readonly,
 		return ERR_PTR(-errno);
 
 	/* blk device ?*/
-	disk = blkdev__probe(filename, flags, &st);
+	disk = blkdev__probe(filename, flags, &st, kvm);
 	if (!IS_ERR_OR_NULL(disk)) {
 		disk->readonly = readonly;
 		return disk;
@@ -129,7 +130,7 @@ static struct disk_image *disk_image__open(const char *filename, bool readonly,
 		return ERR_PTR(fd);
 
 	/* qcow image ?*/
-	disk = qcow_probe(fd, true);
+	disk = qcow_probe(fd, true, kvm);
 	if (!IS_ERR_OR_NULL(disk)) {
 		pr_warning("Forcing read-only support for QCOW");
 		disk->readonly = true;
@@ -137,7 +138,7 @@ static struct disk_image *disk_image__open(const char *filename, bool readonly,
 	}
 
 	/* raw image ?*/
-	disk = raw_image__probe(fd, &st, readonly);
+	disk = raw_image__probe(fd, &st, readonly, kvm);
 	if (!IS_ERR_OR_NULL(disk)) {
 		disk->readonly = readonly;
 		return disk;
@@ -188,7 +189,7 @@ static struct disk_image **disk_image__open_all(struct kvm *kvm)
 		if (!filename)
 			continue;
 
-		disks[i] = disk_image__open(filename, readonly, direct);
+		disks[i] = disk_image__open(filename, readonly, direct, kvm);
 		if (IS_ERR_OR_NULL(disks[i])) {
 			pr_err("Loading disk image '%s' failed", filename);
 			err = disks[i];
diff --git a/disk/qcow.c b/disk/qcow.c
index dd6be62ee183..eed18bb169bb 100644
--- a/disk/qcow.c
+++ b/disk/qcow.c
@@ -1273,8 +1273,13 @@ static void *qcow2_read_header(int fd)
 	return header;
 }
 
-static struct disk_image *qcow2_probe(int fd, bool readonly)
+static struct disk_image *qcow2_probe(int fd, bool readonly, struct kvm *kvm)
 {
+	struct new_disk_image ndi = {
+		.kvm	= kvm,
+		.fd	= fd,
+		.flags	= DISK_IMAGE_REGULAR,
+	};
 	struct disk_image *disk_image;
 	struct qcow_l1_table *l1t;
 	struct qcow_header *h;
@@ -1295,6 +1300,7 @@ static struct disk_image *qcow2_probe(int fd, bool readonly)
 	h = q->header = qcow2_read_header(fd);
 	if (!h)
 		goto free_qcow;
+	ndi.size = h->size;
 
 	q->version = QCOW2_VERSION;
 	q->csize_shift = (62 - (q->header->cluster_bits - 8));
@@ -1329,10 +1335,13 @@ static struct disk_image *qcow2_probe(int fd, bool readonly)
 	/*
 	 * Do not use mmap use read/write instead
 	 */
-	if (readonly)
-		disk_image = disk_image__new(fd, h->size, &qcow_disk_readonly_ops, DISK_IMAGE_REGULAR);
-	else
-		disk_image = disk_image__new(fd, h->size, &qcow_disk_ops, DISK_IMAGE_REGULAR);
+	if (readonly) {
+		ndi.ops = &qcow_disk_readonly_ops;
+		disk_image = disk_image__new(&ndi);
+	} else {
+		ndi.ops = &qcow_disk_ops;
+		disk_image = disk_image__new(&ndi);
+	}
 
 	if (IS_ERR_OR_NULL(disk_image))
 		goto free_refcount_table;
@@ -1418,8 +1427,13 @@ static void *qcow1_read_header(int fd)
 	return header;
 }
 
-static struct disk_image *qcow1_probe(int fd, bool readonly)
+static struct disk_image *qcow1_probe(int fd, bool readonly, struct kvm *kvm)
 {
+	struct new_disk_image ndi = {
+		.kvm	= kvm,
+		.fd	= fd,
+		.flags	= DISK_IMAGE_REGULAR,
+	};
 	struct disk_image *disk_image;
 	struct qcow_l1_table *l1t;
 	struct qcow_header *h;
@@ -1441,6 +1455,7 @@ static struct disk_image *qcow1_probe(int fd, bool readonly)
 	h = q->header = qcow1_read_header(fd);
 	if (!h)
 		goto free_qcow;
+	ndi.size = h->size;
 
 	q->version = QCOW1_VERSION;
 	q->cluster_size = 1 << q->header->cluster_bits;
@@ -1465,10 +1480,13 @@ static struct disk_image *qcow1_probe(int fd, bool readonly)
 	/*
 	 * Do not use mmap use read/write instead
 	 */
-	if (readonly)
-		disk_image = disk_image__new(fd, h->size, &qcow_disk_readonly_ops, DISK_IMAGE_REGULAR);
-	else
-		disk_image = disk_image__new(fd, h->size, &qcow_disk_ops, DISK_IMAGE_REGULAR);
+	if (readonly) {
+		ndi.ops = &qcow_disk_readonly_ops;
+		disk_image = disk_image__new(&ndi);
+	} else {
+		ndi.ops = &qcow_disk_ops;
+		disk_image = disk_image__new(&ndi);
+	}
 
 	if (!disk_image)
 		goto free_l1_table;
@@ -1514,13 +1532,13 @@ static bool qcow1_check_image(int fd)
 	return true;
 }
 
-struct disk_image *qcow_probe(int fd, bool readonly)
+struct disk_image *qcow_probe(int fd, bool readonly, struct kvm *kvm)
 {
 	if (qcow1_check_image(fd))
-		return qcow1_probe(fd, readonly);
+		return qcow1_probe(fd, readonly, kvm);
 
 	if (qcow2_check_image(fd))
-		return qcow2_probe(fd, readonly);
+		return qcow2_probe(fd, readonly, kvm);
 
 	return NULL;
 }
diff --git a/disk/raw.c b/disk/raw.c
index 54b4e7408661..9ad09d20db12 100644
--- a/disk/raw.c
+++ b/disk/raw.c
@@ -83,8 +83,14 @@ struct disk_image_operations ro_ops_nowrite = {
 	.async	= true,
 };
 
-struct disk_image *raw_image__probe(int fd, struct stat *st, bool readonly)
+struct disk_image *raw_image__probe(int fd, struct stat *st, bool readonly,
+				    struct kvm *kvm)
 {
+	struct new_disk_image new = {
+		.kvm	= kvm,
+		.size	= st->st_size,
+		.fd	= fd,
+	};
 	if (readonly) {
 		/*
 		 * Use mmap's MAP_PRIVATE to implement non-persistent write
@@ -92,16 +98,23 @@ struct disk_image *raw_image__probe(int fd, struct stat *st, bool readonly)
 		 */
 		struct disk_image *disk;
 
-		disk = disk_image__new(fd, st->st_size, &ro_ops, DISK_IMAGE_MMAP);
+		new.ops = &ro_ops;
+		new.flags = DISK_IMAGE_MMAP;
+		disk = disk_image__new(&new);
+
 		if (IS_ERR_OR_NULL(disk)) {
-			disk = disk_image__new(fd, st->st_size, &ro_ops_nowrite, DISK_IMAGE_REGULAR);
+			new.ops = &ro_ops_nowrite;
+			new.flags = DISK_IMAGE_REGULAR;
+			disk = disk_image__new(&new);
 		}
 
 		return disk;
 	} else {
+		new.ops = &raw_image_regular_ops;
+		new.flags = DISK_IMAGE_REGULAR;
 		/*
 		 * Use read/write instead of mmap
 		 */
-		return disk_image__new(fd, st->st_size, &raw_image_regular_ops, DISK_IMAGE_REGULAR);
+		return disk_image__new(&new);
 	}
 }
diff --git a/epoll.c b/epoll.c
index 8cb0cee5822e..0ddd235d89a8 100644
--- a/epoll.c
+++ b/epoll.c
@@ -58,7 +58,7 @@ int epoll__init(struct kvm *kvm, struct kvm__epoll *epoll,
 	if (r < 0)
 		goto err_close_all;
 
-	r = pthread_create(&epoll->thread, NULL, epoll__thread, epoll);
+	r = kvm_create_thread(kvm, &epoll->thread, epoll__thread, epoll);
 	if (r < 0)
 		goto err_close_all;
 
diff --git a/include/kvm/disk-image.h b/include/kvm/disk-image.h
index cbe91b07af0b..f1c61f0403e3 100644
--- a/include/kvm/disk-image.h
+++ b/include/kvm/disk-image.h
@@ -26,11 +26,6 @@
 #define SECTOR_SHIFT		9
 #define SECTOR_SIZE		(1UL << SECTOR_SHIFT)
 
-enum {
-	DISK_IMAGE_REGULAR,
-	DISK_IMAGE_MMAP,
-};
-
 #define MAX_DISK_IMAGES         4
 
 struct disk_image;
@@ -54,6 +49,20 @@ struct disk_image_params {
 	bool direct;
 };
 
+enum new_disk_image_flags {
+	DISK_IMAGE_REGULAR,
+	DISK_IMAGE_MMAP,
+};
+
+struct kvm;
+struct new_disk_image {
+	struct disk_image_operations	*ops;
+	struct kvm			*kvm;
+	u64				size;
+	int				fd;
+	enum new_disk_image_flags	flags;
+};
+
 struct disk_image {
 	int				fd;
 	u64				size;
@@ -61,6 +70,7 @@ struct disk_image {
 	void				*priv;
 	void				*disk_req_cb_param;
 	void				(*disk_req_cb)(void *param, long len);
+	struct kvm			*kvm;
 	bool				readonly;
 	bool				async;
 #ifdef CONFIG_HAS_AIO
@@ -76,7 +86,7 @@ struct disk_image {
 int disk_img_name_parser(const struct option *opt, const char *arg, int unset);
 int disk_image__init(struct kvm *kvm);
 int disk_image__exit(struct kvm *kvm);
-struct disk_image *disk_image__new(int fd, u64 size, struct disk_image_operations *ops, int mmap);
+struct disk_image *disk_image__new(struct new_disk_image *ndi);
 int disk_image__flush(struct disk_image *disk);
 int disk_image__wait(struct disk_image *disk);
 ssize_t disk_image__read(struct disk_image *disk, u64 sector, const struct iovec *iov,
@@ -86,8 +96,10 @@ ssize_t disk_image__write(struct disk_image *disk, u64 sector, const struct iove
 ssize_t disk_image__get_serial(struct disk_image *disk, struct iovec *iov,
 			       int iovcount, ssize_t len);
 
-struct disk_image *raw_image__probe(int fd, struct stat *st, bool readonly);
-struct disk_image *blkdev__probe(const char *filename, int flags, struct stat *st);
+struct disk_image *raw_image__probe(int fd, struct stat *st, bool readonly,
+				    struct kvm *kvm);
+struct disk_image *blkdev__probe(const char *filename, int flags, struct stat *st,
+				 struct kvm *kvm);
 
 ssize_t raw_image__read_sync(struct disk_image *disk, u64 sector,
 			     const struct iovec *iov, int iovcount, void *param);
diff --git a/include/kvm/qcow.h b/include/kvm/qcow.h
index f8492462ddaa..c0293ef8bf75 100644
--- a/include/kvm/qcow.h
+++ b/include/kvm/qcow.h
@@ -128,6 +128,7 @@ struct qcow2_header_disk {
 	u64				snapshots_offset;
 };
 
-struct disk_image *qcow_probe(int fd, bool readonly);
+struct kvm;
+struct disk_image *qcow_probe(int fd, bool readonly, struct kvm *kvm);
 
 #endif /* KVM__QCOW_H */
diff --git a/include/kvm/uip.h b/include/kvm/uip.h
index efa508a50f51..39db4b656fd5 100644
--- a/include/kvm/uip.h
+++ b/include/kvm/uip.h
@@ -184,6 +184,7 @@ struct uip_dhcp {
 	u8 option[UIP_DHCP_OPTION_LEN];
 } __attribute__((packed));
 
+struct kvm;
 struct uip_info {
 	struct list_head udp_socket_head;
 	struct list_head tcp_socket_head;
@@ -196,6 +197,7 @@ struct uip_info {
 	struct list_head buf_head;
 	struct mutex buf_lock;
 	pthread_t udp_thread;
+	struct kvm *kvm;
 	u8 *udp_buf;
 	int udp_epollfd;
 	int buf_free_nr;
diff --git a/include/kvm/util.h b/include/kvm/util.h
index 0f5a4bba5714..86edb77edfce 100644
--- a/include/kvm/util.h
+++ b/include/kvm/util.h
@@ -151,6 +151,8 @@ void *mmap_hugetlbfs(struct kvm *kvm, const char *htlbfs_path, u64 size);
 void *mmap_anon_or_hugetlbfs(struct kvm *kvm, const char *hugetlbfs_path, u64 size);
 void *mmap_guest_memfd(struct kvm *kvm, u64 size);
 
+int kvm_create_thread(struct kvm *kvm, pthread_t *thread,
+		      void *(*start_routine)(void *), void *arg);
 int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
 			   void *(*start_routine)(void *), void *arg);
 
diff --git a/net/uip/tcp.c b/net/uip/tcp.c
index 42e6e992cd6a..54252016ac02 100644
--- a/net/uip/tcp.c
+++ b/net/uip/tcp.c
@@ -253,7 +253,7 @@ out:
 	return NULL;
 }
 
-static int uip_tcp_socket_receive(struct uip_tcp_socket *sk)
+static int uip_tcp_socket_receive(struct uip_tcp_socket *sk, struct kvm *kvm)
 {
 	int ret;
 
@@ -261,8 +261,7 @@ static int uip_tcp_socket_receive(struct uip_tcp_socket *sk)
 		sk->buf = malloc(UIP_MAX_TCP_PAYLOAD);
 		if (!sk->buf)
 			return -ENOMEM;
-		ret = pthread_create(&sk->thread, NULL, uip_tcp_socket_thread,
-				     (void *)sk);
+		ret = kvm_create_thread(kvm, &sk->thread, uip_tcp_socket_thread, sk);
 		if (ret)
 			free(sk->buf);
 		return ret;
@@ -324,7 +323,7 @@ int uip_tx_do_ipv4_tcp(struct uip_tx_arg *arg)
 		/*
 		 * Start receive thread for data from remote to guest
 		 */
-		uip_tcp_socket_receive(sk);
+		uip_tcp_socket_receive(sk, arg->info->kvm);
 
 		goto out;
 	}
diff --git a/net/uip/udp.c b/net/uip/udp.c
index d2580d06e851..e63c8ad97c63 100644
--- a/net/uip/udp.c
+++ b/net/uip/udp.c
@@ -238,10 +238,11 @@ int uip_tx_do_ipv4_udp(struct uip_tx_arg *arg)
 		if (!info->udp_buf)
 			return -1;
 
-		pthread_create(&info->udp_thread, NULL, uip_udp_socket_thread, (void *)info);
+		ret = kvm_create_thread(info->kvm, &info->udp_thread,
+					uip_udp_socket_thread, info);
 	}
 
-	return 0;
+	return ret;
 }
 
 void uip_udp_exit(struct uip_info *info)
diff --git a/term.c b/term.c
index b8a70fe2ab7b..e1b31737e0f7 100644
--- a/term.c
+++ b/term.c
@@ -196,8 +196,8 @@ static int term_init(struct kvm *kvm)
 
 
 	/* Use our own blocking thread to read stdin, don't require a tick */
-	if(pthread_create(&term_poll_thread, NULL, term_poll_thread_loop,kvm))
-		die("Unable to create console input poll thread\n");
+	if (kvm_create_thread(kvm, &term_poll_thread, term_poll_thread_loop, kvm))
+		die_perror("Unable to create console input poll thread");
 
 	signal(SIGTERM, term_sig_cleanup);
 	atexit(term_cleanup);
diff --git a/ui/gtk3.c b/ui/gtk3.c
index 1e08a8f6b76a..b6f3a3000b01 100644
--- a/ui/gtk3.c
+++ b/ui/gtk3.c
@@ -277,7 +277,7 @@ static int kvm_gtk_start(struct framebuffer *fb)
 {
 	pthread_t thread;
 
-	if (pthread_create(&thread, NULL, kvm_gtk_thread, fb) != 0)
+	if (kvm_create_thread(fb->kvm, &thread, kvm_gtk_thread, fb) != 0)
 		return -1;
 
 	return 0;
diff --git a/ui/sdl.c b/ui/sdl.c
index 5035405bb488..2014ec5bcfe9 100644
--- a/ui/sdl.c
+++ b/ui/sdl.c
@@ -277,7 +277,7 @@ static int sdl__start(struct framebuffer *fb)
 
 	running = true;
 
-	if (pthread_create(&thread, NULL, sdl__thread, fb) != 0)
+	if (kvm_create_thread(fb->kvm, &thread, sdl__thread, fb) != 0)
 		return -1;
 
 	return 0;
diff --git a/ui/vnc.c b/ui/vnc.c
index 12e4bd53fe0d..8371b3cc9f8c 100644
--- a/ui/vnc.c
+++ b/ui/vnc.c
@@ -205,7 +205,7 @@ static int vnc__start(struct framebuffer *fb)
 {
 	pthread_t thread;
 
-	if (pthread_create(&thread, NULL, vnc__thread, fb) != 0)
+	if (kvm_create_thread(fb->kvm, &thread, vnc__thread, fb) != 0)
 		return -1;
 
 	return 0;
diff --git a/util/threadpool.c b/util/threadpool.c
index 1dc3bf7e7ef2..05722b5c09a4 100644
--- a/util/threadpool.c
+++ b/util/threadpool.c
@@ -97,7 +97,7 @@ static void *thread_pool__threadfunc(void *param)
 	return NULL;
 }
 
-static int thread_pool__addthread(void)
+static int thread_pool__addthread(struct kvm *kvm)
 {
 	int res;
 	void *newthreads;
@@ -111,9 +111,8 @@ static int thread_pool__addthread(void)
 
 	threads = newthreads;
 
-	res = pthread_create(threads + threadcount, NULL,
-			     thread_pool__threadfunc, NULL);
-
+	res = kvm_create_thread(kvm, threads + threadcount,
+				thread_pool__threadfunc, NULL);
 	if (res == 0)
 		threadcount++;
 	mutex_unlock(&thread_mutex);
@@ -129,7 +128,7 @@ int thread_pool__init(struct kvm *kvm)
 	running = true;
 
 	for (i = 0; i < thread_count; i++)
-		if (thread_pool__addthread() < 0)
+		if (thread_pool__addthread(kvm) < 0)
 			return i;
 
 	return i;
diff --git a/util/util.c b/util/util.c
index ff888bf9e061..7238fe57bace 100644
--- a/util/util.c
+++ b/util/util.c
@@ -207,6 +207,18 @@ void *mmap_guest_memfd(struct kvm *kvm, u64 size)
 	return addr;
 }
 
+int kvm_create_thread(struct kvm *kvm, pthread_t *thread,
+		      void *(*start_routine)(void *), void *arg)
+{
+	int ret;
+
+	ret = pthread_create(thread, NULL, start_routine, arg);
+	if (ret)
+		errno = ret;
+
+	return -ret;
+}
+
 int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
 			   void *(*start_routine)(void *), void *arg)
 {
diff --git a/virtio/blk.c b/virtio/blk.c
index b2d6180d118a..0ffafff1663b 100644
--- a/virtio/blk.c
+++ b/virtio/blk.c
@@ -237,7 +237,7 @@ static int init_vq(struct kvm *kvm, void *dev, u32 vq)
 	if (bdev->io_efd < 0)
 		return -errno;
 
-	if (pthread_create(&bdev->io_thread, NULL, virtio_blk_thread, bdev))
+	if (kvm_create_thread(kvm, &bdev->io_thread, virtio_blk_thread, bdev))
 		return -errno;
 
 	return 0;
diff --git a/virtio/net.c b/virtio/net.c
index 492c57675b1f..097bc32fc12b 100644
--- a/virtio/net.c
+++ b/virtio/net.c
@@ -610,17 +610,23 @@ static int init_vq(struct kvm *kvm, void *dev, u32 vq)
 	mutex_init(&net_queue->lock);
 	pthread_cond_init(&net_queue->cond, NULL);
 	if (is_ctrl_vq(ndev, vq)) {
-		pthread_create(&net_queue->thread, NULL, virtio_net_ctrl_thread,
-			       net_queue);
-
+		r = kvm_create_thread(kvm, &net_queue->thread,
+				      virtio_net_ctrl_thread, net_queue);
+		if (r)
+			die_perror("virtio_net_ctrl_thread");
 		return 0;
 	} else if (ndev->vhost_fd == 0 ) {
-		if (vq & 1)
-			pthread_create(&net_queue->thread, NULL,
-				       virtio_net_tx_thread, net_queue);
-		else
-			pthread_create(&net_queue->thread, NULL,
-				       virtio_net_rx_thread, net_queue);
+		if (vq & 1) {
+			r = kvm_create_thread(kvm, &net_queue->thread,
+					      virtio_net_tx_thread, net_queue);
+			if (r)
+				die_perror("virtio_net_tx_thread");
+		} else {
+			r = kvm_create_thread(kvm, &net_queue->thread,
+					      virtio_net_rx_thread, net_queue);
+			if (r)
+				die_perror("virtio_net_rx_thread");
+		}
 
 		return 0;
 	}
@@ -870,6 +876,8 @@ static int virtio_net__init_one(struct virtio_net_params *params)
 	mutex_init(&ndev->mutex);
 	ndev->queue_pairs = max(1, min(VIRTIO_NET_NUM_QUEUES, params->mq));
 
+	ndev->info.kvm = params->kvm;
+
 	for (i = 0 ; i < 6 ; i++) {
 		ndev->config.mac[i]		= params->guest_mac[i];
 		ndev->info.guest_mac.addr[i]	= params->guest_mac[i];
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [RFC PATCH kvmtool 7/7] Don't apply --vcpu-affinity to threads spawned from VCPUs
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
                   ` (5 preceding siblings ...)
  2026-09-17 15:49 ` [RFC PATCH kvmtool 6/7] Introduce kvm_create_thread() Alexandru Elisei
@ 2026-09-17 15:49 ` Alexandru Elisei
  2026-09-18 12:22 ` [PATCH kvmtool 0/7] Fix --vcpu-affinity Marc Zyngier
  7 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-17 15:49 UTC (permalink / raw)
  To: will, julien.thierry.kdev, kvm, maz, oupton, fuad.tabba,
	joey.gouly, seiden, suzuki.poulose, yuzenghui, linux-arm-kernel,
	kvmarm

Threads spawned with pthread_create() inherit the parent's affinity. For
kvmtool, threads spawned from the VCPU threads, like the virtio threads,
will inherit the affinity specified with --vcpu-affinity, instead of
inheriting the main process' affinity. Make kvm_create_thread() use the
main thread's affinity when spawning new threads.

A command like this:

$ taskset -c 0-5 ./lkvm run .. --vcpu-affinity 6-11

will work as expected and all kvmtool threads will be pinned to CPUs 0-5,
instead of some of them being pinned on the same CPUs as the VCPUs.

Signed-off-by: Alexandru Elisei <alexandru.elisei@arm.com>
---
 arm64/pmu.c       | 19 +++----------------
 include/kvm/kvm.h |  1 +
 kvm.c             | 15 +++++++++++++++
 util/util.c       | 33 +++++++++++++++++----------------
 4 files changed, 36 insertions(+), 32 deletions(-)

diff --git a/arm64/pmu.c b/arm64/pmu.c
index bd3f225842be..10d1641b42cc 100644
--- a/arm64/pmu.c
+++ b/arm64/pmu.c
@@ -200,20 +200,10 @@ static int find_pmu(struct kvm *kvm)
 
 	setsize = CPU_ALLOC_SIZE(NR_CPUS);
 
-	if (kvm->vcpu_affinity) {
+	if (kvm->vcpu_affinity)
 		affinity = kvm->vcpu_affinity;
-	} else {
-		affinity = CPU_ALLOC(NR_CPUS);
-		if (!affinity)
-			die_perror("CPU_ALLOC");
-		CPU_ZERO_S(setsize, affinity);
-
-		ret = sched_getaffinity(0, setsize, affinity);
-		if (ret < 0) {
-			ret = -errno;
-			goto out_free;
-		}
-	}
+	else
+		affinity = kvm->main_affinity;
 
 	for (i = 0; i < NR_CPUS; i ++) {
 		if (CPU_ISSET_S(i, setsize, affinity))
@@ -222,10 +212,7 @@ static int find_pmu(struct kvm *kvm)
 
 	ret = find_pmu_cpumask(kvm, cpumask);
 
-out_free:
 	free(cpumask);
-	if (!kvm->vcpu_affinity)
-		CPU_FREE(affinity);
 	return ret;
 }
 
diff --git a/include/kvm/kvm.h b/include/kvm/kvm.h
index 01c4f5fa952d..91a30bbbbfd0 100644
--- a/include/kvm/kvm.h
+++ b/include/kvm/kvm.h
@@ -102,6 +102,7 @@ struct kvm {
 	int                     nr_disks;
 
 	int			vm_state;
+	cpu_set_t		*main_affinity;
 	cpu_set_t		*vcpu_affinity;
 
 #ifdef KVM_BRLOCK_DEBUG
diff --git a/kvm.c b/kvm.c
index e29b73d27f33..7bc91cce97b0 100644
--- a/kvm.c
+++ b/kvm.c
@@ -156,6 +156,9 @@ static int kvm__check_extensions(struct kvm *kvm)
 struct kvm *kvm__new(void)
 {
 	struct kvm *kvm = calloc(1, sizeof(*kvm));
+	size_t setsize;
+	int ret;
+
 	if (!kvm)
 		return ERR_PTR(-ENOMEM);
 
@@ -168,6 +171,17 @@ struct kvm *kvm__new(void)
 	kvm->brlock_sem = (pthread_rwlock_t) PTHREAD_RWLOCK_INITIALIZER;
 #endif
 
+	kvm->main_affinity = CPU_ALLOC(NR_CPUS);
+	if (!kvm->main_affinity)
+		die_perror("CPU_ALLOC");
+
+	setsize = CPU_ALLOC_SIZE(NR_CPUS);
+	CPU_ZERO_S(setsize, kvm->main_affinity);
+
+	ret = sched_getaffinity(0, setsize, kvm->main_affinity);
+	if (ret)
+		die_perror("sched_getaffinity");
+
 	return kvm;
 }
 
@@ -182,6 +196,7 @@ int kvm__exit(struct kvm *kvm)
 		free(bank);
 	}
 
+	CPU_FREE(kvm->main_affinity);
 	if (kvm->vcpu_affinity)
 		CPU_FREE(kvm->vcpu_affinity);
 
diff --git a/util/util.c b/util/util.c
index 7238fe57bace..a779eb59817c 100644
--- a/util/util.c
+++ b/util/util.c
@@ -207,28 +207,15 @@ void *mmap_guest_memfd(struct kvm *kvm, u64 size)
 	return addr;
 }
 
-int kvm_create_thread(struct kvm *kvm, pthread_t *thread,
-		      void *(*start_routine)(void *), void *arg)
-{
-	int ret;
-
-	ret = pthread_create(thread, NULL, start_routine, arg);
-	if (ret)
-		errno = ret;
-
-	return -ret;
-}
-
-int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
-			   void *(*start_routine)(void *), void *arg)
+static int __kvm_create_thread(struct kvm *kvm, cpu_set_t *affinity,
+			       pthread_t *thread, void *(*start_routine)(void *),
+			       void *arg)
 {
 	pthread_attr_t attr;
-	cpu_set_t *affinity;
 	int ret;
 
 	pthread_attr_init(&attr);
 
-	affinity = kvm->vcpu_affinity;
 	if (affinity) {
 		ret = pthread_attr_setaffinity_np(&attr, sizeof(cpu_set_t), affinity);
 		if (ret) {
@@ -247,3 +234,17 @@ int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
 
 	return -ret;
 }
+
+int kvm_create_thread(struct kvm *kvm, pthread_t *thread,
+		      void *(*start_routine)(void *), void *arg)
+{
+	return __kvm_create_thread(kvm, kvm->main_affinity, thread,
+				   start_routine, arg);
+}
+
+int kvm_create_vcpu_thread(struct kvm *kvm, pthread_t *thread,
+			   void *(*start_routine)(void *), void *arg)
+{
+	return __kvm_create_thread(kvm, kvm->vcpu_affinity, thread,
+				   start_routine, arg);
+}
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* Re: [PATCH kvmtool 0/7] Fix --vcpu-affinity
  2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
                   ` (6 preceding siblings ...)
  2026-09-17 15:49 ` [RFC PATCH kvmtool 7/7] Don't apply --vcpu-affinity to threads spawned from VCPUs Alexandru Elisei
@ 2026-09-18 12:22 ` Marc Zyngier
  2026-09-21  8:59   ` Alexandru Elisei
  7 siblings, 1 reply; 10+ messages in thread
From: Marc Zyngier @ 2026-09-18 12:22 UTC (permalink / raw)
  To: Alexandru Elisei
  Cc: will, julien.thierry.kdev, kvm, oupton, fuad.tabba, joey.gouly,
	seiden, suzuki.poulose, yuzenghui, linux-arm-kernel, kvmarm

On Thu, 17 Sep 2026 16:49:25 +0100,
Alexandru Elisei <alexandru.elisei@arm.com> wrote:
> 
> Tested on an Orion board and a x86 machine. On the x86 machine, there's was this
> one thread called 'kvm-nx-lpage-re' which had affinity of the VCPUs, but I
> couldn't figure out where it was created. I grep'ed for each part of the name in
> the source node, no luck. I used strace, and couldn't see a clone call that
> returned the corresponding pid. Any clues would be appreciated.

This is a KVM internal thread, from what I can tell. See
arch/x86/kvm/mmu/mmu.c::kvm_mmu_start_lpage_recovery().

	M.

-- 
Without deviation from the norm, progress is not possible.

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH kvmtool 0/7] Fix --vcpu-affinity
  2026-09-18 12:22 ` [PATCH kvmtool 0/7] Fix --vcpu-affinity Marc Zyngier
@ 2026-09-21  8:59   ` Alexandru Elisei
  0 siblings, 0 replies; 10+ messages in thread
From: Alexandru Elisei @ 2026-09-21  8:59 UTC (permalink / raw)
  To: Marc Zyngier
  Cc: will, julien.thierry.kdev, kvm, oupton, fuad.tabba, joey.gouly,
	seiden, suzuki.poulose, yuzenghui, linux-arm-kernel, kvmarm

Hi Marc,

On Fri, Sep 18, 2026 at 01:22:17PM +0100, Marc Zyngier wrote:
> On Thu, 17 Sep 2026 16:49:25 +0100,
> Alexandru Elisei <alexandru.elisei@arm.com> wrote:
> > 
> > Tested on an Orion board and a x86 machine. On the x86 machine, there's was this
> > one thread called 'kvm-nx-lpage-re' which had affinity of the VCPUs, but I
> > couldn't figure out where it was created. I grep'ed for each part of the name in
> > the source node, no luck. I used strace, and couldn't see a clone call that
> > returned the corresponding pid. Any clues would be appreciated.
> 
> This is a KVM internal thread, from what I can tell. See
> arch/x86/kvm/mmu/mmu.c::kvm_mmu_start_lpage_recovery().

Indeed, it didn't cross my mind that it might KVM that is creating the thread.
Thanks for the hint, and to Joey too, he suggested it offline.

Alex

^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-09-21  8:59 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 15:49 [PATCH kvmtool 0/7] Fix --vcpu-affinity Alexandru Elisei
2026-09-17 15:49 ` [PATCH kvmtool 1/7] arm64: Pass the number of elements as the first argument to calloc() Alexandru Elisei
2026-09-17 15:49 ` [PATCH kvmtool 2/7] arm64: Consistently treat vcpu_affinity_cpuset as dynamically allocated Alexandru Elisei
2026-09-17 15:49 ` [PATCH kvmtool 3/7] arm64: Free the temporary cpumask in vcpu_affinity_parser() Alexandru Elisei
2026-09-17 15:49 ` [PATCH kvmtool 4/7] arm64/pmu: Consider kvmtool's affinity when searching for a PMU Alexandru Elisei
2026-09-17 15:49 ` [PATCH kvmtool 5/7] Correctly apply --vcpu-affinity to the VCPU threads Alexandru Elisei
2026-09-17 15:49 ` [RFC PATCH kvmtool 6/7] Introduce kvm_create_thread() Alexandru Elisei
2026-09-17 15:49 ` [RFC PATCH kvmtool 7/7] Don't apply --vcpu-affinity to threads spawned from VCPUs Alexandru Elisei
2026-09-18 12:22 ` [PATCH kvmtool 0/7] Fix --vcpu-affinity Marc Zyngier
2026-09-21  8:59   ` Alexandru Elisei

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox