From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 60DD2274641; Fri, 7 Aug 2026 17:25:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786123518; cv=none; b=lA2qFDFSe9xBcYnoNsjp0BMVceWt26iAX2CpjI9NKYGp2ntYUzGziOYhQC/TuwsqnY9SZQDBlwka18LdJ3N8vyw+Vh/3WYF4ePSDvoao2I4Ajq6dRtqjF71280zLKWAuULJXPxrIHE0rMAurROXRvz2GhUS0yaFl6+uBPXsnDPw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786123518; c=relaxed/simple; bh=hVPjohWCWcjBwKb9wuXF3Mm1CkRm974eDc4eHr9djKo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Pe//kxvyWvHLFt/g2x2SOAqqbvUC03pb07Rjr96bkGRSCQiADyV7G8o24s00OenP/fJFotAFQuT2L5I6z0THQwrR7E9f6RkM9imt6NIKu+20UZKEINXMBdbwhsVD7/SCIKw5XjrKPgqoTo72OjE9iTFuDzF8cu0a60BOBrcUQ9E= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=DZW1L+n6; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="DZW1L+n6" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 677GlsjJ1937437; Fri, 7 Aug 2026 17:24:53 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=POGGW7D4HqvqqWOHZ qJMmbMxinU2dGjKXW6RNdKVNs0=; b=DZW1L+n6W0Pkd/cgAK9ILfFFkyOdNwT/D YbuuTqCn8Vq+ifEAOUUMvKlz5QH87vGhdkGYq5AOk2NgeiCklJiEAkO8/Vv4bQgR yGN/XH/UZby+caljAxg+X3GuqS+rAREl7p/awAae6o/GqNKue1ILUG7zCsTuG755 Q7IyUBxcD8b0tkJ4hkQfsG7qDmsB3s1XZvI+nPcuvjyxeTErTQ2gCQY3obUTTJIX VySApNsHr6cN6xwN47+JJ7TYFFOF6WlJ3Pl+xMm3ojz75dF9RmWCJQmrR3ojZ+Ri BlStSu/Ntw799vVgAonT7xNmQXQwwKLbDcvQ88wc/5qM7IjS+neEg== Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4fvy024wg1-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 07 Aug 2026 17:24:53 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 677HBGHO009346; Fri, 7 Aug 2026 17:24:52 GMT Received: from smtprelay04.fra02v.mail.ibm.com ([9.218.2.228]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4fswbgrfa2-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 07 Aug 2026 17:24:52 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (smtpav07.fra02v.mail.ibm.com [10.20.54.106]) by smtprelay04.fra02v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 677HOmS430278222 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 7 Aug 2026 17:24:48 GMT Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 5460E2004D; Fri, 7 Aug 2026 17:24:48 +0000 (GMT) Received: from smtpav07.fra02v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 974C720043; Fri, 7 Aug 2026 17:24:44 +0000 (GMT) Received: from localhost.localdomain (unknown [9.124.216.72]) by smtpav07.fra02v.mail.ibm.com (Postfix) with ESMTP; Fri, 7 Aug 2026 17:24:44 +0000 (GMT) From: Amit Machhiwal To: linuxppc-dev@lists.ozlabs.org, Madhavan Srinivasan Cc: Vaibhav Jain , Amit Machhiwal , Anushree Mathur , Paolo Bonzini , Nicholas Piggin , Michael Ellerman , "Christophe Leroy (CS GROUP)" , Jonathan Corbet , Shuah Khan , Ritesh Harjani , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, Gautam Menghani Subject: [PATCH v8 1/4] KVM: PPC: Introduce KVM_CAP_PPC_COMPAT_CAPS and wire up ioctl Date: Fri, 7 Aug 2026 22:54:30 +0530 Message-ID: <20260807172433.82045-2-amachhiw@linux.ibm.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260807172433.82045-1-amachhiw@linux.ibm.com> References: <20260807172433.82045-1-amachhiw@linux.ibm.com> Precedence: bulk X-Mailing-List: kvm@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-TM-AS-GCONF: 00 X-Proofpoint-Reinject: loops=2 maxloops=12 X-Proofpoint-ORIG-GUID: mPSN9XVJfx015M7NbOndJ8BeRBw-Fr06 X-Authority-Analysis: v=2.4 cv=e5k2j6p/ c=1 sm=1 tr=0 ts=6a7614e5 cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=VnNF1IyMAAAA:8 a=ETYFO_97Qo2pusGNp_0A:9 X-Proofpoint-GUID: kurpvY7eqfTKg5IFse8TXSsPxJ94pzWm X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODA3MDEzNSBTYWx0ZWRfXzxapaznAKGiw DPux05OMCDc8NplWRbsqthEbrp3uwexV2KL5LngcDzPooo74XcLyEZ4YrtvQOZadeGFmf1diS0m LWIVC4PgUyicTWH+2qA8g/sQd4Ba/IVPk8yZhCD2dtKkILFelE3lP6aF54FHzvuswFRamyhiGer SbNtKN5FCq9PlucxKr3o9Q/dUTs+7F0Apa8oGZTPRKxp5xM7QHeRhQR3jEvw/WvCHH7jx5z9JIW amK58RgnTdHG72zyDGJgWhljKMMHAvQ/8Y93piWwUPLVCFQwu4fS088ulSP1EEoUHoM4YlTSq66 8TmDxRKgnae/efD4/0v0E16FisRoemE6ChKWFFJ7C7RfdO9gKMSCkq4GLfNVv4J14Svvctl3vVe 2XFyXkRd8xA9QVhZmQ0du/BRH0KxBLJg/7tr8RRLtRJDu7HcPnPK14Kso7uteZeHgjhURPw9BoX TxfUQkpGlcDyjiqy3rQ== X-Proofpoint-Spam-Info: AW1haW4tMjYwODA3MDEzNSBTYWx0ZWRfXxlfrTI0jTcQV FtANJ220pZF65eownEh0yQhwAe8cL7U6RL0jj1ddcBf4UO+9+xmlSiBm+/+PzwPGdWIHTel9wYk 5zgte/tL7ojh7Ml46t+pInouASVpQi4= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-07_03,2026-08-07_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 impostorscore=0 priorityscore=1501 adultscore=0 phishscore=0 clxscore=1015 malwarescore=0 spamscore=0 lowpriorityscore=0 bulkscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608070135 Introduce a new capability and ioctl to expose CPU compatibility modes supported by the host processor for nested guests. On IBM POWER systems, newer processor generations (N) can operate in compatibility modes corresponding to earlier generations, like (N-1) and (N-2). This is particularly relevant for nested virtualization, where nested KVM guests may need to run with a specific processor compatibility level. Introduce KVM_CAP_PPC_COMPAT_CAPS capability and the corresponding KVM_PPC_GET_COMPAT_CAPS vm ioctl. The ioctl returns a bitmap describing the compatibility modes supported by the host in respective bit numbers, allowing userspace (e.g., QEMU) to select an appropriate compatibility level when configuring nested KVM guests. The ioctl handling is added in kvm_arch_vm_ioctl() and retrieves host CPU compatibility capabilities via a PowerPC-specific backend implementation when available. The struct kvm_ppc_compat_caps places the 'size' field first so it can be read alone via get_user() before copy_struct_from_user() is called, avoiding pointer arithmetic to locate the size field. The ioctl is defined using _IO so the ioctl number remains stable even if the struct grows in future versions. It uses copy_struct_from_user() and copy_struct_to_user() to provide forward- and backward-compatible extensibility: older userspace passing a smaller struct to a newer kernel gets zero-padded trailing fields. Newer userspace passing a larger struct to an older kernel (usize > ksize) succeeds if trailing bytes are zero (the kernel reports back min(usize, ksize) as the filled size); if trailing bytes are non-zero, the kernel writes back ksize into host_caps.size and returns -E2BIG so userspace can retry with the correct size. KVM_PPC_COMPAT_CAPS_SIZE_VER0 is defined as a frozen integer constant (24) marking the size of the initial struct version, used as the minimum floor for size field validation, similar to other versioned struct interfaces in the kernel. The 'flags' field is reserved for future use. The kernel rejects any call where flags is non-zero with -EINVAL, preventing garbage values from being baked into ABI permanently. The ioctl returns appropriate error codes: E2BIG if usize exceeds PAGE_SIZE, or if new userspace provides a larger struct with non-zero trailing bytes (with ksize written back into host_caps.size for the retry); EINVAL for an invalid size or non-zero reserved fields; EFAULT for failed copy operations; and ENOTTY if the backend doesn't implement get_compat_caps. Suggested-by: Vaibhav Jain Tested-by: Gautam Menghani Reviewed-by: Gautam Menghani Tested-by: Anushree Mathur Signed-off-by: Amit Machhiwal --- Changes in this version: - Add PAGE_SIZE guard after get_user() to bound the check_zeroed_user() scan in the usize > ksize path [Ritesh] - Drop manual usize > sizeof(host_caps) pre-check; delegate entirely to copy_struct_from_user() which succeeds on zero trailing bytes and returns -E2BIG only on non-zero trailing bytes; handle -E2BIG with ksize writeback and -EFAULT escalation if put_user() fails [Ritesh] - Fix host_caps.size on success path: use min_t(u64, usize, sizeof(host_caps)) so new userspace with zero trailing bytes gets back the number of bytes the kernel actually populated, not usize [Ritesh] arch/powerpc/include/asm/kvm_ppc.h | 1 + arch/powerpc/include/uapi/asm/kvm.h | 8 +++ arch/powerpc/kvm/powerpc.c | 78 +++++++++++++++++++++++++++++ include/uapi/linux/kvm.h | 3 ++ 4 files changed, 90 insertions(+) diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h index 0953f2daa466..169ea6a7fbad 100644 --- a/arch/powerpc/include/asm/kvm_ppc.h +++ b/arch/powerpc/include/asm/kvm_ppc.h @@ -319,6 +319,7 @@ struct kvmppc_ops { bool (*hash_v3_possible)(void); int (*create_vm_debugfs)(struct kvm *kvm); int (*create_vcpu_debugfs)(struct kvm_vcpu *vcpu, struct dentry *debugfs_dentry); + int (*get_compat_caps)(struct kvm_ppc_compat_caps *host_caps); }; extern struct kvmppc_ops *kvmppc_hv_ops; diff --git a/arch/powerpc/include/uapi/asm/kvm.h b/arch/powerpc/include/uapi/asm/kvm.h index 077c5437f521..19e53d5ae540 100644 --- a/arch/powerpc/include/uapi/asm/kvm.h +++ b/arch/powerpc/include/uapi/asm/kvm.h @@ -437,6 +437,14 @@ struct kvm_ppc_cpu_char { __u64 behaviour_mask; /* valid bits in behaviour */ }; +/* For KVM_PPC_GET_COMPAT_CAPS */ +struct kvm_ppc_compat_caps { + __u64 size; /* Size of this structure */ + __u64 flags; /* Reserved for future use */ + __u64 compat_capabilities; /* Capabilities supported by the host */ +}; +#define KVM_PPC_COMPAT_CAPS_SIZE_VER0 24 /* sizeof first published struct */ + /* * Values for character and character_mask. * These are identical to the values used by H_GET_CPU_CHARACTERISTICS. diff --git a/arch/powerpc/kvm/powerpc.c b/arch/powerpc/kvm/powerpc.c index b6b83fe3233f..14a661a88d4e 100644 --- a/arch/powerpc/kvm/powerpc.c +++ b/arch/powerpc/kvm/powerpc.c @@ -703,6 +703,13 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext) } } break; +#if defined(CONFIG_KVM_BOOK3S_HV_POSSIBLE) + case KVM_CAP_PPC_COMPAT_CAPS: + r = 0; + if (hv_enabled && kvmhv_on_pseries()) + r = 1; + break; +#endif /* CONFIG_KVM_BOOK3S_HV_POSSIBLE */ default: r = 0; break; @@ -2469,6 +2476,77 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) r = kvm->arch.kvm_ops->svm_off(kvm); break; } + case KVM_PPC_GET_COMPAT_CAPS: { + struct kvm_ppc_compat_caps host_caps = {}; + u64 usize; + + /* + * Read the size field first to drive copy_struct_from_user. + * size must be the first field of the struct. + */ + r = -EFAULT; + if (get_user(usize, (__u64 __user *)argp)) + goto out; + + r = -E2BIG; + if (unlikely(usize > PAGE_SIZE)) + goto out; + + /* + * Enforce a minimum: reject buffers smaller than the initial + * struct version (VER0). This allows old userspace compiled + * against the original struct to still work on a newer kernel + * that has grown the struct with appended fields. + */ + r = -EINVAL; + if (usize < KVM_PPC_COMPAT_CAPS_SIZE_VER0) + goto out; + + /* + * copy_struct_from_user() handles forward/backward compat: + * usize == ksize: verbatim copy + * usize < ksize: zero-pad trailing (old userspace, new kernel) + * usize > ksize: succeed iff trailing bytes are zero, else -E2BIG + */ + r = copy_struct_from_user(&host_caps, sizeof(host_caps), + argp, usize); + if (r) { + /* + * New userspace with a larger struct called an older + * kernel. Write back ksize in host_caps.size so + * userspace knows which older struct to retry with, + * then fail with -E2BIG. + */ + if (r == -E2BIG) + if (put_user((__u64)sizeof(host_caps), + (__u64 __user *)argp)) + r = -EFAULT; + goto out; + } + + /* Reserved fields must be zero */ + r = -EINVAL; + if (host_caps.flags) + goto out; + + r = -ENOTTY; + if (!kvm->arch.kvm_ops->get_compat_caps) + goto out; + + r = kvm->arch.kvm_ops->get_compat_caps(&host_caps); + if (r) + goto out; + + /* + * Report the number of bytes actually populated by the kernel, + * not usize: if new userspace passed a larger struct with zero + * trailing bytes, we only filled sizeof(host_caps) bytes. + */ + host_caps.size = min_t(u64, usize, sizeof(host_caps)); + r = copy_struct_to_user(argp, usize, &host_caps, + sizeof(host_caps), NULL); + break; + } default: { struct kvm *kvm = filp->private_data; r = kvm->arch.kvm_ops->arch_vm_ioctl(filp, ioctl, arg); diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 419011097fa8..70e36e6a0ad4 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -997,6 +997,7 @@ struct kvm_enable_cap { #define KVM_CAP_S390_KEYOP 247 #define KVM_CAP_S390_VSIE_ESAMODE 248 #define KVM_CAP_S390_HPAGE_2G 249 +#define KVM_CAP_PPC_COMPAT_CAPS 250 struct kvm_irq_routing_irqchip { __u32 irqchip; @@ -1341,6 +1342,8 @@ struct kvm_s390_keyop { /* Available with KVM_CAP_COUNTER_OFFSET */ #define KVM_ARM_SET_COUNTER_OFFSET _IOW(KVMIO, 0xb5, struct kvm_arm_counter_offset) #define KVM_ARM_GET_REG_WRITABLE_MASKS _IOR(KVMIO, 0xb6, struct reg_mask_range) +/* Available with KVM_CAP_PPC_COMPAT_CAPS */ +#define KVM_PPC_GET_COMPAT_CAPS _IO(KVMIO, 0xb8) /* ioctl for vm fd */ #define KVM_CREATE_DEVICE _IOWR(KVMIO, 0xe0, struct kvm_create_device) -- 2.50.1 (Apple Git-155)