From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D1EA939CCF3 for ; Thu, 6 Aug 2026 13:48:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786024145; cv=none; b=eDJTDL9aOuuYYh3mUg3cN9pJ5VQpqeAv9oiXq6vh9nveo1SprIV1lWDmoZuivg0eNYJmSm2CUCVLb+hyxEtQZV8odOYVguHIozVrs4YioIJCpmGCs+HEgmGuQusuoSxaWMXUs3ZcmvV79DSPOIMiBXHaM7Qm4RDBnBoBalteFbM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786024145; c=relaxed/simple; bh=zR4Buq8Cqbhx+BoaM3N9Hxez+9qWEGN61TfM+T+/GOE=; h=From:To:Cc:Subject:In-Reply-To:Date:Message-ID:References: MIME-version:Content-type; b=JuPwlAgP6E34piGyiPHxh4WXp5bMNli6kduaYRPVH/ormDKU7Zw3Pk0DbveoojlnCLyWNFjxvG/0hCm95tcTWxyHBfntYllfUIG3miKX/vMbdkIiZvFQ5ZsPOTtJ9mRvS63K1/cptDfv8bDJAozH3bd7N7I+ndMG2GHEpvNrpZE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=KZ+xoyFP; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="KZ+xoyFP" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2caea3f742bso31528195ad.0 for ; Thu, 06 Aug 2026 06:48:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786024135; x=1786628935; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :message-id:date:in-reply-to:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=QjVIN0UdWGoOCTancp6az5kLjHEYcF1w7gKX5HKdYSA=; b=KZ+xoyFPZc8FZIdY6ETtPzI8oTwr5YDZG/dGkp3gmH2IpSPWxmHiJ5t+jt/x2GElp2 myWLTYENr5Iw5E6/KMmDDjO0q+s7n3z419sQ4e1tE01MO2fUNjSSscqAhT9krsMxdaOm wuNM37tJ6EyIEZF0yYzPj/OqNPAWmBOUQsCm/RwFDyUiuhTOarA/YM7tDmO00vIKuZy0 1BaG86uIsMSUt4dhO+uoHIgXCNJIc9vmH6P3enEQIiWsXUl9UPV21+rsguEW0hB/J/VA 4tg/S6WytwXdaF216xNI6xX/pbp92Rql88fvNla7DFtKXVqhX+a6ir7TagxfLzCJQNN2 Rhlg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786024135; x=1786628935; h=content-transfer-encoding:content-type:mime-version:references :message-id:date:in-reply-to:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=QjVIN0UdWGoOCTancp6az5kLjHEYcF1w7gKX5HKdYSA=; b=GgSdXm7QiCZkqqgokMaw+R/YV+MG82YMFzHTARTvAmUH7UBYZufJ4ures9QgWwXl+1 itmFsHD4R0o5/em+OUiFC6E9lKARcBMvpSu9M+n0IZZh2rCjr6j31WuDRvKDryH5DdLr Go60G2plLcNl1Cg4hFUG07qeuByxOhdqIrOU0Gp1jYurm3QnIqPSt7lv4x3h0rPgNKE9 RQA079HkRqCN4TTFiR5stsPkGABfR55cSF92zl/b9lYqg1pV9dQgGFD0QQfvE/hmFycQ bJ9frWWRC6F7eDLOLm8moIYYsZETPpnws6LrIiGPu0LRisaaP8EuY+HaHcYZsHprVYUA TJWg== X-Forwarded-Encrypted: i=1; AHgh+Rqb3bD6pVJGVZC3/OnL5wnUeLOR7t6nvjQ+hOkQBJYjrog8eJVjx+9YWs5iANYB3sJ429LCbQyG8lhFox8=@vger.kernel.org X-Gm-Message-State: AOJu0Yz60MUFs+euV6AUOjkRVR0uC52ruIAg6Blr2V9HYHqamZfdb11P 4W2xEODQfDO1ih0f9UJE7x8eP3baId6dEplyX31eJroZe0nkpRQ4MX6K X-Gm-Gg: AR+sD12V7uu/xgZJofiZHWIUze87fC5Vxd/2KVpBDIyRs7VIV6LhE492xpoZDj2Fmih o9UQCkGU6PptBCsRApLlJRyDdC8xtb+pj45tZud+8AmdBWTQH7G8TdBA/x5A4WxlNiv6MGxfLWt GH8lsa/gUI7qWRL6DhOfbTbXNiHhav+KkUqgywqC3Wjeu/VTGOEnrB+beY8tHvZToH2Nmgu2feJ v89dG2pH+RjU4uEnsa6+BvMqXgEkZLRjGiVPM4fYbzauLc35I/jp6uJnKQBu5JUJ9uzHeATt52t JCEx5nBsSZhHDf85ab1NpccyIwyb7BVF3n55p+ZyPuvrGc0EWzB2BR2qV3TdFQ1bDaCBgQguN58 EHygYKmVWmdzJ8errYykfoxLRXR0JZg16oOvGyFSXF3xg2fRyFzi9s0+lu16s/L7HaSsBJfomvV LtkcIXXYlTpDTjsHfgYP9pTqOqUsttC0tRawsV6x87LdE8kDvZWYC3hsnf0oQ= X-Received: by 2002:a05:6a21:6013:b0:3c3:7ac4:dac0 with SMTP id adf61e73a8af0-3cb85df2790mr19356193637.13.1786024134598; Thu, 06 Aug 2026 06:48:54 -0700 (PDT) Received: from pve-server ([49.205.216.49]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-315ae6c1fd4sm6563410eec.14.2026.08.06.06.48.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 06 Aug 2026 06:48:53 -0700 (PDT) From: Ritesh Harjani (IBM) To: Amit Machhiwal Cc: Amit Machhiwal , linuxppc-dev@lists.ozlabs.org, Madhavan Srinivasan , Vaibhav Jain , Anushree Mathur , Paolo Bonzini , Nicholas Piggin , Michael Ellerman , "Christophe Leroy (CS GROUP)" , Jonathan Corbet , Shuah Khan , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org Subject: Re: [PATCH v6 0/4] KVM: PPC: Expose CPU compatibility modes for nested guests In-Reply-To: <20260806104222.7e435b28-8d-amachhiw@linux.ibm.com> Date: Thu, 06 Aug 2026 18:35:31 +0530 Message-ID: References: <20260804180705.59160-1-amachhiw@linux.ibm.com> <20260806104222.7e435b28-8d-amachhiw@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-version: 1.0 Content-type: text/plain; charset=utf-8 Content-Transfer-Encoding: 8bit Amit Machhiwal writes: > Hi Ritesh, > > Thanks for taking a look. Please find my response inline. > > On 2026/08/06 12:09 AM, Ritesh Harjani wrote: >> >> Hi Amit, >> >> Amit Machhiwal writes: >> >> > On POWER systems, newer processor generations can operate in compatibility >> > modes corresponding to earlier generations (e.g., a Power11 system running >> > in Power10 compatibility mode). In such cases, the effective CPU level >> > exposed to guests differs from the physical processor generation. >> > >> > This creates a problem for nested virtualization. When booting a nested KVM >> > guest (L2) inside a host KVM guest (L1) running in a compatibility mode, >> > userspace (e.g., QEMU) may derive the CPU model from the raw hardware PVR >> > and attempt to configure the nested guest accordingly. However, the L1 >> > partition is constrained by the compatibility level negotiated with the >> > hypervisor (L0), and requests exceeding that level are rejected, leading to >> > guest boot failures such as: >> > >> > KVM-NESTEDv2: couldn't set guest wide elements >> > >> > This series provides a mechanism for userspace to query the effective CPU >> > compatibility modes supported by the host, so it can select an appropriate >> > CPU model for nested guests. >> > >> > To achieve this, the series introduces a new KVM capability and ioctl >> > (KVM_CAP_PPC_COMPAT_CAPS / KVM_PPC_GET_COMPAT_CAPS) that expose the >> > compatibility modes supported by the host. >> > >> >> Sorry, but I am somehow not convinced on whether we need all of this >> machinary just to get these 3 bits of information, which we are >> returning today. >> >> Since KVM_CHECK_EXTENSION can already return an int, so why can't we use >> KVM_CAP_PPC_COMPAT_CAPS itself and return the bitmap of supported compat >> modes to the user? >> Say if the cap is not supported, we can return 0, otherwise we can >> return the bitmap of supported compat modes. This will easily allow us >> to use 31-bits which as I see would be hardly a problem in the near >> future. In the future if it grows - we can always use KVM_CAP_PPC_COMPAT_CAPS2. >> >> This should reduce the code complexity both in the kernel and >> userspace and we don't even need a new ioctl then. > > Thanks for the suggestion. I considered this approach Then we should have brought that up early on during the design discussion. But for the sake of discussion let's call this as approach-2. > but would like to > go with a dedicated ioctl for the following reasons: > > 1. Intended semantics: The KVM API documentation states: > > ..kvm defines extension identifiers and a facility to query > whether a particular extension identifier is available. If it is, a > set of ioctls is available for application use. > > [...] > > KVM defines many constants of the form KVM_CAP_*, each corresponding > to a set of functionality provided by one or more ioctls. Availability > of these capabilities can be checked with KVM_CHECK_EXTENSION. > > The intended role of KVM_CAP_* is to signal ioctl availability, not > to serve as a data retrieval mechanism itself. That's not entirely true. We do return data as part of check extension for e.g. for getting the SMT modes check KVM_CAP_PPC_SMT_POSSIBLE. > > You may take a look at KVM_CAP_PPC_GET_CPU_CHAR for instance. > > 2. Return type constraint: KVM_CHECK_EXTENSION returns a signed 32-bit int. The > capability bits are defined as (1ULL << 62), (1ULL << 61), and (1ULL << 60) — > 64-bit values that cannot fit in a 32-bit return. Renumbering them to small > integers would be a UAPI change and would lose alignment with the There is no _change_ in the UAPI so far. This is the patch which is defining that in the first place. > H_GUEST_CAP_* values from the hypervisor ABI. No please. Those are 2 different ABIs and there is no need to set a hard dependency among the two. > > 3. Extensibility: The struct-based approach with the size field provides clean > forward and backward ABI versioning via copy_struct_from/to_user(), without > needing a KVM_CAP_PPC_COMPAT_CAPS2 in the future. > This only make sense if we really have a usecase already in mind which you are planning to extend it for. Otherwise, IMO, this is a lot of machinary and I think we should consider the simpler approach. IMO - I think approach-2 is a much simpler for this usecase. I don't see any valid reason on why we should not do that instead. We don't need an extra ioctl and all the struct machinary along with that just for returning a bitmask. The existing check extension ioctl can be easily used for this purpose. Would it be possible for you to give, approach-2 a try? Do you see any geniunine roadblock or limitation with that? -ritesh