From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f43.google.com (mail-wm1-f43.google.com [209.85.128.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5554F39A4DC for ; Wed, 5 Aug 2026 14:57:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.43 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785941832; cv=none; b=H8PW1DVra7cAPQf57tbKfAHj6vQijUXbMZGoheGsxsrp247DQBAk7QT5eN0504fD8J03u21dwoFb+8dRC8TbgoZ7tepNCipZyWTOGnlyg9RtUlykQs00DTmVyplX151ygZK5F+zZ3LLRMj/zowrf614evzazN2PLgWK7bDeBdw4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785941832; c=relaxed/simple; bh=cqzy97cVVOO/aObREG1arsE54P9gBkOM8bNx5Yj95qM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=e2ccgL4+fWH8N6ONE4ilagC9XFD1Hl9CswqAh4SikZMut3xmI3nSU4c+XGspvY1PEwRtfFVSdKIf7S+Dzc8TrbB4xXVqFOK7/w7DX8URVDFmzXaSswxs4teLmRt4i5+is7SqQf48l4SSW8QrDZY5XGrwsB34cnqOiVgsYafNg1k= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Y9Xec/zD; arc=none smtp.client-ip=209.85.128.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Y9Xec/zD" Received: by mail-wm1-f43.google.com with SMTP id 5b1f17b1804b1-4994d67d260so51095e9.1 for ; Wed, 05 Aug 2026 07:57:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1785941827; x=1786546627; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=NNzat26Y19ZrKtjltXLxT2fvXKCHP/AoQJ9WhsEOxAU=; b=Y9Xec/zDQyyLDWBEeWY25AJ+K2D3JhNFXs0Opub8GrDDOBvzBRKS0PkDx5WzHwlAiy niFlFnMbzptdEwPPCFJx9+EXY+m46xZ+GTrPSsufFGqRYMQT5TzCjprEUAvC9kTSnPE9 vaoAxCM03M4pH+MEBo2M7aU20kg7InI/Rxno/WUvy1CBJ18HyDJSvth45PiyV9UgypmR nWyqA20crGp7pEOzZh/fK8uLPal+vTFiPQUiYLuK44+7wZZikiFVAYnu62txngOtDzuO GSMQ7GqhBlVnNVHhTyPH9fpI+2J3BRHZpa4OVokzDq8pUCMQ/zdUNv6Cd/k+Ey2P1NNd Jncw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785941827; x=1786546627; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=NNzat26Y19ZrKtjltXLxT2fvXKCHP/AoQJ9WhsEOxAU=; b=lDn5GorDGheiFUU4QCjSe9QtBh1GqPiPxOCpHAp6G5zhfwRDOyIvYZr26Lfr0f9aCs IXPH2jhHs2VE3DcnvTJtFwT48nO5j8enX9rMRHfEBMJumq0tyRH/zZNM9ltJtihpv9vV XQKi24IUePbw7i0OY6gQTOq/UCybrrq2aZS46YhyRnW2nA3M9jb4YUqqWepvw11FKLnv hUQwktPvvq5oSqB5BJMTGL4K19oz/cLywjknW0Q6HA1hRgU61kBqFK1SKTyKz6C30o0w 1TDr+QuDsNHhMSbyqdrfj0SwZbbcWkfrSAu7hf/7Bo/2q3O8FLY9FN6JKJxO6EU5vtVT jiYw== X-Forwarded-Encrypted: i=1; AHgh+RoejQygMCKrDaX5w61Qc/tKcEWAY82U6qox9QGx14ftLod8AzWYvcqOFEok3/arLsSwYau0+o1eX/Dws2E=@vger.kernel.org X-Gm-Message-State: AOJu0YydgQbtnvP16bqWua7g3QuIUgXJkXJpOG1YjSqK/2CBUbxqDVtJ dLp5Pdrkao2gr+OdYBf9NPCoecXce8/I7dCT5O81PjbzYrN0cuK1Ss7tFQ38FxN7pA== X-Gm-Gg: AR+sD11DsJKRKUAHaWDS4R9DidYIVrJ4GyRhIcMETGY7hq77LZWDyYQE5jPMMmGOSdY LIvgi7szFl1XXI0CY7aLxNiR8zdDMjBtGu4oUMwyXFUHV1pPPzpphz6fKl710vAJMo/cSbej1el Szh1KbjgXbxgcBoQ6ehGW60JKsl+gTend51EYFPQRcOm26ZQ+H4s/g/5rCn8DHKDuRiu6JWs2FX Hj28fE8Hd1o0VUCDyNaUgMgQqTw5YiceEld4fvazJwKhwoZu78qN504TOE5a5+50MzduXIEB3Vk fpB9cAHKm2/bI1PPhFdXw0lnSO94XH5sAhUgzaS8N9LBAcN9V4Lo8dLbvrypEoh3JiULtp3yHzI i3aueCYbZBJGd/lVjQepo7cO4ae0jKymrNu073HrdIsO1OFkdFei8m2D2/huAubd3gQFnFWgVjt SE7aezQMIFjfoFiItregoTxjS71Q8/BLzTApaldTxjeIMNnt2Y2IAo6pLjrPb31Hin301udvBp1 IjLLWQBTpjDNB3/7YUrwooN8vvWAJvT X-Received: by 2002:a05:600c:6554:b0:493:ae5f:d29f with SMTP id 5b1f17b1804b1-4994f0fb66dmr834065e9.3.1785941827064; Wed, 05 Aug 2026 07:57:07 -0700 (PDT) Received: from google.com (63.235.189.35.bc.googleusercontent.com. [35.189.235.63]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-47fec232d74sm9786235f8f.17.2026.08.05.07.57.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 05 Aug 2026 07:57:06 -0700 (PDT) Date: Wed, 5 Aug 2026 14:57:02 +0000 From: Sebastian Ene To: Mostafa Saleh Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, kvmarm@lists.linux.dev, iommu@lists.linux.dev, catalin.marinas@arm.com, will@kernel.org, maz@kernel.org, oliver.upton@linux.dev, joey.gouly@arm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, joro@8bytes.org, jgg@ziepe.ca, mark.rutland@arm.com, qperret@google.com, tabba@google.com, vdonnefort@google.com, keirf@google.com Subject: Re: [PATCH v7 02/24] KVM: arm64: Donate MMIO to the hypervisor Message-ID: References: <20260715115906.2664882-1-smostafa@google.com> <20260715115906.2664882-3-smostafa@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Wed, Aug 05, 2026 at 01:42:10PM +0000, Mostafa Saleh wrote: > On Wed, Aug 05, 2026 at 12:47:03PM +0000, Sebastian Ene wrote: > > On Wed, Jul 15, 2026 at 11:58:43AM +0000, Mostafa Saleh wrote: > > > Add a function to donate MMIO to the hypervisor so IOMMU hypervisor > > > drivers can protect and access the MMIO of IOMMUs. > > > > > > As donating MMIO is very rare, and we don’t need to encode the full > > > state, it’s reasonable to have a separate function to do this. > > > It will init the host s2 page table with an invalid leaf with the owner ID > > > to prevent the host from mapping the page on faults. > > > > > > Also, prevent kvm_pgtable_stage2_unmap() from removing owner ID from > > > stage-2 PTEs, as this can be triggered from recycle logic under memory > > > pressure. There is no code relying on this, as all ownership changes is > > > done via kvm_pgtable_stage2_set_owner() > > > > > > For the error path in IOMMU drivers, add a function to donate MMIO > > > back from hyp to host. However, that leaks the hypervisor virtual > > > address range which should be acceptable as this is quite rare and > > > it matches the behaviour of fix_map/block. > > > > > > Signed-off-by: Mostafa Saleh > > > --- > > > arch/arm64/kvm/hyp/include/nvhe/mem_protect.h | 7 ++ > > > arch/arm64/kvm/hyp/nvhe/mem_protect.c | 91 ++++++++++++++++++- > > > arch/arm64/kvm/hyp/pgtable.c | 11 +-- > > > 3 files changed, 102 insertions(+), 7 deletions(-) > > > > > > diff --git a/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h b/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h > > > index 29935c7da1de..51b0eb3844a9 100644 > > > --- a/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h > > > +++ b/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h > > > @@ -36,6 +36,13 @@ int __pkvm_guest_share_host(struct pkvm_hyp_vcpu *vcpu, u64 gfn); > > > int __pkvm_guest_unshare_host(struct pkvm_hyp_vcpu *vcpu, u64 gfn); > > > int __pkvm_host_unshare_hyp(u64 pfn); > > > int __pkvm_host_donate_hyp(u64 pfn, u64 nr_pages); > > > +/* > > > + * Donate MMIO range to the hypervisor, it will be mapped in the hypervisor's > > > + * private range and unmapped from the host stage-2. > > > + */ > > > +int __pkvm_host_donate_hyp_mmio(phys_addr_t addr, size_t size, unsigned long *haddr); > > > +/* Remaps MMIO range in the host, typically used in error path. */ > > > +int __pkvm_hyp_donate_host_mmio(phys_addr_t addr, size_t size); > > > int __pkvm_hyp_donate_host(u64 pfn, u64 nr_pages); > > > int __pkvm_host_share_ffa(u64 pfn, u64 nr_pages); > > > int __pkvm_host_unshare_ffa(u64 pfn, u64 nr_pages); > > > diff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvhe/mem_protect.c > > > index 4e329e39a695..d803b3dd4cb4 100644 > > > --- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c > > > +++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c > > > @@ -378,7 +378,11 @@ static int host_stage2_unmap_dev_all(void) > > > u64 addr = 0; > > > int i, ret; > > > > > > - /* Unmap all non-memory regions to recycle the pages */ > > > + /* > > > + * Unmap all non-memory regions to recycle the pages. > > > + * That relies on kvm_pgtable_stage2_unmap() not clearing > > > + * counted PTEs which include hypervisor MMIO. > > > + */ > > > for (i = 0; i < hyp_memblock_nr; i++, addr = reg->base + reg->size) { > > > reg = &hyp_memory[i]; > > > ret = kvm_pgtable_stage2_unmap(pgt, addr, reg->base - addr); > > > @@ -1119,6 +1123,91 @@ int __pkvm_host_donate_hyp(u64 pfn, u64 nr_pages) > > > return ret; > > > } > > > > > > +int __pkvm_host_donate_hyp_mmio(phys_addr_t addr, size_t size, unsigned long *haddr) > > > +{ > > > + kvm_pte_t pte; > > > + u64 offset; > > > + int ret; > > > + > > > + /* Only before de-privilege. */ > > > + if (static_branch_unlikely(&kvm_protected_mode_initialized)) > > > + return -EPERM; > > > + > > > + if (!PAGE_ALIGNED(addr | size) || > > > + !pfn_range_is_valid(hyp_phys_to_pfn(addr), size >> PAGE_SHIFT)) > > > + return -EINVAL; > > > + > > > + ret = __pkvm_create_private_mapping(addr, size, PAGE_HYP_DEVICE, haddr); > > > > > + if (ret) > > > + return ret; > > > + > > > + host_lock_component(); > > > + for (offset = 0; offset < size; offset += PAGE_SIZE) { > > > + if (addr_is_memory(addr + offset)) { > > > + ret = -EINVAL; > > > + goto unlock; > > > > If this fails we are left with the mapping inside the hyp because the unlock > > doesn't destroy the private mapping. Is this intended ? > > > > Yes, there is no way to remove a private mapping at the moment, all > the callers to __pkvm_create_private_mapping() will leak it on failure. > > > > + } > > > + ret = kvm_pgtable_get_leaf(&host_mmu.pgt, addr + offset, &pte, NULL); > > > + if (ret) > > > + goto unlock; > > > + if (pte && !kvm_pte_valid(pte)) { > > > + ret = -EPERM; > > > + goto unlock; > > > + } > > > + } > > > + /* > > > + * We set HYP as the owner of the MMIO pages in the host stage-2, for: > > > + * - host aborts: host_stage2_adjust_range() would fail for invalid non zero PTEs. > > > + * - recycle under memory pressure: host_stage2_unmap_dev_all() would call > > > + * kvm_pgtable_stage2_unmap() which will not clear non zero invalid ptes (counted). > > > + * - other MMIO donation: Would fail as we check that the PTE is valid or empty. > > > + */ > > > + ret = host_stage2_try(kvm_pgtable_stage2_annotate, &host_mmu.pgt, > > > + addr, size, &host_s2_pool, > > > + KVM_HOST_INVALID_PTE_TYPE_DONATION, > > > + FIELD_PREP(KVM_HOST_DONATION_PTE_OWNER_MASK, PKVM_ID_HYP)); > > > +unlock: > > > + host_unlock_component(); > > > + return ret; > > > +} > > > + > > > +int __pkvm_hyp_donate_host_mmio(phys_addr_t addr, size_t size) > > > > This function seems to only update the host stage-2 annotation but it > > doesn't destroy the hyp mapping. > > Yes, as mentioned above there is no way to destroy it, and this was > not desgined for frequent use. Typicaly, __pkvm_host_donate_hyp_mmio() > is called at boot per area/device. And __pkvm_hyp_donate_host_mmio() > is only used for failures. > Yes, I saw that you only allow the call during the init pKVM calls, however it seems a bit fragile. > > > > I was looking to make use of this patch in an upcoming posting for the > > v2 ITS hardening but in my case I don't need the private VA range > > creation. > > I believe if you need to map MMIO, private range is the right way to > do it as the linear map was mainly designed around system memory. Is there a reason behind not having MMIO as part of the linear map in the hypervisor ? I would like to avoid holding the hva around and just do a simple addition to resolve the hyp_va. > > Thanks, > Mostafa > Thanks, Sebastian