From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id E1A4DC982FA for ; Tue, 22 Sep 2026 13:48:46 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Type:Cc:To:From: Subject:Message-ID:References:Mime-Version:In-Reply-To:Date:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=LSo/cDlJ7ohD2G3UI5KuRXiO2Rxr2FOcqemAL7pGINk=; b=iZiuxhOXmzaR6nVPPgzJwYQDT7 L9EE/v6qlarDXGBjC4F6JFCGFqeCxu48UqWSyi2CGkmyhcUEPSsTQ7VSXGtXsmbqnpr7SIcNL3U2A X+FhAkyx53sENc+OE6jqUVbjnEo6rwPIaPvRfLrNHXhcwfnIsNz9I5gctQ7n/r/pZH6wo6S8400LQ p24AcJtpKv+BK7vS/KxNdlS48UZZV/qqa21bwUk3qucEDaRQO3S2ubNYi3dE+k3lbdr1e37Fju55D YhI99sGnYKfORhuti4IhvbM86f2DDkZlP1GRWy9HuvlZO2A13Z45nhfe6xaTeOy3KavXGdF+Bymre pKnCawoQ==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1x90Ji-00000005SiT-3Ean; Tue, 22 Sep 2026 13:13:50 +0000 Received: from mail-wm1-x345.google.com ([2a00:1450:4864:20::345]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1x90JP-00000005SQb-3bNc for linux-arm-kernel@lists.infradead.org; Tue, 22 Sep 2026 13:13:33 +0000 Received: by mail-wm1-x345.google.com with SMTP id 5b1f17b1804b1-49953abe51fso26292695e9.1 for ; Tue, 22 Sep 2026 06:13:31 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790082810; x=1790687610; darn=lists.infradead.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=LSo/cDlJ7ohD2G3UI5KuRXiO2Rxr2FOcqemAL7pGINk=; b=CBYKGFwM0QvXAnWfkrIpPjLXU81F8wHleMyXQiXgA62WzLY/Wyne8IOKNQOfd0oU4C 1EYwCZL6fjJrQ63ksIlkANkLAcOqlIdxgvUIat4NCCMOvVnNadLMiU8qGjYHgApAIksD HYkorQqYNK+2RXbT0uM66DEz+MfE2+IUFGRwbJ4wjfBQRbJU6+7w1ngiHTX2WJ144Lkw uZLrqT9wyU69jlzPEG+vRwXV0YE7kRM6Xh0blv7pCMfo6iWhsYC86Mk2W3AvnbfwJvUB +638H/Tn4mUodcTsJyrHo+jblS/F7IniB1b30AmNrbk1UYGfpFP7hScueEe3aYgGaUwT 4LJw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790082810; x=1790687610; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=LSo/cDlJ7ohD2G3UI5KuRXiO2Rxr2FOcqemAL7pGINk=; b=bfMR8NOfoeWXROOMPiUbc4erUgfdTZP9GIrquqDoEM4hPKGZtD6m/s6Jsj24ubSJkF 58GdfdwPbTcyjWdSDXi8Yf6xG+UZYdDElqiwuCXLWj2CJFfXbEiC9l4lLMfWVo67NbjI JnQ0y9Zal+5KuxwhgnWz/7i7wQS4M8pRN0H/gMgO+jJf45WlBtKs2kTr2N7BaTMeSINl AvV+yW6G9+nsSWM3mI/8XW0TClha7EmHapJ5TWVG8GYLY2qMMjWlH0fXfQ33Gyf8ojH7 lS5ZvrpMiR1PZqrkclyhGp6yi19WR0I3VKscP8ly53rKCKokdUN0kKaByH7dtLBHO8Z2 jRiw== X-Gm-Message-State: AFuF++nvSjg2NPH9pj8Sfzi2GgWNupGuK/79+J3/59r3Hukm66KiTUOH JlB9t3FiddzDX81+QLWfllFJEajvrmp8Aqgm3rwCdciUwCoCHIm9COR7VHiP2GGV7jqC9UHGJ2l gBpwNeev6+/F1LHdnytuWhoAUGiWgAoXQUDb4tyF0HMlYX4o0rh8H2RMWc8mvI69LzC6YeR5hf3 EOXPZKHqn5ewQi3KaNr4bIzBm6CykBEuBnlIzo/x7IyvAxhkEQkd2dyhYkCH803BPrCw== X-Received: from wmdd15.prod.google.com ([2002:a05:600c:a20f:b0:49e:78cc:d2e]) (user=smostafa job=prod-delivery.src-stubby-dispatcher) by 2002:a05:600c:3f0b:b0:49d:1df6:2592 with SMTP id 5b1f17b1804b1-49fc5750c61mr171840065e9.21.1790082809523; Tue, 22 Sep 2026 06:13:29 -0700 (PDT) Date: Tue, 22 Sep 2026 13:12:55 +0000 In-Reply-To: <20260922131259.2975334-1-smostafa@google.com> Mime-Version: 1.0 References: <20260922131259.2975334-1-smostafa@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260922131259.2975334-23-smostafa@google.com> Subject: [PATCH v8 22/25] iommu/arm-smmu-v3-kvm: Shadow the CPU stage-2 page table From: Mostafa Saleh To: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, kvmarm@lists.linux.dev, iommu@lists.linux.dev Cc: catalin.marinas@arm.com, will@kernel.org, maz@kernel.org, oliver.upton@linux.dev, joey.gouly@arm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, joro@8bytes.org, jgg@ziepe.ca, mark.rutland@arm.com, qperret@google.com, tabba@google.com, vdonnefort@google.com, sebastianene@google.com, keirf@google.com, Mostafa Saleh Content-Type: text/plain; charset="UTF-8" X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.9.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20260922_061331_968036_25614A8A X-CRM114-Status: GOOD ( 29.49 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org The hypervisor calls back into the driver on every change to the host stage-2 page table. Mirror those changes into an identity mapped stage-2 for the SMMUv3, That would be attached to all active SIDs. Difference between memory and MMIO handling: - Memory is always mapped with PAGE_SIZE. io-pgtable-arm no longer supports split_blk_unmap, so a block cannot be broken into a table once it is mapped, and pages are donated back and forth at PAGE_SIZE granularity. The page table pool is sized to cover all of memory at that granularity. - MMIO is mapped with the largest block that fits, as it is assumed to cover the whole IAS that is not memory while pKVM only reserves 1G for the page table. MMIO is never donated at runtime, so it is never unmapped and never needs a block to be split. The TLB maintenance ops are left as stubs, they are implemented in the next patch. Signed-off-by: Mostafa Saleh --- .../iommu/arm/arm-smmu-v3/pkvm/arm-smmu-v3.c | 153 +++++++++++++++++- 1 file changed, 152 insertions(+), 1 deletion(-) diff --git a/drivers/iommu/arm/arm-smmu-v3/pkvm/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/pkvm/arm-smmu-v3.c index 6e023e968ed3..6f0ea3a4e48d 100644 --- a/drivers/iommu/arm/arm-smmu-v3/pkvm/arm-smmu-v3.c +++ b/drivers/iommu/arm/arm-smmu-v3/pkvm/arm-smmu-v3.c @@ -14,6 +14,9 @@ #include "../arm-smmu-v3.h" #include "../arm-smmu-v3-common-lib.h" +#include +#include "../../../io-pgtable-arm.h" + size_t __ro_after_init kvm_hyp_arm_smmu_v3_count; struct hyp_arm_smmu_v3_device *kvm_hyp_arm_smmu_v3_smmus; @@ -70,6 +73,9 @@ static inline u64 hyp_clock_ns(void) __ret; \ }) +/* Protected by host_mmu.lock from core code. */ +static struct io_pgtable *idmap_pgtable; + static bool is_cmdq_enabled(struct hyp_arm_smmu_v3_device *smmu) { return FIELD_GET(CR0_CMDQEN, smmu->cr0); @@ -215,6 +221,24 @@ static int smmu_send_cmd(struct hyp_arm_smmu_v3_device *smmu, return smmu_sync_cmd(smmu); } +static void smmu_tlb_flush_walk(unsigned long iova, size_t size, + size_t granule, void *cookie) +{ + /* TBD: Invalidate the range in all the SMMUs. */ +} + +static void smmu_tlb_add_page(struct iommu_iotlb_gather *gather, + unsigned long iova, size_t granule, + void *cookie) +{ + /* TBD: Invalidate the granule in all the SMMUs. */ +} + +static const struct iommu_flush_ops smmu_tlb_ops = { + .tlb_flush_walk = smmu_tlb_flush_walk, + .tlb_add_page = smmu_tlb_add_page, +}; + static int smmu_abort_gbpa(struct hyp_arm_smmu_v3_device *smmu) { int ret; @@ -480,6 +504,38 @@ static int smmu_init_device(struct hyp_arm_smmu_v3_device *smmu) return ret; } +static int smmu_init_pgt(void) +{ + /* Default values overridden based on SMMUs common features. */ + struct io_pgtable_cfg cfg = (struct io_pgtable_cfg) { + .tlb = &smmu_tlb_ops, + .pgsize_bitmap = ~0UL, + .ias = 48, + .oas = 48, + .coherent_walk = true, + .quirks = IO_PGTABLE_QUIRK_NO_WARN, + }; + struct hyp_arm_smmu_v3_device *smmu; + struct io_pgtable_ops *ops; + + for_each_smmu(smmu) { + cfg.ias = min(cfg.ias, smmu->oas); + cfg.oas = min(cfg.oas, smmu->oas); + cfg.pgsize_bitmap &= smmu->pgsize_bitmap; + cfg.coherent_walk &= !!(smmu->features & ARM_SMMU_FEAT_COHERENCY); + } + + /* At least PAGE_SIZE must be supported by all SMMUs */ + if ((cfg.pgsize_bitmap & PAGE_SIZE) == 0) + return -EINVAL; + + ops = kvm_alloc_io_pgtable_ops(ARM_64_LPAE_S2, &cfg, NULL); + if (!ops) + return -ENOMEM; + idmap_pgtable = io_pgtable_ops_to_pgtable(ops); + return 0; +} + /* Called while is the host is still trusted. */ static int smmu_init(void) { @@ -509,7 +565,10 @@ static int smmu_init(void) BUILD_BUG_ON(sizeof(hyp_spinlock_t) != sizeof(u32)); - return 0; + ret = smmu_init_pgt(); + if (ret) + goto out_reclaim_smmu; + return ret; out_reclaim_smmu: while (smmu != kvm_hyp_arm_smmu_v3_smmus) @@ -998,8 +1057,100 @@ static bool smmu_dabt_handler(struct user_pt_regs *regs, u64 esr, u64 addr) return false; } +static size_t smmu_pgsize_idmap(size_t size, u64 paddr, size_t pgsize_bitmap) +{ + size_t pgsizes; + + /* Remove page sizes that are larger than the current size */ + pgsizes = pgsize_bitmap & GENMASK_ULL(__fls(size), 0); + + /* Remove page sizes that the address is not aligned to. */ + if (likely(paddr)) + pgsizes &= GENMASK_ULL(__ffs(paddr), 0); + + WARN_ON(!pgsizes); + + /* Return the largest page size that fits. */ + return BIT(__fls(pgsizes)); +} + static int smmu_host_stage2_idmap(phys_addr_t start, phys_addr_t end, int prot) { + size_t pgsize = PAGE_SIZE, pgcount, size; + struct io_pgtable *pgtable = idmap_pgtable; + int ret = 0; + + end = min(end, BIT(pgtable->cfg.oas)); + if (start >= end) + return 0; + + size = end - start; + if (prot) { + size_t mapped; + + if (!(prot & IOMMU_MMIO)) + prot |= IOMMU_CACHE; + + while (size) { + mapped = 0; + /* + * We handle pages size for memory and MMIO differently: + * - memory: Map everything with PAGE_SIZE, that is guaranteed to + * find memory as we allocated enough pages to cover the entire + * memory, we do that as io-pgtable-arm doesn't support + * split_blk_unmap logic any more, so we can't break blocks once + * mapped to tables. + * - MMIO: Unlike memory, pKVM allocates 1G for all MMIO, while + * the MMIO space can be large, as it is assumed to cover the + * whole IAS that is not memory, we have to use block mappings, + * that is fine for MMIO as it is never donated at the moment, + * so we never need to unmap MMIO at the run time triggering + * split block logic. + */ + if (prot & IOMMU_MMIO) + pgsize = smmu_pgsize_idmap(size, start, pgtable->cfg.pgsize_bitmap); + + pgcount = size / pgsize; + ret = pgtable->ops.map_pages(&pgtable->ops, start, start, + pgsize, pgcount, prot, 0, &mapped); + size -= mapped; + start += mapped; + + if (ret == -EEXIST) { + /* + * It is possible to get EEXIST when a VM dies with pages + * in a shared state. + */ + ret = 0; + size -= pgsize; + start += pgsize; + continue; + } + if (!mapped || ret) + break; + } + } else { + struct iommu_iotlb_gather gather; + size_t unmapped; + + while (size) { + pgcount = size / pgsize; + iommu_iotlb_gather_init(&gather); + unmapped = pgtable->ops.unmap_pages(&pgtable->ops, start, + pgsize, pgcount, &gather); + size -= unmapped; + start += unmapped; + if (!unmapped) + break; + } + } + + if (ret) + return ret; + + if (WARN_ON(size)) + return -EINVAL; + return 0; } -- 2.55.0.1082.g2b9226bbc0-goog