From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 1C071C83F1A for ; Thu, 10 Jul 2025 06:39:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Type: Content-Transfer-Encoding:MIME-Version:References:In-Reply-To:Message-ID:Date :Subject:CC:To:From:Reply-To:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=kKhmEjVohefFqifBxcPk8Vy2F7mQIREO9bbjlZadeHc=; b=3ugSHM1sTfA8G96lsCaVaqlQJw JywK4PL6XRc1+Q8o2D7xi/lRVqlE8K3D5nXkl4OwakZ9kEeKE2zwot/HGtO8hJ0oKcJdddTOLhQIy 9xmIJ2yac8sH3Y7Iu7zGgYxIbJQ6U9otpIsHlTlxNy3+Kyfq68FVPBsVyT2+QVbSTYTLNkC3Znfth eXJCHL1Cw5yMBiz2B3Wwe8JIo/yq2HBgepr3XfJATo3AZSBGQg79M9sJHH2vlFm6p9jlazEBRiRQj ahuae+JPh4Bb1OeU8r5wxjkjng/bpt8wKOyokxuwXV+1s7+krH4oAqKdYr2y5TExpi+Sv2evrCMf2 ZTsJW8Rg==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.98.2 #2 (Red Hat Linux)) id 1uZkvx-0000000AtNq-3ei1; Thu, 10 Jul 2025 06:39:05 +0000 Received: from mail-bn1nam02on20606.outbound.protection.outlook.com ([2a01:111:f403:2407::606] helo=NAM02-BN1-obe.outbound.protection.outlook.com) by bombadil.infradead.org with esmtps (Exim 4.98.2 #2 (Red Hat Linux)) id 1uZkKl-0000000Am6l-3eEE for linux-arm-kernel@lists.infradead.org; Thu, 10 Jul 2025 06:00:41 +0000 ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=bj1GpFOHdIA1TMxW3N2jDdtG2k/GGErb+Qndpdrr5cIM3jLYVHf2CWYaxZ+CBkJMkUFUY8WUBhqC9qcX4XsAY/Na1Us2GnoE5WiP1v4Y/AFgpiRQ/qqqwqIXOv14WUDwwOQROkwVeS8NlyAr+Jg0EBEWuPwsFe8FxFNa/GXLARA7aJQOodI8iqSTZtcSNUdw8BB/jVCR9tGDj55gPfx6XpBvLdc4iqXcuNSmCHdhX4ere7HyvrZKfiTihBN6QJm0ozsKfIsqwVzCN9w5UrcStCbZiPM1Q/l0Cx16b4EuFgkhfqpcYwQyfWvXQZYVY5HBQkJW5jQig5IZz9sNprsPhA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=kKhmEjVohefFqifBxcPk8Vy2F7mQIREO9bbjlZadeHc=; b=a63ptSR4gKvjAHuK9FUW/CbTseKhrN8shG6vPb2DkCvPFMkBsicnjIJCN55VWNPWX09D56HW1CsUVw6k/YzAYtEkcby9gbB8SJHmloPr8DlQNAXw1axRiWdgVp4vy1VdjTa1tfMfZCc2ovpQH8f9G/cdGO3VkQXv3VjwBoAVhziPQWf8boFIivahw4pHLUXDtlBzfpq9BDs+O85vORFsiDWlKp9Hmuf60StGQHV8NtLSQqmroQGUgYflvvBeJ0PnvR5V7qL0gdmFE0op4O4q33i0mQibvdHuqs414jHd+hf4JLgpfDKnDCYMV6FdxTnITLiIR5weLl4K5MM+nB0Znw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=amd.com smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=kKhmEjVohefFqifBxcPk8Vy2F7mQIREO9bbjlZadeHc=; b=ZC17FrAG+BHt1E/EDW/V9oGDsIF7v+LYEyeE1+qf272SAoCkZ77g+1GHyQpA7MGIA2bEZIcWk+xaZcT6OS9QYzNeRwzX1ZXxtR69LyAF+J1DsQtWK2rNVP2kEJ7+DuixJjk+AMiIJJwsRQZi02i4h1LT3rnHSOuMAuXvldnvIjywxuruyetmM6ERXagoPm0OjUxxfOWwZfHjXz2+rN4HXtnf8vk3U/6DYKDa0UimXV6OCZenpvSGU+FGd79YtBx0CGZ+j9RfYpC+5p/tCPte82qejUoLpWKGdtiVeZzjzwtjEOJmHFMzibEoUyo29vIKcuDaUasrLurY1xd+/ihaqw== Received: from CH2PR17CA0012.namprd17.prod.outlook.com (2603:10b6:610:53::22) by IA1PR12MB6603.namprd12.prod.outlook.com (2603:10b6:208:3a1::17) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.8901.21; Thu, 10 Jul 2025 06:00:33 +0000 Received: from CH2PEPF00000146.namprd02.prod.outlook.com (2603:10b6:610:53:cafe::10) by CH2PR17CA0012.outlook.office365.com (2603:10b6:610:53::22) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.20.8922.22 via Frontend Transport; Thu, 10 Jul 2025 06:00:32 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by CH2PEPF00000146.mail.protection.outlook.com (10.167.244.103) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.20.8922.22 via Frontend Transport; Thu, 10 Jul 2025 06:00:32 +0000 Received: from rnnvmail205.nvidia.com (10.129.68.10) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.4; Wed, 9 Jul 2025 23:00:12 -0700 Received: from rnnvmail204.nvidia.com (10.129.68.6) by rnnvmail205.nvidia.com (10.129.68.10) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.14; Wed, 9 Jul 2025 23:00:12 -0700 Received: from Asurada-Nvidia.nvidia.com (10.127.8.14) by mail.nvidia.com (10.129.68.6) with Microsoft SMTP Server id 15.2.1544.14 via Frontend Transport; Wed, 9 Jul 2025 23:00:10 -0700 From: Nicolin Chen To: CC: , , , , , , , , , , , , , , , , , , , , , , , , , , , , Subject: [PATCH v9 14/29] iommufd/viommu: Add IOMMUFD_CMD_HW_QUEUE_ALLOC ioctl Date: Wed, 9 Jul 2025 22:59:06 -0700 Message-ID: X-Mailer: git-send-email 2.43.0 In-Reply-To: References: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Content-Type: text/plain X-NV-OnPremToCloud: AnonymousSubmission X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CH2PEPF00000146:EE_|IA1PR12MB6603:EE_ X-MS-Office365-Filtering-Correlation-Id: 61f48136-17a8-4911-9f9c-08ddbf7714ed X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|82310400026|36860700013|1800799024|7416014|376014; X-Microsoft-Antispam-Message-Info: =?us-ascii?Q?2/eSukthBQYJrJrhi5aERUEUy15fk/WghsP/7N/lBI4XRhFNVIEr7NMesd2S?= =?us-ascii?Q?EuUaIHL+rBoWxWkM/0XJJrc46Efzxj9frLaCzsMygExEL9BqdBRLljGCfjJ4?= =?us-ascii?Q?Qwr55SUQdbi1LV8+KgUHYeA4iTwJKExJDGWwrpIV+DZnALYlXfziRerjfKDu?= =?us-ascii?Q?Fv4dMbF1lgWmAGW45OyNXo/StfGAbTO97PPZuadS1K4h6TXMJpat5zI0deSj?= =?us-ascii?Q?P0cT0fii8uMjTHqLx0lE6ebSz81yt1EFtr/6ZYFdJDXWSU13dVrpXlLQpUAT?= =?us-ascii?Q?yfKhNsMMMgnaervCtOMLOFFB/LlwQ1UeBdEFkrVu+QdxUiVx5nszsI6mLdoa?= =?us-ascii?Q?fvOLruLsmOw0fqXJjzKu7YgT6oS0e3o9GcC0vXub0CRgE9eCMlMklyrqhmvs?= =?us-ascii?Q?5bl5QR8Qy3MHskS72qgkChHo/hF6vDBuK4zTNowYKvu/ud14BRJf0SBMcV/F?= =?us-ascii?Q?KnbrSMfDBF+kgvzd81fN3ts/eMNa6jV1ZM3R/0GDSAj/warF4Hb9bNd29LAL?= =?us-ascii?Q?3CjXNU2a8Uyp5xy/sHoWN/OWzRxCtgpZYhsMaT69gRN4R39BpXGT5AFaGhDf?= =?us-ascii?Q?CTg/AEjb3sIUuIBNmXwZuBOJPwftz9gWH+Q+vnk1xD/WspKW3xpPyUJkMUHX?= =?us-ascii?Q?MYeeYHaeZS8kcFjTWMTgUArrMK0CS2RLsWAGwgi+UEp70G5OF84iyF5KbufG?= =?us-ascii?Q?O4nO5Golc4bdzuT3Lj8goL5ELtCont5yXSBJ88y8t8aL+x4WMYC/8w54SeXz?= =?us-ascii?Q?DP5zQ7HGK1zX/eCo5ZuXrHt32wXNqC7QoB7oe3QpVAykiARv+NJ7tJJxENyG?= =?us-ascii?Q?Kd6uX4Tm8LHCXt3DYgrKxZmQ3zWveD4iS94wUzO2f1U4HExQd1EjP6+fEqpn?= =?us-ascii?Q?6dYj8+/M57urqO115jBAdQTT/loGqlDjzLmHCdVQuIT/51qG6nPRX6KnK4+j?= =?us-ascii?Q?lOL9pl9heD/H2aGGS7IQOhxCpyl0T+XSiVZ9yTWv59bKpVP+C96tMFZuTRC6?= =?us-ascii?Q?yj1cgtg2k297y+7PHG4Lghu+p+khmQcwy/CTgR9PfVweZDNlcfo7b308nxal?= =?us-ascii?Q?ZVaM3f86vpRm+mVNmRkMLnSTpEKJvQ0tl+XZZMy+pV6pXNuczD4sEyUlDZPN?= =?us-ascii?Q?P2DCxSH5YdCz6+GR5kBrLjAcFQDyHBRASHu+QSQYFFWib6CZrLkpO7d9c678?= =?us-ascii?Q?N225R0RETutVTS4gOkKYIB16se8YAtzR884DOjI5nO5p2Vj+lmFAQCyRDwnO?= =?us-ascii?Q?ng0JqF2UCdMbhiHPEhlXPYhpiTLbH8B/oQZ/qxJYiulOlgfy4pWZIIrxzXje?= =?us-ascii?Q?MOsTovT+hvXUVJyWEZvIkhgIChuc+9QCsjGRfJ3eNvXTbq3u5UVnsp48oHxr?= =?us-ascii?Q?bbPfEFKG9bLZMRsZJfDkHace9iRKk6O+w5wZlJQYph/Mxjs3xkAQoHBlSVe2?= =?us-ascii?Q?5uvmqpPCSHOKCoeg1t26/bkkhRwfnrJ5ui4Qq/0z2dqmxe8wdYd5Au2HBohw?= =?us-ascii?Q?okWO0e1zoakLZEoUDYYRqkHyCi2spQDb52wu?= X-Forefront-Antispam-Report: CIP:216.228.117.160;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:mail.nvidia.com;PTR:dc6edge1.nvidia.com;CAT:NONE;SFS:(13230040)(82310400026)(36860700013)(1800799024)(7416014)(376014);DIR:OUT;SFP:1101; X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 10 Jul 2025 06:00:32.8232 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 61f48136-17a8-4911-9f9c-08ddbf7714ed X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a;Ip=[216.228.117.160];Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: CH2PEPF00000146.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: IA1PR12MB6603 X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20250709_230039_978516_67C4C493 X-CRM114-Status: GOOD ( 18.45 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org Introduce a new IOMMUFD_CMD_HW_QUEUE_ALLOC ioctl for user space to allocate a HW QUEUE object for a vIOMMU specific HW-accelerated queue, e.g.: - NVIDIA's Virtual Command Queue - AMD vIOMMU's Command Buffer, Event Log Buffers, and PPR Log Buffers Since this is introduced with NVIDIA's VCMDQs that access the guest memory in the physical address space, add an iommufd_hw_queue_alloc_phys() helper that will create an access object to the queue memory in the IOAS, to avoid the mappings of the guest memory from being unmapped, during the life cycle of the HW queue object. AMD's HW will need an hw_queue_init op that is mutually exclusive with the hw_queue_init_phys op, and their case will bypass the access part, i.e. no iommufd_hw_queue_alloc_phys() call. Reviewed-by: Pranjal Shrivastava Reviewed-by: Kevin Tian Reviewed-by: Lu Baolu Signed-off-by: Nicolin Chen --- drivers/iommu/iommufd/iommufd_private.h | 2 + include/linux/iommufd.h | 1 + include/uapi/linux/iommufd.h | 33 +++++ drivers/iommu/iommufd/main.c | 6 + drivers/iommu/iommufd/viommu.c | 180 ++++++++++++++++++++++++ 5 files changed, 222 insertions(+) diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommufd/iommufd_private.h index 06b8c2e2d9e6..dcd609573244 100644 --- a/drivers/iommu/iommufd/iommufd_private.h +++ b/drivers/iommu/iommufd/iommufd_private.h @@ -652,6 +652,8 @@ int iommufd_viommu_alloc_ioctl(struct iommufd_ucmd *ucmd); void iommufd_viommu_destroy(struct iommufd_object *obj); int iommufd_vdevice_alloc_ioctl(struct iommufd_ucmd *ucmd); void iommufd_vdevice_destroy(struct iommufd_object *obj); +int iommufd_hw_queue_alloc_ioctl(struct iommufd_ucmd *ucmd); +void iommufd_hw_queue_destroy(struct iommufd_object *obj); #ifdef CONFIG_IOMMUFD_TEST int iommufd_test(struct iommufd_ucmd *ucmd); diff --git a/include/linux/iommufd.h b/include/linux/iommufd.h index f13f3ca6adb5..ce4011a2fc27 100644 --- a/include/linux/iommufd.h +++ b/include/linux/iommufd.h @@ -123,6 +123,7 @@ struct iommufd_vdevice { struct iommufd_hw_queue { struct iommufd_object obj; struct iommufd_viommu *viommu; + struct iommufd_access *access; u64 base_addr; /* in guest physical address space */ size_t length; diff --git a/include/uapi/linux/iommufd.h b/include/uapi/linux/iommufd.h index 640a8b5147c2..55459b9eee31 100644 --- a/include/uapi/linux/iommufd.h +++ b/include/uapi/linux/iommufd.h @@ -56,6 +56,7 @@ enum { IOMMUFD_CMD_VDEVICE_ALLOC = 0x91, IOMMUFD_CMD_IOAS_CHANGE_PROCESS = 0x92, IOMMUFD_CMD_VEVENTQ_ALLOC = 0x93, + IOMMUFD_CMD_HW_QUEUE_ALLOC = 0x94, }; /** @@ -1156,4 +1157,36 @@ enum iommu_hw_queue_type { IOMMU_HW_QUEUE_TYPE_DEFAULT = 0, }; +/** + * struct iommu_hw_queue_alloc - ioctl(IOMMU_HW_QUEUE_ALLOC) + * @size: sizeof(struct iommu_hw_queue_alloc) + * @flags: Must be 0 + * @viommu_id: Virtual IOMMU ID to associate the HW queue with + * @type: One of enum iommu_hw_queue_type + * @index: The logical index to the HW queue per virtual IOMMU for a multi-queue + * model + * @out_hw_queue_id: The ID of the new HW queue + * @nesting_parent_iova: Base address of the queue memory in the guest physical + * address space + * @length: Length of the queue memory + * + * Allocate a HW queue object for a vIOMMU-specific HW-accelerated queue, which + * allows HW to access a guest queue memory described using @nesting_parent_iova + * and @length. + * + * A vIOMMU can allocate multiple queues, but it must use a different @index per + * type to separate each allocation, e.g. + * Type1 HW queue0, Type1 HW queue1, Type2 HW queue0, ... + */ +struct iommu_hw_queue_alloc { + __u32 size; + __u32 flags; + __u32 viommu_id; + __u32 type; + __u32 index; + __u32 out_hw_queue_id; + __aligned_u64 nesting_parent_iova; + __aligned_u64 length; +}; +#define IOMMU_HW_QUEUE_ALLOC _IO(IOMMUFD_TYPE, IOMMUFD_CMD_HW_QUEUE_ALLOC) #endif diff --git a/drivers/iommu/iommufd/main.c b/drivers/iommu/iommufd/main.c index 778694d7c207..4e8dbbfac890 100644 --- a/drivers/iommu/iommufd/main.c +++ b/drivers/iommu/iommufd/main.c @@ -354,6 +354,7 @@ union ucmd_buffer { struct iommu_destroy destroy; struct iommu_fault_alloc fault; struct iommu_hw_info info; + struct iommu_hw_queue_alloc hw_queue; struct iommu_hwpt_alloc hwpt; struct iommu_hwpt_get_dirty_bitmap get_dirty_bitmap; struct iommu_hwpt_invalidate cache; @@ -396,6 +397,8 @@ static const struct iommufd_ioctl_op iommufd_ioctl_ops[] = { struct iommu_fault_alloc, out_fault_fd), IOCTL_OP(IOMMU_GET_HW_INFO, iommufd_get_hw_info, struct iommu_hw_info, __reserved), + IOCTL_OP(IOMMU_HW_QUEUE_ALLOC, iommufd_hw_queue_alloc_ioctl, + struct iommu_hw_queue_alloc, length), IOCTL_OP(IOMMU_HWPT_ALLOC, iommufd_hwpt_alloc, struct iommu_hwpt_alloc, __reserved), IOCTL_OP(IOMMU_HWPT_GET_DIRTY_BITMAP, iommufd_hwpt_get_dirty_bitmap, @@ -559,6 +562,9 @@ static const struct iommufd_object_ops iommufd_object_ops[] = { [IOMMUFD_OBJ_FAULT] = { .destroy = iommufd_fault_destroy, }, + [IOMMUFD_OBJ_HW_QUEUE] = { + .destroy = iommufd_hw_queue_destroy, + }, [IOMMUFD_OBJ_HWPT_PAGING] = { .destroy = iommufd_hwpt_paging_destroy, .abort = iommufd_hwpt_paging_abort, diff --git a/drivers/iommu/iommufd/viommu.c b/drivers/iommu/iommufd/viommu.c index 081ee6697a11..91339f799916 100644 --- a/drivers/iommu/iommufd/viommu.c +++ b/drivers/iommu/iommufd/viommu.c @@ -201,3 +201,183 @@ int iommufd_vdevice_alloc_ioctl(struct iommufd_ucmd *ucmd) iommufd_put_object(ucmd->ictx, &viommu->obj); return rc; } + +static void iommufd_hw_queue_destroy_access(struct iommufd_ctx *ictx, + struct iommufd_access *access, + u64 base_iova, size_t length) +{ + u64 aligned_iova = PAGE_ALIGN_DOWN(base_iova); + u64 offset = base_iova - aligned_iova; + + iommufd_access_unpin_pages(access, aligned_iova, + PAGE_ALIGN(length + offset)); + iommufd_access_detach_internal(access); + iommufd_access_destroy_internal(ictx, access); +} + +void iommufd_hw_queue_destroy(struct iommufd_object *obj) +{ + struct iommufd_hw_queue *hw_queue = + container_of(obj, struct iommufd_hw_queue, obj); + + if (hw_queue->destroy) + hw_queue->destroy(hw_queue); + if (hw_queue->access) + iommufd_hw_queue_destroy_access(hw_queue->viommu->ictx, + hw_queue->access, + hw_queue->base_addr, + hw_queue->length); + if (hw_queue->viommu) + refcount_dec(&hw_queue->viommu->obj.users); +} + +/* + * When the HW accesses the guest queue via physical addresses, the underlying + * physical pages of the guest queue must be contiguous. Also, for the security + * concern that IOMMUFD_CMD_IOAS_UNMAP could potentially remove the mappings of + * the guest queue from the nesting parent iopt while the HW is still accessing + * the guest queue memory physically, such a HW queue must require an access to + * pin the underlying pages and prevent that from happening. + */ +static struct iommufd_access * +iommufd_hw_queue_alloc_phys(struct iommu_hw_queue_alloc *cmd, + struct iommufd_viommu *viommu, phys_addr_t *base_pa) +{ + u64 aligned_iova = PAGE_ALIGN_DOWN(cmd->nesting_parent_iova); + u64 offset = cmd->nesting_parent_iova - aligned_iova; + struct iommufd_access *access; + struct page **pages; + size_t max_npages; + size_t length; + size_t i; + int rc; + + /* max_npages = DIV_ROUND_UP(offset + cmd->length, PAGE_SIZE) */ + if (check_add_overflow(offset, cmd->length, &length)) + return ERR_PTR(-ERANGE); + if (check_add_overflow(length, PAGE_SIZE - 1, &length)) + return ERR_PTR(-ERANGE); + max_npages = length / PAGE_SIZE; + /* length needs to be page aligned too */ + length = max_npages * PAGE_SIZE; + + /* + * Use kvcalloc() to avoid memory fragmentation for a large page array. + * Set __GFP_NOWARN to avoid syzkaller blowups + */ + pages = kvcalloc(max_npages, sizeof(*pages), GFP_KERNEL | __GFP_NOWARN); + if (!pages) + return ERR_PTR(-ENOMEM); + + access = iommufd_access_create_internal(viommu->ictx); + if (IS_ERR(access)) { + rc = PTR_ERR(access); + goto out_free; + } + + rc = iommufd_access_attach_internal(access, viommu->hwpt->ioas); + if (rc) + goto out_destroy; + + rc = iommufd_access_pin_pages(access, aligned_iova, length, pages, 0); + if (rc) + goto out_detach; + + /* Validate if the underlying physical pages are contiguous */ + for (i = 1; i < max_npages; i++) { + if (page_to_pfn(pages[i]) == page_to_pfn(pages[i - 1]) + 1) + continue; + rc = -EFAULT; + goto out_unpin; + } + + *base_pa = (page_to_pfn(pages[0]) << PAGE_SHIFT) + offset; + kfree(pages); + return access; + +out_unpin: + iommufd_access_unpin_pages(access, aligned_iova, length); +out_detach: + iommufd_access_detach_internal(access); +out_destroy: + iommufd_access_destroy_internal(viommu->ictx, access); +out_free: + kfree(pages); + return ERR_PTR(rc); +} + +int iommufd_hw_queue_alloc_ioctl(struct iommufd_ucmd *ucmd) +{ + struct iommu_hw_queue_alloc *cmd = ucmd->cmd; + struct iommufd_hw_queue *hw_queue; + struct iommufd_viommu *viommu; + struct iommufd_access *access; + size_t hw_queue_size; + phys_addr_t base_pa; + u64 last; + int rc; + + if (cmd->flags || cmd->type == IOMMU_HW_QUEUE_TYPE_DEFAULT) + return -EOPNOTSUPP; + if (!cmd->length) + return -EINVAL; + if (check_add_overflow(cmd->nesting_parent_iova, cmd->length - 1, + &last)) + return -EOVERFLOW; + + viommu = iommufd_get_viommu(ucmd, cmd->viommu_id); + if (IS_ERR(viommu)) + return PTR_ERR(viommu); + + if (!viommu->ops || !viommu->ops->get_hw_queue_size || + !viommu->ops->hw_queue_init_phys) { + rc = -EOPNOTSUPP; + goto out_put_viommu; + } + + hw_queue_size = viommu->ops->get_hw_queue_size(viommu, cmd->type); + if (!hw_queue_size) { + rc = -EOPNOTSUPP; + goto out_put_viommu; + } + + /* + * It is a driver bug for providing a hw_queue_size smaller than the + * core HW queue structure size + */ + if (WARN_ON_ONCE(hw_queue_size < sizeof(*hw_queue))) { + rc = -EOPNOTSUPP; + goto out_put_viommu; + } + + hw_queue = (struct iommufd_hw_queue *)_iommufd_object_alloc_ucmd( + ucmd, hw_queue_size, IOMMUFD_OBJ_HW_QUEUE); + if (IS_ERR(hw_queue)) { + rc = PTR_ERR(hw_queue); + goto out_put_viommu; + } + + access = iommufd_hw_queue_alloc_phys(cmd, viommu, &base_pa); + if (IS_ERR(access)) { + rc = PTR_ERR(access); + goto out_put_viommu; + } + + hw_queue->viommu = viommu; + refcount_inc(&viommu->obj.users); + hw_queue->access = access; + hw_queue->type = cmd->type; + hw_queue->length = cmd->length; + hw_queue->base_addr = cmd->nesting_parent_iova; + + rc = viommu->ops->hw_queue_init_phys(hw_queue, cmd->index, base_pa); + if (rc) + goto out_put_viommu; + + cmd->out_hw_queue_id = hw_queue->obj.id; + rc = iommufd_ucmd_respond(ucmd, sizeof(*cmd)); + +out_put_viommu: + iommufd_put_object(ucmd->ictx, &viommu->obj); + return rc; +} -- 2.43.0