* [PATCH v18 1/7] firmware: arm_rmm: Add SMC definitions for calling the RMM
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 0:28 ` Gavin Shan
2026-09-12 8:36 ` [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init Suzuki K Poulose
` (5 subsequent siblings)
6 siblings, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
The RMM (Realm Management Monitor) provides functionality that can be
accessed by SMC calls from the host.
The SMC definitions are based on DEN0137[1] version 2.0-bet3
[1] https://developer.arm.com/documentation/den0137/2-0bet3/
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
* Use GENMASK()/BIT() for masks consistently
* Rename RMI_{ADDR_RANGE, DONATE}_SIZE => RMI_{*}_BLOCK_SIZE
* Reorder the definitions for MSB to LSB
* Add definions for RMI_OP_MEM_*CONTIG and RMI_OP_CAN*_CANCEL
Changes since v16:
* Updated definitions to RMM specification v2.0-bet3.
Changes since v15:
* Dropped unused symbols REC_MAX_GIC_NUM_LRS and RMI_PERMITTED_GICV3_HCR_BITS.
* Output is now (partially) generated from the spec source.
Changes since v14:
* Updated to RMM spec v2.0-bet2 but without the changes to move
metadata out of individual address range descriptors as this is
expected to be reverted in a future spec release.
Changes since v13:
* Updated to RMM spec v2.0-bet1
Changes since v12:
* Updated to RMM spec v2.0-bet0
Changes since v9:
* Corrected size of 'ripas_value' in struct rec_exit. The spec states
this is an 8-bit type with padding afterwards (rather than a u64).
Changes since v8:
* Added RMI_PERMITTED_GICV3_HCR_BITS to define which bits the RMM
permits to be modified.
Changes since v6:
* Renamed REC_ENTER_xxx defines to include 'FLAG' to make it obvious
these are flag values.
Changes since v5:
* Sorted the SMC #defines by value.
* Renamed SMI_RxI_CALL to SMI_RMI_CALL since the macro is only used for
RMI calls.
* Renamed REC_GIC_NUM_LRS to REC_MAX_GIC_NUM_LRS since the actual
number of available list registers could be lower.
* Provided a define for the reserved fields of FeatureRegister0.
* Fix inconsistent names for padding fields.
Changes since v4:
* Update to point to final released RMM spec.
* Minor rearrangements.
Changes since v3:
* Update to match RMM spec v1.0-rel0-rc1.
Changes since v2:
* Fix specification link.
* Rename rec_entry->rec_enter to match spec.
* Fix size of pmu_ovf_status to match spec.
---
include/linux/arm-smccc-rmi.h | 497 ++++++++++++++++++++++++++++++++++
1 file changed, 497 insertions(+)
create mode 100644 include/linux/arm-smccc-rmi.h
diff --git a/include/linux/arm-smccc-rmi.h b/include/linux/arm-smccc-rmi.h
new file mode 100644
index 0000000000000..214d6228dfc22
--- /dev/null
+++ b/include/linux/arm-smccc-rmi.h
@@ -0,0 +1,497 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (C) 2023-2026 ARM Ltd.
+ *
+ * The values and structures in this file are from the Realm Management Monitor
+ * specification (DEN0137) version 2.0-bet3:
+ * https://developer.arm.com/documentation/den0137/2-0bet3/
+ */
+
+#ifndef __LINUX_ARM_SMCCC_RMI_H_
+#define __LINUX_ARM_SMCCC_RMI_H_
+
+#include <linux/arm-smccc.h>
+#include <linux/bitfield.h>
+#include <linux/bits.h>
+#include <linux/build_bug.h>
+#include <linux/sizes.h>
+
+#include <asm/page.h>
+
+#define SMC_RMI_CALL(func) \
+ ARM_SMCCC_CALL_VAL(ARM_SMCCC_FAST_CALL, \
+ ARM_SMCCC_SMC_64, \
+ ARM_SMCCC_OWNER_STANDARD, \
+ (func))
+
+#define SMC_RMI_VERSION SMC_RMI_CALL(0x0150)
+
+#define SMC_RMI_RTT_DATA_MAP_INIT SMC_RMI_CALL(0x0153)
+
+#define SMC_RMI_REALM_ACTIVATE SMC_RMI_CALL(0x0157)
+#define SMC_RMI_REALM_CREATE SMC_RMI_CALL(0x0158)
+#define SMC_RMI_REALM_DESTROY SMC_RMI_CALL(0x0159)
+#define SMC_RMI_REC_CREATE SMC_RMI_CALL(0x015a)
+#define SMC_RMI_REC_DESTROY SMC_RMI_CALL(0x015b)
+#define SMC_RMI_REC_ENTER SMC_RMI_CALL(0x015c)
+#define SMC_RMI_RTT_CREATE SMC_RMI_CALL(0x015d)
+#define SMC_RMI_RTT_DESTROY SMC_RMI_CALL(0x015e)
+
+#define SMC_RMI_RTT_READ_ENTRY SMC_RMI_CALL(0x0161)
+
+#define SMC_RMI_RTT_DEV_VALIDATE SMC_RMI_CALL(0x0163)
+#define SMC_RMI_PSCI_COMPLETE SMC_RMI_CALL(0x0164)
+#define SMC_RMI_FEATURES SMC_RMI_CALL(0x0165)
+#define SMC_RMI_RTT_FOLD SMC_RMI_CALL(0x0166)
+
+#define SMC_RMI_RTT_INIT_RIPAS SMC_RMI_CALL(0x0168)
+#define SMC_RMI_RTT_SET_RIPAS SMC_RMI_CALL(0x0169)
+#define SMC_RMI_VSMMU_CREATE SMC_RMI_CALL(0x016a)
+#define SMC_RMI_VSMMU_DESTROY SMC_RMI_CALL(0x016b)
+
+#define SMC_RMI_RMM_CONFIG_SET SMC_RMI_CALL(0x016e)
+#define SMC_RMI_PSMMU_IRQ_NOTIFY SMC_RMI_CALL(0x016f)
+#define SMC_RMI_ATTEST_PLAT_TOKEN_REFRESH SMC_RMI_CALL(0x0170)
+
+#define SMC_RMI_PDEV_ABORT SMC_RMI_CALL(0x0174)
+#define SMC_RMI_PDEV_COMMUNICATE SMC_RMI_CALL(0x0175)
+#define SMC_RMI_PDEV_CREATE SMC_RMI_CALL(0x0176)
+#define SMC_RMI_PDEV_DESTROY SMC_RMI_CALL(0x0177)
+#define SMC_RMI_PDEV_GET_STATE SMC_RMI_CALL(0x0178)
+
+#define SMC_RMI_PDEV_STREAM_KEY_REFRESH SMC_RMI_CALL(0x017a)
+#define SMC_RMI_PDEV_SET_PUBKEY SMC_RMI_CALL(0x017b)
+#define SMC_RMI_PDEV_STOP SMC_RMI_CALL(0x017c)
+#define SMC_RMI_RTT_AUX_CREATE SMC_RMI_CALL(0x017d)
+#define SMC_RMI_RTT_AUX_DESTROY SMC_RMI_CALL(0x017e)
+#define SMC_RMI_RTT_AUX_FOLD SMC_RMI_CALL(0x017f)
+
+#define SMC_RMI_VDEV_ABORT SMC_RMI_CALL(0x0185)
+#define SMC_RMI_VDEV_COMMUNICATE SMC_RMI_CALL(0x0186)
+#define SMC_RMI_VDEV_CREATE SMC_RMI_CALL(0x0187)
+#define SMC_RMI_VDEV_DESTROY SMC_RMI_CALL(0x0188)
+#define SMC_RMI_VDEV_GET_STATE SMC_RMI_CALL(0x0189)
+#define SMC_RMI_VDEV_UNLOCK SMC_RMI_CALL(0x018a)
+#define SMC_RMI_RTT_SET_S2AP SMC_RMI_CALL(0x018b)
+
+#define SMC_RMI_VDEV_GET_INTERFACE_REPORT SMC_RMI_CALL(0x01d0)
+#define SMC_RMI_VDEV_GET_MEASUREMENTS SMC_RMI_CALL(0x01d1)
+#define SMC_RMI_VDEV_LOCK SMC_RMI_CALL(0x01d2)
+#define SMC_RMI_VDEV_START SMC_RMI_CALL(0x01d3)
+
+#define SMC_RMI_VSMMU_EVENT_HANDLE SMC_RMI_CALL(0x01d6)
+#define SMC_RMI_PSMMU_ACTIVATE SMC_RMI_CALL(0x01d7)
+#define SMC_RMI_PSMMU_DEACTIVATE SMC_RMI_CALL(0x01d8)
+
+#define SMC_RMI_PSMMU_ST_L2_CREATE SMC_RMI_CALL(0x01db)
+#define SMC_RMI_PSMMU_ST_L2_DESTROY SMC_RMI_CALL(0x01dc)
+#define SMC_RMI_DPT_L0_CREATE SMC_RMI_CALL(0x01dd)
+#define SMC_RMI_DPT_L0_DESTROY SMC_RMI_CALL(0x01de)
+#define SMC_RMI_DPT_L1_CREATE SMC_RMI_CALL(0x01df)
+#define SMC_RMI_DPT_L1_DESTROY SMC_RMI_CALL(0x01e0)
+#define SMC_RMI_GRANULE_TRACKING_GET SMC_RMI_CALL(0x01e1)
+
+#define SMC_RMI_GRANULE_TRACKING_SET SMC_RMI_CALL(0x01e3)
+
+#define SMC_RMI_RMM_CONFIG_GET SMC_RMI_CALL(0x01ec)
+
+#define SMC_RMI_RMM_STATE_GET SMC_RMI_CALL(0x01ee)
+
+#define SMC_RMI_PSMMU_EVENT_CONSUME SMC_RMI_CALL(0x01f0)
+#define SMC_RMI_GRANULE_RANGE_DELEGATE SMC_RMI_CALL(0x01f1)
+#define SMC_RMI_GRANULE_RANGE_UNDELEGATE SMC_RMI_CALL(0x01f2)
+#define SMC_RMI_GPT_L1_CREATE SMC_RMI_CALL(0x01f3)
+#define SMC_RMI_GPT_L1_DESTROY SMC_RMI_CALL(0x01f4)
+#define SMC_RMI_RTT_DATA_MAP SMC_RMI_CALL(0x01f5)
+#define SMC_RMI_RTT_DATA_UNMAP SMC_RMI_CALL(0x01f6)
+#define SMC_RMI_RTT_DEV_MAP SMC_RMI_CALL(0x01f7)
+#define SMC_RMI_RTT_DEV_UNMAP SMC_RMI_CALL(0x01f8)
+#define SMC_RMI_RTT_ARCH_DEV_MAP SMC_RMI_CALL(0x01f9)
+#define SMC_RMI_RTT_ARCH_DEV_UNMAP SMC_RMI_CALL(0x01fa)
+#define SMC_RMI_RTT_UNPROT_MAP SMC_RMI_CALL(0x01fb)
+#define SMC_RMI_RTT_UNPROT_UNMAP SMC_RMI_CALL(0x01fc)
+#define SMC_RMI_RTT_AUX_PROT_MAP SMC_RMI_CALL(0x01fd)
+#define SMC_RMI_RTT_AUX_PROT_UNMAP SMC_RMI_CALL(0x01fe)
+#define SMC_RMI_RTT_AUX_UNPROT_MAP SMC_RMI_CALL(0x01ff)
+#define SMC_RMI_RTT_AUX_UNPROT_UNMAP SMC_RMI_CALL(0x0200)
+#define SMC_RMI_REALM_TERMINATE SMC_RMI_CALL(0x0201)
+#define SMC_RMI_RMM_ACTIVATE SMC_RMI_CALL(0x0202)
+#define SMC_RMI_OP_CONTINUE SMC_RMI_CALL(0x0203)
+#define SMC_RMI_PDEV_STREAM_CONNECT SMC_RMI_CALL(0x0204)
+#define SMC_RMI_PDEV_STREAM_DISCONNECT SMC_RMI_CALL(0x0205)
+#define SMC_RMI_PDEV_STREAM_COMPLETE SMC_RMI_CALL(0x0206)
+#define SMC_RMI_PDEV_STREAM_KEY_PURGE SMC_RMI_CALL(0x0207)
+#define SMC_RMI_OP_MEM_DONATE SMC_RMI_CALL(0x0208)
+#define SMC_RMI_OP_MEM_RECLAIM SMC_RMI_CALL(0x0209)
+#define SMC_RMI_OP_CANCEL SMC_RMI_CALL(0x020a)
+#define SMC_RMI_VSMMU_FEATURES SMC_RMI_CALL(0x020b)
+#define SMC_RMI_VSMMU_CMD_GET SMC_RMI_CALL(0x020c)
+#define SMC_RMI_VSMMU_CMD_COMPLETE SMC_RMI_CALL(0x020d)
+#define SMC_RMI_PSMMU_INFO SMC_RMI_CALL(0x020e)
+#define SMC_RMI_RMM_DEACTIVATE SMC_RMI_CALL(0x020f)
+#define SMC_RMI_PDEV_STREAM_INFO SMC_RMI_CALL(0x0210)
+#define SMC_RMI_GPT_INFO SMC_RMI_CALL(0x0211)
+
+#define RMI_ABI_MAJOR_VERSION 2
+#define RMI_ABI_MINOR_VERSION 0
+
+#define RMI_ABI_VERSION_GET_MAJOR(version) ((version) >> 16)
+#define RMI_ABI_VERSION_GET_MINOR(version) ((version) & 0xFFFF)
+#define RMI_ABI_VERSION(major, minor) (((major) << 16) | (minor))
+
+#define RMI_RETURN_STATUS_MASK GENMASK(7, 0)
+#define RMI_RETURN_INDEX_MASK GENMASK(15, 8)
+#define RMI_RETURN_MEMREQ_MASK GENMASK(9, 8)
+#define RMI_RETURN_CAN_CANCEL_MASK BIT(10)
+
+#define RMI_RETURN_STATUS(ret) FIELD_GET(RMI_RETURN_STATUS_MASK, ret)
+#define RMI_RETURN_INDEX(ret) FIELD_GET(RMI_RETURN_INDEX_MASK, ret)
+#define RMI_RETURN_MEMREQ(ret) FIELD_GET(RMI_RETURN_MEMREQ_MASK, ret)
+#define RMI_RETURN_CAN_CANCEL(ret) FIELD_GET(RMI_RETURN_CAN_CANCEL_MASK, ret)
+
+#define RMI_SUCCESS 0
+#define RMI_ERROR_INPUT 1
+#define RMI_ERROR_REALM 2
+#define RMI_ERROR_REC 3
+#define RMI_ERROR_RTT 4
+#define RMI_ERROR_NOT_SUPPORTED 5
+#define RMI_ERROR_DEVICE 6
+#define RMI_ERROR_RTT_AUX 7
+#define RMI_ERROR_PSMMU_ST 8
+#define RMI_ERROR_DPT 9
+#define RMI_BUSY 10
+#define RMI_ERROR_GLOBAL 11
+#define RMI_ERROR_TRACKING 12
+#define RMI_INCOMPLETE 13
+#define RMI_BLOCKED 14
+#define RMI_ERROR_GPT 15
+#define RMI_ERROR_GRANULE 16
+
+#define RMI_CONTINUE_KEEP_GOING 0
+#define RMI_CONTINUE_STOP 1
+
+#define RMI_OP_MEM_REQ_NONE 0
+#define RMI_OP_MEM_REQ_DONATE 1
+#define RMI_OP_MEM_REQ_RECLAIM 2
+
+#define RMI_OP_CANNOT_CANCEL 0
+#define RMI_OP_CAN_CANCEL 1
+
+#define RMI_DONATE_STATE_MASK GENMASK(18, 17)
+#define RMI_DONATE_CONTIG_MASK BIT(16)
+#define RMI_DONATE_COUNT_MASK GENMASK(15, 2)
+#define RMI_DONATE_BLOCK_SIZE_MASK GENMASK(1, 0)
+
+#define RMI_DONATE_STATE(req) FIELD_GET(RMI_DONATE_STATE_MASK, req)
+#define RMI_DONATE_CONTIG(req) FIELD_GET(RMI_DONATE_CONTIG_MASK, req)
+#define RMI_DONATE_COUNT(req) FIELD_GET(RMI_DONATE_COUNT_MASK, req)
+#define RMI_DONATE_BLOCK_SIZE(req) FIELD_GET(RMI_DONATE_BLOCK_SIZE_MASK, req)
+
+#define RMI_OP_MEM_DELEGATED 0
+#define RMI_OP_MEM_UNDELEGATED 1
+#define RMI_OP_MEM_CONDITIONAL 2
+
+#define RMI_OP_MEM_NON_CONTIG 0
+#define RMI_OP_MEM_CONTIG 1
+
+#define RMI_ADDR_TYPE_NONE 0
+#define RMI_ADDR_TYPE_SINGLE 1
+#define RMI_ADDR_TYPE_LIST 2
+
+#define RMI_ADDR_RANGE_STATE_MASK GENMASK(63, 62)
+#define RMI_ADDR_RANGE_ADDR_MASK GENMASK(51, PAGE_SHIFT)
+#define RMI_ADDR_RANGE_COUNT_MASK GENMASK(PAGE_SHIFT - 1, 2)
+#define RMI_ADDR_RANGE_BLOCK_SIZE_MASK GENMASK(1, 0)
+
+#define RMI_ADDR_RANGE_BLOCK_SIZE(r) FIELD_GET(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, (r))
+#define RMI_ADDR_RANGE_COUNT(r) FIELD_GET(RMI_ADDR_RANGE_COUNT_MASK, (r))
+#define RMI_ADDR_RANGE_ADDR(r) ((r) & RMI_ADDR_RANGE_ADDR_MASK)
+#define RMI_ADDR_RANGE_STATE(r) FIELD_GET(RMI_ADDR_RANGE_STATE_MASK, (r))
+
+enum rmi_ripas {
+ RMI_EMPTY = 0,
+ RMI_RAM = 1,
+ RMI_DESTROYED = 2,
+ RMI_DEV = 3,
+};
+
+#define RMI_NO_MEASURE_CONTENT 0
+#define RMI_MEASURE_CONTENT 1
+
+#define RMI_FEATURE_REGISTER_0_S2OASZ GENMASK(40, 33)
+#define RMI_FEATURE_REGISTER_0_L0GPT_BLOCK_DELEGATE BIT(32)
+#define RMI_FEATURE_REGISTER_0_PMU_NUM_CTRS GENMASK(31, 27)
+#define RMI_FEATURE_REGISTER_0_PMU BIT(26)
+#define RMI_FEATURE_REGISTER_0_NUM_WPS GENMASK(25, 20)
+#define RMI_FEATURE_REGISTER_0_NUM_BPS GENMASK(19, 14)
+#define RMI_FEATURE_REGISTER_0_SVE_VL GENMASK(13, 10)
+#define RMI_FEATURE_REGISTER_0_SVE BIT(9)
+#define RMI_FEATURE_REGISTER_0_LPA2 BIT(8)
+#define RMI_FEATURE_REGISTER_0_S2SZ GENMASK(7, 0)
+
+#define RMI_FEATURE_REGISTER_1_PPS GENMASK(16, 14)
+#define RMI_FEATURE_REGISTER_1_L0GPTSZ GENMASK(13, 10)
+#define RMI_FEATURE_REGISTER_1_MAX_RECS_ORDER GENMASK(9, 6)
+#define RMI_FEATURE_REGISTER_1_HASH_SHA_512 BIT(5)
+#define RMI_FEATURE_REGISTER_1_HASH_SHA_384 BIT(4)
+#define RMI_FEATURE_REGISTER_1_HASH_SHA_256 BIT(3)
+#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_64KB BIT(2)
+#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_16KB BIT(1)
+#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_4KB BIT(0)
+
+#define RMI_FEATURE_REGISTER_2_REALM_MAX_VDEVS_ORDER GENMASK(14, 10)
+#define RMI_FEATURE_REGISTER_2_NON_TEE_STREAM BIT(9)
+#define RMI_FEATURE_REGISTER_2_VDEV_KROU BIT(8)
+#define RMI_FEATURE_REGISTER_2_PDEV_MAX_VDEVS_ORDER GENMASK(7, 4)
+#define RMI_FEATURE_REGISTER_2_ATS BIT(3)
+#define RMI_FEATURE_REGISTER_2_VSMMU BIT(2)
+#define RMI_FEATURE_REGISTER_2_DA_COH BIT(1)
+#define RMI_FEATURE_REGISTER_2_DA BIT(0)
+
+#define RMI_FEATURE_REGISTER_3_RTT_S2AP_INDIRECT BIT(6)
+#define RMI_FEATURE_REGISTER_3_RTT_PLANE GENMASK(5, 4)
+#define RMI_FEATURE_REGISTER_3_MAX_NUM_AUX_PLANES GENMASK(3, 0)
+
+#define RMI_FEATURE_REGISTER_4_MEC_COUNT GENMASK(63, 0)
+
+#define RMI_MEM_CATEGORY_CONVENTIONAL 0
+#define RMI_MEM_CATEGORY_DEV_NCOH 1
+#define RMI_MEM_CATEGORY_DEV_COH 2
+#define RMI_MEM_CATEGORY_NONE 3
+
+#define RMI_TRACKING_RESERVED 0
+#define RMI_TRACKING_NONE 1
+#define RMI_TRACKING_FINE 2
+#define RMI_TRACKING_COARSE 3
+#define RMI_TRACKING_INTERMEDIATE 4
+
+#define RMI_GRANULE_SIZE_4KB 0
+#define RMI_GRANULE_SIZE_16KB 1
+#define RMI_GRANULE_SIZE_64KB 2
+
+#define RMI_GPT_PAR_RESERVED 0U
+#define RMI_GPT_PAR_PLAT 1U
+#define RMI_GPT_PAR_HOST_NOT_CREATED 2U
+#define RMI_GPT_PAR_HOST_CREATED 3U
+
+/*
+ * Note many of these fields are smaller than u64 but all fields have u64
+ * alignment, so use u64 to ensure correct alignment.
+ */
+struct rmm_config {
+ union { /* 0x0 */
+ struct {
+ u64 tracking_region_size;
+ u64 rmi_granule_size;
+ };
+ u8 sizer[SZ_4K];
+ };
+};
+
+static_assert(sizeof(struct rmm_config) == SZ_4K);
+
+#define RMI_REALM_PARAM_FLAG_SVE BIT(1)
+#define RMI_REALM_PARAM_FLAG_PMU BIT(2)
+#define RMI_REALM_PARAM_FLAG_DA BIT(3)
+#define RMI_REALM_PARAM_FLAG_LFA_POLICY GENMASK(6, 5)
+#define RMI_REALM_PARAM_FLAG_MEC_POLICY GENMASK(8, 7)
+
+#define RMI_HASH_SHA_256 0
+#define RMI_HASH_SHA_512 1
+#define RMI_HASH_SHA_384 2
+
+struct realm_params {
+ union { /* 0x0 */
+ struct {
+ u64 flags0;
+ u64 s2sz;
+ u64 sve_vl;
+ u64 num_bps;
+ u64 num_wps;
+ u64 pmu_num_ctrs;
+ u64 hash_algo;
+ u64 num_aux_planes;
+ };
+ u8 padding0[0x400];
+ };
+ union { /* 0x400 */
+ struct {
+ u8 rpv[64];
+ u64 ats_plane;
+ };
+ u8 padding1[0x400];
+ };
+ union { /* 0x800 */
+ struct {
+ u64 padding2;
+ u64 rtt_base;
+ s64 rtt_level_start;
+ u64 rtt_num_start;
+ u64 flags1;
+ u64 max_num_vdevs;
+ };
+ u8 padding3[0x700];
+ };
+ union { /* 0xf00 */
+ struct {
+ u8 padding4[0x80];
+ u64 aux_rtt_base[3];
+ };
+ u8 padding5[0x100];
+ };
+};
+
+static_assert(sizeof(struct realm_params) == SZ_4K);
+
+/*
+ * The number of GPRs (starting from X0) that are
+ * configured by the host when a REC is created.
+ */
+#define REC_CREATE_NR_GPRS 8
+
+#define REC_PARAMS_FLAG_RUNNABLE BIT(0)
+
+struct rec_params {
+ union { /* 0x0 */
+ u64 flags;
+ u8 padding0[0x100];
+ };
+ union { /* 0x100 */
+ u64 mpidr;
+ u8 padding1[0x100];
+ };
+ union { /* 0x200 */
+ u64 pc;
+ u8 padding2[0x100];
+ };
+ union { /* 0x300 */
+ u64 gprs[REC_CREATE_NR_GPRS];
+ u8 padding3[0xd00];
+ };
+};
+
+static_assert(sizeof(struct rec_params) == SZ_4K);
+
+#define REC_ENTER_FLAG_EMULATED_MMIO BIT(0)
+#define REC_ENTER_FLAG_INJECT_SEA BIT(1)
+#define REC_ENTER_FLAG_TRAP_WFI BIT(2)
+#define REC_ENTER_FLAG_TRAP_WFE BIT(3)
+#define REC_ENTER_FLAG_RIPAS_RESPONSE BIT(4)
+#define REC_ENTER_FLAG_S2AP_RESPONSE BIT(5)
+#define REC_ENTER_FLAG_DEV_MEM_RESPONSE BIT(6)
+#define REC_ENTER_FLAG_FORCE_P0 BIT(7)
+
+#define REC_RUN_GPRS 31
+
+struct rec_enter {
+ union { /* 0x000 */
+ u64 flags;
+ u8 padding0[0x200];
+ };
+ union { /* 0x200 */
+ u64 gprs[REC_RUN_GPRS];
+ u8 padding1[0x600];
+ };
+};
+
+static_assert(sizeof(struct rec_enter) == SZ_2K);
+
+#define RMI_EXIT_SYNC 0x00
+#define RMI_EXIT_IRQ 0x01
+#define RMI_EXIT_FIQ 0x02
+#define RMI_EXIT_PSCI 0x03
+#define RMI_EXIT_RIPAS_CHANGE 0x04
+#define RMI_EXIT_HOST_CALL 0x05
+#define RMI_EXIT_SERROR 0x06
+#define RMI_EXIT_S2AP_CHANGE 0x07
+#define RMI_EXIT_VDEV_VALIDATE_MAPPING 0x08
+#define RMI_EXIT_VSMMU_COMMAND 0x0a
+
+struct rec_exit {
+ union { /* 0x000 */
+ u8 exit_reason;
+ u8 padding0[0x100];
+ };
+ union { /* 0x100 */
+ struct {
+ u64 esr;
+ u64 far;
+ u64 hpfar;
+ u64 rtt_tree;
+ };
+ u8 padding1[0x100];
+ };
+ union { /* 0x200 */
+ u64 gprs[REC_RUN_GPRS];
+ u8 padding2[0x100];
+ };
+ union { /* 0x300 */
+ u8 padding3[0x100];
+ };
+ union { /* 0x400 */
+ struct {
+ u64 cntp_ctl;
+ u64 cntp_cval;
+ u64 cntv_ctl;
+ u64 cntv_cval;
+ };
+ u8 padding4[0x100];
+ };
+ union { /* 0x500 */
+ struct {
+ u64 ripas_base;
+ u64 ripas_top;
+ u8 ripas_value;
+ u8 padding5[0xf];
+ u64 s2ap_base;
+ u64 s2ap_top;
+ u64 vdev_id_1;
+ u64 vdev_id_2;
+ u64 dev_mem_base;
+ u64 dev_mem_top;
+ u64 dev_mem_pa;
+ };
+ u8 padding6[0x100];
+ };
+ union { /* 0x600 */
+ struct {
+ u16 imm;
+ u8 padding7[0x6];
+ u64 plane;
+ };
+ u8 padding8[0x100];
+ };
+ union { /* 0x700 */
+ struct {
+ u8 pmu_ovf_status;
+ u8 padding9[0xf];
+ u64 vsmmu;
+ };
+ u8 padding10[0x100];
+ };
+};
+
+static_assert(sizeof(struct rec_exit) == SZ_2K);
+
+struct rec_run {
+ struct rec_enter enter;
+ struct rec_exit exit;
+};
+
+static_assert(sizeof(struct rec_run) == SZ_4K);
+
+/* RMI_RTT_UNPROT_MAP_FLAGS definitions */
+#define RMI_RTT_UNPROT_MAP_FLAGS_OADDR_TYPE GENMASK(1, 0)
+#define RMI_RTT_UNPROT_MAP_FLAGS_LIST_COUNT GENMASK(15, 2)
+#define RMI_RTT_UNPROT_MAP_FLAGS_MEMATTR GENMASK(18, 16)
+#define RMI_RTT_UNPROT_MAP_FLAGS_S2AP GENMASK(22, 19)
+
+/* RMI_RTT_PROT_MAP_FLAGS definitions */
+#define RMI_RTT_PROT_MAP_FLAGS_OADDR_TYPE GENMASK(1, 0)
+#define RMI_RTT_PROT_MAP_FLAGS_LIST_COUNT GENMASK(15, 2)
+
+/* S2AP Direct Encodings, used in RMI_RTT_UNPROT_MAP_FLAGS_S2AP */
+#define RMI_S2AP_DIRECT_WRITE BIT(0)
+#define RMI_S2AP_DIRECT_READ BIT(1)
+
+#endif /* __LINUX_ARM_SMCCC_RMI_H_ */
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 1/7] firmware: arm_rmm: Add SMC definitions for calling the RMM
2026-09-12 8:36 ` [PATCH v18 1/7] firmware: arm_rmm: Add SMC definitions for calling the RMM Suzuki K Poulose
@ 2026-09-14 0:28 ` Gavin Shan
0 siblings, 0 replies; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 0:28 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> The RMM (Realm Management Monitor) provides functionality that can be
> accessed by SMC calls from the host.
>
> The SMC definitions are based on DEN0137[1] version 2.0-bet3
>
> [1] https://developer.arm.com/documentation/den0137/2-0bet3/
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v17:
> * Use GENMASK()/BIT() for masks consistently
> * Rename RMI_{ADDR_RANGE, DONATE}_SIZE => RMI_{*}_BLOCK_SIZE
> * Reorder the definitions for MSB to LSB
> * Add definions for RMI_OP_MEM_*CONTIG and RMI_OP_CAN*_CANCEL
> Changes since v16:
> * Updated definitions to RMM specification v2.0-bet3.
> Changes since v15:
> * Dropped unused symbols REC_MAX_GIC_NUM_LRS and RMI_PERMITTED_GICV3_HCR_BITS.
> * Output is now (partially) generated from the spec source.
> Changes since v14:
> * Updated to RMM spec v2.0-bet2 but without the changes to move
> metadata out of individual address range descriptors as this is
> expected to be reverted in a future spec release.
> Changes since v13:
> * Updated to RMM spec v2.0-bet1
> Changes since v12:
> * Updated to RMM spec v2.0-bet0
> Changes since v9:
> * Corrected size of 'ripas_value' in struct rec_exit. The spec states
> this is an 8-bit type with padding afterwards (rather than a u64).
> Changes since v8:
> * Added RMI_PERMITTED_GICV3_HCR_BITS to define which bits the RMM
> permits to be modified.
> Changes since v6:
> * Renamed REC_ENTER_xxx defines to include 'FLAG' to make it obvious
> these are flag values.
> Changes since v5:
> * Sorted the SMC #defines by value.
> * Renamed SMI_RxI_CALL to SMI_RMI_CALL since the macro is only used for
> RMI calls.
> * Renamed REC_GIC_NUM_LRS to REC_MAX_GIC_NUM_LRS since the actual
> number of available list registers could be lower.
> * Provided a define for the reserved fields of FeatureRegister0.
> * Fix inconsistent names for padding fields.
> Changes since v4:
> * Update to point to final released RMM spec.
> * Minor rearrangements.
> Changes since v3:
> * Update to match RMM spec v1.0-rel0-rc1.
> Changes since v2:
> * Fix specification link.
> * Rename rec_entry->rec_enter to match spec.
> * Fix size of pmu_ovf_status to match spec.
> ---
> include/linux/arm-smccc-rmi.h | 497 ++++++++++++++++++++++++++++++++++
> 1 file changed, 497 insertions(+)
> create mode 100644 include/linux/arm-smccc-rmi.h
>
Two nitpicks below, with them addressed:
Reviewed-by: Gavin Shan <gshan@redhat.com>
> diff --git a/include/linux/arm-smccc-rmi.h b/include/linux/arm-smccc-rmi.h
> new file mode 100644
> index 0000000000000..214d6228dfc22
> --- /dev/null
> +++ b/include/linux/arm-smccc-rmi.h
> @@ -0,0 +1,497 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +/*
> + * Copyright (C) 2023-2026 ARM Ltd.
> + *
> + * The values and structures in this file are from the Realm Management Monitor
> + * specification (DEN0137) version 2.0-bet3:
> + * https://developer.arm.com/documentation/den0137/2-0bet3/
> + */
> +
> +#ifndef __LINUX_ARM_SMCCC_RMI_H_
> +#define __LINUX_ARM_SMCCC_RMI_H_
> +
> +#include <linux/arm-smccc.h>
> +#include <linux/bitfield.h>
> +#include <linux/bits.h>
> +#include <linux/build_bug.h>
> +#include <linux/sizes.h>
> +
> +#include <asm/page.h>
> +
> +#define SMC_RMI_CALL(func) \
> + ARM_SMCCC_CALL_VAL(ARM_SMCCC_FAST_CALL, \
> + ARM_SMCCC_SMC_64, \
> + ARM_SMCCC_OWNER_STANDARD, \
> + (func))
> +
> +#define SMC_RMI_VERSION SMC_RMI_CALL(0x0150)
> +
> +#define SMC_RMI_RTT_DATA_MAP_INIT SMC_RMI_CALL(0x0153)
> +
> +#define SMC_RMI_REALM_ACTIVATE SMC_RMI_CALL(0x0157)
> +#define SMC_RMI_REALM_CREATE SMC_RMI_CALL(0x0158)
> +#define SMC_RMI_REALM_DESTROY SMC_RMI_CALL(0x0159)
> +#define SMC_RMI_REC_CREATE SMC_RMI_CALL(0x015a)
> +#define SMC_RMI_REC_DESTROY SMC_RMI_CALL(0x015b)
> +#define SMC_RMI_REC_ENTER SMC_RMI_CALL(0x015c)
> +#define SMC_RMI_RTT_CREATE SMC_RMI_CALL(0x015d)
> +#define SMC_RMI_RTT_DESTROY SMC_RMI_CALL(0x015e)
> +
> +#define SMC_RMI_RTT_READ_ENTRY SMC_RMI_CALL(0x0161)
> +
> +#define SMC_RMI_RTT_DEV_VALIDATE SMC_RMI_CALL(0x0163)
> +#define SMC_RMI_PSCI_COMPLETE SMC_RMI_CALL(0x0164)
> +#define SMC_RMI_FEATURES SMC_RMI_CALL(0x0165)
> +#define SMC_RMI_RTT_FOLD SMC_RMI_CALL(0x0166)
> +
> +#define SMC_RMI_RTT_INIT_RIPAS SMC_RMI_CALL(0x0168)
> +#define SMC_RMI_RTT_SET_RIPAS SMC_RMI_CALL(0x0169)
> +#define SMC_RMI_VSMMU_CREATE SMC_RMI_CALL(0x016a)
> +#define SMC_RMI_VSMMU_DESTROY SMC_RMI_CALL(0x016b)
> +
> +#define SMC_RMI_RMM_CONFIG_SET SMC_RMI_CALL(0x016e)
> +#define SMC_RMI_PSMMU_IRQ_NOTIFY SMC_RMI_CALL(0x016f)
> +#define SMC_RMI_ATTEST_PLAT_TOKEN_REFRESH SMC_RMI_CALL(0x0170)
> +
> +#define SMC_RMI_PDEV_ABORT SMC_RMI_CALL(0x0174)
> +#define SMC_RMI_PDEV_COMMUNICATE SMC_RMI_CALL(0x0175)
> +#define SMC_RMI_PDEV_CREATE SMC_RMI_CALL(0x0176)
> +#define SMC_RMI_PDEV_DESTROY SMC_RMI_CALL(0x0177)
> +#define SMC_RMI_PDEV_GET_STATE SMC_RMI_CALL(0x0178)
> +
> +#define SMC_RMI_PDEV_STREAM_KEY_REFRESH SMC_RMI_CALL(0x017a)
> +#define SMC_RMI_PDEV_SET_PUBKEY SMC_RMI_CALL(0x017b)
> +#define SMC_RMI_PDEV_STOP SMC_RMI_CALL(0x017c)
> +#define SMC_RMI_RTT_AUX_CREATE SMC_RMI_CALL(0x017d)
> +#define SMC_RMI_RTT_AUX_DESTROY SMC_RMI_CALL(0x017e)
> +#define SMC_RMI_RTT_AUX_FOLD SMC_RMI_CALL(0x017f)
> +
> +#define SMC_RMI_VDEV_ABORT SMC_RMI_CALL(0x0185)
> +#define SMC_RMI_VDEV_COMMUNICATE SMC_RMI_CALL(0x0186)
> +#define SMC_RMI_VDEV_CREATE SMC_RMI_CALL(0x0187)
> +#define SMC_RMI_VDEV_DESTROY SMC_RMI_CALL(0x0188)
> +#define SMC_RMI_VDEV_GET_STATE SMC_RMI_CALL(0x0189)
> +#define SMC_RMI_VDEV_UNLOCK SMC_RMI_CALL(0x018a)
> +#define SMC_RMI_RTT_SET_S2AP SMC_RMI_CALL(0x018b)
> +
> +#define SMC_RMI_VDEV_GET_INTERFACE_REPORT SMC_RMI_CALL(0x01d0)
> +#define SMC_RMI_VDEV_GET_MEASUREMENTS SMC_RMI_CALL(0x01d1)
> +#define SMC_RMI_VDEV_LOCK SMC_RMI_CALL(0x01d2)
> +#define SMC_RMI_VDEV_START SMC_RMI_CALL(0x01d3)
> +
> +#define SMC_RMI_VSMMU_EVENT_HANDLE SMC_RMI_CALL(0x01d6)
> +#define SMC_RMI_PSMMU_ACTIVATE SMC_RMI_CALL(0x01d7)
> +#define SMC_RMI_PSMMU_DEACTIVATE SMC_RMI_CALL(0x01d8)
> +
> +#define SMC_RMI_PSMMU_ST_L2_CREATE SMC_RMI_CALL(0x01db)
> +#define SMC_RMI_PSMMU_ST_L2_DESTROY SMC_RMI_CALL(0x01dc)
> +#define SMC_RMI_DPT_L0_CREATE SMC_RMI_CALL(0x01dd)
> +#define SMC_RMI_DPT_L0_DESTROY SMC_RMI_CALL(0x01de)
> +#define SMC_RMI_DPT_L1_CREATE SMC_RMI_CALL(0x01df)
> +#define SMC_RMI_DPT_L1_DESTROY SMC_RMI_CALL(0x01e0)
> +#define SMC_RMI_GRANULE_TRACKING_GET SMC_RMI_CALL(0x01e1)
> +
> +#define SMC_RMI_GRANULE_TRACKING_SET SMC_RMI_CALL(0x01e3)
> +
> +#define SMC_RMI_RMM_CONFIG_GET SMC_RMI_CALL(0x01ec)
> +
> +#define SMC_RMI_RMM_STATE_GET SMC_RMI_CALL(0x01ee)
> +
> +#define SMC_RMI_PSMMU_EVENT_CONSUME SMC_RMI_CALL(0x01f0)
> +#define SMC_RMI_GRANULE_RANGE_DELEGATE SMC_RMI_CALL(0x01f1)
> +#define SMC_RMI_GRANULE_RANGE_UNDELEGATE SMC_RMI_CALL(0x01f2)
> +#define SMC_RMI_GPT_L1_CREATE SMC_RMI_CALL(0x01f3)
> +#define SMC_RMI_GPT_L1_DESTROY SMC_RMI_CALL(0x01f4)
> +#define SMC_RMI_RTT_DATA_MAP SMC_RMI_CALL(0x01f5)
> +#define SMC_RMI_RTT_DATA_UNMAP SMC_RMI_CALL(0x01f6)
> +#define SMC_RMI_RTT_DEV_MAP SMC_RMI_CALL(0x01f7)
> +#define SMC_RMI_RTT_DEV_UNMAP SMC_RMI_CALL(0x01f8)
> +#define SMC_RMI_RTT_ARCH_DEV_MAP SMC_RMI_CALL(0x01f9)
> +#define SMC_RMI_RTT_ARCH_DEV_UNMAP SMC_RMI_CALL(0x01fa)
> +#define SMC_RMI_RTT_UNPROT_MAP SMC_RMI_CALL(0x01fb)
> +#define SMC_RMI_RTT_UNPROT_UNMAP SMC_RMI_CALL(0x01fc)
> +#define SMC_RMI_RTT_AUX_PROT_MAP SMC_RMI_CALL(0x01fd)
> +#define SMC_RMI_RTT_AUX_PROT_UNMAP SMC_RMI_CALL(0x01fe)
> +#define SMC_RMI_RTT_AUX_UNPROT_MAP SMC_RMI_CALL(0x01ff)
> +#define SMC_RMI_RTT_AUX_UNPROT_UNMAP SMC_RMI_CALL(0x0200)
> +#define SMC_RMI_REALM_TERMINATE SMC_RMI_CALL(0x0201)
> +#define SMC_RMI_RMM_ACTIVATE SMC_RMI_CALL(0x0202)
> +#define SMC_RMI_OP_CONTINUE SMC_RMI_CALL(0x0203)
> +#define SMC_RMI_PDEV_STREAM_CONNECT SMC_RMI_CALL(0x0204)
> +#define SMC_RMI_PDEV_STREAM_DISCONNECT SMC_RMI_CALL(0x0205)
> +#define SMC_RMI_PDEV_STREAM_COMPLETE SMC_RMI_CALL(0x0206)
> +#define SMC_RMI_PDEV_STREAM_KEY_PURGE SMC_RMI_CALL(0x0207)
> +#define SMC_RMI_OP_MEM_DONATE SMC_RMI_CALL(0x0208)
> +#define SMC_RMI_OP_MEM_RECLAIM SMC_RMI_CALL(0x0209)
> +#define SMC_RMI_OP_CANCEL SMC_RMI_CALL(0x020a)
> +#define SMC_RMI_VSMMU_FEATURES SMC_RMI_CALL(0x020b)
> +#define SMC_RMI_VSMMU_CMD_GET SMC_RMI_CALL(0x020c)
> +#define SMC_RMI_VSMMU_CMD_COMPLETE SMC_RMI_CALL(0x020d)
> +#define SMC_RMI_PSMMU_INFO SMC_RMI_CALL(0x020e)
> +#define SMC_RMI_RMM_DEACTIVATE SMC_RMI_CALL(0x020f)
> +#define SMC_RMI_PDEV_STREAM_INFO SMC_RMI_CALL(0x0210)
> +#define SMC_RMI_GPT_INFO SMC_RMI_CALL(0x0211)
> +
> +#define RMI_ABI_MAJOR_VERSION 2
> +#define RMI_ABI_MINOR_VERSION 0
> +
> +#define RMI_ABI_VERSION_GET_MAJOR(version) ((version) >> 16)
> +#define RMI_ABI_VERSION_GET_MINOR(version) ((version) & 0xFFFF)
> +#define RMI_ABI_VERSION(major, minor) (((major) << 16) | (minor))
> +
> +#define RMI_RETURN_STATUS_MASK GENMASK(7, 0)
> +#define RMI_RETURN_INDEX_MASK GENMASK(15, 8)
> +#define RMI_RETURN_MEMREQ_MASK GENMASK(9, 8)
> +#define RMI_RETURN_CAN_CANCEL_MASK BIT(10)
> +
> +#define RMI_RETURN_STATUS(ret) FIELD_GET(RMI_RETURN_STATUS_MASK, ret)
> +#define RMI_RETURN_INDEX(ret) FIELD_GET(RMI_RETURN_INDEX_MASK, ret)
> +#define RMI_RETURN_MEMREQ(ret) FIELD_GET(RMI_RETURN_MEMREQ_MASK, ret)
> +#define RMI_RETURN_CAN_CANCEL(ret) FIELD_GET(RMI_RETURN_CAN_CANCEL_MASK, ret)
> +
> +#define RMI_SUCCESS 0
> +#define RMI_ERROR_INPUT 1
> +#define RMI_ERROR_REALM 2
> +#define RMI_ERROR_REC 3
> +#define RMI_ERROR_RTT 4
> +#define RMI_ERROR_NOT_SUPPORTED 5
> +#define RMI_ERROR_DEVICE 6
> +#define RMI_ERROR_RTT_AUX 7
> +#define RMI_ERROR_PSMMU_ST 8
> +#define RMI_ERROR_DPT 9
> +#define RMI_BUSY 10
> +#define RMI_ERROR_GLOBAL 11
> +#define RMI_ERROR_TRACKING 12
> +#define RMI_INCOMPLETE 13
> +#define RMI_BLOCKED 14
> +#define RMI_ERROR_GPT 15
> +#define RMI_ERROR_GRANULE 16
> +
> +#define RMI_CONTINUE_KEEP_GOING 0
> +#define RMI_CONTINUE_STOP 1
> +
> +#define RMI_OP_MEM_REQ_NONE 0
> +#define RMI_OP_MEM_REQ_DONATE 1
> +#define RMI_OP_MEM_REQ_RECLAIM 2
> +
> +#define RMI_OP_CANNOT_CANCEL 0
> +#define RMI_OP_CAN_CANCEL 1
> +
> +#define RMI_DONATE_STATE_MASK GENMASK(18, 17)
> +#define RMI_DONATE_CONTIG_MASK BIT(16)
> +#define RMI_DONATE_COUNT_MASK GENMASK(15, 2)
> +#define RMI_DONATE_BLOCK_SIZE_MASK GENMASK(1, 0)
> +
> +#define RMI_DONATE_STATE(req) FIELD_GET(RMI_DONATE_STATE_MASK, req)
> +#define RMI_DONATE_CONTIG(req) FIELD_GET(RMI_DONATE_CONTIG_MASK, req)
> +#define RMI_DONATE_COUNT(req) FIELD_GET(RMI_DONATE_COUNT_MASK, req)
> +#define RMI_DONATE_BLOCK_SIZE(req) FIELD_GET(RMI_DONATE_BLOCK_SIZE_MASK, req)
> +
> +#define RMI_OP_MEM_DELEGATED 0
> +#define RMI_OP_MEM_UNDELEGATED 1
> +#define RMI_OP_MEM_CONDITIONAL 2
> +
> +#define RMI_OP_MEM_NON_CONTIG 0
> +#define RMI_OP_MEM_CONTIG 1
> +
> +#define RMI_ADDR_TYPE_NONE 0
> +#define RMI_ADDR_TYPE_SINGLE 1
> +#define RMI_ADDR_TYPE_LIST 2
> +
> +#define RMI_ADDR_RANGE_STATE_MASK GENMASK(63, 62)
> +#define RMI_ADDR_RANGE_ADDR_MASK GENMASK(51, PAGE_SHIFT)
> +#define RMI_ADDR_RANGE_COUNT_MASK GENMASK(PAGE_SHIFT - 1, 2)
> +#define RMI_ADDR_RANGE_BLOCK_SIZE_MASK GENMASK(1, 0)
> +
> +#define RMI_ADDR_RANGE_BLOCK_SIZE(r) FIELD_GET(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, (r))
> +#define RMI_ADDR_RANGE_COUNT(r) FIELD_GET(RMI_ADDR_RANGE_COUNT_MASK, (r))
> +#define RMI_ADDR_RANGE_ADDR(r) ((r) & RMI_ADDR_RANGE_ADDR_MASK)
> +#define RMI_ADDR_RANGE_STATE(r) FIELD_GET(RMI_ADDR_RANGE_STATE_MASK, (r))
> +
Please reorder those 4 definitions so that their orders are same to those
for RMI_ADDR_RANGE_{STATE, ADDR, COUNT, BLOCK_SIZE}_MASK.
#define RMI_ADDR_RANGE_STATE(r) FIELD_GET(RMI_ADDR_RANGE_STATE_MASK, (r))
#define RMI_ADDR_RANGE_ADDR(r) ((r) & RMI_ADDR_RANGE_ADDR_MASK)
#define RMI_ADDR_RANGE_COUNT(r) FIELD_GET(RMI_ADDR_RANGE_COUNT_MASK, (r))
#define RMI_ADDR_RANGE_BLOCK_SIZE(r) FIELD_GET(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, (r))
> +enum rmi_ripas {
> + RMI_EMPTY = 0,
> + RMI_RAM = 1,
> + RMI_DESTROYED = 2,
> + RMI_DEV = 3,
> +};
> +
> +#define RMI_NO_MEASURE_CONTENT 0
> +#define RMI_MEASURE_CONTENT 1
> +
> +#define RMI_FEATURE_REGISTER_0_S2OASZ GENMASK(40, 33)
> +#define RMI_FEATURE_REGISTER_0_L0GPT_BLOCK_DELEGATE BIT(32)
> +#define RMI_FEATURE_REGISTER_0_PMU_NUM_CTRS GENMASK(31, 27)
> +#define RMI_FEATURE_REGISTER_0_PMU BIT(26)
> +#define RMI_FEATURE_REGISTER_0_NUM_WPS GENMASK(25, 20)
> +#define RMI_FEATURE_REGISTER_0_NUM_BPS GENMASK(19, 14)
> +#define RMI_FEATURE_REGISTER_0_SVE_VL GENMASK(13, 10)
> +#define RMI_FEATURE_REGISTER_0_SVE BIT(9)
> +#define RMI_FEATURE_REGISTER_0_LPA2 BIT(8)
> +#define RMI_FEATURE_REGISTER_0_S2SZ GENMASK(7, 0)
> +
> +#define RMI_FEATURE_REGISTER_1_PPS GENMASK(16, 14)
> +#define RMI_FEATURE_REGISTER_1_L0GPTSZ GENMASK(13, 10)
> +#define RMI_FEATURE_REGISTER_1_MAX_RECS_ORDER GENMASK(9, 6)
> +#define RMI_FEATURE_REGISTER_1_HASH_SHA_512 BIT(5)
> +#define RMI_FEATURE_REGISTER_1_HASH_SHA_384 BIT(4)
> +#define RMI_FEATURE_REGISTER_1_HASH_SHA_256 BIT(3)
> +#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_64KB BIT(2)
> +#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_16KB BIT(1)
> +#define RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_4KB BIT(0)
> +
> +#define RMI_FEATURE_REGISTER_2_REALM_MAX_VDEVS_ORDER GENMASK(14, 10)
> +#define RMI_FEATURE_REGISTER_2_NON_TEE_STREAM BIT(9)
> +#define RMI_FEATURE_REGISTER_2_VDEV_KROU BIT(8)
> +#define RMI_FEATURE_REGISTER_2_PDEV_MAX_VDEVS_ORDER GENMASK(7, 4)
> +#define RMI_FEATURE_REGISTER_2_ATS BIT(3)
> +#define RMI_FEATURE_REGISTER_2_VSMMU BIT(2)
> +#define RMI_FEATURE_REGISTER_2_DA_COH BIT(1)
> +#define RMI_FEATURE_REGISTER_2_DA BIT(0)
> +
> +#define RMI_FEATURE_REGISTER_3_RTT_S2AP_INDIRECT BIT(6)
> +#define RMI_FEATURE_REGISTER_3_RTT_PLANE GENMASK(5, 4)
> +#define RMI_FEATURE_REGISTER_3_MAX_NUM_AUX_PLANES GENMASK(3, 0)
> +
> +#define RMI_FEATURE_REGISTER_4_MEC_COUNT GENMASK(63, 0)
> +
> +#define RMI_MEM_CATEGORY_CONVENTIONAL 0
> +#define RMI_MEM_CATEGORY_DEV_NCOH 1
> +#define RMI_MEM_CATEGORY_DEV_COH 2
> +#define RMI_MEM_CATEGORY_NONE 3
> +
> +#define RMI_TRACKING_RESERVED 0
> +#define RMI_TRACKING_NONE 1
> +#define RMI_TRACKING_FINE 2
> +#define RMI_TRACKING_COARSE 3
> +#define RMI_TRACKING_INTERMEDIATE 4
> +
> +#define RMI_GRANULE_SIZE_4KB 0
> +#define RMI_GRANULE_SIZE_16KB 1
> +#define RMI_GRANULE_SIZE_64KB 2
> +
> +#define RMI_GPT_PAR_RESERVED 0U
> +#define RMI_GPT_PAR_PLAT 1U
> +#define RMI_GPT_PAR_HOST_NOT_CREATED 2U
> +#define RMI_GPT_PAR_HOST_CREATED 3U
> +
> +/*
> + * Note many of these fields are smaller than u64 but all fields have u64
> + * alignment, so use u64 to ensure correct alignment.
> + */
> +struct rmm_config {
> + union { /* 0x0 */
> + struct {
> + u64 tracking_region_size;
> + u64 rmi_granule_size;
> + };
> + u8 sizer[SZ_4K];
> + };
> +};
> +
> +static_assert(sizeof(struct rmm_config) == SZ_4K);
> +
> +#define RMI_REALM_PARAM_FLAG_SVE BIT(1)
> +#define RMI_REALM_PARAM_FLAG_PMU BIT(2)
> +#define RMI_REALM_PARAM_FLAG_DA BIT(3)
> +#define RMI_REALM_PARAM_FLAG_LFA_POLICY GENMASK(6, 5)
> +#define RMI_REALM_PARAM_FLAG_MEC_POLICY GENMASK(8, 7)
> +
> +#define RMI_HASH_SHA_256 0
> +#define RMI_HASH_SHA_512 1
> +#define RMI_HASH_SHA_384 2
> +
> +struct realm_params {
> + union { /* 0x0 */
> + struct {
> + u64 flags0;
> + u64 s2sz;
> + u64 sve_vl;
> + u64 num_bps;
> + u64 num_wps;
> + u64 pmu_num_ctrs;
> + u64 hash_algo;
> + u64 num_aux_planes;
> + };
> + u8 padding0[0x400];
> + };
> + union { /* 0x400 */
> + struct {
> + u8 rpv[64];
> + u64 ats_plane;
> + };
> + u8 padding1[0x400];
> + };
> + union { /* 0x800 */
> + struct {
> + u64 padding2;
> + u64 rtt_base;
> + s64 rtt_level_start;
> + u64 rtt_num_start;
> + u64 flags1;
> + u64 max_num_vdevs;
> + };
> + u8 padding3[0x700];
> + };
> + union { /* 0xf00 */
> + struct {
> + u8 padding4[0x80];
> + u64 aux_rtt_base[3];
> + };
> + u8 padding5[0x100];
> + };
> +};
> +
> +static_assert(sizeof(struct realm_params) == SZ_4K);
> +
> +/*
> + * The number of GPRs (starting from X0) that are
> + * configured by the host when a REC is created.
> + */
> +#define REC_CREATE_NR_GPRS 8
> +
There are more space in the first line of the comments.
/*
* The number of GPRs (starting from X0) that are configured by the host when
* a REC is created.
*/
> +#define REC_PARAMS_FLAG_RUNNABLE BIT(0)
> +
> +struct rec_params {
> + union { /* 0x0 */
> + u64 flags;
> + u8 padding0[0x100];
> + };
> + union { /* 0x100 */
> + u64 mpidr;
> + u8 padding1[0x100];
> + };
> + union { /* 0x200 */
> + u64 pc;
> + u8 padding2[0x100];
> + };
> + union { /* 0x300 */
> + u64 gprs[REC_CREATE_NR_GPRS];
> + u8 padding3[0xd00];
> + };
> +};
> +
> +static_assert(sizeof(struct rec_params) == SZ_4K);
> +
> +#define REC_ENTER_FLAG_EMULATED_MMIO BIT(0)
> +#define REC_ENTER_FLAG_INJECT_SEA BIT(1)
> +#define REC_ENTER_FLAG_TRAP_WFI BIT(2)
> +#define REC_ENTER_FLAG_TRAP_WFE BIT(3)
> +#define REC_ENTER_FLAG_RIPAS_RESPONSE BIT(4)
> +#define REC_ENTER_FLAG_S2AP_RESPONSE BIT(5)
> +#define REC_ENTER_FLAG_DEV_MEM_RESPONSE BIT(6)
> +#define REC_ENTER_FLAG_FORCE_P0 BIT(7)
> +
> +#define REC_RUN_GPRS 31
> +
> +struct rec_enter {
> + union { /* 0x000 */
> + u64 flags;
> + u8 padding0[0x200];
> + };
> + union { /* 0x200 */
> + u64 gprs[REC_RUN_GPRS];
> + u8 padding1[0x600];
> + };
> +};
> +
> +static_assert(sizeof(struct rec_enter) == SZ_2K);
> +
> +#define RMI_EXIT_SYNC 0x00
> +#define RMI_EXIT_IRQ 0x01
> +#define RMI_EXIT_FIQ 0x02
> +#define RMI_EXIT_PSCI 0x03
> +#define RMI_EXIT_RIPAS_CHANGE 0x04
> +#define RMI_EXIT_HOST_CALL 0x05
> +#define RMI_EXIT_SERROR 0x06
> +#define RMI_EXIT_S2AP_CHANGE 0x07
> +#define RMI_EXIT_VDEV_VALIDATE_MAPPING 0x08
> +#define RMI_EXIT_VSMMU_COMMAND 0x0a
> +
> +struct rec_exit {
> + union { /* 0x000 */
> + u8 exit_reason;
> + u8 padding0[0x100];
> + };
> + union { /* 0x100 */
> + struct {
> + u64 esr;
> + u64 far;
> + u64 hpfar;
> + u64 rtt_tree;
> + };
> + u8 padding1[0x100];
> + };
> + union { /* 0x200 */
> + u64 gprs[REC_RUN_GPRS];
> + u8 padding2[0x100];
> + };
> + union { /* 0x300 */
> + u8 padding3[0x100];
> + };
> + union { /* 0x400 */
> + struct {
> + u64 cntp_ctl;
> + u64 cntp_cval;
> + u64 cntv_ctl;
> + u64 cntv_cval;
> + };
> + u8 padding4[0x100];
> + };
> + union { /* 0x500 */
> + struct {
> + u64 ripas_base;
> + u64 ripas_top;
> + u8 ripas_value;
> + u8 padding5[0xf];
> + u64 s2ap_base;
> + u64 s2ap_top;
> + u64 vdev_id_1;
> + u64 vdev_id_2;
> + u64 dev_mem_base;
> + u64 dev_mem_top;
> + u64 dev_mem_pa;
> + };
> + u8 padding6[0x100];
> + };
> + union { /* 0x600 */
> + struct {
> + u16 imm;
> + u8 padding7[0x6];
> + u64 plane;
> + };
> + u8 padding8[0x100];
> + };
> + union { /* 0x700 */
> + struct {
> + u8 pmu_ovf_status;
> + u8 padding9[0xf];
> + u64 vsmmu;
> + };
> + u8 padding10[0x100];
> + };
> +};
> +
> +static_assert(sizeof(struct rec_exit) == SZ_2K);
> +
> +struct rec_run {
> + struct rec_enter enter;
> + struct rec_exit exit;
> +};
> +
> +static_assert(sizeof(struct rec_run) == SZ_4K);
> +
> +/* RMI_RTT_UNPROT_MAP_FLAGS definitions */
> +#define RMI_RTT_UNPROT_MAP_FLAGS_OADDR_TYPE GENMASK(1, 0)
> +#define RMI_RTT_UNPROT_MAP_FLAGS_LIST_COUNT GENMASK(15, 2)
> +#define RMI_RTT_UNPROT_MAP_FLAGS_MEMATTR GENMASK(18, 16)
> +#define RMI_RTT_UNPROT_MAP_FLAGS_S2AP GENMASK(22, 19)
> +
> +/* RMI_RTT_PROT_MAP_FLAGS definitions */
> +#define RMI_RTT_PROT_MAP_FLAGS_OADDR_TYPE GENMASK(1, 0)
> +#define RMI_RTT_PROT_MAP_FLAGS_LIST_COUNT GENMASK(15, 2)
> +
> +/* S2AP Direct Encodings, used in RMI_RTT_UNPROT_MAP_FLAGS_S2AP */
> +#define RMI_S2AP_DIRECT_WRITE BIT(0)
> +#define RMI_S2AP_DIRECT_READ BIT(1)
> +
> +#endif /* __LINUX_ARM_SMCCC_RMI_H_ */
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
2026-09-12 8:36 ` [PATCH v18 1/7] firmware: arm_rmm: Add SMC definitions for calling the RMM Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 1:04 ` Gavin Shan
2026-09-14 10:27 ` Sudeep Holla
2026-09-12 8:36 ` [PATCH v18 3/7] firmware: arm_rmm: Configure the RMM with the host's page size Suzuki K Poulose
` (4 subsequent siblings)
6 siblings, 2 replies; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Query the RMI version number and check if it is a compatible version.
The first two feature registers are read and exposed for future code to
use.
We only support this for Little Endian kernels, the Big Endian kernel
support is anyway marked BROKEN and is being removed.
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
v18:
* Always use arm_smccc_1_2_invoke() for all RMIs making sure the unsused
parameters are 0 - Sashiko
* Move rmi_features() calls away from the arm-rmi-cmds.h to rmi.c - Gavin
v17:
* Rename ARM_RMM to ARM_RMM_RMI to make it easier to add Guest facing RSI
support, which is also in progress
v16:
* Update Kconfig text to include PCIe TDISP.
* Export rmi_feat_reg() here rather than in a later commit.
v15:
* The code is moved again, this time into the 'firmware' directory.
v14:
* This moves the basic RMI setup into the 'kernel' directory. This is
because RMI will be used for some features outside of KVM so should
be available even if KVM isn't compiled in.
---
arch/arm64/Kconfig | 1 +
arch/arm64/kernel/cpufeature.c | 1 +
drivers/firmware/Kconfig | 1 +
drivers/firmware/Makefile | 1 +
drivers/firmware/arm_rmm/Kconfig | 26 +++++++
drivers/firmware/arm_rmm/Makefile | 2 +
drivers/firmware/arm_rmm/rmi.c | 122 ++++++++++++++++++++++++++++++
include/linux/arm-rmi-cmds.h | 36 +++++++++
8 files changed, 190 insertions(+)
create mode 100644 drivers/firmware/arm_rmm/Kconfig
create mode 100644 drivers/firmware/arm_rmm/Makefile
create mode 100644 drivers/firmware/arm_rmm/rmi.c
create mode 100644 include/linux/arm-rmi-cmds.h
diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
index b5a51b0ef9440..ff9565d3ffa59 100644
--- a/arch/arm64/Kconfig
+++ b/arch/arm64/Kconfig
@@ -38,6 +38,7 @@ config ARM64
select ARCH_HAS_MEMBARRIER_SYNC_CORE
select ARCH_HAS_MEM_ENCRYPT
select ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS
+ select ARCH_SUPPORTS_RMM
select ARCH_HAS_NMI_SAFE_THIS_CPU_OPS
select ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE
select ARCH_HAS_NONLEAF_PMD_YOUNG if ARM64_HAFT
diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c
index 32102c3912fa7..e8b29983b0021 100644
--- a/arch/arm64/kernel/cpufeature.c
+++ b/arch/arm64/kernel/cpufeature.c
@@ -293,6 +293,7 @@ static const struct arm64_ftr_bits ftr_id_aa64isar3[] = {
static const struct arm64_ftr_bits ftr_id_aa64pfr0[] = {
ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_CSV3_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_CSV2_SHIFT, 4, 0),
+ ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_RME_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_DIT_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_AMU_SHIFT, 4, 0),
ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_MPAM_SHIFT, 4, 0),
diff --git a/drivers/firmware/Kconfig b/drivers/firmware/Kconfig
index b7cc11e4fbfa6..62660bf520a8d 100644
--- a/drivers/firmware/Kconfig
+++ b/drivers/firmware/Kconfig
@@ -310,5 +310,6 @@ source "drivers/firmware/samsung/Kconfig"
source "drivers/firmware/smccc/Kconfig"
source "drivers/firmware/tegra/Kconfig"
source "drivers/firmware/xilinx/Kconfig"
+source "drivers/firmware/arm_rmm/Kconfig"
endmenu
diff --git a/drivers/firmware/Makefile b/drivers/firmware/Makefile
index be46f1e1dc77f..196a650ccf025 100644
--- a/drivers/firmware/Makefile
+++ b/drivers/firmware/Makefile
@@ -39,3 +39,4 @@ obj-y += samsung/
obj-y += smccc/
obj-y += tegra/
obj-y += xilinx/
+obj-y += arm_rmm/
diff --git a/drivers/firmware/arm_rmm/Kconfig b/drivers/firmware/arm_rmm/Kconfig
new file mode 100644
index 0000000000000..82724ee23186d
--- /dev/null
+++ b/drivers/firmware/arm_rmm/Kconfig
@@ -0,0 +1,26 @@
+
+config ARCH_SUPPORTS_RMM
+ bool
+
+config ARM_RMM_RMI
+ bool "Realm Management Interface (RMI) Support"
+ depends on ARCH_SUPPORTS_RMM
+ default y
+ help
+ Support the Realm Management Monitor (RMM) on Arm systems that
+ implement the Realm Management Extension (RME), as defined by the
+ Arm Confidential Compute Architecture.
+
+ The RMM runs at EL2 in the Realm world and provides the Realm
+ Management Interface (RMI) used by a Normal World host to create,
+ manage and run protected virtual machines called Realms. The RMM can
+ also act as a TSM, as defined by the PCIe TDISP and can manage the
+ PCI IDE setup for securing the PCIe links.
+
+ This option builds the host-side RMI support used by KVM to detect a
+ compatible RMM, configure it, manage delegated memory and enable
+ Realm guests.
+
+ Selecting this option does not by itself make Realm guests available:
+ the system must also provide RME-capable hardware and firmware with a
+ compatible RMM implementation.
diff --git a/drivers/firmware/arm_rmm/Makefile b/drivers/firmware/arm_rmm/Makefile
new file mode 100644
index 0000000000000..65171988fdcae
--- /dev/null
+++ b/drivers/firmware/arm_rmm/Makefile
@@ -0,0 +1,2 @@
+
+obj-$(CONFIG_ARM_RMM_RMI) = rmi.o
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
new file mode 100644
index 0000000000000..2fd538c937dca
--- /dev/null
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -0,0 +1,122 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (C) 2023-2026 ARM Ltd.
+ */
+
+#include <linux/cpufeature.h>
+#include <linux/memblock.h>
+#include <linux/arm-rmi-cmds.h>
+#include <linux/slab.h>
+
+#include <asm/memory.h>
+#include <asm/pgtable-hwdef.h>
+
+/* Currently only the first 2 registers are used by Linux */
+#define RMI_FEAT_REG_COUNT 2
+static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT] __ro_after_init;
+
+/**
+ * rmi_features() - Read feature register
+ * @index: Feature register index
+ * @out: Feature register value is written to this pointer
+ *
+ * Return: RMI return code
+ */
+static int rmi_features(unsigned long index, unsigned long *out)
+{
+ struct arm_smccc_1_2_regs args = {
+ SMC_RMI_FEATURES, index
+ };
+
+ rmi_smccc_invoke(&args);
+ if (args.a0 == RMI_SUCCESS && out)
+ *out = args.a1;
+
+ return args.a0;
+}
+
+unsigned long rmi_feat_reg(unsigned long index)
+{
+ if (WARN_ON(index >= RMI_FEAT_REG_COUNT))
+ return 0;
+
+ return rmi_feat_reg_cache[index];
+}
+EXPORT_SYMBOL_GPL(rmi_feat_reg);
+
+static int rmi_check_version(void)
+{
+ unsigned short version_major, version_minor;
+ unsigned long host_version = RMI_ABI_VERSION(RMI_ABI_MAJOR_VERSION,
+ RMI_ABI_MINOR_VERSION);
+ unsigned long aa64pfr0 = read_sanitised_ftr_reg(SYS_ID_AA64PFR0_EL1);
+ struct arm_smccc_1_2_regs res = {
+ SMC_RMI_VERSION, host_version,
+ };
+
+ /* If RME isn't supported, then RMI can't be */
+ if (cpuid_feature_extract_unsigned_field(aa64pfr0, ID_AA64PFR0_EL1_RME_SHIFT) == 0)
+ return -ENXIO;
+
+ rmi_smccc_invoke(&res);
+ if (res.a0 == SMCCC_RET_NOT_SUPPORTED)
+ return -ENXIO;
+
+ version_major = RMI_ABI_VERSION_GET_MAJOR(res.a1);
+ version_minor = RMI_ABI_VERSION_GET_MINOR(res.a1);
+
+ if (res.a0 != RMI_SUCCESS) {
+ unsigned short high_version_major, high_version_minor;
+
+ high_version_major = RMI_ABI_VERSION_GET_MAJOR(res.a2);
+ high_version_minor = RMI_ABI_VERSION_GET_MINOR(res.a2);
+
+ pr_err("Unsupported RMI ABI (v%d.%d - v%d.%d) we want v%d.%d\n",
+ version_major, version_minor,
+ high_version_major, high_version_minor,
+ RMI_ABI_MAJOR_VERSION,
+ RMI_ABI_MINOR_VERSION);
+ return -ENXIO;
+ }
+
+ pr_info("RMI ABI version %d.%d\n", version_major, version_minor);
+
+ return 0;
+}
+
+static int rmi_read_features(void)
+{
+ /*
+ * Since we've negotiated a compatible version these feature registers
+ * should always be available
+ */
+ for (int i = 0; i < RMI_FEAT_REG_COUNT; i++) {
+ if (WARN_ON(rmi_features(i, &rmi_feat_reg_cache[i])))
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+static int __init arm64_init_rmi(void)
+{
+ int ret;
+
+ /* Continue without realm support if we can't agree on a version */
+ ret = rmi_check_version();
+ if (ret)
+ return ret;
+
+ ret = rmi_read_features();
+ if (ret)
+ return ret;
+
+ return 0;
+}
+
+/*
+ * Note arm64_init_rmi() must be called before kvm_init_rmi() otherwise KVM
+ * will not support realm guests. subsys_initcall() is called before
+ * module_init() (used for KVM) so this is OK.
+ */
+subsys_initcall(arm64_init_rmi);
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
new file mode 100644
index 0000000000000..9792bf0e00cb9
--- /dev/null
+++ b/include/linux/arm-rmi-cmds.h
@@ -0,0 +1,36 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (C) 2026 ARM Ltd.
+ */
+
+#ifndef __LINUX_ARM_RMI_CMDS_H_
+#define __LINUX_ARM_RMI_CMDS_H_
+
+#include <linux/arm-smccc-rmi.h>
+#include <linux/bug.h>
+#include <linux/processor.h>
+#include <linux/types.h>
+
+
+/*
+ * rmi_smccc_invoke: Invoke the RMI call and return the results in @regs_out
+ * @regs_in: Registers with the arguments filled in.
+ * @regs_out: Ouptput results from the call.
+ */
+static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
+{
+ struct arm_smccc_1_2_regs args = *regs;
+ unsigned long status;
+
+ while (1) {
+ arm_smccc_1_2_invoke(&args, regs);
+ status = RMI_RETURN_STATUS(regs->a0);
+ if (status != RMI_BUSY && status != RMI_BLOCKED)
+ break;
+ cpu_relax();
+ }
+}
+
+unsigned long rmi_feat_reg(unsigned long index);
+
+#endif
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init
2026-09-12 8:36 ` [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init Suzuki K Poulose
@ 2026-09-14 1:04 ` Gavin Shan
2026-09-14 10:27 ` Sudeep Holla
1 sibling, 0 replies; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 1:04 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> Query the RMI version number and check if it is a compatible version.
> The first two feature registers are read and exposed for future code to
> use.
>
> We only support this for Little Endian kernels, the Big Endian kernel
> support is anyway marked BROKEN and is being removed.
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> v18:
> * Always use arm_smccc_1_2_invoke() for all RMIs making sure the unsused
> parameters are 0 - Sashiko
> * Move rmi_features() calls away from the arm-rmi-cmds.h to rmi.c - Gavin
> v17:
> * Rename ARM_RMM to ARM_RMM_RMI to make it easier to add Guest facing RSI
> support, which is also in progress
> v16:
> * Update Kconfig text to include PCIe TDISP.
> * Export rmi_feat_reg() here rather than in a later commit.
> v15:
> * The code is moved again, this time into the 'firmware' directory.
> v14:
> * This moves the basic RMI setup into the 'kernel' directory. This is
> because RMI will be used for some features outside of KVM so should
> be available even if KVM isn't compiled in.
> ---
> arch/arm64/Kconfig | 1 +
> arch/arm64/kernel/cpufeature.c | 1 +
> drivers/firmware/Kconfig | 1 +
> drivers/firmware/Makefile | 1 +
> drivers/firmware/arm_rmm/Kconfig | 26 +++++++
> drivers/firmware/arm_rmm/Makefile | 2 +
> drivers/firmware/arm_rmm/rmi.c | 122 ++++++++++++++++++++++++++++++
> include/linux/arm-rmi-cmds.h | 36 +++++++++
> 8 files changed, 190 insertions(+)
> create mode 100644 drivers/firmware/arm_rmm/Kconfig
> create mode 100644 drivers/firmware/arm_rmm/Makefile
> create mode 100644 drivers/firmware/arm_rmm/rmi.c
> create mode 100644 include/linux/arm-rmi-cmds.h
>
> diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
> index b5a51b0ef9440..ff9565d3ffa59 100644
> --- a/arch/arm64/Kconfig
> +++ b/arch/arm64/Kconfig
> @@ -38,6 +38,7 @@ config ARM64
> select ARCH_HAS_MEMBARRIER_SYNC_CORE
> select ARCH_HAS_MEM_ENCRYPT
> select ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS
> + select ARCH_SUPPORTS_RMM
> select ARCH_HAS_NMI_SAFE_THIS_CPU_OPS
> select ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE
> select ARCH_HAS_NONLEAF_PMD_YOUNG if ARM64_HAFT
> diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c
> index 32102c3912fa7..e8b29983b0021 100644
> --- a/arch/arm64/kernel/cpufeature.c
> +++ b/arch/arm64/kernel/cpufeature.c
> @@ -293,6 +293,7 @@ static const struct arm64_ftr_bits ftr_id_aa64isar3[] = {
> static const struct arm64_ftr_bits ftr_id_aa64pfr0[] = {
> ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_CSV3_SHIFT, 4, 0),
> ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_CSV2_SHIFT, 4, 0),
> + ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_RME_SHIFT, 4, 0),
> ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_DIT_SHIFT, 4, 0),
> ARM64_FTR_BITS(FTR_HIDDEN, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_AMU_SHIFT, 4, 0),
> ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64PFR0_EL1_MPAM_SHIFT, 4, 0),
> diff --git a/drivers/firmware/Kconfig b/drivers/firmware/Kconfig
> index b7cc11e4fbfa6..62660bf520a8d 100644
> --- a/drivers/firmware/Kconfig
> +++ b/drivers/firmware/Kconfig
> @@ -310,5 +310,6 @@ source "drivers/firmware/samsung/Kconfig"
> source "drivers/firmware/smccc/Kconfig"
> source "drivers/firmware/tegra/Kconfig"
> source "drivers/firmware/xilinx/Kconfig"
> +source "drivers/firmware/arm_rmm/Kconfig"
>
> endmenu
> diff --git a/drivers/firmware/Makefile b/drivers/firmware/Makefile
> index be46f1e1dc77f..196a650ccf025 100644
> --- a/drivers/firmware/Makefile
> +++ b/drivers/firmware/Makefile
> @@ -39,3 +39,4 @@ obj-y += samsung/
> obj-y += smccc/
> obj-y += tegra/
> obj-y += xilinx/
> +obj-y += arm_rmm/
> diff --git a/drivers/firmware/arm_rmm/Kconfig b/drivers/firmware/arm_rmm/Kconfig
> new file mode 100644
> index 0000000000000..82724ee23186d
> --- /dev/null
> +++ b/drivers/firmware/arm_rmm/Kconfig
> @@ -0,0 +1,26 @@
> +
> +config ARCH_SUPPORTS_RMM
> + bool
> +
> +config ARM_RMM_RMI
> + bool "Realm Management Interface (RMI) Support"
> + depends on ARCH_SUPPORTS_RMM
> + default y
> + help
> + Support the Realm Management Monitor (RMM) on Arm systems that
> + implement the Realm Management Extension (RME), as defined by the
> + Arm Confidential Compute Architecture.
> +
> + The RMM runs at EL2 in the Realm world and provides the Realm
> + Management Interface (RMI) used by a Normal World host to create,
> + manage and run protected virtual machines called Realms. The RMM can
> + also act as a TSM, as defined by the PCIe TDISP and can manage the
> + PCI IDE setup for securing the PCIe links.
> +
> + This option builds the host-side RMI support used by KVM to detect a
> + compatible RMM, configure it, manage delegated memory and enable
> + Realm guests.
> +
> + Selecting this option does not by itself make Realm guests available:
> + the system must also provide RME-capable hardware and firmware with a
> + compatible RMM implementation.
> diff --git a/drivers/firmware/arm_rmm/Makefile b/drivers/firmware/arm_rmm/Makefile
> new file mode 100644
> index 0000000000000..65171988fdcae
> --- /dev/null
> +++ b/drivers/firmware/arm_rmm/Makefile
> @@ -0,0 +1,2 @@
> +
> +obj-$(CONFIG_ARM_RMM_RMI) = rmi.o
> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
> new file mode 100644
> index 0000000000000..2fd538c937dca
> --- /dev/null
> +++ b/drivers/firmware/arm_rmm/rmi.c
> @@ -0,0 +1,122 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * Copyright (C) 2023-2026 ARM Ltd.
> + */
> +
> +#include <linux/cpufeature.h>
> +#include <linux/memblock.h>
> +#include <linux/arm-rmi-cmds.h>
> +#include <linux/slab.h>
> +
> +#include <asm/memory.h>
> +#include <asm/pgtable-hwdef.h>
> +
> +/* Currently only the first 2 registers are used by Linux */
> +#define RMI_FEAT_REG_COUNT 2
> +static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT] __ro_after_init;
> +
> +/**
> + * rmi_features() - Read feature register
> + * @index: Feature register index
> + * @out: Feature register value is written to this pointer
> + *
> + * Return: RMI return code
> + */
> +static int rmi_features(unsigned long index, unsigned long *out)
> +{
> + struct arm_smccc_1_2_regs args = {
> + SMC_RMI_FEATURES, index
> + };
> +
> + rmi_smccc_invoke(&args);
> + if (args.a0 == RMI_SUCCESS && out)
> + *out = args.a1;
> +
> + return args.a0;
> +}
> +
The comments for rmi_features() is unnecessary since the code is self-explainning.
Besides, it needs to be moved to right before its only caller rmi_read_features().
I would drop this function by combining the logics into its only caller rmi_read_features(),
seeing below for more details.
> +unsigned long rmi_feat_reg(unsigned long index)
> +{
> + if (WARN_ON(index >= RMI_FEAT_REG_COUNT))
> + return 0;
> +
> + return rmi_feat_reg_cache[index];
> +}
> +EXPORT_SYMBOL_GPL(rmi_feat_reg);
> +
> +static int rmi_check_version(void)
> +{
> + unsigned short version_major, version_minor;
> + unsigned long host_version = RMI_ABI_VERSION(RMI_ABI_MAJOR_VERSION,
> + RMI_ABI_MINOR_VERSION);
> + unsigned long aa64pfr0 = read_sanitised_ftr_reg(SYS_ID_AA64PFR0_EL1);
> + struct arm_smccc_1_2_regs res = {
> + SMC_RMI_VERSION, host_version,
> + };
> +
> + /* If RME isn't supported, then RMI can't be */
> + if (cpuid_feature_extract_unsigned_field(aa64pfr0, ID_AA64PFR0_EL1_RME_SHIFT) == 0)
> + return -ENXIO;
> +
> + rmi_smccc_invoke(&res);
> + if (res.a0 == SMCCC_RET_NOT_SUPPORTED)
> + return -ENXIO;
> +
> + version_major = RMI_ABI_VERSION_GET_MAJOR(res.a1);
> + version_minor = RMI_ABI_VERSION_GET_MINOR(res.a1);
> +
> + if (res.a0 != RMI_SUCCESS) {
> + unsigned short high_version_major, high_version_minor;
> +
> + high_version_major = RMI_ABI_VERSION_GET_MAJOR(res.a2);
> + high_version_minor = RMI_ABI_VERSION_GET_MINOR(res.a2);
> +
> + pr_err("Unsupported RMI ABI (v%d.%d - v%d.%d) we want v%d.%d\n",
> + version_major, version_minor,
> + high_version_major, high_version_minor,
> + RMI_ABI_MAJOR_VERSION,
> + RMI_ABI_MINOR_VERSION);
> + return -ENXIO;
> + }
> +
> + pr_info("RMI ABI version %d.%d\n", version_major, version_minor);
> +
> + return 0;
> +}
> +
> +static int rmi_read_features(void)
> +{
> + /*
> + * Since we've negotiated a compatible version these feature registers
> + * should always be available
> + */
> + for (int i = 0; i < RMI_FEAT_REG_COUNT; i++) {
> + if (WARN_ON(rmi_features(i, &rmi_feat_reg_cache[i])))
> + return -EINVAL;
> + }
> +
> + return 0;
> +}
> +
As above, I would suggest to combine the logics of rmi_features() into this function.
The function name would be rmi_init_feat_regs(), consistent to rmi_feat_reg().
static int rmi_init_feat_regs(void)
{
struct arm_smccc_1_2_regs args;
int i;
/*
* Since we've negotiated a compatible version these feature registers
* should always be available
*/
for (i = 0; i < RMI_FEAT_REG_COUNT; i++) {
args.a0 = SMC_RMI_FEATURES;
args.a1 = i;
rmi_smccc_invoke(&args);
if (WARN_ON(args.a0 != RMI_SUCCESS))
return -EINVAL;
rmi_feat_reg_cache[i] = args.a1;
}
return 0;
}
> +static int __init arm64_init_rmi(void)
> +{
> + int ret;
> +
> + /* Continue without realm support if we can't agree on a version */
> + ret = rmi_check_version();
> + if (ret)
> + return ret;
The comment seems not applicable and can be dropped.
> +
> + ret = rmi_read_features();
> + if (ret)
> + return ret;
> +
> + return 0;
> +}
> +
> +/*
> + * Note arm64_init_rmi() must be called before kvm_init_rmi() otherwise KVM
> + * will not support realm guests. subsys_initcall() is called before
> + * module_init() (used for KVM) so this is OK.
> + */
> +subsys_initcall(arm64_init_rmi);
> diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
> new file mode 100644
> index 0000000000000..9792bf0e00cb9
> --- /dev/null
> +++ b/include/linux/arm-rmi-cmds.h
> @@ -0,0 +1,36 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +/*
> + * Copyright (C) 2026 ARM Ltd.
> + */
> +
> +#ifndef __LINUX_ARM_RMI_CMDS_H_
> +#define __LINUX_ARM_RMI_CMDS_H_
> +
> +#include <linux/arm-smccc-rmi.h>
> +#include <linux/bug.h>
> +#include <linux/processor.h>
> +#include <linux/types.h>
> +
> +
> +/*
> + * rmi_smccc_invoke: Invoke the RMI call and return the results in @regs_out
> + * @regs_in: Registers with the arguments filled in.
> + * @regs_out: Ouptput results from the call.
> + */
> +static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
> +{
> + struct arm_smccc_1_2_regs args = *regs;
> + unsigned long status;
> +
> + while (1) {
> + arm_smccc_1_2_invoke(&args, regs);
> + status = RMI_RETURN_STATUS(regs->a0);
> + if (status != RMI_BUSY && status != RMI_BLOCKED)
> + break;
> + cpu_relax();
> + }
> +}
> +
@regs_in and @regs_out in the comments don't exist. So the comments need to be
improved with something as below.
/**
* rmi_smccc_invoke: Invoke RMI call
* @regs: input and output arguments
*
* Invoke RMI call by taking the input arguments from @regs, and the output of
* the RMI call is also stored to @regs.
*/
> +unsigned long rmi_feat_reg(unsigned long index);
> +
> +#endif
"/* __LINUX_ARM_RMI_CMDS_H_ */" is missed here.
#endif /* __LINUX_ARM_RMI_CMDS_H_ */
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init
2026-09-12 8:36 ` [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init Suzuki K Poulose
2026-09-14 1:04 ` Gavin Shan
@ 2026-09-14 10:27 ` Sudeep Holla
1 sibling, 0 replies; 25+ messages in thread
From: Sudeep Holla @ 2026-09-14 10:27 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, will, catalin.marinas, linux-kernel,
linux-arm-kernel, steven.price, aneesh.kumar, oupton, gshan,
joey.gouly, tabba, yuzenghui, linux-coco, gankulkarni,
sdonthineni, alpergun, fj0570is, WeiLin.Chang, lpieralisi,
enju.kohei
On Sat, Sep 12, 2026 at 09:36:05AM +0100, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> Query the RMI version number and check if it is a compatible version.
> The first two feature registers are read and exposed for future code to
> use.
>
> We only support this for Little Endian kernels, the Big Endian kernel
> support is anyway marked BROKEN and is being removed.
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> v18:
> * Always use arm_smccc_1_2_invoke() for all RMIs making sure the unsused
> parameters are 0 - Sashiko
> * Move rmi_features() calls away from the arm-rmi-cmds.h to rmi.c - Gavin
> v17:
> * Rename ARM_RMM to ARM_RMM_RMI to make it easier to add Guest facing RSI
> support, which is also in progress
> v16:
> * Update Kconfig text to include PCIe TDISP.
> * Export rmi_feat_reg() here rather than in a later commit.
> v15:
> * The code is moved again, this time into the 'firmware' directory.
> v14:
> * This moves the basic RMI setup into the 'kernel' directory. This is
> because RMI will be used for some features outside of KVM so should
> be available even if KVM isn't compiled in.
> ---
> arch/arm64/Kconfig | 1 +
> arch/arm64/kernel/cpufeature.c | 1 +
> drivers/firmware/Kconfig | 1 +
> drivers/firmware/Makefile | 1 +
> drivers/firmware/arm_rmm/Kconfig | 26 +++++++
> drivers/firmware/arm_rmm/Makefile | 2 +
> drivers/firmware/arm_rmm/rmi.c | 122 ++++++++++++++++++++++++++++++
> include/linux/arm-rmi-cmds.h | 36 +++++++++
> 8 files changed, 190 insertions(+)
> create mode 100644 drivers/firmware/arm_rmm/Kconfig
> create mode 100644 drivers/firmware/arm_rmm/Makefile
> create mode 100644 drivers/firmware/arm_rmm/rmi.c
> create mode 100644 include/linux/arm-rmi-cmds.h
>
[...]
> diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
> new file mode 100644
> index 0000000000000..9792bf0e00cb9
> --- /dev/null
> +++ b/include/linux/arm-rmi-cmds.h
> @@ -0,0 +1,36 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +/*
> + * Copyright (C) 2026 ARM Ltd.
> + */
> +
> +#ifndef __LINUX_ARM_RMI_CMDS_H_
> +#define __LINUX_ARM_RMI_CMDS_H_
> +
> +#include <linux/arm-smccc-rmi.h>
> +#include <linux/bug.h>
> +#include <linux/processor.h>
> +#include <linux/types.h>
> +
> +
> +/*
> + * rmi_smccc_invoke: Invoke the RMI call and return the results in @regs_out
> + * @regs_in: Registers with the arguments filled in.
> + * @regs_out: Ouptput results from the call.
> + */
> +static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
> +{
> + struct arm_smccc_1_2_regs args = *regs;
> + unsigned long status;
> +
> + while (1) {
> + arm_smccc_1_2_invoke(&args, regs);
> + status = RMI_RETURN_STATUS(regs->a0);
> + if (status != RMI_BUSY && status != RMI_BLOCKED)
> + break;
> + cpu_relax();
> + }
> +}
I haven't done a detailed review, this is just a drive through comment.
The while(1) gained my attention.
Should RMI_BLOCKED be returned to the caller instead of retried here?
RMM spec defines RMI_BLOCKED as persisting until the Host takes action.
It also says it is returned when another SRO on the same context is
incomplete. You may be running it on different CPUs and hence different
context I assume. But this loop takes no such action and hides the status
from the caller, so a command issued against that context if that can
happen can spin indefinitely while the operation which would unblock it
cannot run. Ignore me if it taken care not to happen elsewhere. I am
just looking at this in isolation.
Could this retry only RMI_BUSY and propagate RMI_BLOCKED so that the caller
can arrange for the incomplete operation to make progress if the above
scenario is possible ?
--
Regards,
Sudeep
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v18 3/7] firmware: arm_rmm: Configure the RMM with the host's page size
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
2026-09-12 8:36 ` [PATCH v18 1/7] firmware: arm_rmm: Add SMC definitions for calling the RMM Suzuki K Poulose
2026-09-12 8:36 ` [PATCH v18 2/7] firmware: arm_rmm: Check for RMI support at init Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 1:21 ` Gavin Shan
2026-09-12 8:36 ` [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
` (3 subsequent siblings)
6 siblings, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
RMM v2.0 brings the ability to set the RMM's granule size. Check the
feature registers and configure the RMM so that it matches the host's
page size. This means that operations can be done with a granularity
equal to PAGE_SIZE.
Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
* Move rmi_config_set() out of the header file.
* Print the error message for rmi_config_set if it fails
Changes since v15:
* Actually check the feature register for the host's page-size support.
Changes since v14:
* Move the implementation into drivers/firmware/arm_rmm.
Changes since v13:
* Moved out of KVM.
---
drivers/firmware/arm_rmm/rmi.c | 79 ++++++++++++++++++++++++++++++++++
1 file changed, 79 insertions(+)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index 2fd538c937dca..5b0e342ce3d58 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -35,6 +35,25 @@ static int rmi_features(unsigned long index, unsigned long *out)
return args.a0;
}
+/**
+ * rmi_rmm_config_set() - Configure the RMM
+ * @cfg_ptr: PA of a struct rmm_config
+ *
+ * Sets configuration options on the RMM.
+ *
+ * Return: RMI return code
+ */
+static int rmi_rmm_config_set(unsigned long cfg_ptr)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RMM_CONFIG_SET, cfg_ptr,
+ };
+
+ rmi_smccc_invoke(®s);
+
+ return regs.a0;
+}
+
unsigned long rmi_feat_reg(unsigned long index)
{
if (WARN_ON(index >= RMI_FEAT_REG_COUNT))
@@ -98,6 +117,62 @@ static int rmi_read_features(void)
return 0;
}
+static int rmi_configure(void)
+{
+ unsigned long granule_feature;
+ unsigned long granule_size;
+ int ret = 0;
+ struct rmm_config *config;
+
+ switch (PAGE_SIZE) {
+ case SZ_4K:
+ granule_size = RMI_GRANULE_SIZE_4KB;
+ granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_4KB;
+ break;
+ case SZ_16K:
+ granule_size = RMI_GRANULE_SIZE_16KB;
+ granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_16KB;
+ break;
+ case SZ_64K:
+ granule_size = RMI_GRANULE_SIZE_64KB;
+ granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_64KB;
+ break;
+ default:
+ BUILD_BUG();
+ }
+
+ if (!(rmi_feat_reg(1) & granule_feature)) {
+ pr_err("RMM does not support %luKB granules\n",
+ PAGE_SIZE >> 10);
+ return -ENXIO;
+ }
+
+ config = (struct rmm_config *)get_zeroed_page(GFP_KERNEL);
+ if (!config) {
+ pr_err("Unable to allocate memory for RMM config\n");
+ return -ENOMEM;
+ }
+
+ config->rmi_granule_size = granule_size;
+
+ /*
+ * For now we set the tracking_region_size to 0 which is the only option
+ * for 4KB PAGE_SIZE (1GB for 4KB PAGE_SIZE, 32MB/512MB for 16KB/64KB).
+ * TODO: Support other tracking sizes via Kconfig option for other
+ * PAGE_SIZES
+ */
+ config->tracking_region_size = 0;
+
+ ret = rmi_rmm_config_set(virt_to_phys(config));
+ if (ret) {
+ pr_err("RMM config set failed (%d)\n", ret);
+ ret = -EINVAL;
+ }
+
+ free_page((unsigned long)config);
+ return ret;
+}
+
static int __init arm64_init_rmi(void)
{
int ret;
@@ -111,6 +186,10 @@ static int __init arm64_init_rmi(void)
if (ret)
return ret;
+ ret = rmi_configure();
+ if (ret)
+ return ret;
+
return 0;
}
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 3/7] firmware: arm_rmm: Configure the RMM with the host's page size
2026-09-12 8:36 ` [PATCH v18 3/7] firmware: arm_rmm: Configure the RMM with the host's page size Suzuki K Poulose
@ 2026-09-14 1:21 ` Gavin Shan
2026-09-14 6:34 ` Suzuki K Poulose
0 siblings, 1 reply; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 1:21 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> RMM v2.0 brings the ability to set the RMM's granule size. Check the
> feature registers and configure the RMM so that it matches the host's
> page size. This means that operations can be done with a granularity
> equal to PAGE_SIZE.
>
> Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v17:
> * Move rmi_config_set() out of the header file.
> * Print the error message for rmi_config_set if it fails
> Changes since v15:
> * Actually check the feature register for the host's page-size support.
> Changes since v14:
> * Move the implementation into drivers/firmware/arm_rmm.
> Changes since v13:
> * Moved out of KVM.
> ---
> drivers/firmware/arm_rmm/rmi.c | 79 ++++++++++++++++++++++++++++++++++
> 1 file changed, 79 insertions(+)
>
> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
> index 2fd538c937dca..5b0e342ce3d58 100644
> --- a/drivers/firmware/arm_rmm/rmi.c
> +++ b/drivers/firmware/arm_rmm/rmi.c
> @@ -35,6 +35,25 @@ static int rmi_features(unsigned long index, unsigned long *out)
> return args.a0;
> }
>
> +/**
> + * rmi_rmm_config_set() - Configure the RMM
> + * @cfg_ptr: PA of a struct rmm_config
> + *
> + * Sets configuration options on the RMM.
> + *
> + * Return: RMI return code
> + */
> +static int rmi_rmm_config_set(unsigned long cfg_ptr)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RMM_CONFIG_SET, cfg_ptr,
> + };
> +
> + rmi_smccc_invoke(®s);
> +
> + return regs.a0;
> +}
> +
The comments for rmi_rmm_config_set() can be dropped since its logic is simply enough and
the code is self-explainning. Besides, I would move this right before its only caller
rmi_configure().
I would suggest drop this function by combining its logics into the only caller
rmi_configure(), seeing below for more details.
> unsigned long rmi_feat_reg(unsigned long index)
> {
> if (WARN_ON(index >= RMI_FEAT_REG_COUNT))
> @@ -98,6 +117,62 @@ static int rmi_read_features(void)
> return 0;
> }
>
> +static int rmi_configure(void)
> +{
> + unsigned long granule_feature;
> + unsigned long granule_size;
> + int ret = 0;
> + struct rmm_config *config;
> +
> + switch (PAGE_SIZE) {
> + case SZ_4K:
> + granule_size = RMI_GRANULE_SIZE_4KB;
> + granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_4KB;
> + break;
> + case SZ_16K:
> + granule_size = RMI_GRANULE_SIZE_16KB;
> + granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_16KB;
> + break;
> + case SZ_64K:
> + granule_size = RMI_GRANULE_SIZE_64KB;
> + granule_feature = RMI_FEATURE_REGISTER_1_RMI_GRAN_SZ_64KB;
> + break;
> + default:
> + BUILD_BUG();
> + }
> +
> + if (!(rmi_feat_reg(1) & granule_feature)) {
> + pr_err("RMM does not support %luKB granules\n",
> + PAGE_SIZE >> 10);
> + return -ENXIO;
> + }
> +
> + config = (struct rmm_config *)get_zeroed_page(GFP_KERNEL);
> + if (!config) {
> + pr_err("Unable to allocate memory for RMM config\n");
> + return -ENOMEM;
> + }
> +
> + config->rmi_granule_size = granule_size;
> +
> + /*
> + * For now we set the tracking_region_size to 0 which is the only option
> + * for 4KB PAGE_SIZE (1GB for 4KB PAGE_SIZE, 32MB/512MB for 16KB/64KB).
> + * TODO: Support other tracking sizes via Kconfig option for other
> + * PAGE_SIZES
> + */
> + config->tracking_region_size = 0;
> +
> + ret = rmi_rmm_config_set(virt_to_phys(config));
> + if (ret) {
> + pr_err("RMM config set failed (%d)\n", ret);
> + ret = -EINVAL;
> + }
> +
I would suggest to drop rmi_rmm_config_set() by combining its logic to rmi_configure().
struct arm_smccc_1_2_regs args = { SMC_RMI_RMM_CONFIG_SET };
args.a1 = virt_to_phys(config);
rmi_smccc_invoke(&args);
if (args.a0 != RMI_SUCCESS) {
pr_err("RMM config set failed (%ld)\n", args.a0);
ret = -EINVAL;
}
> + free_page((unsigned long)config);
> + return ret;
> +}
> +
> static int __init arm64_init_rmi(void)
> {
> int ret;
> @@ -111,6 +186,10 @@ static int __init arm64_init_rmi(void)
> if (ret)
> return ret;
>
> + ret = rmi_configure();
> + if (ret)
> + return ret;
> +
> return 0;
> }
>
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 3/7] firmware: arm_rmm: Configure the RMM with the host's page size
2026-09-14 1:21 ` Gavin Shan
@ 2026-09-14 6:34 ` Suzuki K Poulose
0 siblings, 0 replies; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-14 6:34 UTC (permalink / raw)
To: Gavin Shan, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 14/09/2026 02:21, Gavin Shan wrote:
> On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
>> From: Steven Price <steven.price@arm.com>
>>
>> RMM v2.0 brings the ability to set the RMM's granule size. Check the
>> feature registers and configure the RMM so that it matches the host's
>> page size. This means that operations can be done with a granularity
>> equal to PAGE_SIZE.
>>
>> Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> Signed-off-by: Steven Price <steven.price@arm.com>
>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> ---
>> Changes since v17:
>> * Move rmi_config_set() out of the header file.
>> * Print the error message for rmi_config_set if it fails
>> Changes since v15:
>> * Actually check the feature register for the host's page-size
>> support.
>> Changes since v14:
>> * Move the implementation into drivers/firmware/arm_rmm.
>> Changes since v13:
>> * Moved out of KVM.
>> ---
>> drivers/firmware/arm_rmm/rmi.c | 79 ++++++++++++++++++++++++++++++++++
>> 1 file changed, 79 insertions(+)
>>
>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/
>> arm_rmm/rmi.c
>> index 2fd538c937dca..5b0e342ce3d58 100644
>> --- a/drivers/firmware/arm_rmm/rmi.c
>> +++ b/drivers/firmware/arm_rmm/rmi.c
>> @@ -35,6 +35,25 @@ static int rmi_features(unsigned long index,
>> unsigned long *out)
>> return args.a0;
>> }
>> +/**
>> + * rmi_rmm_config_set() - Configure the RMM
>> + * @cfg_ptr: PA of a struct rmm_config
>> + *
>> + * Sets configuration options on the RMM.
>> + *
>> + * Return: RMI return code
>> + */
>> +static int rmi_rmm_config_set(unsigned long cfg_ptr)
>> +{
>> + struct arm_smccc_1_2_regs regs = {
>> + SMC_RMI_RMM_CONFIG_SET, cfg_ptr,
>> + };
>> +
>> + rmi_smccc_invoke(®s);
>> +
>> + return regs.a0;
>> +}
>> +
>
> The comments for rmi_rmm_config_set() can be dropped since its logic is
> simply enough and
> the code is self-explainning. Besides, I would move this right before
> its only caller
> rmi_configure().
Ack
>
> I would suggest drop this function by combining its logics into the only
> caller
> rmi_configure(), seeing below for more details.
That looks a bit odd in the middle of a function, given the argument
setting. I have moved it closer to the configure().
Cheers
Suzuki
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (2 preceding siblings ...)
2026-09-12 8:36 ` [PATCH v18 3/7] firmware: arm_rmm: Configure the RMM with the host's page size Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 5:04 ` Gavin Shan
2026-09-14 12:50 ` Sudeep Holla
2026-09-12 8:36 ` [PATCH v18 5/7] firmware: arm_rmm: Activate the RMM Suzuki K Poulose
` (2 subsequent siblings)
6 siblings, 2 replies; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
means that an SMC can return with an operation still in progress. The
host is expected to continue the operation until it reaches a conclusion
(either success or failure). During this process the RMM can request
additional memory ('donate') or hand memory back to the host
('reclaim'). The host can request an in progress operation is cancelled,
but still continue the operation until it has completed (otherwise the
incomplete operation may cause future RMM operations to fail).
The SRO is tracked using a struct rmi_sro_state object which keeps track
of any memory which has been allocated but not yet consumed by the RMM
or reclaimed from the RMM. This allows the memory to be reused in a
future request within the same operation. It will also permit an
operation to be done in a context where memory allocation may be
difficult (e.g. atomic context) with the option to abort the operation
and retry the memory allocation outside of the atomic context. The
memory stored in the struct rmi_sro_state object can then be reused on
the subsequent attempt.
Wrappers for SRO RMI commands are also provided here because they depend
on the rmi_sro_execute() implementation added by this patch.
Delegate/undelegate handles are also added here because they now use the
SRO/stateful command infrastructure and are also used for the memory
DONATE/RECLAIM flows.
Signed-off-by: Steven Price <steven.price@arm.com>
Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
v18:
* Prevent overflow for donated_granules output from buggy RMM
* Handle corrupted addr_count in the sro
* Avoid spilling literal pools on stack with sro initialisation
* Handle buggy RMM when the out_top is not changed with RMI_SUCCESS
for delegat/undelegate range calls
* Rename free_delegated_page => rmi_free_delegated_page
* Rename donate_req_to_unit_size => donate_req_to_block_size
* Introduce rmi_addr_block_size_to_bytes() helper to convert a RmiAddrBlockSize
encoding used in RMI_DONATE_REQ and RMI_ADDR_RANGE Descriptors, replaces
donate_req_to_unit_size()
* Rename unit_size => block_size_fld, unit_size_bytes => block_size etc.
* Explicitly check for MEM_CONTIG/CAN_CANCEL fields to match the RMM spec values.
* Rename free_delegated_page => rmi_free_delegated_page()
* Drop RMI_BUSY, RMI_BLOCKED checks from rmi_*delegate_range as they are already
handled by the rmi_smccc_invoke() used by the SRO.
* Add a helper to free an address range entry, which may be partially consumed.
* Ensure RMI_OP_RECLAIM output is valid before consumption
v17:
* Handle buggy RMM firmware to avoid looping forever for non-cancellable SROs.
* Add comment (for the AI agents) to clarify that all memory donating SROs are
cancellable.
v16:
* Wrappers for realm guests split into a separate patch.
* Better support for cancellation - previously a cancelled operation
could be treated as successful.
* Consistently use a signed type for wrapper return values so that
Linux error codes can be returned as well as RMI return values.
v15:
* Wrappers for SRO RMI functions are provided in this patch due to
their dependency on the SRO infrastructure.
* Fold the range delegate/undelegate wrappers into this patch because
they depend on the stateful command infrastructure.
* Add cpu_relax() calls when RMI_BUSY/RMI_BLOCKED is returned.
* Various fixes.
v14:
* SRO support has improved although is still not fully complete. The
infrastructure has been moved out of KVM.
---
drivers/firmware/arm_rmm/rmi.c | 586 +++++++++++++++++++++++++++++++++
include/linux/arm-rmi-cmds.h | 41 +++
2 files changed, 627 insertions(+)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index 5b0e342ce3d58..4f9898ece7547 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -15,6 +15,59 @@
#define RMI_FEAT_REG_COUNT 2
static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT] __ro_after_init;
+/**
+ * rmi_granule_range_delegate() - Delegate granules
+ * @base: PA of the first granule of the range
+ * @top: PA of the first granule after the range
+ * @out_top: PA of the first granule not delegated
+ *
+ * Delegate a range of granule for use by the realm world. If the entire range
+ * was delegated then @out_top == @top, otherwise the function should be called
+ * again with @base == @out_top.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_granule_range_delegate(unsigned long base,
+ unsigned long top,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GRANULE_RANGE_DELEGATE, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_granule_range_undelegate() - Undelegate a range of granules
+ * @base: Base PA of the target range
+ * @top: Top PA of the target range
+ * @out_top: Returns the top PA of range whose state is undelegated
+ *
+ * Undelegate a range of granules to allow use by the normal world. Will fail if
+ * the granules are in use.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_granule_range_undelegate(unsigned long base,
+ unsigned long top,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GRANULE_RANGE_UNDELEGATE, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
/**
* rmi_features() - Read feature register
* @index: Feature register index
@@ -63,6 +116,539 @@ unsigned long rmi_feat_reg(unsigned long index)
}
EXPORT_SYMBOL_GPL(rmi_feat_reg);
+int rmi_undelegate_range(phys_addr_t phys,
+ unsigned long size)
+{
+ long ret = 0;
+ unsigned long top = phys + size;
+ unsigned long out_top;
+
+ while (phys < top) {
+ ret = rmi_granule_range_undelegate(phys, top, &out_top);
+
+ if (ret == RMI_SUCCESS) {
+ /* Buggy RMM ? Let the caller leak the pages */
+ if (WARN_ON(out_top <= phys))
+ return -ENXIO;
+ phys = out_top;
+ } else {
+ break;
+ }
+ }
+
+ return ret;
+}
+EXPORT_SYMBOL_GPL(rmi_undelegate_range);
+
+int rmi_delegate_range(phys_addr_t phys,
+ unsigned long size,
+ phys_addr_t *out_phys)
+{
+ long ret = 0;
+ unsigned long top = phys + size;
+ unsigned long out_top;
+
+ while (phys < top) {
+ ret = rmi_granule_range_delegate(phys, top, &out_top);
+
+ if (ret == RMI_SUCCESS) {
+ /* Buggy RMM ? */
+ if (WARN_ON(out_top <= phys)) {
+ rmi_undelegate_range(top - size, size);
+ return -ENXIO;
+ }
+ phys = out_top;
+ } else {
+ break;
+ }
+ }
+
+ if (out_phys)
+ *out_phys = phys;
+
+ return ret;
+}
+EXPORT_SYMBOL_GPL(rmi_delegate_range);
+
+/*
+ * Convert the RmiAddrBlockSize to actual size. This is used in RmiDonateReq
+ * and RmiAddrRangeDesc*.
+ */
+static unsigned long rmi_addr_block_size_to_bytes(unsigned long block_size_fld)
+{
+ return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(3 - block_size_fld));
+}
+
+/*
+ * free_addr_range: Free memory described by the address range entry, which may
+ * be partially consumed by RMM.
+ *
+ * @entry: RMI_ADDR_RANGE descriptor
+ * @consumed_size: Page aligned size consumed by the RMM from the address range.
+ *
+ * If the state of the address is DELEGATED, undelegate it back, before freeing.
+ * Leaks the memory if we cannot undelegate the range.
+ */
+static void free_addr_range(unsigned long entry, unsigned long consumed_size)
+{
+ unsigned long phys = RMI_ADDR_RANGE_ADDR(entry);
+ unsigned long block_size_fld = RMI_ADDR_RANGE_BLOCK_SIZE(entry);
+ unsigned long count = RMI_ADDR_RANGE_COUNT(entry);
+ unsigned long state = RMI_ADDR_RANGE_STATE(entry);
+ unsigned long size = rmi_addr_block_size_to_bytes(block_size_fld) * count;
+
+ WARN_ON(!PAGE_ALIGNED(phys) || !PAGE_ALIGNED(consumed_size));
+
+ /* Adjust the address and size for partially consumed entry */
+ phys += consumed_size;
+ size -= consumed_size;
+ /*
+ * Undelegate the pages back if required. If we can't
+ * change them back, leak the pages.
+ */
+ if (state == RMI_OP_MEM_DELEGATED &&
+ WARN_ON(rmi_undelegate_range(phys, size)))
+ return;
+ free_pages_exact(phys_to_virt(phys), size);
+}
+
+static void rmi_op_continue(unsigned long sro_handle, unsigned long flags,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_CONTINUE, sro_handle, flags
+ };
+
+ rmi_smccc_invoke(out_regs);
+}
+
+static void rmi_op_cancel(unsigned long sro_handle,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_CANCEL, sro_handle
+ };
+
+ rmi_smccc_invoke(out_regs);
+}
+
+static void rmi_op_mem_donate(unsigned long sro_handle, unsigned long list_addr,
+ unsigned long list_count, unsigned long flags,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_MEM_DONATE, sro_handle, list_addr, list_count, flags
+ };
+
+ /*
+ * The output donated count (a1) is always valid, irrespective
+ * of the return result. i.e., 0 if there was an error
+ */
+ rmi_smccc_invoke(out_regs);
+}
+
+static void rmi_op_mem_reclaim(unsigned long sro_handle,
+ unsigned long list_addr,
+ unsigned long list_count,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ *out_regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_OP_MEM_RECLAIM, sro_handle, list_addr, list_count
+ };
+
+ rmi_smccc_invoke(out_regs);
+}
+
+int rmi_free_delegated_page(phys_addr_t phys)
+{
+ if (WARN_ON_ONCE(rmi_undelegate_page(phys))) {
+ /* Undelegate failed: leak the page */
+ return -EBUSY;
+ }
+
+ free_page((unsigned long)phys_to_virt(phys));
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(rmi_free_delegated_page);
+
+static int rmi_sro_ensure_capacity(struct rmi_sro_state *sro,
+ unsigned long count)
+{
+ if (WARN_ON_ONCE(sro->addr_count > RMI_MAX_ADDR_LIST))
+ return -EOVERFLOW;
+
+ if (count > RMI_MAX_ADDR_LIST - sro->addr_count)
+ return -ENOSPC;
+
+ return 0;
+}
+
+static int rmi_sro_donate_contig(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ unsigned long donatereq,
+ struct arm_smccc_1_2_regs *out_regs,
+ gfp_t gfp)
+{
+ unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
+ unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
+ unsigned long count = RMI_DONATE_COUNT(donatereq);
+ unsigned long state = RMI_DONATE_STATE(donatereq);
+ unsigned long size = block_size * count;
+ unsigned long addr_range;
+ unsigned long donated_granules;
+ unsigned long donated_size;
+ int ret;
+ void *virt;
+ phys_addr_t phys;
+
+ /*
+ * The RMM specification requires contiguous allocations are always a
+ * power of 2
+ */
+ if (WARN_ON_ONCE(!is_power_of_2(size)))
+ return -EINVAL;
+
+ /* Reuse the cached address range if we have one */
+ for (int i = 0; i < sro->addr_count; i++) {
+ unsigned long entry = sro->addr_list[i];
+
+ if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
+ RMI_ADDR_RANGE_COUNT(entry) == count &&
+ RMI_ADDR_RANGE_STATE(entry) == state &&
+ IS_ALIGNED(RMI_ADDR_RANGE_ADDR(entry), size)) {
+ sro->addr_count--;
+ swap(sro->addr_list[sro->addr_count],
+ sro->addr_list[i]);
+
+ goto out;
+ }
+ }
+
+ ret = rmi_sro_ensure_capacity(sro, 1);
+ if (ret)
+ return ret;
+
+ virt = alloc_pages_exact(size, gfp);
+ if (!virt)
+ return -ENOMEM;
+ phys = virt_to_phys(virt);
+
+ if (state == RMI_OP_MEM_DELEGATED) {
+ phys_addr_t delegated_phys;
+
+ if (rmi_delegate_range(phys, size, &delegated_phys)) {
+ if (!rmi_undelegate_range(phys, delegated_phys - phys))
+ free_pages_exact(virt, size);
+ return -ENXIO;
+ }
+ }
+
+ addr_range = phys & RMI_ADDR_RANGE_ADDR_MASK;
+ FIELD_MODIFY(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, &addr_range, block_size_fld);
+ FIELD_MODIFY(RMI_ADDR_RANGE_COUNT_MASK, &addr_range, count);
+ FIELD_MODIFY(RMI_ADDR_RANGE_STATE_MASK, &addr_range, state);
+
+ sro->addr_list[sro->addr_count] = addr_range;
+
+out:
+ rmi_op_mem_donate(sro_handle,
+ virt_to_phys(&sro->addr_list[sro->addr_count]), 1,
+ 0, out_regs);
+ donated_granules = out_regs->a1;
+
+ if (WARN_ON(donated_granules > (size >> PAGE_SHIFT)))
+ donated_granules = (size >> PAGE_SHIFT);
+
+ donated_size = donated_granules << PAGE_SHIFT;
+
+ /* All granules consumed by the RMM */
+ if (donated_size == size)
+ return 0;
+ /* No granules were consumed by the RMM, cache them */
+ if (donated_granules == 0) {
+ sro->addr_count++;
+ return 0;
+ }
+
+ /* The granules were partially consumed, reclaim the unused ones. */
+ free_addr_range(sro->addr_list[sro->addr_count], donated_size);
+
+ return 0;
+}
+
+static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ unsigned long donatereq,
+ struct arm_smccc_1_2_regs *out_regs,
+ gfp_t gfp)
+{
+ unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
+ unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
+ unsigned long count = RMI_DONATE_COUNT(donatereq);
+ unsigned long state = RMI_DONATE_STATE(donatereq);
+ unsigned long found = 0;
+ unsigned long donated_granules;
+ unsigned long granules_per_block = block_size >> PAGE_SHIFT;
+ unsigned long consumed_blocks;
+ int addr_list_start = sro->addr_count;
+
+ int ret;
+
+ for (int i = 0; i < addr_list_start && found < count; i++) {
+ unsigned long entry = sro->addr_list[i];
+
+ if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
+ RMI_ADDR_RANGE_COUNT(entry) == 1 &&
+ RMI_ADDR_RANGE_STATE(entry) == state) {
+ addr_list_start--;
+ swap(sro->addr_list[addr_list_start],
+ sro->addr_list[i]);
+ found++;
+ i--;
+ }
+ }
+
+ ret = rmi_sro_ensure_capacity(sro, count - found);
+ if (ret)
+ return ret;
+
+ while (found < count) {
+ unsigned long addr_range;
+ void *virt = alloc_pages_exact(block_size, gfp);
+ phys_addr_t phys;
+
+ if (!virt)
+ return -ENOMEM;
+
+ phys = virt_to_phys(virt);
+
+ if (state == RMI_OP_MEM_DELEGATED) {
+ phys_addr_t delegated_phys;
+
+ if (rmi_delegate_range(phys, block_size,
+ &delegated_phys)) {
+ if (!rmi_undelegate_range(phys, delegated_phys - phys))
+ free_pages_exact(virt, block_size);
+ return -ENXIO;
+ }
+ }
+
+ addr_range = phys & RMI_ADDR_RANGE_ADDR_MASK;
+ FIELD_MODIFY(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, &addr_range, block_size_fld);
+ FIELD_MODIFY(RMI_ADDR_RANGE_COUNT_MASK, &addr_range, 1);
+ FIELD_MODIFY(RMI_ADDR_RANGE_STATE_MASK, &addr_range, state);
+
+ sro->addr_list[sro->addr_count++] = addr_range;
+ found++;
+ }
+
+ rmi_op_mem_donate(sro_handle,
+ virt_to_phys(&sro->addr_list[addr_list_start]),
+ count, 0, out_regs);
+
+ donated_granules = out_regs->a1;
+ /*
+ * The RMM shouldn't report more granules than we provided, but clamp
+ * just in case.
+ */
+ if (WARN_ON_ONCE(donated_granules > found * granules_per_block))
+ donated_granules = count * granules_per_block;
+
+ /*
+ * The RMM reports the consumed memory in terms of granules, but we
+ * track in the address lists in block-sized ranges. So divide to get
+ * the number of (complete) consumed blocks.
+ */
+ consumed_blocks = donated_granules / granules_per_block;
+ if (donated_granules % granules_per_block) {
+ /*
+ * A block has been partially consumed, the start is owned by
+ * the RMM, the tail is owned by the host
+ */
+ unsigned long entry =
+ sro->addr_list[addr_list_start + consumed_blocks];
+ unsigned long donated_size =
+ (donated_granules % granules_per_block) << PAGE_SHIFT;
+
+ free_addr_range(entry, donated_size);
+ /*
+ * This block is now fully 'consumed' (either held by the RMM or
+ * freed)
+ */
+ consumed_blocks++;
+ }
+
+ /* Keep just the blocks the RMM didn't use in addr_list */
+ for (int i = consumed_blocks; i < count; i++)
+ sro->addr_list[addr_list_start + i - consumed_blocks] =
+ sro->addr_list[addr_list_start + i];
+
+ sro->addr_count -= consumed_blocks;
+
+ return 0;
+}
+
+static int rmi_sro_donate(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ unsigned long donatereq,
+ struct arm_smccc_1_2_regs *regs,
+ gfp_t gfp)
+{
+ if (WARN_ON_ONCE(!RMI_DONATE_COUNT(donatereq)))
+ return -EINVAL;
+
+ if (RMI_DONATE_CONTIG(donatereq) == RMI_OP_MEM_CONTIG) {
+ return rmi_sro_donate_contig(sro, sro_handle, donatereq,
+ regs, gfp);
+ } else {
+ return rmi_sro_donate_noncontig(sro, sro_handle, donatereq,
+ regs, gfp);
+ }
+}
+
+static int rmi_sro_reclaim(struct rmi_sro_state *sro,
+ unsigned long sro_handle,
+ struct arm_smccc_1_2_regs *out_regs)
+{
+ unsigned long capacity;
+
+ if (rmi_sro_ensure_capacity(sro, 1))
+ rmi_sro_free(sro);
+
+ capacity = RMI_MAX_ADDR_LIST - sro->addr_count;
+
+ rmi_op_mem_reclaim(sro_handle,
+ virt_to_phys(&sro->addr_list[sro->addr_count]),
+ capacity, out_regs);
+
+ /*
+ * RMI_OP_MEM_RECLAIM always return RMI_INCOMPLETE, except when the
+ * input parameters were invalid.
+ */
+ if (WARN_ON_ONCE(RMI_RETURN_STATUS(out_regs->a0) != RMI_INCOMPLETE))
+ return -EINVAL;
+ if (WARN_ON_ONCE(out_regs->a1 > capacity))
+ out_regs->a1 = capacity;
+
+ sro->addr_count += out_regs->a1;
+
+ return 0;
+}
+
+void rmi_sro_free(struct rmi_sro_state *sro)
+{
+ /* Handle the worse */
+ if (WARN_ON(sro->addr_count < 0))
+ return;
+
+ if (WARN_ON(sro->addr_count > RMI_MAX_ADDR_LIST))
+ sro->addr_count = RMI_MAX_ADDR_LIST;
+
+ for (int i = 0; i < sro->addr_count; i++)
+ free_addr_range(sro->addr_list[i], 0);
+
+ sro->addr_count = 0;
+}
+EXPORT_SYMBOL_GPL(rmi_sro_free);
+
+long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp)
+{
+ struct arm_smccc_1_2_regs *regs = &sro->regs;
+ bool cancelled = false;
+ unsigned long sro_handle;
+
+ rmi_smccc_invoke(regs);
+
+ sro_handle = regs->a1;
+ while (RMI_RETURN_STATUS(regs->a0) == RMI_INCOMPLETE) {
+ bool can_cancel = RMI_RETURN_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
+ int ret = 0;
+
+ switch (RMI_RETURN_MEMREQ(regs->a0)) {
+ case RMI_OP_MEM_REQ_NONE:
+ rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
+ regs);
+ break;
+ case RMI_OP_MEM_REQ_DONATE:
+ ret = rmi_sro_donate(sro, sro_handle, regs->a2, regs,
+ gfp);
+ break;
+ case RMI_OP_MEM_REQ_RECLAIM:
+ ret = rmi_sro_reclaim(sro, sro_handle, regs);
+ break;
+ default:
+ ret = WARN_ON_ONCE(1);
+ break;
+ }
+
+ if (ret) {
+ /*
+ * All memory donating SROs must be cancellable. So a
+ * failure in memory allocation shouldn't be an issue.
+ * However, if we encounter a random failure (e.g.,
+ * buggy RMM), don't loop forever, just give up.
+ */
+ if (WARN_ON_ONCE(!can_cancel))
+ return ret;
+ /*
+ * If we have already cancelled, and came back here due
+ * to an error in MEMREQ, then there is no point
+ * in going in loops.
+ */
+ if (WARN_ON_ONCE(cancelled))
+ break;
+ rmi_op_cancel(sro_handle, regs);
+ cancelled = true;
+
+ if (WARN_ON_ONCE(RMI_RETURN_STATUS(regs->a0) != RMI_INCOMPLETE))
+ return ret;
+ }
+ }
+
+ if (cancelled)
+ return -ECANCELED;
+
+ return regs->a0;
+}
+EXPORT_SYMBOL_GPL(rmi_sro_memxfer_execute);
+
+/* For RMI commands that are stateful but not memory-transferring */
+long rmi_sro_execute(struct arm_smccc_1_2_regs *regs)
+{
+ bool cancelled = false;
+ unsigned long sro_handle = regs->a1;
+
+ rmi_smccc_invoke(regs);
+
+ sro_handle = regs->a1;
+ while (RMI_RETURN_STATUS(regs->a0) == RMI_INCOMPLETE) {
+ bool can_cancel = RMI_RETURN_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
+
+ switch (RMI_RETURN_MEMREQ(regs->a0)) {
+ case RMI_OP_MEM_REQ_NONE:
+ rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
+ regs);
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ if (!can_cancel)
+ return regs->a0;
+ /* If we have already cancelled, don't retry this */
+ if (cancelled)
+ return -ECANCELED;
+ rmi_op_cancel(sro_handle, regs);
+ cancelled = true;
+ }
+ }
+
+ if (cancelled)
+ return -ECANCELED;
+
+ return regs->a0;
+}
+EXPORT_SYMBOL_GPL(rmi_sro_execute);
+
static int rmi_check_version(void)
{
unsigned short version_major, version_minor;
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
index 9792bf0e00cb9..cd6e309ddf2d8 100644
--- a/include/linux/arm-rmi-cmds.h
+++ b/include/linux/arm-rmi-cmds.h
@@ -8,10 +8,20 @@
#include <linux/arm-smccc-rmi.h>
#include <linux/bug.h>
+#include <linux/gfp.h>
#include <linux/processor.h>
+#include <linux/string.h>
#include <linux/types.h>
+#define RMI_MAX_ADDR_LIST 256
+
+struct rmi_sro_state {
+ struct arm_smccc_1_2_regs regs;
+ int addr_count;
+ unsigned long addr_list[RMI_MAX_ADDR_LIST];
+};
+
/*
* rmi_smccc_invoke: Invoke the RMI call and return the results in @regs_out
* @regs_in: Registers with the arguments filled in.
@@ -33,4 +43,35 @@ static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
unsigned long rmi_feat_reg(unsigned long index);
+int rmi_delegate_range(phys_addr_t phys, unsigned long size,
+ phys_addr_t *out_phys);
+int rmi_undelegate_range(phys_addr_t phys, unsigned long size);
+int rmi_free_delegated_page(phys_addr_t phys);
+
+static inline int rmi_delegate_page(phys_addr_t phys)
+{
+ return rmi_delegate_range(phys, PAGE_SIZE, NULL);
+}
+
+static inline int rmi_undelegate_page(phys_addr_t phys)
+{
+ return rmi_undelegate_range(phys, PAGE_SIZE);
+}
+
+long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp);
+void rmi_sro_free(struct rmi_sro_state *sro);
+long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
+
+/*
+ * Resetting the addr_count is sufficient to ignore the addr_list contents.
+ */
+#define rmi_sro_memxfer_cmd(sro, gfp, ...) ({ \
+ struct rmi_sro_state *__sro = (sro); \
+ __sro->addr_count = 0; \
+ __sro->regs = (struct arm_smccc_1_2_regs){ __VA_ARGS__ }; \
+ long __ret = rmi_sro_memxfer_execute(__sro, gfp); \
+ rmi_sro_free(__sro); \
+ __ret; \
+})
+
#endif
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-12 8:36 ` [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
@ 2026-09-14 5:04 ` Gavin Shan
2026-09-14 6:22 ` Suzuki K Poulose
2026-09-14 12:50 ` Sudeep Holla
1 sibling, 1 reply; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 5:04 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
> means that an SMC can return with an operation still in progress. The
> host is expected to continue the operation until it reaches a conclusion
> (either success or failure). During this process the RMM can request
> additional memory ('donate') or hand memory back to the host
> ('reclaim'). The host can request an in progress operation is cancelled,
> but still continue the operation until it has completed (otherwise the
> incomplete operation may cause future RMM operations to fail).
>
> The SRO is tracked using a struct rmi_sro_state object which keeps track
> of any memory which has been allocated but not yet consumed by the RMM
> or reclaimed from the RMM. This allows the memory to be reused in a
> future request within the same operation. It will also permit an
> operation to be done in a context where memory allocation may be
> difficult (e.g. atomic context) with the option to abort the operation
> and retry the memory allocation outside of the atomic context. The
> memory stored in the struct rmi_sro_state object can then be reused on
> the subsequent attempt.
>
> Wrappers for SRO RMI commands are also provided here because they depend
> on the rmi_sro_execute() implementation added by this patch.
> Delegate/undelegate handles are also added here because they now use the
> SRO/stateful command infrastructure and are also used for the memory
> DONATE/RECLAIM flows.
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> v18:
> * Prevent overflow for donated_granules output from buggy RMM
> * Handle corrupted addr_count in the sro
> * Avoid spilling literal pools on stack with sro initialisation
> * Handle buggy RMM when the out_top is not changed with RMI_SUCCESS
> for delegat/undelegate range calls
> * Rename free_delegated_page => rmi_free_delegated_page
> * Rename donate_req_to_unit_size => donate_req_to_block_size
> * Introduce rmi_addr_block_size_to_bytes() helper to convert a RmiAddrBlockSize
> encoding used in RMI_DONATE_REQ and RMI_ADDR_RANGE Descriptors, replaces
> donate_req_to_unit_size()
> * Rename unit_size => block_size_fld, unit_size_bytes => block_size etc.
> * Explicitly check for MEM_CONTIG/CAN_CANCEL fields to match the RMM spec values.
> * Rename free_delegated_page => rmi_free_delegated_page()
> * Drop RMI_BUSY, RMI_BLOCKED checks from rmi_*delegate_range as they are already
> handled by the rmi_smccc_invoke() used by the SRO.
> * Add a helper to free an address range entry, which may be partially consumed.
> * Ensure RMI_OP_RECLAIM output is valid before consumption
> v17:
> * Handle buggy RMM firmware to avoid looping forever for non-cancellable SROs.
> * Add comment (for the AI agents) to clarify that all memory donating SROs are
> cancellable.
> v16:
> * Wrappers for realm guests split into a separate patch.
> * Better support for cancellation - previously a cancelled operation
> could be treated as successful.
> * Consistently use a signed type for wrapper return values so that
> Linux error codes can be returned as well as RMI return values.
> v15:
> * Wrappers for SRO RMI functions are provided in this patch due to
> their dependency on the SRO infrastructure.
> * Fold the range delegate/undelegate wrappers into this patch because
> they depend on the stateful command infrastructure.
> * Add cpu_relax() calls when RMI_BUSY/RMI_BLOCKED is returned.
> * Various fixes.
> v14:
> * SRO support has improved although is still not fully complete. The
> infrastructure has been moved out of KVM.
> ---
> drivers/firmware/arm_rmm/rmi.c | 586 +++++++++++++++++++++++++++++++++
> include/linux/arm-rmi-cmds.h | 41 +++
> 2 files changed, 627 insertions(+)
>
> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
> index 5b0e342ce3d58..4f9898ece7547 100644
> --- a/drivers/firmware/arm_rmm/rmi.c
> +++ b/drivers/firmware/arm_rmm/rmi.c
> @@ -15,6 +15,59 @@
> #define RMI_FEAT_REG_COUNT 2
> static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT] __ro_after_init;
>
> +/**
> + * rmi_granule_range_delegate() - Delegate granules
> + * @base: PA of the first granule of the range
> + * @top: PA of the first granule after the range
> + * @out_top: PA of the first granule not delegated
> + *
> + * Delegate a range of granule for use by the realm world. If the entire range
> + * was delegated then @out_top == @top, otherwise the function should be called
> + * again with @base == @out_top.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_granule_range_delegate(unsigned long base,
> + unsigned long top,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_GRANULE_RANGE_DELEGATE, base, top
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_top)
> + *out_top = regs.a1;
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_granule_range_undelegate() - Undelegate a range of granules
> + * @base: Base PA of the target range
> + * @top: Top PA of the target range
> + * @out_top: Returns the top PA of range whose state is undelegated
> + *
> + * Undelegate a range of granules to allow use by the normal world. Will fail if
> + * the granules are in use.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_granule_range_undelegate(unsigned long base,
> + unsigned long top,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_GRANULE_RANGE_UNDELEGATE, base, top
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_top)
> + *out_top = regs.a1;
> +
> + return ret;
> +}
> +
I would drop rmi_granule_range_{delegate, undelegate}() by combining their logics to
their only callers rmi_{delegate, undelegate}_range(). More details are provided for
rmi_{delegate, undelegate}_range() in the below.
> /**
> * rmi_features() - Read feature register
> * @index: Feature register index
> @@ -63,6 +116,539 @@ unsigned long rmi_feat_reg(unsigned long index)
> }
> EXPORT_SYMBOL_GPL(rmi_feat_reg);
>
> +int rmi_undelegate_range(phys_addr_t phys,
> + unsigned long size)
> +{
> + long ret = 0;
> + unsigned long top = phys + size;
> + unsigned long out_top;
> +
> + while (phys < top) {
> + ret = rmi_granule_range_undelegate(phys, top, &out_top);
> +
> + if (ret == RMI_SUCCESS) {
> + /* Buggy RMM ? Let the caller leak the pages */
> + if (WARN_ON(out_top <= phys))
> + return -ENXIO;
> + phys = out_top;
> + } else {
> + break;
> + }
> + }
> +
> + return ret;
> +}
> +EXPORT_SYMBOL_GPL(rmi_undelegate_range);
> +
I would move rmi_undelegate_range() after rmi_delegate_range().
> +int rmi_delegate_range(phys_addr_t phys,
> + unsigned long size,
> + phys_addr_t *out_phys)
> +{
> + long ret = 0;
> + unsigned long top = phys + size;
> + unsigned long out_top;
> +
> + while (phys < top) {
> + ret = rmi_granule_range_delegate(phys, top, &out_top);
> +
> + if (ret == RMI_SUCCESS) {
> + /* Buggy RMM ? */
> + if (WARN_ON(out_top <= phys)) {
> + rmi_undelegate_range(top - size, size);
[top - size, size] is incorrect because we may be delegating a sub-range of the
range of granules. It's actually the caller's responsibility to undelegate the
graunles that have been delegated.
if (WARN_ON(out_top <= phys)) {
ret = -ENXIO;
break;
}
> + return -ENXIO;
> + }
> + phys = out_top;
> + } else {
> + break;
> + }
> + }
> +
> + if (out_phys)
> + *out_phys = phys;
> +
> + return ret;
> +}
> +EXPORT_SYMBOL_GPL(rmi_delegate_range);
> +
As suggested above, I would drop rmi_granule_range_{delegate, undelegate}() by combining
their logics to their only callers rmi_{delegate, undelegate}_range().
/**
* rmi_delegate_range - Delegate a range of granules
* @phys: PA of the granule range
* @size: Size of the granule range
* @out_phys: PA of the first undelegated granule
*
* Delegate a range of granules for use by the realm world. If the entire range
* is delegated, then @out_phys == (@phys + @size). Otherwise, @out_phys points
* to the first undelegated granule.
*/
int rmi_delegate_range(phys_addr_t phys,
unsigned long size,
phys_addr_t *out_phys)
{
unsigned long start = phys;
unsigned long end = start + size;
struct arm_smccc_1_2_regs args;
long ret = 0;
while (start < end) {
args.a0 = SMC_RMI_GRANULE_RANGE_DELEGATE;
args.a1 = start;
args.a2 = end;
ret = rmi_sro_execute(&args);
if (ret != RMI_SUCCESS)
break;
/* Buggy RMM ? */
if (WARN_ON(args.a1 <= start)) {
ret = -ENXIO;
break;
}
start = args.a1;
}
if (out_phys)
*out_phys = start;
return ret;
}
EXPORT_SYMBOL_GPL(rmi_delegate_range);
/**
* rmi_undelegate_range - Undelegate a range of granules
* @phys: PA of the granule range
* @size: Size of the granule range
*
* Undelegate a range of granules for use by the normal world.
*
* Return: 0 on success, positive RMI result code or negative Linux error code
*/
int rmi_undelegate_range(phys_addr_t phys,
unsigned long size)
{
unsigned long start = phys;
unsigned long end = start + size;
struct arm_smccc_1_2_regs args;
long ret = 0;
while (start < end) {
args.a0 = SMC_RMI_GRANULE_RANGE_UNDELEGATE;
args.a1 = start;
args.a2 = end;
ret = rmi_sro_execute(&args);
if (ret != RMI_SUCCESS)
break;
/* Buggy RMM ? Let the caller leak the pages */
if (WARN_ON(args.a1 <= phys)) {
ret = -ENXIO;
break;
}
start = args.a1;
}
return ret;
}
EXPORT_SYMBOL_GPL(rmi_undelegate_range);
> +/*
> + * Convert the RmiAddrBlockSize to actual size. This is used in RmiDonateReq
> + * and RmiAddrRangeDesc*.
> + */
> +static unsigned long rmi_addr_block_size_to_bytes(unsigned long block_size_fld)
> +{
> + return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(3 - block_size_fld));
> +}
> +
> +/*
> + * free_addr_range: Free memory described by the address range entry, which may
> + * be partially consumed by RMM.
> + *
> + * @entry: RMI_ADDR_RANGE descriptor
> + * @consumed_size: Page aligned size consumed by the RMM from the address range.
> + *
> + * If the state of the address is DELEGATED, undelegate it back, before freeing.
> + * Leaks the memory if we cannot undelegate the range.
> + */
> +static void free_addr_range(unsigned long entry, unsigned long consumed_size)
> +{
> + unsigned long phys = RMI_ADDR_RANGE_ADDR(entry);
> + unsigned long block_size_fld = RMI_ADDR_RANGE_BLOCK_SIZE(entry);
> + unsigned long count = RMI_ADDR_RANGE_COUNT(entry);
> + unsigned long state = RMI_ADDR_RANGE_STATE(entry);
> + unsigned long size = rmi_addr_block_size_to_bytes(block_size_fld) * count;
> +
> + WARN_ON(!PAGE_ALIGNED(phys) || !PAGE_ALIGNED(consumed_size));
> +
> + /* Adjust the address and size for partially consumed entry */
> + phys += consumed_size;
> + size -= consumed_size;
> + /*
> + * Undelegate the pages back if required. If we can't
> + * change them back, leak the pages.
> + */
> + if (state == RMI_OP_MEM_DELEGATED &&
> + WARN_ON(rmi_undelegate_range(phys, size)))
> + return;
> + free_pages_exact(phys_to_virt(phys), size);
> +}
> +
> +static void rmi_op_continue(unsigned long sro_handle, unsigned long flags,
> + struct arm_smccc_1_2_regs *out_regs)
> +{
> + *out_regs = (struct arm_smccc_1_2_regs) {
> + SMC_RMI_OP_CONTINUE, sro_handle, flags
> + };
> +
> + rmi_smccc_invoke(out_regs);
> +}
> +
The pattern 'regs' is used in some of the 'struct arm_smccc_1_2_regs' arguments
or variables in this series, which is incosistent to the existing patterns which
is either 'args' or 'res' by searching the source files using 'git grep arm_smccc_1_2_regs'.
So I would suggest we have the fixed the pattern 'args' :-)
> +static void rmi_op_cancel(unsigned long sro_handle,
> + struct arm_smccc_1_2_regs *out_regs)
> +{
> + *out_regs = (struct arm_smccc_1_2_regs) {
> + SMC_RMI_OP_CANCEL, sro_handle
> + };
> +
> + rmi_smccc_invoke(out_regs);
> +}
> +
> +static void rmi_op_mem_donate(unsigned long sro_handle, unsigned long list_addr,
> + unsigned long list_count, unsigned long flags,
> + struct arm_smccc_1_2_regs *out_regs)
> +{
> + *out_regs = (struct arm_smccc_1_2_regs) {
> + SMC_RMI_OP_MEM_DONATE, sro_handle, list_addr, list_count, flags
> + };
> +
> + /*
> + * The output donated count (a1) is always valid, irrespective
> + * of the return result. i.e., 0 if there was an error
> + */
> + rmi_smccc_invoke(out_regs);
> +}
> +
> +static void rmi_op_mem_reclaim(unsigned long sro_handle,
> + unsigned long list_addr,
> + unsigned long list_count,
> + struct arm_smccc_1_2_regs *out_regs)
> +{
> + *out_regs = (struct arm_smccc_1_2_regs) {
> + SMC_RMI_OP_MEM_RECLAIM, sro_handle, list_addr, list_count
> + };
> +
> + rmi_smccc_invoke(out_regs);
> +}
> +
> +int rmi_free_delegated_page(phys_addr_t phys)
> +{
> + if (WARN_ON_ONCE(rmi_undelegate_page(phys))) {
> + /* Undelegate failed: leak the page */
> + return -EBUSY;
> + }
> +
> + free_page((unsigned long)phys_to_virt(phys));
> +
> + return 0;
> +}
> +EXPORT_SYMBOL_GPL(rmi_free_delegated_page);
> +
I would move rmi_free_delegated_page() right after rmi_undelegate_range().
> +static int rmi_sro_ensure_capacity(struct rmi_sro_state *sro,
> + unsigned long count)
> +{
> + if (WARN_ON_ONCE(sro->addr_count > RMI_MAX_ADDR_LIST))
> + return -EOVERFLOW;
> +
> + if (count > RMI_MAX_ADDR_LIST - sro->addr_count)
> + return -ENOSPC;
> +
> + return 0;
> +}
> +
> +static int rmi_sro_donate_contig(struct rmi_sro_state *sro,
> + unsigned long sro_handle,
> + unsigned long donatereq,
> + struct arm_smccc_1_2_regs *out_regs,
> + gfp_t gfp)
> +{
> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
> + unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
> + unsigned long count = RMI_DONATE_COUNT(donatereq);
> + unsigned long state = RMI_DONATE_STATE(donatereq);
> + unsigned long size = block_size * count;
> + unsigned long addr_range;
> + unsigned long donated_granules;
> + unsigned long donated_size;
> + int ret;
> + void *virt;
> + phys_addr_t phys;
> +
> + /*
> + * The RMM specification requires contiguous allocations are always a
> + * power of 2
> + */
> + if (WARN_ON_ONCE(!is_power_of_2(size)))
> + return -EINVAL;
> +
> + /* Reuse the cached address range if we have one */
> + for (int i = 0; i < sro->addr_count; i++) {
> + unsigned long entry = sro->addr_list[i];
> +
> + if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
> + RMI_ADDR_RANGE_COUNT(entry) == count &&
> + RMI_ADDR_RANGE_STATE(entry) == state &&
> + IS_ALIGNED(RMI_ADDR_RANGE_ADDR(entry), size)) {
> + sro->addr_count--;
> + swap(sro->addr_list[sro->addr_count],
> + sro->addr_list[i]);
> +
> + goto out;
> + }
> + }
> +
> + ret = rmi_sro_ensure_capacity(sro, 1);
> + if (ret)
> + return ret;
> +
> + virt = alloc_pages_exact(size, gfp);
> + if (!virt)
> + return -ENOMEM;
> + phys = virt_to_phys(virt);
> +
> + if (state == RMI_OP_MEM_DELEGATED) {
> + phys_addr_t delegated_phys;
> +
> + if (rmi_delegate_range(phys, size, &delegated_phys)) {
> + if (!rmi_undelegate_range(phys, delegated_phys - phys))
> + free_pages_exact(virt, size);
> + return -ENXIO;
> + }
> + }
> +
> + addr_range = phys & RMI_ADDR_RANGE_ADDR_MASK;
> + FIELD_MODIFY(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, &addr_range, block_size_fld);
> + FIELD_MODIFY(RMI_ADDR_RANGE_COUNT_MASK, &addr_range, count);
> + FIELD_MODIFY(RMI_ADDR_RANGE_STATE_MASK, &addr_range, state);
> +
> + sro->addr_list[sro->addr_count] = addr_range;
> +
> +out:
> + rmi_op_mem_donate(sro_handle,
> + virt_to_phys(&sro->addr_list[sro->addr_count]), 1,
> + 0, out_regs);
> + donated_granules = out_regs->a1;
> +
> + if (WARN_ON(donated_granules > (size >> PAGE_SHIFT)))
> + donated_granules = (size >> PAGE_SHIFT);
> +
> + donated_size = donated_granules << PAGE_SHIFT;
> +
> + /* All granules consumed by the RMM */
> + if (donated_size == size)
> + return 0;
> + /* No granules were consumed by the RMM, cache them */
> + if (donated_granules == 0) {
> + sro->addr_count++;
> + return 0;
> + }
> +
> + /* The granules were partially consumed, reclaim the unused ones. */
> + free_addr_range(sro->addr_list[sro->addr_count], donated_size);
> +
> + return 0;
> +}
> +
> +static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
> + unsigned long sro_handle,
> + unsigned long donatereq,
> + struct arm_smccc_1_2_regs *out_regs,
> + gfp_t gfp)
> +{
> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
> + unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
> + unsigned long count = RMI_DONATE_COUNT(donatereq);
> + unsigned long state = RMI_DONATE_STATE(donatereq);
> + unsigned long found = 0;
> + unsigned long donated_granules;
> + unsigned long granules_per_block = block_size >> PAGE_SHIFT;
> + unsigned long consumed_blocks;
> + int addr_list_start = sro->addr_count;
> +
^^^^^^
Unecessary blank line.
> + int ret;
> +
> + for (int i = 0; i < addr_list_start && found < count; i++) {
> + unsigned long entry = sro->addr_list[i];
> +
> + if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
> + RMI_ADDR_RANGE_COUNT(entry) == 1 &&
> + RMI_ADDR_RANGE_STATE(entry) == state) {
> + addr_list_start--;
> + swap(sro->addr_list[addr_list_start],
> + sro->addr_list[i]);
> + found++;
> + i--;
> + }
> + }
> +
> + ret = rmi_sro_ensure_capacity(sro, count - found);
> + if (ret)
> + return ret;
> +
> + while (found < count) {
> + unsigned long addr_range;
> + void *virt = alloc_pages_exact(block_size, gfp);
> + phys_addr_t phys;
> +
> + if (!virt)
> + return -ENOMEM;
> +
> + phys = virt_to_phys(virt);
> +
> + if (state == RMI_OP_MEM_DELEGATED) {
> + phys_addr_t delegated_phys;
> +
> + if (rmi_delegate_range(phys, block_size,
> + &delegated_phys)) {
> + if (!rmi_undelegate_range(phys, delegated_phys - phys))
> + free_pages_exact(virt, block_size);
> + return -ENXIO;
> + }
> + }
> +
> + addr_range = phys & RMI_ADDR_RANGE_ADDR_MASK;
> + FIELD_MODIFY(RMI_ADDR_RANGE_BLOCK_SIZE_MASK, &addr_range, block_size_fld);
> + FIELD_MODIFY(RMI_ADDR_RANGE_COUNT_MASK, &addr_range, 1);
> + FIELD_MODIFY(RMI_ADDR_RANGE_STATE_MASK, &addr_range, state);
> +
> + sro->addr_list[sro->addr_count++] = addr_range;
> + found++;
> + }
> +
> + rmi_op_mem_donate(sro_handle,
> + virt_to_phys(&sro->addr_list[addr_list_start]),
> + count, 0, out_regs);
> +
> + donated_granules = out_regs->a1;
> + /*
> + * The RMM shouldn't report more granules than we provided, but clamp
> + * just in case.
> + */
> + if (WARN_ON_ONCE(donated_granules > found * granules_per_block))
> + donated_granules = count * granules_per_block;
> +
> + /*
> + * The RMM reports the consumed memory in terms of granules, but we
> + * track in the address lists in block-sized ranges. So divide to get
> + * the number of (complete) consumed blocks.
> + */
> + consumed_blocks = donated_granules / granules_per_block;
> + if (donated_granules % granules_per_block) {
> + /*
> + * A block has been partially consumed, the start is owned by
> + * the RMM, the tail is owned by the host
> + */
> + unsigned long entry =
> + sro->addr_list[addr_list_start + consumed_blocks];
> + unsigned long donated_size =
> + (donated_granules % granules_per_block) << PAGE_SHIFT;
> +
> + free_addr_range(entry, donated_size);
> + /*
> + * This block is now fully 'consumed' (either held by the RMM or
> + * freed)
> + */
> + consumed_blocks++;
> + }
> +
> + /* Keep just the blocks the RMM didn't use in addr_list */
> + for (int i = consumed_blocks; i < count; i++)
> + sro->addr_list[addr_list_start + i - consumed_blocks] =
> + sro->addr_list[addr_list_start + i];
> +
> + sro->addr_count -= consumed_blocks;
> +
> + return 0;
> +}
> +
> +static int rmi_sro_donate(struct rmi_sro_state *sro,
> + unsigned long sro_handle,
> + unsigned long donatereq,
> + struct arm_smccc_1_2_regs *regs,
> + gfp_t gfp)
> +{
> + if (WARN_ON_ONCE(!RMI_DONATE_COUNT(donatereq)))
> + return -EINVAL;
> +
> + if (RMI_DONATE_CONTIG(donatereq) == RMI_OP_MEM_CONTIG) {
> + return rmi_sro_donate_contig(sro, sro_handle, donatereq,
> + regs, gfp);
> + } else {
> + return rmi_sro_donate_noncontig(sro, sro_handle, donatereq,
> + regs, gfp);
> + }
> +}
> +
> +static int rmi_sro_reclaim(struct rmi_sro_state *sro,
> + unsigned long sro_handle,
> + struct arm_smccc_1_2_regs *out_regs)
> +{
> + unsigned long capacity;
> +
> + if (rmi_sro_ensure_capacity(sro, 1))
> + rmi_sro_free(sro);
> +
> + capacity = RMI_MAX_ADDR_LIST - sro->addr_count;
> +
> + rmi_op_mem_reclaim(sro_handle,
> + virt_to_phys(&sro->addr_list[sro->addr_count]),
> + capacity, out_regs);
> +
> + /*
> + * RMI_OP_MEM_RECLAIM always return RMI_INCOMPLETE, except when the
> + * input parameters were invalid.
> + */
> + if (WARN_ON_ONCE(RMI_RETURN_STATUS(out_regs->a0) != RMI_INCOMPLETE))
> + return -EINVAL;
> + if (WARN_ON_ONCE(out_regs->a1 > capacity))
> + out_regs->a1 = capacity;
> +
> + sro->addr_count += out_regs->a1;
> +
> + return 0;
> +}
> +
> +void rmi_sro_free(struct rmi_sro_state *sro)
> +{
> + /* Handle the worse */
> + if (WARN_ON(sro->addr_count < 0))
> + return;
> +
> + if (WARN_ON(sro->addr_count > RMI_MAX_ADDR_LIST))
> + sro->addr_count = RMI_MAX_ADDR_LIST;
> +
> + for (int i = 0; i < sro->addr_count; i++)
> + free_addr_range(sro->addr_list[i], 0);
> +
> + sro->addr_count = 0;
> +}
> +EXPORT_SYMBOL_GPL(rmi_sro_free);
> +
> +long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp)
> +{
> + struct arm_smccc_1_2_regs *regs = &sro->regs;
> + bool cancelled = false;
> + unsigned long sro_handle;
> +
> + rmi_smccc_invoke(regs);
> +
> + sro_handle = regs->a1;
> + while (RMI_RETURN_STATUS(regs->a0) == RMI_INCOMPLETE) {
> + bool can_cancel = RMI_RETURN_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
> + int ret = 0;
> +
> + switch (RMI_RETURN_MEMREQ(regs->a0)) {
> + case RMI_OP_MEM_REQ_NONE:
> + rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
> + regs);
> + break;
> + case RMI_OP_MEM_REQ_DONATE:
> + ret = rmi_sro_donate(sro, sro_handle, regs->a2, regs,
> + gfp);
> + break;
> + case RMI_OP_MEM_REQ_RECLAIM:
> + ret = rmi_sro_reclaim(sro, sro_handle, regs);
> + break;
> + default:
> + ret = WARN_ON_ONCE(1);
> + break;
"ret = WARN_ON_ONCE(1)" is same to "ret = true". I guess we would return
-EINVAL here.
WARN_ON_ONCE(1);
ret = -EINVAL;
break;
> + }
> +
> + if (ret) {
> + /*
> + * All memory donating SROs must be cancellable. So a
> + * failure in memory allocation shouldn't be an issue.
> + * However, if we encounter a random failure (e.g.,
> + * buggy RMM), don't loop forever, just give up.
> + */
> + if (WARN_ON_ONCE(!can_cancel))
> + return ret;
> + /*
> + * If we have already cancelled, and came back here due
> + * to an error in MEMREQ, then there is no point
> + * in going in loops.
> + */
> + if (WARN_ON_ONCE(cancelled))
> + break;
> + rmi_op_cancel(sro_handle, regs);
> + cancelled = true;
> +
> + if (WARN_ON_ONCE(RMI_RETURN_STATUS(regs->a0) != RMI_INCOMPLETE))
> + return ret;
> + }
> + }
> +
> + if (cancelled)
> + return -ECANCELED;
> +
> + return regs->a0;
> +}
> +EXPORT_SYMBOL_GPL(rmi_sro_memxfer_execute);
> +
> +/* For RMI commands that are stateful but not memory-transferring */
> +long rmi_sro_execute(struct arm_smccc_1_2_regs *regs)
> +{
> + bool cancelled = false;
> + unsigned long sro_handle = regs->a1;
> +
> + rmi_smccc_invoke(regs);
> +
> + sro_handle = regs->a1;
> + while (RMI_RETURN_STATUS(regs->a0) == RMI_INCOMPLETE) {
> + bool can_cancel = RMI_RETURN_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
> +
> + switch (RMI_RETURN_MEMREQ(regs->a0)) {
> + case RMI_OP_MEM_REQ_NONE:
> + rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
> + regs);
> + break;
> + default:
> + WARN_ON_ONCE(1);
> + if (!can_cancel)
> + return regs->a0;
> + /* If we have already cancelled, don't retry this */
> + if (cancelled)
> + return -ECANCELED;
> + rmi_op_cancel(sro_handle, regs);
> + cancelled = true;
> + }
> + }
> +
> + if (cancelled)
> + return -ECANCELED;
> +
> + return regs->a0;
> +}
> +EXPORT_SYMBOL_GPL(rmi_sro_execute);
> +
> static int rmi_check_version(void)
> {
> unsigned short version_major, version_minor;
> diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
> index 9792bf0e00cb9..cd6e309ddf2d8 100644
> --- a/include/linux/arm-rmi-cmds.h
> +++ b/include/linux/arm-rmi-cmds.h
> @@ -8,10 +8,20 @@
>
> #include <linux/arm-smccc-rmi.h>
> #include <linux/bug.h>
> +#include <linux/gfp.h>
> #include <linux/processor.h>
> +#include <linux/string.h>
> #include <linux/types.h>
>
>
> +#define RMI_MAX_ADDR_LIST 256
> +
> +struct rmi_sro_state {
> + struct arm_smccc_1_2_regs regs;
> + int addr_count;
> + unsigned long addr_list[RMI_MAX_ADDR_LIST];
> +};
> +
> /*
> * rmi_smccc_invoke: Invoke the RMI call and return the results in @regs_out
> * @regs_in: Registers with the arguments filled in.
> @@ -33,4 +43,35 @@ static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *regs)
>
> unsigned long rmi_feat_reg(unsigned long index);
>
> +int rmi_delegate_range(phys_addr_t phys, unsigned long size,
> + phys_addr_t *out_phys);
> +int rmi_undelegate_range(phys_addr_t phys, unsigned long size);
> +int rmi_free_delegated_page(phys_addr_t phys);
> +
> +static inline int rmi_delegate_page(phys_addr_t phys)
> +{
> + return rmi_delegate_range(phys, PAGE_SIZE, NULL);
> +}
> +
> +static inline int rmi_undelegate_page(phys_addr_t phys)
> +{
> + return rmi_undelegate_range(phys, PAGE_SIZE);
> +}
> +
> +long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp);
> +void rmi_sro_free(struct rmi_sro_state *sro);
> +long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
> +
> +/*
> + * Resetting the addr_count is sufficient to ignore the addr_list contents.
> + */
> +#define rmi_sro_memxfer_cmd(sro, gfp, ...) ({ \
> + struct rmi_sro_state *__sro = (sro); \
> + __sro->addr_count = 0; \
> + __sro->regs = (struct arm_smccc_1_2_regs){ __VA_ARGS__ }; \
> + long __ret = rmi_sro_memxfer_execute(__sro, gfp); \
> + rmi_sro_free(__sro); \
> + __ret; \
> +})
> +
> #endif
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-14 5:04 ` Gavin Shan
@ 2026-09-14 6:22 ` Suzuki K Poulose
2026-09-14 8:19 ` Suzuki K Poulose
2026-09-14 9:50 ` Gavin Shan
0 siblings, 2 replies; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-14 6:22 UTC (permalink / raw)
To: Gavin Shan, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
Hi Gavin
Thank you for the review, I will address most of them. Responses inline.
On 14/09/2026 06:04, Gavin Shan wrote:
> On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
>> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
>> means that an SMC can return with an operation still in progress. The
>> host is expected to continue the operation until it reaches a conclusion
>> (either success or failure). During this process the RMM can request
>> additional memory ('donate') or hand memory back to the host
>> ('reclaim'). The host can request an in progress operation is cancelled,
>> but still continue the operation until it has completed (otherwise the
>> incomplete operation may cause future RMM operations to fail).
>>
>> The SRO is tracked using a struct rmi_sro_state object which keeps track
>> of any memory which has been allocated but not yet consumed by the RMM
>> or reclaimed from the RMM. This allows the memory to be reused in a
>> future request within the same operation. It will also permit an
>> operation to be done in a context where memory allocation may be
>> difficult (e.g. atomic context) with the option to abort the operation
>> and retry the memory allocation outside of the atomic context. The
>> memory stored in the struct rmi_sro_state object can then be reused on
>> the subsequent attempt.
>>
>> Wrappers for SRO RMI commands are also provided here because they depend
>> on the rmi_sro_execute() implementation added by this patch.
>> Delegate/undelegate handles are also added here because they now use the
>> SRO/stateful command infrastructure and are also used for the memory
>> DONATE/RECLAIM flows.
>>
>> Signed-off-by: Steven Price <steven.price@arm.com>
>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> ---
>> drivers/firmware/arm_rmm/rmi.c | 586 +++++++++++++++++++++++++++++++++
>> include/linux/arm-rmi-cmds.h | 41 +++
>> 2 files changed, 627 insertions(+)
>>
>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/
>> arm_rmm/rmi.c
>> index 5b0e342ce3d58..4f9898ece7547 100644
>> --- a/drivers/firmware/arm_rmm/rmi.c
>> +++ b/drivers/firmware/arm_rmm/rmi.c
>
> I would drop rmi_granule_range_{delegate, undelegate}() by combining
> their logics to
> their only callers rmi_{delegate, undelegate}_range(). More details are
> provided for
> rmi_{delegate, undelegate}_range() in the below.
Ack
>
>> +int rmi_delegate_range(phys_addr_t phys,
>> + unsigned long size,
>> + phys_addr_t *out_phys)
>> +{
>> + long ret = 0;
>> + unsigned long top = phys + size;
>> + unsigned long out_top;
>> +
>> + while (phys < top) {
>> + ret = rmi_granule_range_delegate(phys, top, &out_top);
>> +
>> + if (ret == RMI_SUCCESS) {
>> + /* Buggy RMM ? */
>> + if (WARN_ON(out_top <= phys)) {
>> + rmi_undelegate_range(top - size, size);
>
> [top - size, size] is incorrect because we may be delegating a sub-range
> of the
> range of granules. It's actually the caller's responsibility to
> undelegate the
> graunles that have been delegated.
>
> if (WARN_ON(out_top <= phys)) {
> ret = -ENXIO;
> break;
> }
Agree, I have done this already based on Sashiko review, and added a
comment too.
>
>> +/*
>> + * Convert the RmiAddrBlockSize to actual size. This is used in
>> RmiDonateReq
>> + * and RmiAddrRangeDesc*.
>> + */
>> +static unsigned long rmi_addr_block_size_to_bytes(unsigned long
>> block_size_fld)
>> +{
>> + return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(3 - block_size_fld));
>> +}
>> +
>> +/*
>> + * free_addr_range: Free memory described by the address range entry,
>> which may
>> + * be partially consumed by RMM.
>> + *
>> + * @entry: RMI_ADDR_RANGE descriptor
>> + * @consumed_size: Page aligned size consumed by the RMM from the
>> address range.
>> + *
>> + * If the state of the address is DELEGATED, undelegate it back,
>> before freeing.
>> + * Leaks the memory if we cannot undelegate the range.
>> + */
>> +static void free_addr_range(unsigned long entry, unsigned long
>> consumed_size)
>> +{
>> + unsigned long phys = RMI_ADDR_RANGE_ADDR(entry);
>> + unsigned long block_size_fld = RMI_ADDR_RANGE_BLOCK_SIZE(entry);
>> + unsigned long count = RMI_ADDR_RANGE_COUNT(entry);
>> + unsigned long state = RMI_ADDR_RANGE_STATE(entry);
>> + unsigned long size = rmi_addr_block_size_to_bytes(block_size_fld)
>> * count;
>> +
>> + WARN_ON(!PAGE_ALIGNED(phys) || !PAGE_ALIGNED(consumed_size));
>> +
>> + /* Adjust the address and size for partially consumed entry */
>> + phys += consumed_size;
>> + size -= consumed_size;
>> + /*
>> + * Undelegate the pages back if required. If we can't
>> + * change them back, leak the pages.
>> + */
>> + if (state == RMI_OP_MEM_DELEGATED &&
>> + WARN_ON(rmi_undelegate_range(phys, size)))
>> + return;
>> + free_pages_exact(phys_to_virt(phys), size);
>> +}
>> +
>> +static void rmi_op_continue(unsigned long sro_handle, unsigned long
>> flags,
>> + struct arm_smccc_1_2_regs *out_regs)
>> +{
>> + *out_regs = (struct arm_smccc_1_2_regs) {
>> + SMC_RMI_OP_CONTINUE, sro_handle, flags
>> + };
>> +
>> + rmi_smccc_invoke(out_regs);
>> +}
>> +
>
> The pattern 'regs' is used in some of the 'struct arm_smccc_1_2_regs'
> arguments
> or variables in this series, which is incosistent to the existing
> patterns which
> is either 'args' or 'res' by searching the source files using 'git grep
> arm_smccc_1_2_regs'.
> So I would suggest we have the fixed the pattern 'args' :-)
I would prefer to keep it "regs" as, unlike the smccc_1_1 calls, we
pass "arm_smccc_1_2_regs" for both arguments and results. In this case
we are using a single structure, so, to avoid the confusion, I
intentionally used regs
>
>> +
>> +int rmi_free_delegated_page(phys_addr_t phys)
>> +{
>> + if (WARN_ON_ONCE(rmi_undelegate_page(phys))) {
>> + /* Undelegate failed: leak the page */
>> + return -EBUSY;
>> + }
>> +
>> + free_page((unsigned long)phys_to_virt(phys));
>> +
>> + return 0;
>> +}
>> +EXPORT_SYMBOL_GPL(rmi_free_delegated_page);
>> +
>
> I would move rmi_free_delegated_page() right after rmi_undelegate_range().
Ack
>> +
>> +static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
>> + unsigned long sro_handle,
>> + unsigned long donatereq,
>> + struct arm_smccc_1_2_regs *out_regs,
>> + gfp_t gfp)
>> +{
>> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
>> + unsigned long block_size =
>> rmi_addr_block_size_to_bytes(block_size_fld);
>> + unsigned long count = RMI_DONATE_COUNT(donatereq);
>> + unsigned long state = RMI_DONATE_STATE(donatereq);
>> + unsigned long found = 0;
>> + unsigned long donated_granules;
>> + unsigned long granules_per_block = block_size >> PAGE_SHIFT;
>> + unsigned long consumed_blocks;
>> + int addr_list_start = sro->addr_count;
>> +
> ^^^^^^
>
> Unecessary blank line.
Removed
...
>> +
>> +void rmi_sro_free(struct rmi_sro_state *sro)
>> +{
>> + /* Handle the worse */
>> + if (WARN_ON(sro->addr_count < 0))
>> + return;
>> +
>> + if (WARN_ON(sro->addr_count > RMI_MAX_ADDR_LIST))
>> + sro->addr_count = RMI_MAX_ADDR_LIST;
>> +
>> + for (int i = 0; i < sro->addr_count; i++)
>> + free_addr_range(sro->addr_list[i], 0);
>> +
>> + sro->addr_count = 0;
>> +}
>> +EXPORT_SYMBOL_GPL(rmi_sro_free);
>> +
>> +long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp)
>> +{
>> + struct arm_smccc_1_2_regs *regs = &sro->regs;
>> + bool cancelled = false;
>> + unsigned long sro_handle;
>> +
>> + rmi_smccc_invoke(regs);
>> +
>> + sro_handle = regs->a1;
>> + while (RMI_RETURN_STATUS(regs->a0) == RMI_INCOMPLETE) {
>> + bool can_cancel = RMI_RETURN_CAN_CANCEL(regs->a0) ==
>> RMI_OP_CAN_CANCEL;
>> + int ret = 0;
>> +
>> + switch (RMI_RETURN_MEMREQ(regs->a0)) {
>> + case RMI_OP_MEM_REQ_NONE:
>> + rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
>> + regs);
>> + break;
>> + case RMI_OP_MEM_REQ_DONATE:
>> + ret = rmi_sro_donate(sro, sro_handle, regs->a2, regs,
>> + gfp);
>> + break;
>> + case RMI_OP_MEM_REQ_RECLAIM:
>> + ret = rmi_sro_reclaim(sro, sro_handle, regs);
>> + break;
>> + default:
>> + ret = WARN_ON_ONCE(1);
>> + break;
>
> "ret = WARN_ON_ONCE(1)" is same to "ret = true". I guess we would return
> -EINVAL here.
>
> WARN_ON_ONCE(1);
> ret = -EINVAL;
> break;
Ack, this should be -ENXIO
Cheers
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-14 6:22 ` Suzuki K Poulose
@ 2026-09-14 8:19 ` Suzuki K Poulose
2026-09-14 9:59 ` Gavin Shan
2026-09-14 9:50 ` Gavin Shan
1 sibling, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-14 8:19 UTC (permalink / raw)
To: Gavin Shan, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 14/09/2026 07:22, Suzuki K Poulose wrote:
> Hi Gavin
>
> Thank you for the review, I will address most of them. Responses inline.
>
> On 14/09/2026 06:04, Gavin Shan wrote:
>> On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
>>> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
>>> means that an SMC can return with an operation still in progress. The
>>> host is expected to continue the operation until it reaches a conclusion
>>> (either success or failure). During this process the RMM can request
>>> additional memory ('donate') or hand memory back to the host
>>> ('reclaim'). The host can request an in progress operation is cancelled,
>>> but still continue the operation until it has completed (otherwise the
>>> incomplete operation may cause future RMM operations to fail).
>>>
>>> The SRO is tracked using a struct rmi_sro_state object which keeps track
>>> of any memory which has been allocated but not yet consumed by the RMM
>>> or reclaimed from the RMM. This allows the memory to be reused in a
>>> future request within the same operation. It will also permit an
>>> operation to be done in a context where memory allocation may be
>>> difficult (e.g. atomic context) with the option to abort the operation
>>> and retry the memory allocation outside of the atomic context. The
>>> memory stored in the struct rmi_sro_state object can then be reused on
>>> the subsequent attempt.
>>>
>>> Wrappers for SRO RMI commands are also provided here because they depend
>>> on the rmi_sro_execute() implementation added by this patch.
>>> Delegate/undelegate handles are also added here because they now use the
>>> SRO/stateful command infrastructure and are also used for the memory
>>> DONATE/RECLAIM flows.
>>>
>>> Signed-off-by: Steven Price <steven.price@arm.com>
>>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>
>
>
>>> ---
>>> drivers/firmware/arm_rmm/rmi.c | 586 +++++++++++++++++++++++++++++++++
>>> include/linux/arm-rmi-cmds.h | 41 +++
>>> 2 files changed, 627 insertions(+)
>>>
>>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/
>>> arm_rmm/rmi.c
>>> index 5b0e342ce3d58..4f9898ece7547 100644
>>> --- a/drivers/firmware/arm_rmm/rmi.c
>>> +++ b/drivers/firmware/arm_rmm/rmi.c
>
>>
>> I would drop rmi_granule_range_{delegate, undelegate}() by combining
>> their logics to
>> their only callers rmi_{delegate, undelegate}_range(). More details
>> are provided for
>> rmi_{delegate, undelegate}_range() in the below.
>
> Ack
I have moved them closer, but kept the logic separate, since this is not
a simple RMI call.
Cheers
Suzuki
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-14 8:19 ` Suzuki K Poulose
@ 2026-09-14 9:59 ` Gavin Shan
0 siblings, 0 replies; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 9:59 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/14/26 6:19 PM, Suzuki K Poulose wrote:
> On 14/09/2026 07:22, Suzuki K Poulose wrote:
>> On 14/09/2026 06:04, Gavin Shan wrote:
>>> On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
>>>> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
>>>> means that an SMC can return with an operation still in progress. The
>>>> host is expected to continue the operation until it reaches a conclusion
>>>> (either success or failure). During this process the RMM can request
>>>> additional memory ('donate') or hand memory back to the host
>>>> ('reclaim'). The host can request an in progress operation is cancelled,
>>>> but still continue the operation until it has completed (otherwise the
>>>> incomplete operation may cause future RMM operations to fail).
>>>>
>>>> The SRO is tracked using a struct rmi_sro_state object which keeps track
>>>> of any memory which has been allocated but not yet consumed by the RMM
>>>> or reclaimed from the RMM. This allows the memory to be reused in a
>>>> future request within the same operation. It will also permit an
>>>> operation to be done in a context where memory allocation may be
>>>> difficult (e.g. atomic context) with the option to abort the operation
>>>> and retry the memory allocation outside of the atomic context. The
>>>> memory stored in the struct rmi_sro_state object can then be reused on
>>>> the subsequent attempt.
>>>>
>>>> Wrappers for SRO RMI commands are also provided here because they depend
>>>> on the rmi_sro_execute() implementation added by this patch.
>>>> Delegate/undelegate handles are also added here because they now use the
>>>> SRO/stateful command infrastructure and are also used for the memory
>>>> DONATE/RECLAIM flows.
>>>>
>>>> Signed-off-by: Steven Price <steven.price@arm.com>
>>>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>>>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>>
>>
>>
>>>> ---
>>>> drivers/firmware/arm_rmm/rmi.c | 586 +++++++++++++++++++++++++++++++++
>>>> include/linux/arm-rmi-cmds.h | 41 +++
>>>> 2 files changed, 627 insertions(+)
>>>>
>>>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/ arm_rmm/rmi.c
>>>> index 5b0e342ce3d58..4f9898ece7547 100644
>>>> --- a/drivers/firmware/arm_rmm/rmi.c
>>>> +++ b/drivers/firmware/arm_rmm/rmi.c
>>
>>>
>>> I would drop rmi_granule_range_{delegate, undelegate}() by combining their logics to
>>> their only callers rmi_{delegate, undelegate}_range(). More details are provided for
>>> rmi_{delegate, undelegate}_range() in the below.
>>
>> Ack
>
> I have moved them closer, but kept the logic separate, since this is not
> a simple RMI call.
>
Yeah, it's fine by moving rmi_granule_range_{delegate, undelegate}() to their only
callers. However, the unnecessary if statements can be avoided in their only callers
rmi_{delegate, undelegate)_range(). Besides, rmi_delegate_range() would come before
rmi_undelegate_range().
while (...) {
ret = rmi_granule_range_undelegate(phys, top, &next);
if (ret != RMI_SUCCESS)
break;
/* Buggy RMM ? Let the caller leak the pages */
if (next <= phys) {
ret = -ENXIO;
break;
}
phys = next;
}
The variable 'out_top' in rmi_{delegate, undelegate)_range() may be renamed to 'next',
indicating it's the next (starting) granule for the RMI calls.
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-14 6:22 ` Suzuki K Poulose
2026-09-14 8:19 ` Suzuki K Poulose
@ 2026-09-14 9:50 ` Gavin Shan
1 sibling, 0 replies; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 9:50 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
Hi Suzuki,
On 9/14/26 4:22 PM, Suzuki K Poulose wrote:
> On 14/09/2026 06:04, Gavin Shan wrote:
>> On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
>>> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
>>> means that an SMC can return with an operation still in progress. The
>>> host is expected to continue the operation until it reaches a conclusion
>>> (either success or failure). During this process the RMM can request
>>> additional memory ('donate') or hand memory back to the host
>>> ('reclaim'). The host can request an in progress operation is cancelled,
>>> but still continue the operation until it has completed (otherwise the
>>> incomplete operation may cause future RMM operations to fail).
>>>
>>> The SRO is tracked using a struct rmi_sro_state object which keeps track
>>> of any memory which has been allocated but not yet consumed by the RMM
>>> or reclaimed from the RMM. This allows the memory to be reused in a
>>> future request within the same operation. It will also permit an
>>> operation to be done in a context where memory allocation may be
>>> difficult (e.g. atomic context) with the option to abort the operation
>>> and retry the memory allocation outside of the atomic context. The
>>> memory stored in the struct rmi_sro_state object can then be reused on
>>> the subsequent attempt.
>>>
>>> Wrappers for SRO RMI commands are also provided here because they depend
>>> on the rmi_sro_execute() implementation added by this patch.
>>> Delegate/undelegate handles are also added here because they now use the
>>> SRO/stateful command infrastructure and are also used for the memory
>>> DONATE/RECLAIM flows.
>>>
>>> Signed-off-by: Steven Price <steven.price@arm.com>
>>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
scripts/checkpatch.pl recommends s/Co-Developed-by/Co-developed-by, the same
format issue exists in other patches and please double check.
>>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>
>
>
>>> ---
>>> drivers/firmware/arm_rmm/rmi.c | 586 +++++++++++++++++++++++++++++++++
>>> include/linux/arm-rmi-cmds.h | 41 +++
>>> 2 files changed, 627 insertions(+)
>>>
>>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/ arm_rmm/rmi.c
>>> index 5b0e342ce3d58..4f9898ece7547 100644
>>> --- a/drivers/firmware/arm_rmm/rmi.c
>>> +++ b/drivers/firmware/arm_rmm/rmi.c
[...]
>>> +
>>> +static void rmi_op_continue(unsigned long sro_handle, unsigned long flags,
>>> + struct arm_smccc_1_2_regs *out_regs)
>>> +{
>>> + *out_regs = (struct arm_smccc_1_2_regs) {
>>> + SMC_RMI_OP_CONTINUE, sro_handle, flags
>>> + };
>>> +
>>> + rmi_smccc_invoke(out_regs);
>>> +}
>>> +
>>
>> The pattern 'regs' is used in some of the 'struct arm_smccc_1_2_regs' arguments
>> or variables in this series, which is incosistent to the existing patterns which
>> is either 'args' or 'res' by searching the source files using 'git grep arm_smccc_1_2_regs'.
>> So I would suggest we have the fixed the pattern 'args' :-)
>
>
> I would prefer to keep it "regs" as, unlike the smccc_1_1 calls, we
> pass "arm_smccc_1_2_regs" for both arguments and results. In this case
> we are using a single structure, so, to avoid the confusion, I
> intentionally used regs
>
It's fine to keep "regs" pattern, then the only place using "args" is
rmi_smccc_invoke(). I think the variable or argument names in rmi_smccc_invoke()
can be improved there to use "regs" pattern. With this, we have the unified
pattern "regs".
static inline void rmi_smccc_invoke(struct arm_smccc_1_2_regs *pregs)
{
struct arm_smccc_1_2_regs regs = *pregs;
:
}
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-12 8:36 ` [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
2026-09-14 5:04 ` Gavin Shan
@ 2026-09-14 12:50 ` Sudeep Holla
2026-09-14 14:02 ` Suzuki K Poulose
1 sibling, 1 reply; 25+ messages in thread
From: Sudeep Holla @ 2026-09-14 12:50 UTC (permalink / raw)
To: Suzuki K Poulose
Cc: kvm, kvmarm, maz, Sudeep Holla, will, catalin.marinas,
linux-kernel, linux-arm-kernel, steven.price, aneesh.kumar,
oupton, gshan, joey.gouly, tabba, yuzenghui, linux-coco,
gankulkarni, sdonthineni, alpergun, fj0570is, WeiLin.Chang,
lpieralisi, enju.kohei
On Sat, Sep 12, 2026 at 09:36:07AM +0100, Suzuki K Poulose wrote:
> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
> means that an SMC can return with an operation still in progress. The
> host is expected to continue the operation until it reaches a conclusion
> (either success or failure). During this process the RMM can request
> additional memory ('donate') or hand memory back to the host
> ('reclaim'). The host can request an in progress operation is cancelled,
> but still continue the operation until it has completed (otherwise the
> incomplete operation may cause future RMM operations to fail).
>
> The SRO is tracked using a struct rmi_sro_state object which keeps track
> of any memory which has been allocated but not yet consumed by the RMM
> or reclaimed from the RMM. This allows the memory to be reused in a
> future request within the same operation. It will also permit an
> operation to be done in a context where memory allocation may be
> difficult (e.g. atomic context) with the option to abort the operation
> and retry the memory allocation outside of the atomic context. The
> memory stored in the struct rmi_sro_state object can then be reused on
> the subsequent attempt.
>
> Wrappers for SRO RMI commands are also provided here because they depend
> on the rmi_sro_execute() implementation added by this patch.
> Delegate/undelegate handles are also added here because they now use the
> SRO/stateful command infrastructure and are also used for the memory
> DONATE/RECLAIM flows.
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
[...]
> +static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
> + unsigned long sro_handle,
> + unsigned long donatereq,
> + struct arm_smccc_1_2_regs *out_regs,
> + gfp_t gfp)
> +{
> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
> + unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
> + unsigned long count = RMI_DONATE_COUNT(donatereq);
> + unsigned long state = RMI_DONATE_STATE(donatereq);
> + unsigned long found = 0;
> + unsigned long donated_granules;
> + unsigned long granules_per_block = block_size >> PAGE_SHIFT;
> + unsigned long consumed_blocks;
> + int addr_list_start = sro->addr_count;
> +
> + int ret;
> +
> + for (int i = 0; i < addr_list_start && found < count; i++) {
> + unsigned long entry = sro->addr_list[i];
> +
> + if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
> + RMI_ADDR_RANGE_COUNT(entry) == 1 &&
> + RMI_ADDR_RANGE_STATE(entry) == state) {
> + addr_list_start--;
> + swap(sro->addr_list[addr_list_start],
> + sro->addr_list[i]);
> + found++;
> + i--;
> + }
> + }
> +
> + ret = rmi_sro_ensure_capacity(sro, count - found);
> + if (ret)
> + return ret;
> +
> + while (found < count) {
> + unsigned long addr_range;
> + void *virt = alloc_pages_exact(block_size, gfp);
> + phys_addr_t phys;
> +
> + if (!virt)
> + return -ENOMEM;
> +
> + phys = virt_to_phys(virt);
> +
> + if (state == RMI_OP_MEM_DELEGATED) {
Based on my understanding, rmi_sro_memxfer_execute() is an exported function
and can be invoked by any module. The donatereq argument appears to accept one
of three operations:
RMI_OP_MEM_DELEGATED
RMI_OP_MEM_UNDELEGATED
RMI_OP_MEM_CONDITIONAL
Currently, the check confirming the state is RMI_OP_MEM_DELEGATED occurs
relatively late in the function execution. It seems this function is
explicitly designed to handle only RMI_OP_MEM_DELEGATED.
Given that this is an exported interface, would it make sense to
fail-fast by moving this validation to the very beginning of the function?
Even if RMI_OP_MEM_CONDITIONAL is intended for future use, it should
probably be rejected as invalid for now. Also, it is not clear why the
current check is inside the loop while the state itself doesn't get
modified.
If this is a valid concern, the same architectural pattern should likely be
applied to rmi_sro_donate_contig().
--
Regards,
Sudeep
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-14 12:50 ` Sudeep Holla
@ 2026-09-14 14:02 ` Suzuki K Poulose
2026-09-14 14:47 ` Suzuki K Poulose
0 siblings, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-14 14:02 UTC (permalink / raw)
To: Sudeep Holla
Cc: kvm, kvmarm, maz, will, catalin.marinas, linux-kernel,
linux-arm-kernel, steven.price, aneesh.kumar, oupton, gshan,
joey.gouly, tabba, yuzenghui, linux-coco, gankulkarni,
sdonthineni, alpergun, fj0570is, WeiLin.Chang, lpieralisi,
enju.kohei
On 14/09/2026 13:50, Sudeep Holla wrote:
> On Sat, Sep 12, 2026 at 09:36:07AM +0100, Suzuki K Poulose wrote:
>> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
>> means that an SMC can return with an operation still in progress. The
>> host is expected to continue the operation until it reaches a conclusion
>> (either success or failure). During this process the RMM can request
>> additional memory ('donate') or hand memory back to the host
>> ('reclaim'). The host can request an in progress operation is cancelled,
>> but still continue the operation until it has completed (otherwise the
>> incomplete operation may cause future RMM operations to fail).
>>
>> The SRO is tracked using a struct rmi_sro_state object which keeps track
>> of any memory which has been allocated but not yet consumed by the RMM
>> or reclaimed from the RMM. This allows the memory to be reused in a
>> future request within the same operation. It will also permit an
>> operation to be done in a context where memory allocation may be
>> difficult (e.g. atomic context) with the option to abort the operation
>> and retry the memory allocation outside of the atomic context. The
>> memory stored in the struct rmi_sro_state object can then be reused on
>> the subsequent attempt.
>>
>> Wrappers for SRO RMI commands are also provided here because they depend
>> on the rmi_sro_execute() implementation added by this patch.
>> Delegate/undelegate handles are also added here because they now use the
>> SRO/stateful command infrastructure and are also used for the memory
>> DONATE/RECLAIM flows.
>>
>> Signed-off-by: Steven Price <steven.price@arm.com>
>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>
> [...]
>
>> +static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
>> + unsigned long sro_handle,
>> + unsigned long donatereq,
>> + struct arm_smccc_1_2_regs *out_regs,
>> + gfp_t gfp)
>> +{
>> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
>> + unsigned long block_size = rmi_addr_block_size_to_bytes(block_size_fld);
>> + unsigned long count = RMI_DONATE_COUNT(donatereq);
>> + unsigned long state = RMI_DONATE_STATE(donatereq);
>> + unsigned long found = 0;
>> + unsigned long donated_granules;
>> + unsigned long granules_per_block = block_size >> PAGE_SHIFT;
>> + unsigned long consumed_blocks;
>> + int addr_list_start = sro->addr_count;
>> +
>> + int ret;
>> +
>> + for (int i = 0; i < addr_list_start && found < count; i++) {
>> + unsigned long entry = sro->addr_list[i];
>> +
>> + if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
>> + RMI_ADDR_RANGE_COUNT(entry) == 1 &&
>> + RMI_ADDR_RANGE_STATE(entry) == state) {
>> + addr_list_start--;
>> + swap(sro->addr_list[addr_list_start],
>> + sro->addr_list[i]);
>> + found++;
>> + i--;
>> + }
>> + }
>> +
>> + ret = rmi_sro_ensure_capacity(sro, count - found);
>> + if (ret)
>> + return ret;
>> +
>> + while (found < count) {
>> + unsigned long addr_range;
>> + void *virt = alloc_pages_exact(block_size, gfp);
>> + phys_addr_t phys;
>> +
>> + if (!virt)
>> + return -ENOMEM;
>> +
>> + phys = virt_to_phys(virt);
>> +
>> + if (state == RMI_OP_MEM_DELEGATED) {
>
> Based on my understanding, rmi_sro_memxfer_execute() is an exported function
> and can be invoked by any module. The donatereq argument appears to accept one
> of three operations:
>
> RMI_OP_MEM_DELEGATED
> RMI_OP_MEM_UNDELEGATED
> RMI_OP_MEM_CONDITIONAL
Ack
>
> Currently, the check confirming the state is RMI_OP_MEM_DELEGATED occurs
> relatively late in the function execution. It seems this function is
> explicitly designed to handle only RMI_OP_MEM_DELEGATED.
No, that is not correct. The function handles both OP_MEM_DELEGATED and
OP_MEM_UNDELEGATED. In the former case, we explicitly "delegate" the
pages before donating. The "UNDELEGATED" case doesn't need to do that
extra step.
>
> Given that this is an exported interface, would it make sense to
> fail-fast by moving this validation to the very beginning of the function?
> Even if RMI_OP_MEM_CONDITIONAL is intended for future use, it should
Yep, agree. We can rejec the CONDITIONAL ones.
> probably be rejected as invalid for now. Also, it is not clear why the
> current check is inside the loop while the state itself doesn't get
> modified.
As above, that check is additionally preparing the memory for RMM
consumption.
>
> If this is a valid concern, the same architectural pattern should likely be
> applied to rmi_sro_donate_contig().
>
Ack, we can reject the CONDITIONAL ones.
Suzuki
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO
2026-09-14 14:02 ` Suzuki K Poulose
@ 2026-09-14 14:47 ` Suzuki K Poulose
0 siblings, 0 replies; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-14 14:47 UTC (permalink / raw)
To: Sudeep Holla
Cc: kvm, kvmarm, maz, will, catalin.marinas, linux-kernel,
linux-arm-kernel, steven.price, aneesh.kumar, oupton, gshan,
joey.gouly, tabba, yuzenghui, linux-coco, gankulkarni,
sdonthineni, alpergun, fj0570is, WeiLin.Chang, lpieralisi,
enju.kohei
On 14/09/2026 15:02, Suzuki K Poulose wrote:
> On 14/09/2026 13:50, Sudeep Holla wrote:
>> On Sat, Sep 12, 2026 at 09:36:07AM +0100, Suzuki K Poulose wrote:
>>> RMM v2.0 introduces the concept of "Stateful RMI Operations" (SRO). This
>>> means that an SMC can return with an operation still in progress. The
>>> host is expected to continue the operation until it reaches a conclusion
>>> (either success or failure). During this process the RMM can request
>>> additional memory ('donate') or hand memory back to the host
>>> ('reclaim'). The host can request an in progress operation is cancelled,
>>> but still continue the operation until it has completed (otherwise the
>>> incomplete operation may cause future RMM operations to fail).
>>>
>>> The SRO is tracked using a struct rmi_sro_state object which keeps track
>>> of any memory which has been allocated but not yet consumed by the RMM
>>> or reclaimed from the RMM. This allows the memory to be reused in a
>>> future request within the same operation. It will also permit an
>>> operation to be done in a context where memory allocation may be
>>> difficult (e.g. atomic context) with the option to abort the operation
>>> and retry the memory allocation outside of the atomic context. The
>>> memory stored in the struct rmi_sro_state object can then be reused on
>>> the subsequent attempt.
>>>
>>> Wrappers for SRO RMI commands are also provided here because they depend
>>> on the rmi_sro_execute() implementation added by this patch.
>>> Delegate/undelegate handles are also added here because they now use the
>>> SRO/stateful command infrastructure and are also used for the memory
>>> DONATE/RECLAIM flows.
>>>
>>> Signed-off-by: Steven Price <steven.price@arm.com>
>>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>>
>> [...]
>>
>>> +static int rmi_sro_donate_noncontig(struct rmi_sro_state *sro,
>>> + unsigned long sro_handle,
>>> + unsigned long donatereq,
>>> + struct arm_smccc_1_2_regs *out_regs,
>>> + gfp_t gfp)
>>> +{
>>> + unsigned long block_size_fld = RMI_DONATE_BLOCK_SIZE(donatereq);
>>> + unsigned long block_size =
>>> rmi_addr_block_size_to_bytes(block_size_fld);
>>> + unsigned long count = RMI_DONATE_COUNT(donatereq);
>>> + unsigned long state = RMI_DONATE_STATE(donatereq);
>>> + unsigned long found = 0;
>>> + unsigned long donated_granules;
>>> + unsigned long granules_per_block = block_size >> PAGE_SHIFT;
>>> + unsigned long consumed_blocks;
>>> + int addr_list_start = sro->addr_count;
>>> +
>>> + int ret;
>>> +
>>> + for (int i = 0; i < addr_list_start && found < count; i++) {
>>> + unsigned long entry = sro->addr_list[i];
>>> +
>>> + if (RMI_ADDR_RANGE_BLOCK_SIZE(entry) == block_size_fld &&
>>> + RMI_ADDR_RANGE_COUNT(entry) == 1 &&
>>> + RMI_ADDR_RANGE_STATE(entry) == state) {
>>> + addr_list_start--;
>>> + swap(sro->addr_list[addr_list_start],
>>> + sro->addr_list[i]);
>>> + found++;
>>> + i--;
>>> + }
>>> + }
>>> +
>>> + ret = rmi_sro_ensure_capacity(sro, count - found);
>>> + if (ret)
>>> + return ret;
>>> +
>>> + while (found < count) {
>>> + unsigned long addr_range;
>>> + void *virt = alloc_pages_exact(block_size, gfp);
>>> + phys_addr_t phys;
>>> +
>>> + if (!virt)
>>> + return -ENOMEM;
>>> +
>>> + phys = virt_to_phys(virt);
>>> +
>>> + if (state == RMI_OP_MEM_DELEGATED) {
>>
>> Based on my understanding, rmi_sro_memxfer_execute() is an exported
>> function
>> and can be invoked by any module. The donatereq argument appears to
>> accept one
>> of three operations:
>>
>> RMI_OP_MEM_DELEGATED
>> RMI_OP_MEM_UNDELEGATED
>> RMI_OP_MEM_CONDITIONAL
>
> Ack
>
>>
>> Currently, the check confirming the state is RMI_OP_MEM_DELEGATED occurs
>> relatively late in the function execution. It seems this function is
>> explicitly designed to handle only RMI_OP_MEM_DELEGATED.
>
> No, that is not correct. The function handles both OP_MEM_DELEGATED and
> OP_MEM_UNDELEGATED. In the former case, we explicitly "delegate" the
> pages before donating. The "UNDELEGATED" case doesn't need to do that
> extra step.
>
>>
>> Given that this is an exported interface, would it make sense to
>> fail-fast by moving this validation to the very beginning of the
>> function?
>> Even if RMI_OP_MEM_CONDITIONAL is intended for future use, it should
>
> Yep, agree. We can rejec the CONDITIONAL ones.
For the record, the CONDITIONAL ones are required for self-describing
cases for L1_GPT_CREATE and TRACKING_GRANULE_SET, where the RMM could
accept DELEGATED granules for the objects (if not self describing) or
UNDELEGATED granules (if they are self-describing).
For now, we don't support such RMMs, so will reject the type for now.
Cheers
Suzuki
>
>> probably be rejected as invalid for now. Also, it is not clear why the
>> current check is inside the loop while the state itself doesn't get
>> modified.
>
> As above, that check is additionally preparing the memory for RMM
> consumption.
>
>
>>
>> If this is a valid concern, the same architectural pattern should
>> likely be
>> applied to rmi_sro_donate_contig().
>>
>
> Ack, we can reject the CONDITIONAL ones.
>
> Suzuki
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v18 5/7] firmware: arm_rmm: Activate the RMM
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (3 preceding siblings ...)
2026-09-12 8:36 ` [PATCH v18 4/7] firmware: arm_rmm: Add support for SRO Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 5:06 ` Gavin Shan
2026-09-12 8:36 ` [PATCH v18 6/7] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
2026-09-12 8:36 ` [PATCH v18 7/7] firmware: arm_rmm: Add wrappers for Realm related RMI commands Suzuki K Poulose
6 siblings, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Activate the RMM after the basic configuration. This is a memory transferring,
stateful operation.
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
* Inline RMM_ACTIVATE command and remove the definitions from arm-rmi-cmds.h
* Use scope-based cleanup to free sro object
Changes since v16:
* Split into a new patch
---
drivers/firmware/arm_rmm/rmi.c | 14 +++++++++++++-
1 file changed, 13 insertions(+), 1 deletion(-)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index 4f9898ece7547..ecc89e91d264d 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -762,6 +762,7 @@ static int rmi_configure(void)
static int __init arm64_init_rmi(void)
{
int ret;
+ struct rmi_sro_state *sro __free(kfree) = NULL;
/* Continue without realm support if we can't agree on a version */
ret = rmi_check_version();
@@ -776,7 +777,18 @@ static int __init arm64_init_rmi(void)
if (ret)
return ret;
- return 0;
+ /* Activate the RMM */
+ sro = kmalloc_obj(*sro);
+ if (!sro)
+ return -ENOMEM;
+
+ ret = rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_RMM_ACTIVATE);
+ if (ret) {
+ pr_err("RMM activate failed\n");
+ ret = ret < 0 ? ret : -ENXIO;
+ }
+
+ return ret;
}
/*
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 5/7] firmware: arm_rmm: Activate the RMM
2026-09-12 8:36 ` [PATCH v18 5/7] firmware: arm_rmm: Activate the RMM Suzuki K Poulose
@ 2026-09-14 5:06 ` Gavin Shan
0 siblings, 0 replies; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 5:06 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> Activate the RMM after the basic configuration. This is a memory transferring,
> stateful operation.
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v17:
> * Inline RMM_ACTIVATE command and remove the definitions from arm-rmi-cmds.h
> * Use scope-based cleanup to free sro object
> Changes since v16:
> * Split into a new patch
> ---
> drivers/firmware/arm_rmm/rmi.c | 14 +++++++++++++-
> 1 file changed, 13 insertions(+), 1 deletion(-)
>
Reviewed-by: Gavin Shan <gshan@redhat.com>
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v18 6/7] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (4 preceding siblings ...)
2026-09-12 8:36 ` [PATCH v18 5/7] firmware: arm_rmm: Activate the RMM Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 5:41 ` Gavin Shan
2026-09-12 8:36 ` [PATCH v18 7/7] firmware: arm_rmm: Add wrappers for Realm related RMI commands Suzuki K Poulose
6 siblings, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
The RMM maintains the state of all the granules in the system to make
sure that the host is abiding by the rules. This state can be maintained
at different granularity, per page (TRACKING_FINE) or per region
(TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on the
underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region, all pages
in the region must be of the same state, this implies we need to have "fine"
tracking for DRAM, so that we can delegate individual pages.
For now we only support a statically carved out memory for tracking
granules for the "fine" regions. This can be extended in the future to
allow modifying the tracking granularity and remove the need for a
static allocation by the firmware.
Similarly, the firmware may create L0 GPT entries describing the total
address space. But if we change the "PAS" (Physical Address Space) of a
granule, then the firmware may need to create L1 tables to track the PAS
at a finer granularity. Linux therefore checks if the platform firmware manages
the PAR region. i.e., the firmware is in charge of managing the L1 GPTs
(creation and the required memory for the GPT tables - via static carveouts)
without host intervention. Support for dynamic GPT creation by the host will be
added later.
If the firmware requires us to manage the tracking or GPT memory, Deactivate
the RMM and reclaim any memory donated at RMM activation.
Apply the same checks when hotplugged memory is brought online.
Signed-off-by: Steven Price <steven.price@arm.com>
[ Switch to RMI_GPT_L1_INFO for checking GPTs and deactivate RMM ]
Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
* Move wrappers that may not be used elsewhere, out of arm-rmi-cmds.h
Changes since v16:
* Check fine tracking and create L1 GPTs for hotplug-added memory.
* Clarify the L1 GPT setup and move the explanatory comment.
* Switch to using RMI_GPT_INFO command for checking the GPTs.
* Deactivate the RMM and reclaim the memory if we can't proceed.
Changes since v15:
* Skip firmware-reserved NOMAP memory in rmi_init_metadata()
* Handle negative error codes from wrappers.
Changes since v14:
* Move the implementation into drivers/firmware/arm_rmm.
Changes since v13:
* Moved out of KVM
---
drivers/firmware/arm_rmm/rmi.c | 200 +++++++++++++++++++++++++++++++++
include/linux/arm-rmi-cmds.h | 2 +
2 files changed, 202 insertions(+)
diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
index ecc89e91d264d..583e1aca9b15a 100644
--- a/drivers/firmware/arm_rmm/rmi.c
+++ b/drivers/firmware/arm_rmm/rmi.c
@@ -5,12 +5,15 @@
#include <linux/cpufeature.h>
#include <linux/memblock.h>
+#include <linux/memory.h>
#include <linux/arm-rmi-cmds.h>
#include <linux/slab.h>
#include <asm/memory.h>
#include <asm/pgtable-hwdef.h>
+static bool arm64_rmi_is_available;
+
/* Currently only the first 2 registers are used by Linux */
#define RMI_FEAT_REG_COUNT 2
static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT] __ro_after_init;
@@ -68,6 +71,69 @@ static inline long rmi_granule_range_undelegate(unsigned long base,
return ret;
}
+/**
+ * rmi_granule_tracking_get() - Get configuration of a Granule tracking region
+ * @start: Base PA of the tracking region
+ * @end: End of the PA region
+ * @out_category: Memory category
+ * @out_state: Tracking region state
+ * @out_top: Top of the memory region
+ *
+ * Return: RMI return code
+ */
+static inline int rmi_granule_tracking_get(unsigned long start,
+ unsigned long end,
+ unsigned long *out_category,
+ unsigned long *out_state,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GRANULE_TRACKING_GET, start, end,
+ };
+
+ rmi_smccc_invoke(®s);
+
+ if (regs.a0 != RMI_SUCCESS)
+ return regs.a0;
+
+ if (out_category)
+ *out_category = regs.a1;
+ if (out_state)
+ *out_state = regs.a2;
+ if (out_top)
+ *out_top = regs.a3;
+
+ return RMI_SUCCESS;
+}
+
+/*
+ * rmi_gpt_info - Query the GPT info for the given PAR.
+ * @start: Base of the physical address region
+ * @top: Top of the physical address region
+ * @out_top: Top of the phyiscal address region for which
+ * the GPT @out_gpt_par_state is valid
+ * @out_gpt_par_state: State of the GPT covered by [start, out_top)
+ */
+static inline long rmi_gpt_info(unsigned long start, unsigned long end,
+ unsigned long *out_top,
+ unsigned long *out_gpt_par_state)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_GPT_INFO, start, end,
+ };
+
+ rmi_smccc_invoke(®s);
+ if (regs.a0 != RMI_SUCCESS)
+ return regs.a0;
+
+ if (out_top)
+ *out_top = regs.a1;
+ if (out_gpt_par_state)
+ *out_gpt_par_state = regs.a2;
+
+ return RMI_SUCCESS;
+}
+
/**
* rmi_features() - Read feature register
* @index: Feature register index
@@ -759,6 +825,124 @@ static int rmi_configure(void)
return ret;
}
+/*
+ * Make sure the area is tracked by RMM at FINE granularity.
+ * We do not support changing the tracking yet.
+ */
+static int rmi_verify_memory_tracking(phys_addr_t start, phys_addr_t end)
+{
+ while (start < end) {
+ unsigned long ret, category, state, next;
+
+ ret = rmi_granule_tracking_get(start, end, &category, &state, &next);
+ if (ret != RMI_SUCCESS)
+ return -ENOMEM;
+
+ if (state != RMI_TRACKING_FINE ||
+ category != RMI_MEM_CATEGORY_CONVENTIONAL) {
+ /* TODO: Set granule tracking in this case */
+ pr_err("Granule tracking for region isn't fine/conventional: %llx-%lx\n",
+ start, next);
+ return -ENODEV;
+ }
+ start = next;
+ }
+
+ return 0;
+}
+
+/*
+ * We do not support creating L1 GPTs yet. So, make sure that
+ * all the regions are managed by the firmware.
+ */
+static int rmi_verify_gpt_firmware_managed(phys_addr_t start, phys_addr_t end)
+{
+ unsigned long l0gpt_sz;
+ unsigned long next, par_state;
+
+ l0gpt_sz = 1UL << (30 + FIELD_GET(RMI_FEATURE_REGISTER_1_L0GPTSZ,
+ rmi_feat_reg(1)));
+ start = ALIGN_DOWN(start, l0gpt_sz);
+ end = ALIGN(end, l0gpt_sz);
+
+ while (start < end) {
+ long ret = rmi_gpt_info(start, end, &next, &par_state);
+
+ if (ret != RMI_SUCCESS)
+ return -ENOMEM;
+
+ if (par_state != RMI_GPT_PAR_PLAT) {
+ pr_err("GPT for the region is not managed by firmware %llx-%lx\n",
+ start, next);
+ return -ENOMEM;
+ }
+ start = next;
+ }
+
+ return 0;
+}
+
+static int rmi_prepare_memory(phys_addr_t start, phys_addr_t end)
+{
+ int ret;
+
+ ret = rmi_verify_memory_tracking(start, end);
+ if (ret)
+ return ret;
+
+ return rmi_verify_gpt_firmware_managed(start, end);
+}
+
+static int rmi_init_metadata(void)
+{
+ phys_addr_t start, end;
+ struct memblock_region *r;
+
+ for_each_mem_region(r) {
+ int ret;
+
+ /* Firmware-reserved NOMAP regions are not usable system RAM */
+ if (memblock_is_nomap(r))
+ continue;
+
+ start = memblock_region_memory_base_pfn(r) << PAGE_SHIFT;
+ end = memblock_region_memory_end_pfn(r) << PAGE_SHIFT;
+
+ ret = rmi_prepare_memory(start, end);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
+static int rmi_memory_notifier(struct notifier_block *nb,
+ unsigned long action, void *data)
+{
+ struct memory_notify *arg = data;
+ phys_addr_t start, end;
+ int ret;
+
+ if (action != MEM_GOING_ONLINE)
+ return NOTIFY_DONE;
+
+ start = PFN_PHYS(arg->start_pfn);
+ end = PFN_PHYS(arg->start_pfn + arg->nr_pages);
+ ret = rmi_prepare_memory(start, end);
+
+ return notifier_from_errno(ret);
+}
+
+static struct notifier_block rmi_memory_nb = {
+ .notifier_call = rmi_memory_notifier,
+};
+
+bool is_rmi_available(void)
+{
+ return arm64_rmi_is_available;
+}
+EXPORT_SYMBOL_GPL(is_rmi_available);
+
static int __init arm64_init_rmi(void)
{
int ret;
@@ -786,8 +970,24 @@ static int __init arm64_init_rmi(void)
if (ret) {
pr_err("RMM activate failed\n");
ret = ret < 0 ? ret : -ENXIO;
+ return ret;
}
+ ret = rmi_init_metadata();
+ if (ret)
+ goto out_deactivate;
+
+ ret = register_memory_notifier(&rmi_memory_nb);
+ if (ret)
+ goto out_deactivate;
+
+ arm64_rmi_is_available = true;
+ pr_info("RMI configured\n");
+
+ return 0;
+
+out_deactivate:
+ WARN_ON(rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_RMM_DEACTIVATE));
return ret;
}
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
index cd6e309ddf2d8..27a0c56944976 100644
--- a/include/linux/arm-rmi-cmds.h
+++ b/include/linux/arm-rmi-cmds.h
@@ -58,6 +58,8 @@ static inline int rmi_undelegate_page(phys_addr_t phys)
return rmi_undelegate_range(phys, PAGE_SIZE);
}
+bool is_rmi_available(void);
+
long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp);
void rmi_sro_free(struct rmi_sro_state *sro);
long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 6/7] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-12 8:36 ` [PATCH v18 6/7] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
@ 2026-09-14 5:41 ` Gavin Shan
2026-09-14 8:38 ` Suzuki K Poulose
0 siblings, 1 reply; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 5:41 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> The RMM maintains the state of all the granules in the system to make
> sure that the host is abiding by the rules. This state can be maintained
> at different granularity, per page (TRACKING_FINE) or per region
> (TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on the
> underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region, all pages
> in the region must be of the same state, this implies we need to have "fine"
> tracking for DRAM, so that we can delegate individual pages.
>
> For now we only support a statically carved out memory for tracking
> granules for the "fine" regions. This can be extended in the future to
> allow modifying the tracking granularity and remove the need for a
> static allocation by the firmware.
>
> Similarly, the firmware may create L0 GPT entries describing the total
> address space. But if we change the "PAS" (Physical Address Space) of a
> granule, then the firmware may need to create L1 tables to track the PAS
> at a finer granularity. Linux therefore checks if the platform firmware manages
> the PAR region. i.e., the firmware is in charge of managing the L1 GPTs
> (creation and the required memory for the GPT tables - via static carveouts)
> without host intervention. Support for dynamic GPT creation by the host will be
> added later.
>
> If the firmware requires us to manage the tracking or GPT memory, Deactivate
> the RMM and reclaim any memory donated at RMM activation.
>
> Apply the same checks when hotplugged memory is brought online.
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> [ Switch to RMI_GPT_L1_INFO for checking GPTs and deactivate RMM ]
> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v17:
> * Move wrappers that may not be used elsewhere, out of arm-rmi-cmds.h
> Changes since v16:
> * Check fine tracking and create L1 GPTs for hotplug-added memory.
> * Clarify the L1 GPT setup and move the explanatory comment.
> * Switch to using RMI_GPT_INFO command for checking the GPTs.
> * Deactivate the RMM and reclaim the memory if we can't proceed.
> Changes since v15:
> * Skip firmware-reserved NOMAP memory in rmi_init_metadata()
> * Handle negative error codes from wrappers.
> Changes since v14:
> * Move the implementation into drivers/firmware/arm_rmm.
> Changes since v13:
> * Moved out of KVM
> ---
> drivers/firmware/arm_rmm/rmi.c | 200 +++++++++++++++++++++++++++++++++
> include/linux/arm-rmi-cmds.h | 2 +
> 2 files changed, 202 insertions(+)
>
> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/arm_rmm/rmi.c
> index ecc89e91d264d..583e1aca9b15a 100644
> --- a/drivers/firmware/arm_rmm/rmi.c
> +++ b/drivers/firmware/arm_rmm/rmi.c
> @@ -5,12 +5,15 @@
>
> #include <linux/cpufeature.h>
> #include <linux/memblock.h>
> +#include <linux/memory.h>
> #include <linux/arm-rmi-cmds.h>
> #include <linux/slab.h>
>
> #include <asm/memory.h>
> #include <asm/pgtable-hwdef.h>
>
> +static bool arm64_rmi_is_available;
> +
> /* Currently only the first 2 registers are used by Linux */
> #define RMI_FEAT_REG_COUNT 2
> static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT] __ro_after_init;
> @@ -68,6 +71,69 @@ static inline long rmi_granule_range_undelegate(unsigned long base,
> return ret;
> }
>
> +/**
> + * rmi_granule_tracking_get() - Get configuration of a Granule tracking region
> + * @start: Base PA of the tracking region
> + * @end: End of the PA region
> + * @out_category: Memory category
> + * @out_state: Tracking region state
> + * @out_top: Top of the memory region
> + *
> + * Return: RMI return code
> + */
> +static inline int rmi_granule_tracking_get(unsigned long start,
> + unsigned long end,
> + unsigned long *out_category,
> + unsigned long *out_state,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_GRANULE_TRACKING_GET, start, end,
> + };
> +
> + rmi_smccc_invoke(®s);
> +
> + if (regs.a0 != RMI_SUCCESS)
> + return regs.a0;
> +
> + if (out_category)
> + *out_category = regs.a1;
> + if (out_state)
> + *out_state = regs.a2;
> + if (out_top)
> + *out_top = regs.a3;
> +
> + return RMI_SUCCESS;
> +}
> +
> +/*
> + * rmi_gpt_info - Query the GPT info for the given PAR.
> + * @start: Base of the physical address region
> + * @top: Top of the physical address region
> + * @out_top: Top of the phyiscal address region for which
> + * the GPT @out_gpt_par_state is valid
> + * @out_gpt_par_state: State of the GPT covered by [start, out_top)
> + */
> +static inline long rmi_gpt_info(unsigned long start, unsigned long end,
> + unsigned long *out_top,
> + unsigned long *out_gpt_par_state)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_GPT_INFO, start, end,
> + };
> +
> + rmi_smccc_invoke(®s);
> + if (regs.a0 != RMI_SUCCESS)
> + return regs.a0;
> +
> + if (out_top)
> + *out_top = regs.a1;
> + if (out_gpt_par_state)
> + *out_gpt_par_state = regs.a2;
> +
> + return RMI_SUCCESS;
> +}
> +
> /**
> * rmi_features() - Read feature register
> * @index: Feature register index
> @@ -759,6 +825,124 @@ static int rmi_configure(void)
> return ret;
> }
>
> +/*
> + * Make sure the area is tracked by RMM at FINE granularity.
> + * We do not support changing the tracking yet.
> + */
> +static int rmi_verify_memory_tracking(phys_addr_t start, phys_addr_t end)
> +{
> + while (start < end) {
> + unsigned long ret, category, state, next;
> +
> + ret = rmi_granule_tracking_get(start, end, &category, &state, &next);
> + if (ret != RMI_SUCCESS)
> + return -ENOMEM;
> +
> + if (state != RMI_TRACKING_FINE ||
> + category != RMI_MEM_CATEGORY_CONVENTIONAL) {
> + /* TODO: Set granule tracking in this case */
> + pr_err("Granule tracking for region isn't fine/conventional: %llx-%lx\n",
> + start, next);
> + return -ENODEV;
> + }
> + start = next;
> + }
> +
> + return 0;
> +}
> +
I would suggest to drop rmi_granule_tracking_get() by combining its logics into
rmi_verify_memory_tracking().
/*
* Make sure the area is tracked by RMM at FINE granularity.
* We do not support changing the tracking yet.
*/
static int rmi_verify_memory_tracking(phys_addr_t start, phys_addr_t end)
{
struct arm_smccc_1_2_regs args;
while (start < end) {
args.a0 = SMC_RMI_GRANULE_TRACKING_GET;
args.a1 = start;
args.a2 = end;
rmi_smccc_invoke(&args);
if (args.a0 != RMI_SUCCESS)
return -ENOMEM;
if (args.a1 != RMI_MEM_CATEGORY_CONVENTIONAL ||
args.a2 != RMI_TRACKING_FINE) {
/* TODO: Set granule tracking in this case */
pr_err("Granule tracking for region isn't fine/conventional: %llx-%lx\n",
start, args.a3);
return -ENODEV;
}
start = args.a3;
}
return 0;
}
> +/*
> + * We do not support creating L1 GPTs yet. So, make sure that
> + * all the regions are managed by the firmware.
> + */
> +static int rmi_verify_gpt_firmware_managed(phys_addr_t start, phys_addr_t end)
> +{
> + unsigned long l0gpt_sz;
> + unsigned long next, par_state;
> +
> + l0gpt_sz = 1UL << (30 + FIELD_GET(RMI_FEATURE_REGISTER_1_L0GPTSZ,
> + rmi_feat_reg(1)));
> + start = ALIGN_DOWN(start, l0gpt_sz);
> + end = ALIGN(end, l0gpt_sz);
> +
> + while (start < end) {
> + long ret = rmi_gpt_info(start, end, &next, &par_state);
> +
> + if (ret != RMI_SUCCESS)
> + return -ENOMEM;
> +
> + if (par_state != RMI_GPT_PAR_PLAT) {
> + pr_err("GPT for the region is not managed by firmware %llx-%lx\n",
> + start, next);
> + return -ENOMEM;
> + }
> + start = next;
> + }
> +
> + return 0;
> +}
I would suggest to drop rmi_gpt_info() by combining its logics into rmi_verify_gpt_firmware_managed().
/*
* We do not support creating L1 GPTs yet. So, make sure that
* all the regions are managed by the firmware.
*/
static int rmi_verify_gpt_firmware_managed(phys_addr_t start, phys_addr_t end)
{
struct arm_smccc_1_2_regs args;
unsigned long l0gpt_sz;
l0gpt_sz = 1UL << (30 + FIELD_GET(RMI_FEATURE_REGISTER_1_L0GPTSZ,
rmi_feat_reg(1)));
start = ALIGN_DOWN(start, l0gpt_sz);
end = ALIGN(end, l0gpt_sz);
while (start < end) {
args.a0 = SMC_RMI_GPT_INFO;
args.a1 = start;
args.a2 = end;
rmi_smccc_invoke(&args);
if (args.a0 != RMI_SUCCESS)
return -ENOMEM;
if (args.a2 != RMI_GPT_PAR_PLAT) {
pr_err("GPT for the region is not managed by firmware %llx-%lx\n",
start, args.a1);
return -ENODEV;
}
start = args.a2;
}
return 0;
}
> +
> +static int rmi_prepare_memory(phys_addr_t start, phys_addr_t end)
> +{
> + int ret;
> +
> + ret = rmi_verify_memory_tracking(start, end);
> + if (ret)
> + return ret;
> +
> + return rmi_verify_gpt_firmware_managed(start, end);
> +}
> +
> +static int rmi_init_metadata(void)
> +{
> + phys_addr_t start, end;
> + struct memblock_region *r;
> +
> + for_each_mem_region(r) {
> + int ret;
> +
> + /* Firmware-reserved NOMAP regions are not usable system RAM */
> + if (memblock_is_nomap(r))
> + continue;
> +
> + start = memblock_region_memory_base_pfn(r) << PAGE_SHIFT;
> + end = memblock_region_memory_end_pfn(r) << PAGE_SHIFT;
> +
nit: unnecessary to covert the physical address (struct memblock_region::base and (base + size))
to PFN and then convert PFN to physical address:
start = PAGE_ALIGN(r->base);
end = PAGE_ALIGN_DOWN(r->base + r->size);
if (start >= end)
continue;
> + ret = rmi_prepare_memory(start, end);
> + if (ret)
> + return ret;
> + }
> +
> + return 0;
> +}
> +
> +static int rmi_memory_notifier(struct notifier_block *nb,
> + unsigned long action, void *data)
> +{
> + struct memory_notify *arg = data;
> + phys_addr_t start, end;
> + int ret;
> +
> + if (action != MEM_GOING_ONLINE)
> + return NOTIFY_DONE;
> +
> + start = PFN_PHYS(arg->start_pfn);
> + end = PFN_PHYS(arg->start_pfn + arg->nr_pages);
> + ret = rmi_prepare_memory(start, end);
> +
> + return notifier_from_errno(ret);
> +}
> +
> +static struct notifier_block rmi_memory_nb = {
> + .notifier_call = rmi_memory_notifier,
> +};
> +
> +bool is_rmi_available(void)
> +{
> + return arm64_rmi_is_available;
> +}
> +EXPORT_SYMBOL_GPL(is_rmi_available);
> +
> static int __init arm64_init_rmi(void)
> {
> int ret;
> @@ -786,8 +970,24 @@ static int __init arm64_init_rmi(void)
> if (ret) {
> pr_err("RMM activate failed\n");
> ret = ret < 0 ? ret : -ENXIO;
> + return ret;
> }
>
> + ret = rmi_init_metadata();
> + if (ret)
> + goto out_deactivate;
> +
> + ret = register_memory_notifier(&rmi_memory_nb);
> + if (ret)
> + goto out_deactivate;
> +
> + arm64_rmi_is_available = true;
> + pr_info("RMI configured\n");
> +
> + return 0;
> +
> +out_deactivate:
> + WARN_ON(rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_RMM_DEACTIVATE));
> return ret;
> }
>
> diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
> index cd6e309ddf2d8..27a0c56944976 100644
> --- a/include/linux/arm-rmi-cmds.h
> +++ b/include/linux/arm-rmi-cmds.h
> @@ -58,6 +58,8 @@ static inline int rmi_undelegate_page(phys_addr_t phys)
> return rmi_undelegate_range(phys, PAGE_SIZE);
> }
>
> +bool is_rmi_available(void);
> +
> long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp);
> void rmi_sro_free(struct rmi_sro_state *sro);
> long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v18 6/7] firmware: arm_rmm: Ensure the RMM has GPT entries for memory
2026-09-14 5:41 ` Gavin Shan
@ 2026-09-14 8:38 ` Suzuki K Poulose
0 siblings, 0 replies; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-14 8:38 UTC (permalink / raw)
To: Gavin Shan, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 14/09/2026 06:41, Gavin Shan wrote:
> On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
>> From: Steven Price <steven.price@arm.com>
>>
>> The RMM maintains the state of all the granules in the system to make
>> sure that the host is abiding by the rules. This state can be maintained
>> at different granularity, per page (TRACKING_FINE) or per region
>> (TRACKING_COARSE or TRACKING_INTERMEDIATE). The region size depends on
>> the
>> underlying "RMI_GRANULE_SIZE". For a "coarse"/"intermediate" region,
>> all pages
>> in the region must be of the same state, this implies we need to have
>> "fine"
>> tracking for DRAM, so that we can delegate individual pages.
>>
>> For now we only support a statically carved out memory for tracking
>> granules for the "fine" regions. This can be extended in the future to
>> allow modifying the tracking granularity and remove the need for a
>> static allocation by the firmware.
>>
>> Similarly, the firmware may create L0 GPT entries describing the total
>> address space. But if we change the "PAS" (Physical Address Space) of a
>> granule, then the firmware may need to create L1 tables to track the PAS
>> at a finer granularity. Linux therefore checks if the platform
>> firmware manages
>> the PAR region. i.e., the firmware is in charge of managing the L1 GPTs
>> (creation and the required memory for the GPT tables - via static
>> carveouts)
>> without host intervention. Support for dynamic GPT creation by the
>> host will be
>> added later.
>>
>> If the firmware requires us to manage the tracking or GPT memory,
>> Deactivate
>> the RMM and reclaim any memory donated at RMM activation.
>>
>> Apply the same checks when hotplugged memory is brought online.
>>
>> Signed-off-by: Steven Price <steven.price@arm.com>
>> [ Switch to RMI_GPT_L1_INFO for checking GPTs and deactivate RMM ]
>> Co-Developed-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
>> ---
>> Changes since v17:
>> * Move wrappers that may not be used elsewhere, out of arm-rmi-
>> cmds.h
>> Changes since v16:
>> * Check fine tracking and create L1 GPTs for hotplug-added memory.
>> * Clarify the L1 GPT setup and move the explanatory comment.
>> * Switch to using RMI_GPT_INFO command for checking the GPTs.
>> * Deactivate the RMM and reclaim the memory if we can't proceed.
>> Changes since v15:
>> * Skip firmware-reserved NOMAP memory in rmi_init_metadata()
>> * Handle negative error codes from wrappers.
>> Changes since v14:
>> * Move the implementation into drivers/firmware/arm_rmm.
>> Changes since v13:
>> * Moved out of KVM
>> ---
>> drivers/firmware/arm_rmm/rmi.c | 200 +++++++++++++++++++++++++++++++++
>> include/linux/arm-rmi-cmds.h | 2 +
>> 2 files changed, 202 insertions(+)
>>
>> diff --git a/drivers/firmware/arm_rmm/rmi.c b/drivers/firmware/
>> arm_rmm/rmi.c
>> index ecc89e91d264d..583e1aca9b15a 100644
>> --- a/drivers/firmware/arm_rmm/rmi.c
>> +++ b/drivers/firmware/arm_rmm/rmi.c
>> @@ -5,12 +5,15 @@
>> #include <linux/cpufeature.h>
>> #include <linux/memblock.h>
>> +#include <linux/memory.h>
>> #include <linux/arm-rmi-cmds.h>
>> #include <linux/slab.h>
>> #include <asm/memory.h>
>> #include <asm/pgtable-hwdef.h>
>> +static bool arm64_rmi_is_available;
>> +
>> /* Currently only the first 2 registers are used by Linux */
>> #define RMI_FEAT_REG_COUNT 2
>> static unsigned long rmi_feat_reg_cache[RMI_FEAT_REG_COUNT]
>> __ro_after_init;
>> @@ -68,6 +71,69 @@ static inline long
>> rmi_granule_range_undelegate(unsigned long base,
>> return ret;
>> }
>> +/**
>> + * rmi_granule_tracking_get() - Get configuration of a Granule
>> tracking region
>> + * @start: Base PA of the tracking region
>> + * @end: End of the PA region
>> + * @out_category: Memory category
>> + * @out_state: Tracking region state
>> + * @out_top: Top of the memory region
>> + *
>> + * Return: RMI return code
>> + */
>> +static inline int rmi_granule_tracking_get(unsigned long start,
>> + unsigned long end,
>> + unsigned long *out_category,
>> + unsigned long *out_state,
>> + unsigned long *out_top)
>> +{
>> + struct arm_smccc_1_2_regs regs = {
>> + SMC_RMI_GRANULE_TRACKING_GET, start, end,
>> + };
>> +
>> + rmi_smccc_invoke(®s);
>> +
>> + if (regs.a0 != RMI_SUCCESS)
>> + return regs.a0;
>> +
>> + if (out_category)
>> + *out_category = regs.a1;
>> + if (out_state)
>> + *out_state = regs.a2;
>> + if (out_top)
>> + *out_top = regs.a3;
>> +
>> + return RMI_SUCCESS;
>> +}
>> +
>> +/*
>> + * rmi_gpt_info - Query the GPT info for the given PAR.
>> + * @start: Base of the physical address region
>> + * @top: Top of the physical address region
>> + * @out_top: Top of the phyiscal address region for which
>> + * the GPT @out_gpt_par_state is valid
>> + * @out_gpt_par_state: State of the GPT covered by [start, out_top)
>> + */
>> +static inline long rmi_gpt_info(unsigned long start, unsigned long end,
>> + unsigned long *out_top,
>> + unsigned long *out_gpt_par_state)
>> +{
>> + struct arm_smccc_1_2_regs regs = {
>> + SMC_RMI_GPT_INFO, start, end,
>> + };
>> +
>> + rmi_smccc_invoke(®s);
>> + if (regs.a0 != RMI_SUCCESS)
>> + return regs.a0;
>> +
>> + if (out_top)
>> + *out_top = regs.a1;
>> + if (out_gpt_par_state)
>> + *out_gpt_par_state = regs.a2;
>> +
>> + return RMI_SUCCESS;
>> +}
>> +
>> /**
>> * rmi_features() - Read feature register
>> * @index: Feature register index
>> @@ -759,6 +825,124 @@ static int rmi_configure(void)
>> return ret;
>> }
>> +/*
>> + * Make sure the area is tracked by RMM at FINE granularity.
>> + * We do not support changing the tracking yet.
>> + */
>> +static int rmi_verify_memory_tracking(phys_addr_t start, phys_addr_t
>> end)
>> +{
>> + while (start < end) {
>> + unsigned long ret, category, state, next;
>> +
>> + ret = rmi_granule_tracking_get(start, end, &category, &state,
>> &next);
>> + if (ret != RMI_SUCCESS)
>> + return -ENOMEM;
>> +
>> + if (state != RMI_TRACKING_FINE ||
>> + category != RMI_MEM_CATEGORY_CONVENTIONAL) {
>> + /* TODO: Set granule tracking in this case */
>> + pr_err("Granule tracking for region isn't fine/
>> conventional: %llx-%lx\n",
>> + start, next);
>> + return -ENODEV;
>> + }
>> + start = next;
>> + }
>> +
>> + return 0;
>> +}
>> +
>
> I would suggest to drop rmi_granule_tracking_get() by combining its
> logics into
> rmi_verify_memory_tracking().
>
> /*
> * Make sure the area is tracked by RMM at FINE granularity.
> * We do not support changing the tracking yet.
> */
> static int rmi_verify_memory_tracking(phys_addr_t start, phys_addr_t end)
> {
> struct arm_smccc_1_2_regs args;
>
> while (start < end) {
> args.a0 = SMC_RMI_GRANULE_TRACKING_GET;
> args.a1 = start;
> args.a2 = end;
> rmi_smccc_invoke(&args);
>
> if (args.a0 != RMI_SUCCESS)
> return -ENOMEM;
>
> if (args.a1 != RMI_MEM_CATEGORY_CONVENTIONAL ||
> args.a2 != RMI_TRACKING_FINE) {
> /* TODO: Set granule tracking in this case */
> pr_err("Granule tracking for region isn't fine/
> conventional: %llx-%lx\n",
> start, args.a3);
> return -ENODEV;
I have moved the wrappers closer to the caller, but retained them to
make it easier to read the code.
e.g., error conditions around next >= start etc.
> }
>
> start = args.a3;
> }
>
> return 0;
> }
...
>
> I would suggest to drop rmi_gpt_info() by combining its logics into
> rmi_verify_gpt_firmware_managed().
>
Same as above.
>> +
>> +static int rmi_init_metadata(void)
>> +{
>> + phys_addr_t start, end;
>> + struct memblock_region *r;
>> +
>> + for_each_mem_region(r) {
>> + int ret;
>> +
>> + /* Firmware-reserved NOMAP regions are not usable system RAM */
>> + if (memblock_is_nomap(r))
>> + continue;
>> +
>> + start = memblock_region_memory_base_pfn(r) << PAGE_SHIFT;
>> + end = memblock_region_memory_end_pfn(r) << PAGE_SHIFT;
>> +
>
> nit: unnecessary to covert the physical address (struct
> memblock_region::base and (base + size))
> to PFN and then convert PFN to physical address:
>
> start = PAGE_ALIGN(r->base);
> end = PAGE_ALIGN_DOWN(r->base + r->size);
> if (start >= end)
> continue;
Yes, I have handled this already per Sashiko review comment.
Thank you for the review!
Cheers
Suzuki
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v18 7/7] firmware: arm_rmm: Add wrappers for Realm related RMI commands
2026-09-12 8:36 [PATCH v18 0/7] firmware: arm_rmm: Add RMM v2.0 base RMI support Suzuki K Poulose
` (5 preceding siblings ...)
2026-09-12 8:36 ` [PATCH v18 6/7] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Suzuki K Poulose
@ 2026-09-12 8:36 ` Suzuki K Poulose
2026-09-14 5:46 ` Gavin Shan
6 siblings, 1 reply; 25+ messages in thread
From: Suzuki K Poulose @ 2026-09-12 8:36 UTC (permalink / raw)
To: kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, gshan, joey.gouly, tabba,
yuzenghui, linux-coco, gankulkarni, sdonthineni, alpergun,
fj0570is, WeiLin.Chang, lpieralisi, enju.kohei, Suzuki K Poulose
From: Steven Price <steven.price@arm.com>
Introduce wrappers for the RMI functions needed for creating and
managing realm guests. This will be used by the KVM to manage the
Realms
Signed-off-by: Steven Price <steven.price@arm.com>
Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
---
Changes since v17:
* Avoid nesting if conditions for output populating RMI calls
* Clean up comments
* Always use arm_smccc_1_2_invoke() for RMI calls.
Changes since v16:
* Split into a separate patch and move away from arch/arm64 to
include/linux/.
* Also moved into the firmware_rmm series from the KVM CCA support.
This is done in a hope to reduce the merge conflicts and make
the KVM CCA upstreaming in independent parallel chunks
---
include/linux/arm-rmi-cmds.h | 461 +++++++++++++++++++++++++++++++++++
1 file changed, 461 insertions(+)
diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
index 27a0c56944976..fd7ef7e82b246 100644
--- a/include/linux/arm-rmi-cmds.h
+++ b/include/linux/arm-rmi-cmds.h
@@ -76,4 +76,465 @@ long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
__ret; \
})
+/**
+ * rmi_rtt_data_map_init() - Create a mapping at protected IPA, copying contents
+ * from a given non-secure source granule.
+ * @rd: PA of the RD
+ * @data: PA of the target granule mapped in the guest
+ * @ipa: IPA at which the granule @data will be mapped in the guest
+ * @src: PA of the source granule with contents
+ * @flags: RMI_MEASURE_CONTENT if the contents should be measured
+ *
+ * Create a mapping from Protected IPA space to conventional memory, copying
+ * contents from a Non-secure Granule provided by the caller.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_data_map_init(unsigned long rd, unsigned long data,
+ unsigned long ipa, unsigned long src,
+ unsigned long flags)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DATA_MAP_INIT, rd, data, ipa, src, flags
+ };
+
+ return rmi_sro_execute(®s);
+}
+
+/**
+ * rmi_rtt_data_map() - Create mappings in protected IPA range with unknown contents
+ * @rd: PA of the RD
+ * @base: Base of the target IPA range
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Top address of range which was processed.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_data_map(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DATA_MAP, rd, base, top, flags, oaddr
+ };
+ long ret;
+
+ ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_data_unmap() - Remove mappings to conventional memory at a protected
+ * IPA range
+ * @rd: PA of the RD
+ * @base: Base of the target IPA range
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Returns top IPA of range which has been unmapped
+ * @out_range: Output address range
+ * @out_count: Number of entries in output address list
+ *
+ * Removes mappings to convention memory with a target Protected IPA range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_data_unmap(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top,
+ unsigned long *out_range,
+ unsigned long *out_count)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DATA_UNMAP, rd, base, top, flags, oaddr
+ };
+ long ret;
+
+ ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS) {
+ if (out_top)
+ *out_top = regs.a1;
+ if (out_range)
+ *out_range = regs.a2;
+ if (out_count)
+ *out_count = regs.a3;
+ }
+
+ return ret;
+}
+
+/**
+ * rmi_psci_complete() - Complete pending PSCI command
+ * @calling_rec: PA of the calling REC
+ * @status: Status of the PSCI request
+ *
+ * Completes a pending PSCI command.
+ *
+ * Return: RMI return code
+ */
+static inline long rmi_psci_complete(unsigned long calling_rec,
+ unsigned long status)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_PSCI_COMPLETE, calling_rec, status,
+ };
+
+ rmi_smccc_invoke(®s);
+ return regs.a0;
+}
+
+/**
+ * rmi_realm_activate() - Activate a realm
+ * @rd: PA of the RD
+ *
+ * Mark a realm as Active, signalling that creation is completed, allowing
+ * execution of the realm.
+ *
+ * Return: RMI return code
+ */
+static inline long rmi_realm_activate(unsigned long rd)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_REALM_ACTIVATE, rd,
+ };
+
+ rmi_smccc_invoke(®s);
+ return regs.a0;
+}
+
+/**
+ * rmi_realm_create() - Create a realm
+ * @rd: PA of the RD
+ * @params: PA of realm parameters
+ * @sro: Preallocated SRO context
+ *
+ * Create a new realm using the given parameters.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_realm_create(unsigned long rd, unsigned long params,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REALM_CREATE, rd, params);
+}
+
+/**
+ * rmi_realm_terminate() - Terminate a realm
+ * @rd: PA of the RD
+ * @sro: Preallocated SRO context
+ *
+ * Terminates a realm, moving it into a ZOMBIE state
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_realm_terminate(unsigned long rd,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REALM_TERMINATE, rd);
+}
+
+/**
+ * rmi_realm_destroy() - Destroy a realm
+ * @rd: PA of the RD
+ * @sro: Preallocated SRO context
+ *
+ * Destroys a realm, all objects belonging to the realm must be destroyed first.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_realm_destroy(unsigned long rd,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
+ SMC_RMI_REALM_DESTROY, rd);
+}
+
+/**
+ * rmi_rec_create() - Create a REC
+ * @rd: PA of the RD
+ * @rec: PA of the target REC
+ * @params: PA of REC parameters
+ * @sro: Allocated SRO context to be used
+ *
+ * Create a REC using the parameters specified in the struct rec_params pointed
+ * to by @params.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rec_create(unsigned long rd,
+ unsigned long rec,
+ unsigned long params,
+ struct rmi_sro_state *sro)
+{
+ long ret;
+
+ sro->addr_count = 0;
+ sro->regs = (struct arm_smccc_1_2_regs) {
+ SMC_RMI_REC_CREATE, rd, rec, params
+ };
+ ret = rmi_sro_memxfer_execute(sro, GFP_KERNEL);
+ rmi_sro_free(sro);
+
+ return ret;
+}
+
+/**
+ * rmi_rec_destroy() - Destroy a REC
+ * @rec: PA of the target REC
+ * @sro: Allocated SRO context to be used
+ *
+ * Destroys a REC. The REC must not be running.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rec_destroy(unsigned long rec,
+ struct rmi_sro_state *sro)
+{
+ return rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_REC_DESTROY, rec);
+}
+
+/**
+ * rmi_rec_enter() - Enter a REC
+ * @rec: PA of the target REC
+ * @run_ptr: PA of RecRun structure
+ *
+ * Starts (or continues) execution within a REC.
+ *
+ * Return: RMI return code
+ */
+static inline long rmi_rec_enter(unsigned long rec, unsigned long run_ptr)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_REC_ENTER, rec, run_ptr,
+ };
+
+ rmi_smccc_invoke(®s);
+ return regs.a0;
+}
+
+/**
+ * rmi_rtt_create() - Creates an RTT
+ * @rd: PA of the RD
+ * @rtt: PA of the target RTT
+ * @ipa: Base of the IPA range described by the RTT
+ * @level: Depth of the RTT within the tree
+ *
+ * Creates an RTT (Realm Translation Table) at the specified level for the
+ * translation of the specified address within the realm.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_create(unsigned long rd, unsigned long rtt,
+ unsigned long ipa, long level)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_CREATE, rd, rtt, ipa, level
+ };
+
+ return rmi_sro_execute(®s);
+}
+
+/**
+ * rmi_rtt_destroy() - Destroy an RTT
+ * @rd: PA of the RD
+ * @ipa: Base of the IPA range described by the RTT
+ * @level: RTT level
+ * @out_rtt: Pointer to write the PA of the RTT which was destroyed
+ * @out_top: Pointer to write the top IPA of non-live RTT entries, from entry
+ * at which the RTT walk terminated.
+ *
+ * Destroys an RTT. The RTT must be non-live, i.e. none of the entries in the
+ * table are in ASSIGNED or TABLE state.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code.
+ */
+static inline long rmi_rtt_destroy(unsigned long rd,
+ unsigned long ipa,
+ long level,
+ unsigned long *out_rtt,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_DESTROY, rd, ipa, level
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret != RMI_SUCCESS)
+ return ret;
+
+ if (out_rtt)
+ *out_rtt = regs.a1;
+ if (out_top)
+ *out_top = regs.a2;
+
+ return RMI_SUCCESS;
+}
+
+/**
+ * rmi_rtt_fold() - Fold an RTT
+ * @rd: PA of the RD
+ * @ipa: Base of the IPA range described by the RTT
+ * @level: Depth of the RTT within the tree
+ * @out_rtt: Pointer to write the PA of the RTT which was destroyed
+ *
+ * Folds an RTT. If all entries with the RTT are 'homogeneous' the RTT can be
+ * folded into the parent and the RTT destroyed.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_fold(unsigned long rd, unsigned long ipa,
+ long level, unsigned long *out_rtt)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_FOLD, rd, ipa, level
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_rtt)
+ *out_rtt = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_init_ripas() - Set RIPAS for new realm
+ * @rd: PA of the RD
+ * @base: Base of target IPA region
+ * @top: Top of target IPA region
+ * @out_top: Top IPA of range whose RIPAS was modified
+ *
+ * Sets the RIPAS of a target IPA range to RAM, for a realm in the NEW state.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_init_ripas(unsigned long rd, unsigned long base,
+ unsigned long top, unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_INIT_RIPAS, rd, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_unprot_map() - Map unprotected granules into a realm
+ * @rd: PA of the RD
+ * @base: Base IPA of the mapping
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Top IPA of range which has been mapped
+ *
+ * Create mappings to memory within a target unprotected IPA range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_unprot_map(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_UNPROT_MAP, rd, base, top, flags, oaddr
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_set_ripas() - Set RIPAS for an running realm
+ * @rd: PA of the RD
+ * @rec: PA of the REC making the request
+ * @base: Base of target IPA region
+ * @top: Top of target IPA region
+ * @out_top: Pointer to write top IPA of range whose RIPAS was modified
+ *
+ * Completes a request made by the realm to change the RIPAS of a target IPA
+ * range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_set_ripas(unsigned long rd, unsigned long rec,
+ unsigned long base, unsigned long top,
+ unsigned long *out_top)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_SET_RIPAS, rd, rec, base, top
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret == RMI_SUCCESS && out_top)
+ *out_top = regs.a1;
+
+ return ret;
+}
+
+/**
+ * rmi_rtt_unprot_unmap() - Remove mappings within an unprotected IPA range
+ * @rd: PA of the RD
+ * @base: Base IPA of the mapping
+ * @top: Top of the target IPA range
+ * @flags: Flags
+ * @oaddr: Output address set descriptor
+ * @out_top: Top IPA which has been unmapped
+ * @out_range: Output address range
+ * @out_count: Number of entries in output address list
+ *
+ * Removes mappings to memory within a target unprotected IPA range.
+ *
+ * Return: 0 on success, positive RMI result code or negative Linux error code
+ */
+static inline long rmi_rtt_unprot_unmap(unsigned long rd,
+ unsigned long base,
+ unsigned long top,
+ unsigned long flags,
+ unsigned long oaddr,
+ unsigned long *out_top,
+ unsigned long *out_range,
+ unsigned long *out_count)
+{
+ struct arm_smccc_1_2_regs regs = {
+ SMC_RMI_RTT_UNPROT_UNMAP, rd, base, top, flags, oaddr
+ };
+ long ret = rmi_sro_execute(®s);
+
+ if (ret != RMI_SUCCESS)
+ return ret;
+
+ if (out_top)
+ *out_top = regs.a1;
+ if (out_range)
+ *out_range = regs.a2;
+ if (out_count)
+ *out_count = regs.a3;
+
+ return RMI_SUCCESS;
+}
+
#endif
--
2.43.0
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v18 7/7] firmware: arm_rmm: Add wrappers for Realm related RMI commands
2026-09-12 8:36 ` [PATCH v18 7/7] firmware: arm_rmm: Add wrappers for Realm related RMI commands Suzuki K Poulose
@ 2026-09-14 5:46 ` Gavin Shan
0 siblings, 0 replies; 25+ messages in thread
From: Gavin Shan @ 2026-09-14 5:46 UTC (permalink / raw)
To: Suzuki K Poulose, kvm, kvmarm
Cc: maz, will, catalin.marinas, linux-kernel, linux-arm-kernel,
steven.price, aneesh.kumar, oupton, joey.gouly, tabba, yuzenghui,
linux-coco, gankulkarni, sdonthineni, alpergun, fj0570is,
WeiLin.Chang, lpieralisi, enju.kohei
On 9/12/26 6:36 PM, Suzuki K Poulose wrote:
> From: Steven Price <steven.price@arm.com>
>
> Introduce wrappers for the RMI functions needed for creating and
> managing realm guests. This will be used by the KVM to manage the
> Realms
>
> Signed-off-by: Steven Price <steven.price@arm.com>
> Signed-off-by: Suzuki K Poulose <suzuki.poulose@arm.com>
> ---
> Changes since v17:
> * Avoid nesting if conditions for output populating RMI calls
> * Clean up comments
> * Always use arm_smccc_1_2_invoke() for RMI calls.
> Changes since v16:
> * Split into a separate patch and move away from arch/arm64 to
> include/linux/.
> * Also moved into the firmware_rmm series from the KVM CCA support.
> This is done in a hope to reduce the merge conflicts and make
> the KVM CCA upstreaming in independent parallel chunks
> ---
> include/linux/arm-rmi-cmds.h | 461 +++++++++++++++++++++++++++++++++++
> 1 file changed, 461 insertions(+)
>
One nitpick for rmi_rtt_data_unmap() below. With it addressed:
Reviewed-by: Gavin Shan <gshan@redhat.com>
> diff --git a/include/linux/arm-rmi-cmds.h b/include/linux/arm-rmi-cmds.h
> index 27a0c56944976..fd7ef7e82b246 100644
> --- a/include/linux/arm-rmi-cmds.h
> +++ b/include/linux/arm-rmi-cmds.h
> @@ -76,4 +76,465 @@ long rmi_sro_execute(struct arm_smccc_1_2_regs *regs);
> __ret; \
> })
>
> +/**
> + * rmi_rtt_data_map_init() - Create a mapping at protected IPA, copying contents
> + * from a given non-secure source granule.
> + * @rd: PA of the RD
> + * @data: PA of the target granule mapped in the guest
> + * @ipa: IPA at which the granule @data will be mapped in the guest
> + * @src: PA of the source granule with contents
> + * @flags: RMI_MEASURE_CONTENT if the contents should be measured
> + *
> + * Create a mapping from Protected IPA space to conventional memory, copying
> + * contents from a Non-secure Granule provided by the caller.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_data_map_init(unsigned long rd, unsigned long data,
> + unsigned long ipa, unsigned long src,
> + unsigned long flags)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_DATA_MAP_INIT, rd, data, ipa, src, flags
> + };
> +
> + return rmi_sro_execute(®s);
> +}
> +
> +/**
> + * rmi_rtt_data_map() - Create mappings in protected IPA range with unknown contents
> + * @rd: PA of the RD
> + * @base: Base of the target IPA range
> + * @top: Top of the target IPA range
> + * @flags: Flags
> + * @oaddr: Output address set descriptor
> + * @out_top: Top address of range which was processed.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_data_map(unsigned long rd,
> + unsigned long base,
> + unsigned long top,
> + unsigned long flags,
> + unsigned long oaddr,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_DATA_MAP, rd, base, top, flags, oaddr
> + };
> + long ret;
> +
> + ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_top)
> + *out_top = regs.a1;
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_rtt_data_unmap() - Remove mappings to conventional memory at a protected
> + * IPA range
> + * @rd: PA of the RD
> + * @base: Base of the target IPA range
> + * @top: Top of the target IPA range
> + * @flags: Flags
> + * @oaddr: Output address set descriptor
> + * @out_top: Returns top IPA of range which has been unmapped
> + * @out_range: Output address range
> + * @out_count: Number of entries in output address list
> + *
> + * Removes mappings to convention memory with a target Protected IPA range.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_data_unmap(unsigned long rd,
> + unsigned long base,
> + unsigned long top,
> + unsigned long flags,
> + unsigned long oaddr,
> + unsigned long *out_top,
> + unsigned long *out_range,
> + unsigned long *out_count)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_DATA_UNMAP, rd, base, top, flags, oaddr
> + };
> + long ret;
> +
> + ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS) {
> + if (out_top)
> + *out_top = regs.a1;
> + if (out_range)
> + *out_range = regs.a2;
> + if (out_count)
> + *out_count = regs.a3;
> + }
> +
> + return ret;
> +}
> +
The unnecessary if statements could be avoided by:
if (ret != RMI_SUCCESS)
return ret;
if (out_top)
*out_top = regs.a1;
if (out_range)
*out_range = regs.a2;
if (out_count)
*out_count = regs.a3;
return RMI_SUCCESS;
> +/**
> + * rmi_psci_complete() - Complete pending PSCI command
> + * @calling_rec: PA of the calling REC
> + * @status: Status of the PSCI request
> + *
> + * Completes a pending PSCI command.
> + *
> + * Return: RMI return code
> + */
> +static inline long rmi_psci_complete(unsigned long calling_rec,
> + unsigned long status)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_PSCI_COMPLETE, calling_rec, status,
> + };
> +
> + rmi_smccc_invoke(®s);
> + return regs.a0;
> +}
> +
> +/**
> + * rmi_realm_activate() - Activate a realm
> + * @rd: PA of the RD
> + *
> + * Mark a realm as Active, signalling that creation is completed, allowing
> + * execution of the realm.
> + *
> + * Return: RMI return code
> + */
> +static inline long rmi_realm_activate(unsigned long rd)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_REALM_ACTIVATE, rd,
> + };
> +
> + rmi_smccc_invoke(®s);
> + return regs.a0;
> +}
> +
> +/**
> + * rmi_realm_create() - Create a realm
> + * @rd: PA of the RD
> + * @params: PA of realm parameters
> + * @sro: Preallocated SRO context
> + *
> + * Create a new realm using the given parameters.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_realm_create(unsigned long rd, unsigned long params,
> + struct rmi_sro_state *sro)
> +{
> + return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
> + SMC_RMI_REALM_CREATE, rd, params);
> +}
> +
> +/**
> + * rmi_realm_terminate() - Terminate a realm
> + * @rd: PA of the RD
> + * @sro: Preallocated SRO context
> + *
> + * Terminates a realm, moving it into a ZOMBIE state
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_realm_terminate(unsigned long rd,
> + struct rmi_sro_state *sro)
> +{
> + return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
> + SMC_RMI_REALM_TERMINATE, rd);
> +}
> +
> +/**
> + * rmi_realm_destroy() - Destroy a realm
> + * @rd: PA of the RD
> + * @sro: Preallocated SRO context
> + *
> + * Destroys a realm, all objects belonging to the realm must be destroyed first.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_realm_destroy(unsigned long rd,
> + struct rmi_sro_state *sro)
> +{
> + return rmi_sro_memxfer_cmd(sro, GFP_KERNEL,
> + SMC_RMI_REALM_DESTROY, rd);
> +}
> +
> +/**
> + * rmi_rec_create() - Create a REC
> + * @rd: PA of the RD
> + * @rec: PA of the target REC
> + * @params: PA of REC parameters
> + * @sro: Allocated SRO context to be used
> + *
> + * Create a REC using the parameters specified in the struct rec_params pointed
> + * to by @params.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rec_create(unsigned long rd,
> + unsigned long rec,
> + unsigned long params,
> + struct rmi_sro_state *sro)
> +{
> + long ret;
> +
> + sro->addr_count = 0;
> + sro->regs = (struct arm_smccc_1_2_regs) {
> + SMC_RMI_REC_CREATE, rd, rec, params
> + };
> + ret = rmi_sro_memxfer_execute(sro, GFP_KERNEL);
> + rmi_sro_free(sro);
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_rec_destroy() - Destroy a REC
> + * @rec: PA of the target REC
> + * @sro: Allocated SRO context to be used
> + *
> + * Destroys a REC. The REC must not be running.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rec_destroy(unsigned long rec,
> + struct rmi_sro_state *sro)
> +{
> + return rmi_sro_memxfer_cmd(sro, GFP_KERNEL, SMC_RMI_REC_DESTROY, rec);
> +}
> +
> +/**
> + * rmi_rec_enter() - Enter a REC
> + * @rec: PA of the target REC
> + * @run_ptr: PA of RecRun structure
> + *
> + * Starts (or continues) execution within a REC.
> + *
> + * Return: RMI return code
> + */
> +static inline long rmi_rec_enter(unsigned long rec, unsigned long run_ptr)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_REC_ENTER, rec, run_ptr,
> + };
> +
> + rmi_smccc_invoke(®s);
> + return regs.a0;
> +}
> +
> +/**
> + * rmi_rtt_create() - Creates an RTT
> + * @rd: PA of the RD
> + * @rtt: PA of the target RTT
> + * @ipa: Base of the IPA range described by the RTT
> + * @level: Depth of the RTT within the tree
> + *
> + * Creates an RTT (Realm Translation Table) at the specified level for the
> + * translation of the specified address within the realm.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_create(unsigned long rd, unsigned long rtt,
> + unsigned long ipa, long level)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_CREATE, rd, rtt, ipa, level
> + };
> +
> + return rmi_sro_execute(®s);
> +}
> +
> +/**
> + * rmi_rtt_destroy() - Destroy an RTT
> + * @rd: PA of the RD
> + * @ipa: Base of the IPA range described by the RTT
> + * @level: RTT level
> + * @out_rtt: Pointer to write the PA of the RTT which was destroyed
> + * @out_top: Pointer to write the top IPA of non-live RTT entries, from entry
> + * at which the RTT walk terminated.
> + *
> + * Destroys an RTT. The RTT must be non-live, i.e. none of the entries in the
> + * table are in ASSIGNED or TABLE state.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code.
> + */
> +static inline long rmi_rtt_destroy(unsigned long rd,
> + unsigned long ipa,
> + long level,
> + unsigned long *out_rtt,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_DESTROY, rd, ipa, level
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret != RMI_SUCCESS)
> + return ret;
> +
> + if (out_rtt)
> + *out_rtt = regs.a1;
> + if (out_top)
> + *out_top = regs.a2;
> +
> + return RMI_SUCCESS;
> +}
> +
> +/**
> + * rmi_rtt_fold() - Fold an RTT
> + * @rd: PA of the RD
> + * @ipa: Base of the IPA range described by the RTT
> + * @level: Depth of the RTT within the tree
> + * @out_rtt: Pointer to write the PA of the RTT which was destroyed
> + *
> + * Folds an RTT. If all entries with the RTT are 'homogeneous' the RTT can be
> + * folded into the parent and the RTT destroyed.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_fold(unsigned long rd, unsigned long ipa,
> + long level, unsigned long *out_rtt)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_FOLD, rd, ipa, level
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_rtt)
> + *out_rtt = regs.a1;
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_rtt_init_ripas() - Set RIPAS for new realm
> + * @rd: PA of the RD
> + * @base: Base of target IPA region
> + * @top: Top of target IPA region
> + * @out_top: Top IPA of range whose RIPAS was modified
> + *
> + * Sets the RIPAS of a target IPA range to RAM, for a realm in the NEW state.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_init_ripas(unsigned long rd, unsigned long base,
> + unsigned long top, unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_INIT_RIPAS, rd, base, top
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_top)
> + *out_top = regs.a1;
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_rtt_unprot_map() - Map unprotected granules into a realm
> + * @rd: PA of the RD
> + * @base: Base IPA of the mapping
> + * @top: Top of the target IPA range
> + * @flags: Flags
> + * @oaddr: Output address set descriptor
> + * @out_top: Top IPA of range which has been mapped
> + *
> + * Create mappings to memory within a target unprotected IPA range.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_unprot_map(unsigned long rd,
> + unsigned long base,
> + unsigned long top,
> + unsigned long flags,
> + unsigned long oaddr,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_UNPROT_MAP, rd, base, top, flags, oaddr
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_top)
> + *out_top = regs.a1;
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_rtt_set_ripas() - Set RIPAS for an running realm
> + * @rd: PA of the RD
> + * @rec: PA of the REC making the request
> + * @base: Base of target IPA region
> + * @top: Top of target IPA region
> + * @out_top: Pointer to write top IPA of range whose RIPAS was modified
> + *
> + * Completes a request made by the realm to change the RIPAS of a target IPA
> + * range.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_set_ripas(unsigned long rd, unsigned long rec,
> + unsigned long base, unsigned long top,
> + unsigned long *out_top)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_SET_RIPAS, rd, rec, base, top
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret == RMI_SUCCESS && out_top)
> + *out_top = regs.a1;
> +
> + return ret;
> +}
> +
> +/**
> + * rmi_rtt_unprot_unmap() - Remove mappings within an unprotected IPA range
> + * @rd: PA of the RD
> + * @base: Base IPA of the mapping
> + * @top: Top of the target IPA range
> + * @flags: Flags
> + * @oaddr: Output address set descriptor
> + * @out_top: Top IPA which has been unmapped
> + * @out_range: Output address range
> + * @out_count: Number of entries in output address list
> + *
> + * Removes mappings to memory within a target unprotected IPA range.
> + *
> + * Return: 0 on success, positive RMI result code or negative Linux error code
> + */
> +static inline long rmi_rtt_unprot_unmap(unsigned long rd,
> + unsigned long base,
> + unsigned long top,
> + unsigned long flags,
> + unsigned long oaddr,
> + unsigned long *out_top,
> + unsigned long *out_range,
> + unsigned long *out_count)
> +{
> + struct arm_smccc_1_2_regs regs = {
> + SMC_RMI_RTT_UNPROT_UNMAP, rd, base, top, flags, oaddr
> + };
> + long ret = rmi_sro_execute(®s);
> +
> + if (ret != RMI_SUCCESS)
> + return ret;
> +
> + if (out_top)
> + *out_top = regs.a1;
> + if (out_range)
> + *out_range = regs.a2;
> + if (out_count)
> + *out_count = regs.a3;
> +
> + return RMI_SUCCESS;
> +}
> +
> #endif
Thanks,
Gavin
^ permalink raw reply [flat|nested] 25+ messages in thread