From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from PH8PR06CU001.outbound.protection.outlook.com (mail-westus3azon11012017.outbound.protection.outlook.com [40.107.209.17]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DB51D47143B for ; Tue, 21 Jul 2026 20:02:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.107.209.17 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784664180; cv=fail; b=SS52lU3bwjgD9Jk6DLGAlw/pzmdJIVZ5JKwCnR+45qYwozDyUa5mAy/tI528nMS3Jhy16swlAGr+s4YHoUC+fMnAp6VbutD90QeVfynmtQLRc5EkhkpkzaOcTmPBM4VudC6ZzIchMhNKjmlSeoDM4dHJIKESn+PBerepG9VId1o= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784664180; c=relaxed/simple; bh=p3MWuAj/T3KG66ZhSi+IGeEhwFoxZWuNkyD8aQbcoXk=; h=Message-ID:Date:Subject:To:Cc:References:From:In-Reply-To: Content-Type:MIME-Version; b=WPUWeUyugNONUknD7cfdrtUdHerxC2kP5Ld1PXzeOtDw/F2b5MhyCjV7AG2/QpRdzRHgBatp/Zma0lR7/c3CWNC5Ieum6uFPpysKGalbY712m7YnKPn4j7ZaPrCjPFK6bFqtmleSk1EzXEwcvO8Q9HARLnYS+4G6sVG8fQ27KMo= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com; spf=fail smtp.mailfrom=amd.com; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b=1XNIjdP2; arc=fail smtp.client-ip=40.107.209.17 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=amd.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=amd.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=amd.com header.i=@amd.com header.b="1XNIjdP2" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=sJZzjaDW09AfKQiagXlJ+027wcLZQYZFmbBuSaMpH9Pk9DvGdnJkXPMXeEiFUWzsBCC460xAGBQ27KyedBBBNL/FqQM2J1+MWKRXgFWG4pRKDpkL3LsYaM9qx+3V7CfEK/QvgD1zXklPX7t1gq3PhKl4BjtVgg/yzPORjrJ+/+i41GyairqHvvof+HBaagSajRTlUZfB/wG7VKhFdNEjIEetlaw3oSnD5X9wUMaY5MeS9HIAsEYi0s6GjmWRditWp9txu/rMHcUje4aCIz4cvnvsldI+p9EQIa7a2wncHqDvFq93+nT+7PrhEca4FYGe85UNpiTyKPvWm6JkFH88zA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=+iZrtZKrmuONao6k0TI4Gsl9SGYguQUd48WKD/Mg9kQ=; b=HK4x+fVcP42d0viAw67ck2hXRgqRwQs1EEeHpHGTYMkcVlRNnOK1taRQQm0RjCqe+kJ/+PjRQaZoCFfHMsN6IWje3S7SthM13R0nntg3MGpZAa7F+yfeaIL0XN+233rF7lsSMU6UMaIZ1MEjE3HmT1JaVGVqWx31xn0SgeVMn/id/FaMwWMi2di0TU92AcxnEdwLNMUNAZRn4xsd7BqETho6oVe/P3/aALDugfX+fi57KbsIHWsiOqBF+N9eHEUvunWHNmAoRSV8QFkti9VXfCdsJhuAoznK4Uhdhr7gMInBxMeaTWp6UvTR2m7DyCX0l/nlPxONzv4vmKvNXlTrrg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=amd.com; dmarc=pass action=none header.from=amd.com; dkim=pass header.d=amd.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=amd.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=+iZrtZKrmuONao6k0TI4Gsl9SGYguQUd48WKD/Mg9kQ=; b=1XNIjdP2wmiyJBRxGzbBJZHqRNwKbNCVpoqaOVhRQXeqBMDeVGvkU8jRxPFnwRY29XN8b7Duo1eVPgMSpIUJzVV0VnXyx+YizdfoJLFiiqRNmEqUnLFHrK6NL1kUOS+WiIf4LBlcv27dFPxbezD1XAHtKB1Yxoy2oeofTRGtNH4= Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=amd.com; Received: from BL1PR12MB5320.namprd12.prod.outlook.com (2603:10b6:208:314::17) by DSVPR12MB999191.namprd12.prod.outlook.com (2603:10b6:8:496::9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.202.27; Tue, 21 Jul 2026 20:02:54 +0000 Received: from BL1PR12MB5320.namprd12.prod.outlook.com ([fe80::1876:4a6d:2cf5:b8d1]) by BL1PR12MB5320.namprd12.prod.outlook.com ([fe80::1876:4a6d:2cf5:b8d1%5]) with mapi id 15.21.0245.009; Tue, 21 Jul 2026 20:02:54 +0000 Message-ID: <21e614b8-50fa-49e4-87c3-e6bdb4e83ab1@amd.com> Date: Tue, 21 Jul 2026 15:02:52 -0500 User-Agent: Mozilla Thunderbird Subject: Re: [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept To: Reinette Chatre , Ben Horgan , Fenghua Yu , Tony Luck , James Morse , Dave Martin , Drew Fustini , Chen Yu Cc: Borislav Petkov , Thomas Gleixner , Dave Hansen , Peter Newman , "x86@kernel.org" , "linux-kernel@vger.kernel.org" References: <5ee87762-1898-4b62-94da-85b3e9917ecc@intel.com> <62701203-c4a3-4ec2-a9af-602e1fc15863@nvidia.com> <8f9f78dd-e3f5-4b35-bc72-0eb5dafdcedf@nvidia.com> <36163a81-9737-49e3-93ef-6c392f7272f0@intel.com> <0fc6df54-26c7-43fa-948a-528cd94937f1@arm.com> <9049378c-699a-4155-b1e4-737a1d7265d5@intel.com> <57740b97-80ee-4632-bca3-dc43cd7776c2@arm.com> <44f26cd4-be79-476e-b002-7ccfb7705179@intel.com> <749bd904-523d-4e9d-8493-0e8cfd79949e@arm.com> <9db33feb-cf04-420c-a99a-e31e4b8e4954@arm.com> <8fd6caed-820f-457a-a1ef-a0a006fa52aa@intel.com> <4ef15dde-2fbb-4763-93b6-4333b02d6859@arm.com> <7b751c28-2f04-42b7-b957-af6447e7f824@intel.com> <34b95afb-8b60-4680-9ad1-90c5b24e8fb7@arm.com> <08f016bc-2ba6-439e-bb3e-20061166402c@intel.com> Content-Language: en-US From: Babu Moger In-Reply-To: <08f016bc-2ba6-439e-bb3e-20061166402c@intel.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit X-ClientProxiedBy: CH3P221CA0013.NAMP221.PROD.OUTLOOK.COM (2603:10b6:610:1e7::21) To BL1PR12MB5320.namprd12.prod.outlook.com (2603:10b6:208:314::17) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BL1PR12MB5320:EE_|DSVPR12MB999191:EE_ X-MS-Office365-Filtering-Correlation-Id: 4493ef64-39cc-4a68-d7f2-08dee7630d24 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|23010399003|7416014|366016|1800799024|6133799003|4143699003|11063799006|56012099006|5023799004|10067099003|18002099003|22082099003|3023799007; X-Microsoft-Antispam-Message-Info: 7g3Eqrzko4b0Wsm/W+kO1jMFZkKmGQi2/6yNLPx3n0fLqvrpKRejRYe/+ONEysIh9Zk3nYLUxrIjyWjSmduxyegWukGyDwDH1oEglz6IA9NISHb6hwSwt6ZldEcQPasOub4vaaExgQl56yWck/zi1aZgvM7r7gAYCiZ7uxcK504dfp8dunsT1U8icfvbuOOvFGHsZKyaj7CAPnfo0rUCt3OZ92vsvV4331E2q9DaIR6pBgAGNwThCw5smMtnpbw0LII/JFDf32irJWmLISqu9eqcV9pnw0zbunkDZ923FSSPhM1JH2HBNZyGJ0JGopHSFGURVEfhTQJNop3KuEBchFWrW3RW77A6Is+mphkwT+VUOBpo0sbQD5Zn7cD/j3So799EXoAY8eXCfWxH36FPC62BxIDxzlKJ0Znc3xKJd/s08ZDcpetWFEOKaWNQcCqNJMM7M0wc0fjfJ1iDX0wLuZaqCYkJ344BVZVNkNMcWj4q//wdnfVabv31BlnYji3uizBgr/2ioJDwbF1FWP34hRmwOKTMxraeKisIKLnZ5vUfJKqQXu8p02cp3gc20h1D5doILOGcISWDJKEFck2amevbI2enBFAJP7UMZagLvugxs5tjvD5sqhCON+8NH3lC X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:BL1PR12MB5320.namprd12.prod.outlook.com;PTR:;CAT:NONE;SFS:(13230040)(376014)(23010399003)(7416014)(366016)(1800799024)(6133799003)(4143699003)(11063799006)(56012099006)(5023799004)(10067099003)(18002099003)(22082099003)(3023799007);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?utf-8?B?d3J6NkU0ejU4VUFyUng3NDE1S3RmZUpRRnlLRWZpKzR0dFRIYTNJZE8zVzB1?= =?utf-8?B?dkdvNktQS0RMbFhkMVVDbXcrcm1tdU9nVWtSVkg2L05VVFdtOEtRV2pzNFJG?= =?utf-8?B?Tk1RV3BpWGx6RWhrM1JFdXFacTBMVXdQQ3Q0bEFhbVkxRlRZV0RNRHU5RUdE?= =?utf-8?B?VmRWZmpxTjliKzZVaGxaT0JwU1hvbW5BUDdhWkdKR244cGRpcUlWRDRZK2Mv?= =?utf-8?B?blV0ZmVyWHVrTk9KcHJTVUVNRzY0YjdMdDJNRGF4TW84Y0wzQUFGTm5xSFdh?= =?utf-8?B?T0c2MWdITFNpVmpiL2ErUTd5T1krNk9DWmVhZmNlYWd2WnR0UUtmSEVSWHdy?= =?utf-8?B?ZHRkYXhRa0FxbUVaL0IwR3dJVy9kYnZ3Vk9MQldYa2lnTlAwak9IempoTVBs?= =?utf-8?B?RW84TXZubGcyWFQwRWhMQitUQTBNVEw0OGxkeVNxVk1vUkRkTjROS2RnS1Yw?= =?utf-8?B?TVZsbEFITE9tdDBlWnY3NGYwUk5samhLV1ZOeklJZmpVQzVWMXlVZ2Z0NHpL?= =?utf-8?B?ajl3UkJsa0NRU3JURW13OGdUaHFlS2N0bTlGWlhFQkp6b0FWdkVzbjhLU2NI?= =?utf-8?B?V3ZCRm83dlBWSG9leGNxNTVyZGI2ZkM2OXBCWm0zOUVFaDVBSDdROGtDYUNV?= =?utf-8?B?OU5uRG9hWGN0YkQ5UzgxaDNqZCtaaHhBSjJyRXVZUElEUkJQbnEvTTB4OXFX?= =?utf-8?B?Z3RzV25NamM4UWNsNGxiUUlSSnRuY1N3azRvOEJCVWh1TTFDYlUzTzl5M1pz?= =?utf-8?B?MlB3MXlxcktpbFp4VGdTUEFDMUh1cVpjV2tiUHM4eXQyNTJDT2Zkc3g2VnNr?= =?utf-8?B?dUVVdGFJR241aHZucmtMSUxUM2UxWTkzZHMzb01uanpadkpmOHRKU1g3WnRx?= =?utf-8?B?YVZOcjJtd1pZcTZYQzhiM1pWa0lpT1JDNTJ3aWEzdm5yazBmVEtJa0ZpcmlW?= =?utf-8?B?T2RuUGRTeGdzcU13NFZDbGVmRW1PWDZZSTlOblFHOUhnSGI2amZDUGQyQkZZ?= =?utf-8?B?aTZjc0ViVTk5MHVrRE1BRVFEbXJmQUF4ZTl2VUVxVzdEcGJlWjJ6RVUyaFlK?= =?utf-8?B?VlB2aGZKSkdYdlRMeE0vY3NqbGV1QUNjOTZtdFJ5aWhjTktYVTdPNE9ibThy?= =?utf-8?B?bWFwbyt5Mkl4Wmh5K1FBdGFJYTgzTHdJUWRtdDJFN2doRmcvc09MeG9zMDMv?= =?utf-8?B?Qm44ZytyT04reFVFT1VINk5hVTVyQWtjcDUySWliaVpVVENTUzI0Rk0vUUZy?= =?utf-8?B?cHQ3cmc4NXo5RFJlcE9iVUpCY1pNUkgzYm5WdmFaY0xpNlZEWHpjTklRNjVM?= =?utf-8?B?cXFzSkxmM2NnanZIRkM5Y3JFTVRSOWhJVWYxQlRQNGNtR1VRWHBmYUhYM0FX?= =?utf-8?B?ZHg1UzA2Q2FaZWxGWVRuWllvaDVCRncvSXVzUnRDVEFEQW9GMkd3OXRpU3l1?= =?utf-8?B?WlVaZ1Zad1REVG1IVXkyS3hUZmRlNjFldklUSmpudjJySWtLUU5tMjJxczY1?= =?utf-8?B?WEFSQTdlbzFOeStpR0VaYWw3QkRWUVU3VWpVV0tVcUs1Y29MYm9Pb3VPQWNZ?= =?utf-8?B?VXVQN0JxNVg5Q21HRmNtVi9BZlNRc2dDL2hKR0pBdUQyNHdZM0F5cUVja1V3?= =?utf-8?B?V2VjNlhPcFZucW5aZ2ZTWGZuN3JxUnJVZStZdHB1T0JxaTRrYU94bWY5SmZl?= =?utf-8?B?ZzZvU3lLdmFiQXZUMUJSYlJENGJVTnN0S0htWUZoZjd4ckZ0OXY5eFJBbTNQ?= =?utf-8?B?S2lYMkJBU0pGQW5vU2haVkNrS1ptZkptWnhtb2tDd3FwREVZYlNXODZlWElW?= =?utf-8?B?d0JRekovV2Zrc2g1MlBvNGplQm1xaGJBS0hSVUxrTmhCU3IxNGg1Q1BqeVZv?= =?utf-8?B?VmFSY3VVaWhmTU9ReU5sbUp4bHFmZjBxbTJ3YlhKemkxbmF4M21OMEtXemhj?= =?utf-8?B?SHE4RzcrOFRWZ1JHZGtkM2RBREtWMGtnRjdGc002aVEyRm4wd0ZWWUhWNVRC?= =?utf-8?B?dnN0ZDl3NGxxbjlScW5lUC9wamFLNzdkWldvUFgrUGtRRWw3TXF3MG5tSDNn?= =?utf-8?B?Sks4aVVaQWtmenNuWjlRWGdZN0tWTGJzSG03WEFRU0hnNFpMT284VEVmcE52?= =?utf-8?B?RTVxRk1pMytnZ25uRm1ZK2xpS0hWakpRZmRBNTdQcXBOWFFPT3EzaWxuL3Zs?= =?utf-8?B?V0ljeUtyOUJEbEhocXZwZmNoRXYzYU14Z2JxcG0rR0g1cjZlTXA5ZXZBbjVj?= =?utf-8?B?d2ZmMWoxWFVHZ293M2RwVXBsU2tia3g4Y0RRSXlta0ZzOThmcmpFakh0Q3Zj?= =?utf-8?Q?Fq/K9Dzm4vPhEf9Ks3?= X-OriginatorOrg: amd.com X-MS-Exchange-CrossTenant-Network-Message-Id: 4493ef64-39cc-4a68-d7f2-08dee7630d24 X-MS-Exchange-CrossTenant-AuthSource: BL1PR12MB5320.namprd12.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 21 Jul 2026 20:02:54.2761 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 3dd8961f-e488-4e60-8e11-a82d994e183d X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: sDKNne2j/MfWTn0zV3TTzN7p9PeKH93oHsqL30fqbKShZqZsZwmxA1YqFriIaQpq X-MS-Exchange-Transport-CrossTenantHeadersStamped: DSVPR12MB999191 Hi Ben/Reinette, On 7/21/26 12:30, Reinette Chatre wrote: > Hi Ben, > > On 7/21/26 6:23 AM, Ben Horgan wrote: >> On 7/20/26 23:54, Reinette Chatre wrote: >>> On 7/20/26 6:30 AM, Ben Horgan wrote: > ...> >> The former, info/ contains a directory for each allocation scope of each resource. >> >> >> >> info >> ├── L2 >> │   ├── resource_schemata >> │   │   ├── L2 >> │   │   ├── L2_CMAX >> │   │   └── L2_CMIN >> │   └── scope : L2 >> ├── L3 >> │   ├── resource_schemata >> │   │   ├── L3 >> │   │   ├── L3_CMAX >> │   │   └── L3_CMIN >> │   └── scope : L3 >> ├── MB >> │   ├── resource_schemata >> │   │   ├── MB >> │   │   │   └── MB_MAX >> │   │   ├── MB_MIN >> │   │   ├── MB_PBM >> │   │   └── MB_PROP >> │   └── scope : L3 >> └── MB_NODE >> ├── resource_schemata >> │   ├── MB_NODE_MAX >> │   ├── MB_NODE_MIN >> │   ├── MB_NODE_PBM >> │   └── MB_NODE_PROP >> └── scope : NUMA NODE >> > > At first glance this looks good to me. As you highlight below the nuances of how "Global S/MBA" can fit > in here still needs to be worked out. One thing that resctrl may need to highlight when documenting this > new capability is that while historically a resource had a matching schemata entry, this is no longer the > case. As you show above some systems may have MB_NODE resource but no MB_NODE control while I expect that > "Global MBA" (if it adopts this) may indeed have a MB_NODE resource with a MB_NODE control. > > >> Please just consider this a mistake. MB_MAX should be a child of MB. >> >> info/ >> ├── MB >> │   ├── resource_schemata >> │   │   └── MB >> │   │   └── MB_MAX >> │   └── scope >> └── MB_NODE >> ├── resource_schemata >> │   └── MB_NODE_MAX (Changed from MB_NODE) >> └── scope > > ack. > >> >>> >>> When thinking about MPAM, what would the underlying hardware control of "MB_NODE" be? It looks >>> from above that it would either start out by itself having the properties of the underlying >>> "MAX" control or is the plan to have it be a percentage based control backed by the >>> underlying "MAX" hardware control? >> >> The underlying hardware of MB_NODE would be essentially the same hardware as that backing MB_MAX, >> but at a different location in the SoC, at the memory controller rather than in the L3. >> >> To correct myself slightly, I don't think we should have a control called MB_NODE, rather, it should >> be MB_NODE_MAX. > > ack. > >> >> My understanding of previous discussions is that __ it the >> pattern for control names and the pattern for resource names being _. >> (allowing for or being missing to match existing naming.) > > This is where discussions are from my view also. > >> >> I don't think we should introduce more percent based controls and for new controls we can introduce >> a new format to describe them. Perhaps just the positive integer with a resolution supplied in info/ > > No, we should not introduce more percentage based controls per se. With the new schema format a > percentage based control is just a variant of a proportional scalar control. > >> as discussed previously. Although, I have been pondering on whether we can do a bit better. >> >> We could use hexadecimal point based format for controls which are a proportion of a resource and >> have a resolution which is a power of 2. The advantage of this is that the meaning of the value is >> independent of the granularity of the control (number of parts). >> >> 0 is represented as 0x0 >> 1 as 0x1 >> 1/2 as 0x0.8 >> 7/256 as 0x0.07 >> 1/2**28 0x0.00000001 >> etc >> >> This maps well to the MPAM fixed-point fraction point format without having the weirdness of having >> values forced to 1 or 0 not being really 0. These MPAM h/w oddities can be hidden just by using the >> mbw_min mbw_max of a control. In MPAM this could be used in CMIN, CMAX, MB_MAX, MB_MIN and I would >> hope this would be useful for other architectures too. I am preparing some RFC patches on top of >> your PoC for consideration of this idea and to explore some of the proposals discussed relating to >> generic schemata and how they land in practice from the MPAM side. > > I'm going to stand with Dave Martin [1] on this point with a preference to avoid floating point in > resctrl input and output. > Could you please elaborate where values are forced to 1 or 0? The new schema format was intentionally > created to *avoid* rounding errors (and parsing complexity). > > [1] https://lore.kernel.org/lkml/aNFliMZTTUiXyZzd@e133380.arm.com/ > >>> I also understand MPAM to support more memory bandwidth controls ("MIN", "HARDMAX"/"HARDLIM", etc.). >>> Do you envision them to exist within info/MB/resource_schemata/ as well as within >>> info/MB_NODE/resource_schemata/? >> >> Yes, at least for MIN, see the info/ tree above. For HARDLIM, perhaps, but HARDLIM has the added >> complications that it is a property of the MBW_MAX control and that it may be configurable for each >> PARTID or a fixed property of the h/w. When HARDLIM is configurable the control name could be of the >> form ___ where is HARDLIM and the full >> name for the HARDLIM configuration on the MB_NODE resource is MB_NODE_MAX_HARDLIM. There can also be >> an info//resource_schemata//lim file which has values, soft, hard, configurable. > > ack. HARDLIM sounds like it would be a new control type. Perhaps a "boolean" type for which new control > files need to be decided on? Sounds like you are headed in this direction and already have one control file > in mind for this new type. > > With this in mind the control could be built on top of what is being developed at the moment, possibly > be presented to user space following Dave Martin's suggestion in > https://lore.kernel.org/lkml/aO0Oazuxt54hQFbx@e133380.arm.com/: > > | MB_HARDMAX: 0=0, 1=1, 2=1, 3=0 [...] > > or > > | MB_HARDMAX: 0=off, 1=on, 2=on, 3=off [...] > >> >> For CMAX, maximum cache capacity, there is an equivalent control SOFTLIM, which behaves as HARDLIM >> except the meaning of the bit is reversed. We can just use a consistent name in s/w though. >> >>> >>>> >>>> On an x86 system: >>>> >>>> info >>>> ├── MB >>>> │   ├── resource_schemata >>>> │   │   ├── MB >>>> │   │   └── MB_MAX (Finer grained MB, more below) >>>> │   └── scope >>>> └── MB_REGION >>>> ├── resource_schemata >>>> │   └── MB_REGION >>>> └── scope >>>> >>>> >>>> Do you think this helps? >>> >>> "REGION" is not a new scope but instead region-aware MBA is controlled and manages bandwidth at L3 scope. >>> Combine that with up to (currently) four regions each with three controls I find an interface like above >>> potentially confusing to document in an intuitive way. Unless you are perhaps saying that we should introduce >>> a new separate "MB_REGION" L3 scope resource (so let resctrl support multiple "MB" resources at the same >>> scope?) and then *it* contains the twelve new controls within its resource_schemata directory? >>> >>> Since the region-aware controls are orthogonal to the MSR based legacy control resctrl would still need a >>> way for user space to switch from one to the other which implies a dependency between "MB" and "MB_REGION" >>> that is not presented in above hierarchy. >> >> OK. The /sys/fs/resctrl/info/MB/schemata/mode and the MB_REGION controls a child of MB you described >> previously seem s better fit than what I suggested. > > ok, I'll keep following that. > >> >>> >>> I think I am missing quite a bit here as I try to navigate an interface so different from what we have >>> discussed so far. >>> I would like to explore with more detail how this interface can handle the different scenarios we have >>> discussed so far. >> >> Certainly, I don't think we have got to the bottom of this yet. >> >>> >>> The other x86 feature to consider is AMD's upcoming "Global" MBA/SMBA that exposes memory bandwidth allocation >>> in "groups of L3" that I understand could usually be mapped to NODE scope (but it remains controlled at L3 scope), >>> except for one configuration where it is "SYSTEM"(?) scope. >>> Ref.: https://lore.kernel.org/lkml/8f77f498b1c77fa8fd8f5d5687f03ae598068544.1776980182.git.babu.moger@amd.com/ >> >> Hmmm, I'm not sure that the scope can be considered to be NODE scope for GMBA. To me it seems to be >> accidental that it maps to the NUMA node but really the scope is just a grouping of L3 instances. >> For a control to NUMA scope I would expect the resctrl domains to go offline and online in sync with >> the NUMA nodes. For GMBA it looks like it would just going offline/online based on whether any of >> the CPUs and so L3 instances in the group are online. Am I correct here? >> >> Assuming the domains are on L3 groups rather than NUMA also changes which end of the link the >> traffic is regulated and so how cross-NUMA traffic behaves differently. If the domain is an L3 group >> then a task running on a CPU affine to that L3 group won't be throttled unless that particular >> domain is throttled but with NUMA node domains it may be throttled if it has traffic going to that >> domain. > > I'll defer to Babu for accurate answers about this hardware capability. > To me, Global MBA should be considered a NODE-scoped resource. In some configurations it may appear as SYSTEM-scoped, but that is effectively equivalent to a single-node encompassing the entire system. In such cases, there is only one schemata entry controlling the whole system. Yes, multiple L3 instances are grouped together to form a NODE. Internally, programming is still performed at the L3 level, but that implementation detail can be hidden from users and does not need to be exposed through the interface. Thanks Babu >>> >>>> >>>> This also brings another question. On MPAM systems the 'MB_MAX' is backed by the same MSC h/w as MB >>>> but it exposed a different interface to the user. If I understand correctly intel have an option to >>>> have finer grained control of MB (delay) as well and so it would make sense to use a common name >>>> rather than just going for the MPAM centric name of MB_MAX. >>> >>> Apologies but I was not able to parse above. >> >> Ok, let me try to explain again (although it's probably not what we want to do). My intent here was >> to try and explore whether we can reuse naming and controls across architectures in the same way we >> already have for the L2/L3 cache portion bitmap and the existing MB control. >> >> To quote from a previous mail of yours: >> https://lore.kernel.org/lkml/a84af037-6439-4362-be07-d45143e06309@intel.com/ >> """ >> For example, on an MPAM system (if I understand correctly) the user may see: >> info/ >> └── MB/ >> └── resource_schemata/ >> ├── MB/ >> │ └── MB_MAX/ >> └── MB_MIN/ >> >> Compared with a possible implementation on Intel that looks like: >> info/ >> └── MB/ >> └── resource_schemata/ >> ├── MB/ >> │ └── MB_OPT/ >> ├── MB_MAX/ >> └── MB_MIN/ >> """ >> >> In the two setups MPAM MB_MAX and intel MB_OPT play the same role, a finer grained control of the >> legacy MB control. I was thinking these could share a name (MB-PRECISE), but it probably doesn't >> make sense as MB_OPT and MB_MAX have different relationships to MB_MIN. >> > Apologies, I neglected to follow up on this after I clarified internally which RDT control should > actually be used to emulate the percentage based control. My original thoughts that you highlight > above is not correct the RDT also plans to use the "MAX" underlying hardware control for the > percentage based MB control. With this RDT and MPAM should look more similar with RDT simplified > (by dropping the region-aware terms) as: > > info/ > └── MB/ > └── resource_schemata/ > ├── MB/ > │ └── MB_MAX/ > ├── MB_OPT/ > └── MB_MIN/ > > > Reinette >