From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from DM5PR21CU001.outbound.protection.outlook.com (mail-centralusazon11011020.outbound.protection.outlook.com [52.101.62.20]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 44D441D555 for ; Thu, 23 Jul 2026 00:17:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.62.20 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784765849; cv=fail; b=sbH8nV2UqlxuCFZmRYOQuNbDH/HJZsNa2E48tim+doFa1zEIM2mY18umNz6QNeUjSwwL8noyVGxXhy57nos+l9yH+S+KWzwy/LH/j1eqlNUXWhnA1W9kCGGT+YsjZWQLr8yBUzfH9YW5kttnvsvso1EiZXFpfygf0c2u3pXzJtE= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784765849; c=relaxed/simple; bh=KWseaRRuN7LXgGL98CjfBSZu7SIh4ZLFZHYOU/8A/Qw=; h=Message-ID:Date:Subject:To:Cc:References:From:In-Reply-To: Content-Type:MIME-Version; b=Ai0xI5HciwsuK9GP+NEllK2++qjBQJrgbMLjeTOqI6QTADpDZNimUAdQd8AZa9XBpmhwElFDJyHhmFMixkdMqg7JMDI2KzEZWiUP0ah7BvbGrE/iwG3zMhJjEODWVFJ55UyUkFG+VQ3tjRk1AMIkn6m8msus4eAU3Qpb8HF0FU0= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=iJgFzffQ; arc=fail smtp.client-ip=52.101.62.20 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="iJgFzffQ" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=rEI2xwVcC5p/w0GoOhekfwdOQCe7XJaOWBt7OkUY3SaqSxjfOnrH4DW/uIGqpEBbvOfpkB9GQaKh48gpbOqSnl2PkwljgApcNGLwdr5bhUAZU4J8BJNTB31XbzAQeI1I4Id8j0ciyLz7EBo3Ixm204ciyg0OX8F0vdhyoUlQFD3jXKlCp6sEkvrHjCa2srr6WpiMedIal9iBa4ZlkpefE7Do3n8OodZSPIdGZ4iT/bBQzwDp9gVwbtypJn5ZR7v2yVx0nDALZqQmY9pSy0heFhOIPdg5xhZTYs0T7Tco1LJs3OfrChcPjHf69dNIVoQUbUGPIuD0XohKeNrF5LzTrw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=MlhZsUcKkmqtmeLpSYNcU1GHQB2ZsCi/OAkwVhaggqU=; b=RsRrMHCfaPgWE+jXbnc2BnF4ZWKyj5kjGP4ssaIlyH2DcSha9Lz8qnEXAoj5Noz4C3Pr+PTMrkU9tc9I0WvCXJT9lypLzqJCC9NSjd64ZIcUnNgM1rMjIB/5m818KS4hLWHoBrUT7oTlwkdsitnF55NVNwI57G3wN1s+Ip5J77qT4taRMTIvXNixQUa/KK1ZNapevm6xsH0AHNJjvvHfMS0ciJ+cghF88DTLVCmZc1xLFuiynkNsZylVV5zwvCOIPJlrPAbn8cZaJ9tAmjZLVE+9wrQf12Sb+3yqF2ROQqUkaZ8gELz8+9oP1696kcgyJJJ2zIIxN4ZO6zlhgu+kWg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=nvidia.com; dmarc=pass action=none header.from=nvidia.com; dkim=pass header.d=nvidia.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=MlhZsUcKkmqtmeLpSYNcU1GHQB2ZsCi/OAkwVhaggqU=; b=iJgFzffQqJmcD8uqEgnOD3lDmR0fQb9JL0ufvq0fMplMqa0sqVjDWiozSkgnRkMP2cTiBi436WFOHcYzzOPrqd1y5uRjMMxVpoSaKgSXUzwlUUaMHHgjEjSk+RQ+N6A8deRfr51Vs8tyxOD4NoJ5tS1uOHmW+j0ohuLX2lwQwW8wN+808pk9SSTLloGFUhmuWkQ/yaynshCEEkIAmv+KxDNZ2xfK5NFH/312vl241440JgjNAK28AV5cbhwZPzCCRfqULVCZGyFraZDwIhfMV8NDAJ1Ek6UA9SDyIsIIMnWhedTR5vbtkUcJxO39a/wB9ZjOwJ3oxjZiGeOSy7JBNg== Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=nvidia.com; Received: from DM4PR12MB5230.namprd12.prod.outlook.com (2603:10b6:5:399::11) by SJ0PR12MB6781.namprd12.prod.outlook.com (2603:10b6:a03:44b::22) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.245.10; Thu, 23 Jul 2026 00:17:22 +0000 Received: from DM4PR12MB5230.namprd12.prod.outlook.com ([fe80::6e87:1bde:1853:3b73]) by DM4PR12MB5230.namprd12.prod.outlook.com ([fe80::6e87:1bde:1853:3b73%5]) with mapi id 15.21.0245.009; Thu, 23 Jul 2026 00:17:22 +0000 Message-ID: <9425e9fc-36cf-44cd-b6fd-88b76d106eea@nvidia.com> Date: Wed, 22 Jul 2026 17:17:20 -0700 User-Agent: Mozilla Thunderbird Subject: Re: [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept To: Reinette Chatre , Tony Luck , Ben Horgan , James Morse , Dave Martin , Babu Moger , Drew Fustini , Chen Yu Cc: Borislav Petkov , Thomas Gleixner , Dave Hansen , Peter Newman , "x86@kernel.org" , "linux-kernel@vger.kernel.org" References: <5ee87762-1898-4b62-94da-85b3e9917ecc@intel.com> <62701203-c4a3-4ec2-a9af-602e1fc15863@nvidia.com> <8f9f78dd-e3f5-4b35-bc72-0eb5dafdcedf@nvidia.com> <36163a81-9737-49e3-93ef-6c392f7272f0@intel.com> Content-Language: en-US From: Fenghua Yu In-Reply-To: <36163a81-9737-49e3-93ef-6c392f7272f0@intel.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit X-ClientProxiedBy: SJ0PR05CA0065.namprd05.prod.outlook.com (2603:10b6:a03:332::10) To DM4PR12MB5230.namprd12.prod.outlook.com (2603:10b6:5:399::11) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DM4PR12MB5230:EE_|SJ0PR12MB6781:EE_ X-MS-Office365-Filtering-Correlation-Id: 763a3887-c719-430b-9327-08dee84fc3d7 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|366016|1800799024|23010399003|22082099003|18002099003|56012099006|5023799004|11063799006|3023799007|4143699003|10067099003|6133799003; X-Microsoft-Antispam-Message-Info: l/wYZN+95xQAHj3gJB7oZOIytVndfRhaK+2pKa8vEZtwN+t9SrP/lXjhO24HrX7GFLrb9azmAIehSjYtIb2aOt1Wa0q1EY0u+ZZKQFB3Pcu0A9aEslfm7CuCrAKQS52gwcMKwl4FLCkMCuOdETU7jY6hZICVdVUe0jzWN1HzqGDN5tOPQoGNdrvnn7DXP05E1U6nkAtbyIY//Q0zazjD/xlf6mXjJn+qy2e9IEQkF174IM/MRWqAPTM5BTwcmbeOwWFvKF7X49kfbnHMW48HjvdRca2MMTLybalBSXpVnWWimaP58ORFVdy1FXK68hQN/Rq82BtKJ5MrjLA4InI68nyniMxhoIt7b9u+p05pkLV3o1r4M7nBPTImROpY+i9j6YiICW/hpIQBUs8GQWMuSRqHv3Xak0MLhoBTKYb6P18KUrn2Uo3FX8NuDAaWz+FJ+U5BQMnaPh15tSXYbODt+PXoJ541XISvMnKWT6AqmikcMKDRioTn9g2d4gVDm1KNZwP1DbaIzCjA4Z+rA2mwkQ1/L2U3/DbM1XGkhufBiekagVZPEuOrhvznxc9Uwi70IN/c12x48bM2dVNaH5myJT6cD12Lc5HsMcvbkQOPvKf/kJlWLSQbfU+2qcXdiX5Nl8Q7HbwNBTamuL7nS527dUHVrboL7W7Ydqi7QWPaRs0= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:DM4PR12MB5230.namprd12.prod.outlook.com;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(366016)(1800799024)(23010399003)(22082099003)(18002099003)(56012099006)(5023799004)(11063799006)(3023799007)(4143699003)(10067099003)(6133799003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?utf-8?B?YThmSWQ2VmpUUDNjYzVXTENKUUQyOVZuem9PKzJwcnJsQkJ4bmJuL3l0bHVY?= =?utf-8?B?WnlTMkJlUU9jN0tqcUIyMks5OFc5aHZnYUREdHJuczExaHYvS3pjTzRURkFO?= =?utf-8?B?TzJzNGM1Qk83U1d0R0hLYzVXbFdXdHViRDg4cWFzcVV3MTNOR1BlMEd0d0J5?= =?utf-8?B?OW9qYUJIVVZTT0ZKclB3RFhPTGxRd2E5d3lOVFNKREdtOWNtNkExWHlaUW42?= =?utf-8?B?TkR1bkFZTWhxMThUcVZFNG42eDZBVnFVcmhsaENsWElqWW5EM05ZRTVJSEVt?= =?utf-8?B?Y0RYdEpZZFpMKzBLczdrZ2dFNFFXVFdvMk9QQTlDc3Y0SDFQNGJPSVRLOWc2?= =?utf-8?B?WE1KZkxmWU5VUjVFd3Z4NUlNeW5xOXYwVC9mZDNSL2VTSWR4N01SbStveEh1?= =?utf-8?B?UmVxL2RadmJWa0NURVBNdjFkbGI4aG5ZQnBGaU82VklZZHVOTmcvV2ZlWFE3?= =?utf-8?B?cCtwS2FjdGxWOTNLck1hRC9MV0I4RGhkVjZaQkU0VDVQRHhuNjVMOHljUCtB?= =?utf-8?B?c0ZxdGN2RGRja3ZrcXZOdnlvUjRwTWQzeEs2VVNpWnhxamNXcmpqNkdSMDFh?= =?utf-8?B?SXFzT3hLNzRKVGtHQmRiL3haYlptQTZYR1JicERQUjNXeVBrOFFhRVRldDNC?= =?utf-8?B?MldyejgwLzFPUGh5SktRRlovc2FTSlcrYlJqbG1TRStyUVNUU0dySGEwT1BF?= =?utf-8?B?WnRSZlh0NWNNU3VxVlk2dW5KcHRtU0FzaTBVMXRtcUMxM1lGNDhMbkh0eERQ?= =?utf-8?B?b1gySUpSazEwbGsycVN5dFA4dU9pTWdUbXNJL0llT1JXTXpiTnM1STR0YUs1?= =?utf-8?B?dmIvdytPY0ZVVTdEMnZ0YnhPSU5VZ2s1aXhGQy9KeDdZRkJqQXRhbDA5YUNs?= =?utf-8?B?VU5MMlNlUm1VT08rcW8rMWRQNkg1T25seW02cW04T0JneVZyV2NMckxPSWJB?= =?utf-8?B?V0gvYk9xSzBMVTJBaDN5aXlOV2E3Q29EY0M4K04yOG82cWN0TUMwTHFzTHlQ?= =?utf-8?B?QWQvd1N2b09NcUVaWHAvQWt2aXlSRUZObnVjZE5FS1VSVFVkT2hTY3ZHeDlW?= =?utf-8?B?RDQ0bEVuZ1BZUHl0c3BmWUlaSE9vMi9EVFZEVnU3ZFg1bEllK2ZiS2ViTjlT?= =?utf-8?B?L3dCR1hRN3h0dGwwZkZYZFRzSko4WEN1VmFUUkZRRnVQTDRZbkNVOEhkeWo5?= =?utf-8?B?NjQycm95WGY4YW5MWUU0LzhpM0pHNlhkTlZmYWVOQnBraEVrc0dBYWJsZGNx?= =?utf-8?B?cmRwYWNnek9ldmNndXBCbE9wZTJwK2tmYlJWclZJaldDN2syNDRYdjBuM0Vu?= =?utf-8?B?TS9ocW5LU0VqeXNaQk5YSmdFaFJuaEVOQXlwK1FTd2JRaENpOTg3QkFVZHpj?= =?utf-8?B?QWh2WHlMampiQ3k5dkNncUZmYTBlWUVYNms3K2t0WEVGejZQZUVHQ3dQZGtN?= =?utf-8?B?cnZPN3hIR3hQY1BmbkR6N2JVZVFVQzFzWUxsY0xjRHhyaW5jSUc4TVNuaVFk?= =?utf-8?B?N0NXSllBUGNBMWtDSGZIUWdZYWo2QXZESUhJRll5NllDVVFDMjR2YThZRGh0?= =?utf-8?B?MTJNdGZTekxjSlZwQStvWlNDcWl4cmplVmVxZEFWcXV5NzBEbWs1N1NSVXRX?= =?utf-8?B?NjZ2eXg1dFgzOUNXcHFtSGw2MzhrWlBibm9FbzBaaGxMekgyU3pFQmRCdTVn?= =?utf-8?B?b0N1ZU9QYUwrQmhWL2JSVThjZDFJT3ZFdmh4a1B4SFlwYjJBTWd4QTRTcUJU?= =?utf-8?B?bEo5a3kzUk82TXFIQi8vYlFBZ2hjSEpaNzJZL2hyVVVVL25rblJTMmdNVkI2?= =?utf-8?B?USsrbitNTklrYjRQaHZlMko0eUVmN010MlNHeXltNVpxUGtiQnFwMG10c3hs?= =?utf-8?B?Q0FGNjhqbEtMMTM5WnJid0s0Mmh6alBGQVBMeG1KcUhVM0VqcGpMOVF4cVBQ?= =?utf-8?B?N0NRUlJHdWQ4cEVYVlZrZzlxSkp4dmxQOVVFRUZYb0tPVmt0ZmluVVUzeit5?= =?utf-8?B?cGpsL3hVYTN3cmhvZ2xuUjV3a3lUK1F4T3pHZ1NwdldwWXpSOTJHZUZIZ1lz?= =?utf-8?B?clhseG5LbnhSeklneHo3eFpEUldTRXh1ZWVwanhjRUp1NExvUDQ5YU5SV0JR?= =?utf-8?B?WG1CN2o0ZjFuYnYyYVdhWUFIcDNVVFB5QkRrZjZOU1JtMStqSnJpcGpDdkhI?= =?utf-8?B?aDRPclR1SXA2ZGNqdTNhS3hjMkVDYWVMZFVnSFRiQTRoNXBNdlR0bFdQSXNZ?= =?utf-8?B?d3B3Sk4xU2hjOUFlVWRZbWVTWVk5YVBzUjZlZjFvYXEyN2lqdVdmSmhIS1VO?= =?utf-8?B?WVl5YytZTGpBN0RGYmFVUXBSWmF3TVVsWmtnLzhEenpSL2gxLzRUUT09?= X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-Network-Message-Id: 763a3887-c719-430b-9327-08dee84fc3d7 X-MS-Exchange-CrossTenant-AuthSource: DM4PR12MB5230.namprd12.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 23 Jul 2026 00:17:22.1088 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: sJUCE57uvg+2NqfAHPMrDxyru0sqtecDts4jRkzrPPCoKqFQLH07xsEP+256wb9vg1FoRcquOuUxEsNCKz+FaA== X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ0PR12MB6781 Hi, Reinette, On 7/14/26 15:06, Reinette Chatre wrote: > Hi Fenghua, > > On 7/10/26 1:59 PM, Fenghua Yu wrote: >> On 6/25/26 08:43, Reinette Chatre wrote: >>> On 6/24/26 6:26 PM, Fenghua Yu wrote: >>>> On 6/24/26 15:22, Reinette Chatre wrote: >>>>> On 6/24/26 12:08 PM, Fenghua Yu wrote: >>>>>> On 5/29/26 11:06, Reinette Chatre wrote: >>>>>> >>>>>> As Shaopen and Ben mentioned earlier, we are working on two MPAM >>>>>> features that may need to change schemata interface. The CPU-less >>>>>> feature was discussed on LPC (although the interfaces will be >>>>>> slightly different from the LPC). >>>>> >>>>> I know. Here is where I tried to engage with you on needed interfaces after LPC: >>>>> https://lore.kernel.org/lkml/fb1e2686-237b-4536-acd6-15159abafcba@intel.com/ >>>> >>>> MPAM ACPI defines MSC (Memory System Control) is defined in one of two ways (not both) on one platform: >>>> 1. L3 and memory together on each processor MSC >>>> 2. L3 in processor MSC and memory control/monitoring in different memory MSCs. >> >> Ben said there is type 3 platform: >> 3. L3 cache and memory bandwidth in processor MSCs and memory bandwidth in different memory MSCs. >> >>>> >>>> On type 1 platform, schemata is legacy: >>>> MB:1=100;2=100  <-- cache id 1 and 2 as domain id >>>> >>>> On type 2 platform, I will not reuse "MB:" name. Instead, define new resource name "MBN:" for numa node and schemata is: >>>> MBN:0=100;1=100;2=100;10=100;18=100;26=100 <-- numa id 0, 1, 2, 10, 18, >>>>                             26 as domain id >>>> On type 2 platform, there won't be "MB:" line. Numa 0 and 1 >>>> are for mbm allocation on socket 0 and 1. 2,10, 18 and 26 are for GPU >>>> memory nodes allocation. >> >> On type 3 platform, there could be "MB:" line for L3 cache and "MB_NODE:" for numa node. Example schemata is: >> >>      MB:1=100;2=100                 <-- cache id 1 and 2 as domain id >> MB_NODE:0=100;1=100;2=100;10=100;18=100;26=100 <-- numa id 0, 1, 2, 10, >>                                                    18, 26 as domain id >>> >>> (to help make things explicit I will refer to what you call "MBN" as "MB_NODE" to make it >>> explicit that it is memory bandwidth allocation at node scope) >>> >>> I am trying to consider how this can be accomplished while also considering all the other >>> new hardware features that resctrl need to support. Consider, for example, AMD's "Global >>> MBA" (https://lore.kernel.org/lkml/cover.1776980182.git.babu.moger@amd.com/) that throttles >>> memory bandwidth at L3 scope but the user configures allocations at NODE scope. At this time >>> the plan is to support this with a second control associated with the MB resource that can >>> allocate memory bandwidth at node scope. See >>> https://lore.kernel.org/lkml/430ffb48-29f4-44d9-9164-9f8b743b2739@amd.com/ >>> >>> If resctrl creates a new resource for node scoped memory bandwidth allocations to support these >>> "type 2" systems then that will result in an inconsistent interface between architectures that >>> we should avoid. >>> >>> Have you been listening in on the discussions surrounding emulated controls? Considering that, >>> would it be possible to support the "MB" control on a type "2" system but have it be backed by >>> (emulated by) the underlying "MB_NODE" control? >>> >>> resctrl could expose both controls on these "type 2" systems but make it clear that "MB" >>> is emulated by "MB_NODE". For example: >>> >>> info/ >>> └── MB/ >>>      └── resource_schemata/ >>>          └── MB/ >>>              └── MB_NODE/ >>> >>> User will see both controls in schemata file but when changes are made to "MB" control it >>> will show in the "MB_NODE" control and vice-versa. User could also disable the "MB" control >>> that will establish familiarity with the interface at which point resctrl can drop the >>> "MB" control from the schemata file on these "type 2" systems. >>> >>> Having the MB resource available with an MB control will keep resctrl backward compatible >>> if there are any tools that expect that. If backward compatibility is not of concern then >>> resctrl could initialize with the emulated control disabled by default. See discussion at >>> https://lore.kernel.org/lkml/5e575bc2-e67f-4696-9332-33c54023c057@intel.com/ >>> that describes a new resctrl capability in support of RISC-V and RDT. >>> With this resctrl could initialize with: >>> >>> info/ >>> └── MB/ >>>      └── resource_schemata/ >>>          ├── MB/ >>>          │   ├── MB_NODE/ >>>          │   │   └── status:enabled >>>          │   └── status:disabled >>>          └── mode:legacy [native] >>> >>> With above a "type 2" system will boot with its schemata file just containing the "MB_NODE" >>> control while info/MB describes the memory bandwidth resource. >>> >> >> On type 3 machine, schemata has both MB in legacy mode with cache id as domain id and MB_NODE with numa id as domain id. >> >> Is this directory OK? >> >>  info/ >>  └── MB/ >>       └── resource_schemata/ >>           ├── MB/ >>           │   ├── MB_NODE/ >>           │   │   └── status:disabled >>           │   └── status:enabled >>           ├── MB_NODE/ >>           └── mode:node >> >> 1. MB and MB_NODE are shown in parallel in inf/MB/resource_schemata/ > > I do not think there is a need to expose an emulated MB_NODE control if the actual MB_NODE > hardware control exists. > >> 2. mode is set as "node" meaning "MB" is for L3 and "MB_NODE" is for numa node > > I assume you mean "scope" instead of "mode"? (more below) > >> 3. Emulation "MB_NODE" is disabled (or should the "MB_NODE" sub-dir be invisible?) > > Right, I do not think emulation is needed here. No need to make it invisible since it should not exist. > >>>From what I understand these "type 3" machines could be simplified to: > > info/ > └── MB/ > └── resource_schemata/ > ├── MB/ > │   └── scope:L3 > └── MB_NODE/ > └── scope:NODE > > Beyond this I believe that MPAM currently emulates the MB control with its "MB_MAX" control and users may want > to make bandwidth allocations at the fine granularity that it supports. Taking this into account the interface > may end up looking like: > > info/ > └── MB/ > └── resource_schemata/ > ├── MB/ > │   ├── MB_MAX/ > │   │   └── scope:L3 > │   └── scope:L3 > └── MB_NODE/ > └── scope:NODE > > A system like above will thus have three schemata file entries: > MB > MB_MAX > MB_NODE > > Three schemata file entries would be unnecessary for users familiar with the finer granularity MB_MAX control > so that is where the "mode" file can be used to disable the legacy MB control to just expose MB_MAX and MB_NODE > on these systems. > > Would that work for these systems? > > ... > >>>>>> There is another MPAM feature called MBW Max hardlimit which sets >>>>>> "MB:" allocation as hardlimit (i.e. MBW throttling percentage must >>>>>> be satisfied) per domain. Adding a new "MB_HLIM:" line in schemata. >>>>>> It's 1:1 mapped to "MB:" to control hardlimit of MB throttling >>>>>> percentage on each domain. By default hardlimit is off (0) and can >>>>>> be turned on to set MBW Max hardlimit on a domain. >>>>> >>>>> ack. This sounds like a new control associated with the MB resource. >>>>> This is a boolean control as Dave highlighted in previous discussion so >>>>> resctrl would need to know its properties. >>>>> See https://lore.kernel.org/lkml/aO0Oazuxt54hQFbx@e133380.arm.com/ >>>>> >>>> >>>> Right. ("MB_HLIM" name may be adjusted accordingly when "MB_MAX" is available.) >>>> >>>>>> For exmple: >>>>>> MB_HLIM: 0=0;1=0;2=1;10=0;18=0;26=0 >>>>>> MB:0=100;1=100;2=80;10=100;18=100;26=100 >>>>>> >>>>>> On GPU memory numa node 2: cannot use more than 80% of total max mbw even if there is still idle mem bandwidth on this node). >>>>>> >>>>>> MBW allocations on all other domains are soft limited, meaning MBW can be used more than specified if mem is idle. >>>>>> >>>>> >>>>> ack. >>>>> >>>>>>>            L3:0=fff;1=fff >>>>>>> # echo 'MB_MIN:0=50' > schemata >>>>>>> # cat schemata >>>>>>>            MB_MAX:0=100;1=100 >>>>>>>            MB_MIN:0=50;1=100 >>>>>>>            MB:0=100;1=100 >>>>>>>            L3:0=fff;1=fff >>>>>>> >>>>>>> Writing to the dummy control will call a dummy callback that just prints to the >>>>>>> kernel log: >>>>>>> "resctrl: Updata temporary MIN control on domain 0 with user value 50" >>>>>>> >>>>>>> >>>>>>> Example output of info/MB/: >>>>>>> /sys/fs/resctrl/info/MB/thread_throttle_mode:max >>>>>>> /sys/fs/resctrl/info/MB/num_closids:15 >>>>>>> /sys/fs/resctrl/info/MB/delay_linear:1 >>>>>>> /sys/fs/resctrl/info/MB/min_bandwidth:10 >>>>>> >>>>>> Add two new MB info RO files: >>>>>> 1. /sys/fs/resctrl/info/MB/domain_id >>>>>> It shows "numa" for using numa id in "MB:" or "cache" for using legacy cache id. >>>>> >>>>> This proposal introduces a *global* property to the MB *resource*? It does not seem as though >>>>> this takes into account *anything* about how resctrl can support new hardware that has been >>>>> discussed before, during, or after LPC. You have not participated in these discussions and >>>>> now make an orthogonal proposal that does not take into account *any* of the requirements >>>>> that we have been struggling with for months. >>>>> >>>>> Why should this proposal be taken seriously? In your absence folks have been trying to >>>>> accommodate how these upcoming products and be supported and the "scope" file associated with >>>>> a control is intended to communicate to user space how the domain ID should be interpreted. >>>>> >>>>> Why are you proposing something entirely different here without even acknowledging current >>>>> approach and explaining why it does not work for you? >>>>> >>>> >>>> So can I change this part to adding the following files in info dirctory? >>>> >>>> 1. For numa memory bw allocation (MBN): >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/ >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/resolution:100 >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/tolerance:5 >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/type:scalar >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/min:10 >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/scale:1 >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/scope:NUMA >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/unit:all >>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/max:100 >>> >>> This is not just about adding files to the info directory. The files, directories, their relationships, >>> and content have meaning. All I see from these proposals is an attempt to slap some new files into >>> resctrl without any consideration to present consistent interface to users and without consideration of >>> other architectures that need to be supported by resctrl. >>> >>> resctrl needs to provide a generic and consistent interface to user space irrespective of the >>> underlying architecture. Architectures cannot just slap some new files for their convenience. >>> >>>> >>>>>> 2. /sys/fs/resctrl/info/MB/max_lim >>>>>> It shows number 0-3 for MPAM MBW max limit behaviors: 0 for supporting both softlimit and hardlimit, etc. >>>>> >>>>> Again this adds another *global* property to the MB resource but then above you >>>>> describe the new "MB_HLIM" schemata file entry that implies that it is a new control >>>>> for the MB resource. Having it be a new control for the MB resource matches earlier >>>>> discussions. To support this I thus expect it to be exposed as a new control with >>>>> potentially a new type if any of the existing planned types do not suffice. >>>>> >>>> >>>> How about adding these MB_HLIM dir and files in info? >>>> >>>> /sys/fs/resctrl/info/MB_HLIM/resource_schemata/MB_HLIM/type: boolean >>>> /sys/fs/resctrl/info/MB_HLIM/resource_schemata/MB_HLIM/max_lim: 0 >>> >>> This presents "MB_HLIM" as a *resource* to user space. It is not a resource >>> but a *control* of a resource, no? I thus expect it to instead look something like >>> below that makes it clear that MB_HARDMAX is a control of the MB resource. >>> >>> info >>> └── MB >>>      └── resource_schemata >>>          ├── MB >>>          └── MB_HARDMAX >> >> Yes, this makes sense. I have changed to this hierarchy. > > Thank you very much for considering this approach. [ MB_MAXHLIM: I use this name for MBW_MAX hard limit feature as Dave Martin suggested before. He also suggested MB_HARDMAX. Either name is good for me. I use MB_MAXHLIM to explain MBW_MAX hard limit for now.] Some implementation thoughts: MBW_MAX hard limit itself is not a MB control. Rather, it configures MB control, i.e. turn on MB control's hard limit or turn off its hard limit. So MBW_MAX hard limit doesn't have properties like bandwidth_gran, delay_linear, etc. MBW_MAX hard limit's property is only a boolean type. So I would think it maybe a configuration inside a control. Similar configurations could be hard limit for cache capacity in MPAM. Maybe can add "configs" inside resctrl_ctrl. Schemata and info/MB/resource_schemata/MB will show/write the configurations per control? For this configuration or future configurations, add "configs" list in: struct resctrl_ctrl { struct list_head entry; enum resctrl_scope scope; struct list_head domains; enum resctrl_ctrl_type type; enum resctrl_ctrl_name name; struct resctrl_ctrl *emulated_by; struct list_head configs; <--- Add configs for this control union { struct resctrl_cache cache; struct resctrl_membw membw; }; }; A resctrl control can have one or multiple configurations. Currently MBW_MAX hard limit is the only one. But the infrastrucutre supports multiple configurations per control. schemata: MB:1=100 <-- MBW_MAX on L3 id 1 MB_MAXHLIM:1=0 <-- turn on/off MBW_MAX hardlimit on L3 id 1 L3:1=fff info/ ├── MB │   ├── bandwidth_gran │   ├── delay_linear │   ├── min_bandwidth │   ├── num_closids │   └── resource_schemata │   ├── MB │   │   ├── configs │   │   │   └── MB_MAXHLIM │   │   │   └── type <--- bool │   │   ├── max │   │   ├── min │   │   ├── resolution │   │   ├── scale │   │   ├── scope │   │   ├── status │   │   ├── tolerance │   │   ├── type │   │   └── unit │   └── mode Is this a valid way to handle MB_MAX hard limit (and future more configurations per control)? Thanks. -Fenghua