The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Fenghua Yu <fenghuay@nvidia.com>
To: Reinette Chatre <reinette.chatre@intel.com>,
	Tony Luck <tony.luck@intel.com>, Ben Horgan <ben.horgan@arm.com>,
	James Morse <james.morse@arm.com>,
	Dave Martin <Dave.Martin@arm.com>,
	Babu Moger <babu.moger@amd.com>,
	Drew Fustini <fustini@kernel.org>, Chen Yu <yu.c.chen@intel.com>
Cc: Borislav Petkov <bp@alien8.de>,
	Thomas Gleixner <tglx@linutronix.de>,
	Dave Hansen <dave.hansen@linux.intel.com>,
	Peter Newman <peternewman@google.com>,
	"x86@kernel.org" <x86@kernel.org>,
	"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: Re: [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept
Date: Wed, 22 Jul 2026 17:17:20 -0700	[thread overview]
Message-ID: <9425e9fc-36cf-44cd-b6fd-88b76d106eea@nvidia.com> (raw)
In-Reply-To: <36163a81-9737-49e3-93ef-6c392f7272f0@intel.com>

Hi, Reinette,

On 7/14/26 15:06, Reinette Chatre wrote:
> Hi Fenghua,
> 
> On 7/10/26 1:59 PM, Fenghua Yu wrote:
>> On 6/25/26 08:43, Reinette Chatre wrote:
>>> On 6/24/26 6:26 PM, Fenghua Yu wrote:
>>>> On 6/24/26 15:22, Reinette Chatre wrote:
>>>>> On 6/24/26 12:08 PM, Fenghua Yu wrote:
>>>>>> On 5/29/26 11:06, Reinette Chatre wrote:
>>>>>>
>>>>>> As Shaopen and Ben mentioned earlier, we are working on two MPAM
>>>>>> features that may need to change schemata interface. The CPU-less
>>>>>> feature was discussed on LPC (although the interfaces will be
>>>>>> slightly different from the LPC).
>>>>>
>>>>> I know. Here is where I tried to engage with you on needed interfaces after LPC:
>>>>> https://lore.kernel.org/lkml/fb1e2686-237b-4536-acd6-15159abafcba@intel.com/
>>>>
>>>> MPAM ACPI defines MSC (Memory System Control) is defined in one of two ways (not both) on one platform:
>>>> 1. L3 and memory together on each processor MSC
>>>> 2. L3 in processor MSC and memory control/monitoring in different memory MSCs.
>>
>> Ben said there is type 3 platform:
>> 3. L3 cache and memory bandwidth in processor MSCs and memory bandwidth in different memory MSCs.
>>
>>>>
>>>> On type 1 platform, schemata is legacy:
>>>> MB:1=100;2=100  <-- cache id 1 and 2 as domain id
>>>>
>>>> On type 2 platform, I will not reuse "MB:" name. Instead, define new resource name "MBN:" for numa node and schemata is:
>>>> MBN:0=100;1=100;2=100;10=100;18=100;26=100 <-- numa id 0, 1, 2, 10, 18,
>>>>                              26 as domain id
>>>> On type 2 platform, there won't be "MB:" line. Numa 0 and 1
>>>> are for mbm allocation on socket 0 and 1. 2,10, 18 and 26 are for GPU
>>>> memory nodes allocation.
>>
>> On type 3 platform, there could be "MB:" line for L3 cache and "MB_NODE:" for numa node. Example schemata is:
>>
>>       MB:1=100;2=100                 <-- cache id 1 and 2 as domain id
>> MB_NODE:0=100;1=100;2=100;10=100;18=100;26=100 <-- numa id 0, 1, 2, 10,
>>                                                     18, 26 as domain id
>>>
>>> (to help make things explicit I will refer to what you call "MBN" as "MB_NODE" to make it
>>> explicit that it is memory bandwidth allocation at node scope)
>>>
>>> I am trying to consider how this can be accomplished while also considering all the other
>>> new hardware features that resctrl need to support. Consider, for example, AMD's "Global
>>> MBA" (https://lore.kernel.org/lkml/cover.1776980182.git.babu.moger@amd.com/) that throttles
>>> memory bandwidth at L3 scope but the user configures allocations at NODE scope. At this time
>>> the plan is to support this with a second control associated with the MB resource that can
>>> allocate memory bandwidth at node scope. See
>>> https://lore.kernel.org/lkml/430ffb48-29f4-44d9-9164-9f8b743b2739@amd.com/
>>>
>>> If resctrl creates a new resource for node scoped memory bandwidth allocations to support these
>>> "type 2" systems then that will result in an inconsistent interface between architectures that
>>> we should avoid.
>>>
>>> Have you been listening in on the discussions surrounding emulated controls? Considering that,
>>> would it be possible to support the "MB" control on a type "2" system but have it be backed by
>>> (emulated by) the underlying "MB_NODE" control?
>>>
>>> resctrl could expose both controls on these "type 2" systems but make it clear that "MB"
>>> is emulated by "MB_NODE". For example:
>>>
>>> info/
>>> └── MB/
>>>       └── resource_schemata/
>>>           └── MB/
>>>               └── MB_NODE/
>>>
>>> User will see both controls in schemata file but when changes are made to "MB" control it
>>> will show in the "MB_NODE" control and vice-versa. User could also disable the "MB" control
>>> that will establish familiarity with the interface at which point resctrl can drop the
>>> "MB" control from the schemata file on these "type 2" systems.
>>>
>>> Having the MB resource available with an MB control will keep resctrl backward compatible
>>> if there are any tools that expect that. If backward compatibility is not of concern then
>>> resctrl could initialize with the emulated control disabled by default. See discussion at
>>> https://lore.kernel.org/lkml/5e575bc2-e67f-4696-9332-33c54023c057@intel.com/
>>> that describes a new resctrl capability in support of RISC-V and RDT.
>>> With this resctrl could initialize with:
>>>
>>> info/
>>> └── MB/
>>>       └── resource_schemata/
>>>           ├── MB/
>>>           │   ├── MB_NODE/
>>>           │   │   └── status:enabled
>>>           │   └── status:disabled
>>>           └── mode:legacy [native]
>>>
>>> With above a "type 2" system will boot with its schemata file just containing the "MB_NODE"
>>> control while info/MB describes the memory bandwidth resource.
>>>
>>
>> On type 3 machine, schemata has both MB in legacy mode with cache id as domain id and MB_NODE with numa id as domain id.
>>
>> Is this directory OK?
>>
>>   info/
>>   └── MB/
>>        └── resource_schemata/
>>            ├── MB/
>>            │   ├── MB_NODE/
>>            │   │   └── status:disabled
>>            │   └── status:enabled
>>            ├── MB_NODE/
>>            └── mode:node
>>
>> 1. MB and MB_NODE are shown in parallel in inf/MB/resource_schemata/
> 
> I do not think there is a need to expose an emulated MB_NODE control if the actual MB_NODE
> hardware control exists.
> 
>> 2. mode is set as "node" meaning "MB" is for L3 and "MB_NODE" is for numa node
> 
> I assume you mean "scope" instead of "mode"? (more below)
> 
>> 3. Emulation "MB_NODE" is disabled (or should the "MB_NODE" sub-dir be invisible?)
> 
> Right, I do not think emulation is needed here. No need to make it invisible since it should not exist.
> 
>>From what I understand these "type 3" machines could be simplified to:
> 
> info/
> └── MB/
>      └── resource_schemata/
>          ├── MB/
>          │   └── scope:L3
>          └── MB_NODE/
>              └── scope:NODE
> 
> Beyond this I believe that MPAM currently emulates the MB control with its "MB_MAX" control and users may want
> to make bandwidth allocations at the fine granularity that it supports. Taking this into account the interface
> may end up looking like:
> 
> info/
> └── MB/
>      └── resource_schemata/
>          ├── MB/
>          │   ├── MB_MAX/
>          │   │   └── scope:L3
>          │   └── scope:L3
>          └── MB_NODE/
>              └── scope:NODE
> 
> A system like above will thus have three schemata file entries:
> MB
> MB_MAX
> MB_NODE
> 
> Three schemata file entries would be unnecessary for users familiar with the finer granularity MB_MAX control
> so that is where the "mode" file can be used to disable the legacy MB control to just expose MB_MAX and MB_NODE
> on these systems.
> 
> Would that work for these systems?
> 
> ...
> 
>>>>>> There is another MPAM feature called MBW Max hardlimit which sets
>>>>>> "MB:" allocation as hardlimit (i.e. MBW throttling percentage must
>>>>>> be satisfied) per domain. Adding a new "MB_HLIM:" line in schemata.
>>>>>> It's 1:1 mapped to "MB:" to control hardlimit of MB throttling
>>>>>> percentage on each domain. By default hardlimit is off (0) and can
>>>>>> be turned on to set MBW Max hardlimit on a domain.
>>>>>
>>>>> ack. This sounds like a new control associated with the MB resource.
>>>>> This is a boolean control as Dave highlighted in previous discussion so
>>>>> resctrl would need to know its properties.
>>>>> See https://lore.kernel.org/lkml/aO0Oazuxt54hQFbx@e133380.arm.com/
>>>>>
>>>>
>>>> Right. ("MB_HLIM" name may be adjusted accordingly when "MB_MAX" is available.)
>>>>
>>>>>> For exmple:
>>>>>> MB_HLIM: 0=0;1=0;2=1;10=0;18=0;26=0
>>>>>> MB:0=100;1=100;2=80;10=100;18=100;26=100
>>>>>>
>>>>>> On GPU memory numa node 2: cannot use more than 80% of total max mbw even if there is still idle mem bandwidth on this node).
>>>>>>
>>>>>> MBW allocations on all other domains are soft limited, meaning MBW can be used more than specified if mem is idle.
>>>>>>
>>>>>
>>>>> ack.
>>>>>
>>>>>>>             L3:0=fff;1=fff
>>>>>>> # echo 'MB_MIN:0=50' > schemata
>>>>>>> # cat schemata
>>>>>>>             MB_MAX:0=100;1=100
>>>>>>>             MB_MIN:0=50;1=100
>>>>>>>             MB:0=100;1=100
>>>>>>>             L3:0=fff;1=fff
>>>>>>>
>>>>>>> Writing to the dummy control will call a dummy callback that just prints to the
>>>>>>> kernel log:
>>>>>>> "resctrl: Updata temporary MIN control on domain 0 with user value 50"
>>>>>>>
>>>>>>>
>>>>>>> Example output of info/MB/:
>>>>>>> /sys/fs/resctrl/info/MB/thread_throttle_mode:max
>>>>>>> /sys/fs/resctrl/info/MB/num_closids:15
>>>>>>> /sys/fs/resctrl/info/MB/delay_linear:1
>>>>>>> /sys/fs/resctrl/info/MB/min_bandwidth:10
>>>>>>
>>>>>> Add two new MB info RO files:
>>>>>> 1. /sys/fs/resctrl/info/MB/domain_id
>>>>>> It shows "numa" for using numa id in "MB:" or "cache" for using legacy cache id.
>>>>>
>>>>> This proposal introduces a *global* property to the MB *resource*? It does not seem as though
>>>>> this takes into account *anything* about how resctrl can support new hardware that has been
>>>>> discussed before, during, or after LPC. You have not participated in these discussions and
>>>>> now make an orthogonal proposal that does not take into account *any* of the requirements
>>>>> that we have been struggling with for months.
>>>>>
>>>>> Why should this proposal be taken seriously? In your absence folks have been trying to
>>>>> accommodate how these upcoming products and be supported and the "scope" file associated with
>>>>> a control is intended to communicate to user space how the domain ID should be interpreted.
>>>>>
>>>>> Why are you proposing something entirely different here without even acknowledging current
>>>>> approach and explaining why it does not work for you?
>>>>>
>>>>
>>>> So can I change this part to adding the following files in info dirctory?
>>>>
>>>> 1. For numa memory bw allocation (MBN):
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/resolution:100
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/tolerance:5
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/type:scalar
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/min:10
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/scale:1
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/scope:NUMA
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/unit:all
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/max:100
>>>
>>> This is not just about adding files to the info directory. The files, directories, their relationships,
>>> and content have meaning. All I see from these proposals is an attempt to slap some new files into
>>> resctrl without any consideration to present consistent interface to users and without consideration of
>>> other architectures that need to be supported by resctrl.
>>>
>>> resctrl needs to provide a generic and consistent interface to user space irrespective of the
>>> underlying architecture. Architectures cannot just slap some new files for their convenience.
>>>
>>>>
>>>>>> 2. /sys/fs/resctrl/info/MB/max_lim
>>>>>> It shows number 0-3 for MPAM MBW max limit behaviors: 0 for supporting both softlimit and hardlimit, etc.
>>>>>
>>>>> Again this adds another *global* property to the MB resource but then above you
>>>>> describe the new "MB_HLIM" schemata file entry that implies that it is a new control
>>>>> for the MB resource. Having it be a new control for the MB resource matches earlier
>>>>> discussions. To support this I thus expect it to be exposed as a new control with
>>>>> potentially a new type if any of the existing planned types do not suffice.
>>>>>
>>>>
>>>> How about adding these MB_HLIM dir and files in info?
>>>>
>>>> /sys/fs/resctrl/info/MB_HLIM/resource_schemata/MB_HLIM/type: boolean
>>>> /sys/fs/resctrl/info/MB_HLIM/resource_schemata/MB_HLIM/max_lim: 0
>>>
>>> This presents "MB_HLIM" as a *resource* to user space. It is not a resource
>>> but a *control* of a resource, no? I thus expect it to instead look something like
>>> below that makes it clear that MB_HARDMAX is a control of the MB resource.
>>>
>>> info
>>> └── MB
>>>       └── resource_schemata
>>>           ├── MB
>>>           └── MB_HARDMAX
>>
>> Yes, this makes sense. I have changed to this hierarchy.
> 
> Thank you very much for considering this approach.

[ MB_MAXHLIM: I use this name for MBW_MAX hard limit feature as Dave 
Martin suggested before. He also suggested MB_HARDMAX. Either name is 
good for me. I use MB_MAXHLIM to explain MBW_MAX hard limit for now.]

Some implementation thoughts:

MBW_MAX hard limit itself is not a MB control. Rather, it configures MB 
control, i.e. turn on MB control's hard limit or turn off its hard 
limit. So MBW_MAX hard limit doesn't have properties like 
bandwidth_gran, delay_linear, etc. MBW_MAX hard limit's property is only 
a boolean type.

So I would think it maybe a configuration inside a control.

Similar configurations could be hard limit for cache capacity in MPAM.

Maybe can add "configs" inside resctrl_ctrl. Schemata and 
info/MB/resource_schemata/MB will show/write the configurations per control?

For this configuration or future configurations, add "configs" list in:
struct resctrl_ctrl {
         struct list_head        entry;
         enum resctrl_scope      scope;
         struct list_head        domains;
         enum resctrl_ctrl_type  type;
         enum resctrl_ctrl_name  name;
         struct resctrl_ctrl     *emulated_by;
         struct list_head        configs; <--- Add configs for this control
         union {
                 struct resctrl_cache    cache;
                 struct resctrl_membw    membw;
         };
};

A resctrl control can have one or multiple configurations. Currently 
MBW_MAX hard limit is the only one. But the infrastrucutre supports 
multiple configurations per control.

schemata:
         MB:1=100   <-- MBW_MAX on L3 id 1
MB_MAXHLIM:1=0     <-- turn on/off MBW_MAX hardlimit on L3 id 1
         L3:1=fff

info/
├── MB
│   ├── bandwidth_gran
│   ├── delay_linear
│   ├── min_bandwidth
│   ├── num_closids
│   └── resource_schemata
│       ├── MB
│       │   ├── configs
│       │   │   └── MB_MAXHLIM
│       │   │       └── type   <--- bool
│       │   ├── max
│       │   ├── min
│       │   ├── resolution
│       │   ├── scale
│       │   ├── scope
│       │   ├── status
│       │   ├── tolerance
│       │   ├── type
│       │   └── unit
│       └── mode


Is this a valid way to handle MB_MAX hard limit (and future more 
configurations per control)?

Thanks.

-Fenghua

  parent reply	other threads:[~2026-07-23  0:17 UTC|newest]

Thread overview: 108+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-05-29 18:06 [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept Reinette Chatre
2026-06-02 20:23 ` Babu Moger
2026-06-02 22:56   ` Reinette Chatre
2026-06-03  1:14     ` Moger, Babu
2026-06-03  3:55       ` Reinette Chatre
2026-06-03 14:40         ` Babu Moger
2026-06-02 23:32 ` Chen, Yu C
2026-06-03  3:45   ` Reinette Chatre
2026-06-03 11:53     ` Chen, Yu C
2026-06-04 16:37       ` Reinette Chatre
2026-06-05 15:43         ` Chen, Yu C
2026-06-05 16:20           ` Reinette Chatre
2026-06-03 15:15 ` Ben Horgan
2026-06-03 19:34   ` Drew Fustini
2026-06-04 11:24     ` Ben Horgan
2026-06-04 17:38       ` Drew Fustini
2026-06-12  1:30         ` Shaopeng Tan (Fujitsu)
2026-06-17 15:29           ` Reinette Chatre
2026-06-19  1:42             ` Shaopeng Tan (Fujitsu)
2026-06-22 16:10               ` Reinette Chatre
2026-06-23  5:04                 ` Shaopeng Tan (Fujitsu)
2026-06-04 21:05     ` Reinette Chatre
2026-06-05 19:35       ` Drew Fustini
2026-06-06  5:10         ` Drew Fustini
2026-06-06  5:23           ` Drew Fustini
2026-06-04 17:43   ` Reinette Chatre
2026-06-05 14:53     ` Ben Horgan
2026-06-05 15:39       ` Reinette Chatre
2026-06-05 16:37         ` Ben Horgan
2026-06-08 16:16           ` Reinette Chatre
2026-06-09 10:10             ` Ben Horgan
2026-06-09 15:28               ` Reinette Chatre
2026-06-09 16:37                 ` Ben Horgan
2026-06-09 17:41                   ` Reinette Chatre
2026-06-10  7:09                     ` Chen, Yu C
2026-06-10 14:27                       ` Chen, Yu C
2026-06-10 16:13                         ` Reinette Chatre
2026-06-10 17:57                           ` Chen, Yu C
2026-06-10 18:10                             ` Reinette Chatre
2026-06-10 15:59                       ` Reinette Chatre
2026-06-10 18:05                         ` Chen, Yu C
2026-06-11  3:26                         ` Chen, Yu C
2026-06-11 15:45                           ` Reinette Chatre
2026-06-26 15:46                             ` Chen, Yu C
2026-07-02 14:27                               ` Ben Horgan
2026-07-03  9:01                                 ` Chen, Yu C
2026-07-14 21:37                               ` Reinette Chatre
2026-07-15  2:49                                 ` Chen, Yu C
2026-06-10  4:31                 ` Drew Fustini
2026-06-10 15:14                   ` Reinette Chatre
2026-06-03 18:46 ` Luck, Tony
2026-06-04 10:02   ` Ben Horgan
2026-06-04 21:42   ` Reinette Chatre
2026-07-08 12:56     ` Chen, Yu C
2026-07-14 21:39       ` Reinette Chatre
2026-06-03 22:14 ` Drew Fustini
2026-06-04 21:47   ` Reinette Chatre
2026-06-05 19:48     ` Drew Fustini
2026-06-15 21:05 ` Moger, Babu
2026-06-17 17:18   ` Reinette Chatre
2026-06-17 20:29     ` Babu Moger
2026-06-24 19:08 ` Fenghua Yu
2026-06-24 22:22   ` Reinette Chatre
2026-06-25  1:26     ` Fenghua Yu
2026-06-25 15:43       ` Reinette Chatre
2026-07-10 20:59         ` Fenghua Yu
2026-07-14 22:06           ` Reinette Chatre
2026-07-15  8:34             ` Ben Horgan
2026-07-15 15:41               ` Reinette Chatre
2026-07-16 14:59                 ` Ben Horgan
2026-07-16 16:02                   ` Luck, Tony
2026-07-16 16:22                     ` Ben Horgan
2026-07-16 17:50                       ` Reinette Chatre
2026-07-17 10:27                         ` Ben Horgan
2026-07-16 16:04                   ` Reinette Chatre
2026-07-16 16:44                     ` Ben Horgan
2026-07-16 17:07                       ` Reinette Chatre
2026-07-17 12:20                         ` Ben Horgan
2026-07-17 16:00                           ` Reinette Chatre
2026-07-20 13:30                             ` Ben Horgan
2026-07-20 22:54                               ` Reinette Chatre
2026-07-21 13:23                                 ` Ben Horgan
2026-07-21 17:30                                   ` Reinette Chatre
2026-07-21 20:02                                     ` Babu Moger
2026-07-22 10:47                                       ` Ben Horgan
2026-07-22 17:02                                         ` Babu Moger
2026-08-04  9:11                                           ` Ben Horgan
2026-08-04 14:09                                             ` Babu Moger
2026-08-04 15:11                                               ` Ben Horgan
2026-08-04 19:51                                                 ` Babu Moger
2026-08-05  6:06                                                 ` Reinette Chatre
2026-08-06  9:24                                                   ` Ben Horgan
2026-08-07 22:53                                                     ` Reinette Chatre
2026-07-22 10:03                                     ` Ben Horgan
2026-07-22 16:32                                       ` Reinette Chatre
2026-07-23  7:58                                         ` Ben Horgan
2026-07-23 15:52                                           ` Reinette Chatre
2026-07-23 22:08                                   ` Fenghua Yu
2026-07-27  9:13                                     ` Ben Horgan
2026-07-17 16:02                     ` Chen, Yu C
2026-07-17 16:55                       ` Reinette Chatre
2026-07-23  0:17             ` Fenghua Yu [this message]
2026-07-23 16:18               ` Reinette Chatre
2026-07-23 22:27                 ` Fenghua Yu
2026-07-23 23:31                   ` Reinette Chatre
2026-07-02 13:37       ` Ben Horgan
2026-07-02 15:16         ` Fenghua Yu
2026-07-03 13:42           ` Ben Horgan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=9425e9fc-36cf-44cd-b6fd-88b76d106eea@nvidia.com \
    --to=fenghuay@nvidia.com \
    --cc=Dave.Martin@arm.com \
    --cc=babu.moger@amd.com \
    --cc=ben.horgan@arm.com \
    --cc=bp@alien8.de \
    --cc=dave.hansen@linux.intel.com \
    --cc=fustini@kernel.org \
    --cc=james.morse@arm.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=peternewman@google.com \
    --cc=reinette.chatre@intel.com \
    --cc=tglx@linutronix.de \
    --cc=tony.luck@intel.com \
    --cc=x86@kernel.org \
    --cc=yu.c.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox