From: Fenghua Yu <fenghuay@nvidia.com>
To: Reinette Chatre <reinette.chatre@intel.com>,
Tony Luck <tony.luck@intel.com>, Ben Horgan <ben.horgan@arm.com>,
James Morse <james.morse@arm.com>,
Dave Martin <Dave.Martin@arm.com>,
Babu Moger <babu.moger@amd.com>,
Drew Fustini <fustini@kernel.org>, Chen Yu <yu.c.chen@intel.com>
Cc: Borislav Petkov <bp@alien8.de>,
Thomas Gleixner <tglx@linutronix.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
Peter Newman <peternewman@google.com>,
"x86@kernel.org" <x86@kernel.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: Re: [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept
Date: Wed, 22 Jul 2026 17:17:20 -0700 [thread overview]
Message-ID: <9425e9fc-36cf-44cd-b6fd-88b76d106eea@nvidia.com> (raw)
In-Reply-To: <36163a81-9737-49e3-93ef-6c392f7272f0@intel.com>
Hi, Reinette,
On 7/14/26 15:06, Reinette Chatre wrote:
> Hi Fenghua,
>
> On 7/10/26 1:59 PM, Fenghua Yu wrote:
>> On 6/25/26 08:43, Reinette Chatre wrote:
>>> On 6/24/26 6:26 PM, Fenghua Yu wrote:
>>>> On 6/24/26 15:22, Reinette Chatre wrote:
>>>>> On 6/24/26 12:08 PM, Fenghua Yu wrote:
>>>>>> On 5/29/26 11:06, Reinette Chatre wrote:
>>>>>>
>>>>>> As Shaopen and Ben mentioned earlier, we are working on two MPAM
>>>>>> features that may need to change schemata interface. The CPU-less
>>>>>> feature was discussed on LPC (although the interfaces will be
>>>>>> slightly different from the LPC).
>>>>>
>>>>> I know. Here is where I tried to engage with you on needed interfaces after LPC:
>>>>> https://lore.kernel.org/lkml/fb1e2686-237b-4536-acd6-15159abafcba@intel.com/
>>>>
>>>> MPAM ACPI defines MSC (Memory System Control) is defined in one of two ways (not both) on one platform:
>>>> 1. L3 and memory together on each processor MSC
>>>> 2. L3 in processor MSC and memory control/monitoring in different memory MSCs.
>>
>> Ben said there is type 3 platform:
>> 3. L3 cache and memory bandwidth in processor MSCs and memory bandwidth in different memory MSCs.
>>
>>>>
>>>> On type 1 platform, schemata is legacy:
>>>> MB:1=100;2=100 <-- cache id 1 and 2 as domain id
>>>>
>>>> On type 2 platform, I will not reuse "MB:" name. Instead, define new resource name "MBN:" for numa node and schemata is:
>>>> MBN:0=100;1=100;2=100;10=100;18=100;26=100 <-- numa id 0, 1, 2, 10, 18,
>>>> 26 as domain id
>>>> On type 2 platform, there won't be "MB:" line. Numa 0 and 1
>>>> are for mbm allocation on socket 0 and 1. 2,10, 18 and 26 are for GPU
>>>> memory nodes allocation.
>>
>> On type 3 platform, there could be "MB:" line for L3 cache and "MB_NODE:" for numa node. Example schemata is:
>>
>> MB:1=100;2=100 <-- cache id 1 and 2 as domain id
>> MB_NODE:0=100;1=100;2=100;10=100;18=100;26=100 <-- numa id 0, 1, 2, 10,
>> 18, 26 as domain id
>>>
>>> (to help make things explicit I will refer to what you call "MBN" as "MB_NODE" to make it
>>> explicit that it is memory bandwidth allocation at node scope)
>>>
>>> I am trying to consider how this can be accomplished while also considering all the other
>>> new hardware features that resctrl need to support. Consider, for example, AMD's "Global
>>> MBA" (https://lore.kernel.org/lkml/cover.1776980182.git.babu.moger@amd.com/) that throttles
>>> memory bandwidth at L3 scope but the user configures allocations at NODE scope. At this time
>>> the plan is to support this with a second control associated with the MB resource that can
>>> allocate memory bandwidth at node scope. See
>>> https://lore.kernel.org/lkml/430ffb48-29f4-44d9-9164-9f8b743b2739@amd.com/
>>>
>>> If resctrl creates a new resource for node scoped memory bandwidth allocations to support these
>>> "type 2" systems then that will result in an inconsistent interface between architectures that
>>> we should avoid.
>>>
>>> Have you been listening in on the discussions surrounding emulated controls? Considering that,
>>> would it be possible to support the "MB" control on a type "2" system but have it be backed by
>>> (emulated by) the underlying "MB_NODE" control?
>>>
>>> resctrl could expose both controls on these "type 2" systems but make it clear that "MB"
>>> is emulated by "MB_NODE". For example:
>>>
>>> info/
>>> └── MB/
>>> └── resource_schemata/
>>> └── MB/
>>> └── MB_NODE/
>>>
>>> User will see both controls in schemata file but when changes are made to "MB" control it
>>> will show in the "MB_NODE" control and vice-versa. User could also disable the "MB" control
>>> that will establish familiarity with the interface at which point resctrl can drop the
>>> "MB" control from the schemata file on these "type 2" systems.
>>>
>>> Having the MB resource available with an MB control will keep resctrl backward compatible
>>> if there are any tools that expect that. If backward compatibility is not of concern then
>>> resctrl could initialize with the emulated control disabled by default. See discussion at
>>> https://lore.kernel.org/lkml/5e575bc2-e67f-4696-9332-33c54023c057@intel.com/
>>> that describes a new resctrl capability in support of RISC-V and RDT.
>>> With this resctrl could initialize with:
>>>
>>> info/
>>> └── MB/
>>> └── resource_schemata/
>>> ├── MB/
>>> │ ├── MB_NODE/
>>> │ │ └── status:enabled
>>> │ └── status:disabled
>>> └── mode:legacy [native]
>>>
>>> With above a "type 2" system will boot with its schemata file just containing the "MB_NODE"
>>> control while info/MB describes the memory bandwidth resource.
>>>
>>
>> On type 3 machine, schemata has both MB in legacy mode with cache id as domain id and MB_NODE with numa id as domain id.
>>
>> Is this directory OK?
>>
>> info/
>> └── MB/
>> └── resource_schemata/
>> ├── MB/
>> │ ├── MB_NODE/
>> │ │ └── status:disabled
>> │ └── status:enabled
>> ├── MB_NODE/
>> └── mode:node
>>
>> 1. MB and MB_NODE are shown in parallel in inf/MB/resource_schemata/
>
> I do not think there is a need to expose an emulated MB_NODE control if the actual MB_NODE
> hardware control exists.
>
>> 2. mode is set as "node" meaning "MB" is for L3 and "MB_NODE" is for numa node
>
> I assume you mean "scope" instead of "mode"? (more below)
>
>> 3. Emulation "MB_NODE" is disabled (or should the "MB_NODE" sub-dir be invisible?)
>
> Right, I do not think emulation is needed here. No need to make it invisible since it should not exist.
>
>>From what I understand these "type 3" machines could be simplified to:
>
> info/
> └── MB/
> └── resource_schemata/
> ├── MB/
> │ └── scope:L3
> └── MB_NODE/
> └── scope:NODE
>
> Beyond this I believe that MPAM currently emulates the MB control with its "MB_MAX" control and users may want
> to make bandwidth allocations at the fine granularity that it supports. Taking this into account the interface
> may end up looking like:
>
> info/
> └── MB/
> └── resource_schemata/
> ├── MB/
> │ ├── MB_MAX/
> │ │ └── scope:L3
> │ └── scope:L3
> └── MB_NODE/
> └── scope:NODE
>
> A system like above will thus have three schemata file entries:
> MB
> MB_MAX
> MB_NODE
>
> Three schemata file entries would be unnecessary for users familiar with the finer granularity MB_MAX control
> so that is where the "mode" file can be used to disable the legacy MB control to just expose MB_MAX and MB_NODE
> on these systems.
>
> Would that work for these systems?
>
> ...
>
>>>>>> There is another MPAM feature called MBW Max hardlimit which sets
>>>>>> "MB:" allocation as hardlimit (i.e. MBW throttling percentage must
>>>>>> be satisfied) per domain. Adding a new "MB_HLIM:" line in schemata.
>>>>>> It's 1:1 mapped to "MB:" to control hardlimit of MB throttling
>>>>>> percentage on each domain. By default hardlimit is off (0) and can
>>>>>> be turned on to set MBW Max hardlimit on a domain.
>>>>>
>>>>> ack. This sounds like a new control associated with the MB resource.
>>>>> This is a boolean control as Dave highlighted in previous discussion so
>>>>> resctrl would need to know its properties.
>>>>> See https://lore.kernel.org/lkml/aO0Oazuxt54hQFbx@e133380.arm.com/
>>>>>
>>>>
>>>> Right. ("MB_HLIM" name may be adjusted accordingly when "MB_MAX" is available.)
>>>>
>>>>>> For exmple:
>>>>>> MB_HLIM: 0=0;1=0;2=1;10=0;18=0;26=0
>>>>>> MB:0=100;1=100;2=80;10=100;18=100;26=100
>>>>>>
>>>>>> On GPU memory numa node 2: cannot use more than 80% of total max mbw even if there is still idle mem bandwidth on this node).
>>>>>>
>>>>>> MBW allocations on all other domains are soft limited, meaning MBW can be used more than specified if mem is idle.
>>>>>>
>>>>>
>>>>> ack.
>>>>>
>>>>>>> L3:0=fff;1=fff
>>>>>>> # echo 'MB_MIN:0=50' > schemata
>>>>>>> # cat schemata
>>>>>>> MB_MAX:0=100;1=100
>>>>>>> MB_MIN:0=50;1=100
>>>>>>> MB:0=100;1=100
>>>>>>> L3:0=fff;1=fff
>>>>>>>
>>>>>>> Writing to the dummy control will call a dummy callback that just prints to the
>>>>>>> kernel log:
>>>>>>> "resctrl: Updata temporary MIN control on domain 0 with user value 50"
>>>>>>>
>>>>>>>
>>>>>>> Example output of info/MB/:
>>>>>>> /sys/fs/resctrl/info/MB/thread_throttle_mode:max
>>>>>>> /sys/fs/resctrl/info/MB/num_closids:15
>>>>>>> /sys/fs/resctrl/info/MB/delay_linear:1
>>>>>>> /sys/fs/resctrl/info/MB/min_bandwidth:10
>>>>>>
>>>>>> Add two new MB info RO files:
>>>>>> 1. /sys/fs/resctrl/info/MB/domain_id
>>>>>> It shows "numa" for using numa id in "MB:" or "cache" for using legacy cache id.
>>>>>
>>>>> This proposal introduces a *global* property to the MB *resource*? It does not seem as though
>>>>> this takes into account *anything* about how resctrl can support new hardware that has been
>>>>> discussed before, during, or after LPC. You have not participated in these discussions and
>>>>> now make an orthogonal proposal that does not take into account *any* of the requirements
>>>>> that we have been struggling with for months.
>>>>>
>>>>> Why should this proposal be taken seriously? In your absence folks have been trying to
>>>>> accommodate how these upcoming products and be supported and the "scope" file associated with
>>>>> a control is intended to communicate to user space how the domain ID should be interpreted.
>>>>>
>>>>> Why are you proposing something entirely different here without even acknowledging current
>>>>> approach and explaining why it does not work for you?
>>>>>
>>>>
>>>> So can I change this part to adding the following files in info dirctory?
>>>>
>>>> 1. For numa memory bw allocation (MBN):
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/resolution:100
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/tolerance:5
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/type:scalar
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/min:10
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/scale:1
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/scope:NUMA
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/unit:all
>>>> /sys/fs/resctrl/info/MBN/resource_schemata/MBN/max:100
>>>
>>> This is not just about adding files to the info directory. The files, directories, their relationships,
>>> and content have meaning. All I see from these proposals is an attempt to slap some new files into
>>> resctrl without any consideration to present consistent interface to users and without consideration of
>>> other architectures that need to be supported by resctrl.
>>>
>>> resctrl needs to provide a generic and consistent interface to user space irrespective of the
>>> underlying architecture. Architectures cannot just slap some new files for their convenience.
>>>
>>>>
>>>>>> 2. /sys/fs/resctrl/info/MB/max_lim
>>>>>> It shows number 0-3 for MPAM MBW max limit behaviors: 0 for supporting both softlimit and hardlimit, etc.
>>>>>
>>>>> Again this adds another *global* property to the MB resource but then above you
>>>>> describe the new "MB_HLIM" schemata file entry that implies that it is a new control
>>>>> for the MB resource. Having it be a new control for the MB resource matches earlier
>>>>> discussions. To support this I thus expect it to be exposed as a new control with
>>>>> potentially a new type if any of the existing planned types do not suffice.
>>>>>
>>>>
>>>> How about adding these MB_HLIM dir and files in info?
>>>>
>>>> /sys/fs/resctrl/info/MB_HLIM/resource_schemata/MB_HLIM/type: boolean
>>>> /sys/fs/resctrl/info/MB_HLIM/resource_schemata/MB_HLIM/max_lim: 0
>>>
>>> This presents "MB_HLIM" as a *resource* to user space. It is not a resource
>>> but a *control* of a resource, no? I thus expect it to instead look something like
>>> below that makes it clear that MB_HARDMAX is a control of the MB resource.
>>>
>>> info
>>> └── MB
>>> └── resource_schemata
>>> ├── MB
>>> └── MB_HARDMAX
>>
>> Yes, this makes sense. I have changed to this hierarchy.
>
> Thank you very much for considering this approach.
[ MB_MAXHLIM: I use this name for MBW_MAX hard limit feature as Dave
Martin suggested before. He also suggested MB_HARDMAX. Either name is
good for me. I use MB_MAXHLIM to explain MBW_MAX hard limit for now.]
Some implementation thoughts:
MBW_MAX hard limit itself is not a MB control. Rather, it configures MB
control, i.e. turn on MB control's hard limit or turn off its hard
limit. So MBW_MAX hard limit doesn't have properties like
bandwidth_gran, delay_linear, etc. MBW_MAX hard limit's property is only
a boolean type.
So I would think it maybe a configuration inside a control.
Similar configurations could be hard limit for cache capacity in MPAM.
Maybe can add "configs" inside resctrl_ctrl. Schemata and
info/MB/resource_schemata/MB will show/write the configurations per control?
For this configuration or future configurations, add "configs" list in:
struct resctrl_ctrl {
struct list_head entry;
enum resctrl_scope scope;
struct list_head domains;
enum resctrl_ctrl_type type;
enum resctrl_ctrl_name name;
struct resctrl_ctrl *emulated_by;
struct list_head configs; <--- Add configs for this control
union {
struct resctrl_cache cache;
struct resctrl_membw membw;
};
};
A resctrl control can have one or multiple configurations. Currently
MBW_MAX hard limit is the only one. But the infrastrucutre supports
multiple configurations per control.
schemata:
MB:1=100 <-- MBW_MAX on L3 id 1
MB_MAXHLIM:1=0 <-- turn on/off MBW_MAX hardlimit on L3 id 1
L3:1=fff
info/
├── MB
│ ├── bandwidth_gran
│ ├── delay_linear
│ ├── min_bandwidth
│ ├── num_closids
│ └── resource_schemata
│ ├── MB
│ │ ├── configs
│ │ │ └── MB_MAXHLIM
│ │ │ └── type <--- bool
│ │ ├── max
│ │ ├── min
│ │ ├── resolution
│ │ ├── scale
│ │ ├── scope
│ │ ├── status
│ │ ├── tolerance
│ │ ├── type
│ │ └── unit
│ └── mode
Is this a valid way to handle MB_MAX hard limit (and future more
configurations per control)?
Thanks.
-Fenghua
next prev parent reply other threads:[~2026-07-23 0:17 UTC|newest]
Thread overview: 95+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-05-29 18:06 [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept Reinette Chatre
2026-06-02 20:23 ` Babu Moger
2026-06-02 22:56 ` Reinette Chatre
2026-06-03 1:14 ` Moger, Babu
2026-06-03 3:55 ` Reinette Chatre
2026-06-03 14:40 ` Babu Moger
2026-06-02 23:32 ` Chen, Yu C
2026-06-03 3:45 ` Reinette Chatre
2026-06-03 11:53 ` Chen, Yu C
2026-06-04 16:37 ` Reinette Chatre
2026-06-05 15:43 ` Chen, Yu C
2026-06-05 16:20 ` Reinette Chatre
2026-06-03 15:15 ` Ben Horgan
2026-06-03 19:34 ` Drew Fustini
2026-06-04 11:24 ` Ben Horgan
2026-06-04 17:38 ` Drew Fustini
2026-06-12 1:30 ` Shaopeng Tan (Fujitsu)
2026-06-17 15:29 ` Reinette Chatre
2026-06-19 1:42 ` Shaopeng Tan (Fujitsu)
2026-06-22 16:10 ` Reinette Chatre
2026-06-23 5:04 ` Shaopeng Tan (Fujitsu)
2026-06-04 21:05 ` Reinette Chatre
2026-06-05 19:35 ` Drew Fustini
2026-06-06 5:10 ` Drew Fustini
2026-06-06 5:23 ` Drew Fustini
2026-06-04 17:43 ` Reinette Chatre
2026-06-05 14:53 ` Ben Horgan
2026-06-05 15:39 ` Reinette Chatre
2026-06-05 16:37 ` Ben Horgan
2026-06-08 16:16 ` Reinette Chatre
2026-06-09 10:10 ` Ben Horgan
2026-06-09 15:28 ` Reinette Chatre
2026-06-09 16:37 ` Ben Horgan
2026-06-09 17:41 ` Reinette Chatre
2026-06-10 7:09 ` Chen, Yu C
2026-06-10 14:27 ` Chen, Yu C
2026-06-10 16:13 ` Reinette Chatre
2026-06-10 17:57 ` Chen, Yu C
2026-06-10 18:10 ` Reinette Chatre
2026-06-10 15:59 ` Reinette Chatre
2026-06-10 18:05 ` Chen, Yu C
2026-06-11 3:26 ` Chen, Yu C
2026-06-11 15:45 ` Reinette Chatre
2026-06-26 15:46 ` Chen, Yu C
2026-07-02 14:27 ` Ben Horgan
2026-07-03 9:01 ` Chen, Yu C
2026-07-14 21:37 ` Reinette Chatre
2026-07-15 2:49 ` Chen, Yu C
2026-06-10 4:31 ` Drew Fustini
2026-06-10 15:14 ` Reinette Chatre
2026-06-03 18:46 ` Luck, Tony
2026-06-04 10:02 ` Ben Horgan
2026-06-04 21:42 ` Reinette Chatre
2026-07-08 12:56 ` Chen, Yu C
2026-07-14 21:39 ` Reinette Chatre
2026-06-03 22:14 ` Drew Fustini
2026-06-04 21:47 ` Reinette Chatre
2026-06-05 19:48 ` Drew Fustini
2026-06-15 21:05 ` Moger, Babu
2026-06-17 17:18 ` Reinette Chatre
2026-06-17 20:29 ` Babu Moger
2026-06-24 19:08 ` Fenghua Yu
2026-06-24 22:22 ` Reinette Chatre
2026-06-25 1:26 ` Fenghua Yu
2026-06-25 15:43 ` Reinette Chatre
2026-07-10 20:59 ` Fenghua Yu
2026-07-14 22:06 ` Reinette Chatre
2026-07-15 8:34 ` Ben Horgan
2026-07-15 15:41 ` Reinette Chatre
2026-07-16 14:59 ` Ben Horgan
2026-07-16 16:02 ` Luck, Tony
2026-07-16 16:22 ` Ben Horgan
2026-07-16 17:50 ` Reinette Chatre
2026-07-17 10:27 ` Ben Horgan
2026-07-16 16:04 ` Reinette Chatre
2026-07-16 16:44 ` Ben Horgan
2026-07-16 17:07 ` Reinette Chatre
2026-07-17 12:20 ` Ben Horgan
2026-07-17 16:00 ` Reinette Chatre
2026-07-20 13:30 ` Ben Horgan
2026-07-20 22:54 ` Reinette Chatre
2026-07-21 13:23 ` Ben Horgan
2026-07-21 17:30 ` Reinette Chatre
2026-07-21 20:02 ` Babu Moger
2026-07-22 10:47 ` Ben Horgan
2026-07-22 17:02 ` Babu Moger
2026-07-22 10:03 ` Ben Horgan
2026-07-22 16:32 ` Reinette Chatre
2026-07-23 7:58 ` Ben Horgan
2026-07-17 16:02 ` Chen, Yu C
2026-07-17 16:55 ` Reinette Chatre
2026-07-23 0:17 ` Fenghua Yu [this message]
2026-07-02 13:37 ` Ben Horgan
2026-07-02 15:16 ` Fenghua Yu
2026-07-03 13:42 ` Ben Horgan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9425e9fc-36cf-44cd-b6fd-88b76d106eea@nvidia.com \
--to=fenghuay@nvidia.com \
--cc=Dave.Martin@arm.com \
--cc=babu.moger@amd.com \
--cc=ben.horgan@arm.com \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=fustini@kernel.org \
--cc=james.morse@arm.com \
--cc=linux-kernel@vger.kernel.org \
--cc=peternewman@google.com \
--cc=reinette.chatre@intel.com \
--cc=tglx@linutronix.de \
--cc=tony.luck@intel.com \
--cc=x86@kernel.org \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.