Linux Documentation
 help / color / mirror / Atom feed
From: Babu Moger <babu.moger@amd.com>
To: Reinette Chatre <reinette.chatre@intel.com>,
	tony.luck@intel.com, bp@alien8.de
Cc: x86@kernel.org, Dave.Martin@arm.com, james.morse@arm.com,
	corbet@lwn.net, skhan@linuxfoundation.org, tglx@kernel.org,
	mingo@redhat.com, dave.hansen@linux.intel.com, hpa@zytor.com,
	linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
	eranian@google.com, peternewman@google.com
Subject: Re: [PATCH 1/2] x86/resctrl, Documentation: Keep mbm_assign_mode "default" on boot
Date: Mon, 20 Jul 2026 14:15:56 -0500	[thread overview]
Message-ID: <78996219-a8de-4dc3-90ee-db4c19e0d66a@amd.com> (raw)
In-Reply-To: <f3fd774d-0662-4094-90b6-f0c5a4bc5d26@intel.com>

Hi Reinette,


On 7/20/26 13:27, Reinette Chatre wrote:
> Hi Babu,
> 
> On 7/20/26 10:00 AM, Babu Moger wrote:
>> Hi Reinette,
>>
>> Thanks for quick response.
>>
>> On 7/17/26 17:56, Reinette Chatre wrote:
>>> Hi Babu,
>>>
>>> On 7/17/26 2:13 PM, Babu Moger wrote:
>>>> The kernel currently enables the ABMC-based "mbm_event" mode by default on
>>>> hardware that supports it. However, this can cause bandwidth monitoring
>>>> failures with existing userspace tools such as pqos.
>>>>
>>>> The pqos tool mounts the resctrl filesystem and creates 16 or more resctrl
>>>> groups by default. On systems with 32 or fewer ABMC counters, this default
>>>> configuration can consume all available counters, since each group requires
>>>> one counter for local MBM and another for total MBM. If additional
>>>> monitoring groups are created, counter resources are exhausted and pqos
>>>> tool reports memory bandwidth counters as zero for those groups.
>>>
>>> It is not obvious to me that this is a problem. If I understand correctly
>>> there are two scenarios possible with this pqos behavior:
>>>
>>> - ABMC is not in use ("mbm_assign_mode" is set to "default")
>>>     - pqos can create 16 or more monitor groups
>>>     - hardware still supports a limited number of counters with consequence that
>>>       underlying counters reset at any time as the different monitoring groups
>>>       need to be tracked.
>>>     - pqos can read monitoring data of all 16 monitor groups, sometimes reading the
>>>       events would return "Unavailable", sometimes reading the events return data.
>>>     - *None* of the monitoring numbers returned are guaranteed to be accurate.
>>>
>>> - ABMC is in use ("mbm_assign_mode" is set to "mbm_event"):
>>>     - pqos can create 16 or more monitor groups
>>>     - only a subset of monitoring groups have counters assigned and these counters
>>>       are guaranteed to only track the monitor groups/events they are assigned to
>>>     - pqos can read monitoring data of all 16 monitor groups with two possibilities:
>>>       - monitor group/event has counter assigned: monitoring numbers are guaranteed to be accurate
>>>       - monitor group/event does not have counter assigned: monitoring numbers return 0
>>>
>>> If my understanding is correct then the preference is to rather have wrong data than
>>> see 0? This does not sound right. What am I missing?
>>
>>
>> This hardware can monitor up to 64 RMIDs without any counter resets.
> 
> How many RMIDs does the hardware claim to support via CPUID that ends up being shown to user
> space via "num_rmids"?

#cat /sys/fs/resctrl/info/L3_MON/num_rmids
4096

> 
> I understood from original ABMC enabling that the underlying hardware counters of "default"
> and "mbm_event" mode on AMD are the same. That is, in "default" mode the hardware does a
> "best effort" assignment of hardware counters to events while "mbm_event" mode lets the user
> control the assignment. It instead sounds like this is not the case and there are actually
> two distinct underlying hardware counter mechanisms?

That is correct. They are two different counters.

> 
> 
>> As you know, the pqos tool creates COS1 through COS15 regardless of
>> the command-line options used, resulting in a total of 16 groups
>> including the default group. With ABMC enabled, this consumes all
>> available ABMC counters(32 counters, 2 counters for each group).
> 
>>
>> When pqos is invoked with the -m option, it creates additional monitoring groups. For example:
>>
>> pqos -m all:0     -> creates 1 monitoring group
>> pqos -m all:0,1   -> creates 2 monitoring groups
>>
>>
>> Since all ABMC counters have already been allocated to the default
>> set of groups, no counters remain for these additional monitoring
>> groups. As a result, the monitoring commands report zero values,
>> effectively making monitoring unusable.
>>
>> In contrast, the default monitoring mode can still support up to 48
>> additional monitoring groups (64 total RMIDs minus the 16 default
>> groups created by pqos).
>>
>> For this reason, I still believe keeping the default monitoring mode as the default is the better option.
> 
> It is not clear to me where the "64" number comes from. Even if resctrl sets the "default"

The count of 64 is known from internal information. It can also be 
determined by allocating monitoring counters in a loop until the 
hardware starts returning an "unavailable" response, which occurs after 
all 64 counters have been assigned. This information is not documented.

> mode as default, what will happen to the scenario you describe and 49, instead of 48,
> additional groups are created? From what I understand the moment the 65th group is created user will
> transition from "accurate per monitor group monitoring data for all 64 monitoring groups" to "inaccurate
> per monitor group monitoring data for all 65 monitoring groups" with no indication that this is happening?

Yes. That is correct.

> 
> I believe AMD supports more than 64 RMIDs and in this case there seems to be three ranges:
> [supported by ABMC, depends on events but lets say RMID 0 to 16] < [RMID 17 to 64] < [RMID 65 to total number of RMIDs supported]
> 
> Current default "mbm_event" mode uses ABMC so as you state this always results in:
> - accurate counts for 16 monitor groups
> - zero for all other monitor groups up to total number of RMIDs supported

That is correct. With ABMC, there are 32 available counters, allowing up 
to 32 monitoring events to be tracked simultaneously. Since each 
monitoring group requires two counters—one for local MBM and one for 
total MBM—the system can support only 16 monitoring groups at a time.

> 
> As I see it switching to the "default" mode would result in:
> Scenario 1, 64 or fewer monitor groups are created:
> - accurate counts for all monitor groups
> Scenario 2, 65 or more monitor groups are created:
> - inaccurate counts for all monitor groups

That is true.

> 
> I do not believe there is any way for user space to know when or if system switches from "scenario 1" to
> "scenario 2" on these systems and this unpredictable behavior that results in wrong data does not
> sound ideal to me.
  I agree, it's not an ideal situation, but that's how it has always 
worked. At the moment, I'm not sure about a better alternative.

Thanks,
Babu

  reply	other threads:[~2026-07-20 19:16 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-17 21:13 [PATCH 1/2] x86/resctrl, Documentation: Keep mbm_assign_mode "default" on boot Babu Moger
2026-07-17 21:13 ` [PATCH 2/2] x86/resctrl: Fix ABMC counter programming for extended counter ranges Babu Moger
2026-07-17 22:56 ` [PATCH 1/2] x86/resctrl, Documentation: Keep mbm_assign_mode "default" on boot Reinette Chatre
2026-07-20 17:00   ` Babu Moger
2026-07-20 18:27     ` Reinette Chatre
2026-07-20 19:15       ` Babu Moger [this message]
2026-07-21 20:54         ` Reinette Chatre
2026-07-20 20:12 ` Babu Moger

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=78996219-a8de-4dc3-90ee-db4c19e0d66a@amd.com \
    --to=babu.moger@amd.com \
    --cc=Dave.Martin@arm.com \
    --cc=bp@alien8.de \
    --cc=corbet@lwn.net \
    --cc=dave.hansen@linux.intel.com \
    --cc=eranian@google.com \
    --cc=hpa@zytor.com \
    --cc=james.morse@arm.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mingo@redhat.com \
    --cc=peternewman@google.com \
    --cc=reinette.chatre@intel.com \
    --cc=skhan@linuxfoundation.org \
    --cc=tglx@kernel.org \
    --cc=tony.luck@intel.com \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox