From: Reinette Chatre <reinette.chatre@intel.com>
To: Tony Luck <tony.luck@intel.com>, Fenghua Yu <fenghuay@nvidia.com>,
"Maciej Wieczor-Retman" <maciej.wieczor-retman@intel.com>,
Peter Newman <peternewman@google.com>,
James Morse <james.morse@arm.com>,
Babu Moger <babu.moger@amd.com>,
Drew Fustini <dfustini@baylibre.com>,
Dave Martin <Dave.Martin@arm.com>, Chen Yu <yu.c.chen@intel.com>,
David E Box <david.e.box@intel.com>, <x86@kernel.org>
Cc: Christoph Hellwig <hch@infradead.org>,
<linux-kernel@vger.kernel.org>, <patches@lists.linux.dev>
Subject: Re: [PATCH v11 12/23] arm,x86,fs/resctrl: Handle change in number of RMIDs on each mount
Date: Wed, 9 Sep 2026 21:01:35 -0700 [thread overview]
Message-ID: <0ef770ec-ac65-4a80-b676-a8bd92caea88@intel.com> (raw)
In-Reply-To: <20260831174421.13921-13-tony.luck@intel.com>
Hi Tony,
On 8/31/26 10:44 AM, Tony Luck wrote:
> Application Energy Telemetry (AET) event enumeration takes place
> asynchronously. Linux builds the pmt_telemetry module into the kernel to
> kick off enumeration early enough that it completes before first mount of
> the resctrl file system.
>
> Allowing pmt_telemetry to be a loadable module means that it is possible
> for different numbers of RMIDs to be supported on each mount, depending
> on whether pmt_telemetry module is loaded.
>
> For simplicity, calculate the maximum possible number of RMIDs and use
> that value to allocate the rmid_ptrs[] array just once. Use this same
> calculated value for all references to rmid_ptrs[] instead of calling
> resctrl_arch_system_max_rmid_idx() in multiple places.
>
> Add resctrl_arch_get_num_rmid_idx(r) to report the maximum RMID index
> for a resource. Use it to allocate the rdt_l3_mon_domain::rmid_busy_llc
> bitmap and rdt_l3_mon_domain::mbm_states and when operating on these
> structures.
>
> The limbo code must deal with changes in the number of RMIDs from one
> mount to the next because some RMIDs may still be "busy" when the file
> system is unmounted, but be above resctrl_arch_system_num_rmid_idx()
> for the remount. In this case RMIDs that can be released are not put
> onto the rmid_free_lru list.
(last sentence needs imperative)
The changelog is reading more like a list of changes. Can some of these
be split out?
>
> Signed-off-by: Tony Luck <tony.luck@intel.com>
> ---
> v11:
> Add resctrl_arch_get_num_rmid_idx(r) and use it for allocating
> an operating on L3 per-domain dynamically allocated structures.
> Update kernel doc comment for max_idx_limit.
> Update comment to explain why all RMIDs need to be checked
> for LLC cache occupancy.
>
> include/linux/resctrl.h | 8 ++-
> arch/x86/kernel/cpu/resctrl/core.c | 30 +++++++++++
> drivers/resctrl/mpam_resctrl.c | 14 +++++
> fs/resctrl/monitor.c | 87 +++++++++++++++++++++---------
> fs/resctrl/rdtgroup.c | 6 +--
> 5 files changed, 114 insertions(+), 31 deletions(-)
>
> diff --git a/include/linux/resctrl.h b/include/linux/resctrl.h
> index e9094d886ba7..4fb06d434c85 100644
> --- a/include/linux/resctrl.h
> +++ b/include/linux/resctrl.h
> @@ -183,10 +183,12 @@ struct mbm_cntr_cfg {
> * struct rdt_l3_mon_domain - group of CPUs sharing RDT_RESOURCE_L3 monitoring
> * @hdr: common header for different domain types
> * @ci_id: cache info id for this domain
> - * @rmid_busy_llc: bitmap of which limbo RMIDs are above threshold
> + * @rmid_busy_llc: bitmap of which limbo RMIDs are above threshold. Sized for
> + * maximum supported RMIDs in L3 resource.
> * @mbm_states: Per-event pointer to the MBM event's saved state.
> * An MBM event's state is an array of struct mbm_state
> * indexed by RMID on x86 or combined CLOSID, RMID on Arm.
> + * Sized same as @rmid_busy_llc.
> * @mbm_over: worker to periodically read MBM h/w counters
> * @cqm_limbo: worker to periodically read CQM h/w counters
> * @mbm_work_cpu: worker CPU for MBM h/w counters
> @@ -440,9 +442,11 @@ static inline u32 resctrl_get_default_ctrl(struct rdt_resource *r)
> return WARN_ON_ONCE(1);
> }
>
> -/* The number of closid supported by this resource regardless of CDP */
> +/* The number of closid/rmid supported by this resource regardless of CDP */
> u32 resctrl_arch_get_num_closid(struct rdt_resource *r);
> +u32 resctrl_arch_get_num_rmid_idx(struct rdt_resource *r);
> u32 resctrl_arch_system_num_rmid_idx(void);
> +u32 resctrl_arch_system_max_rmid_idx(void);
> int resctrl_arch_update_domains(struct rdt_resource *r, u32 closid);
>
> /**
> diff --git a/arch/x86/kernel/cpu/resctrl/core.c b/arch/x86/kernel/cpu/resctrl/core.c
> index e851da431dd9..ef37fbb586d3 100644
> --- a/arch/x86/kernel/cpu/resctrl/core.c
> +++ b/arch/x86/kernel/cpu/resctrl/core.c
> @@ -124,6 +124,31 @@ u32 resctrl_arch_system_num_rmid_idx(void)
> return num_rmids == U32_MAX ? 0 : num_rmids;
> }
>
> +/**
> + * resctrl_arch_system_max_rmid_idx - Largest possible number of RMIDs
> + *
> + * Return: Maximum possible number of RMIDs used for boot time allocations.
> + */
> +u32 resctrl_arch_system_max_rmid_idx(void)
> +{
> + struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3].r_resctrl;
> + u32 ret;
> +
> + /* CPUID enumerates maximum value that can be written to IA32_PQR_ASSOC.RMID */
> + ret = cpuid_ebx(0xf) + 1;
This patch with its one caller of resctrl_arch_system_max_rmid_idx() seems ok but this
series adds more callers and with that the repeated CPUID does not seem necessary.
Looking back I wonder if it will not make this work easier to consume if this RMID
count is instead enumerated from get_rdt_mon_resources() and stored in a global
variable. I think doing so would help to understand the dependencies among and capabilities
of the related feature bits. Something like:
get_rdt_mon_resources()
{
if (!cpu_feature_enabled(X86_FEATURE_RDT_M)) /* or boot_cpu_has() */
return false;
/* Maximum value that can be written to IA32_PQR_ASSOC.RMID */
pqr_assoc_max_rmid = cpuid_ebx(0xf) + 1;
if (!cpu_feature_enabled(X86_FEATURE_L3_MON)) /* or boot_cpu_has() */
return pqr_assoc_max_rmid > 0;
/* L3 monitoring enumeration */
return resctrl_arch_system_max_rmid_idx() > 0;
}
With something like above resctrl_arch_system_max_rmid_idx() could use pqr_assoc_max_rmid
instead of calling CPUID every time?
> +
> + /*
> + * If the system is capable of L3 monitoring the maximum RMID value may
> + * be lower than the system maximum. Either because the L3 monitoring
> + * feature supports fewer RMIDs, or because SNC (Sub-NUMA Cluster)
> + * is enabled and divides RMIDs per cluster.
> + */
> + if (r->mon_capable)
> + ret = r->mon.num_rmid;
> +
> + return ret;
> +}
> +
> struct rdt_resource *resctrl_arch_get_resource(enum resctrl_res_level l)
> {
> if (l >= RDT_NUM_RESOURCES)
> @@ -360,6 +385,11 @@ u32 resctrl_arch_get_num_closid(struct rdt_resource *r)
> return resctrl_to_arch_res(r)->num_closid;
> }
>
> +u32 resctrl_arch_get_num_rmid_idx(struct rdt_resource *r)
> +{
> + return r->mon.num_rmid;
> +}
> +
> void rdt_ctrl_update(void *arg)
> {
> struct rdt_hw_resource *hw_res;
> diff --git a/drivers/resctrl/mpam_resctrl.c b/drivers/resctrl/mpam_resctrl.c
> index 0db62dd2a71c..a117aa98ae90 100644
> --- a/drivers/resctrl/mpam_resctrl.c
> +++ b/drivers/resctrl/mpam_resctrl.c
> @@ -247,11 +247,25 @@ u32 resctrl_arch_get_num_closid(struct rdt_resource *ignored)
> return mpam_partid_max + 1;
> }
>
> +/*
> + * File system calls this for one-time allocation of structures
> + * during initialization. Return the largest possible value.
> + */
Above comment seems more appropriate for resctrl_arch_system_max_rmid_idx().
> +u32 resctrl_arch_get_num_rmid_idx(struct rdt_resource *ignored)
> +{
> + return resctrl_arch_system_num_rmid_idx();
> +}
> +
> u32 resctrl_arch_system_num_rmid_idx(void)
> {
> return (mpam_pmg_max + 1) * (mpam_partid_max + 1);
> }
>
> +u32 resctrl_arch_system_max_rmid_idx(void)
> +{
> + return resctrl_arch_system_num_rmid_idx();
> +}
> +
> u32 resctrl_arch_rmid_idx_encode(u32 closid, u32 rmid)
> {
> return closid * (mpam_pmg_max + 1) + rmid;
> diff --git a/fs/resctrl/monitor.c b/fs/resctrl/monitor.c
> index 2a28fe04284b..5340c764bf7f 100644
> --- a/fs/resctrl/monitor.c
> +++ b/fs/resctrl/monitor.c
> @@ -75,6 +75,11 @@ static unsigned int rmid_limbo_count;
> */
> static struct rmid_entry *rmid_ptrs;
>
> +/*
> + * @max_idx_limit - The number of elements in rmid_ptrs[].
> + */
> +static u32 max_idx_limit;
Please change this patch to avoid local variables shadow this global.
For example, this patch adds a max_idx_limit local variable to
domain_setup_l3_mon_state.
While the name is technically accurate, could this variable perhaps be
more descriptive with a name like: "num_rmid_ptrs" ?
Reinette
next prev parent reply other threads:[~2026-09-10 4:01 UTC|newest]
Thread overview: 52+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 17:43 [PATCH v11 00/23] Allow AET to use PMT as loadable module Tony Luck
2026-08-31 17:43 ` [PATCH v11 01/23] x86/resctrl: Give better names to X86_FEATURE flags for monitoring Tony Luck
2026-09-10 3:48 ` Reinette Chatre
2026-09-11 0:06 ` Luck, Tony
2026-09-11 15:56 ` Reinette Chatre
2026-09-11 18:21 ` Luck, Tony
2026-09-11 22:51 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 02/23] x86/resctrl: Check if monitoring features are enabled Tony Luck
2026-09-10 3:52 ` Reinette Chatre
2026-09-11 19:11 ` Luck, Tony
2026-09-11 23:08 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 03/23] x86/resctrl: Enumerate monitor features in rdt_get_l3_mon_config() Tony Luck
2026-09-10 3:54 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 04/23] x86/resctrl: Apply Intel MBM quirk from rdt_get_l3_mon_config() Tony Luck
2026-09-10 3:55 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 05/23] x86/resctrl: Delete resctrl_cpu_detect() Tony Luck
2026-08-31 17:44 ` [PATCH v11 06/23] arm,x86,fs/resctrl: Replace architecture resctrl_arch_{alloc,mon}_capable() Tony Luck
2026-09-10 3:56 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 07/23] x86/resctrl: Add special case for Intel Haswell enumeration Tony Luck
2026-09-10 3:56 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 08/23] x86/resctrl: Delete rdt_alloc_capable and rdt_mon_capable Tony Luck
2026-09-10 3:57 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 09/23] fs/resctrl: Remove redundant calls to resctrl_mon_capable() Tony Luck
2026-09-10 3:57 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 10/23] x86/resctrl: Honor rdt=perf option to force enable AET perf events Tony Luck
2026-08-31 17:44 ` [PATCH v11 11/23] fs/resctrl: Add interface to disable a monitor event Tony Luck
2026-09-10 3:58 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 12/23] arm,x86,fs/resctrl: Handle change in number of RMIDs on each mount Tony Luck
2026-09-10 4:01 ` Reinette Chatre [this message]
2026-08-31 17:44 ` [PATCH v11 13/23] x86/resctrl: Handle systems when AET is the only resource Tony Luck
2026-09-10 4:04 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 14/23] x86/resctrl: Enforce system RMID limit on AET event groups Tony Luck
2026-09-10 4:05 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 15/23] x86/resctrl: Add PMT registration API for AET enumeration callbacks Tony Luck
2026-08-31 17:44 ` [PATCH v11 16/23] platform/x86/intel/pmt: Register enumeration functions with resctrl Tony Luck
2026-09-01 11:11 ` Ilpo Järvinen
2026-08-31 17:44 ` [PATCH v11 17/23] arm,x86/resctrl: Resolve INTEL_PMT_TELEMETRY symbols at runtime Tony Luck
2026-09-10 4:07 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 18/23] fs/resctrl: Call arch code for every mount Tony Luck
2026-09-10 4:07 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 19/23] x86/resctrl: Export interface to report telemetry unbind/remove Tony Luck
2026-09-10 4:08 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 20/23] platform/x86/intel/pmt: Inform resctrl when MMIO maps are being removed Tony Luck
2026-09-01 11:09 ` Ilpo Järvinen
2026-09-10 4:09 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 21/23] x86/resctrl: Require 64-bit x86 for resctrl support Tony Luck
2026-09-10 4:09 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 22/23] x86/resctrl: Simplify Kconfig options for resctrl Tony Luck
2026-09-10 4:09 ` Reinette Chatre
2026-08-31 17:44 ` [PATCH v11 23/23] x86/resctrl: Document telemetry mount timing caveat Tony Luck
2026-09-10 4:10 ` Reinette Chatre
2026-09-01 19:53 ` [PATCH v11 00/23] Allow AET to use PMT as loadable module Luck, Tony
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=0ef770ec-ac65-4a80-b676-a8bd92caea88@intel.com \
--to=reinette.chatre@intel.com \
--cc=Dave.Martin@arm.com \
--cc=babu.moger@amd.com \
--cc=david.e.box@intel.com \
--cc=dfustini@baylibre.com \
--cc=fenghuay@nvidia.com \
--cc=hch@infradead.org \
--cc=james.morse@arm.com \
--cc=linux-kernel@vger.kernel.org \
--cc=maciej.wieczor-retman@intel.com \
--cc=patches@lists.linux.dev \
--cc=peternewman@google.com \
--cc=tony.luck@intel.com \
--cc=x86@kernel.org \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.