From: Tony Luck <tony.luck@intel.com>
To: Reinette Chatre <reinette.chatre@intel.com>
Cc: Thomas Gleixner <tglx@linutronix.de>,
Borislav Petkov <bp@alien8.de>, James Morse <james.morse@arm.com>,
"x86@kernel.org" <x86@kernel.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"Yu, Fenghua" <fenghua.yu@intel.com>,
Ingo Molnar <mingo@redhat.com>, H Peter Anvin <hpa@zytor.com>,
Babu Moger <Babu.Moger@amd.com>,
"shameerali.kolothum.thodi@huawei.com"
<shameerali.kolothum.thodi@huawei.com>,
D Scott Phillips OS <scott@os.amperecomputing.com>,
"carl@os.amperecomputing.com" <carl@os.amperecomputing.com>,
"lcherian@marvell.com" <lcherian@marvell.com>,
"bobo.shaobowang@huawei.com" <bobo.shaobowang@huawei.com>,
"tan.shaopeng@fujitsu.com" <tan.shaopeng@fujitsu.com>,
"baolin.wang@linux.alibaba.com" <baolin.wang@linux.alibaba.com>,
Jamie Iles <quic_jiles@quicinc.com>,
Xin Hao <xhao@linux.alibaba.com>,
"peternewman@google.com" <peternewman@google.com>,
"dfustini@baylibre.com" <dfustini@baylibre.com>,
"amitsinght@marvell.com" <amitsinght@marvell.com>,
David Hildenbrand <david@redhat.com>
Subject: Re: [PATCH v2] x86/resctrl: Fix WARN in get_domain_from_cpu()
Date: Wed, 21 Feb 2024 15:56:58 -0800 [thread overview]
Message-ID: <ZdaNymyCFAXMn18k@agluck-desk3> (raw)
In-Reply-To: <89745b84-5ead-4694-b14c-341ca5a688f4@intel.com>
On Wed, Feb 21, 2024 at 02:59:43PM -0800, Reinette Chatre wrote:
> Hi Tony,
>
> On 2/21/2024 11:31 AM, Tony Luck wrote:
> > reset_all_ctrls() and resctrl_arch_update_domains() use on_each_cpu_mask()
> > to call rdt_ctrl_update() on potentially one CPU from each domain.
> >
> > But this means rdt_ctrl_update() needs to figure out which domain to
> > apply changes to. Doing so requires a search of all domains in a resource,
> > which can only be done safely if cpus_lock is held. Both callers do hold
> > this lock, but there isn't a way for a function called on another CPU
> > via IPI to verify this.
> >
> > Fix by adding the target domain to the msr_param structure and passing an
> > array with CDP_NUM_TYPES entries. Then calling for each domain separately
> > using smp_call_function_single()
>
> This work contains no changes to get_domain_from_cpu(). I expect the WARN
> within it to be removed as intended with [1] and then this work can build
> on that without urgency. As I understand, to support the stated goal of this
> work, I expect get_domain_from_cpu() in the end to not have any WARN or
> IS_ENABLED checks, but just a lockdep_assert_cpus_held().
>
> Do you have different expectations?
Same expectations. Boris should apply the simple fix (delete the WARN
that is giving a false positive) for this current cycle.
If there is support for my patch (with changes/fixes you point out
below), then it could be added in the future and get_domain_from_cpu()
can use lockdep_assert_cpus_held().
> > Change the low level cat_wrmsr(), mba_wrmsr_intel(), and mba_wrmsr_amd()
> > functions to just take a msr_param structure since it contains the
> > rdt_resource and rdt_domain information.
>
> Could moving the rdt_domain into msr_param be done in a separate patch?
I can break it into more pieces if there is enthusiam to apply it.
> ...
>
> > diff --git a/arch/x86/kernel/cpu/resctrl/ctrlmondata.c b/arch/x86/kernel/cpu/resctrl/ctrlmondata.c
> > index 7997b47743a2..09f6e624f1bb 100644
> > --- a/arch/x86/kernel/cpu/resctrl/ctrlmondata.c
> > +++ b/arch/x86/kernel/cpu/resctrl/ctrlmondata.c
> > @@ -272,22 +272,6 @@ static u32 get_config_index(u32 closid, enum resctrl_conf_type type)
> > }
> > }
> >
> > -static bool apply_config(struct rdt_hw_domain *hw_dom,
> > - struct resctrl_staged_config *cfg, u32 idx,
> > - cpumask_var_t cpu_mask)
> > -{
> > - struct rdt_domain *dom = &hw_dom->d_resctrl;
> > -
> > - if (cfg->new_ctrl != hw_dom->ctrl_val[idx]) {
> > - cpumask_set_cpu(cpumask_any(&dom->cpu_mask), cpu_mask);
> > - hw_dom->ctrl_val[idx] = cfg->new_ctrl;
> > -
> > - return true;
> > - }
> > -
> > - return false;
> > -}
> > -
> > int resctrl_arch_update_one(struct rdt_resource *r, struct rdt_domain *d,
> > u32 closid, enum resctrl_conf_type t, u32 cfg_val)
> > {
> > @@ -304,59 +288,50 @@ int resctrl_arch_update_one(struct rdt_resource *r, struct rdt_domain *d,
> > msr_param.res = r;
> > msr_param.low = idx;
> > msr_param.high = idx + 1;
> > - hw_res->msr_update(d, &msr_param, r);
> > + hw_res->msr_update(&msr_param);
> >
>
> Is this missing setting the domain in msr_param?
Indeed yes.
> > return 0;
> > }
> >
> > int resctrl_arch_update_domains(struct rdt_resource *r, u32 closid)
> > {
> > + struct msr_param msr_param[CDP_NUM_TYPES];
> > struct resctrl_staged_config *cfg;
> > struct rdt_hw_domain *hw_dom;
> > - struct msr_param msr_param;
> > enum resctrl_conf_type t;
> > - cpumask_var_t cpu_mask;
> > struct rdt_domain *d;
> > + bool need_update;
> > + int cpu;
> > u32 idx;
> >
> > /* Walking r->domains, ensure it can't race with cpuhp */
> > lockdep_assert_cpus_held();
> >
> > - if (!zalloc_cpumask_var(&cpu_mask, GFP_KERNEL))
> > - return -ENOMEM;
> > -
> > - msr_param.res = NULL;
> > + memset(msr_param, 0, sizeof(msr_param));
> > list_for_each_entry(d, &r->domains, list) {
> > hw_dom = resctrl_to_arch_dom(d);
> > + need_update = false;
> > for (t = 0; t < CDP_NUM_TYPES; t++) {
> > cfg = &hw_dom->d_resctrl.staged_config[t];
> > if (!cfg->have_new_ctrl)
> > continue;
> >
> > idx = get_config_index(closid, t);
> > - if (!apply_config(hw_dom, cfg, idx, cpu_mask))
> > + if (cfg->new_ctrl == hw_dom->ctrl_val[idx])
> > continue;
> > -
> > - if (!msr_param.res) {
> > - msr_param.low = idx;
> > - msr_param.high = msr_param.low + 1;
> > - msr_param.res = r;
> > - } else {
> > - msr_param.low = min(msr_param.low, idx);
> > - msr_param.high = max(msr_param.high, idx + 1);
> > - }
> > + hw_dom->ctrl_val[idx] = cfg->new_ctrl;
> > + cpu = cpumask_any(&d->cpu_mask);
> > +
> > + msr_param[t].low = idx;
> > + msr_param[t].high = msr_param[t].low + 1;
> > + msr_param[t].res = r;
> > + msr_param[t].dom = d;
> > + need_update = true;
> > }
> > + if (need_update)
> > + smp_call_function_single(cpu, rdt_ctrl_update, &msr_param, 1);
>
> It is not clear to me why it is needed to pass this additional data. Why not
> just use the original mechanism of letting the low and high of msr_param span the
> multiple indices that need updating? There can still be a "need_update" but it
> can be set when msr_param gets its first data. Any other index needing updating
> can just update low/high and a single msr_param can be used.
For some reason this morning I thought that the domain needed to be
different. It isn't, so keeping the code that just adjusts the range
of MSRs will work just fine.
The "need_update" variable isn't required. I will move the
msr_param.res = NULL;
inside the for each domain loop, and can use non-NULL value to
decide whether to IPI a CPU.
-Tony
next prev parent reply other threads:[~2024-02-21 23:57 UTC|newest]
Thread overview: 104+ messages / expand[flat|nested] mbox.gz Atom feed top
2024-02-13 18:44 [PATCH v9 00/24] x86/resctrl: monitored closid+rmid together, separate arch/fs locking James Morse
2024-02-13 18:44 ` [PATCH v9 01/24] tick/nohz: Move tick_nohz_full_mask declaration outside the #ifdef James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 02/24] x86/resctrl: kfree() rmid_ptrs from resctrl_exit() James Morse
2024-02-13 23:14 ` Reinette Chatre
2024-02-19 16:53 ` James Morse
2024-02-19 18:37 ` [tip: x86/cache] x86/resctrl: Free " tip-bot2 for James Morse
2024-02-20 15:27 ` [PATCH v9 02/24] x86/resctrl: kfree() " David Hildenbrand
2024-02-20 15:46 ` James Morse
2024-02-20 15:54 ` Thomas Gleixner
2024-02-20 16:01 ` James Morse
2024-02-20 16:12 ` Thomas Gleixner
2024-02-13 18:44 ` [PATCH v9 03/24] x86/resctrl: Create helper for RMID allocation and mondata dir creation James Morse
2024-02-19 15:49 ` David Hildenbrand
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 04/24] x86/resctrl: Move rmid allocation out of mkdir_rdt_prepare() James Morse
2024-02-19 18:37 ` [tip: x86/cache] x86/resctrl: Move RMID " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 05/24] x86/resctrl: Track the closid with the rmid James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 06/24] x86/resctrl: Access per-rmid structures by index James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 07/24] x86/resctrl: Allow RMID allocation to be scoped by CLOSID James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 08/24] x86/resctrl: Track the number of dirty RMID a CLOSID has James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 09/24] x86/resctrl: Use __set_bit()/__clear_bit() instead of open coding James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-20 16:00 ` [PATCH v9 09/24] " David Hildenbrand
2024-02-20 16:27 ` James Morse
2024-02-20 16:44 ` David Hildenbrand
2024-02-13 18:44 ` [PATCH v9 10/24] x86/resctrl: Allocate the cleanest CLOSID by searching closid_num_dirty_rmid James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 11/24] x86/resctrl: Move CLOSID/RMID matching and setting to use helpers James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 12/24] x86/resctrl: Add cpumask_any_housekeeping() for limbo/overflow James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 13/24] x86/resctrl: Queue mon_event_read() instead of sending an IPI James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 14/24] x86/resctrl: Allow resctrl_arch_rmid_read() to sleep James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 15/24] x86/resctrl: Allow arch to allocate memory needed in resctrl_arch_rmid_read() James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 16/24] x86/resctrl: Make resctrl_mounted checks explicit James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 17/24] x86/resctrl: Move alloc/mon static keys into helpers James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 18/24] x86/resctrl: Make rdt_enable_key the arch's decision to switch James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 19/24] x86/resctrl: Add helpers for system wide mon/alloc capable James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 20/24] x86/resctrl: Add CPU online callback for resctrl work James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 21/24] x86/resctrl: Allow overflow/limbo handlers to be scheduled on any-but cpu James Morse
2024-02-19 18:37 ` [tip: x86/cache] x86/resctrl: Allow overflow/limbo handlers to be scheduled on any-but CPU tip-bot2 for James Morse
2024-03-25 23:14 ` [PATCH v9 21/24] x86/resctrl: Allow overflow/limbo handlers to be scheduled on any-but cpu Tony Luck
2024-02-13 18:44 ` [PATCH v9 22/24] x86/resctrl: Add CPU offline callback for resctrl work James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 23/24] x86/resctrl: Move domain helper migration into resctrl_offline_cpu() James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 18:44 ` [PATCH v9 24/24] x86/resctrl: Separate arch and fs resctrl locks James Morse
2024-02-19 18:37 ` [tip: x86/cache] " tip-bot2 for James Morse
2024-02-13 23:26 ` [PATCH v9 00/24] x86/resctrl: monitored closid+rmid together, separate arch/fs locking Reinette Chatre
2024-02-14 15:01 ` Moger, Babu
2024-02-19 16:53 ` James Morse
2024-02-17 0:28 ` Tony Luck
2024-02-20 18:18 ` Reinette Chatre
2024-02-20 18:48 ` Luck, Tony
2024-02-17 10:55 ` Borislav Petkov
2024-02-19 16:49 ` Thomas Gleixner
2024-02-19 16:53 ` James Morse
2024-02-19 17:51 ` Borislav Petkov
2024-02-19 17:55 ` James Morse
2024-02-20 20:59 ` Tony Luck
2024-02-20 22:58 ` Reinette Chatre
2024-02-20 23:25 ` Luck, Tony
2024-02-21 0:34 ` [PATCH] x86/resctrl: Fix WARN in get_domain_from_cpu() Tony Luck
2024-02-21 5:10 ` Reinette Chatre
2024-02-21 12:06 ` James Morse
2024-02-21 19:31 ` [PATCH v2] " Tony Luck
2024-02-21 22:59 ` Reinette Chatre
2024-02-21 23:56 ` Tony Luck [this message]
2024-02-22 18:50 ` [PATCH v3 0/2] x86/resctrl: Pass domain to target CPU Tony Luck
2024-02-22 18:50 ` [PATCH v3 1/2] " Tony Luck
2024-02-27 22:05 ` Reinette Chatre
2024-02-22 18:50 ` [PATCH v3 2/2] x86/resctrl: Simply call convention for MSR update functions Tony Luck
2024-02-27 22:05 ` Reinette Chatre
2024-02-22 23:18 ` [PATCH v3 0/2] x86/resctrl: Pass domain to target CPU Reinette Chatre
2024-02-22 23:26 ` Luck, Tony
2024-03-08 21:38 ` [PATCH v5 " Tony Luck
2024-03-08 21:38 ` [PATCH v5 1/2] " Tony Luck
2024-03-12 20:06 ` Moger, Babu
2024-03-12 20:08 ` Luck, Tony
2024-04-24 12:01 ` [tip: x86/cache] " tip-bot2 for Tony Luck
2024-03-08 21:38 ` [PATCH v5 2/2] x86/resctrl: Simplify call convention for MSR update functions Tony Luck
2024-03-12 20:06 ` Moger, Babu
2024-04-24 12:01 ` [tip: x86/cache] " tip-bot2 for Tony Luck
2024-03-08 23:08 ` [PATCH v5 0/2] x86/resctrl: Pass domain to target CPU Reinette Chatre
2024-03-08 23:26 ` Luck, Tony
2024-03-29 4:55 ` Reinette Chatre
2024-02-21 12:06 ` [PATCH v9 00/24] x86/resctrl: monitored closid+rmid together, separate arch/fs locking James Morse
2024-02-21 16:48 ` Reinette Chatre
2024-02-21 17:30 ` Luck, Tony
2024-02-22 15:26 ` [tip: x86/cache] x86/resctrl: Remove lockdep annotation that triggers false positive tip-bot2 for James Morse
2024-02-19 16:54 ` [PATCH v9 00/24] x86/resctrl: monitored closid+rmid together, separate arch/fs locking James Morse
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ZdaNymyCFAXMn18k@agluck-desk3 \
--to=tony.luck@intel.com \
--cc=Babu.Moger@amd.com \
--cc=amitsinght@marvell.com \
--cc=baolin.wang@linux.alibaba.com \
--cc=bobo.shaobowang@huawei.com \
--cc=bp@alien8.de \
--cc=carl@os.amperecomputing.com \
--cc=david@redhat.com \
--cc=dfustini@baylibre.com \
--cc=fenghua.yu@intel.com \
--cc=hpa@zytor.com \
--cc=james.morse@arm.com \
--cc=lcherian@marvell.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=peternewman@google.com \
--cc=quic_jiles@quicinc.com \
--cc=reinette.chatre@intel.com \
--cc=scott@os.amperecomputing.com \
--cc=shameerali.kolothum.thodi@huawei.com \
--cc=tan.shaopeng@fujitsu.com \
--cc=tglx@linutronix.de \
--cc=x86@kernel.org \
--cc=xhao@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox