All of lore.kernel.org
 help / color / mirror / Atom feed
From: Waiman Long <llong@redhat.com>
To: Chen Ridong <chenridong@huaweicloud.com>,
	tj@kernel.org, hannes@cmpxchg.org, mkoutny@suse.com,
	peterz@infradead.org
Cc: cgroups@vger.kernel.org, linux-kernel@vger.kernel.org,
	lujialin4@huawei.com, chenridong@huawei.com
Subject: Re: [PATCH v2 -next] cpuset: fix warning when attaching tasks with offline CPUs
Date: Tue, 15 Jul 2025 12:00:43 -0400	[thread overview]
Message-ID: <eec0c87e-b5bf-45b7-9eff-b84d53784678@redhat.com> (raw)
In-Reply-To: <20250715023340.3617147-1-chenridong@huaweicloud.com>

On 7/14/25 10:33 PM, Chen Ridong wrote:
> From: Chen Ridong <chenridong@huawei.com>
>
> A kernel warning was observed in the cpuset migration path:
>
>       WARNING: CPU: 3 PID: 123 at kernel/cgroup/cpuset.c:3130
>       cgroup_migrate_execute+0x8df/0xf30
>       Call Trace:
>        cgroup_transfer_tasks+0x2f3/0x3b0
>        cpuset_migrate_tasks_workfn+0x146/0x3b0
>        process_one_work+0x5ba/0xda0
>        worker_thread+0x788/0x1220
>
> The issue can be reliably reproduced with:
>
>       # Setup test cpuset
>       mkdir /sys/fs/cgroup/cpuset/test
>       echo 2-3 > /sys/fs/cgroup/cpuset/test/cpuset.cpus
>       echo 0 > /sys/fs/cgroup/cpuset/test/cpuset.mems
>
>       # Start test process
>       sleep 100 &
>       pid=$!
>       echo $pid > /sys/fs/cgroup/cpuset/test/cgroup.procs
>       taskset -p 0xC $pid  # Bind to CPUs 2-3
>
>       # Take CPUs offline
>       echo 0 > /sys/devices/system/cpu/cpu3/online
>       echo 0 > /sys/devices/system/cpu/cpu2/online
>
> Root cause analysis:
> When tasks are migrated to top_cpuset due to CPUs going offline,
> cpuset_attach_task() sets the CPU affinity using cpus_attach which
> is initialized from cpu_possible_mask. This mask may include offline
> CPUs. When __set_cpus_allowed_ptr() computes the intersection between:
> 1. cpus_attach (possible CPUs, may include offline)
> 2. p->user_cpus_ptr (original user-set mask)
> The resulting new_mask may contain only offline CPUs, causing the
> operation to fail.
>
> To resolve this issue, if the call to set_cpus_allowed_ptr fails, retry
> using the intersection of cpus_attach and cpu_active_mask.
>
> Fixes: da019032819a ("sched: Enforce user requested affinity")
> Suggested-by: Waiman Long <llong@redhat.com>
> Reported-by: Yang Lijin <yanglijin@huawei.com>
> Signed-off-by: Chen Ridong <chenridong@huawei.com>
> ---
>   kernel/cgroup/cpuset.c | 15 ++++++++++++++-
>   1 file changed, 14 insertions(+), 1 deletion(-)

Thanking further about this problem, the cpuset patch that I proposed is 
just a bandage. It is better to fix the problem at its origin in 
kernel/sched/core.c. I have posted a new patch to do that.

Cheers,
Longman

>
> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
> index f74d04429a29..2cf788a8982a 100644
> --- a/kernel/cgroup/cpuset.c
> +++ b/kernel/cgroup/cpuset.c
> @@ -3114,6 +3114,10 @@ static void cpuset_cancel_attach(struct cgroup_taskset *tset)
>   static cpumask_var_t cpus_attach;
>   static nodemask_t cpuset_attach_nodemask_to;
>   
> +/*
> + * Note that tasks in the top cpuset won't get update to their cpumasks when
> + * a hotplug event happens. So we include offline CPUs as well.
> + */
>   static void cpuset_attach_task(struct cpuset *cs, struct task_struct *task)
>   {
>   	lockdep_assert_held(&cpuset_mutex);
> @@ -3127,7 +3131,16 @@ static void cpuset_attach_task(struct cpuset *cs, struct task_struct *task)
>   	 * can_attach beforehand should guarantee that this doesn't
>   	 * fail.  TODO: have a better way to handle failure here
>   	 */
> -	WARN_ON_ONCE(set_cpus_allowed_ptr(task, cpus_attach));
> +	if (unlikely(set_cpus_allowed_ptr(task, cpus_attach))) {
> +		/*
> +		 * Since offline CPUs are included for top_cpuset,
> +		 * set_cpus_allowed_ptr() can fail if user_cpus_ptr contains
> +		 * only offline CPUs. Take out the offline CPUs and retry.
> +		 */
> +		if (cs == &top_cpuset)
> +			cpumask_and(cpus_attach, cpus_attach, cpu_active_mask);
> +		WARN_ON_ONCE(set_cpus_allowed_ptr(task, cpus_attach));
> +	}
>   
>   	cpuset_change_task_nodemask(task, &cpuset_attach_nodemask_to);
>   	cpuset1_update_task_spread_flags(cs, task);


      reply	other threads:[~2025-07-15 16:00 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-07-15  2:33 [PATCH v2 -next] cpuset: fix warning when attaching tasks with offline CPUs Chen Ridong
2025-07-15 16:00 ` Waiman Long [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=eec0c87e-b5bf-45b7-9eff-b84d53784678@redhat.com \
    --to=llong@redhat.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chenridong@huawei.com \
    --cc=chenridong@huaweicloud.com \
    --cc=hannes@cmpxchg.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lujialin4@huawei.com \
    --cc=mkoutny@suse.com \
    --cc=peterz@infradead.org \
    --cc=tj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.