From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3ABEC305679 for ; Sun, 21 Jun 2026 03:29:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782012559; cv=none; b=uzBXehRcapKqGT0Qup9XtSWsDfY+aKMSDTTKoDxXK2qkvIQwCNmDXxvc0MlP4f8Uyv4p1+FxlhY8Ysu3F+kwzVvcF5NtKQINYLaNYyBMr0Esc3OjLaEcXH9QcDWVNqm7AqANvDn/L8aEyzJLW91uAK/pcOlc9BrPcN1bphLO0Ss= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782012559; c=relaxed/simple; bh=YbHHqXLqguG1H7WJAgdz+3M/zVzPlGDZiCuLwAoAlm8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NcchBK7puxe+j9vBYQ4DMs8jYDczq1nUuvmUcOqbcrVSYXNqzYm6AUduC5VbmOtIamdhi9gPQJRLZ3PTgdtW1tsUHo30cUCmQCNUoHPa9oi2Uu+m3wAvKpiCch+cRkPhpNJUCJFWDKKRgnRG9Bb7FBCQXJFVsQyfUOzmsnLvXlY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=PRuDzNpN; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="PRuDzNpN" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1782012557; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=r0nitzex4gAqxO/5li9BTyWCaz0H8fP6g/UzjMknWgw=; b=PRuDzNpN7N60wySB7CVwGWoZTBtAft5v3aA5QpW+YdZw+3EkmubkvK+SlZtnzYkO8cxll9 2BnZv26M7B3oBhI/MTgvPjGNkDs9HIQky4Gy2ZUSJgC6cWDzapojL7KeWad6mrxMHripmA zrj1bMMqx1VPY7yBMmIuXeVdRumXHqQ= Received: from mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-501-a1F2j42WPIyMUEkXpy85Kw-1; Sat, 20 Jun 2026 23:29:13 -0400 X-MC-Unique: a1F2j42WPIyMUEkXpy85Kw-1 X-Mimecast-MFC-AGG-ID: a1F2j42WPIyMUEkXpy85Kw_1782012552 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-03.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id A95DA1955DE3; Sun, 21 Jun 2026 03:29:11 +0000 (UTC) Received: from llong-thinkpadp16vgen1.westford.csb (unknown [10.22.88.8]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id 8CE37195604D; Sun, 21 Jun 2026 03:29:09 +0000 (UTC) From: Waiman Long To: Ridong Chen , Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Li Zefan , Farhad Alemi , Andrew Morton Cc: cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, Aaron Tomlin , Guopeng Zhang , Gregory Price , David Hildenbrand , Waiman Long Subject: [PATCH v7 6/9] cgroup/cpuset: Make cpuset_attach_old_cs track task group leaders Date: Sat, 20 Jun 2026 23:28:13 -0400 Message-ID: <20260621032816.1806773-7-longman@redhat.com> In-Reply-To: <20260621032816.1806773-1-longman@redhat.com> References: <20260621032816.1806773-1-longman@redhat.com> Precedence: bulk X-Mailing-List: cgroups@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 There are two possible ways that migration of tasks from multiple source cpusets to a target cpuset can happen. Either a multithread application with threads in different cpusets is wholely moved to a new cpuset or disabling of v2 cpuset controller will move all the tasks in child cpusets to the parent cpuset. In the former case, it is the mm setting of the group leader that really matters. So cpuset_attach_old_cs should track the oldcs of the thread leader. In the latter case, effective_mems of child cpusets must always be a subset of the parent. So no real page migration will be necessary no matter which child cpuset is selected as cpuset_attach_old_cs. IOW, cpuset_attach_old_cs should be updated to match the latest task group leader in cpuset_can_attach(), but fall back to that of the first task if there is no group leader in the taskset. Suggested-by: Ridong Chen Reviewed-by: Ridong Chen Signed-off-by: Waiman Long --- kernel/cgroup/cpuset.c | 25 +++++++++++++++++++++++++ 1 file changed, 25 insertions(+) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index b7b5072f2fdd..0375dae26d0b 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -2978,6 +2978,10 @@ static int update_prstate(struct cpuset *cs, int new_prs) return 0; } +/* + * cpuset_can_attach() and cpuset_attach() specific internal data + * Protected by cpuset_mutex + */ static struct cpuset *cpuset_attach_old_cs; /* @@ -3068,11 +3072,32 @@ static int cpuset_can_attach(struct cgroup_taskset *tset) if (ret) goto out_unlock; + /* + * The cpuset_attach_old_cs is used mainly by cpuset_migrate_mm() to get + * the old_mems_allowed value. There are two ways that many-to-one + * cpuset migration can happen: + * 1) A multithread application with threads in different cpusets is + * wholely migrated to a new cpuset. + * 2) Disabling v2 cpuset controller will move all the tasks in child + * cpusets to the parent cpuset. + * + * In the former case, it is the mm setting of the group leader that + * really matters. So cpuset_attach_old_cs should track the oldcs of the + * group leader. It falls back to the oldcs of the first task if there + * is no group leader in the taskset. In the latter case, effective_mems + * of child cpusets must always be a subset of the parent. So no real + * page migration will be necessary no matter which child cpuset is + * selected as cpuset_attach_old_cs. + */ cgroup_taskset_for_each(task, css, tset) { ret = task_can_attach(task); if (ret) goto out_unlock; + /* Update cpuset_attach_old_cs to the latest group leader */ + if (task == task->group_leader) + cpuset_attach_old_cs = task_cs(task); + if (setsched_check) { ret = security_task_setscheduler(task); if (ret) -- 2.54.0