From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qv1-f52.google.com (mail-qv1-f52.google.com [209.85.219.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BFBD341DDE0 for ; Mon, 20 Jul 2026 19:34:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.52 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784576100; cv=none; b=uWEIW/d6sLZyAEJooxpll+RmJV4cZws0W16kO9FWBSK8rCGknrh5BbEKWxD4wDlK0f2/XGZQJZNWg7Z/nDUKSGflNAXvBB9PjHVAnmXuFo9lJCMo7xWwqtCYoX8OtCUEACR0oYORRyCMZWZEMtGKcG9NbjMVdFOcvOvGfDpjnag= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784576100; c=relaxed/simple; bh=YsoKo3cXHgH93neLpy3o6IszKmoV0YxgYYAN+UFPkrY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=qbS7+FgTTQxF/I/zmYr8a0J+yG1Nd2p6+pxtgYISc34FGlm+YBQo1XKh/kApvX5yoQp7fdEddWsBztzRwaiqQh3VWOLQ0lW5B75qUmz3XJ+mep5jCUa5oq4gx0TTnFW24Xnue+ZdhM+zRU+92+bvoNkl+ZvQRPTMfDlVxfhM0TE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=K3xGMZw8; arc=none smtp.client-ip=209.85.219.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="K3xGMZw8" Received: by mail-qv1-f52.google.com with SMTP id 6a1803df08f44-8ee88fce572so82095496d6.1 for ; Mon, 20 Jul 2026 12:34:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1784576094; x=1785180894; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9m4X8mdOzY8Hd25SBKdvyEitcu3reaQ/sG147Spw2gM=; b=K3xGMZw8L+kWA3YlWFQnjsHoCM8sw1LuKZSy6EmLVta9wCzYDt3z4X6mImpOqrxve/ C/dTkKKfxuL14uZnknAwpQnfyaiSV2Lls428W2CTBMj2l9NWEh+EhnW7V3wDp1BdAnxf MOOb/Harjei9fux3cVEUkG++D5e8QehfmEz2N9cwJOy8W/gL/NJr62gGJjJdUivrCgPT VXrF/p4xbsYceK6EboxAoV4HPRZJ+OIoYcXNnTmIOfDCNDS47nxAY5TJfie1qn6T5+pK 1kQGvq91sgGrRBsaIlK82BAXoNbs3bEW+fkwHRZwyJUQBMbGMrQ9VyBK2eGWDxxg9PRB Sb0Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784576094; x=1785180894; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=9m4X8mdOzY8Hd25SBKdvyEitcu3reaQ/sG147Spw2gM=; b=hum2EV+xtqLJTeaHXnIUwLqGs9oIukajwS9jzRUnMdN42r2E2L5fAnMNN0nxKo9A5Z ujFXcwOAyYohDoWFg0CLuPPt4OBDmi5VMF65uurKTzrU1u8KR2VuaijOsPuz4morP/iP DGAxRs8opemoX8ArLVy+oBPFPQQH0RFcacSITHOgWDb3Exyo901h4AOzkRDN0n5piAOU 48izENIb3MkCiYKJu6NYv0HmX2SfSbuFPQBP/8RHzGFjgKBmqc4EUXoOiZ5175AZKWl4 Ft903dcu1e9hZV8rXr9zreTZ4RslQRlEGpL33lQhA/rOuhWp5bZJqvl04h4wUcajVCbO bBLg== X-Forwarded-Encrypted: i=1; AHgh+RpbnRWIhaH/tqwAx4/cwMkfP1FbQHQ05Aa21E260ruKNVnmK0sLlnpgW4U0CAouUuLrlsI=@vger.kernel.org X-Gm-Message-State: AOJu0YxYufQPW0FfbLLFOP5Grj5+ypb3pNQmo7TYj+QLVU6EqESZ5hiT fCpSifGeBwtpJNX4fvkQvkoL7hT+Qd8sdNFmiK7wjvIBYc6/ybX1mEY7L5UbruIVOkg= X-Gm-Gg: AfdE7cmW63x651sntoXAY/tBU9kUJPTEVHlvB6BeUYzudWNpq8HeTcxAAlHsPBb+Zhq gTvJn3FYR9oEQRdb+pA0bz/8rehthym+zok/1Y6CMm4Urv6IvavoJwzx4iVylZE1K8tHJasqoAq 4x1xgxYTbX2NxfLh/bil0Kq1rZTzhPX3ezH9zLG1lUNfyWzo5hS5d+6wVkbXuNj0IzPYuEyAE0V MEioFqrs5atTX5yX+V9stZmc6vNdK5vsl73Fvfj/oxQWCQQv3RdEhEgObSutSstLmsD+EFFtaA/ RL8OhmUCwxxw+xC6MFHLxvEO5zQ7ps75ksWUBN85KNlmURVMmLHRvagQ5YNsuJiM43go/N1WeDZ EsOVADV9mtejXR6c15DOta0x18jIGjHcZPObva0PRH2tZCEyF5nSbQpXzit6aZxAF6XpS/km3+3 JBAnN3EBMczS3dvWxkf8SDvpl48rQsrnto+ABfR5W+x9u4yT0XFlD8SXbveDeDlb8= X-Received: by 2002:a05:620a:2b9c:b0:915:83f3:780d with SMTP id af79cd13be357-930b3ec55e3mr1402108585a.32.1784576094094; Mon, 20 Jul 2026 12:34:54 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-930b545e47bsm957792285a.35.2026.07.20.12.34.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 20 Jul 2026 12:34:53 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: Zhigang.Luo@amd.com, arun.george@samsung.com, balbirs@nvidia.com, brendan.jackman@linux.dev, yuzenghui@huawei.com, apopple@nvidia.com, alucerop@amd.com, matthew.brost@intel.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, corbet@lwn.net, skhan@linuxfoundation.org, gregkh@linuxfoundation.org, rafael@kernel.org, dakr@kernel.org, djbw@kernel.org, vishal.l.verma@intel.com, dave.jiang@intel.com, alison.schofield@intel.com, osandov@osandov.com, jannh@google.com, pfalcato@suse.de, jackmanb@google.com, hannes@cmpxchg.org, ziy@nvidia.com, pbonzini@redhat.com, osalvador@suse.de, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev, baohua@kernel.org, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, yury.norov@gmail.com, linux@rasmusvillemoes.dk, longman@redhat.com, ridong.chen@linux.dev, tj@kernel.org, mkoutny@suse.com, sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com, peterx@redhat.com, baolin.wang@linux.alibaba.com, npache@redhat.com, ryan.roberts@arm.com, dev.jain@arm.com, lance.yang@linux.dev, usama.arif@linux.dev, xu.xin16@zte.com.cn, chengming.zhou@linux.dev, roman.gushchin@linux.dev, muchun.song@linux.dev, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, driver-core@lists.linux.dev, nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org, linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org, kvm@vger.kernel.org, cgroups@vger.kernel.org, damon@lists.linux.dev, linux-kselftest@vger.kernel.org, kernel-team@meta.com Subject: [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Date: Mon, 20 Jul 2026 15:34:00 -0400 Message-ID: <20260720193431.3841992-7-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260720193431.3841992-1-gourry@gourry.net> References: <20260720193431.3841992-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: kvm@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit N_MEMORY_PRIVATE node access is gated by per-node capabilities and the zonelist. Isolating by cpuset offers no functionality and creates issues on rebind when a cpuset loses access to that node (in particular: it generates migrations that may not be supported). Do not partition private nodes via cpuset.mem, instead treat them as globally accessible resources. Do not engage in silent background operations on these nodes due to cpuset.mem membership changing. The result: cpuset operations checking for node validity always allow N_MEMORY_PRIVATE nodes. On hot-unplug, we still need to rebind memory policies if the private node has left N_MEMORY_PRIVATE (i.e. no more memory). Signed-off-by: Gregory Price --- kernel/cgroup/cpuset.c | 26 ++++++++++++++++++++++---- mm/memcontrol.c | 2 ++ 2 files changed, 24 insertions(+), 4 deletions(-) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index dfd0f827e3b92..05468f95c10bd 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -26,6 +26,7 @@ #include #include #include +#include #include #include #include @@ -3867,7 +3868,8 @@ static void cpuset_handle_hotplug(void) static DECLARE_WORK(hk_sd_work, hk_sd_workfn); static cpumask_t new_cpus; static nodemask_t new_mems; - bool cpus_updated, mems_updated; + static nodemask_t prev_priv_mems; + bool cpus_updated, mems_updated, priv_shrank; bool on_dfl = is_in_v2_mode(); struct tmpmasks tmp, *ptmp = NULL; @@ -3890,6 +3892,14 @@ static void cpuset_handle_hotplug(void) !cpumask_empty(subpartitions_cpus); mems_updated = !nodes_equal(top_cpuset.effective_mems, new_mems); + /* + * Private nodes are not partitioned by cpuset, but if one leaves + * N_MEMORY_PRIVATE we still need to run mpol_rebind_* to clean up + * mempolicies that are binding them. + */ + priv_shrank = !nodes_subset(prev_priv_mems, node_states[N_MEMORY_PRIVATE]); + prev_priv_mems = node_states[N_MEMORY_PRIVATE]; + /* For v1, synchronize cpus_allowed to cpu_active_mask */ if (cpus_updated) { cpuset_force_rebuild(); @@ -3922,9 +3932,12 @@ static void cpuset_handle_hotplug(void) top_cpuset.mems_allowed = new_mems; top_cpuset.effective_mems = new_mems; spin_unlock_irq(&callback_lock); - cpuset_update_tasks_nodemask(&top_cpuset); } + /* Rebind task mempolicies if any memory node changed state */ + if (mems_updated || priv_shrank) + cpuset_update_tasks_nodemask(&top_cpuset); + mutex_unlock(&cpuset_mutex); /* if cpus or mems changed, we need to propagate to descendants */ @@ -4155,11 +4168,13 @@ nodemask_t cpuset_mems_allowed(struct task_struct *tsk) * cpuset_nodemask_valid_mems_allowed - check nodemask vs. current mems_allowed * @nodemask: the nodemask to be checked * - * Are any of the nodes in the nodemask allowed in current->mems_allowed? + * Are any of the nodes in the nodemask usable? N_MEMORY nodes must be in + * current->mems_allowed, while N_MEMORY_PRIVATE nodes are always valid. */ int cpuset_nodemask_valid_mems_allowed(const nodemask_t *nodemask) { - return nodes_intersects(*nodemask, current->mems_allowed); + return nodes_intersects(*nodemask, current->mems_allowed) || + nodes_intersects(*nodemask, node_states[N_MEMORY_PRIVATE]); } /* @@ -4223,6 +4238,9 @@ bool cpuset_current_node_allowed(int node, gfp_t gfp_mask) if (in_interrupt()) return true; + /* N_MEMORY_PRIVATE nodes are not partitioned by cpusets.mems */ + if (node_is_private(node)) + return true; if (node_isset(node, current->mems_allowed)) return true; /* diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 8319ad8c5c23a..f0dde52dc9e0e 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -6069,6 +6069,8 @@ void mem_cgroup_node_filter_allowed(struct mem_cgroup *memcg, nodemask_t *mask) * mask is acceptable. */ cpuset_nodes_allowed(memcg->css.cgroup, &allowed); + /* N_MEMORY_PRIVATE nodes are not partitioned by cpuset, include them */ + nodes_or(allowed, allowed, node_states[N_MEMORY_PRIVATE]); nodes_and(*mask, *mask, allowed); } -- 2.53.0-Meta