From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qk1-f177.google.com (mail-qk1-f177.google.com [209.85.222.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B511541A57E for ; Mon, 20 Jul 2026 19:34:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.177 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784576099; cv=none; b=XhF5DLCOxjY5FP9hGjBsflZLoHiBL+nwXL8A9ujqaI0NsGi1VXRkQ+HLTpg2RqMEYeE6ri5khGr6e6MjakF1qtDy9322czidoMgEchdgNMAoI5/xZ2cpIyk/d9nbFKgpyuPzt4Wl8N1uWrZcMib+nxCgpRD5GdGldMHB7T8LdKE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784576099; c=relaxed/simple; bh=YsoKo3cXHgH93neLpy3o6IszKmoV0YxgYYAN+UFPkrY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=EXYuBJPJT3dTSR1CDE54TklgBlnWLpkccd3O3GKya1vA7zEuozhB4FR5ITDRJgU//unzqT/wErqteeyJP7TR9TmPcGy+LNZMFEZ94eP5n7Rdckz9i2N7+wHZhoUXavH60uNxeGLxNSqRL1QT/XkYyOI769HOIKo2ljbLB/OW43k= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=IGKWV6gA; arc=none smtp.client-ip=209.85.222.177 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="IGKWV6gA" Received: by mail-qk1-f177.google.com with SMTP id af79cd13be357-92e85499ffbso577722685a.0 for ; Mon, 20 Jul 2026 12:34:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1784576094; x=1785180894; darn=lists.linux.dev; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9m4X8mdOzY8Hd25SBKdvyEitcu3reaQ/sG147Spw2gM=; b=IGKWV6gAIsiYkVYoXYXff9MsHJq5NXUEFEx+HgBjXwy4KQ3gdcH8j4VppYd/wdH7bF rwrNjooHJ5ljBNvHMjDx4iZaCbD5sjdS3Zp3YJC5yTaUrHvBPZhEAZ61YcPgjDkFvjud +OQiMgJeRE9Apmr7YAPwT2LdATeBoS6WioXBY0Iyy/66aaD3nAFMEmYX9VeDJy9jC46d 6vbKYp7+OjNNKF8WmHYiptPtWTkK+M/kKTcJXHPRAxygbSmofMEe6e5xjX6QlZUBYoM5 hVQeE/4MyinF6TFIVSsFn69XgG1co+S/PhHKM+vUghsqJX0ZTfYISmNVnfHNtzSDMGR7 towA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784576094; x=1785180894; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=9m4X8mdOzY8Hd25SBKdvyEitcu3reaQ/sG147Spw2gM=; b=Iy4kmAsyF5xRIKtdMxGCkMOZfXSkQ8LvESQyUqRelHEBAv1OPHNkxAtQSTi1XvkAmS VXTd6Ua3sRNXA1kKLLZ3s401n+L3TTozXVhG91TSK46gSRfuncMRHOfMVO963b0ZaJHX DIqysyUPEtrZ37HYhag7FzHr16bXzcxLR6PlmO+YOC8lmMiAHLRAvL3goHKDH3AbmVoR 90iX94zlqPvhMKDER1jp0l+ac3n5GB6kXpjKNN+e1Hx6dzXPh/Hyq4+OoX9j2LiA2p+1 sxN3zo/6OF+kUzFNDc3a+gYJv4+4YIrdJvGoZO2KayWjVqZb8R04M0sp+8qNmorc65rx a+fg== X-Forwarded-Encrypted: i=1; AHgh+RrymITiD+KS4gVcuQiRJKLP9qRZuUGUS/4Md0NmAUdXI3PDj8JKrWtjKGUgz9+5gZPuHh7baVX6K3VVRQ==@lists.linux.dev X-Gm-Message-State: AOJu0YyTjMbs9GBGkkgdg/g5t25cUxCPdTIDORcDM2uGPNAy1Fl9dZsS iHNLH0iJndCnuKj9O5HwBHil79apCDMntqAVm3AE5ccmPZuWCUmuVy99AFaP9bat6ew= X-Gm-Gg: AfdE7cml8z2fgyig4NEXPQ77WREHpAJ3s1rnh2VlK3b0FNqCkrLbh+ElSi1f82gnGqy YxUvFotC/Wu0FkinqI559l/vmTudEAIe8n4uiKZgifhXMCeGx78OBwtDr3R6q2ChUcZmtQ/RKF5 Bp9IoLn/mSSYJ+kHID501PfLEYv3cvgRaFhTCar5MhI3Ui2HUmg6uF0vy4gWg7vltpqbE/9Jnpp 1mPnNLm6Zu3Z+nfQGxfutOApDYptGnjqRWEJrd5Z6CrY/AWklzc58Ax0a3t/dwai0m+KyJt4vMj 52BqgdxXW/y+GB5z0vvsU40nh/eVLAKO9HrV5gXWRWJZQXFBGA09hHxxro8DtBEUH/v/OzwRF1I L1xbWkmyiZa2cMY8j5zJt9ucWlZaNThBbTdTM/NeEXvwt0HFIhd/Y/WZQrOowxxVHjp7+QZtWMI wAqFsmR2zg4bUNiK1Jsz7xAZJCAku0kst8JZjyx6+sse/A5ONJoVH1Du+Lk7Pt7yM= X-Received: by 2002:a05:620a:2b9c:b0:915:83f3:780d with SMTP id af79cd13be357-930b3ec55e3mr1402108585a.32.1784576094094; Mon, 20 Jul 2026 12:34:54 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-930b545e47bsm957792285a.35.2026.07.20.12.34.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 20 Jul 2026 12:34:53 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: Zhigang.Luo@amd.com, arun.george@samsung.com, balbirs@nvidia.com, brendan.jackman@linux.dev, yuzenghui@huawei.com, apopple@nvidia.com, alucerop@amd.com, matthew.brost@intel.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, corbet@lwn.net, skhan@linuxfoundation.org, gregkh@linuxfoundation.org, rafael@kernel.org, dakr@kernel.org, djbw@kernel.org, vishal.l.verma@intel.com, dave.jiang@intel.com, alison.schofield@intel.com, osandov@osandov.com, jannh@google.com, pfalcato@suse.de, jackmanb@google.com, hannes@cmpxchg.org, ziy@nvidia.com, pbonzini@redhat.com, osalvador@suse.de, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev, baohua@kernel.org, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, yury.norov@gmail.com, linux@rasmusvillemoes.dk, longman@redhat.com, ridong.chen@linux.dev, tj@kernel.org, mkoutny@suse.com, sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com, peterx@redhat.com, baolin.wang@linux.alibaba.com, npache@redhat.com, ryan.roberts@arm.com, dev.jain@arm.com, lance.yang@linux.dev, usama.arif@linux.dev, xu.xin16@zte.com.cn, chengming.zhou@linux.dev, roman.gushchin@linux.dev, muchun.song@linux.dev, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, driver-core@lists.linux.dev, nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org, linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org, kvm@vger.kernel.org, cgroups@vger.kernel.org, damon@lists.linux.dev, linux-kselftest@vger.kernel.org, kernel-team@meta.com Subject: [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Date: Mon, 20 Jul 2026 15:34:00 -0400 Message-ID: <20260720193431.3841992-7-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260720193431.3841992-1-gourry@gourry.net> References: <20260720193431.3841992-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: driver-core@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit N_MEMORY_PRIVATE node access is gated by per-node capabilities and the zonelist. Isolating by cpuset offers no functionality and creates issues on rebind when a cpuset loses access to that node (in particular: it generates migrations that may not be supported). Do not partition private nodes via cpuset.mem, instead treat them as globally accessible resources. Do not engage in silent background operations on these nodes due to cpuset.mem membership changing. The result: cpuset operations checking for node validity always allow N_MEMORY_PRIVATE nodes. On hot-unplug, we still need to rebind memory policies if the private node has left N_MEMORY_PRIVATE (i.e. no more memory). Signed-off-by: Gregory Price --- kernel/cgroup/cpuset.c | 26 ++++++++++++++++++++++---- mm/memcontrol.c | 2 ++ 2 files changed, 24 insertions(+), 4 deletions(-) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index dfd0f827e3b92..05468f95c10bd 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -26,6 +26,7 @@ #include #include #include +#include #include #include #include @@ -3867,7 +3868,8 @@ static void cpuset_handle_hotplug(void) static DECLARE_WORK(hk_sd_work, hk_sd_workfn); static cpumask_t new_cpus; static nodemask_t new_mems; - bool cpus_updated, mems_updated; + static nodemask_t prev_priv_mems; + bool cpus_updated, mems_updated, priv_shrank; bool on_dfl = is_in_v2_mode(); struct tmpmasks tmp, *ptmp = NULL; @@ -3890,6 +3892,14 @@ static void cpuset_handle_hotplug(void) !cpumask_empty(subpartitions_cpus); mems_updated = !nodes_equal(top_cpuset.effective_mems, new_mems); + /* + * Private nodes are not partitioned by cpuset, but if one leaves + * N_MEMORY_PRIVATE we still need to run mpol_rebind_* to clean up + * mempolicies that are binding them. + */ + priv_shrank = !nodes_subset(prev_priv_mems, node_states[N_MEMORY_PRIVATE]); + prev_priv_mems = node_states[N_MEMORY_PRIVATE]; + /* For v1, synchronize cpus_allowed to cpu_active_mask */ if (cpus_updated) { cpuset_force_rebuild(); @@ -3922,9 +3932,12 @@ static void cpuset_handle_hotplug(void) top_cpuset.mems_allowed = new_mems; top_cpuset.effective_mems = new_mems; spin_unlock_irq(&callback_lock); - cpuset_update_tasks_nodemask(&top_cpuset); } + /* Rebind task mempolicies if any memory node changed state */ + if (mems_updated || priv_shrank) + cpuset_update_tasks_nodemask(&top_cpuset); + mutex_unlock(&cpuset_mutex); /* if cpus or mems changed, we need to propagate to descendants */ @@ -4155,11 +4168,13 @@ nodemask_t cpuset_mems_allowed(struct task_struct *tsk) * cpuset_nodemask_valid_mems_allowed - check nodemask vs. current mems_allowed * @nodemask: the nodemask to be checked * - * Are any of the nodes in the nodemask allowed in current->mems_allowed? + * Are any of the nodes in the nodemask usable? N_MEMORY nodes must be in + * current->mems_allowed, while N_MEMORY_PRIVATE nodes are always valid. */ int cpuset_nodemask_valid_mems_allowed(const nodemask_t *nodemask) { - return nodes_intersects(*nodemask, current->mems_allowed); + return nodes_intersects(*nodemask, current->mems_allowed) || + nodes_intersects(*nodemask, node_states[N_MEMORY_PRIVATE]); } /* @@ -4223,6 +4238,9 @@ bool cpuset_current_node_allowed(int node, gfp_t gfp_mask) if (in_interrupt()) return true; + /* N_MEMORY_PRIVATE nodes are not partitioned by cpusets.mems */ + if (node_is_private(node)) + return true; if (node_isset(node, current->mems_allowed)) return true; /* diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 8319ad8c5c23a..f0dde52dc9e0e 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -6069,6 +6069,8 @@ void mem_cgroup_node_filter_allowed(struct mem_cgroup *memcg, nodemask_t *mask) * mask is acceptable. */ cpuset_nodes_allowed(memcg->css.cgroup, &allowed); + /* N_MEMORY_PRIVATE nodes are not partitioned by cpuset, include them */ + nodes_or(allowed, allowed, node_states[N_MEMORY_PRIVATE]); nodes_and(*mask, *mask, allowed); } -- 2.53.0-Meta