From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D26CF4D179C for ; Fri, 2 Oct 2026 16:54:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790960093; cv=none; b=hoJSI0k6EbXxXRfnNHL6eZH04YtWLFba4x3ERbWY6V2zgtHaUpDZ2OMOjmwNQ24Gku+1X1hLaD++gRxqnprVIl2nALYMd7nP9TSmSCh87GzFVxKTv7EIokdlb24ewn+9h0zNZsYLKtwwon1AGq8/U+umR6jnYEKTXQwwKbWJjPw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790960093; c=relaxed/simple; bh=F3lX++ECEDnCerWaE/Qm2cVLVgUL0h3RSsZEjxjpA7g=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=pJw5bHjHUCYb51URYNts5Aca0SvuAbv7mpiTtGzT15hmFP8v5rlAuZJrB87EaYQ0+Wo0F8VNWf5HbkSVgGr8uMwi8j1aAOm48+zqjXde2yENF4K2uSKlI03JPFgKWPmsid+EPF2iWtDTs8OnncX9BhslvfEZHT3DUk5OU4/DbYI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=CWVOIwqt; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="CWVOIwqt" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1790960090; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding; bh=3venwHttFa+LgHQOSujImduuESDWjIxfZmH5qy3F8h4=; b=CWVOIwqtkKk1zciziW/8Mr5fsRDLOwQo8D5LlQHibVCQr2DRSxiajKaKG5JxeiG8yX85GB ZpskFbg/+bpUXvsQVhuydG5SMshiml7N0cMoSvVS5HSSzyXqnjWxUBekxLjU2d9LR4CfhF 5yZekb8gyx4PCrZc2qMFTLwkaUyWHxc= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-669-0DeL34JjN4OoD5f_MtJ0XA-1; Fri, 02 Oct 2026 12:54:49 -0400 X-MC-Unique: 0DeL34JjN4OoD5f_MtJ0XA-1 X-Mimecast-MFC-AGG-ID: 0DeL34JjN4OoD5f_MtJ0XA_1790960088 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id A26681954121; Fri, 2 Oct 2026 16:54:47 +0000 (UTC) Received: from llong-thinkpadp16vgen1.rmtusnh.csb (headnet03.pony-001.prod.iad2.dc.redhat.com [10.2.32.114]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id CA8CA1956042; Fri, 2 Oct 2026 16:54:45 +0000 (UTC) From: Waiman Long To: Ridong Chen , Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= Cc: cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, Farhad Alemi , Waiman Long Subject: [PATCH v2] cgroup/cpuset: Handle cpu hotplug race in guarantee_active_cpus() Date: Fri, 2 Oct 2026 12:54:38 -0400 Message-ID: <20261002165438.951550-1-longman@redhat.com> Precedence: bulk X-Mailing-List: cgroups@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 With commit 2125c0034c5d ("cgroup/cpuset: Make cpuset hotplug processing synchronous"), the cpuset hotplug operation becomes synchronous. That commit also removes the code that handles race between cpuset_hotplug_work and cpu hotplug notifier with the assumption that race is now gone. Later commit 7a0aabd9ce69 ("cgroup/cpuset: Always use cpu_active_mask") updates the cpuset code to always use cpu_active_mask instead of cpu_ohline_mask in some places including guarantee_online_cpus() which is renamed to guarantee_active_cpus() in that commit. In the case of CPU offline operation, cpuset_active_mask is updated first in sched_cpu_deactivate() to remove the offline CPU before cpuset_handle_hotplug() is called to update the effective_cpus of the affected cpusets. The cpu_online_mask is updated after that near the end of the offline operation to remove the offline CPU. As a result, the race comes back and the top cpuset may not have any active CPU leading to NULL pointer dereference during the race window when guarantee_active_cpus() is called after cpu_active_mask is updated to remove the CPU to be torn down but before cpuset_handle_hotplug() is able to properly update the effective_cpus of the top cpuset. Fix this by adding back the NULL cs check to avoid this problem. Fixes: 7a0aabd9ce69 ("cgroup/cpuset: Always use cpu_active_mask") Reported-by: Farhad Alemi Link: https://lore.kernel.org/lkml/CA+0ovChh3VjsKN1g+ZGjwwY2fGTpP7uD+aCCByLj5Qbymw=bfQ@mail.gmail.com Tested-by: Farhad Alemi Signed-off-by: Waiman Long --- kernel/cgroup/cpuset.c | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index ddc4c9c9b4fb..92559fc4c444 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -517,10 +517,26 @@ static void guarantee_active_cpus(struct task_struct *tsk, rcu_read_lock(); cs = task_cs(tsk); - while (!cpumask_intersects(cs->effective_cpus, pmask)) + while (!cpumask_intersects(cs->effective_cpus, pmask)) { cs = parent_cs(cs); - + if (unlikely(!cs)) { + /* + * The top cpuset doesn't have any active cpu as a + * consequence of a race between its caller and the cpu + * hotplug operation where cpu_active_mask is updated + * asynchronously before cpuset_handle_hotplug() is + * being called to adjust the effective_cpus of the + * affected cpusets. But we know the top cpuset's + * effective_cpus is on its way to be identical to + * cpu_active_mask minus the exclusive CPUs dedicated + * to other valid cpuset partitions. Just pass back + * the filtered cpu_active_mask in this case. + */ + goto out_unlock; + } + } cpumask_and(pmask, pmask, cs->effective_cpus); +out_unlock: rcu_read_unlock(); } -- 2.55.0