* [BUG] cgroup/cpuset: Concurrent WARN_ON_ONCE triggers during remote partition stress-test
@ 2026-08-31 9:05 Zhang Qiao
2026-08-31 14:18 ` Waiman Long
0 siblings, 1 reply; 3+ messages in thread
From: Zhang Qiao @ 2026-08-31 9:05 UTC (permalink / raw)
To: Waiman Long, ridong.chen, Tejun Heo, Johannes Weiner,
Michal Koutný
Cc: jifa, Hui Tang, cgroups
Hi,
While stress-testing cpuset partitions under KASAN (kernel 7.2-rc1,
dc59e4fea9d83 "Linux 7.2-rc1"), four distinct WARN_ON_ONCE() in
kernel/cgroup/cpuset.c fire within a ~1s window. They all concern the
remote partition / effective-cpumask invariants and are triggered by
concurrent CPU hotplug, cpuset.cpus/cpuset.cpus.exclusive/
cpuset.cpus.partition writes, task migration and (un)partitioning of
nested subgroups.
Environment
-----------
- Kernel : Linux 7.2-rc1, generic KASAN enabled, x86_64, PREEMPT
- server : Intel Xeon Platinum 8380 @ 2.30GHz, 2 NUMA nodes, 160 CPUs
- Cmdline: ... cgroup_disable=files apparmor=0 systemd.unified_cgroup_hierarchy=1
Reproduction
------------
A pure-shell self-contained script (attached below) runs several concurrent
"disturbance" loops. The issue is extremely easy to reproduce; the script
consistently triggers the WARNs within seconds of execution.
- CPU hotplug toggle of a helper pool (CPU online/offline)
- random writes to cpuset.cpus / cpuset.cpus.exclusive
(single CPU, empty, or range)
- random root <-> member switching of cpuset.cpus.partition
- migrating burner tasks between subgroups via cgroup.procs
- creating and removing nested sub-partitions under S0-S3
Run as: bash fn_warn_repro_008_shell.sh
The parent cgroup TEST_CGROOT is set as "member" and four children
S0-S3 as remote partition roots, giving the invariant checks above a
reason to run concurrently.
Observed
--------
Four WARN_ON_ONCE fire (trimmed traces below; register dumps omitted). No KASAN
use-after-free or list corruption was observed on this run.
1) kernel/cgroup/cpuset.c:867 generate_sched_domains+0x489/0x900, CPU#155:
bash/9129
WARN_ON_ONCE(1) when two partition-root cpusets overlap.
---------------[ cut here ]---------------
WARNING: kernel/cgroup/cpuset.c:867 at generate_sched_domains+0x489/0x900,
CPU#155: bash/9129
CPU: 155 UID: 0 PID: 9129 Comm: bash Kdump: loaded Tainted: G S 7.2.0+
#11 PREEMPTLAZY
Tainted: [S]=CPU_OUT_OF_SPEC
RIP: 0010:generate_sched_domains+0x489/0x900
Call Trace:
<TASK>
? update_cpumask+0x617/0x7c0
rebuild_sched_domains_locked+0x8c/0x1c0
? __pfx_rebuild_sched_domains_locked+0x10/0x10
? kasan_save_track+0x10/0x30
? cpuset_write_resmask+0x1fc/0x880
? kfree+0x162/0x440
cpuset_update_sd_hk_unlock+0xc7/0xf0
cpuset_write_resmask+0x201/0x880
cgroup_file_write+0x1b4/0x600
kernfs_fop_write_iter+0x31c/0x500
vfs_write+0x58c/0xca0
ksys_write+0xef/0x1c0
do_syscall_64+0xaf/0x550
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
---[ end trace 0000000000000000 ]---
2) kernel/cgroup/cpuset.c:1599 remote_cpus_update+0x6d5/0x990, CPU#150: bash/9125
WARN_ON_ONCE(!cpumask_subset(cs->effective_xcpus, subpartitions_cpus)).
---------------[ cut here ]---------------
WARNING: kernel/cgroup/cpuset.c:1599 at remote_cpus_update+0x6d5/0x990, CPU#150:
bash/9125
CPU: 150 UID: 0 PID: 9125 Comm: bash Kdump: loaded Tainted: G S W
7.2.0+ #11 PREEMPTLAZY
Tainted: [S]=CPU_OUT_OF_SPEC, [W]=WARN
RIP: 0010:remote_cpus_update+0x6d5/0x990
Call Trace:
<TASK>
? partition_cpus_change.part.0+0x249/0x470
update_exclusive_cpumask+0x438/0x810
? __pfx_update_exclusive_cpumask+0x10/0x10
cpuset_write_resmask+0x33b/0x880
cgroup_file_write+0x1b4/0x600
kernfs_fop_write_iter+0x31c/0x500
vfs_write+0x58c/0xca0
ksys_write+0xef/0x1c0
do_syscall_64+0xaf/0x550
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
---[ end trace 0000000000000000 ]---
3) kernel/cgroup/cpuset.c:1513 remote_partition_enable+0x3ac/0x4b0, CPU#149:
bash/9124
WARN_ON_ONCE(cpumask_intersects(tmp->new_cpus, subpartitions_cpus)).
---------------[ cut here ]---------------
WARNING: kernel/cgroup/cpuset.c:1513 at remote_partition_enable+0x3ac/0x4b0,
CPU#149: bash/9124
CPU: 149 UID: 0 PID: 9124 Comm: bash Kdump: loaded Tainted: G S W
7.2.0+ #11 PREEMPTLAZY
Tainted: [S]=CPU_OUT_OF_SPEC, [W]=WARN
RIP: 0010:remote_partition_enable+0x3ac/0x4b0
Call Trace:
<TASK>
update_prstate+0x998/0xb60
? __pfx_update_prstate+0x10/0x10
? entry_SYSCALL_64_after_hwframe+0x76/0x7e
cpuset_partition_write+0xcb/0x120
cgroup_file_write+0x1b4/0x600
kernfs_fop_write_iter+0x31c/0x500
vfs_write+0x58c/0xca0
ksys_write+0xef/0x1c0
do_syscall_64+0xaf/0x550
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
---[ end trace 0000000000000000 ]---
4) kernel/cgroup/cpuset.c:2050 compute_partition_effective_cpumask+0x710/0x900,
CPU#44: bash/9124
WARN_ON_ONCE(is_remote_partition(child)) — remote partition underneath
another partition root.
---------------[ cut here ]---------------
WARNING: kernel/cgroup/cpuset.c:2050 at
compute_partition_effective_cpumask+0x710/0x900, CPU#44: bash/9124
CPU: 44 UID: 0 PID: 9124 Comm: bash Kdump: loaded Tainted: G S W
7.2.0+ #11 PREEMPTLAZY
Tainted: [S]=CPU_OUT_OF_SPEC, [W]=WARN
RIP: 0010:compute_partition_effective_cpumask+0x710/0x900
Call Trace:
<TASK>
update_cpumasks_hier+0x9df/0x1200
? remote_partition_enable+0x353/0x4b0
update_prstate+0x8ff/0xb60
? __pfx_update_prstate+0x10/0x10
? entry_SYSCALL_64_after_hwframe+0x76/0x7e
cpuset_partition_write+0xcb/0x120
cgroup_file_write+0x1b4/0x600
kernfs_fop_write_iter+0x31c/0x500
vfs_write+0x58c/0xca0
ksys_write+0xef/0x1c0
do_syscall_64+0xaf/0x550
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
---[ end trace 0000000000000000 ]---
The line numbers above map to the exact 7.2-rc1 source
These WARNs suggest the remote-partition / subpartitions_cpus bookkeeping
can become momentarily inconsistent when partition *, exclusive and
cpuset.cpus writes race with CPU hotplug and parallel subgroup
(un)partitioning. I haven't pinpointed the exact race or prepared a patch yet.
Please let me know if you need full dmesg logs, specific ftrace outputs, or if I
should test any debugging patches.
Thanks,
Zhang Qiao
======================================================================
Reproduction script (fn_warn_repro_008_shell.sh)
======================================================================
#!/bin/bash
#
# Copyright (c) Huawei Technologies Co., Ltd. 2026-2026. All rights reserved.
# Author: jifa@huawei.com
# Create: 2026/08/31
# TestCase Description: FN-WARN-REPRO-008 (PURE SHELL) — reproduce
slab-use-after-free under KASAN
#
# This is a pure-shell, self-contained single-file version of
fn_warn_repro_008.sh.
# It has no dependency on:
# - common_base.sh / sched_common.sh / common.sh
# - the cpu_sim binary (replaced by a bash `while` busy-loop sub-shell)
# - base64 / xz decoding (the script no longer embeds any binary)
#
# The key to triggering the cpuset_mutex race is not CPU utilization itself,
# but the cgroup.procs write. So a simple bash busy-loop sub-process is enough
# to create task-migration pressure.
#
# Goals:
# 1. KASAN: slab-use-after-free in remote_partition_check /
remote_partition_enable
# 2. cpuset.c:1753 (remote_partition_disable operating on a corrupted list)
# 3. lib/list_debug.c:29 / 35 (list_add operating on a corrupted list)
# 4. soft lockup / NULL deref / Oops / panic (as a side-effect of the race)
#
# Verification: look for KASAN / list_debug / cpuset.c:1753 / NULL deref / soft
lockup in dmesg
#
# Usage:
# bash fn_warn_repro_008_shell.sh [duration_sec]
# default duration=60s
#
# Exit codes:
# 0 = completed (inspect the log output to determine whether it reproduced)
# 1+ = environment error
set -u
# =============================================================================
# Global configuration
# =============================================================================
STANDALONE_DIR="$(cd "$(dirname "$0")" && pwd)"
CASE_NAME="${0/.sh/}"
LOG_FILE="${CASE_NAME}.log"
REPRO_FILE="${CASE_NAME}.repro"
: > "${LOG_FILE}"
: > "${REPRO_FILE}"
DURATION="${1:-${ST_LONG_DURATION:-60}}"
export DURATION
MAX_TASKS=2
HELPER_CPU_POOL=""
ROOT_CG="/sys/fs/cgroup"
TEST_CGROOT=""
declare -a task_pid task_cg_idx task_alive
task_count=0
ALL_CGS=()
SUB_PATHS=()
# =============================================================================
# Logging
# =============================================================================
function logger_info() { printf "\033[0;37m[INFO] %s\033[0m\n" "$*"; }
function logger_error() { printf "\033[0;31m[ERROR] %s\033[0m\n" "$*"; }
function logger_setup() { printf "\033[0;34m[SETUP] %s\033[0m\n" "$*"; }
function logger_pass() { printf "\033[0;32m[PASS] %s\033[0m\n" "$*"; }
function logger_failed(){ printf "\033[0;31m[FAILED] %s\033[0m\n" "$*"; }
ret=0
# =============================================================================
# 1. cgroup / environment helpers (self-contained implementation,
# equivalent to common.sh / sched_common.sh)
# =============================================================================
function is_cgroup_v2() {
[ -f /sys/fs/cgroup/cgroup.controllers ] && echo true || echo false
}
function set_subtree_control_repeat() {
local cg="$1" val="$2"
echo "${val}" > "${cg}/cgroup.subtree_control" 2>/dev/null
}
# CPU pool (avoid the system's main CPUs 0-1; use 2-5 on a 6-CPU machine)
function helper_init_cpu_pool() {
local nr_cpus
nr_cpus=$(lscpu 2>/dev/null | awk '/^CPU\(s\):/ {gsub(/[^0-9]/,"",$2); print
$2; exit}')
if [[ -n "$nr_cpus" && "$nr_cpus" -lt 8 ]]; then
if [[ "$nr_cpus" -ge 4 ]]; then
HELPER_CPU_POOL=$(seq 2 $((nr_cpus - 1)) | tr '\n' ' ')
fi
fi
[[ -z "$HELPER_CPU_POOL" ]] && HELPER_CPU_POOL="4 5 6 7"
export HELPER_CPU_POOL
}
function helper_pick_cpus() {
local n=${1:-1}
[[ -z "$HELPER_CPU_POOL" ]] && helper_init_cpu_pool
local pool=($HELPER_CPU_POOL)
[[ ${#pool[@]} -lt $n ]] && return 1
echo "${pool[@]:0:n}"
}
# Create an isolated CGROOT and set it as the partition root
function helper_setup_cgroot() {
local suffix=${1:-$$}
TEST_CGROOT="${ROOT_CG}/test_cpuset_partition_${suffix}"
if [[ -d "$TEST_CGROOT" ]]; then
helper_cleanup_cgroot "$TEST_CGROOT"
fi
set_subtree_control_repeat "$ROOT_CG" "+cpuset"
mkdir -p "$TEST_CGROOT"
set_subtree_control_repeat "$TEST_CGROOT" "+cpuset"
[[ -z "$HELPER_CPU_POOL" ]] && helper_init_cpu_pool
if [[ -n "$HELPER_CPU_POOL" ]]; then
echo 0 > "$TEST_CGROOT/cpuset.mems" 2>/dev/null
echo "$HELPER_CPU_POOL" > "$TEST_CGROOT/cpuset.cpus" 2>/dev/null
echo "root" > "$TEST_CGROOT/cpuset.cpus.partition" 2>/dev/null
echo "$HELPER_CPU_POOL" > "$TEST_CGROOT/cpuset.cpus.exclusive" 2>/dev/null
fi
export TEST_CGROOT
}
# Enable +cpuset top-down one level at a time so child cgroups can write cpuset.cpus
function helper_mkdir_cg() {
local cg
for cg in "$@"; do
mkdir -p "$cg"
local chain=()
local p="$cg"
while [[ -n "$p" && "$p" != "/" ]]; do
[[ "$p" == "$ROOT_CG" ]] && break
chain=("$p" "${chain[@]}")
p=$(dirname "$p")
done
local d
for d in "${chain[@]}"; do
[[ -f "$d/cgroup.subtree_control" ]] && \
set_subtree_control_repeat "$d" "+cpuset" 2>/dev/null
done
done
}
function helper_cleanup_partition() {
local cg="$1"
[[ -z "$cg" || ! -d "$cg" ]] && return 0
echo "member" > "$cg/cpuset.cpus.partition" 2>/dev/null
echo "" > "$cg/cpuset.cpus" 2>/dev/null
echo "" > "$cg/cpuset.cpus.exclusive" 2>/dev/null
}
function helper_cleanup_cgroot() {
local cgroot="$1"
[[ -z "$cgroot" || ! -d "$cgroot" ]] && return 0
if [[ -f "$cgroot/cgroup.procs" ]]; then
while read -r pid; do
[[ -n "$pid" && "$pid" -gt 100 ]] && echo "$pid" >
/sys/fs/cgroup/cgroup.procs 2>/dev/null
done < "$cgroot/cgroup.procs"
fi
find "$cgroot" -depth -type d 2>/dev/null | while read -r d; do
helper_cleanup_partition "$d"
rmdir "$d" 2>/dev/null
done
}
# Kill the burner child processes started by this script (isolated via process
group to avoid killing unrelated processes)
function cleanup_all_tasks() {
local p
for p in "${task_pid[@]:-}"; do
[[ -n "$p" ]] && kill -9 "$p" 2>/dev/null
done
# Fallback: kill all bash burners started by this script (tagged via env var
BURNER_PARENT=<ppid>)
pkill -9 -f "BURNER_PARENT=$$" 2>/dev/null
}
# Recover all CPUs to online by enumerating sysfs (cannot use nproc — it only
counts online CPUs)
function recover_hotplug_cpus() {
local cpu_sysfs="/sys/devices/system/cpu"
local cpu_dir
for cpu_dir in ${cpu_sysfs}/cpu[0-9]*; do
[[ -d "$cpu_dir" ]] || continue
local n=${cpu_dir##*cpu}
[[ "$n" == "0" ]] && continue
local f="${cpu_dir}/online"
if [[ -f "$f" ]]; then
local s
s=$(cat "$f" 2>/dev/null)
if [[ "$s" == "0" ]]; then
echo 1 > "$f" 2>/dev/null
[[ $? -eq 0 ]] && logger_info "Recovered cpu${n} online=1" ||
logger_error "Failed to recover cpu${n}"
fi
fi
done
}
function helper_full_cleanup() {
cleanup_all_tasks
recover_hotplug_cpus
[[ -n "$TEST_CGROOT" ]] && helper_cleanup_cgroot "$TEST_CGROOT"
}
# =============================================================================
# 2. CPU burner — a pure-bash sub-process, replacing cpu_sim
# Uses a `while true; do :; done` busy-loop to consume CPU, tagged with
# BURNER_PARENT=<ppid> on the command line so cleanup can pkill -f it
# precisely without killing other bash processes
# =============================================================================
function start_burner() {
# $1 = runtime in seconds (self-terminates after that to avoid leaking)
local secs=${1:-$((DURATION + 30))}
BURNER_PARENT=$$ bash -c "
trap 'exit 0' TERM
end=\$((SECONDS + ${secs}))
while [[ \$SECONDS -lt \$end ]]; do
:
done
" &
echo $!
}
# =============================================================================
# 3. Disturbance functions (ported directly from the original fn_warn_repro_008.sh)
# =============================================================================
function disturb_hotplug() {
local pool=($HELPER_CPU_POOL)
local end=$((SECONDS + DURATION))
while [[ $SECONDS -lt $end ]]; do
local c="${pool[$((RANDOM % ${#pool[@]}))]}"
echo 0 > "/sys/devices/system/cpu/cpu$c/online" 2>/dev/null
sleep 0.02
echo 1 > "/sys/devices/system/cpu/cpu$c/online" 2>/dev/null
sleep 0.03
if (( RANDOM % 2 == 0 )); then
for c2 in "${pool[@]}"; do
echo 0 > "/sys/devices/system/cpu/cpu$c2/online" 2>/dev/null
done
sleep 0.05
for c2 in "${pool[@]}"; do
echo 1 > "/sys/devices/system/cpu/cpu$c2/online" 2>/dev/null
done
sleep 0.08
fi
done
}
function disturb_partition_switch() {
local end=$((SECONDS + DURATION))
while [[ $SECONDS -lt $end ]]; do
local n=${#ALL_CGS[@]}
[[ $n -eq 0 ]] && { sleep 0.1; continue; }
local idx=$((RANDOM % n))
local cg="${ALL_CGS[idx]}"
[[ ! -d "$cg" ]] && { sleep 0.02; continue; }
local pr="root"
(( RANDOM % 2 == 0 )) && pr="member"
echo "$pr" > "$cg/cpuset.cpus.partition" 2>/dev/null
sleep 0.02
done
}
function disturb_exclusive_write() {
local pool=($HELPER_CPU_POOL)
local end=$((SECONDS + DURATION))
while [[ $SECONDS -lt $end ]]; do
local n=${#ALL_CGS[@]}
[[ $n -eq 0 ]] && { sleep 0.1; continue; }
local idx=$((RANDOM % n))
local cg="${ALL_CGS[idx]}"
[[ ! -d "$cg" ]] && { sleep 0.02; continue; }
local r=$((RANDOM % 3))
case $r in
0)
local c="${pool[$((RANDOM % ${#pool[@]}))]}"
echo "$c" > "$cg/cpuset.cpus.exclusive" 2>/dev/null
;;
1)
echo "" > "$cg/cpuset.cpus.exclusive" 2>/dev/null
;;
2)
local c1="${pool[$((RANDOM % ${#pool[@]}))]}"
local c2="${pool[$((RANDOM % ${#pool[@]}))]}"
[[ "$c1" -gt "$c2" ]] && { local t=$c1; c1=$c2; c2=$t; }
echo "$c1-$c2" > "$cg/cpuset.cpus.exclusive" 2>/dev/null
;;
esac
sleep 0.02
done
}
function disturb_cpus_write() {
local pool=($HELPER_CPU_POOL)
local end=$((SECONDS + DURATION))
while [[ $SECONDS -lt $end ]]; do
local n=${#ALL_CGS[@]}
[[ $n -eq 0 ]] && { sleep 0.1; continue; }
local idx=$((RANDOM % n))
local cg="${ALL_CGS[idx]}"
[[ ! -d "$cg" ]] && { sleep 0.02; continue; }
local r=$((RANDOM % 2))
case $r in
0)
local c="${pool[$((RANDOM % ${#pool[@]}))]}"
echo "$c" > "$cg/cpuset.cpus" 2>/dev/null
;;
1)
echo "" > "$cg/cpuset.cpus" 2>/dev/null
;;
esac
sleep 0.02
done
}
# Task-migration disturbance: use a bash burner instead of cpu_sim
function disturb_task_migrate() {
local i
for ((i = 0; i < MAX_TASKS; i++)); do
local pid
pid=$(start_burner $((DURATION + 30)))
task_pid[task_count]=$pid
task_cg_idx[task_count]=0
task_alive[task_count]=1
((task_count++))
sleep 0.1
done
for ((i = 0; i < task_count; i++)); do
[[ ${#ALL_CGS[@]} -gt 0 ]] && echo "${task_pid[i]}" >
"${ALL_CGS[0]}/cgroup.procs" 2>/dev/null
done
local end=$((SECONDS + DURATION))
while [[ $SECONDS -lt $end ]]; do
local alive=()
for ((i = 0; i < task_count; i++)); do
[[ "${task_alive[i]}" == "1" ]] && alive+=($i)
done
[[ ${#alive[@]} -eq 0 ]] && break
local ti=${alive[$((RANDOM % ${#alive[@]}))]}
local n=${#ALL_CGS[@]}
[[ $n -eq 0 ]] && { sleep 0.2; continue; }
local idx=$((RANDOM % n))
local cg="${ALL_CGS[idx]}"
[[ ! -d "$cg" ]] && { sleep 0.05; continue; }
echo "${task_pid[ti]}" > "$cg/cgroup.procs" 2>/dev/null
task_cg_idx[ti]=$idx
sleep 0.1
done
for ((i = 0; i < task_count; i++)); do
[[ "${task_alive[i]}" == "1" ]] && kill -9 "${task_pid[i]}" 2>/dev/null
done
}
function disturb_grandson_ops() {
local pool=($HELPER_CPU_POOL)
local end=$((SECONDS + DURATION))
local sons=("${SUB_PATHS[@]}")
while [[ $SECONDS -lt $end ]]; do
local k
for ((k = 0; k < 5; k++)); do
[[ ${#ALL_CGS[@]} -ge 80 ]] && break
[[ ${#sons[@]} -eq 0 ]] && break
local sidx=$((RANDOM % ${#sons[@]}))
local parent="${sons[sidx]}"
[[ ! -d "$parent" ]] && continue
local suffix=$((RANDOM % 100000))
local new="${parent}/L1_${suffix}"
[[ -d "$new" ]] && continue
helper_mkdir_cg "$new" 2>/dev/null
if [[ -d "$new" ]]; then
echo 0 > "$new/cpuset.mems" 2>/dev/null
local c="${pool[$((RANDOM % ${#pool[@]}))]}"
echo "$c" > "$new/cpuset.cpus" 2>/dev/null
echo "$c" > "$new/cpuset.cpus.exclusive" 2>/dev/null
echo "root" > "$new/cpuset.cpus.partition" 2>/dev/null
ALL_CGS+=("$new")
fi
done
local rmdone=0
local i
for ((i = 4; i < ${#ALL_CGS[@]} && rmdone < 3; i++)); do
if [[ "${ALL_CGS[i]}" == *"/L1_"* ]]; then
local cg="${ALL_CGS[i]}"
rmdir "$cg" 2>/dev/null
unset 'ALL_CGS[i]'
((rmdone++))
fi
done
ALL_CGS=("${ALL_CGS[@]}")
sleep 0.01
done
}
# =============================================================================
# 4. setup / do_test / cleanup
# =============================================================================
function setup() {
[[ "$(is_cgroup_v2)" != "true" ]] && { logger_error "need cgroup v2";
((ret+=1)); exit 1; }
ROOT_CG=/sys/fs/cgroup
helper_init_cpu_pool
helper_setup_cgroot
}
function do_test() {
local pool
pool=$(helper_pick_cpus 4)
[[ -z "$pool" || $(echo "$pool" | wc -w) -lt 4 ]] && { logger_error "need ≥4
CPUs in pool, skip"; ((ret+=1)); return; }
local pool_arr=($pool)
logger_info "pool=${pool[*]} duration=${DURATION}s"
local dmesg_lines_before
dmesg_lines_before=$(dmesg 2>/dev/null | wc -l)
# TEST_CGROOT as member + 4 remote-partition subgroups S0-S3
echo "member" > "$TEST_CGROOT/cpuset.cpus.partition" 2>/dev/null
sleep 0.2
echo "$pool" > "$TEST_CGROOT/cpuset.cpus" 2>/dev/null
echo "$pool" > "$TEST_CGROOT/cpuset.cpus.exclusive" 2>/dev/null
sleep 0.3
local i
for ((i = 0; i < 4; i++)); do
local cg="${TEST_CGROOT}/S${i}"
local c="${pool_arr[i % 4]}"
helper_mkdir_cg "$cg"
echo 0 > "$cg/cpuset.mems" 2>/dev/null
echo "$c" > "$cg/cpuset.cpus" 2>/dev/null
echo "$c" > "$cg/cpuset.cpus.exclusive" 2>/dev/null
echo "root" > "$cg/cpuset.cpus.partition" 2>/dev/null
SUB_PATHS+=("$cg")
ALL_CGS+=("$cg")
sleep 0.03
done
logger_info "created ${#SUB_PATHS[@]} remote partition subgroups
(parent=member)"
# 8 concurrent disturbances: hotplug + partition_switch + 3x exclusive_write
+ cpus_write + task_migrate + grandson_ops
disturb_hotplug & local p1=$!
disturb_partition_switch & local p2=$!
disturb_exclusive_write & local p3=$!
disturb_exclusive_write & local p4=$!
disturb_exclusive_write & local p5=$!
disturb_cpus_write & local p6=$!
disturb_task_migrate & local p7=$!
disturb_grandson_ops & local p8=$!
# Check dmesg every 2s; stop as soon as an issue is hit
local warned=0
local check_end=$((SECONDS + DURATION))
while [[ $SECONDS -lt $check_end ]]; do
sleep 2
local new_warnings
new_warnings=$(dmesg 2>/dev/null | tail -n +$((dmesg_lines_before + 1)) | \
grep -cE
"WARNING.*cpuset|WARNING.*list_debug|KASAN|cpuset.c:(1753|1793|2195)|list_add_valid_or_report|list_del
corruption|soft lockup|general protection fault|Unable to handl
e|oops|BUG:|panic|segfault|slab-use-after-free|use-after-free")
if [[ $new_warnings -gt 0 ]]; then
warned=1
logger_info "issue detected in dmesg — stopping disturbances early"
break
fi
done
kill -9 $p1 $p2 $p3 $p4 $p5 $p6 $p7 $p8 2>/dev/null
pkill -9 -f "BASH_BURNER_${$}" 2>/dev/null
cleanup_all_tasks
sleep 1
# Force-clean any leftovers in TEST_CGROOT
if [[ -n "$TEST_CGROOT" && -d "$TEST_CGROOT" ]]; then
if [[ -f "$TEST_CGROOT/cgroup.procs" ]]; then
while read -r p; do
[[ -n "$p" && "$p" -gt 100 ]] && echo "$p" >
/sys/fs/cgroup/cgroup.procs 2>/dev/null
done < "$TEST_CGROOT/cgroup.procs"
fi
find "$TEST_CGROOT" -depth -mindepth 1 -type d 2>/dev/null | while read
-r d; do
echo "member" > "$d/cpuset.cpus.partition" 2>/dev/null
echo "" > "$d/cpuset.cpus" 2>/dev/null
echo "" > "$d/cpuset.cpus.exclusive" 2>/dev/null
rmdir "$d" 2>/dev/null
done
echo "member" > "$TEST_CGROOT/cpuset.cpus.partition" 2>/dev/null
echo "" > "$TEST_CGROOT/cpuset.cpus" 2>/dev/null
echo "" > "$TEST_CGROOT/cpuset.cpus.exclusive" 2>/dev/null
fi
# Extract all trace segments + other dmesg anomalies
local dmesg_full dmesg_inc
dmesg_full=$(dmesg 2>/dev/null)
dmesg_inc=$(echo "$dmesg_full" | tail -n +$((dmesg_lines_before + 1)))
local trace_count=0
local current_trace=""
local in_trace=0
local line
while IFS= read -r line; do
if [[ "$line" == *"------------[ cut here ]------------"* ]]; then
in_trace=1
current_trace="$line"$'\n'
elif [[ "$in_trace" == "1" ]]; then
current_trace+="$line"$'\n'
if [[ "$line" == *"---[ end trace "* ]]; then
in_trace=0
((trace_count++))
echo "=== Reproduced Issue #${trace_count} ===" >> "${REPRO_FILE}"
echo "Timestamp: $(date '+%Y-%m-%d %H:%M:%S')" >> "${REPRO_FILE}"
local warn_line
warn_line=$(echo "$current_trace" | grep -E
"WARNING:.*cpuset\.c:|WARNING:.*list_debug" | head -1)
[[ -n "$warn_line" ]] && echo "Warning: ${warn_line}" >>
"${REPRO_FILE}"
local kasan_line
kasan_line=$(echo "$current_trace" | grep -iE
"KASAN|slab-use-after-free|use-after-free" | head -1)
[[ -n "$kasan_line" ]] && echo "KASAN: ${kasan_line}" >>
"${REPRO_FILE}"
local oops_line
oops_line=$(echo "$current_trace" | grep -iE "general protection
fault|Unable to handle|oops|BUG:|panic|segfault|NULL pointer" | head -1)
[[ -n "$oops_line" ]] && echo "Oops: ${oops_line}" >>
"${REPRO_FILE}"
echo "--- Full Trace ---" >> "${REPRO_FILE}"
echo "$current_trace" >> "${REPRO_FILE}"
echo "" >> "${REPRO_FILE}"
current_trace=""
fi
fi
done <<< "$dmesg_inc"
local other_issues
other_issues=$(echo "$dmesg_inc" | grep -iE "soft lockup|list_del
corruption|list_add_valid_or_report|KASAN|slab-use-after-free|use-after-free|BUG:|oops:|panic|Unable
to handle|genera
l protection fault|segfault|NULL pointer|kernel page fault" | grep -v "WARNING:"
| head -30)
if [[ -n "$other_issues" ]]; then
echo "=== Other Kernel Issues ===" >> "${REPRO_FILE}"
echo "$other_issues" >> "${REPRO_FILE}"
echo "" >> "${REPRO_FILE}"
fi
if [[ $trace_count -gt 0 || -n "$other_issues" ]]; then
logger_info "==== REPRODUCED ${trace_count} traces + other issues ===="
logger_info "Detailed traces saved to: ${REPRO_FILE}"
awk '
/^=== Reproduced Issue #/ { if (cur) print cur; cur=$0 }
/^Warning: / { cur = cur " | " $0 }
/^KASAN: / { cur = cur " | " $0 }
/^Oops: / { cur = cur " | " $0 }
END { if (cur) print cur }
' "${REPRO_FILE}" | while IFS= read -r summary; do
logger_info "$summary"
done
logger_pass "reproduced ${trace_count} traces (see ${REPRO_FILE})"
if echo "$dmesg_inc" | grep -qE
"KASAN|slab-use-after-free|use-after-free"; then
logger_pass "KASAN UAF detected! (target of fn_warn_repro_008)"
fi
if echo "$dmesg_inc" | grep -qE "BUG:|oops:|panic|Unable to
handle|general protection fault|segfault"; then
logger_info "kernel panic/oops detected (see ${REPRO_FILE})"
fi
else
logger_info "no issues this run (probabilistic) — try more iterations"
fi
if echo "$dmesg_inc" | grep -qE "KASAN|slab-use-after-free|use-after-free"; then
logger_pass "KASAN UAF reproduced!"
fi
if echo "$dmesg_inc" | grep -qE "BUG:|oops:|panic|Unable to handle|general
protection fault|segfault|kernel page fault"; then
logger_info "kernel panic/crash detected (details in ${REPRO_FILE})"
fi
[ $trace_count -gt 0 ] && logger_pass "PASS: reproduced kernel issues!" ||
logger_info "INFO: no issue reproduced this iteration (run again to increase
chance)"
}
function cleanup() {
helper_full_cleanup
}
# =============================================================================
# main
# =============================================================================
logger_setup "Starting ${CASE_NAME} (pure-shell, duration=${DURATION}s)"
setup
do_test
cleanup
exit ${ret}
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [BUG] cgroup/cpuset: Concurrent WARN_ON_ONCE triggers during remote partition stress-test
2026-08-31 9:05 [BUG] cgroup/cpuset: Concurrent WARN_ON_ONCE triggers during remote partition stress-test Zhang Qiao
@ 2026-08-31 14:18 ` Waiman Long
2026-09-01 2:25 ` Zhang Qiao
0 siblings, 1 reply; 3+ messages in thread
From: Waiman Long @ 2026-08-31 14:18 UTC (permalink / raw)
To: Zhang Qiao, ridong.chen, Tejun Heo, Johannes Weiner,
Michal Koutný
Cc: jifa, Hui Tang, cgroups
On 8/31/26 5:05 AM, Zhang Qiao wrote:
> Hi,
>
> While stress-testing cpuset partitions under KASAN (kernel 7.2-rc1,
> dc59e4fea9d83 "Linux 7.2-rc1"), four distinct WARN_ON_ONCE() in
> kernel/cgroup/cpuset.c fire within a ~1s window. They all concern the
> remote partition / effective-cpumask invariants and are triggered by
> concurrent CPU hotplug, cpuset.cpus/cpuset.cpus.exclusive/
> cpuset.cpus.partition writes, task migration and (un)partitioning of
> nested subgroups.
>
> Environment
> -----------
> - Kernel : Linux 7.2-rc1, generic KASAN enabled, x86_64, PREEMPT
> - server : Intel Xeon Platinum 8380 @ 2.30GHz, 2 NUMA nodes, 160 CPUs
> - Cmdline: ... cgroup_disable=files apparmor=0 systemd.unified_cgroup_hierarchy=1
>
> Reproduction
> ------------
> A pure-shell self-contained script (attached below) runs several concurrent
> "disturbance" loops. The issue is extremely easy to reproduce; the script
> consistently triggers the WARNs within seconds of execution.
>
> - CPU hotplug toggle of a helper pool (CPU online/offline)
> - random writes to cpuset.cpus / cpuset.cpus.exclusive
> (single CPU, empty, or range)
> - random root <-> member switching of cpuset.cpus.partition
> - migrating burner tasks between subgroups via cgroup.procs
> - creating and removing nested sub-partitions under S0-S3
Thank for reporting the bug. It is recently known that there are bugs in
the handling of nested sub-partitions. We are in the process of getting
it fixed.
Cheers,
Longman
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [BUG] cgroup/cpuset: Concurrent WARN_ON_ONCE triggers during remote partition stress-test
2026-08-31 14:18 ` Waiman Long
@ 2026-09-01 2:25 ` Zhang Qiao
0 siblings, 0 replies; 3+ messages in thread
From: Zhang Qiao @ 2026-09-01 2:25 UTC (permalink / raw)
To: Waiman Long
Cc: jifa, Hui Tang, cgroups, Johannes Weiner, ridong.chen,
Michal Koutný, Tejun Heo
在 2026/8/31 22:18, Waiman Long 写道:
> On 8/31/26 5:05 AM, Zhang Qiao wrote:
>> Hi,
>>
>> While stress-testing cpuset partitions under KASAN (kernel 7.2-rc1,
>> dc59e4fea9d83 "Linux 7.2-rc1"), four distinct WARN_ON_ONCE() in
>> kernel/cgroup/cpuset.c fire within a ~1s window. They all concern the
>> remote partition / effective-cpumask invariants and are triggered by
>> concurrent CPU hotplug, cpuset.cpus/cpuset.cpus.exclusive/
>> cpuset.cpus.partition writes, task migration and (un)partitioning of
>> nested subgroups.
>>
>> Environment
>> -----------
>> - Kernel : Linux 7.2-rc1, generic KASAN enabled, x86_64, PREEMPT
>> - server : Intel Xeon Platinum 8380 @ 2.30GHz, 2 NUMA nodes, 160 CPUs
>> - Cmdline: ... cgroup_disable=files apparmor=0
>> systemd.unified_cgroup_hierarchy=1
>>
>> Reproduction
>> ------------
>> A pure-shell self-contained script (attached below) runs several concurrent
>> "disturbance" loops. The issue is extremely easy to reproduce; the script
>> consistently triggers the WARNs within seconds of execution.
>>
>> - CPU hotplug toggle of a helper pool (CPU online/offline)
>> - random writes to cpuset.cpus / cpuset.cpus.exclusive
>> (single CPU, empty, or range)
>> - random root <-> member switching of cpuset.cpus.partition
>> - migrating burner tasks between subgroups via cgroup.procs
>> - creating and removing nested sub-partitions under S0-S3
>
> Thank for reporting the bug. It is recently known that there are bugs in the
> handling of nested sub-partitions. We are in the process of getting it fixed.
>
> Cheers,
> Longman
Hi Waiman,
Thanks for the update!
I'm looking forward to the fix. Please CC me when the patch is ready, and I'll
be happy to test it on my end.
Thanks,
Zhang Qiao
>
>
> .
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-01 2:26 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-31 9:05 [BUG] cgroup/cpuset: Concurrent WARN_ON_ONCE triggers during remote partition stress-test Zhang Qiao
2026-08-31 14:18 ` Waiman Long
2026-09-01 2:25 ` Zhang Qiao
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox