From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from va-1-111.ptr.blmpb.com (va-1-111.ptr.blmpb.com [209.127.230.111]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9660D3D6473 for ; Wed, 26 Aug 2026 10:01:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.230.111 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787738469; cv=none; b=gyTrE8WJTUptOuDSQNZexySvWrOQqnSwIW22TVjZlXDUerpUiy3J33JdkaJV9ihZLmuIMpbu3z7xMTwkOzV3sOgE0dqrdPOIEq2xInh4odHXm3sca46mkMAtac1bYwSjJzWuESpnrG4KHAlXSPXAuy7U7ehnJ0ZB5ibBxyOvf2w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787738469; c=relaxed/simple; bh=PGt/hEM1AfGXk8Nq3YoL1cnwNwCVkjwdL73bo8w6Qtk=; h=Subject:Date:Cc:From:Message-Id:References:To:Mime-Version: Content-Disposition:In-Reply-To:Content-Type; b=ufECj2oUbNKw1fSpkqfmU9FILfghK/YBGYV+9TD8OcbaARU5boTUDE5kFc0vCBuN66hBCvnixZjhywCWtRB0Sbk/F0pp8xqBvEw7fBmZ3kHTI2PT5K63uKaB3b2MpVS0EUEn0DkMj8BA70lrKtcdMIEeXNO1lBQIHSM63OaVDK4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=foNPH9Ci; arc=none smtp.client-ip=209.127.230.111 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="foNPH9Ci" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=2212171451; d=bytedance.com; t=1787738463; h=from:subject: mime-version:from:date:message-id:subject:to:cc:reply-to:content-type: mime-version:in-reply-to:message-id; bh=ebjhkF6zcd6itLAobqPH6pMiOaXkDg99ledLDKWzank=; b=foNPH9CiOJ5mOlyoVHX9t0O7sz6/ifnVgJvZTD6kHGk1SsRzgWjzS4CuSYxdOoQdRl3YCU LsOyKfuyRQ5/1teMcMr5CnXObfOGuxsFqa0yukFkIDIfcn+XDlUXJBOQIOGU02tjG/GFvN gHIa8MrWBrX3bk4nbmvBO2dXZ9yk0UnHul2pIWOIOn7yqwkzke0jWrs0kWupVJUGpMLx4U hlKA5R8uCZbhw9Z2nhyNVTAtDSm1E71D5AUR6DNnRLrtFObm3KJnz8esxrE5oCHsfpDH3Z B69/+eUUVPC8Tcri5/ZjFYIPXVYgxrLbvrIAz1gYTGeWjG4MMqLB5sdhn1baDw== Subject: Re: [Question] Userspace throttling + "sched/fair: Combine detach into dequeue when migrating task" causes guest boot hang Date: Wed, 26 Aug 2026 18:00:09 +0800 Cc: "Ingo Molnar" , "Peter Zijlstra" , "Juri Lelli" , "Vincent Guittot" , "Dietmar Eggemann" , "Steven Rostedt" , "Ben Segall" , "Mel Gorman" , "Valentin Schneider" , "K Prateek Nayak" , From: "Aaron Lu" Message-Id: <20260826100009.GA3616635@bytedance.com> X-Original-From: Aaron Lu X-Lms-Return-Path: Content-Transfer-Encoding: quoted-printable References: <20260825120629.2472938-1-chenjinghuang2@huawei.com> To: "Chen Jinghuang" Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Disposition: inline In-Reply-To: <20260825120629.2472938-1-chenjinghuang2@huawei.com> Content-Type: text/plain; charset=UTF-8 On Tue, Aug 25, 2026 at 12:06:29PM +0000, Chen Jinghuang wrote: > Hi, I'm seeing a VM boot hang on mainline, and I'd like to understand the > interaction between userspace throttling and the scheduler patch > "e1f078f50478 sched/fair: Combine detach into dequeue when migrating > task". >=20 > Host: > aarch64, 96 CPUs (0-95), 4 NUMA nodes: > node0: 0-23, node1: 24-47, node2: 48-71, node3: 72-95 > Mainline kernel tag: 7.2-rc1. >=20 > Guest(libvirt/KVM) - described in words: > An aarch64 (virt-6.2) UEFI VM launched with `virsh create`; key config: >=20 > - 128 vCPUs (statically placed, oversubscribed =E2=80=94 the host has onl= y 96 > physical CPUs). > - host-passthrough CPU model; GICv3; 64 GiB RAM. > - has set to 400000; all and > entries are commented out, so there is no vCPU pinning. > - Storage: qcow2 on virtio-scsi (cache=3Dnone, io=3Dnative). HPET disable= d. >=20 > Userspace throttling: > The VM runs under a CPU-quota cap applied on the host. The actual values > from the cgroup controller are: >=20 > cpu.cfs_period_us =3D 100000 > cpu.cfs_quota_us =3D 400000 >=20 > I also found that if I set cpu.cfs_quota_us to -1, or enlarge it beyond a > certain point, the guest boots fine. >=20 > Symptom: > The guest hangs at some command early in boot and never reaches the login > prompt. >=20 I tried this on an x86 machine with v7.2-rc1 kernel and with quota set to 4 cpus, the VM booted fine; when I further reduced quota to 1 cpu, the guest kernel would dump a ton of soft lockups during boot. I also tried running an old 5.10 kernel(which doesn't have per-task throttle) and it behaved the same as v7.2-rc1. The x86 machine has 64cores/128cpus and the VM I created has 128cpus and 128G memory. > Observations: > Only reverting both of the following together makes it boot (neither one > alone suffices): >=20 > 1. The kernel patch for userspace throttling. > 2. The scheduler patch: > e1f078f50478 ("sched/fair: Combine detach into dequeue when migrating > task") I'm curious how you found e1f078f50478, just because it touched pelt? >=20 > Reverting only one of them still hangs; reverting both together boots fin= e. >=20 On top of v7.2-rc1, right? > Question: > I don't fully understand how these two interact. My rough guess: e1f078f5= 0478 > ("sched/fair: Combine detach into dequeue when migrating task") affects t= he > PELT accounting, and the userspace throttling also has logic that affects= PELT > accounting. When both are combined, load balancing and subsequent schedul= ing > behavior may end up misbehaving, stalling the guest. Is the host busy? If the host has many idle cpus, even the pelt is wrecked(which I doubt), it should not cause the qemu task being starved. The PELT accounting matters when tasks have to compet the same CPU, but if your host system has many idle cpus, that should not happen. And from the log you posted for the cpu usage, it appears that task group is getting cpu time.