From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail3-165.sinamail.sina.com.cn (mail3-165.sinamail.sina.com.cn [202.108.3.165]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6A6893EE1DB for ; Fri, 7 Aug 2026 05:18:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.108.3.165 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786079899; cv=none; b=T726N2xR60BX6GSDdWy+cUNNLvlGT7ts0dTdmQWRDnvnCtXkmS570xSUmaaxrVHKKHKq0QqfgboFMU1Wp6vQlH8BzlB1rL35AFi+XE3trTBfHRkL1IS5AfPxBISmGBK9lZ5+eHLtVA5O1/zgTEe3CtU601LI64n3xN+yJmRtgEQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786079899; c=relaxed/simple; bh=vhhv7++YMC1qYhuhyPUcpqwCsymqvGzdE4IrKjIXHWs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=mXMHZMUHxnsE86ZZrbLD3UJ+9RP10v5IvBRarDb2BZAdsKWOUydu2QcaCqHvN9IlFSfcDd9PbHvkdwYFVUg+Nugc3Q/yjuGev2f7fa9oemz3Cvb+ygkm9cmyuTYT70+sWqZwZ8uO1DKxmy/81z1CNNJBopeub3PzyZag9zR9C/g= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sina.com; spf=pass smtp.mailfrom=sina.com; dkim=pass (1024-bit key) header.d=sina.com header.i=@sina.com header.b=dTgRjiKx; arc=none smtp.client-ip=202.108.3.165 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=sina.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sina.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=sina.com header.i=@sina.com header.b="dTgRjiKx" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=sina.com; s=201208; t=1786079891; bh=f0bfWVqT+bcAH4fkOsT9V7cn7uO/spMt5sfzfCX6oIY=; h=From:Subject:Date:Message-ID; b=dTgRjiKxzbkdUUuxW12lCqjdgDutN6YfxB1cy7J3ezgjSaBaPdDgMs62ynHeTLXVz 6rYmckRk5qevWjEYcWFv4s2LzjWLdGoYeiEuuLhE6nf3PASQRAWEprqWtwWq+F4sJE FkjbDMF6VgNo9AtSnHWbE8Gjgu4zCklqtLjWJHj8= X-SMAIL-HELO: localhost.localdomain Received: from unknown (HELO localhost.localdomain)([221.216.154.155]) by sina.com (10.54.253.33) with ESMTP id 6A756A88000076C5; Fri, 7 Aug 2026 13:18:03 +0800 (CST) X-Sender: hdanton@sina.com X-Auth-ID: hdanton@sina.com Authentication-Results: sina.com; spf=none smtp.mailfrom=hdanton@sina.com; dkim=none header.i=none; dmarc=none action=none header.from=hdanton@sina.com X-SMAIL-MID: 5263806685307 X-SMAIL-UIID: 2F79B77C788F4AD986EDD6918A561BE5-20260807-131803-1 From: Hillf Danton To: Vinicius Costa Gomes Cc: Peter Zijlstra , K Prateek Nayak , Christoph Lameter , linux-kernel@vger.kernel.org Subject: Re: [PATCH RFC] sched/fair: decline WF_SYNC stacking when waker LLC is the busier share Date: Fri, 7 Aug 2026 13:17:48 +0800 Message-ID: <20260807051750.1018-1-hdanton@sina.com> In-Reply-To: <87cxvu25sp.fsf@intel.com> References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit On Thu, 06 Aug 2026 16:22:14 -0700 Vinicius Costa Gomes wrote: >Hillf Danton writes: >> On Thu, 06 Aug 2026 10:44:18 -0700 Vinicius Costa Gomes wrote: >>>Hillf Danton writes: >>>> On Tue, 04 Aug 2026 16:14:05 -0700 Vinicius Costa Gomes wrote: >>>>> Since commit 900bbaae67e9 ("epoll: Add synchronous wakeup support for >>>>> ep_poll_callback"), epoll driven WF_SYNC wakeups have been "too >>>>> strong" and could cause tasks to stack on a busy NUMA node while other >>>>> nodes are relatively idle. >>>>> >>>>> As commit 900bbaae67e9 ("epoll: Add synchronous wakeup support for >>>>> ep_poll_callback") improves real workloads a revert is not the answer. >>>>> The fix is to make the WF_SYNC "stack on waker" shortcut take into >>>>> account the load on this and prev's LLC, rejecting the shortcut only >>>>> when the waker (this) LLC is fully loaded and prev's LLC is less >>>>> loaded than the waker's. >>>>> >>>>> Signed-off-by: Vinicius Costa Gomes >>>>> --- >>>>> We received a report of a regression on a openresty based >>>>> workload (the main metric being tail latencies) on a CWF SNC3 single >>>>> socket system, the main symptom that we could measure was one node >>>>> being overloaded while the other nodes were relatively idle. >>>>> >>>>> Further investigation showed that spreading the NIC RX interrupts over >>>>> all NUMA nodes helped. Reverting commit 900bbaae67e9 ("epoll: Add >>>>> synchronous wakeup support for ep_poll_callback") also helped. >>>>> >>>> The irq approach is prefered because anything that gets the eevdf offloaded >>>> is good, you see it is near to the knowall point, needless to say that they >>>> lie in different layers and from the scheduling-cpu pov irq is a gray rhino >>>> in the room while WF_SYNC is a tiger mosquito in the corner at best in your >>>> case where the mosquito failed to understand your workload. >>> >>> I don't think irq spreading across NUMA nodes is that good of an idea on >>> low loads/by default, as it loses the locality that the kernel (even >>> with irqbalance) try to maintain. I used that as a hackish way of >>> testing "if I spread tasks, does it improve the tail latencies?". >>> >> Can you specify the root cause of the tail latencies, particularly after >> spending two minutes thinking if changing config in user space to solve the >> issue is a light year better than adding a couple lines of code in the >> wakeup path? >> > On a machine running the customer workload, I ran a bpftrace script that > tracks the time "openresty" tasks stay on the runqueue waiting to be run. > > (this was a using an earlier version of the patch, on top of v7.1-rc7, I > don't have easy access to the machine anymore) > > Before: > > @runq_us: > [0] 5376 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| > [1] 5122 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ | > [2, 4) 5073 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ | > [4, 8) 3821 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ | > [8, 16) 1649 |@@@@@@@@@@@@@@@ | > [16, 32) 747 |@@@@@@@ | > [32, 64) 552 |@@@@@ | > [64, 128) 668 |@@@@@@ | > [128, 256) 907 |@@@@@@@@ | > [256, 512) 1235 |@@@@@@@@@@@ | > [512, 1K) 1946 |@@@@@@@@@@@@@@@@@@ | > [1K, 2K) 1832 |@@@@@@@@@@@@@@@@@ | > [2K, 4K) 1130 |@@@@@@@@@@ | > [4K, 8K) 561 |@@@@@ | > [8K, 16K) 340 |@@@ | > [16K, 32K) 250 |@@ | > [32K, 64K) 108 |@ | > [64K, 128K) 33 | | > [128K, 256K) 0 | | > [256K, 512K) 0 | | > [512K, 1M) 1 | | > > After: > > @runq_us: > [0] 1781353 |@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@| > [1] 450223 |@@@@@@@@@@@@@ | > [2, 4) 71505 |@@ | > [4, 8) 21345 | | > [8, 16) 22406 | | > [16, 32) 14792 | | > [32, 64) 7336 | | > [64, 128) 3867 | | > [128, 256) 4607 | | > [256, 512) 6886 | | > [512, 1K) 5862 | | > [1K, 2K) 2829 | | > [2K, 4K) 876 | | > [4K, 8K) 11 | | > > This made me think that the almost unconditional stacking shortcut that > WF_SYNC promotes was causing tasks to wait on already busy CPUs, while > there were idle CPUs around. > Then like the line-speed ether switch below, the CPU cycles per tick, CPT, instead of LLC is the critical point, and increasing CPT is the correct pill, no? What is not unusual from Monday to Friday is code is added in kernel for mis-configured boxes. > That was as close to a root cause that I got. > >>> Note that the "local-only reproducer" workload (memcached + >>> memtier_benchmark) runs over loopback (no NIC irqs here), I pin memtier >>> (the client) to one NUMA node, leave the server unpinned and I am able >>> to reproduce the issue: with the RFC patch the tail latencies reduce by >>> 2-3x. (on the customer workload the impact is even higher) >>> >>> My expectation was that the scheduler would be able to say: "ugh, even >>> though respecting WF_SYNC is good most of the cases, this isn't one of >> >> No comment before root cause specified, even if I suspect it sounds like >> that two CPU cores could not provide line speed 24-port 1000MB ether switch >> before 2010 while 12 cores could. >> >>> them". And this is the spirit of the RFC, showing that those cases >>> exist and their impact. >>> >>> >>>Cheers, >>>-- >>>Vinicius