From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 57C0DC76196 for ; Mon, 10 Apr 2023 08:23:25 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 9A1236B007B; Mon, 10 Apr 2023 04:23:24 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 95216280008; Mon, 10 Apr 2023 04:23:24 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 840BE280002; Mon, 10 Apr 2023 04:23:24 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0014.hostedemail.com [216.40.44.14]) by kanga.kvack.org (Postfix) with ESMTP id 764E36B007B for ; Mon, 10 Apr 2023 04:23:24 -0400 (EDT) Received: from smtpin30.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay06.hostedemail.com (Postfix) with ESMTP id 3A39CABF52 for ; Mon, 10 Apr 2023 08:23:24 +0000 (UTC) X-FDA: 80664791928.30.9D67CFE Received: from mail78-36.sinamail.sina.com.cn (mail78-36.sinamail.sina.com.cn [219.142.78.36]) by imf03.hostedemail.com (Postfix) with ESMTP id DB9FB20019 for ; Mon, 10 Apr 2023 08:23:20 +0000 (UTC) Authentication-Results: imf03.hostedemail.com; dkim=none; dmarc=none; spf=pass (imf03.hostedemail.com: domain of hdanton@sina.com designates 219.142.78.36 as permitted sender) smtp.mailfrom=hdanton@sina.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1681115002; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=yTFV2Wy4ewVI6KVnXXZitg4K0O2ZPw+D09t1v0i0bVc=; b=sVac9nZ4E9zlyX76Wb3jnIA21D5NtxME7+z02XLFexIgNR4gpmQt5Co3v/3x5WGpSmh6dH 0YqbjLEOxhJTwZujYoALhRySnYlgSbVn3YXjeYdnf7nzzFlvVM4nh8C3e++KEKO22WhgVS rgCQlOGhh1uvDLPtmG2s3L2Vf80kanY= ARC-Authentication-Results: i=1; imf03.hostedemail.com; dkim=none; dmarc=none; spf=pass (imf03.hostedemail.com: domain of hdanton@sina.com designates 219.142.78.36 as permitted sender) smtp.mailfrom=hdanton@sina.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1681115002; a=rsa-sha256; cv=none; b=6XYCOncyqRD/tQbgCHDnZ3pHveSEti2Ciadtca3LqLJlzUzOq2HKkXlYuBelqePPWJ59M8 dP9RC9LRVH35hWpxAOpVx0aW0F9S9A/BgDJSWuFP5zpeuQtUOHa4qOA3dfWUn9/+j6WKcP 4IWwzz964CnYOyh3te5QkBkZCUsKbas= X-SMAIL-HELO: localhost.localdomain Received: from 106.23.202.1.static.bjtelecom.net (HELO localhost.localdomain)([1.202.23.106]) by sina.com (172.16.235.25) with ESMTP id 6433C77100000565; Mon, 10 Apr 2023 16:23:16 +0800 (CST) X-Sender: hdanton@sina.com X-Auth-ID: hdanton@sina.com X-SMAIL-MID: 56788034210207 From: Hillf Danton To: David Vernet Cc: Peter Zijlstra , mingo@kernel.org, vincent.guittot@linaro.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, mgorman@suse.de, bristot@redhat.com, corbet@lwn.net, kprateek.nayak@amd.com, youssefesmat@chromium.org, joel@joelfernandes.org, efault@gmx.de Subject: Re: [PATCH 00/17] sched: EEVDF using latency-nice Date: Mon, 10 Apr 2023 16:23:07 +0800 Message-Id: <20230410082307.1327-1-hdanton@sina.com> In-Reply-To: <20230410031350.GA49280@maniforge> References: <20230328092622.062917921@infradead.org> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Rspamd-Server: rspam05 X-Rspamd-Queue-Id: DB9FB20019 X-Stat-Signature: onj4xdg6afxp4r8pb7y1fej947jdboep X-Rspam-User: X-HE-Tag: 1681115000-979115 X-HE-Meta: U2FsdGVkX1/zfbNppCO1IImyFEtoNgAoQGj7sgTXxR0DDvQYs8YzZNEy9hctADd5E4IL7/lH9gbYvW+T7dmV7GfWZa2EH7Tyb0HPPWS7fmNQC+XCC+AOPCJ2X0l/pkJxz21vZB4L8FlyS0cbexxq/c5/GJPzrrePuUCt5xnEpKJOp8XE24sjirX3lqfVvime/hikQ13Ng5oSULHUzyOLsdTWYCIJIXk3f60Vh1m+hn1lADsDSFc2JqUorwyCZounLs8oq6hWgb5fZWhlZqLodXYKQeJCwb4saOc5Gbk0UTkUNaHJDYlRDpLGaP+DxmeanotPUsn3c9jHtI0ptrewkgRLPQJpNYDEDR+PVUCy+ZDgAdFO0DRoti+7zBa9wBRUfXjlIBkpPD3xkoda0/42i6aiiZzi8rtxQ+yQS+b7Hr/3+8jzT0s+zKmmvejzm8oErIX2J0b/NCzfz8b3b0gMUrF1VB4jHJtypU/gRkk1gUzRinEGTeucnxaAVB7WRGObCM3UKehLVZlKYEmTD/mDSEvIZEizbrDW4o56tIcNvDfB1SE/xCPuINnIjkWOkDtvhD6YLNsLtXLDwdim0qKppSwfb2RmcQuOl07Wl/h4y/A9u2plErKNlzbP1r1nRNChJDUsNR3c3cNgp3Kkhi+n8xZDC0qD9MVkJy9GlzbKEw7L9pm74+kmLI/V3XcFf3W9hhTeerUXfDTz0Z5kuLUwALm/ttI4tJEM9ux3hly8kiAZjKo4iENbkJqrqCnyiJseeYU6wPlrDOslVUrvttL1kliG648jjeCDswHqoOFYo7Y05beD9gJH8Rg2eEQhu6VYjzkz6tU2+8v/iGODR7xfIswY8zEbBh7Z66pQmyt1eLXeh+IbBFHl6rlTNqBV6QypCYGJWuH2g1Jw69CjiaFIA3shMEe+tNSM3NnULFjbSrCU1fa1pLWMJONPrK276bf0/ORKRnNlL6Mc54S1s3n 5SZTE9Pj EtXd9EOZBuTVwMERmLmn28fqXNeJ0bjDLrp9axK4YAHUnTnRm9tCKc+lZhwDKGh7Da6gnXZt24jga4qZ55O5pbNuveWVWresbORt+fdV4/X6GDOKOtq4eCaNXGBm90tVQ/4W5lvBgX89x1qZOYvxe48lxOmZCWdyvnp9Z X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: On 9 Apr 2023 22:13:50 -0500 David Vernet > > Hi Peter, > > I used the EEVDF scheduler to run workloads on one of Meta's largest > services (our main HHVM web server), and I wanted to share my > observations with you. Thanks for your testing. > > 3. Low latency + long slice are not mutually exclusive for us > > An interesting quality of web workloads running JIT engines is that they > require both low-latency, and long slices on the CPU. The reason we need > the tasks to be low latency is they're on the critical path for > servicing web requests (for most of their runtime, at least), and the > reasons we need them to have long slices are enumerated above -- they > thrash the icache / DSB / iTLB, more aggressive context switching causes > us to thrash on paging from disk, and in general, these tasks are on the > critical path for servicing web requests and we want to encourage them > to run to completion. > > This causes EEVDF to perform poorly for workloads with these > characteristics. If we decrease latency nice for our web workers then Take a look at the diff below. > they'll have lower latency, but only because their slices are smaller. > This in turn causes the increase in context switches, which causes the > thrashing described above. > > Worth noting -- I did try and increase the default base slice length by > setting sysctl_sched_base_slice to 35ms, and these were the results: > > With EEVDF slice 35ms and latency_nice 0 > ---------------------------------------- > - .5 - 2.25% drop in throughput > - 2.5 - 4.5% increase in p95 latencies > - 2.5 - 5.25% increase in p99 latencies > - Context switch per minute increase: 9.5 - 12.4% > - Involuntary context switch increase: ~320 - 330% > - Major fault delta: -3.6% to 37.6% > - IPC decrease .5 - .9% > > With EEVDF slice 35ms and latency_nice -8 for web workers > --------------------------------------------------------- > - .5 - 2.5% drop in throughput > - 1.7 - 4.75% increase in p95 latencies > - 2.5 - 5% increase in p99 latencies > - Context switch per minute increase: 10.5 - 15% > - Involuntary context switch increase: ~327 - 350% > - Major fault delta: -1% to 45% > - IPC decrease .4 - 1.1% > > I was expecting the increase in context switches and involuntary context > switches to be lower what than they ended up being with the increased > default slice length. Regardless, it still seems to tell a relatively > consistent story with the numbers from above. The improvement in IPC is > expected, though also less improved than I was anticipating (presumably > due to the still-high context switch rate). There were also fewer major > faults per minute compared to runs with a shorter default slice. > > Note that even if increasing the slice length did cause fewer context > switches and major faults, I still expect that it would hurt throughput > and latency for HHVM given that when latency-nicer tasks are eventually > given the CPU, the web workers will have to wait around for longer than > we'd like for those tasks to burn through their longer slices. > > In summary, I must admit that this patch set makes me a bit nervous. > Speaking for Meta at least, the patch set in its current form exceeds > the performance regressions (generally < .5% at the very most) that > we're able to tolerate in production. More broadly, it will certainly > cause us to have to carefully consider how it affects our model for > server capacity. > > Thanks, > David > In order to only narrow down the poor performance reported, make a tradeoff between runtime and latency simply by restoring sysctl_sched_min_granularity at tick preempt, given the known order on the runqueue. --- x/kernel/sched/fair.c +++ y/kernel/sched/fair.c @@ -5172,6 +5172,12 @@ dequeue_entity(struct cfs_rq *cfs_rq, st static void check_preempt_tick(struct cfs_rq *cfs_rq, struct sched_entity *curr) { + unsigned int sysctl_sched_latency = 1000000ULL; + unsigned long delta_exec; + + delta_exec = curr->sum_exec_runtime - curr->prev_sum_exec_runtime; + if (delta_exec < sysctl_sched_latency) + return; if (pick_eevdf(cfs_rq) != curr) { resched_curr(rq_of(cfs_rq)); /*