From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f47.google.com (mail-pj1-f47.google.com [209.85.216.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE9022628D for ; Mon, 6 Apr 2026 01:46:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.47 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775439987; cv=none; b=MAn03zBizbBBhfC/oLUNLcBFe5djONXAcgHWaMpdsPNJRa46SVsiblfpWMGl+xvQDmFXXn9E4UoChW5fq6lQxhNPqCuhT2gMxdMQPB/Tea0+ehre4MfwYaMGJF9M4hLIqNjxSuSlnxqjrriMu34sJ31V/uluqXH2L89wbsWYB0I= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775439987; c=relaxed/simple; bh=QTXrP1pvQw71VJBE+VmNyw4t82T4fe9nzuUGD2PSgUI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=J0dbtc7GGIri4hJOjgWmP8m6TipkiYRaROGNKhugQFB46E1G2ZiEhW3DpAXDNIes4d6Z9ePD0iJE1JNtJJF2WsXedadOy77MV2/KBs9zfbaPxBuZ+apTEcS1/jzr1ozdjJgI7rLbUh2BimwzXbIu3Aw4TO/2yDJdfz0o2UAYP4s= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=QLHm0pA8; arc=none smtp.client-ip=209.85.216.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="QLHm0pA8" Received: by mail-pj1-f47.google.com with SMTP id 98e67ed59e1d1-35d9f68d011so2185272a91.2 for ; Sun, 05 Apr 2026 18:46:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1775439985; x=1776044785; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=tzEp5QH/ekVv892eabdZ8H0/BwVcTHUkE7wpU0Cr564=; b=QLHm0pA8d3zEKGtRkZp+IF0Ab+N4xvYRNbQXhjZbBDBa3LSU+IfVQz4OylswXVM0P/ ksXrbjIF8qHABClKFefWWZYlPy8HJVfFNVrn0veYoMaL3K7NybCDaB9S233fTAXBqGq+ +urPvMsV0Mmu7Np+/BSP9/B9YWCPeNCiO7EersW1eQvaGhGe8HUABFRb8HI8GZ8PG33h 1kDoBXalB7vYcZv7vEdvDgXAcDB6Tild4BEUGlh3oua30zeah4f9OtI6YzLRDAY4vxG9 tJqvAo79+vDhhoTU0IB9pONvWfrej2PR+E1W2yuZb2Farn3bup9dOyHusAyUfnLAdI3+ MSzQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1775439985; x=1776044785; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=tzEp5QH/ekVv892eabdZ8H0/BwVcTHUkE7wpU0Cr564=; b=o/Igtup2NsDDUAvws+lFcwSQxaX/i/2C6zrdS1vx/Cb1Qgun5PYd3GB3Urwydg22Gb dX2BFyzuyERASfS6LVj7oCgeP/9srJVP7KGbKBjw4+MMp9ri8oWsA1r2y/x1rOdWlcrk 9D5Z0HEznFNwlI68nFUHIOHMlEibfm8+loMy4xyR9+qYPC1MgPSzoKA9Id1OapQgIrp3 UT0RAItuVjuTDgHkIcIZ0I57oYoqiaOoWd+A8oCXWNLKyoAWHmZ0m7P9QHtBUR5fKWaR BS3DZ/i6fqoL49JTMTg6nNRPbX5pZnyuoyZOy4DltEQuEbJnPJtvW4okvEd4GDazwB0v 0Tvw== X-Gm-Message-State: AOJu0Yw77Gb2gAuC0S2JadDjkFYrBqwnAcWQo9cheaOIfsqfSazuDin7 qgJxngW9p7gAgEVMqSiyzm4MGCGqqlQfICSZ5wZRQaJiPk4x1fw2NkMO X-Gm-Gg: AeBDiessFZgeokj1r2ZwLYTlYJYq5HKZe8Ryzo8zxoNn4Smjc5mxRr/6v8CmJKQ8zy0 M+p4+nlbj7STGIMUDTZg1R5xDkE58igzg+eaQJghGIgJ5pwcCNmwN5HbpwTKvF7MzmXxj7KJyLo pq9soW1XM8bPmt6KizrPlkz8ch6IXZVnJLC7gRBIRaXzta0dJhwGy56/B/9bTylnaGljCON6fh6 JsCTj+7D24MJojlMYb4xxi+E3iHk+UFbcsTmFq5lUO4yEWIXlz7s2eKbmswNsGARezAfkOe1aEV yfncO81JjuQLEOPkXInUOEF4wRUdvXwTHq7kxztyvZcChhVLAxQaDwzXkEvZa/2JpdjmilzyK4P WY5fSxCN8h+Nj59q+zEWDR7pNkryFY3C26M6mCDC562/W+CB7lEndXHsenmPs4NGFm0tsfuxUFZ dv4i+Ojvab9xv6zvKJNeBGcAH0ti7HROn/sHWVPq204OHb3qdaHTW9Iyai0TqXgSyKK/uL9pdbS MjkYQ== X-Received: by 2002:a17:90b:3809:b0:35d:9560:3efc with SMTP id 98e67ed59e1d1-35de68ce84bmr9254190a91.14.1775439985103; Sun, 05 Apr 2026 18:46:25 -0700 (PDT) Received: from localhost.localdomain ([2404:7a85:2900:3f0:6514:327:b8bf:af2d]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-35dbe624756sm21458032a91.5.2026.04.05.18.46.23 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sun, 05 Apr 2026 18:46:24 -0700 (PDT) From: Mitsumasa KONDO To: andres@anarazel.de Cc: linux-kernel@vger.kernel.org, dipiets@amazon.it, peterz@infradead.org, tglx@kernel.org, kondo.mitsumasa@gmail.com Subject: Re: [PATCH 0/1] sched: Restore PREEMPT_NONE as default Date: Mon, 6 Apr 2026 10:46:21 +0900 Message-ID: <20260406014621.38487-1-kondo.mitsumasa@gmail.com> X-Mailer: git-send-email 2.49.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Hi Andres, Thank you for testing this. On 2026-04-06, Andres Freund wrote: > It's not sustained, the spinning just lasts a between 10 and 1000 > iterations, after that there's randomized exponential backoff using > nanosleep. > Which actually will happen after a smaller number of cycles of with a > shorter SPIN_DELAY. > If I remove the rep nop on x86-64, the performance of the 4kB pages > workload is basically unaffected, even with PREEMPT_LAZY. The fact that removing rep nop made no difference suggests that the spinlock is not the bottleneck in your environment. Could you share your storage configuration? Salvatore's setup uses 12x 1TB AWS io2 at 32000 IOPS each (384K IOPS total in RAID0), which effectively eliminates WAL fsync as a bottleneck. In a storage-limited environment, changes to spin delay behavior would naturally be invisible because throughput is capped by I/O before spinlock contention becomes material. Also worth noting: Salvatore's environment is an EC2 instance (m8g.24xlarge), not bare metal. Hypervisor-level vCPU scheduling adds another layer on top of PREEMPT_LAZY -- a lock holder can be descheduled not only by the kernel scheduler but also by the hypervisor, and the guest kernel has no visibility into this. This could amplify the regression in ways that are not reproducible on bare-metal systems, regardless of architecture. If you want to isolate the effect of SPIN_DELAY on throughput under PREEMPT_LAZY, I would suggest: 1. Use synchronous_commit = off or unlogged tables to remove I/O from the critical path entirely. 2. Use a read-only workload (pgbench -S) with shared_buffers sized to force buffer eviction contention. 3. Run on a high-core-count system with all CPUs saturated under PREEMPT_LAZY. This should expose the pure impact of spin loop behavior without I/O or WAL masking the results. > The spinning helps with workloads that are contended for very short > amounts of time. But that's not the case in this workload without > huge pages, instead of low 10s of cycles, we regularly spend a few > orders of magnitude more cycles holding the lock. I agree that the 4kB page / huge page difference is significant. But even when individual spin durations are short, the cumulative effect across hundreds of backends matters. Small per-iteration overhead in the spin loop, multiplied by high concurrency, can add up to measurable throughput loss -- the effect that becomes visible only when I/O is not the dominant bottleneck. Regards, -- Mitsumasa KONDO NTT Software Innovation Center