From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qv1-f53.google.com (mail-qv1-f53.google.com [209.85.219.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6B99B380FCC for ; Fri, 24 Jul 2026 18:07:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784916462; cv=none; b=dMw3f41MhKLzXztHPPMqJU/Xj+ckXHc1ScViN10IfEKOaw4DMJUq3FdyB0ooYk68DZDtmisszEaxuTPdEN1eZUejrYBRfwoMpiUmhP2+mD3OuJW+pW7KNtXrizwDQsldosOHjbxCN/kyWiOQyqtd/M6hUqRGt1i+5McPTtPZaQI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784916462; c=relaxed/simple; bh=tmYmyf97sfoiQLEEAY36DBZCB6pTgGKz0lACx3VRkJg=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=neJTxBAZUmYDnBg907dgYZUm6YIxw/LVJYsrRikUUv5Tg0Jw6v/cEgGq4zwfdWMWEOh5Cf+/Fx4OHauLmPRI+hCAqiNo7oAD23d3DToDEb0CjcGD8OX6ASMXiz9N4JS4aNFmHXs5ADIFokmNE6KFTkzoWDNMDMuLzp3p35UKFiw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=j8sKTQrG; arc=none smtp.client-ip=209.85.219.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="j8sKTQrG" Received: by mail-qv1-f53.google.com with SMTP id 6a1803df08f44-8ff88549786so7768946d6.3 for ; Fri, 24 Jul 2026 11:07:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1784916456; x=1785521256; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=nJ7PDP3fl+p5a28n/jRHcjrCN1+1Z9szAdk1WdxxOqM=; b=j8sKTQrG8FpVgQbzYha5FTO6MfIaiOFpy9ieq75ImgIAWE8MM8/Rzh6NwvqgK8s+VA m7sUKywuFLifToJCmXsy4GvlOdYlibNckEKAeCLsTqt3ErNrCmLAOo3A4uVSr7kwsLxz LxevfLSG8jS36uXFUK8XvLGBE4S3NjauMS+4yptzOJldGu4S7MvS0ElvV6mR8Vwjj5Gg KM7e9CIqCwALDZ6TUwTWHZcVN9qjzMZPxx/B6R30wYZCCy0JRcDpaCeoJCVVBKafUybM J/l32Btdq9EV46ajqgQlzZyqmE1HIk98qFpFVzbwaIhhLG401Jz/QXVbZQkkQglARB4k AUTQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784916456; x=1785521256; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=nJ7PDP3fl+p5a28n/jRHcjrCN1+1Z9szAdk1WdxxOqM=; b=SqK3VH0skvHgj0/QLwQ/z+Hbbb60oHu4bgu9WSiIVlGvHgb66UGhNeqvKC4CIdlPn7 IcOJ12aHX6lRwvJRoaSSd1jdqV/qzdeQ0ZNdAdtjN1EalTqYvIJgzIE/MtpZ1V7p1h8i CABUvo/h+Fd3K/bcdj6E61t5VQqdKBIjCKsaTwZ67qd5Gs9OnCbe5V//JhrwKS5PhDKR FB4MMzlfdy8LTGWvTKVwKBqYN313jJ65rXeg5Htqq/t2b7Ze+ClyqMv+VcyQ4sg8LiJ9 lrxkliR5kYf2dRId1NHFEuBQed7IxuiOXtwYOxHxFj5YIXWsu7OidI4foiu/ajyNIEkg ySew== X-Forwarded-Encrypted: i=1; AHgh+Ron6Q0kWaJV+aOk1BDZqsZr3i4NEILLDKdyoPHZCNWXR9HvULDfkISD+z4wYUooBiUy1t704PcOfNntDZo=@vger.kernel.org X-Gm-Message-State: AOJu0YxZXMYBW6IzcMdsddg6vJZ0ajPzM+IcczA1H4CcofUIhq6BGO9I NMrX2bJHb4ti9L5/nXfPgQIEzMzaccIMc3FToMOEC+LjKOaed9JIyTzbAkf6fO59y7A= X-Gm-Gg: AR+sD10nBWwSQJp/X8pF+DKyrqytAlZXnK3FHFQzg/RnNP/hOQLmlEOD128fWWJ+e8t 8c5F9LvQcVwnzZ0KCuO7Apn1hVpKn6UKylTjwq0agGPWIt76JwIaCLj5NBYv+hW4eHzIUFvi310 VZcVMEQ8NmA93PAw6UQc6N6t+TRMLBQIZXfd9wvkr44WtXqDsBypwg1vO6pjDQhrwuwhUJmQsWF ZTVBjYqmiI6+3vpseB0hTAi2HBMs+GIM6cqi1RIJR0zQ1iTjpWcUnR+yYem94MDWThFtA3vZRma 2E2qf91ks9XQ7HUkmzLgP1I4Qc8dD5FtFdTCjtd82ggHuo6r/JkiqHcRYnWn5dBdLzrNBNESnsy 3FKc31nsjdzVKEQLy8mkXEwBVgJZln6lqVY62D3uPPMFluU479je2dnf4ilF9RigZQgoMltgouR kcn62j5bCatI0= X-Received: by 2002:a05:6214:5c01:b0:907:562d:e3 with SMTP id 6a1803df08f44-907ca33efbbmr100210026d6.39.1784916456394; Fri, 24 Jul 2026 11:07:36 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-907e854feadsm3525256d6.15.2026.07.24.11.07.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 11:07:35 -0700 (PDT) Date: Fri, 24 Jul 2026 14:07:31 -0400 From: Johannes Weiner To: Brendan Jackman Cc: Andrew Morton , Vlastimil Babka , Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Shakeel Butt , linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: Re: [PATCH v2 4/4] mm: page_alloc: fix non-movable reclaim storm in defrag_mode Message-ID: References: <20260722150006.3848560-1-hannes@cmpxchg.org> <20260722150006.3848560-5-hannes@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Fri, Jul 24, 2026 at 03:09:49PM +0000, Brendan Jackman wrote: > On Wed Jul 22, 2026 at 2:56 PM UTC, Johannes Weiner wrote: > > As we deployed defrag_mode into Meta production, pressure spikes and > > excessive swapping were observed on some workloads. Tracing confirmed > > that this is unmovable/reclaimable requests spinning in the allocator > > and direct reclaim, causing excessive amounts of swap. > > > > The initial plan for defrag_mode was to rely on kswapd/kcompactd to > > produce blocks, and if those are overwhelmed under high pressure, let > > the allocator fall back (__rmqueue_steal()) after its retry loops. > > However, that retrying results in more reclaim on some of these > > workloads than we'd hoped, sometimes excessively so, spurred on by the > > !costly order conditions in should_reclaim_retry(). > > > > The storms are dependent on the request type. Reclaim will inevitably > > make room in existing movable blocks, since that's where the LRU pages > > live. So if movable requests retry on reclaim, they make progress. > > > > When non-movable requests spin in reclaim that isn't productive. They > > cannot use the individually freed pages, and the process is unlikely > > to accidentally free whole blocks to meet the ALLOC_NOFRAGMENT bar. > > They spin and overreclaim excessively, which tanks performance and > > triggers userspace guards like swap exhaustion or pressure based OOM. > > > > To fix this, send non-movable requests, regardless of order, into > > pageblock reclaim/compaction. This way, they help move things along to > > meet the ALLOC_NOFRAGMENT bar. After this patch, the reclaim storms > > and excess OOM rates are no longer observed in production. > > > > The longer-term plan is still to have all requests, including the > > movable ones, help make blocks to spread the cost of defragmenting > > more evenly and fairly; combined with proper watermarking to reduce > > allocation latencies in the common case. However, doing this naively > > unearths scaling and concurrency limitations in compaction that need > > to be addressed first. Promoting just non-movables for now is the > > minimally viable bug fix for the above issue. > > Please forgive the noob question, I'm still struggling to get a really > good mental handle on this stuff. But what about compact_first here? > When we're promoting the order for compaction does it also make sense to > promote compaction itself? My understanding is that this is just an optimization for requests that are more likely shape-limited (block contiguity) rather than capacity-limited (free order-0). In that case we do a quick async compaction attempt first and skip a potentially unnecessary but still disruptive reclaim invocation. However, it does seem to me there is a bit of redundancy between the allocator and reclaim: direct reclaim also has that early bailout condition on compaction_ready() for costly orders. In defrag_mode, since reclaim is promoted to pageblock_order (costly), we always bail reclaim on compaction_ready(). So I would assume there is not much difference from which we call first - although I have not tested it. > I'm aware you said "minimally viable bug fix" so it's fine if this falls > outside of that, I'm just trying to poke around to improve my > understanding. It works in production, but I'm not entirely happy yet with the retry logic. I also find it somewhat difficult to understand, especially the split between should_reclaim_retry() and should_compact_retry(). I'm working on unifying them and put the defrag_mode order promotion directly into the slowpath function. This way we can use "work_order" and "order" throughout the functions and its callees as appropriate. But yeah this is follow-up work to the fix here.