From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3AE18E56A for ; Sat, 29 Aug 2026 02:11:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787969485; cv=none; b=CTNPf2d1UxAKOolMgquf2Lo222xIqNLLWoIG0WXFYZhuGBPHteCnO/eC8SjSZTrNh5FRIjSrKmZwthLgNl7KAGuNsJgpdoay1+bwzekrCh5o9haf72/lNT0jd3TsN6EJQJ69urljp6lLenqRwOOLnBYUB5U7G8w76GjKDDCIwlk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787969485; c=relaxed/simple; bh=oelBEjhHl1hxMjbFWrByIYOyj96Vax0A+e/kCuSj454=; h=Date:To:From:Subject:Message-Id; b=LHMsMmI+24hnc/40giorwg5KgRdB59qhHpgR2xHLL7o/xIUBymS7a41Nhf0g7HJjQsC+ow3jRJy8FjNhh4csk6vNp9Hw4CTmFlUUqpiIeYnIZiDDg+brbwTnGEKdfOyD+cy5MjYO/GTxy28g0XbX1AJID9OoKaVXFo/J68CPf8w= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=Vzl0n8g1; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="Vzl0n8g1" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 95D3D1F000E9; Sat, 29 Aug 2026 02:11:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1787969482; bh=Az4dgFQH72iotHxlcw0tcRMaNk9E4sl7eI4sKJh6Cwo=; h=Date:To:From:Subject; b=Vzl0n8g1HUyWKfG0cxqvikKQo5f+tPiojmjVyVWe8ZAFO3MRfdRmp0dCkk8D0/Xa6 5vKEZAl66Sv94asqUeFYYOXE323faEOOKfQPs/vXymfGjqeB3NXxknPY4iFGFNOeJ2 6PSIOG8IVJ0Df0YySpqYMyd4+OTgRRI/elSTYMoo= Date: Fri, 28 Aug 2026 19:11:22 -0700 To: mm-commits@vger.kernel.org,roman.gushchin@linux.dev,muchun.song@linux.dev,mhocko@suse.com,kuba@kernel.org,joshua.hahnjy@gmail.com,hannes@cmpxchg.org,cxiong@meta.com,shakeel.butt@linux.dev,akpm@linux-foundation.org From: Andrew Morton Subject: + memcg-trim-the-per-cpu-charge-stock-instead-of-draining-it.patch added to mm-new branch Message-Id: <20260829021122.95D3D1F000E9@smtp.kernel.org> Precedence: bulk X-Mailing-List: mm-commits@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: The patch titled Subject: memcg: trim the per-cpu charge stock instead of draining it has been added to the -mm mm-new branch. Its filename is memcg-trim-the-per-cpu-charge-stock-instead-of-draining-it.patch This patch will shortly appear at https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/memcg-trim-the-per-cpu-charge-stock-instead-of-draining-it.patch This patch will later appear in the mm-new branch at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Note, mm-new is a provisional staging ground for work-in-progress patches, and acceptance into mm-new is a notification for others take notice and to finish up reviews. Please do not hesitate to respond to review feedback and post updated versions to replace or incrementally fixup patches in mm-new. The mm-new branch of mm.git is not included in linux-next If a few days of testing in mm-new is successful, the patch will me moved into mm.git's mm-unstable branch, which is included in linux-next Before you just go and hit "reply", please: a) Consider who else should be cc'ed b) Prefer to cc a suitable mailing list as well c) Ideally: find the original patch on the mailing list and do a reply-to-all to that, adding suitable additional cc's *** Remember to use Documentation/process/submit-checklist.rst when testing your code *** The -mm tree is included into linux-next via various branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm and is updated there most days ------------------------------------------------------ From: Shakeel Butt Subject: memcg: trim the per-cpu charge stock instead of draining it Date: Wed, 19 Aug 2026 18:20:10 -0700 Joy reported that an application generating a request/response traffic pattern spends 44.6% to 57.0% of CPU in the memcg charge/uncharge path for a range of message sizes, against 0.27% to 0.71% outside that range. Running from the root memcg, where socket memory accounting is skipped, recovers the performance. Tracing the charge path showed that the application generates a pattern where the write syscall charges one page and the read syscall uncharges two pages on the same CPU. This hits a corner case in the memcg percpu stock code that thrashes the stock continuously. In the memcg percpu stock code, MEMCG_CHARGE_BATCH (64) is both the high watermark and the emptying target, i.e. on a request to charge one page the kernel charges MEMCG_CHARGE_BATCH pages and caches (MEMCG_CHARGE_BATCH - 1) of them in the percpu stock. The following uncharge of 2 pages takes the cached count to (MEMCG_CHARGE_BATCH + 1), and refill_stock() then empties the cache completely. With such a pattern the percpu stock becomes completely ineffective. Instead of a single boundary point for charges, use the technique the page allocator uses for its own percpu caches, which keeps the watermark and the emptying target apart: nr_pcp_free() frees between batch and high - batch pages, leaving at least pcp->batch on the list. Add a high watermark MEMCG_STOCK_HIGH and, once the cached count goes over it, return only the pages above MEMCG_STOCK_LOW. The watermarks are MEMCG_CHARGE_BATCH apart, so a page_counter update still covers a full batch. For now, keep MEMCG_STOCK_HIGH same as MEMCG_CHARGE_BATCH and in future we will reevaluate if it makes sense to increase it. Link: https://lore.kernel.org/20260820012010.2016086-1-shakeel.butt@linux.dev Signed-off-by: Shakeel Butt Reported-by: Joy Chaoyue Xiong Acked-by: Michal Hocko Cc: Jakub Kacinski Cc: Johannes Weiner Cc: Joshua Hahn Cc: Muchun Song Cc: Roman Gushchin Signed-off-by: Andrew Morton --- mm/memcontrol.c | 25 +++++++++++++++++++------ 1 file changed, 19 insertions(+), 6 deletions(-) --- a/mm/memcontrol.c~memcg-trim-the-per-cpu-charge-stock-instead-of-draining-it +++ a/mm/memcontrol.c @@ -2032,6 +2032,15 @@ void mem_cgroup_print_oom_group(struct m * nr_pages in a single cacheline. This may change in future. */ #define NR_MEMCG_STOCK 7 + +/* + * Watermarks for a charge stock slot, in the spirit of pcp->high and + * pcp->batch: MEMCG_STOCK_HIGH is the high watermark at which a slot is + * trimmed, and it is trimmed down to MEMCG_STOCK_LOW rather than emptied. + */ +#define MEMCG_STOCK_LOW (MEMCG_CHARGE_BATCH / 2) +#define MEMCG_STOCK_HIGH (MEMCG_CHARGE_BATCH) + #define FLUSHING_CACHED_CHARGE 0 struct memcg_stock_pcp { local_trylock_t lock; @@ -2212,17 +2221,18 @@ static void refill_stock(struct mem_cgro { struct memcg_stock_pcp *stock; struct mem_cgroup *cached; - uint8_t stock_pages; + unsigned int stock_pages; bool success = false; int empty_slot = -1; int i; /* - * For now limit MEMCG_CHARGE_BATCH to 127 and less. In future if we - * decide to increase it more than 127 then we will need more careful - * handling of nr_pages[] in struct memcg_stock_pcp. + * nr_pages[] is a uint8_t and a slot's count is capped at + * MEMCG_STOCK_HIGH. Raising MEMCG_CHARGE_BATCH beyond 127 would need + * more careful handling of nr_pages[] in struct memcg_stock_pcp. */ BUILD_BUG_ON(MEMCG_CHARGE_BATCH > S8_MAX); + BUILD_BUG_ON(MEMCG_STOCK_HIGH > U8_MAX); VM_WARN_ON_ONCE(mem_cgroup_is_root(memcg)); @@ -2243,9 +2253,12 @@ static void refill_stock(struct mem_cgro empty_slot = i; if (memcg == READ_ONCE(stock->cached[i])) { stock_pages = READ_ONCE(stock->nr_pages[i]) + nr_pages; + if (stock_pages > MEMCG_STOCK_HIGH) { + memcg_uncharge(memcg, + stock_pages - MEMCG_STOCK_LOW); + stock_pages = MEMCG_STOCK_LOW; + } WRITE_ONCE(stock->nr_pages[i], stock_pages); - if (stock_pages > MEMCG_CHARGE_BATCH) - drain_stock(stock, i); success = true; break; } _ Patches currently in -mm which might be from shakeel.butt@linux.dev are memcg-clear-flushing_cached_charge-on-cpu-offline.patch memcg-trim-the-per-cpu-charge-stock-instead-of-draining-it.patch